Thursday, October 8, 2026
Evoke Is Statistical Retrieval in Postgres — One Index for Keywords and Meaning
Posted by

Intelligent Internet shipped Evoke today — an Apache-2.0 Postgres extension that adds semantic search without giving you a second index to babysit. A ~30M-parameter model runs inside Postgres on ordinary CPUs, and its output lands in the same inverted index as your BM25 terms. One lookup. One score. One ranked list. No embedding API, no vector column, no sync job, no fusion query.
If you're an agent or RAG builder whose knowledge lives in Postgres, this is the retirement notice for a surprising amount of your infrastructure. The launch tweet drew 246 likes and 256 bookmarks in a few hours — a near 1:1 bookmark-to-like ratio, which is what a "save this, I'll need it" tool looks like on X.
What Evoke Actually Is
The common Postgres search stack today has five moving parts. Your app calls an embedding model or API outside the database. A job keeps vectors in sync with rows. You run keyword search (tsvector/pg_trgm/BM25) and a vector search (pgvector, VectorChord). Then — this is the part everyone underestimates — you merge two ranked lists whose scores sit on different scales.
Evoke (the extension and its SQL functions are named ii42, repo Intelligent-Internet/II-42, v0.2.5) compresses that stack using learned-sparse retrieval, the SPLADE-family idea industrialized by IBM's Granite-Embedding-30M-Sparse. Every document and every query is encoded into weighted terms from a fixed vocabulary — terms that can include related words the text never uses. "heart attack warning signs" produces not just heart, attack, signs but myocardial infarction, chest pain, cardiac.
Those expansion terms go into the same inverted index as your keyword terms, in a namespace of their own. Both kinds of evidence add up to one score per document:
CREATE INDEX docs_semantic_idx ON docs USING ii42 (body) WITH (sae = true);
SELECT d.id, d.title,
ii42_query('docs_semantic_idx'::regclass, 'database search architecture') AS score
FROM docs AS d
ORDER BY score DESC
LIMIT 10;
That's the whole API surface. Plain SQL, nothing to fuse. That is the product — Evoke's own reply thread said it straightest: "One lookup, one score, one ranked list. Nothing to fuse."
The Numbers, Read Honestly
Evaluated on two fixed English suites against BM25 and a 0.6B-parameter dense model (Perplexity's pplx-embed-v1-0.6B via VectorChord):
| Suite | BM25 only | Evoke P2.1 (30M) | Dense 0.6B |
|---|---|---|---|
| BEIR15 Recall@100 | 0.563 | 0.667 | 0.671 |
| MTEB10 Recall@100 | 0.595 | 0.703 | (≈same top-100) |
The marketing shorthand — "matches a 0.6B dense embedding model to within 0.004, with 20× fewer parameters" — is accurate, and the more interesting half is under-tweeted: Evoke placed more relevant documents in the top 1,000 than the dense model on both suites (higher CUB@1000). As a first-stage retriever feeding a reranker, that's the metric that matters — a reranker can only reorder what it receives.
Keep three caveats in view: English-only text for now; the numbers come from fixed eval suites, not your corpus; and this is first-stage retrieval, not final ordering.
Why Databases Love Learned-Sparse and Hate Vectors
The design choices follow from one insight: semantic signal expressed as vocabulary terms fits an index structure Postgres already has. A dense vector doesn't. Dense retrieval drags in an ANN index with its own tuning, its own memory profile, and its own failure mode: the vector drifting out of step with the row it belongs to.
Evoke instead inherits Postgres's existing plumbing whole. Writes commit immediately — keyword evidence is stored, the row queues, and shared background workers encode it afterward with ONNX Runtime on CPU (one shared runtime, not one copy per connection). Deleted rows can't resurface because every result is re-checked against current row visibility. Crash recovery and physical replication come for free because the index follows Postgres's own rules.
The ops-level quote from the thread that best captures why people bookmarked it: "the small model is interesting, but having less to set up and maintain is what catches my eye." That's the honest headline. It's not that 30M parameters beat 600M — it's that one index and one score replace two indexes, two searches, a merge step, an embedding service, and a sync job.
Use Cases the Thread and Docs Pointed At
- Agent / RAG search over internal docs and policies — the flagship. Agents search repeatedly per task, so first-stage recall is the ceiling on everything downstream. An agent's retrieval step becoming one SQL call against a table it already reads is a real simplification. For the worked pattern, see our Evoke guide and agent-facing use cases.
- "Exact code OR plain words" lookups — support and incident tooling where a query is an error string or "users can't log in after the rename." In conventional hybrid stacks, fusing those is exactly where quality dies. In Evoke the literal match still leads because keyword evidence is never given up.
- Air-gapped and regulated deployments — model runs in-database, no text leaves, and the release ships a checksummed offline Docker archive. Clinical guidelines, case law, regulatory rulebooks: places whose only prior option was "no semantic search at all."
- Review-style work — anywhere recall beats precision: e-discovery, compliance sweeps, deduplication passes.
The full blog post covers the scoring internals and the write/query path in more depth, including two nice interactive diagrams.
What to Watch
The thread surfaced two open questions worth tracking, both still unanswered:
- Index size vs plain BM25 — expansion terms are extra postings; nobody has published a growth ratio yet. Ask it on the repo.
- Model swaps — one commenter's guidance is good ops hygiene regardless: queue refreshes, and after any model swap measure recall@k on a held-out set before you trust the new vectors.
Also reasonable to wait before betting big on a v0.2.5 — but the architecture claim it defends is conservative and checkable: you should not need a vector stack to get meaning-aware search out of a Postgres instance you already run.
Related
- Evoke Tool Guide — setup, use cases, and the exact-vs-meaning pattern
- Manticore Search Made Embeddings 14× Faster — the same "in-place inference beats the pipeline" thesis, different engine
- Agent Memory Architectures — where retrieval sits in the agent loop
- Structured CodeAgent — agents that search, read, and refine repeatedly
Related Articles & Deep Dives
#deepseekDeepSeek Harness: Everything Is a Plugin
DeepSeek Harness (dsh) is DeepSeek's MIT-licensed agent runtime where everything is a plugin — the four presets, setup, and what it means for benchmark trust.
#geminiGemini 3.7 Flash: The Workhorse Gets Smarter and Cheaper
Gemini 3.7 Flash ships three weeks after 3.6 Flash with big coding and agent gains at half the price. FrontierCode, DeepSWE, AutomationBench, and what changed.
#harness-ifHarness-IF: Are Coding Agents Following Rules, or Just Doing What They'd Do Anyway?
A benchmark that separates compliance from coincidence: every model is worse at rules opposing its defaults — and where you put the rule changes everything.