Evoke — The Exact-vs-Paraphrase Pattern
How Evoke scores keyword evidence and learned-sparse expansion terms in one inverted index — why exact matches keep leading, when paraphrases enter, and how to hand the result to a reranker.
The Exact-vs-Paraphrase Pattern
Every hybrid-search design has to answer one question: when a literal match and a meaning match disagree, who wins? Most Postgres stacks punt — they run two searches on two scales and merge, and the merge step is where quality quietly dies. Evoke answers it inside one index. This page is the mechanism.
Two kinds of evidence, one score
Evoke's encoder (built on IBM's Granite-Embedding-30M-Sparse) turns text into weighted terms from a fixed vocabulary — terms that include related words the text never uses. At query time:
"heart attack warning signs" → keyword terms: heart, attack, warning, signs
meaning terms: myocardial infarction, symptoms,
chest pain, cardiac
Both lists live in the same inverted index. Keyword terms and expansion terms are namespaced separately, and at query time each document accumulates a single score from both kinds of evidence. There is no second ranked list, no score normalization, no fusion step to tune.
Why exact matches still lead
Expansion terms add evidence; they never subtract it. A query that contains the literal product code gains keyword evidence for that code plus whatever semantic neighbors it has. Because keyword evidence is never given up, the row containing the exact identifier keeps leading — the "Exact codes" demo on the launch post illustrates this directly. Meaning terms only add documents that keyword search would have missed entirely.
That asymmetry matters for real workloads: incident consoles where a query may be ERR_CONN_RESET or "server keeps dropping connections mid-upload." You don't have to choose a query mode in advance — the query itself sorts out which kind of evidence dominates by how your data scores.
The reranker handoff
Evoke is deliberately first-stage. The intended architecture:
- Evoke, top ~1,000 — maximize recall cheaply on CPU (one look-up)
- Cross-encoder reranker, top ~20 — spend quality where a human or agent looks
The benchmark behind that split: on BEIR15, Recall@100 improved 0.563 → 0.667 over BM25-only, within 0.004 of a ~20×-bigger dense encoder, with more relevant documents than the dense model in the top 1,000. A reranker can only reorder what stage one hands it — recall is the ceiling, ranking is a different job.
Failure modes worth knowing
saeoff — withoutWITH (sae = true)the index behaves as keyword-only. If you see purely keyword-quality results, check the index definition (see Getting Started).- Fresh rows aren't semantically searchable yet — expansion terms publish via background workers after commit. Queries still get keyword evidence immediately; meaning evidence follows. Design read-your-write paths for identifiers, not paraphrases.
- Model swap drift — after any model change, re-encoding happens in the background; measure recall@k on a held-out set before trusting the new postings end-to-end.
- Non-English text — evaluated on English suites only today; test on your corpus before committing.
When to reach for this pattern
- Users send both literal identifiers and plain-language descriptions to the same search box
- You can't afford an embedding pipeline, a GPU, or a second index to keep warm
- You're feeding an agent that retries searches with sharpened questions, so recall beats precision
- Someone on your team has to maintain the current hybrid stack, and the fusion query is the thing nobody wants to touch
If none of those hold, BM25 alone may still be the right answer — Evoke's value shows up exactly where keyword-only search is silently losing meaning matches.
Related
- Evoke overview
- Getting Started
- Use Cases
- Jev Confidence Gating — the same "cheap evidence first, judgment second" split on the decision side
Related Articles & Guides
Evoke — Semantic Search Inside Postgres
Evoke (II-42 / ii42) is an open-source Postgres extension for learned-sparse retrieval: keyword and meaning evidence share one inverted index and one score. A 30M-parameter model runs on CPUs inside the database — no embedding pipeline, no vector index, no fusion step.
Evoke — Getting Started
Install Evoke (II-42/ii42) with Docker or package managers, create your first hybrid keyword-plus-semantic index, run your first ii42_query, and understand the async write path.
Evoke — Use Cases
Where Evoke's one-index, CPU-only, inside-Postgres retrieval fits: agent doc search, exact-code-or-plain-words lookups, air-gapped deployments, reranker handoff, async event logs, and MCP-backed knowledge retrieval — plus where it doesn't.