Evoke — Semantic Search Inside Postgres

Evoke (II-42 / ii42) is an open-source Postgres extension for learned-sparse retrieval: keyword and meaning evidence share one inverted index and one score. A 30M-parameter model runs on CPUs inside the database — no embedding pipeline, no vector index, no fusion step.

October 7, 2026
evokeii42postgressemantic-searchlearned-sparsebm25ragagents

Evoke — Semantic Search Inside Postgres

Evoke is an open-source (Apache-2.0) Postgres extension from Intelligent Internet that adds search by meaning without adding machinery. Its ~30M-parameter model (built on IBM's Granite-Embedding-30M-Sparse) turns documents and queries into weighted terms — including related words the text never uses — and puts them in the same inverted index as your BM25 keyword terms, in a namespace of their own. At query time both kinds of evidence add up to one score per document.

The common Postgres hybrid-search stack has five parts: an embedding model or API to call, a job keeping vectors in sync with rows, a vector index beside the keyword index, and a query that fuses two ranked lists sitting on different score scales. Evoke replaces all of it with one index and one lookup, in plain SQL:

CREATE INDEX docs_semantic_idx ON docs USING ii42 (body) WITH (sae = true);

SELECT d.id, d.title,
       ii42_query('docs_semantic_idx'::regclass, 'database search architecture') AS score
FROM docs AS d
ORDER BY score DESC
LIMIT 10;

The extension is named ii42 in SQL. The model runs on ordinary CPUs in shared background workers — writes commit straight away and are encoded afterward, so nothing waits on the model, no GPU is needed, and no text leaves the database. Shipped as v0.2.5 with a PostgreSQL 18 Docker image, PG17/18 Linux x86-64 packages, and a checksummed offline archive for air-gapped installs.

How It Compares

BM25 onlyPostgres hybrid (BM25 + pgvector)Evoke
Indexes121
Query steps13 (2 searches + merge)1
Finds paraphrasesNoYesYes
Score scales to reconcile121
Embedding pipeline/sync jobNoneRequiredNone
Model hardware—GPU (typical)CPU only
Text leaves the DB?NoYes (to embed)No
BenchmarksBEIR15 R@100 0.563similar to dense column0.667 (within 0.004 of a 20× bigger dense model)

The last row is the honest trade table. Evoke gives up almost nothing in first-stage recall against a ~20× larger dense model, while deleting the entire pipeline that made dense retrieval expensive to operate. What it does not give you: multilingual coverage (English only for now), final ordering (it's a first-stage retriever — pair it with a reranker), or RAG out of the box (it returns ranked rows; the agent part is yours).

What Makes It Different

  • Learned-sparse, not dense. Semantic signal is stored as vocabulary terms, not vectors — which is why it fits the inverted index Postgres already has, why there is no ANN index to tune, and why vectors can't drift out of step with rows.
  • One model runtime, shared. Every connection hands encoding to background workers instead of loading its own copy. Writes queue and encode asynchronously; every returned row is re-checked against current row visibility, so deleted documents don't resurface.
  • Keyword evidence is never given up. Exact matches — product codes, case numbers, error strings — keep leading when they exist; meaning terms only add missed paraphrases.
  • Postgres-native ops. Crash recovery and physical replication follow PostgreSQL's own rules for the index. Air-gapped installs are supported out of the box.

Section Contents

  • Getting Started — Docker/package install, first index, first query, async write path.
  • Use Cases — six grounded fit patterns with don't-use-it boundaries, plus honest limits.
  • Exact vs Paraphrase Pattern — the mechanism behind one-index hybrid scoring and when keywords win.

Frequently Asked Questions