NEWPromptev MCP Server: one URL for every tool, every agent and your team’s knowledge. Connect Claude Code, Cursor, or any MCP client.→

Context Engine benchmarks

What was measured, under which conditions, and what was not. Every figure on this page is copied from the benchmark files in the repository, which carry the raw results and the commands to run them again.

Overview

The repository carries four sets of measurements. Each one states its corpus, its queries, its method and its limits, and each one is the library compared with itself: no other product appears in any of them.

BenchmarkWhat it answersMeasured on
Permissions under narrow accessDoes a caller who may see 10% or 1% of a corpus get what a corpus holding only their documents would return?100,000 documents, 200 to 400 searches per leg, synthetic permissions
Search timeWhat a default hybrid search of a sentence-length question costs, and what one setting changes.The same corpus, 200 questions, one laptop
Retrieval qualityHybrid search against the vector leg alone, and what each default was set from.Five public datasets, 300 queries each
Graph searchWhat the graph leg adds to hybrid search, on both graph backends.Four capped multi-hop datasets, 300 queries each

Read the conditions with the numbers. The permissions are synthetic. Recall in the permission benchmark is agreement with a corpus that holds only the caller’s documents, not relevance. The vector leg is exact only at or below 50,000 visible chunks. At the shipped defaults hybrid search does not beat the vector leg alone on most of the public datasets measured. The section What these benchmarks do not show lists the rest.

Permissions under narrow access

The library applies permissions inside the SQL of every retrieval leg. The alternative is to search first and filter afterwards, and it loses results: the search spends its limit on rows the caller may not see, and what is left after the filter is short or empty, with no error. This benchmark measures how much.

On a 100,000-document corpus with synthetic team-shaped permissions, a caller who may see 10% or 1% of the documents gets from every leg measured exactly the top ten that a corpus holding only their documents returns, while the same search filtered afterwards returns nothing for 64% and 99% of hybrid searches, recovers 0.42 and 0.02 with a tenfold fetch, and 0.64 and 0.08 with a fiftyfold one.

What was compared

  • Permissions inside the query. engine.search(query, top_k=10, principals=[the caller's group]). The predicate runs inside each leg’s SQL.
  • Filter after, same limit. The same search with no permissions, then the top 10 filtered to what the caller may see.
  • Filter after, 10x and 50x fetch. The top 100 or the top 500 unscoped, filtered, cut to 10. A leg returns 500 candidates at most, so fifty times is the deepest fetch measured.
  • The reference. The same search, with the same settings, on a copy of the database that holds only the documents the caller may see, with index scans switched off so that its vector leg is an exact scan. Recall@10 is the share of that copy’s top ten a method returned. “Empty” is the share of searches where a method returned nothing although the copy had results.

Hybrid search as shipped

100,646 chunks, clustered (team-shaped) permissions, the first 200 questions of the dataset. The 95% interval, a percentile bootstrap over queries, is under each recall figure.

SeesMethodRecall@10Empty
10%Permissions inside the query1.000[1.000, 1.000]0.0%
10%Filter after, same limit0.084[0.065, 0.103]63.5%
10%Filter after, 10x fetch0.419[0.375, 0.470]13.0%
10%Filter after, 50x fetch0.638[0.598, 0.679]1.0%
1%Permissions inside the query1.000[1.000, 1.000]0.0%
1%Filter after, same limit0.001[0.000, 0.003]99.0%
1%Filter after, 10x fetch0.021[0.012, 0.033]77.0%
1%Filter after, 50x fetch0.083[0.063, 0.106]38.0%

Where every search agreed with the reference the bootstrap interval has no width. The bound that does is the rule of three: with 200 searches and no disagreement, up to 1.5% of searches could still differ. “The same top ten” on this page means the same set of ten results, not the same order. Order was checked separately, for the full-text and trigram legs only, against the oracle described under Conditions.

Recall@10 of hybrid search for a caller who may see 10% or 1% of the corpusAt 10% visibility: permissions inside the query 1.000 with 0.0% of searches empty, filter after, same limit 0.084 with 63.5% of searches empty, filter after, 10x fetch 0.419 with 13.0% of searches empty, filter after, 50x fetch 0.638 with 1.0% of searches empty. At 1% visibility: permissions inside the query 1.000 with 0.0% of searches empty, filter after, same limit 0.001 with 99.0% of searches empty, filter after, 10x fetch 0.021 with 77.0% of searches empty, filter after, 50x fetch 0.083 with 38.0% of searches empty.Caller sees 10% of the corpusPermissions inside the query1.000 · 0.0% emptyPermissions inside the query, 10% visible: recall@10 1.000, 0.0% of searches emptyFilter after, same limit0.084 · 63.5% emptyFilter after, same limit, 10% visible: recall@10 0.084, 63.5% of searches emptyFilter after, 10x fetch0.419 · 13.0% emptyFilter after, 10x fetch, 10% visible: recall@10 0.419, 13.0% of searches emptyFilter after, 50x fetch0.638 · 1.0% emptyFilter after, 50x fetch, 10% visible: recall@10 0.638, 1.0% of searches emptyCaller sees 1% of the corpusPermissions inside the query1.000 · 0.0% emptyPermissions inside the query, 1% visible: recall@10 1.000, 0.0% of searches emptyFilter after, same limit0.001 · 99.0% emptyFilter after, same limit, 1% visible: recall@10 0.001, 99.0% of searches emptyFilter after, 10x fetch0.021 · 77.0% emptyFilter after, 10x fetch, 1% visible: recall@10 0.021, 77.0% of searches emptyFilter after, 50x fetch0.083 · 38.0% emptyFilter after, 50x fetch, 1% visible: recall@10 0.083, 38.0% of searches empty
Recall@10 of hybrid search, and the share of searches returned empty. 100,000 documents (100,646 chunks), 200 queries, synthetic team-shaped permissions. Recall is agreement with a corpus holding only the caller’s documents, not relevance. The vector leg is exact here because each caller sees fewer than 50,000 chunks. Bars are drawn to scale, except that the 0.001 bar is drawn at a minimum visible width. Its label carries the value.

Every leg

Each cell is recall@10 with the share of searches returned empty under it. “In the query” is permissions inside the query; “After” is the same search filtered afterwards at the same limit (1x), with a tenfold fetch and with a fiftyfold fetch. The first four legs ran on the 100,646-chunk corpus. The graph rows ran on a separate corpus of 2,500 documents (2,515 chunks), because graph ingest calls a model for every document, and they are scored only over the searches where the graph leg ran: 322 of 400 at 10% visibility and 66 of 400 at 1%. The Postgres and Neo4j backends returned the same figures in every cell.

Caller sees 10%

LegIn the queryAfter, 1xAfter, 10xAfter, 50x
vector (ann)200 searches1.0000.0%0.08562.5%0.46721.5%0.8173.0%
full-text, "any"200 searches1.0000.0%0.09246.5%0.6796.0%0.9720.0%
trigram195 searches1.0000.0%0.10652.3%0.6853.6%0.9960.0%
hybrid, as shipped200 searches1.0000.0%0.08463.5%0.41913.0%0.6381.0%
graph leg322 searches1.0000.0%0.20739.4%0.59010.2%0.59610.2%
hybrid + graph400 searches1.0000.0%0.08853.3%0.5061.5%0.6800.0%

Caller sees 1%

LegIn the queryAfter, 1xAfter, 10xAfter, 50x
vector (ann)200 searches1.0000.0%0.00298.5%0.01692.0%0.07274.0%
full-text, "any"199 searches1.0000.0%0.01591.0%0.06079.9%0.18547.2%
trigram188 searches1.0000.0%0.02888.8%0.21849.5%0.50916.0%
hybrid, as shipped200 searches1.0000.0%0.00199.0%0.02177.0%0.08338.0%
graph leg66 searches1.0000.0%0.12878.8%0.65322.7%0.67718.2%
hybrid + graph400 searches1.0000.0%0.00893.5%0.06359.0%0.29414.0%

The full-text leg at its default, lexical_match="all", has no row: it matched nothing these callers may see for these questions, under either method, so there was nothing to score. The "any" row is the same leg with lexical_match="any". The graph rows hold for both the Postgres and the Neo4j backend. With no disagreement, up to 1.5% of searches could still differ at 200 searches, 0.75% at 400 and 4.5% at 66. The hybrid + graph rows include the searches where the graph leg did not run, so at 1% they are mostly hybrid results.

Random against clustered permissions

The same number of visible documents, drawn uniformly at random instead of as topic clusters. With permissions inside the query the layout makes no difference: recall against the visible-only corpus is 1.000 either way, under the conditions above. Filtering after the search depends on it. Each cell is recall@10 with the share of searches returned empty under it.

hybrid, as shipped, filtered after the search

SeesFetchClusteredRandom
10%same limit0.08463.5%0.08644.0%
10%10x0.41913.0%0.7300.0%
10%50x0.6381.0%0.8120.0%
1%same limit0.00199.0%0.01189.0%
1%10x0.02177.0%0.08639.5%
1%50x0.08338.0%0.4180.0%

vector leg, filtered after the search

SeesFetchClusteredRandom
10%same limit0.08562.5%0.08842.5%
10%10x0.46721.5%0.8730.0%
10%50x0.8173.0%0.9950.0%
1%same limit0.00298.5%0.01288.5%
1%10x0.01692.0%0.10235.5%
1%50x0.07274.0%0.5030.5%

At the same limit the two layouts lose about the same recall. What clustering changes there is how the losses are spread: more searches come back with nothing at all. Fetching more is where the layout decides the outcome. A benchmark that assigns permissions at random and over-fetches will report that filtering after the search is nearly safe. The longer argument is in Your ACL benchmark is measuring the easy case.

Above 50,000 visible chunks

The vector leg is the only one whose index is approximate. The engine counts the chunks a search may see and, at or below storage.ann_exact_threshold (annExactThreshold, default 50,000), sorts them exactly. Every 10% and 1% caller above is under that. Above it the leg is an iterative scan of the pgvector index. That case was measured on a second ingest of the same 100,000 documents, with clustered permissions that make 48%, 60%, 90% and 100% of them visible. Recall is against the visible-only corpus, as above.

SeesVector routeVector legHybridAfter, 1xAfter, 10x
48%48,378 chunksexact sort1.000not measured0.5765.0%0.9700.0%
60%60,396 chunksindex scan0.9790.9790.5964.0%0.9690.0%
90%90,605 chunksindex scan0.9880.9870.9270.5%0.9880.0%
100%100,646 chunksindex scan0.9880.9880.9880.0%0.9880.0%

Recall@10 with engine defaults. “Vector leg” and “Hybrid” are with permissions inside the query; “After” is the vector leg filtered after the search at the same limit and with a tenfold fetch. 95% intervals: vector leg [1.000, 1.000] at 48%, [0.970, 0.987] at 60%, [0.982, 0.993] at 90% and at 100%; hybrid [0.970, 0.987], [0.981, 0.992] and [0.982, 0.993]. Hybrid was not measured at 48%.

  • Over the threshold the vector leg is approximate, and close. 0.979 at 60% visible, 0.988 at 90% and at 100%, with no empty results. The same index searched with no predicate at all scores 0.988, so that is the approximation of the index itself.
  • With most of the corpus visible, filtering afterwards catches up once it over-fetches. The difference between the two methods is a matter of narrow access.
  • The exact route is the slow one. Just under the threshold (48,378 visible chunks) the vector leg took 337 ms at the median on the exact route, against 18 ms for the approximate index scan on the same caller. The threshold is where the engine stops paying that for the last one or two hundredths of recall.
  • With the exact route switched off at narrow access, the same leg scored 0.943 for the clustered 10% caller (0.978 with random permissions). With the index forced as well it scored 0.652 for the clustered 1% caller (0.961 random), returning a full list of ten with nothing to say it was incomplete. That is the loss the exact route exists to remove.
  • Narrow access above the threshold is not measured. A caller who sees 1% of more than five million chunks also takes the approximate path, with a far more selective predicate than any measured here. Nothing is claimed for that case.

Conditions

  • Synthetic permissions. No real permission data was used. For the clustered layout, 40 seed documents are drawn at random, every document joins its nearest seed by embedding similarity, and whole clusters are taken until the caller sees 10% or 1% of the corpus. The clusters are cut in the same embedding space the vector leg searches, which is the hardest layout for that leg. Real permissions follow topic less cleanly.
  • Agreement, not relevance. A recall of 1.000 here means the caller got what a deployment holding only their documents would return. It does not say those results answer the question.
  • Single-group callers, no public documents. Each caller holds one group and every document carries an ACL.
  • The reference shares the engine’s ranking code, so a ranking fault common to both would not show. The vector reference was checked against a brute-force numpy search and agreed on every query. The full-text and trigram legs were checked against a brute-force oracle that uses none of the engine’s statements and no index: the same top ten in the same order on every query it covers (172 of 200 for full-text with "any", all 200 for trigram). That is the only place order is claimed: every other comparison on this page is of the set of ten. That oracle reuses the engine’s stopword list and trigram threshold rule. The graph leg has no independent oracle.
  • Setup. BEIR HotpotQA, 100,000 documents (the judged documents of 500 questions, the rest sampled with seed 0), BAAI/bge-small-en-v1.5 at 384 dimensions, PostgreSQL 16.15, pgvector 0.8.6, HNSW m=16, ef_construction=64, engine defaults, library commit 08f55b7. The queries are the dataset’s questions in dataset order and were not chosen to match the visible topics. The benchmark runs the Python package.

Method, every table, the raw CSVs and the per-query rows: benchmarks/acl_recall.

Search time and the trigram setting

At the shipped defaults a hybrid search of one of these questions takes about 2.6 seconds at the median on the 100,646-chunk corpus (one query at a time, one laptop, query embedding not counted). Nearly all of that is the trigram leg: about 2.6 s, against about 0.1 s for the vector leg and a few milliseconds for the full-text leg. The trigram leg compares the query’s trigrams with every chunk’s, so its cost grows with the length of the query, and these questions are long: 16 words at the median, 95% of them longer than 8 words and 47% longer than 16.

search.trgm_max_query_words (search.trgmMaxQueryWords) is a setting for this. A query with more words than the setting runs without the trigram leg, and shorter queries run as before. It is unset by default, which means the trigram leg always runs, and that default is unchanged. You choose.

Caller sees 10%

SettingSearch p50Leg ranSame tenShared
unset (default)2,563 msp90 2,955 ms100.0%100.0%100.0%
162,160 msp90 2,692 ms53.0%60.5%95.1%
8113 msp90 152 ms5.0%34.0%92.5%
trigram weight 0110 msp90 129 ms0.0%33.5%92.4%

Caller sees 1%

SettingSearch p50Leg ranSame tenShared
unset (default)2,674 msp90 3,354 ms100.0%100.0%100.0%
162,119 msp90 2,547 ms53.0%59.0%94.8%
896 msp90 112 ms5.0%35.5%91.9%
trigram weight 097 msp90 110 ms0.0%35.0%91.8%

Caller with no restriction

SettingSearch p50Leg ranSame tenShared
unset (default)2,658 msp90 3,115 ms100.0%100.0%100.0%
162,188 msp90 2,650 ms53.0%62.0%96.8%
811 msp90 17 ms5.0%33.0%95.1%
trigram weight 011 msp90 13 ms0.0%29.5%94.9%

“Setting” is the value of trgm_max_query_words; the last row switches the trigram leg off with a fusion weight of 0. 200 questions per row, top_k=10, clustered permissions, median search time with the 90th percentile under it. “Leg ran” is the share of searches that still ran the trigram leg. “Same ten” is the share of searches whose ten results are the same set as the defaults’ ten for that caller, and “shared” is the share of the defaults’ results still in the top ten.

  • What the setting buys. At 8 words, 5% of these questions still run the trigram leg and the median search falls from 2,563 ms to 113 ms for the caller who sees 10% (about 2.6 s to about 0.11 s). At 16 words, 53% of the questions still run the leg, so the median search is still one that runs it and barely moves. A gate helps a workload in proportion to the share of queries it skips.
  • What it costs. The results change. At 8 words about a third of the searches return the same top ten as the defaults, and 92% to 95% of the default’s results are still in the top ten. Whether the different results are better or worse is not measured here. On the public datasets below, a gate at 8 or 16 words raised hybrid nDCG@10 on three of five and lowered it by 0.025 on one, which is why it ships unset.
  • What it does not change. Permissions. Under every profile the engine returned the same set of ten as the visible-only corpus searched under the same profile, on all 200 searches at both visibilities.
  • What the trigram leg is for. Names, codes, identifiers and misspelled words, which are short queries. A gate at 8 or 16 words leaves those untouched.

The times are one machine with one query at a time, and other work was running on that machine during these runs. They show the size of the effect and are not a latency benchmark. The setting is documented in the configuration reference.

Retrieval quality

Hybrid search at the shipped defaults (trigram weight 0.4, RRF k 20) against the vector leg alone, on five public datasets, the first 300 queries of each, with BAAI/bge-small-en-v1.5 embeddings. Every comparison is paired query by query on the same ingest, and the interval is a 95% bootstrap of the per-query differences.

At the shipped defaults hybrid search beats the vector leg alone on 1 of 5 datasets, ties on 1 and trails on 3. The fusion defaults alone do not close the gap with this embedding model. Measure it on your own queries.

DatasetVector legHybridHybrid minus vector
MuSiQue0.5330.508-0.025[-0.039, -0.011] trails
2WikiMultihopQA0.7820.725-0.057[-0.071, -0.044] trails
MultiHop-RAG0.6580.686+0.029[+0.015, +0.042] ahead
ArguAna0.5730.513-0.060[-0.085, -0.032] trails
CQADupStack English0.5430.523-0.020[-0.042, +0.005] tie

nDCG@10, with the 95% interval of the paired difference under each delta. “Tie” means the interval spans zero. Recall@10, vector leg then hybrid: MuSiQue 0.558 / 0.556; 2WikiMultihopQA 0.752 / 0.745; MultiHop-RAG 0.766 / 0.790; ArguAna 0.833 / 0.817; CQADupStack English 0.581 / 0.584.

Which rows these are. The figures are the dense rows and the “RRF k 20” variant rows of round two, whose engine still defaulted to k 60. The BM25 page later ran the same settings as the defaults on a fresh ingest, and its figures are within 0.002 of these with the same verdict on every dataset. A fresh run of the command under Reproduce corresponds to that later run.

Corpus sizes, documents and chunks: MuSiQue 11,656 / 12,211; 2WikiMultihopQA 6,119 / 6,817; MultiHop-RAG 609 / 6,976; ArguAna 8,674 / 12,142; CQADupStack English 10,000 of 40,221 / 10,924.

nDCG@10 of the vector leg alone and of hybrid search at the shipped defaults, on five public datasetsMuSiQue: vector leg alone 0.533, hybrid 0.508. 2WikiMultihopQA: vector leg alone 0.782, hybrid 0.725. MultiHop-RAG: vector leg alone 0.658, hybrid 0.686. ArguAna: vector leg alone 0.573, hybrid 0.513. CQADupStack English: vector leg alone 0.543, hybrid 0.523.Vector leg aloneHybrid, shipped defaultsMuSiQuevectorMuSiQue, vector leg alone: nDCG@10 0.5330.533hybridMuSiQue, hybrid at the shipped defaults: nDCG@10 0.5080.5082WikiMultihopQAvector2WikiMultihopQA, vector leg alone: nDCG@10 0.7820.782hybrid2WikiMultihopQA, hybrid at the shipped defaults: nDCG@10 0.7250.725MultiHop-RAGvectorMultiHop-RAG, vector leg alone: nDCG@10 0.6580.658hybridMultiHop-RAG, hybrid at the shipped defaults: nDCG@10 0.6860.686ArguAnavectorArguAna, vector leg alone: nDCG@10 0.5730.573hybridArguAna, hybrid at the shipped defaults: nDCG@10 0.5130.513CQADupStack EnglishvectorCQADupStack English, vector leg alone: nDCG@10 0.5430.543hybridCQADupStack English, hybrid at the shipped defaults: nDCG@10 0.5230.523
nDCG@10 on the first 300 queries of each dataset, the vector leg alone and hybrid search at the shipped defaults (trigram weight 0.4, RRF k 20). One small embedding model, bge-small-en-v1.5. Retrieval only, document-level judgments, a trusted caller. CQADupStack English is capped at 10,000 of its 40,221 documents. The scale runs from 0 to 1.

What each default was set from

The decision rule was written down before any run: change a default only where the paired delta on hybrid nDCG@10 is positive with a 95% confidence interval above 0 on the majority of datasets (3 of 5), and negative beyond its interval on none. Two defaults changed under it. Each cell is the change in hybrid nDCG@10 against hybrid as it shipped at that commit: (+) marks an interval above zero and (-) one below.

Setting triedSourceMuSiQue2WikiMultiHop-RAGArguAnaCQADupStackDecision
Trigram weight 0.8 to 0.4round 1+0.069 (+)+0.062 (+)-0.001+0.062 (+)+0.042 (+)Default changed to 0.4
RRF k 60 to 20round 2+0.049 (+)+0.049 (+)+0.003+0.042 (+)+0.039 (+)Default changed to 20
RRF k 60 to 100round 2-0.027 (-)-0.027 (-)-0.000-0.029 (-)-0.022 (-)Rejected
lexical_match="any"round 2-0.010+0.004+0.027 (+)-0.122 (-)-0.053 (-)"all" stays
trgm_max_query_words=8round 2+0.071 (+)+0.083 (+)-0.025 (-)+0.097 (+)+0.012Ships unset
trgm_max_query_words=16round 2+0.032 (+)+0.008 (+)-0.025 (-)+0.097 (+)-0.002Ships unset
lexical_candidates=100round 2+0.002+0.006 (+)-0.002-0.001-0.000Ships unset
lexical_rank="bm25" with "any"BM25 page+0.039 (+)+0.016 (+)+0.062 (+)+0.004-0.043 (-)Opt-in, experimental
  • Each round has its own baseline. Round 1 measured against hybrid with the trigram weight at 0.8. Round 2 and the BM25 page ran on fresh ingests, round 2 against hybrid at trigram weight 0.4 and RRF k 60, and the BM25 page against the current defaults. The rows are not additive across rounds.
  • The trigram gate fails the rule on one dataset. MultiHop-RAG questions are long and full of names, so every gate removes the trigram leg there for nearly every query. Elsewhere it is the largest single gain measured, and it cut median latency from 2.4 seconds to 44 milliseconds on ArguAna. It ships as a setting for deployments whose queries are long sentences and whose corpus is not name-heavy.
  • The experimental BM25 rank does not change a default. With lexical_match="any" it raised hybrid nDCG@10 on three datasets and lowered it on CQADupStack, so it stays opt-in. Against the vector leg alone it is ahead on 1 dataset, level on 1 and behind on 3, the same count as the shipped defaults.
  • No value outside each grid was tried, deliberately, so the defaults are not tuned to these queries.

Tables, intervals and commands: round 1 (engine commit ad984580f216), round 2 (7ac727491ff5) and the BM25 page (98d4590eabc2).

Graph search

Each dataset was ingested once in graph mode and searched four ways over that one ingest: the vector leg alone, hybrid search as shipped, hybrid plus the graph leg in Postgres, and hybrid plus the graph leg in Neo4j. Every configuration answered the same 300 queries. Graph mode is opt-in and nothing here changes a default.

The graph leg raised nDCG@10 over hybrid search on two of the four datasets (2WikiMultihopQA by +0.044 and GraphRAG-Bench novel by +0.046, each with a 95% interval above zero), made no measurable difference on MuSiQue, and lowered it slightly on MultiHop-RAG (-0.016). The Postgres and Neo4j backends returned the same ranking for every one of the 1,200 queries.

DatasetChange in nDCG@10Change in Recall@10
MuSiQue-0.000[-0.022, +0.024]+0.008[-0.016, +0.032]
2WikiMultihopQA+0.044[+0.029, +0.061]+0.065[+0.045, +0.086]
GraphRAG-Bench novel+0.046[+0.019, +0.074]+0.076[+0.039, +0.116]
MultiHop-RAG-0.016[-0.033, -0.001]+0.010[-0.006, +0.027]

Hybrid plus the graph leg, minus hybrid, on the same 300 queries, with the 95% interval under each figure. The Neo4j backend returned a ranking identical to the Postgres backend on 300 of 300 queries on every dataset.

DatasetDocumentsChunksHybridHybrid + graph
MuSiQue1,000 of 11,6561,0440.7360.736
2WikiMultihopQA1,000 of 6,1191,1820.8010.845
GraphRAG-Bench novel600 of 4,401 passages1,0770.5870.633
MultiHop-RAG300 of 6093,6180.7320.716

nDCG@10 on the capped corpus of each dataset. “Documents” is the cap out of the dataset’s full size.

Change in nDCG@10 from adding the graph leg to hybrid search, with 95% intervals, on four capped multi-hop datasetsMuSiQue: -0.000 [-0.022, +0.024], interval spans zero. 2WikiMultihopQA: +0.044 [+0.029, +0.061], interval above zero. GraphRAG-Bench novel: +0.046 [+0.019, +0.074], interval above zero. MultiHop-RAG: -0.016 [-0.033, -0.001], interval below zero.MuSiQue-0.000 [-0.022, +0.024]2WikiMultihopQA+0.044 [+0.029, +0.061]2WikiMultihopQA: change in nDCG@10 from adding the graph leg +0.044 [+0.029, +0.061], interval above zeroGraphRAG-Bench novel+0.046 [+0.019, +0.074]GraphRAG-Bench novel: change in nDCG@10 from adding the graph leg +0.046 [+0.019, +0.074], interval above zeroMultiHop-RAG-0.016 [-0.033, -0.001]MultiHop-RAG: change in nDCG@10 from adding the graph leg -0.016 [-0.033, -0.001], interval below zero-0.040+0.04+0.08
Change in nDCG@10 from adding the graph leg to hybrid search, with the 95% interval as a whisker. Blue: interval above zero. Orange: interval below zero. An interval that spans zero is no measurable difference: MuSiQue’s change rounds to zero, so it has no bar, only its interval. Corpora capped at 300 to 1,000 documents to fit a model budget, 300 multi-hop queries each, engine defaults, community summaries off. Both graph backends gave these same figures.
  • The graph leg widens what is found more than it sharpens the top of the list. Recall@10 went up on all four datasets, beyond its interval on the same two. The rank of the first relevant document did not improve in the same way.
  • Cost. Graph ingest cost between $2.48 and $3.26 of model spend per 1,000 chunks with the model used here, and a graph search makes no model call. The model was Gemini 2.5 Flash with thinking switched off, at its list price at the time of the run ($0.30 per million input tokens, $2.50 per million output tokens): $3.23 on MuSiQue, $3.26 on 2WikiMultihopQA, $2.49 on GraphRAG-Bench novel, $2.48 on MultiHop-RAG.
  • The corpora are capped, and the cap was set by cost. A cap keeps every judged document of the 300 queries and fills the rest with distractors drawn with seed 0. With fewer distractors every configuration scores higher than it would on the full corpus, so these absolute numbers are not comparable with the retrieval-quality section above.
  • Setup. Engine commit b98ef202e89d, PostgreSQL 16.15 with pgvector 0.8.6, Neo4j 5.26.31, bge-small-en-v1.5, trigram weight 0.4, RRF k 20, graph weight 1.0.

Tables, the per-dataset figures and commands: benchmarks/retrieval/results/graph.md.

What these benchmarks do not show

  • Hybrid search does not beat vector search here. On five public datasets with a small embedding model, hybrid at the shipped defaults is ahead of the vector leg alone on one, level on one and behind on three. A stronger embedder moves the vector leg, and with it the balance between the legs. A corpus of part numbers, error codes, names or misspellings may need the keyword legs more than these datasets do.
  • The graph leg does not always help. Two of four datasets, all built from multi-hop questions, where a graph is expected to matter most. Single-hop questions and full-size corpora were not measured. On MuSiQue the vector leg alone had the highest nDCG@10 of the four configurations.
  • Not answer quality. Every figure is retrieval only. No answer was generated or graded.
  • Not real permissions. The permission layouts are generated, each caller holds a single group, and every document carries an ACL. Callers with several groups, documents visible to everyone, and scoping by source or document were not measured.
  • Nothing beyond the measured scale. The permission run is 100,000 documents. Above the exact threshold the vector leg was measured only with 60% to 100% of that corpus visible. Narrow access above the threshold is not measured. Over-fetching beyond fifty times is not measured either.
  • One dataset, one embedding model, one machine for the permission benchmark, and it runs the Python package. The TypeScript client was not measured. Longer documents, another model or another dimension move the numbers for filtering after the search.
  • Queries are not aimed at the visible topics. A caller who searches only for what their own team wrote will find filtering after the search less harmful than these numbers. A caller who searches for anything else will not.
  • Not latency. Times are one machine, one query at a time, with other work running on it for part of the runs. They show the size of an effect. They are not a claim about either graph backend or about any other system.
  • The graph leg under permissions has no independent oracle, and its corpus is small: at 1% visibility 26 chunks are visible and the rows rest on 66 to 70 queries.
  • Not a comparison with any other product. Both sides of every comparison are this library’s own legs and settings.

Each benchmark page in the repository carries its own longer list.

Reproduce

The benchmarks run from a clone of the repository: benchmarks/ is not in the installed package. They need Docker, uv and a Postgres with pgvector 0.8 or newer. A fresh ingest gives chunks new ids and builds a new HNSW graph, and extraction by a model is not deterministic, so a rerun reproduces these numbers within their intervals, not to the last digit.

Permissions under narrow access

Shell
# from a clone of the repository: benchmarks/ does not ship in the package
docker run -d --name acl-bench-pg -p 55432:5432 -e POSTGRES_PASSWORD=PASSWORD \
    pgvector/pgvector:0.8.6-pg16
export BENCH_DATABASE_URL=postgresql://postgres:PASSWORD@localhost:55432/postgres

# The hybrid corpus: vector, full-text, trigram and hybrid at 10% and 1%.
# About 50 minutes of ingest and three to four hours of searches on a laptop.
uv run --group bench --extra graph python -m benchmarks.acl_recall.legs \
    --limit-docs 100000 --limit-queries 500 --max-queries 200 --concurrency 12 \
    --out benchmarks/acl_recall/results-legs-main.csv

# The vector leg alone on the same ingest: its plan and its own latency.
uv run --group bench --extra graph python -m benchmarks.acl_recall.legs \
    --limit-docs 100000 --limit-queries 500 --max-queries 200 --reuse-ingest \
    --dense-only --out benchmarks/acl_recall/results-legs-dense.csv

The run creates and drops databases named ce_acl_legs_* and touches nothing else on the server. The second ingest (above the threshold, the search-time profiles, the oracle) and the graph corpus have their own commands in the benchmark’s README.

Retrieval quality

Shell
docker run -d --name ce-bench -e POSTGRES_PASSWORD=bench -p 55433:5432 \
  pgvector/pgvector:0.8.6-pg16 -c max_connections=300
export BENCH_DATABASE_URL=postgresql://postgres:bench@localhost:55433/postgres

# Hybrid at the shipped defaults against the vector leg alone.
CONFIGS="dense,hybrid"

for ds in 2wiki musique multihop-rag arguana; do
  uv run --group bench python -m benchmarks.retrieval.run --dataset $ds \
    --configs "$CONFIGS" --baseline hybrid --seed 0 --limit-queries 300
done
uv run --group bench python -m benchmarks.retrieval.run --dataset cqadupstack-english \
  --configs "$CONFIGS" --baseline hybrid --seed 0 --limit-queries 300 --limit-docs 10000

The full list of configurations each round ran is on its results page. Harness and datasets: benchmarks/retrieval.

Graph search

Shell
# Uses the Postgres container and BENCH_DATABASE_URL of the previous block.
# Entity extraction needs an OpenAI-compatible chat endpoint: set BENCH_LLM_BASE_URL
# and BENCH_LLM_API_KEY, and put your model's name in place of YOUR_MODEL. The run wipes its
# Neo4j database, so give it one of its own on a port that is not Neo4j's default.
export NEO4J_PASSWORD=choose-at-least-8-characters
docker run -d --name ce-bench-neo4j -e NEO4J_AUTH=neo4j/$NEO4J_PASSWORD -p 57687:7687 neo4j:5
export NEO4J_USER=neo4j NEO4J_PASSWORD

uv run --group bench --extra graph python -m benchmarks.retrieval.run --dataset 2wiki \
  --configs dense,hybrid,hybrid+graph,hybrid+graph-neo4j --baseline hybrid \
  --model YOUR_MODEL --neo4j bolt://localhost:57687 \
  --limit-docs 1000 --limit-queries 300 --seed 0 --concurrency 8

The other three datasets use the same command with musique and --limit-docs 1000, graphrag-bench-novel and 600, and multihop-rag and 300. Before it writes or deletes anything the harness checks that the Neo4j database is empty or already carries this run’s marker, and stops otherwise.

Your own corpus

A benchmark on public data says nothing about the shape of your permissions. context-engine check-acl-exposure ships with the package, reads your own database and reports whether filtered retrieval is losing results there. See the CLI reference.

2,500 free credits · No card required · No subscription