What was measured, under which conditions, and what was not. Every figure on this page is copied from the benchmark files in the repository, which carry the raw results and the commands to run them again.
Overview
The repository carries four sets of measurements. Each one states its corpus, its queries, its method and its limits, and each one is the library compared with itself: no other product appears in any of them.
What the graph leg adds to hybrid search, on both graph backends.
Four capped multi-hop datasets, 300 queries each
Read the conditions with the numbers. The permissions are synthetic. Recall in the permission benchmark is agreement with a corpus that holds only the caller’s documents, not relevance. The vector leg is exact only at or below 50,000 visible chunks. At the shipped defaults hybrid search does not beat the vector leg alone on most of the public datasets measured. The section What these benchmarks do not show lists the rest.
Permissions under narrow access
The library applies permissions inside the SQL of every retrieval leg. The alternative is to search first and filter afterwards, and it loses results: the search spends its limit on rows the caller may not see, and what is left after the filter is short or empty, with no error. This benchmark measures how much.
On a 100,000-document corpus with synthetic team-shaped permissions, a caller who may see 10% or 1% of the documents gets from every leg measured exactly the top ten that a corpus holding only their documents returns, while the same search filtered afterwards returns nothing for 64% and 99% of hybrid searches, recovers 0.42 and 0.02 with a tenfold fetch, and 0.64 and 0.08 with a fiftyfold one.
What was compared
Permissions inside the query.engine.search(query, top_k=10, principals=[the caller's group]). The predicate runs inside each leg’s SQL.
Filter after, same limit. The same search with no permissions, then the top 10 filtered to what the caller may see.
Filter after, 10x and 50x fetch. The top 100 or the top 500 unscoped, filtered, cut to 10. A leg returns 500 candidates at most, so fifty times is the deepest fetch measured.
The reference. The same search, with the same settings, on a copy of the database that holds only the documents the caller may see, with index scans switched off so that its vector leg is an exact scan. Recall@10 is the share of that copy’s top ten a method returned. “Empty” is the share of searches where a method returned nothing although the copy had results.
Hybrid search as shipped
100,646 chunks, clustered (team-shaped) permissions, the first 200 questions of the dataset. The 95% interval, a percentile bootstrap over queries, is under each recall figure.
Sees
Method
Recall@10
Empty
10%
Permissions inside the query
1.000[1.000, 1.000]
0.0%
10%
Filter after, same limit
0.084[0.065, 0.103]
63.5%
10%
Filter after, 10x fetch
0.419[0.375, 0.470]
13.0%
10%
Filter after, 50x fetch
0.638[0.598, 0.679]
1.0%
1%
Permissions inside the query
1.000[1.000, 1.000]
0.0%
1%
Filter after, same limit
0.001[0.000, 0.003]
99.0%
1%
Filter after, 10x fetch
0.021[0.012, 0.033]
77.0%
1%
Filter after, 50x fetch
0.083[0.063, 0.106]
38.0%
Where every search agreed with the reference the bootstrap interval has no width. The bound that does is the rule of three: with 200 searches and no disagreement, up to 1.5% of searches could still differ. “The same top ten” on this page means the same set of ten results, not the same order. Order was checked separately, for the full-text and trigram legs only, against the oracle described under Conditions.
Recall@10 of hybrid search, and the share of searches returned empty. 100,000 documents (100,646 chunks), 200 queries, synthetic team-shaped permissions. Recall is agreement with a corpus holding only the caller’s documents, not relevance. The vector leg is exact here because each caller sees fewer than 50,000 chunks. Bars are drawn to scale, except that the 0.001 bar is drawn at a minimum visible width. Its label carries the value.
Every leg
Each cell is recall@10 with the share of searches returned empty under it. “In the query” is permissions inside the query; “After” is the same search filtered afterwards at the same limit (1x), with a tenfold fetch and with a fiftyfold fetch. The first four legs ran on the 100,646-chunk corpus. The graph rows ran on a separate corpus of 2,500 documents (2,515 chunks), because graph ingest calls a model for every document, and they are scored only over the searches where the graph leg ran: 322 of 400 at 10% visibility and 66 of 400 at 1%. The Postgres and Neo4j backends returned the same figures in every cell.
Caller sees 10%
Leg
In the query
After, 1x
After, 10x
After, 50x
vector (ann)200 searches
1.0000.0%
0.08562.5%
0.46721.5%
0.8173.0%
full-text, "any"200 searches
1.0000.0%
0.09246.5%
0.6796.0%
0.9720.0%
trigram195 searches
1.0000.0%
0.10652.3%
0.6853.6%
0.9960.0%
hybrid, as shipped200 searches
1.0000.0%
0.08463.5%
0.41913.0%
0.6381.0%
graph leg322 searches
1.0000.0%
0.20739.4%
0.59010.2%
0.59610.2%
hybrid + graph400 searches
1.0000.0%
0.08853.3%
0.5061.5%
0.6800.0%
Caller sees 1%
Leg
In the query
After, 1x
After, 10x
After, 50x
vector (ann)200 searches
1.0000.0%
0.00298.5%
0.01692.0%
0.07274.0%
full-text, "any"199 searches
1.0000.0%
0.01591.0%
0.06079.9%
0.18547.2%
trigram188 searches
1.0000.0%
0.02888.8%
0.21849.5%
0.50916.0%
hybrid, as shipped200 searches
1.0000.0%
0.00199.0%
0.02177.0%
0.08338.0%
graph leg66 searches
1.0000.0%
0.12878.8%
0.65322.7%
0.67718.2%
hybrid + graph400 searches
1.0000.0%
0.00893.5%
0.06359.0%
0.29414.0%
The full-text leg at its default, lexical_match="all", has no row: it matched nothing these callers may see for these questions, under either method, so there was nothing to score. The "any" row is the same leg with lexical_match="any". The graph rows hold for both the Postgres and the Neo4j backend. With no disagreement, up to 1.5% of searches could still differ at 200 searches, 0.75% at 400 and 4.5% at 66. The hybrid + graph rows include the searches where the graph leg did not run, so at 1% they are mostly hybrid results.
Random against clustered permissions
The same number of visible documents, drawn uniformly at random instead of as topic clusters. With permissions inside the query the layout makes no difference: recall against the visible-only corpus is 1.000 either way, under the conditions above. Filtering after the search depends on it. Each cell is recall@10 with the share of searches returned empty under it.
hybrid, as shipped, filtered after the search
Sees
Fetch
Clustered
Random
10%
same limit
0.08463.5%
0.08644.0%
10%
10x
0.41913.0%
0.7300.0%
10%
50x
0.6381.0%
0.8120.0%
1%
same limit
0.00199.0%
0.01189.0%
1%
10x
0.02177.0%
0.08639.5%
1%
50x
0.08338.0%
0.4180.0%
vector leg, filtered after the search
Sees
Fetch
Clustered
Random
10%
same limit
0.08562.5%
0.08842.5%
10%
10x
0.46721.5%
0.8730.0%
10%
50x
0.8173.0%
0.9950.0%
1%
same limit
0.00298.5%
0.01288.5%
1%
10x
0.01692.0%
0.10235.5%
1%
50x
0.07274.0%
0.5030.5%
At the same limit the two layouts lose about the same recall. What clustering changes there is how the losses are spread: more searches come back with nothing at all. Fetching more is where the layout decides the outcome. A benchmark that assigns permissions at random and over-fetches will report that filtering after the search is nearly safe. The longer argument is in Your ACL benchmark is measuring the easy case.
Above 50,000 visible chunks
The vector leg is the only one whose index is approximate. The engine counts the chunks a search may see and, at or below storage.ann_exact_threshold (annExactThreshold, default 50,000), sorts them exactly. Every 10% and 1% caller above is under that. Above it the leg is an iterative scan of the pgvector index. That case was measured on a second ingest of the same 100,000 documents, with clustered permissions that make 48%, 60%, 90% and 100% of them visible. Recall is against the visible-only corpus, as above.
Sees
Vector route
Vector leg
Hybrid
After, 1x
After, 10x
48%48,378 chunks
exact sort
1.000
not measured
0.5765.0%
0.9700.0%
60%60,396 chunks
index scan
0.979
0.979
0.5964.0%
0.9690.0%
90%90,605 chunks
index scan
0.988
0.987
0.9270.5%
0.9880.0%
100%100,646 chunks
index scan
0.988
0.988
0.9880.0%
0.9880.0%
Recall@10 with engine defaults. “Vector leg” and “Hybrid” are with permissions inside the query; “After” is the vector leg filtered after the search at the same limit and with a tenfold fetch. 95% intervals: vector leg [1.000, 1.000] at 48%, [0.970, 0.987] at 60%, [0.982, 0.993] at 90% and at 100%; hybrid [0.970, 0.987], [0.981, 0.992] and [0.982, 0.993]. Hybrid was not measured at 48%.
Over the threshold the vector leg is approximate, and close. 0.979 at 60% visible, 0.988 at 90% and at 100%, with no empty results. The same index searched with no predicate at all scores 0.988, so that is the approximation of the index itself.
With most of the corpus visible, filtering afterwards catches up once it over-fetches. The difference between the two methods is a matter of narrow access.
The exact route is the slow one. Just under the threshold (48,378 visible chunks) the vector leg took 337 ms at the median on the exact route, against 18 ms for the approximate index scan on the same caller. The threshold is where the engine stops paying that for the last one or two hundredths of recall.
With the exact route switched off at narrow access, the same leg scored 0.943 for the clustered 10% caller (0.978 with random permissions). With the index forced as well it scored 0.652 for the clustered 1% caller (0.961 random), returning a full list of ten with nothing to say it was incomplete. That is the loss the exact route exists to remove.
Narrow access above the threshold is not measured. A caller who sees 1% of more than five million chunks also takes the approximate path, with a far more selective predicate than any measured here. Nothing is claimed for that case.
Conditions
Synthetic permissions. No real permission data was used. For the clustered layout, 40 seed documents are drawn at random, every document joins its nearest seed by embedding similarity, and whole clusters are taken until the caller sees 10% or 1% of the corpus. The clusters are cut in the same embedding space the vector leg searches, which is the hardest layout for that leg. Real permissions follow topic less cleanly.
Agreement, not relevance. A recall of 1.000 here means the caller got what a deployment holding only their documents would return. It does not say those results answer the question.
Single-group callers, no public documents. Each caller holds one group and every document carries an ACL.
The reference shares the engine’s ranking code, so a ranking fault common to both would not show. The vector reference was checked against a brute-force numpy search and agreed on every query. The full-text and trigram legs were checked against a brute-force oracle that uses none of the engine’s statements and no index: the same top ten in the same order on every query it covers (172 of 200 for full-text with "any", all 200 for trigram). That is the only place order is claimed: every other comparison on this page is of the set of ten. That oracle reuses the engine’s stopword list and trigram threshold rule. The graph leg has no independent oracle.
Setup. BEIR HotpotQA, 100,000 documents (the judged documents of 500 questions, the rest sampled with seed 0), BAAI/bge-small-en-v1.5 at 384 dimensions, PostgreSQL 16.15, pgvector 0.8.6, HNSW m=16, ef_construction=64, engine defaults, library commit 08f55b7. The queries are the dataset’s questions in dataset order and were not chosen to match the visible topics. The benchmark runs the Python package.
At the shipped defaults a hybrid search of one of these questions takes about 2.6 seconds at the median on the 100,646-chunk corpus (one query at a time, one laptop, query embedding not counted). Nearly all of that is the trigram leg: about 2.6 s, against about 0.1 s for the vector leg and a few milliseconds for the full-text leg. The trigram leg compares the query’s trigrams with every chunk’s, so its cost grows with the length of the query, and these questions are long: 16 words at the median, 95% of them longer than 8 words and 47% longer than 16.
search.trgm_max_query_words (search.trgmMaxQueryWords) is a setting for this. A query with more words than the setting runs without the trigram leg, and shorter queries run as before. It is unset by default, which means the trigram leg always runs, and that default is unchanged. You choose.
Caller sees 10%
Setting
Search p50
Leg ran
Same ten
Shared
unset (default)
2,563 msp90 2,955 ms
100.0%
100.0%
100.0%
16
2,160 msp90 2,692 ms
53.0%
60.5%
95.1%
8
113 msp90 152 ms
5.0%
34.0%
92.5%
trigram weight 0
110 msp90 129 ms
0.0%
33.5%
92.4%
Caller sees 1%
Setting
Search p50
Leg ran
Same ten
Shared
unset (default)
2,674 msp90 3,354 ms
100.0%
100.0%
100.0%
16
2,119 msp90 2,547 ms
53.0%
59.0%
94.8%
8
96 msp90 112 ms
5.0%
35.5%
91.9%
trigram weight 0
97 msp90 110 ms
0.0%
35.0%
91.8%
Caller with no restriction
Setting
Search p50
Leg ran
Same ten
Shared
unset (default)
2,658 msp90 3,115 ms
100.0%
100.0%
100.0%
16
2,188 msp90 2,650 ms
53.0%
62.0%
96.8%
8
11 msp90 17 ms
5.0%
33.0%
95.1%
trigram weight 0
11 msp90 13 ms
0.0%
29.5%
94.9%
“Setting” is the value of trgm_max_query_words; the last row switches the trigram leg off with a fusion weight of 0. 200 questions per row, top_k=10, clustered permissions, median search time with the 90th percentile under it. “Leg ran” is the share of searches that still ran the trigram leg. “Same ten” is the share of searches whose ten results are the same set as the defaults’ ten for that caller, and “shared” is the share of the defaults’ results still in the top ten.
What the setting buys. At 8 words, 5% of these questions still run the trigram leg and the median search falls from 2,563 ms to 113 ms for the caller who sees 10% (about 2.6 s to about 0.11 s). At 16 words, 53% of the questions still run the leg, so the median search is still one that runs it and barely moves. A gate helps a workload in proportion to the share of queries it skips.
What it costs. The results change. At 8 words about a third of the searches return the same top ten as the defaults, and 92% to 95% of the default’s results are still in the top ten. Whether the different results are better or worse is not measured here. On the public datasets below, a gate at 8 or 16 words raised hybrid nDCG@10 on three of five and lowered it by 0.025 on one, which is why it ships unset.
What it does not change. Permissions. Under every profile the engine returned the same set of ten as the visible-only corpus searched under the same profile, on all 200 searches at both visibilities.
What the trigram leg is for. Names, codes, identifiers and misspelled words, which are short queries. A gate at 8 or 16 words leaves those untouched.
The times are one machine with one query at a time, and other work was running on that machine during these runs. They show the size of the effect and are not a latency benchmark. The setting is documented in the configuration reference.
Retrieval quality
Hybrid search at the shipped defaults (trigram weight 0.4, RRF k 20) against the vector leg alone, on five public datasets, the first 300 queries of each, with BAAI/bge-small-en-v1.5 embeddings. Every comparison is paired query by query on the same ingest, and the interval is a 95% bootstrap of the per-query differences.
At the shipped defaults hybrid search beats the vector leg alone on 1 of 5 datasets, ties on 1 and trails on 3. The fusion defaults alone do not close the gap with this embedding model. Measure it on your own queries.
Dataset
Vector leg
Hybrid
Hybrid minus vector
MuSiQue
0.533
0.508
-0.025[-0.039, -0.011] trails
2WikiMultihopQA
0.782
0.725
-0.057[-0.071, -0.044] trails
MultiHop-RAG
0.658
0.686
+0.029[+0.015, +0.042] ahead
ArguAna
0.573
0.513
-0.060[-0.085, -0.032] trails
CQADupStack English
0.543
0.523
-0.020[-0.042, +0.005] tie
nDCG@10, with the 95% interval of the paired difference under each delta. “Tie” means the interval spans zero. Recall@10, vector leg then hybrid: MuSiQue 0.558 / 0.556; 2WikiMultihopQA 0.752 / 0.745; MultiHop-RAG 0.766 / 0.790; ArguAna 0.833 / 0.817; CQADupStack English 0.581 / 0.584.
Which rows these are. The figures are the dense rows and the “RRF k 20” variant rows of round two, whose engine still defaulted to k 60. The BM25 page later ran the same settings as the defaults on a fresh ingest, and its figures are within 0.002 of these with the same verdict on every dataset. A fresh run of the command under Reproduce corresponds to that later run.
Corpus sizes, documents and chunks: MuSiQue 11,656 / 12,211; 2WikiMultihopQA 6,119 / 6,817; MultiHop-RAG 609 / 6,976; ArguAna 8,674 / 12,142; CQADupStack English 10,000 of 40,221 / 10,924.
nDCG@10 on the first 300 queries of each dataset, the vector leg alone and hybrid search at the shipped defaults (trigram weight 0.4, RRF k 20). One small embedding model, bge-small-en-v1.5. Retrieval only, document-level judgments, a trusted caller. CQADupStack English is capped at 10,000 of its 40,221 documents. The scale runs from 0 to 1.
What each default was set from
The decision rule was written down before any run: change a default only where the paired delta on hybrid nDCG@10 is positive with a 95% confidence interval above 0 on the majority of datasets (3 of 5), and negative beyond its interval on none. Two defaults changed under it. Each cell is the change in hybrid nDCG@10 against hybrid as it shipped at that commit: (+) marks an interval above zero and (-) one below.
Setting tried
Source
MuSiQue
2Wiki
MultiHop-RAG
ArguAna
CQADupStack
Decision
Trigram weight 0.8 to 0.4
round 1
+0.069 (+)
+0.062 (+)
-0.001
+0.062 (+)
+0.042 (+)
Default changed to 0.4
RRF k 60 to 20
round 2
+0.049 (+)
+0.049 (+)
+0.003
+0.042 (+)
+0.039 (+)
Default changed to 20
RRF k 60 to 100
round 2
-0.027 (-)
-0.027 (-)
-0.000
-0.029 (-)
-0.022 (-)
Rejected
lexical_match="any"
round 2
-0.010
+0.004
+0.027 (+)
-0.122 (-)
-0.053 (-)
"all" stays
trgm_max_query_words=8
round 2
+0.071 (+)
+0.083 (+)
-0.025 (-)
+0.097 (+)
+0.012
Ships unset
trgm_max_query_words=16
round 2
+0.032 (+)
+0.008 (+)
-0.025 (-)
+0.097 (+)
-0.002
Ships unset
lexical_candidates=100
round 2
+0.002
+0.006 (+)
-0.002
-0.001
-0.000
Ships unset
lexical_rank="bm25" with "any"
BM25 page
+0.039 (+)
+0.016 (+)
+0.062 (+)
+0.004
-0.043 (-)
Opt-in, experimental
Each round has its own baseline. Round 1 measured against hybrid with the trigram weight at 0.8. Round 2 and the BM25 page ran on fresh ingests, round 2 against hybrid at trigram weight 0.4 and RRF k 60, and the BM25 page against the current defaults. The rows are not additive across rounds.
The trigram gate fails the rule on one dataset. MultiHop-RAG questions are long and full of names, so every gate removes the trigram leg there for nearly every query. Elsewhere it is the largest single gain measured, and it cut median latency from 2.4 seconds to 44 milliseconds on ArguAna. It ships as a setting for deployments whose queries are long sentences and whose corpus is not name-heavy.
The experimental BM25 rank does not change a default. With lexical_match="any" it raised hybrid nDCG@10 on three datasets and lowered it on CQADupStack, so it stays opt-in. Against the vector leg alone it is ahead on 1 dataset, level on 1 and behind on 3, the same count as the shipped defaults.
No value outside each grid was tried, deliberately, so the defaults are not tuned to these queries.
Tables, intervals and commands: round 1 (engine commit ad984580f216), round 2 (7ac727491ff5) and the BM25 page (98d4590eabc2).
Graph search
Each dataset was ingested once in graph mode and searched four ways over that one ingest: the vector leg alone, hybrid search as shipped, hybrid plus the graph leg in Postgres, and hybrid plus the graph leg in Neo4j. Every configuration answered the same 300 queries. Graph mode is opt-in and nothing here changes a default.
The graph leg raised nDCG@10 over hybrid search on two of the four datasets (2WikiMultihopQA by +0.044 and GraphRAG-Bench novel by +0.046, each with a 95% interval above zero), made no measurable difference on MuSiQue, and lowered it slightly on MultiHop-RAG (-0.016). The Postgres and Neo4j backends returned the same ranking for every one of the 1,200 queries.
Dataset
Change in nDCG@10
Change in Recall@10
MuSiQue
-0.000[-0.022, +0.024]
+0.008[-0.016, +0.032]
2WikiMultihopQA
+0.044[+0.029, +0.061]
+0.065[+0.045, +0.086]
GraphRAG-Bench novel
+0.046[+0.019, +0.074]
+0.076[+0.039, +0.116]
MultiHop-RAG
-0.016[-0.033, -0.001]
+0.010[-0.006, +0.027]
Hybrid plus the graph leg, minus hybrid, on the same 300 queries, with the 95% interval under each figure. The Neo4j backend returned a ranking identical to the Postgres backend on 300 of 300 queries on every dataset.
Dataset
Documents
Chunks
Hybrid
Hybrid + graph
MuSiQue
1,000 of 11,656
1,044
0.736
0.736
2WikiMultihopQA
1,000 of 6,119
1,182
0.801
0.845
GraphRAG-Bench novel
600 of 4,401 passages
1,077
0.587
0.633
MultiHop-RAG
300 of 609
3,618
0.732
0.716
nDCG@10 on the capped corpus of each dataset. “Documents” is the cap out of the dataset’s full size.
Change in nDCG@10 from adding the graph leg to hybrid search, with the 95% interval as a whisker. Blue: interval above zero. Orange: interval below zero. An interval that spans zero is no measurable difference: MuSiQue’s change rounds to zero, so it has no bar, only its interval. Corpora capped at 300 to 1,000 documents to fit a model budget, 300 multi-hop queries each, engine defaults, community summaries off. Both graph backends gave these same figures.
The graph leg widens what is found more than it sharpens the top of the list. Recall@10 went up on all four datasets, beyond its interval on the same two. The rank of the first relevant document did not improve in the same way.
Cost. Graph ingest cost between $2.48 and $3.26 of model spend per 1,000 chunks with the model used here, and a graph search makes no model call. The model was Gemini 2.5 Flash with thinking switched off, at its list price at the time of the run ($0.30 per million input tokens, $2.50 per million output tokens): $3.23 on MuSiQue, $3.26 on 2WikiMultihopQA, $2.49 on GraphRAG-Bench novel, $2.48 on MultiHop-RAG.
The corpora are capped, and the cap was set by cost. A cap keeps every judged document of the 300 queries and fills the rest with distractors drawn with seed 0. With fewer distractors every configuration scores higher than it would on the full corpus, so these absolute numbers are not comparable with the retrieval-quality section above.
Hybrid search does not beat vector search here. On five public datasets with a small embedding model, hybrid at the shipped defaults is ahead of the vector leg alone on one, level on one and behind on three. A stronger embedder moves the vector leg, and with it the balance between the legs. A corpus of part numbers, error codes, names or misspellings may need the keyword legs more than these datasets do.
The graph leg does not always help. Two of four datasets, all built from multi-hop questions, where a graph is expected to matter most. Single-hop questions and full-size corpora were not measured. On MuSiQue the vector leg alone had the highest nDCG@10 of the four configurations.
Not answer quality. Every figure is retrieval only. No answer was generated or graded.
Not real permissions. The permission layouts are generated, each caller holds a single group, and every document carries an ACL. Callers with several groups, documents visible to everyone, and scoping by source or document were not measured.
Nothing beyond the measured scale. The permission run is 100,000 documents. Above the exact threshold the vector leg was measured only with 60% to 100% of that corpus visible. Narrow access above the threshold is not measured. Over-fetching beyond fifty times is not measured either.
One dataset, one embedding model, one machine for the permission benchmark, and it runs the Python package. The TypeScript client was not measured. Longer documents, another model or another dimension move the numbers for filtering after the search.
Queries are not aimed at the visible topics. A caller who searches only for what their own team wrote will find filtering after the search less harmful than these numbers. A caller who searches for anything else will not.
Not latency. Times are one machine, one query at a time, with other work running on it for part of the runs. They show the size of an effect. They are not a claim about either graph backend or about any other system.
The graph leg under permissions has no independent oracle, and its corpus is small: at 1% visibility 26 chunks are visible and the rows rest on 66 to 70 queries.
Not a comparison with any other product. Both sides of every comparison are this library’s own legs and settings.
Each benchmark page in the repository carries its own longer list.
Reproduce
The benchmarks run from a clone of the repository: benchmarks/ is not in the installed package. They need Docker, uv and a Postgres with pgvector 0.8 or newer. A fresh ingest gives chunks new ids and builds a new HNSW graph, and extraction by a model is not deterministic, so a rerun reproduces these numbers within their intervals, not to the last digit.
Permissions under narrow access
Shell
# from a clone of the repository: benchmarks/ does not ship in the package
docker run -d --name acl-bench-pg -p 55432:5432 -e POSTGRES_PASSWORD=PASSWORD \
pgvector/pgvector:0.8.6-pg16
export BENCH_DATABASE_URL=postgresql://postgres:PASSWORD@localhost:55432/postgres
# The hybrid corpus: vector, full-text, trigram and hybrid at 10% and 1%.
# About 50 minutes of ingest and three to four hours of searches on a laptop.
uv run --group bench --extra graph python -m benchmarks.acl_recall.legs \
--limit-docs 100000 --limit-queries 500 --max-queries 200 --concurrency 12 \
--out benchmarks/acl_recall/results-legs-main.csv
# The vector leg alone on the same ingest: its plan and its own latency.
uv run --group bench --extra graph python -m benchmarks.acl_recall.legs \
--limit-docs 100000 --limit-queries 500 --max-queries 200 --reuse-ingest \
--dense-only --out benchmarks/acl_recall/results-legs-dense.csv
The run creates and drops databases named ce_acl_legs_* and touches nothing else on the server. The second ingest (above the threshold, the search-time profiles, the oracle) and the graph corpus have their own commands in the benchmark’s README.
Retrieval quality
Shell
docker run -d --name ce-bench -e POSTGRES_PASSWORD=bench -p 55433:5432 \
pgvector/pgvector:0.8.6-pg16 -c max_connections=300
export BENCH_DATABASE_URL=postgresql://postgres:bench@localhost:55433/postgres
# Hybrid at the shipped defaults against the vector leg alone.
CONFIGS="dense,hybrid"
for ds in 2wiki musique multihop-rag arguana; do
uv run --group bench python -m benchmarks.retrieval.run --dataset $ds \
--configs "$CONFIGS" --baseline hybrid --seed 0 --limit-queries 300
done
uv run --group bench python -m benchmarks.retrieval.run --dataset cqadupstack-english \
--configs "$CONFIGS" --baseline hybrid --seed 0 --limit-queries 300 --limit-docs 10000
The full list of configurations each round ran is on its results page. Harness and datasets: benchmarks/retrieval.
Graph search
Shell
# Uses the Postgres container and BENCH_DATABASE_URL of the previous block.
# Entity extraction needs an OpenAI-compatible chat endpoint: set BENCH_LLM_BASE_URL
# and BENCH_LLM_API_KEY, and put your model's name in place of YOUR_MODEL. The run wipes its
# Neo4j database, so give it one of its own on a port that is not Neo4j's default.
export NEO4J_PASSWORD=choose-at-least-8-characters
docker run -d --name ce-bench-neo4j -e NEO4J_AUTH=neo4j/$NEO4J_PASSWORD -p 57687:7687 neo4j:5
export NEO4J_USER=neo4j NEO4J_PASSWORD
uv run --group bench --extra graph python -m benchmarks.retrieval.run --dataset 2wiki \
--configs dense,hybrid,hybrid+graph,hybrid+graph-neo4j --baseline hybrid \
--model YOUR_MODEL --neo4j bolt://localhost:57687 \
--limit-docs 1000 --limit-queries 300 --seed 0 --concurrency 8
The other three datasets use the same command with musique and --limit-docs 1000, graphrag-bench-novel and 600, and multihop-rag and 300. Before it writes or deletes anything the harness checks that the Neo4j database is empty or already carries this run’s marker, and stops otherwise.
Your own corpus
A benchmark on public data says nothing about the shape of your permissions. context-engine check-acl-exposure ships with the package, reads your own database and reports whether filtered retrieval is losing results there. See the CLI reference.
2,500 free credits · No card required · No subscription