Most retrieval benchmarks measure the index. Users never talk to the index. They talk to a chat path that sits on top of it: query classification, filters, thresholds, re-ranking, formatting. Each of those steps can quietly change what the model gets to read.
So we measured both, on the same questions, and compared them.
The setup
We used SciFact from the BEIR benchmark: 5,183 scientific abstracts and 300 test claims, each with human relevance labels. Nothing is graded by an AI judge. A result is relevant because an annotator said so.
We wrote the protocol down before building anything and recorded every deviation before looking at a result. Before reading a single score, we checked the harness itself:
| Check | Result |
|---|---|
| Every record ingested, one chunk each | 5,183 of 5,183 in every index; no duplicates, none unmapped |
| Each index can find its own records | 200 sampled records ranked first for their own text, in every index |
| Our scorer agrees with the reference implementation | identical to pytrec_eval on every row |
| Our scorer reproduces a published number | BM25 scored 0.676 [0.631, 0.721] against BEIR’s published 0.665 |
| Results are deterministic | top 10 identical across two passes, for all 300 questions |
The pilot run also caught an ingestion problem worth mentioning: long records were being stored cut off at 2,000 characters. That was fixed before the full load.
The indexes tie
Queried directly, three production vector indexes and an exact nearest-neighbour search are a statistical tie:
| Queried directly | nDCG@10 [95% CI] | Recall@10 |
|---|---|---|
| Vertex AI Vector Search, Gemini Embedding 2 | 0.900 [0.873, 0.925] | 0.977 |
| Qdrant, Gemini Embedding 001 | 0.898 [0.872, 0.923] | 0.972 |
| Vertex AI Vector Search, Gemini Embedding 001 | 0.898 [0.872, 0.922] | 0.975 |
| Exact search, Gemini Embedding 001 (reference) | 0.898 [0.871, 0.924] | 0.972 |
Approximate search cost nothing measurable here, and the choice of vector database made no difference. If you stopped at this table, you would conclude retrieval is solved.
The chat path doesn’t match
Then we sent the same 300 questions through the path a chat turn actually uses:
| Through the chat path | nDCG@10 | Questions with no context |
|---|---|---|
| Vertex AI Vector Search, Gemini Embedding 2 | 0.834 | 23 of 300 |
| Qdrant, Gemini Embedding 001 | 0.822 | 27 of 300 |
| Vertex AI Vector Search, Gemini Embedding 001 | 0.051 | 284 of 300 |
On the first two configurations, the chat path gave up about seven points. On the third it returned almost nothing.
The loss had one shape. When the chat path returned passages, they largely matched the index’s own ranking. The missing points were questions for which it returned nothing at all, so the model answered with no context. None of this raised an error. Every request returned 200.
Two defects explained all of it.
A keyword rule that matched inside words. The path checks each question for shopping intent, so that product questions go to a product search. The keywords were matched as substrings. “Production” contains “product”. “Hematopoietic” contains “top”. Ordinary scientific claims were classified as shopping queries, sent to a product search, and came back empty from a corpus that has no products.
A threshold of zero that became 0.62. The index was configured with a minimum similarity of 0, meaning “keep everything”. The code read the setting in a way that treats 0 as missing and substitutes a default of 0.62. With one embedding model on one backend, similarity scores top out near 0.46, so every result fell below the threshold and 284 of 300 questions came back empty. The same index, queried directly, scores 0.898.
Neither defect shows up in uptime, latency or error-rate monitoring. Both show up the moment the chat path is scored against labelled questions.
After the fix
We re-ran the same 300 questions against the same indexes, with the same seed and bootstrap draws:
| Through the chat path | Before | After | Change [95% CI] | No-context questions | Index alone |
|---|---|---|---|---|---|
| Vertex, Embedding 2 | 0.834 | 0.900 | +0.066 [+0.040, +0.094] | 23 → 0 | 0.900 |
| Qdrant, Embedding 001 | 0.822 | 0.898 | +0.077 [+0.049, +0.106] | 27 → 0 | 0.898 |
| Vertex, Embedding 001 | 0.051 | 0.898 | +0.847 [+0.812, +0.880] | 284 → 0 | 0.898 |
No question got worse on any configuration. The chat path now returns the same top 20 passages as the index on all 300 questions for all three, so the whole gap was the empty contexts.
What we take from it
- Measure the path, not only the index. Every number in the first table was true, and none of it described what a user received. Retrieval quality lives in the code between the index and the prompt as much as in the index.
- A 200 is not an answer. Both defects produced fast, successful, empty responses. Only a score against labelled questions tells an empty context apart from a good one.
- Small rules have large reach. A substring match and a falsy zero are ordinary bugs. Here they decided whether one question in thirteen, or nearly every question, reached the model with anything to read.
- Score every release. A check like this belongs on every change to the retrieval path, so a change that empties contexts shows up as a score before it reaches anyone.
Caveats
- The absolute scores are probably inflated. Strong public embedding models usually score around 0.75 to 0.80 on SciFact, and SciFact’s training claims share this corpus, which modern embedding models have likely seen. Every row here uses the same models, so the comparisons hold, but read 0.90 as an upper bound rather than a leaderboard number. That is why we also keep a sealed, private question set that no model can have trained on.
- This covers content retrieval only. The index held abstracts, not products, so product ranking was not exercised.
- One index is missing. Cloudflare Vectorize was excluded because the account was at its index limit.
- These runs are not yet signed. Our public leaderboards publish signed results anyone can verify; this benchmark has not been added to them yet.
Ready to Build Your Custom AI Solution?
Discover how Divinci AI can help you implement RAG systems, automate quality assurance, and streamline your AI development process.
Get Started Today
