How we evaluate retrieval: a golden set, a row with no leakage, and a benchmark that refuses to run
Three design decisions in our retrieval benchmark that are not obvious, and that anyone who inherits it needs to understand. Where the test set comes from, which configuration is allowed to decide, and when the benchmark says no.
Our internal knowledge system, Cortex, serves context to the coding agents that work across our repositories: rules, past decisions, things we learned the hard way. Every change to how it searches goes through one benchmark before it stays. This post is the manual for that benchmark: not the numbers, but the three decisions that make the numbers mean something.
If you only remember one thing: most of the work in a retrieval benchmark is deciding when not to trust it.
1. The golden set comes from judgments we already had
We did not write test queries by hand. Every time Cortex serves context to an
agent, it logs the task the agent was working on. When the agent finishes, it
marks each item it received as helped, noise or wrong. A logged task plus
those verdicts is a relevance judgment, so the golden set is built by reading
the log as one.
The current set has 433 queries across 35 projects, with 738 documents marked useful and 252 marked noise. Getting there took four filters, and each one exists because the metric lies without it:
- Items that are always served are left out. Rules go into every brief without passing through the ranking. Counting them as hits would give the retrieval credit for something it did not do.
- A query needs at least one positive. A serve where everything was noise says nothing about where the good document should have been.
- The positive must still exist. A document that was retired is no longer in the index. Asking the search to return it subtracts recall for something it cannot do.
- Repeated tasks are counted once. Agents run the same task across sessions. Keeping the repeats inflates the sample without adding information, and gives double weight to whatever is most frequent.
One bias stays, and we keep it written down so we do not fool ourselves later: only what was served at some point has a label. If a better document existed and never made it into a brief, nobody judged it, and this benchmark cannot reward finding it. It is the classic problem with click data. The set is good for comparing two rankings on the same pairs, which is what deciding on a change requires, and not for claiming an absolute quality number.
2. Only one row is allowed to decide
The benchmark prints several configurations of the same search, and only one of them is the headline metric.
Cortex has two features that make production results better: a graph hop
that pulls in documents related to the ones found, and boosts that rank up
documents that helped before. Both are learned from the same helped / noise
verdicts the golden set is built from.
Measuring with them on is asking the exam with the answers inside. A document that helped on a query gets boosted, and then the benchmark asks that same query and rewards the boost for finding it.
So the rows are:
| row | graph | boosts | used for |
|---|---|---|---|
| core | off | off | deciding. Blind to the labels by construction |
| + graph | on | off | reference |
| + graph + boosts | on | on | what users see today — reference only |
| + graph + boosts, honest | on | on, minus the verdicts from this query's own serve | reference |
The last row is a partial leave-one-out: for each query, boosts are computed without the judgments that came from that query's serve. It is partial, and we say so: the weights of the graph edges are stored in a table and cannot be unlearned per query, so some diffuse leakage remains on that side. The direct leak, which was the strong one, is gone.
A change to the search is judged on core. If it only looks good in the rows
with boosts, it has not been shown to be good.
3. The benchmark refuses to measure
This is the part we would build first if we started over. There are four checks. In two of them the benchmark stops and exits with an error instead of printing a number; in the other two it prints the number with a warning on top. Each one comes from a time it printed a number and the number was wrong.
A search component is down. Cortex is hybrid: a lexical side (BM25) and a dense side (embeddings), fused. If one side fails, the search keeps answering with the other. That is good behavior in production and poison in a benchmark: we once measured half an hour of "hybrid retrieval" with the embedding server off, which was lexical search only. Before every run the benchmark sends a probe query and checks that every side answered. If one did not, it stops (the message, translated):
COMPONENT DOWN: dense
The system degrades silently, so the measurement would be of something else.
Measuring anyway requires a flag you have to type on purpose.
There was a second lesson inside this one. The check only works if a component can say it is down. Our lexical side caught its own errors and returned an empty list, which the search cannot tell apart from "I searched and found nothing". It kept reporting itself as healthy while contributing zero. We found out by accident while switching it off for an experiment: the run was labeled healthy, reported top-24 recall of 0.443, and was really measuring embeddings alone, at 0.371. A fallback that returns an empty result is not graceful degradation. It is a false green.
The golden set is old. The test set was judged against the documents that existed on the day it was built. When the corpus grows and the test set does not, new documents add competition without adding answers that can count as correct. We learned this when bringing the embedding index up to date added 238 documents and top-24 recall dropped, from 0.374 to 0.362, without anything getting worse. The golden set now carries a file saying when it was built and how large the corpus was. Past 30 days the benchmark refuses to run. If the corpus has moved more than 5% since, it warns. Rebuilding the set is part of reindexing.
The corpus moved between two runs. Documents get approved into Cortex while someone is measuring. Once, three were approved between two runs (193 to 196), top-5 recall moved from 0.332 to 0.377, and we spent an afternoon blaming the vector database and the embedding server before looking at the corpus. One of the new documents was about the very change being measured. Every run now saves a fingerprint of the corpus, and a comparison between runs with different fingerprints comes with a warning at the top.
The comparison is not paired. Two runs are compared query by query, with a 95% confidence interval on the differences. If the interval crosses zero, the result is "indistinguishable", whatever the averages say. This one we learned by being wrong: we announced a lexical-search upgrade as "+27% recall@5, +22% recall@24, −36% noise" by comparing means. With 56 queries at the time, the paired test said only one of those held up: top-5 recall, +0.102, interval +0.033 to +0.172. The decision was still right. The size we announced was not.
What this costs, and what it buys
None of these checks is clever. Each is a few lines: a probe query, a date check, a count, a paired difference. What they buy is that a number coming out of the benchmark has already survived the ways it has fooled us before.
Two limits apply to all of it. The test set measures what agents found useful for their tasks, not what is similar to the query, which is a stricter and different target than most public retrieval benchmarks. And the confidence intervals are only as tight as the sample: at 56 queries only effects above roughly 0.07 were detectable, which is why growing the golden set was the single change with the most leverage on everything else.
The benchmark says whether a change worked. Which change to try next, and which result would make us drop it, is a separate process, written up in How we decide what to try next in retrieval.
If you are building your own: write the gates before the first result you want to believe.