·7 min read·Mounaji Studio

How we decide what to try next in retrieval: a queue of written predictions

Our retrieval roadmap is a queue of hypotheses, each with its prediction and the result that would falsify it written down before measuring. How one number ordered the queue, and how the queue reordered itself when the first bet lost.

engineeringretrievalragevaluation

Once a retrieval system works, the list of things you could try is long: rerankers, a better embedding model, chunking, query rewriting, stemming, tuning the fusion. Every one of them has a blog post saying it helped someone. The hard part is not finding ideas. It is choosing the order, and knowing when to stop believing in one.

For Cortex, the knowledge system that serves context to our coding agents, we keep that list as a queue of hypotheses. This post is how the queue works: what goes into an entry, what decides the order, and what happens when an entry is proven wrong. The benchmark behind the numbers is described in How we evaluate retrieval.

An entry is an experiment, not a task

"Add a reranker" is a task. It is done when the code is merged, and nothing about it can be wrong. An entry in the queue has four parts instead:

  1. The hypothesis, in one sentence, with the reason behind it.
  2. The prediction: which metric moves, by how much, and which metrics must not move.
  3. What would falsify it, written as a threshold, not as a mood.
  4. What happens to the queue if it is falsified: which entries move up, decided now, while nobody has a result to defend.

The prediction is committed to the repository before the first measurement runs. The commit history is the proof: the file with the prediction has an earlier timestamp than the file with the result. Writing the prediction after seeing the number is not research; it is narration.

One number set the order

Our benchmark has 433 real queries, each with the documents an agent marked as useful for its task. On the configuration we use to decide, it measured:

metric value
recall@5 0.233
recall@24 0.473

recall@k is the share of useful documents that appear in the first k results.

Context is served as a window of 24 results, but the top 5 is where an agent pays attention. Recall at 24 was roughly double recall at 5. Read plainly: half of what was useful was already being found, and was sitting too low. That describes an ordering problem, not a finding problem.

That one reading ordered the whole queue. The technique built for ordering problems is a cross-encoder reranker, which reads the query and each candidate together and reorders the window. It went first. Changing the embedding model went last on purpose: when we turned each side of our hybrid search off separately, the embedding side added depth but nothing measurable in the top 5, so a better model was the most expensive experiment with the least evidence behind it.

The queue, as it was written that day:

# hypothesis attacks cost
1 Cross-encoder reranker over the top 24 order medium
2 Spanish stemming (propuesta ≠ propuestas today) what gets found low
3 Chunking long documents what gets found low
4 Query expansion with an LLM what gets found medium
5 A different embedding model what gets found high
6 Sweep the fusion constant and weights order low

The column that matters most is attacks. Every technique on the list either changes which documents enter the window or changes the order inside it. The number told us which of those two problems we had, so it told us which half of the list to start with.

The first entry, as written

The reranker's prediction file said, before any run:

  • recall@24 does not move at all. A reranker reorders the window without changing who is in it. If this number moves, the experiment is wired wrong and nothing else in the run counts. This is a sanity check, not a result.
  • recall@5 goes from 0.233 to 0.34, with a paired 95% confidence interval that does not cross zero. This is the number that decides.
  • MRR and nDCG@10 rise to about 0.30 and 0.29.
  • Latency is the real cost, and even with every metric in favor, it ships behind a flag that is off by default.
  • Falsified if recall@5 rises less than 0.05 or its interval crosses zero. In that case the premise "good documents are found but badly ordered" is false or insufficient, and entries 2 and 4, which attack what gets found, move up.

What the measurement said

Same day, same corpus, same 433 queries, compared query by query:

The reranker's written prediction against the measurement: recall@5, MRR and nDCG@10 were predicted to rise and all three fell

metric before with reranker difference 95% interval
recall@5 0.233 0.192 −0.041 −0.070 to −0.011
MRR 0.208 0.180 −0.028 −0.048 to −0.007
nDCG@10 0.203 0.178 −0.025 −0.044 to −0.007
recall@24 0.473 0.473 0.000 identical

The sanity check held: recall@24 did not move, so the run was valid. The decisive metric moved the wrong way, and the interval says the drop is real. Each query also took about 6.9 seconds instead of 0.3.

The entry was closed as false. Nobody had to argue about what came next: the prediction file had already said it. Stemming and query expansion moved up, the reranker code stayed in the repository switched off, with its tests and its measurement next to it, so nobody rebuilds it later thinking it is the obvious fix.

Why the queue gets better when an entry loses

A falsified entry leaves something behind that a confirmed one does not: a new question, sharper than the old one. Our best explanation for the reranker's result is that our labels mean useful for the task, not similar to the query, and a generic reranker is trained on the second. That explanation is not verified yet, so it went into the queue as its own entry, cheap to test: the same reranker, reading the same enriched text the index sees instead of the raw document. If it still loses, the finding is larger than any reranker: on this corpus, ranking by similarity has a ceiling, and what pays is learning from the judgments. One data point already leans that way: the configuration that ranks up documents that helped before scores +0.056 recall@5 over the one that does not.

That is the behavior we want from a roadmap. A list of tasks only grows. A queue of predictions gets reordered by its own results, and each result narrows what the next experiment has to answer.

If you want to keep one

  • Measure before you order. Find the one reading that splits your list in two. For us it was recall@24 against recall@5: found-but-low versus not found. Yours may be different; the point is that a number orders the list, not a preference.
  • Tag every idea with what it attacks. Order or what gets found. Cost goes next to it, so an expensive idea with little evidence sinks on its own.
  • Write the falsifying threshold before running. "Rises less than 0.05, or the interval crosses zero" is a threshold. "Doesn't help much" is not.
  • Write the reorder in advance. The decision about what comes next is cheapest before anyone has a result they would like to be true.
  • Include one metric that must not move. It is the only way to tell a bad idea from a broken experiment.
  • Do not change the corpus while measuring. Between two of our runs, one document was approved into the knowledge base: the note about the change being measured. The benchmark flagged it and we reran both sides back to back. The system that learns and the system under test were the same system.

Where these numbers come from

The prediction, the runs and the per-query results live in our repository: the prediction file was committed before the measurement, and each run is saved with its corpus fingerprint so the intervals can be recomputed. The same limits as in our other retrieval posts apply: one corpus, heavy on identifiers, with labels that mean useful rather than similar. The queue will come out differently on your data. The method should not.

Chat with us