Tagged “retrieval”
2 posts
How we decide what to try next in retrieval: a queue of written predictions
Our retrieval roadmap is a queue of hypotheses, each with its prediction and the result that would falsify it written down before measuring. How one number ordered the queue, and how the queue reordered itself when the first bet lost.
How we evaluate retrieval: a golden set, a row with no leakage, and a benchmark that refuses to run
Three design decisions in our retrieval benchmark that are not obvious, and that anyone who inherits it needs to understand. Where the test set comes from, which configuration is allowed to decide, and when the benchmark says no.