Search: three decisions, measured first
From the energy client project. Every change was measured on the same test questions before it went live.
- 55% → 70%Keyword and vector search together found the right source more often than keyword alone. On questions in the user’s own words: 36% → 58%.
- 70% vs 69%A data rule forced a new embedding model. We measured both first. The switch cost nothing.
- 66% → 77%The cloud’s built-in reranker added nothing. A different one lifted the hit rate by 11 points.
How we measure search: recall, MRR, nDCG and the new judge-based metrics. Read below ↓
How we measure search
One number is not enough. Each metric asks a different question, and each one has a blind spot. So we use them together.
-
1
Did we find it at all?
Recall@10 · hit rateIs the right passage in the top 10 results? The simplest number, and the one that tells you if a change helps.
- Needs
- A gold set: test questions with their known source passages.
- Blind spot
- The order inside the top 10. And a good passage that is not in the gold set counts as wrong.
-
2
Is it near the top?
MRR · nDCG@10MRR: how high is the first right passage? nDCG: are the right passages near the top, and do very relevant ones come before slightly relevant ones?
- Needs
- A gold set with grades: very relevant, a bit relevant, not relevant.
- Blind spot
- The same as recall. A passage that nobody labelled counts as wrong.
-
3
Is what we found really useful?
AI judge · RCP-nDCGNewAn AI judge grades every passage the search returns, also the ones nobody labelled. RCP-nDCG, from Cohere, makes that judge consistent: five yes/no questions for every passage, plus head-to-head comparisons. The comparisons set the order. The yes/no questions put it on one scale for all questions.
Why it matters: in Cohere’s study, 28% of the passages that a benchmark labelled “not relevant” were useful to people.
- Needs
- A judge that is first checked against people, on your questions.
- Blind spot
- It is only as good as the judge. That is why the judge is checked first.
The gold set tells you if a change helps. The judge tells you what the gold set missed. For the answer itself there are two more checks: does every cited source exist and say what the answer claims, and which of two answers is better, judged blind.
Levels 1 and 2 and both answer checks are in use today. We are adding RCP-nDCG now. Source: Cohere, RCP-nDCG. In their study, 46 people compared search results in 289 contests; RCP-nDCG picked the result people preferred 77% of the time, classic nDCG 52%.
Ask for a 30-minute call
No slides. No demo. Five questions on your system, and an honest answer on fit.
Ask for a 30-minute call