How we measure search

One number is not enough. Each metric asks a different question, and each one has a blind spot. So we use them together.

  1. 1

    Did we find it at all?

    Recall@10 · hit rate

    Is the right passage in the top 10 results? The simplest number, and the one that tells you if a change helps.

    Needs
    A gold set: test questions with their known source passages.
    Blind spot
    The order inside the top 10. And a good passage that is not in the gold set counts as wrong.
  2. 2

    Is it near the top?

    MRR · nDCG@10

    MRR: how high is the first right passage? nDCG: are the right passages near the top, and do very relevant ones come before slightly relevant ones?

    Needs
    A gold set with grades: very relevant, a bit relevant, not relevant.
    Blind spot
    The same as recall. A passage that nobody labelled counts as wrong.
  3. 3

    Is what we found really useful?

    AI judge · RCP-nDCGNew

    An AI judge grades every passage the search returns, also the ones nobody labelled. RCP-nDCG, from Cohere, makes that judge consistent: five yes/no questions for every passage, plus head-to-head comparisons. The comparisons set the order. The yes/no questions put it on one scale for all questions.

    Why it matters: in Cohere’s study, 28% of the passages that a benchmark labelled “not relevant” were useful to people.

    Needs
    A judge that is first checked against people, on your questions.
    Blind spot
    It is only as good as the judge. That is why the judge is checked first.

The gold set tells you if a change helps. The judge tells you what the gold set missed. For the answer itself there are two more checks: does every cited source exist and say what the answer claims, and which of two answers is better, judged blind.

Levels 1 and 2 and both answer checks are in use today. We are adding RCP-nDCG now. Source: Cohere, RCP-nDCG. In their study, 46 people compared search results in 289 contests; RCP-nDCG picked the result people preferred 77% of the time, classic nDCG 52%.

Ask for a 30-minute call

No slides. No demo. Five questions on your system, and an honest answer on fit.

Ask for a 30-minute call