Cohere proposes RCP-nDCG@10: a calibrated AI judge replaces fixed answer keys, matching human preference in 77% of 289 blind contests versus 52% for conventional nDCG
Synopsis
Cohere introduces RCP-nDCG@10, a retrieval evaluation methodology in which a calibrated AI judge grades every retrieved document against the same explicit relevance rubric and combines rubric answers with group-wise comparisons and calibration, so relevant documents absent from the answer key can still earn credit; in a blind study with 46 annotators, 273 queries and 289 head-to-head contests, reviewers sided with RCP-nDCG 70% of the time where the two metrics named different winners, and across all contests RCP-nDCG matched reviewer preference 77% of the time versus 52% for conventional nDCG, with the gain driven mainly by crediting relevant documents the answer key missed (19 percentage points).
Interpretation
RCP-nDCG@10 uses a calibrated AI judge to assess every retrieved document against the same explicit relevance rubric, letting relevant documents not in the answer key receive credit while preserving the ordering objective of traditional nDCG. Traditional nDCG only compares results against existing qrels, typically built by human assessors reviewing a subset of documents surfaced by earlier retrieval systems; RCP-nDCG@10 instead scores each retrieved result on its actual relevance. The text illustrates this with a real ViDoRe v3 query: the answer key grades only two documents and only one appears in the system's top five; the document at rank 3 quotes the answer directly but was never judged, so it is not considered relevant, and the resulting nDCG@5 of 0.760 reflects just one of the five results.
Each score is built from two signals: the judge answers five yes/no questions of varying difficulty, identical for every query, and also compares documents against each other in groups to estimate how likely each is to beat the others; comparisons set the order and the rubric places it on a scale shared by every query. This goes beyond asking an LLM to grade each document, since a single grade from an LLM judge cannot yield a relevance score that means the same thing for every query. The text describes calibration as combining the two signals so the relevance score carries the same meaning across queries; this is a design-level account and no separate ablation numbers for the calibration step are given.
In the blind human study, RCP-nDCG tracked human preference more closely than conventional nDCG: where the two named different winners, reviewers sided with RCP-nDCG 70% of the time, and across all contests RCP-nDCG picked the system reviewers preferred 77% of the time versus 52% for conventional nDCG. The text attributes the advantage to giving credit to relevant documents the answer key missed (19 percentage points), with grading how relevant each document is adding a further 6; the two effects interact, since grading alone added only 3.3 points because it could only re-weight documents the key already lists. 46 contracted annotators, each with at least a BSc in the relevant field, graded the top-five results of competing systems on NanoBEIR, BRIGHT and ViDoRe v3 queries without seeing system names, rankings or the answer key; three annotators reviewed every contest with the majority deciding the winner, and 289 contests across 273 queries reached a verdict; the pool deliberately over-samples contests where the two metrics disagree, and where they agree both match reviewers 87% of the time.
The margin RCP-nDCG reports is informative: the bigger the margin between two systems, the more often reviewers agree, reaching 97% for the widest margins; above a margin of about 0.02 reviewers agree 82% of the time, and below that agreement is no better than a coin toss. A bigger margin under conventional nDCG does not help in the same way: at the widest margins reviewers agree with it only 53% of the time. This comes from the same set of 289 blind contests (Figure 5), and where the metrics disagree reviewers back RCP-nDCG in 64% to 78% of contests at every level of answer-key coverage (Figure 6).
Perspective
The work targets enterprise retrieval quality evaluation, suited to settings where several retrieval systems must be compared and answer-key coverage may be incomplete, such as MTEB- and BEIR-style benchmarks and model selection comparisons. It lets relevant documents absent from the answer key earn credit, so systems can still be separated once model capability exceeds what the old labels can express; on this basis the authors chose RCP-nDCG@10 rather than traditional nDCG as the optimization target for their next-generation Embed and Rerank models, and note those models may not always look strongest under legacy nDCG. For readers, this means that when an evaluation conclusion conflicts with an older metric, it is worth first checking whether the metric in use covers relevant documents that were never judged.
Readers should still note that RCP-nDCG depends on a calibrated AI judge, and the text does not provide a separate assessment of that judge's quality or calibration stability; the human study deliberately over-samples contests where the two metrics disagree, so 77% and 52% should not be read as expected agreement rates in general settings. The text mentions other studies reporting inter-annotator agreement as low as 17% when re-annotating a standard benchmark, and this study reports exact agreement of only 42% between two reviewers grading the same document on a 0-4 scale, indicating that relevance judgments are themselves subjective and any single metric has an expressive ceiling. In addition, this is a methodology-introduction article that does not lay out the technical details of the underlying paper, so reproduction should go through that paper and the linked code and data.
