RESEARCH · RESEARCH · #1272
RCP-nDCG@10: rubric-calibrated nDCG metric for enterprise search
Researchers introduced RCP-nDCG@10 (Rubric-Calibrated Preferences nDCG@10), an evaluation method that uses a calibrated LLM to grade every retrieved document against explicit relevance criteria rather than relying solely on sparse qrels. In a 46-person annotation study the calibrated AI judge increased AUC for predicting human relevance from ~0.65 (qrels) to ~0.91, and the metric is being used to optimize the team's next-generation search models.
KEY POINTS
- Researchers introduced RCP-nDCG@10 (Rubric-Calibrated Preferences nDCG@10), an evaluation method that uses a calibrated LLM to grade every retrieved document against explicit relevance criteria rather than relying solely on sparse qrels.
- In a 46-person annotation study the calibrated AI judge increased AUC for predicting human relevance from ~0.65 (qrels) to ~0.91, and the metric is being used to optimize the team's next-generation search models.
- It addresses coverage gaps in traditional nDCG caused by sparse qrels, giving a more complete, reproducible signal for evaluating and tuning modern enterprise retrieval systems.
WHY IT MATTERS
It addresses coverage gaps in traditional nDCG caused by sparse qrels, giving a more complete, reproducible signal for evaluating and tuning modern enterprise retrieval systems.