NEWS · RESEARCH · #234
From Preferences to Principles: Rubric-based reward framework for grounded QA
The paper proposes a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training for open-domain question answering. Averaged across three evaluation axes (composition, grounding, and instruction-following), the approach yields improvements compared to a holistic scalar objective.
KEY POINTS
- The paper proposes a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training for open-domain question answering.
- Averaged across three evaluation axes (composition, grounding, and instruction-following), the approach yields improvements compared to a holistic scalar objective.
- Rubric-based rewards matter because they provide finer-grained, evidence-grounded supervision that can better capture multiple aspects of answer quality than a single scalar objective, improving alignment of generated answers.
WHY IT MATTERS
Rubric-based rewards matter because they provide finer-grained, evidence-grounded supervision that can better capture multiple aspects of answer quality than a single scalar objective, improving alignment of generated answers.