Tech Meridian ← LIVE FEED
RU

NEWS · RESEARCH · #234

From Preferences to Principles: Rubric-based reward framework for grounded QA

The paper proposes a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training for open-domain question answering. Averaged across three evaluation axes (composition, grounding, and instruction-following), the approach yields improvements compared to a holistic scalar objective.

KEY POINTS

  1. The paper proposes a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training for open-domain question answering.
  2. Averaged across three evaluation axes (composition, grounding, and instruction-following), the approach yields improvements compared to a holistic scalar objective.
  3. Rubric-based rewards matter because they provide finer-grained, evidence-grounded supervision that can better capture multiple aspects of answer quality than a single scalar objective, improving alignment of generated answers.

WHY IT MATTERS

Rubric-based rewards matter because they provide finer-grained, evidence-grounded supervision that can better capture multiple aspects of answer quality than a single scalar objective, improving alignment of generated answers.

SOURCES & TIMELINE

1