RESEARCH · RESEARCH · #1627
RISED: Rubric-guided data selection and self-distillation for multi-environment LLM agents
Authors introduce RISED, a method that uses a shared rubric vocabulary and an LLM judge to tag rollouts across diverse interactive environments; rubric profiles guide online data selection while positive rubrics provide privileged context for on-policy self-distillation and negative rubrics steer future rollout generation away from recurring failures. The paper reports that, across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, and includes rubric-based analyses of behavioural change.
KEY POINTS
- Authors introduce RISED, a method that uses a shared rubric vocabulary and an LLM judge to tag rollouts across diverse interactive environments; rubric profiles guide online data selection while positive rubrics provide privileged context for on-policy self-distillation and negative rubrics steer future rollout generation away from recurring failures.
- The paper reports that, across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, and includes rubric-based analyses of behavioural change.
- This matters because rubric-based textual feedback supplies richer cross-environment and within-group supervision than scalar rewards, enabling better data selection and policy learning for generalist LLM agents.
WHY IT MATTERS
This matters because rubric-based textual feedback supplies richer cross-environment and within-group supervision than scalar rewards, enabling better data selection and policy learning for generalist LLM agents.