Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #8441

self-distillation

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · Apple Machine Learning Research

RISED: Rubric-guided data selection and self-distillation for multi-environment LLM agents

Authors introduce RISED, a method that uses a shared rubric vocabulary and an LLM judge to tag rollouts across diverse interactive environments; rubric profiles guide online data selection while positive rubrics provide privileged context for on-policy self-distillation and negative rubrics steer future rollout generation away from recurring failures. The paper reports that, across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, and includes rubric-based analyses of behavioural change.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Privileged Self-Practice (PSP) for multi-turn LLM agents (arXiv:2609.29051v1)

This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP). PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.

7.0