RESEARCH · RESEARCH · #1001
Privileged Self-Practice (PSP) for multi-turn LLM agents (arXiv:2609.29051v1)
This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP). PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.
KEY POINTS
- This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP).
- PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.
- This matters because it exposes a critical limitation of OPSD in multi-turn settings and offers a simple, empirically effective fix that consistently outperforms prior self-distillation and plain RL baselines on benchmarks.
WHY IT MATTERS
This matters because it exposes a critical limitation of OPSD in multi-turn settings and offers a simple, empirically effective fix that consistently outperforms prior self-distillation and plain RL baselines on benchmarks.