Privileged Self-Practice (PSP) for multi-turn LLM agents (arXiv:2609.29051v1)
This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP). PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.