Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1001

Privileged Self-Practice (PSP) for multi-turn LLM agents (arXiv:2609.29051v1)

This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP). PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.

KEY POINTS

  1. This arXiv preprint identifies a failure mode of on-policy self-distillation (OPSD) in multi-turn agents—training yields overconfident behavior without the underlying information—and proposes Privileged Self-Practice (PSP).
  2. PSP keeps privileged information (PI) in the prompt/sampler rather than the loss: when a student fails rollouts, an analyzer model injects a short per-task instruction into the prompt, the task is re-sampled, and the result is trained with the same GRPO objective; across AppWorld and SWE-bench Verified and three student models PSP consistently outperforms OPSD and plain GRPO, improving task-goal completion by up to 65% on AppWorld and resolved rate by up to 61% on SWE-bench Verified.
  3. This matters because it exposes a critical limitation of OPSD in multi-turn settings and offers a simple, empirically effective fix that consistently outperforms prior self-distillation and plain RL baselines on benchmarks.

WHY IT MATTERS

This matters because it exposes a critical limitation of OPSD in multi-turn settings and offers a simple, empirically effective fix that consistently outperforms prior self-distillation and plain RL baselines on benchmarks.

SOURCES & TIMELINE

1