RESEARCH · RESEARCH · #1360
RLTL;DR — self-improvement by internalizing self-generated feedback
RLTL;DR is a reinforcement-learning method where, after each failed attempt, the agent sees the verifier output and writes a single TL;DR insight; subsequent rollouts are conditioned on accumulated insights and the training pipeline backpropagates through those in-context insights to internalize a task→insight mapping. On hard tool-calling and coding datasets filtered to Pass@128 = 0, GRPO training of a Qwen 3.5 9B Thinking policy stays at Pass@1 ≈ 0–1%, while RLTL;DR achieves Pass@1 of 14–31% when insights are in context during training and still 12–13% at evaluation with no insight; a reduced variant (SFTL;DR) trained only on 4k (task, insight) tuples recovers almost the full performance, suggesting a compact “keep this in mind” training paradigm.
KEY POINTS
- RLTL;DR is a reinforcement-learning method where, after each failed attempt, the agent sees the verifier output and writes a single TL;DR insight; subsequent rollouts are conditioned on accumulated insights and the training pipeline backpropagates through those in-context insights to internalize a task→insight mapping.
- On hard tool-calling and coding datasets filtered to Pass@128 = 0, GRPO training of a Qwen 3.5 9B Thinking policy stays at Pass@1 ≈ 0–1%, while RLTL;DR achieves Pass@1 of 14–31% when insights are in context during training and still 12–13% at evaluation with no insight; a reduced variant (SFTL;DR) trained only on 4k (task, insight) tuples recovers almost the full performance, suggesting a compact “keep this in mind” training paradigm.
- Shows that agents can learn to self-generate and internalize compact feedback that unlocks substantial gains on tasks with near-zero initial success, pointing to a practical compact training route for self-improvement.
WHY IT MATTERS
Shows that agents can learn to self-generate and internalize compact feedback that unlocks substantial gains on tasks with near-zero initial success, pointing to a practical compact training route for self-improvement.