Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #11960

Eleonora Gualdoni

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · Apple Machine Learning Research

RLTL;DR — self-improvement by internalizing self-generated feedback

RLTL;DR is a reinforcement-learning method where, after each failed attempt, the agent sees the verifier output and writes a single TL;DR insight; subsequent rollouts are conditioned on accumulated insights and the training pipeline backpropagates through those in-context insights to internalize a task→insight mapping. On hard tool-calling and coding datasets filtered to Pass@128 = 0, GRPO training of a Qwen 3.5 9B Thinking policy stays at Pass@1 ≈ 0–1%, while RLTL;DR achieves Pass@1 of 14–31% when insights are in context during training and still 12–13% at evaluation with no insight; a reduced variant (SFTL;DR) trained only on 4k (task, insight) tuples recovers almost the full performance, suggesting a compact “keep this in mind” training paradigm.

8.0