RESEARCH · RESEARCH · #1511
When Terminal-Agent Training Stalls — arXiv:2610.02405v1
This arXiv paper studies failures in training terminal agents using frontier models (e.g., Claude Opus) as meta-agents to generate tasks and verifiers, identifying three failure modes: benchmark invalidity, harness brittleness, and reward misalignment. The authors show prompt redesign and context extension can increase baseline solvability 5.6×, report that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus–generated tasks, and that adding hard tasks can drop mean pass@2 to 20.6% without changing training configuration, arguing for solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria.
KEY POINTS
- This arXiv paper studies failures in training terminal agents using frontier models (e.g., Claude Opus) as meta-agents to generate tasks and verifiers, identifying three failure modes: benchmark invalidity, harness brittleness, and reward misalignment.
- The authors show prompt redesign and context extension can increase baseline solvability 5.6×, report that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus–generated tasks, and that adding hard tasks can drop mean pass@2 to 20.6% without changing training configuration, arguing for solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria.
- The paper highlights that task/verifier generation and test harness issues can systematically distort terminal-agent training results, so evaluation must include solvability calibration and infrastructure checks.
WHY IT MATTERS
The paper highlights that task/verifier generation and test harness issues can systematically distort terminal-agent training results, so evaluation must include solvability calibration and infrastructure checks.