Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #13390

benchmark invalidity

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

When Terminal-Agent Training Stalls — arXiv:2610.02405v1

This arXiv paper studies failures in training terminal agents using frontier models (e.g., Claude Opus) as meta-agents to generate tasks and verifiers, identifying three failure modes: benchmark invalidity, harness brittleness, and reward misalignment. The authors show prompt redesign and context extension can increase baseline solvability 5.6×, report that a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus–generated tasks, and that adding hard tasks can drop mean pass@2 to 20.6% without changing training configuration, arguing for solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria.

7.0