Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1487

Google researchers propose RRSI to stop agent harnesses memorizing tests

Google Cloud AI Research and academic collaborators introduce RRSI (Regularized Recursive Self-Improvement), a method that constrains automated harness rewrites so agentic systems avoid overfitting to benchmark tasks. Tested with a frozen Claude Opus 4.8 across eight benchmarks (including JobBench and ARC-AGI-3), RRSI improved training-task scores by up to 14.1 points, boosted performance on unseen tests by up to 4.7 points, and cut runtime token use by about 30%; the paper and code are on GitHub.

KEY POINTS

  1. Google Cloud AI Research and academic collaborators introduce RRSI (Regularized Recursive Self-Improvement), a method that constrains automated harness rewrites so agentic systems avoid overfitting to benchmark tasks.
  2. Tested with a frozen Claude Opus 4.8 across eight benchmarks (including JobBench and ARC-AGI-3), RRSI improved training-task scores by up to 14.1 points, boosted performance on unseen tests by up to 4.7 points, and cut runtime token use by about 30%; the paper and code are on GitHub.
  3. Constraining self-improvement reduces benchmark memorization and improves generalization and efficiency for agent harness optimization, addressing a key failure mode of automated recursive tuning.

WHY IT MATTERS

Constraining self-improvement reduces benchmark memorization and improves generalization and efficiency for agent harness optimization, addressing a key failure mode of automated recursive tuning.

SOURCES & TIMELINE

1