Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1336

ArXiv preprint: agent edits its own harness via multi-task self-evolution

The paper (arXiv:2609.38372v1) proposes a framework in which a frozen language-model both solves tasks and, using the same harness, acts as a proposer that directly edits the harness that runs it; evolution draws tasks from five diverse benchmarks with strict train/held-out separation and evaluates on five additional out-of-distribution benchmarks. Starting from a 49-line seed harness and using multi-task pretraining followed by continual training, the evolved harness improves average scores by 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex in-distribution and matching it out-of-distribution; continued evolution on Claw-Eval raises that benchmark from 66.17 to 68.06, exceeding Codex; the paper analyzes emergent mechanisms such as output truncation, history compaction, and independent review.

KEY POINTS

  1. The paper (arXiv:2609.38372v1) proposes a framework in which a frozen language-model both solves tasks and, using the same harness, acts as a proposer that directly edits the harness that runs it; evolution draws tasks from five diverse benchmarks with strict train/held-out separation and evaluates on five additional out-of-distribution benchmarks.
  2. Starting from a 49-line seed harness and using multi-task pretraining followed by continual training, the evolved harness improves average scores by 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex in-distribution and matching it out-of-distribution; continued evolution on Claw-Eval raises that benchmark from 66.17 to 68.06, exceeding Codex; the paper analyzes emergent mechanisms such as output truncation, history compaction, and independent review.
  3. Demonstrates a scalable approach for agents to autonomously improve their own execution harness across diverse tasks, advancing research on recursive self-improvement and agent robustness.

WHY IT MATTERS

Demonstrates a scalable approach for agents to autonomously improve their own execution harness across diverse tasks, advancing research on recursive self-improvement and agent robustness.

SOURCES & TIMELINE

1