Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #11728

continual training

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

ArXiv preprint: agent edits its own harness via multi-task self-evolution

The paper (arXiv:2609.38372v1) proposes a framework in which a frozen language-model both solves tasks and, using the same harness, acts as a proposer that directly edits the harness that runs it; evolution draws tasks from five diverse benchmarks with strict train/held-out separation and evaluates on five additional out-of-distribution benchmarks. Starting from a 49-line seed harness and using multi-task pretraining followed by continual training, the evolved harness improves average scores by 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex in-distribution and matching it out-of-distribution; continued evolution on Claw-Eval raises that benchmark from 66.17 to 68.06, exceeding Codex; the paper analyzes emergent mechanisms such as output truncation, history compaction, and independent review.

6.0