Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #7744

language-model agent

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

ArXiv preprint: agent edits its own harness via multi-task self-evolution

The paper (arXiv:2609.38372v1) proposes a framework in which a frozen language-model both solves tasks and, using the same harness, acts as a proposer that directly edits the harness that runs it; evolution draws tasks from five diverse benchmarks with strict train/held-out separation and evaluates on five additional out-of-distribution benchmarks. Starting from a 49-line seed harness and using multi-task pretraining followed by continual training, the evolved harness improves average scores by 4.48 points in-distribution and 12.64 points out-of-distribution, surpassing Codex in-distribution and matching it out-of-distribution; continued evolution on Claw-Eval raises that benchmark from 66.17 to 68.06, exceeding Codex; the paper analyzes emergent mechanisms such as output truncation, history compaction, and independent review.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Propose, Don't Judge (arXiv:2609.27051v1): introduce a frozen, anytime-valid referee for LLM agents in factor research

The paper proposes an architecture where language-model agents freely propose investment factors while a frozen statistical referee—untouchable by the agent—judges candidates using betting on market outcomes revealed after submission, yielding an anytime-valid false-discovery guarantee. In synthetic tests and a ten-year walk-forward on the CSI 500 the frozen referee admitted 5–11× fewer sub-threshold factors under a scripted proposer, language-model proposers produced higher yield and custom probes, but certified true factors typically waited ~500 trading days and certified portfolios had lower realized Sharpe than ungated portfolios.

6.0