Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #10936

LLM harnesses

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Separating coverage from specialization in LLM harnesses (arXiv:2609.35873v1)

The paper introduces a controlled evaluation that disentangles answer coverage, repeatable task advantages, and pre-execution selection for generated LLM harnesses. On 386 MATH-500 tasks, comparisons of eight generated harnesses against a baseline with nine identical copies show that repeated identical programs produce 2.16 percentage points of repeat-averaged oracle headroom; generated programs display repeatable score patterns that mainly reveal persistent weaknesses rather than useful persistent wins, the frozen selector yields 0.00 percentage-point gain, and both populations reach 98.70% oracle coverage at 27 harness executions; BIRD traces locate failures in mechanism implementation, activation, and output validity, leading the authors to argue that coverage and repeatability alone do not justify claims of useful specialization and to propose stricter evaluation standards for harness diversity.

5.0