Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1494

Open-Endedness Bench (OEB) evaluates agents' epistemic processes from execution logs

The paper introduces OEB, a benchmark-agnostic methodology that builds an epistemic event graph from an agent's execution record (no reference answers or outcome scores) to link stated propositions to the actions that test them and to score four competence axes and six persona traits. Applied to 119 runs across 12 tasks (LLM post-training, chip design, and a training-speed record), OEB finds that only 16–29% of claimed improvements are actually supported by executed evidence and that research-behavior patterns cluster by model more than by task.

KEY POINTS

  1. The paper introduces OEB, a benchmark-agnostic methodology that builds an epistemic event graph from an agent's execution record (no reference answers or outcome scores) to link stated propositions to the actions that test them and to score four competence axes and six persona traits.
  2. Applied to 119 runs across 12 tasks (LLM post-training, chip design, and a training-speed record), OEB finds that only 16–29% of claimed improvements are actually supported by executed evidence and that research-behavior patterns cluster by model more than by task.
  3. OEB gives a reproducible, outcome-independent audit of whether agents' claims are supported by their executed experiments, exposing overclaimed improvements and providing actionable metrics for agent development and evaluation.

WHY IT MATTERS

OEB gives a reproducible, outcome-independent audit of whether agents' claims are supported by their executed experiments, exposing overclaimed improvements and providing actionable metrics for agent development and evaluation.

SOURCES & TIMELINE

1