Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #12594

APEX Accounting Benchmark

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

MODELS · 1 SOURCE · The Decoder

Mercor benchmark: Claude Opus 5.5 and Fable 5.1 outpace CPAs on structured bookkeeping but can't close the books

A Mercor study using tasks from the APEX Accounting Benchmark found modern LLMs much faster, more accurate, and cheaper than 12 licensed CPAs on simplified structured bookkeeping tasks. Claude Opus 5.5 led with 61.8% of grading criteria met, followed by Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%; however Mercor reports no model fully solved nearly 60% of the full-task set and models still require oversight for closing the books, with the study also excluding client interaction and long-term contextual work.

6.0