Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

NEWS · MODELS · #1435

Mercor benchmark: Claude Opus 5.5 and Fable 5.1 outpace CPAs on structured bookkeeping but can't close the books

A Mercor study using tasks from the APEX Accounting Benchmark found modern LLMs much faster, more accurate, and cheaper than 12 licensed CPAs on simplified structured bookkeeping tasks. Claude Opus 5.5 led with 61.8% of grading criteria met, followed by Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%; however Mercor reports no model fully solved nearly 60% of the full-task set and models still require oversight for closing the books, with the study also excluding client interaction and long-term contextual work.

KEY POINTS

  1. A Mercor study using tasks from the APEX Accounting Benchmark found modern LLMs much faster, more accurate, and cheaper than 12 licensed CPAs on simplified structured bookkeeping tasks.
  2. Claude Opus 5.5 led with 61.8% of grading criteria met, followed by Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%; however Mercor reports no model fully solved nearly 60% of the full-task set and models still require oversight for closing the books, with the study also excluding client interaction and long-term contextual work.
  3. The results signal major productivity gains for structured accounting tasks and growing model parity with human accountants, while underscoring that oversight and non-structured work still limit full automation.

WHY IT MATTERS

The results signal major productivity gains for structured accounting tasks and growing model parity with human accountants, while underscoring that oversight and non-structured work still limit full automation.

SOURCES & TIMELINE

1