Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #527

ERPBench evaluates screenshot-based agents on a live ERP system (arXiv:2609.17885v1)

ERPBench is a new benchmark and production-grade harness that evaluates screenshot-only computer-use agents on a live, reproducible ERP system and verifies task success against ground-truth database records. The authors run six closed- and open-source agents and find major enterprise-specific failure modes: agents often complete and save forms (up to 85% of runs) but frequently write incorrect database values (as low as 3% correct), and the paper documents a human-approval gating harness for safer autonomous runs.

KEY POINTS

  1. ERPBench is a new benchmark and production-grade harness that evaluates screenshot-only computer-use agents on a live, reproducible ERP system and verifies task success against ground-truth database records.
  2. The authors run six closed- and open-source agents and find major enterprise-specific failure modes: agents often complete and save forms (up to 85% of runs) but frequently write incorrect database values (as low as 3% correct), and the paper documents a human-approval gating harness for safer autonomous runs.
  3. ERPBench matters because it provides a live, reproducible evaluation and a human-gated harness that expose severe reliability gaps of GUI agents in enterprise workflows, which is critical for safe deployment and benchmarking.

WHY IT MATTERS

ERPBench matters because it provides a live, reproducible evaluation and a human-gated harness that expose severe reliability gaps of GUI agents in enterprise workflows, which is critical for safe deployment and benchmarking.

SOURCES & TIMELINE

1