RESEARCH · RESEARCH · #527
ERPBench evaluates screenshot-based agents on a live ERP system (arXiv:2609.17885v1)
ERPBench is a new benchmark and production-grade harness that evaluates screenshot-only computer-use agents on a live, reproducible ERP system and verifies task success against ground-truth database records. The authors run six closed- and open-source agents and find major enterprise-specific failure modes: agents often complete and save forms (up to 85% of runs) but frequently write incorrect database values (as low as 3% correct), and the paper documents a human-approval gating harness for safer autonomous runs.
KEY POINTS
- ERPBench is a new benchmark and production-grade harness that evaluates screenshot-only computer-use agents on a live, reproducible ERP system and verifies task success against ground-truth database records.
- The authors run six closed- and open-source agents and find major enterprise-specific failure modes: agents often complete and save forms (up to 85% of runs) but frequently write incorrect database values (as low as 3% correct), and the paper documents a human-approval gating harness for safer autonomous runs.
- ERPBench matters because it provides a live, reproducible evaluation and a human-gated harness that expose severe reliability gaps of GUI agents in enterprise workflows, which is critical for safe deployment and benchmarking.
WHY IT MATTERS
ERPBench matters because it provides a live, reproducible evaluation and a human-gated harness that expose severe reliability gaps of GUI agents in enterprise workflows, which is critical for safe deployment and benchmarking.