RESEARCH · RESEARCH · #1325
VAmoS Energy benchmark: tougher, more realistic voice-agent simulation for billing calls
The VAmoS Part Deux paper (arXiv:2609.38512v1) introduces VAmoS Energy, a benchmark of 100 simulated utility-billing calls where each caller makes 2–4 requests and agents use 16 tools (including a Stripe billing twin and Apache Fineract). An LLM-based verifier (99.1% agreement with a code verifier) checks agent actions and spoken figures; across 14 voice stacks, completion rates ranged 17.3%–44.7% (Grok Voice highest), and background television cut pooled completion from 38.7% to 8.6%.
KEY POINTS
- The VAmoS Part Deux paper (arXiv:2609.38512v1) introduces VAmoS Energy, a benchmark of 100 simulated utility-billing calls where each caller makes 2–4 requests and agents use 16 tools (including a Stripe billing twin and Apache Fineract).
- An LLM-based verifier (99.1% agreement with a code verifier) checks agent actions and spoken figures; across 14 voice stacks, completion rates ranged 17.3%–44.7% (Grok Voice highest), and background television cut pooled completion from 38.7% to 8.6%.
- This shows voice-agent performance must be evaluated across the whole call (speech, actions, verification) using realistic multi-request scenarios and tool-backed state, revealing large failure modes and sensitivity to background speech.
WHY IT MATTERS
This shows voice-agent performance must be evaluated across the whole call (speech, actions, verification) using realistic multi-request scenarios and tool-backed state, revealing large failure modes and sensitivity to background speech.