Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1325

VAmoS Energy benchmark: tougher, more realistic voice-agent simulation for billing calls

The VAmoS Part Deux paper (arXiv:2609.38512v1) introduces VAmoS Energy, a benchmark of 100 simulated utility-billing calls where each caller makes 2–4 requests and agents use 16 tools (including a Stripe billing twin and Apache Fineract). An LLM-based verifier (99.1% agreement with a code verifier) checks agent actions and spoken figures; across 14 voice stacks, completion rates ranged 17.3%–44.7% (Grok Voice highest), and background television cut pooled completion from 38.7% to 8.6%.

KEY POINTS

  1. The VAmoS Part Deux paper (arXiv:2609.38512v1) introduces VAmoS Energy, a benchmark of 100 simulated utility-billing calls where each caller makes 2–4 requests and agents use 16 tools (including a Stripe billing twin and Apache Fineract).
  2. An LLM-based verifier (99.1% agreement with a code verifier) checks agent actions and spoken figures; across 14 voice stacks, completion rates ranged 17.3%–44.7% (Grok Voice highest), and background television cut pooled completion from 38.7% to 8.6%.
  3. This shows voice-agent performance must be evaluated across the whole call (speech, actions, verification) using realistic multi-request scenarios and tool-backed state, revealing large failure modes and sensitivity to background speech.

WHY IT MATTERS

This shows voice-agent performance must be evaluated across the whole call (speech, actions, verification) using realistic multi-request scenarios and tool-backed state, revealing large failure modes and sensitivity to background speech.

SOURCES & TIMELINE

1