NEWS · RESEARCH · #139
τ-Elicitation: Benchmarking multi-turn entity extraction in voice agents
The paper introduces τ-Elicitation, a 200-task benchmark for evaluating exact multi-turn entity capture in voice agents across 10 entity types, controlled difficulty, caller realisms, and three environments. A matched text agent succeeds on all tasks, while four voice configurations achieve robust exact success rates of 0.14–0.41; agents verify more for hard or unfamiliar entities and sometimes for incorrect captures, but only 24–37% of verified errors are repaired, and a scaffold enforcing spelling, read-back, correction, and confirmation raises Pass^3 by 14–31 points at a 21–28 second per-call time cost.
KEY POINTS
- The paper introduces τ-Elicitation, a 200-task benchmark for evaluating exact multi-turn entity capture in voice agents across 10 entity types, controlled difficulty, caller realisms, and three environments.
- A matched text agent succeeds on all tasks, while four voice configurations achieve robust exact success rates of 0.14–0.41; agents verify more for hard or unfamiliar entities and sometimes for incorrect captures, but only 24–37% of verified errors are repaired, and a scaffold enforcing spelling, read-back, correction, and confirmation raises Pass^3 by 14–31 points at a 21–28 second per-call time cost.
- This matters because it provides a focused benchmark and measurements showing that strategy selection and recovery—not raw recognition—are primary bottlenecks for exact spoken entity collection in voice agents.
WHY IT MATTERS
This matters because it provides a focused benchmark and measurements showing that strategy selection and recovery—not raw recognition—are primary bottlenecks for exact spoken entity collection in voice agents.