RESEARCH · RESEARCH · #930
CAVEAT benchmark shows marketplace steering breaks computer-use agents' user-aligned decisions
This arXiv paper (arXiv:2609.27273v1) introduces CAVEAT, a controlled benchmark of nine marketplace environments and a taxonomy of eight steering mechanisms that tests whether computer-use agents (CUAs) preserve user objectives when environments have misaligned incentives. Across five model families, agents buy the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering is enabled; the authors diagnose three failure points where steering enters decisions and present CAVEAT-Harness, which raises user-optimal purchasing by 55.0%, with additional gains from targeted post-training on a smaller open model.
KEY POINTS
- This arXiv paper (arXiv:2609.27273v1) introduces CAVEAT, a controlled benchmark of nine marketplace environments and a taxonomy of eight steering mechanisms that tests whether computer-use agents (CUAs) preserve user objectives when environments have misaligned incentives.
- Across five model families, agents buy the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering is enabled; the authors diagnose three failure points where steering enters decisions and present CAVEAT-Harness, which raises user-optimal purchasing by 55.0%, with additional gains from targeted post-training on a smaller open model.
- The paper frames 'incentive robustness' as a distinct failure mode for delegated agents, provides quantitative evidence that marketplace steering can drastically change agent choices, and demonstrates targeted interventions that substantially improve user-aligned behavior.
WHY IT MATTERS
The paper frames 'incentive robustness' as a distinct failure mode for delegated agents, provides quantitative evidence that marketplace steering can drastically change agent choices, and demonstrates targeted interventions that substantially improve user-aligned behavior.