NEWS · RESEARCH · #358
Blindspot benchmark for safety and refusal calibration in long-horizon tool-using agents (arXiv:2609.16305v1)
Blindspot is a new benchmark and extensible live-simulation framework for trajectory-level safety and refusal calibration of long-horizon tool-using LLM agents; its current instantiation includes 22 attack families, 35 scenarios and over 2,500 trajectories (avg. 14.7 turns) labeled with five outcomes. The paper evaluates 13 proprietary and open-weight models across eight metrics and reports substantial differences in safety–utility calibration, including failures that only emerge after several initially safe steps.
KEY POINTS
- Blindspot is a new benchmark and extensible live-simulation framework for trajectory-level safety and refusal calibration of long-horizon tool-using LLM agents; its current instantiation includes 22 attack families, 35 scenarios and over 2,500 trajectories (avg.
- 14.7 turns) labeled with five outcomes.
- The paper evaluates 13 proprietary and open-weight models across eight metrics and reports substantial differences in safety–utility calibration, including failures that only emerge after several initially safe steps.
WHY IT MATTERS
It reframes agent safety as a trajectory-level calibration problem and provides a standardized, extensible evaluation that reveals multi-turn safety failures and trade-offs between utility and refusal.