Tech Meridian ← LIVE FEED
RU

NEWS · RESEARCH · #358

Blindspot benchmark for safety and refusal calibration in long-horizon tool-using agents (arXiv:2609.16305v1)

Blindspot is a new benchmark and extensible live-simulation framework for trajectory-level safety and refusal calibration of long-horizon tool-using LLM agents; its current instantiation includes 22 attack families, 35 scenarios and over 2,500 trajectories (avg. 14.7 turns) labeled with five outcomes. The paper evaluates 13 proprietary and open-weight models across eight metrics and reports substantial differences in safety–utility calibration, including failures that only emerge after several initially safe steps.

KEY POINTS

  1. Blindspot is a new benchmark and extensible live-simulation framework for trajectory-level safety and refusal calibration of long-horizon tool-using LLM agents; its current instantiation includes 22 attack families, 35 scenarios and over 2,500 trajectories (avg.
  2. 14.7 turns) labeled with five outcomes.
  3. The paper evaluates 13 proprietary and open-weight models across eight metrics and reports substantial differences in safety–utility calibration, including failures that only emerge after several initially safe steps.

WHY IT MATTERS

It reframes agent safety as a trajectory-level calibration problem and provides a standardized, extensible evaluation that reveals multi-turn safety failures and trade-offs between utility and refusal.

SOURCES & TIMELINE

1