RESEARCH · RESEARCH · #528
arXiv paper introduces SAFE benchmark to test whether frontier models seek safety evidence before acting
The paper (arXiv:2609.17865v1) introduces SAFE, a controlled benchmark where models decide whether to retrieve optional safety-relevant evidence that varies in cost, probability, severity, and presentation. Evaluating GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, the authors find distinct evidence-acquisition policies (Opus inspects by default, o3 skips most, GPT-5.5 and Sonnet intermediate), strong sensitivity to severity and retrieval cost, weaker sensitivity to probability, and a mismatch between stated rationales and actual causal influences.
KEY POINTS
- The paper (arXiv:2609.17865v1) introduces SAFE, a controlled benchmark where models decide whether to retrieve optional safety-relevant evidence that varies in cost, probability, severity, and presentation.
- Evaluating GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, the authors find distinct evidence-acquisition policies (Opus inspects by default, o3 skips most, GPT-5.5 and Sonnet intermediate), strong sensitivity to severity and retrieval cost, weaker sensitivity to probability, and a mismatch between stated rationales and actual causal influences.
- Shows that deployment-time safety depends not just on model responses to known risks but on whether models actively acquire the evidence needed to recognize those risks, which affects evaluation and mitigation strategies.
WHY IT MATTERS
Shows that deployment-time safety depends not just on model responses to known risks but on whether models actively acquire the evidence needed to recognize those risks, which affects evaluation and mitigation strategies.