arXiv paper introduces SAFE benchmark to test whether frontier models seek safety evidence before acting
The paper (arXiv:2609.17865v1) introduces SAFE, a controlled benchmark where models decide whether to retrieve optional safety-relevant evidence that varies in cost, probability, severity, and presentation. Evaluating GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, the authors find distinct evidence-acquisition policies (Opus inspects by default, o3 skips most, GPT-5.5 and Sonnet intermediate), strong sensitivity to severity and retrieval cost, weaker sensitivity to probability, and a mismatch between stated rationales and actual causal influences.