RESEARCH · RESEARCH · #1320
Sense and Sensitivity: benchmark shows LLMs more likely than physicians to recommend unnecessary care and are sensitive to text perturbations
The arXiv paper 'Sense and Sensitivity' (arXiv:2609.38600v1) introduces a benchmark of >6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses to compare LLM triage recommendations with practicing physicians under text perturbations. The study finds LLMs are more likely than physicians to recommend unnecessary care at baseline and that their recommendations are more sensitive to clinically irrelevant changes such as gender and tone perturbations.
KEY POINTS
- The arXiv paper 'Sense and Sensitivity' (arXiv:2609.38600v1) introduces a benchmark of >6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses to compare LLM triage recommendations with practicing physicians under text perturbations.
- The study finds LLMs are more likely than physicians to recommend unnecessary care at baseline and that their recommendations are more sensitive to clinically irrelevant changes such as gender and tone perturbations.
- This matters because it shows LLM clinical triage recommendations can be systematically noisier and more sensitive to irrelevant text variations than physicians, underscoring the need for deployment-oriented evaluations and safety checks.
WHY IT MATTERS
This matters because it shows LLM clinical triage recommendations can be systematically noisier and more sensitive to irrelevant text variations than physicians, underscoring the need for deployment-oriented evaluations and safety checks.