NEWS · MODELS · #393
LatchBio analysis finds Grok 4.6 best at refusing disguised biohazard tasks on BioSecBench
LatchBio published an independent analysis of Grok 4.6 using its BioSecBench suites. On BioSecBench-Refusal Grok 4.6 achieved the top results across harnesses (trial-weighted harmonic mean 62.1%), refusing 59.2% of red-team tasks while completing 64.8% of routine tasks (the only model >50% on both); on BioSecBench-Surveillance it averaged 53.5% success, behind Opus 5 and ahead of GPT-5.6 Sol. The report says evaluations used multiple agent harnesses and effort levels and includes additional routine biological capability results (e.g., SpatialBench, TxBench-PP) at benchmarks.bio.
KEY POINTS
- LatchBio published an independent analysis of Grok 4.6 using its BioSecBench suites.
- On BioSecBench-Refusal Grok 4.6 achieved the top results across harnesses (trial-weighted harmonic mean 62.1%), refusing 59.2% of red-team tasks while completing 64.8% of routine tasks (the only model >50% on both); on BioSecBench-Surveillance it averaged 53.5% success, behind Opus 5 and ahead of GPT-5.6 Sol.
- The report says evaluations used multiple agent harnesses and effort levels and includes additional routine biological capability results (e.g., SpatialBench, TxBench-PP) at benchmarks.bio.
WHY IT MATTERS
These findings indicate a frontier model that is comparatively well-calibrated to refuse concealed hazardous biological requests while remaining capable on routine biosurveillance tasks, which is directly relevant to AI biosecurity and deployment safeguards.