Tech Meridian ← LIVE FEED
RU

NEWS · MODELS · #393

LatchBio analysis finds Grok 4.6 best at refusing disguised biohazard tasks on BioSecBench

LatchBio published an independent analysis of Grok 4.6 using its BioSecBench suites. On BioSecBench-Refusal Grok 4.6 achieved the top results across harnesses (trial-weighted harmonic mean 62.1%), refusing 59.2% of red-team tasks while completing 64.8% of routine tasks (the only model >50% on both); on BioSecBench-Surveillance it averaged 53.5% success, behind Opus 5 and ahead of GPT-5.6 Sol. The report says evaluations used multiple agent harnesses and effort levels and includes additional routine biological capability results (e.g., SpatialBench, TxBench-PP) at benchmarks.bio.

KEY POINTS

  1. LatchBio published an independent analysis of Grok 4.6 using its BioSecBench suites.
  2. On BioSecBench-Refusal Grok 4.6 achieved the top results across harnesses (trial-weighted harmonic mean 62.1%), refusing 59.2% of red-team tasks while completing 64.8% of routine tasks (the only model >50% on both); on BioSecBench-Surveillance it averaged 53.5% success, behind Opus 5 and ahead of GPT-5.6 Sol.
  3. The report says evaluations used multiple agent harnesses and effort levels and includes additional routine biological capability results (e.g., SpatialBench, TxBench-PP) at benchmarks.bio.

WHY IT MATTERS

These findings indicate a frontier model that is comparatively well-calibrated to refuse concealed hazardous biological requests while remaining capable on routine biosurveillance tasks, which is directly relevant to AI biosecurity and deployment safeguards.

SOURCES & TIMELINE

1