Biosecurity at the frontier
XAI’s Grok 4.6 performed exceptionally well in the biosecurity benchmark tests released by LatchBio. In the BioSecBench-Refusal suite, Grok 4.6 was the only model that both rejected disguised dangerous tasks and successfully completed regular biological tasks, with scores above 50% in both categories. Its refusal rate for red-team tasks was 59.2%, and its completion rate for regular tasks was 64.8%, with an average score of 62.1%. In the BioSecBench-Surveillance suite, Grok 4.6 achieved a success rate of 53.5% in biological monitoring, placing it after Opus 5 and before GPT-5.6 Sol. The tests demonstrated that the model could distinguish between concealed hazards and regular scientific research, and its security was verified through multiple layers of defense mechanisms before its release.