Grok 4.6 is now the only model LatchBio found that both refuses disguised bio-risk work and still completes routine biology, scoring above 50% on each. LatchBio conducts security evaluations of AI models around biology and bio-security.
LatchBio published that eval on 1 September 2026, and xAI restated the scores the same day in an article titled: Biosecurity at the frontier, pointing readers at LatchBio’s methods rather than an in-house leaderboard.
On LatchBio’s BioSecBench-Refusal suite, Grok 4.6 was the strongest model tested at refusing disguised and hazardous tasks while still completing routine biological work. It was the only system to score above 50% on both measures.
xAI, Biosecurity at the frontier
LatchBio’s write-up of of the evaluation says Grok 4.6 now refuses more disguised bio-risk requests without tanking routine biology requests, a lift from earlier 4.6 checkpoints the lab had already tested.
Above 50 on Both
With agents run at their highest offered effort levels across several harnesses, LatchBio put Grok 4.6 at a 62.1% trial-weighted harmonic mean, taking the top three spots, including 59.2% red-team refusal and 64.8% routine completion.
BioSecBench-Refusal pairs 61 routine tasks adapted from published literature with 46 red-team tasks that look like ordinary research but hide a hazard in attached scientific data or mislabelled files. A system that only reacts to surface words will block legitimate work and miss the concealed risk; one that inspects the files can tell the two apart. LatchBio also published a public GitHub subset of the pairing.
On BioSecBench-Surveillance, which asks models to carry out pathogen genomic monitoring workflows, LatchBio recorded a 53.5% success rate for Grok 4.6, behind Opus 5 and ahead of GPT-5.6 Sol.
Reasoning, Not Filters
LatchBio says those refusals are mostly model reasoning, not API classifiers. Evaluation traces show Grok 4.6 inspecting the task and the files on disk, spotting a mismatch between a benign cover story and what is actually there, then refusing. On obviously low-risk work it uses the same inspection and proceeds.
This behavior is driven primarily by model intelligence, rather than input classifiers or system flags, unlike other models we have assessed.
LatchBio, Testing Grok 4.6’s Enhanced Biology Safeguards
Nearly all of Grok’s refusals in the lab’s runs were model-driven rather than system- or API-level blocks, regardless of harness. Other frontier systems LatchBio has assessed lean far more on provider filters, which often fire on wording before an agent can inspect the files or work it has been handed.
Routine Work Still Runs
The extra refusal did not degrade general biological performance. Grok 4.6 stayed near the top of LatchBio’s broader suite, including therapeutics, variant discovery, spatial transcriptomics, and pathogen surveillance, with no erroneous refusals on benign tasks.
Over-refusal that blocks outbreak monitoring or legitimate research is framed as equally serious as failing to stop adversarial use, and the same note records a material improvement for Grok 4.6 over Grok 4.5 and Grok 4.3 in both refusal and biosecurity work.
Making Science Faster
For labs doing cutting-edge research, this type of information can define which models they could be using. LLMs are immensely capable at pattern-recognition and repetitive data analysis, two tasks which can be exceptionally time-consuming for humans.
Using models like Grok 4.6, scientists can cut down on the time needed for them to conduct menial analysis tasks, and instead focus on what actually matters. For the average consumer, this might not seem important, but these safeguards are the line between genuine work on solving cancer, or harmful work to building bio-agents.
All-in-all, agentic AI is a tool that can better our research into biosciences, but only with the correct application.

