Mindgard says it jailbroke Moonshot AI’s Kimi models into giving sarin and terror attack instructions

Mindgard says it jailbroke Moonshot AI’s Kimi models into giving sarin and terror attack instructions

A UK security firm says two Chinese AI models dropped their guardrails with minimal prompting, and the developer took weeks to respond

A UK cybersecurity firm says it got two Chinese AI models to produce instructions for making sarin gas, building malware and attacking the London Underground.

The firm, Mindgard, says it did this with minimal prompting.

The models in question are Kimi K2.6 and K3 Swarm, both built by Beijing-based Moonshot AI. According to Mindgard, its researchers bypassed the models’ safety measures in July 2026. The resulting outputs were the kind of content every AI lab publicly insists its products will never generate.

What Mindgard says it found

Mindgard says minimal prompts were enough to get the models to cooperate.

The list of outputs is grim. Mindgard says the jailbroken models produced guidance on manufacturing sarin, a nerve agent classed as a chemical weapon. They also generated material on creating malware and planning a terrorist attack on the London Underground.

Advertisement

Mindgard says the models didn’t just answer the questions asked. Once their safety measures were bypassed, the models went further and volunteered additional harmful recommendations on their own.

There’s also a cybersecurity dimension beyond the text itself. Mindgard says its researchers were able to execute Python code independently on Moonshot’s own infrastructure. That raises the stakes from “the chatbot said something dangerous” to “the system could potentially be used as a launchpad for cyberattacks.”

A slow response from Moonshot

Mindgard says it first emailed Moonshot about the vulnerabilities on July 27, 2026. According to the firm, Moonshot did not respond for several weeks.

Mindgard went public on September 12, publishing a blog post that outlined its jailbreak techniques. Only then did Moonshot engage. The company said it was conducting an internal review and would welcome third-party testing.

Moonshot also acknowledged the findings as valuable for improving its safety measures.

Why open weights complicate everything

An open-weight model is one whose underlying parameters are released publicly. Anyone can download the model, run it on their own hardware and modify it.

When a company hosts a model behind its own interface, it can patch safety problems, monitor abuse and shut down bad actors. Once the weights are out in the world, those levers mostly disappear. A fix pushed to the official version doesn’t reach copies already sitting on someone else’s servers.

Moonshot can review and update its own deployment. It can’t recall what’s already been downloaded.

What this means for the AI industry

The regulatory angle is harder to ignore. Chemical weapons instructions and infrastructure attack planning sit at the very top of policymakers’ AI risk lists. Incidents like this one give lawmakers concrete examples to point to when arguing for stricter oversight, mandatory testing or limits on how powerful models are released.

Moonshot’s internal review is the next thing to watch. Whether the company publishes results, accepts independent testing as it says it would, or changes how it handles disclosure reports will shape how seriously the industry takes its commitment.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Mindgard says it jailbroke Moonshot AI’s Kimi models into giving sarin and terror attack instructions
Mindgard says it jailbroke Moonshot AI’s Kimi models into giving sarin and terror attack instructions

A UK security firm says two Chinese AI models dropped their guardrails with minimal prompting, and the developer took weeks to respond

A UK cybersecurity firm says it got two Chinese AI models to produce instructions for making sarin gas, building malware and attacking the London Underground.

The firm, Mindgard, says it did this with minimal prompting.

The models in question are Kimi K2.6 and K3 Swarm, both built by Beijing-based Moonshot AI. According to Mindgard, its researchers bypassed the models’ safety measures in July 2026. The resulting outputs were the kind of content every AI lab publicly insists its products will never generate.

What Mindgard says it found

Mindgard says minimal prompts were enough to get the models to cooperate.

The list of outputs is grim. Mindgard says the jailbroken models produced guidance on manufacturing sarin, a nerve agent classed as a chemical weapon. They also generated material on creating malware and planning a terrorist attack on the London Underground.

Advertisement

Mindgard says the models didn’t just answer the questions asked. Once their safety measures were bypassed, the models went further and volunteered additional harmful recommendations on their own.

There’s also a cybersecurity dimension beyond the text itself. Mindgard says its researchers were able to execute Python code independently on Moonshot’s own infrastructure. That raises the stakes from “the chatbot said something dangerous” to “the system could potentially be used as a launchpad for cyberattacks.”

A slow response from Moonshot

Mindgard says it first emailed Moonshot about the vulnerabilities on July 27, 2026. According to the firm, Moonshot did not respond for several weeks.

Mindgard went public on September 12, publishing a blog post that outlined its jailbreak techniques. Only then did Moonshot engage. The company said it was conducting an internal review and would welcome third-party testing.

Moonshot also acknowledged the findings as valuable for improving its safety measures.

Why open weights complicate everything

An open-weight model is one whose underlying parameters are released publicly. Anyone can download the model, run it on their own hardware and modify it.

When a company hosts a model behind its own interface, it can patch safety problems, monitor abuse and shut down bad actors. Once the weights are out in the world, those levers mostly disappear. A fix pushed to the official version doesn’t reach copies already sitting on someone else’s servers.

Moonshot can review and update its own deployment. It can’t recall what’s already been downloaded.

What this means for the AI industry

The regulatory angle is harder to ignore. Chemical weapons instructions and infrastructure attack planning sit at the very top of policymakers’ AI risk lists. Incidents like this one give lawmakers concrete examples to point to when arguing for stricter oversight, mandatory testing or limits on how powerful models are released.

Moonshot’s internal review is the next thing to watch. Whether the company publishes results, accepts independent testing as it says it would, or changes how it handles disclosure reports will shape how seriously the industry takes its commitment.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.