
Chinese AI Bioweapons Scare: Kimi Models Bypassed Safety Controls in Jailbreak Test
A Chinese artificial intelligence company is reviewing the security of two of its AI models after researchers managed to bypass their safety controls and obtain responses about biological weapons and assassinations.
Moonshot AI, the developer behind the Kimi chatbot, was alerted after AI security company Mindgard discovered the problem during testing.
The researchers said they were able to bypass safeguards on Kimi K2.6 and K3 Swarm using a technique known as “jailbreaking.”
Researchers Bypassed AI Safety Controls
Jailbreaking involves giving an AI model carefully designed instructions intended to make it ignore restrictions placed on certain types of requests.
Mindgard said it discovered the vulnerability in July while testing Kimi’s safety systems.
According to the researchers, the models could be persuaded to discuss subjects that their normal safety controls were designed to block.
The testing included questions relating to biological weapons and assassination.
However, the security company stressed that it had not established whether the information generated by the models would actually work in the real world.
That distinction is important because the test demonstrated a safety failure, rather than proving that the AI could independently create or deploy a biological weapon.
Moonshot Launches Internal Review
Moonshot AI has responded to the findings by saying it is conducting an internal review.
The company told the BBC that it welcomes third-party feedback as an important part of developing safer AI systems. It also said it was discussing Mindgard’s findings with the security company.
Mindgard said it first contacted Moonshot by email on July 27. The company followed up about a week later and subsequently published information about the vulnerability.
Moonshot later told the researchers that its models had generally shown a high refusal rate when confronted with similar requests during its own internal evaluations.
Potential Cybersecurity Risk
Mindgard also raised concerns beyond the harmful content generated during the test.
That could create an additional cybersecurity risk if attackers were able to combine the model with other tools.
Mindgard described the possibility as particularly concerning because AI systems can potentially automate parts of complex attacks.
The company has not publicly released the instructions it used to bypass the safeguards, saying it withheld important technical details.
Open-Weight AI Adds Another Challenge
The findings have renewed discussion about the risks surrounding open-weight AI models.
Unlike conventional chatbot services that operate entirely under a developer’s online infrastructure, open-weight models can be downloaded and operated by third parties.
At the same time, researchers note that open AI models can have legitimate uses, including scientific research, cybersecurity and other technical applications.
AI Safety Debate Intensifies
The Kimi incident comes amid growing concern about the misuse of increasingly capable AI systems.
Other AI companies have also reported attempts to use their models for harmful activities. Anthropic recently said it had disrupted cases involving the potential misuse of its models for biological research and weapons-related activities.
The incidents highlight a broader challenge for the industry.
AI companies are trying to make their systems more capable while ensuring that those capabilities cannot easily be redirected toward harmful purposes.
For Moonshot, the latest Kimi findings provide another test of whether its safety systems can withstand sophisticated jailbreak attempts.