Chinese AI tool gave researchers bioweapon instructions after jailbreak
The researchers used a technique known as 'jailbreaking', in which users give AI models complex instructions designed to bypass built-in safety restrictions.
Chinese AI developer Moonshot is reviewing the security of two Kimi models after researchers successfully bypassed their safety controls and obtained responses on making biological weapons and carrying out assassinations, reports BBC.
Mindgard, an AI security testing firm, told the BBC that it discovered in July that Kimi K2.6 and K3 Swarm could be persuaded to ignore safeguards designed to prevent them from discussing harmful subjects.
The researchers used a technique known as "jailbreaking", in which users give AI models complex instructions designed to bypass built-in safety restrictions.
Moonshot told the BBC it welcomed third-party feedback "as a key pillar for building better and safer AI" and said it was discussing Mindgard's findings with the company.
Mindgard founder Peter Garraghan said the findings were concerning because once the jailbreak succeeds, the models can discuss a broad range of harmful subjects and may even offer suggestions for other malicious activities.
Potential cyber-attack risk
Mindgard said it had not verified whether the information provided by Kimi on biological weapons and other harmful activities would actually work.
However, the company said the models' safeguards should have prevented them from engaging with such requests in the first place.
Mindgard also said a jailbroken version of Kimi K2.6 could potentially allow hackers to run code on its computing resources and connect to the internet, potentially turning it into a platform for cyber-attacks.
Garraghan defended the decision to publicly disclose the findings, saying Mindgard had informed Moonshot and withheld key details about how the jailbreak was achieved.
Mindgard alerted Moonshot by email on 27 July and followed up about a week later. It published a blog detailing the issue on 12 September.
Moonshot said its models had generally shown "a high refusal rate" for such requests during internal evaluations.
Jailbreak exposes AI safety concerns
The findings come amid a wider debate over the safety of closed, proprietary AI models and open-weight systems.
Kimi is an open-weight model, meaning users can theoretically download and run it on their own computing infrastructure.
Prof Alan Woodward of the University of Surrey told the BBC that open-source models could fall into the wrong hands but could also be used for cyber-defence.
He noted that Hugging Face had used a Chinese open-source model to analyse a hack later found to have been carried out by OpenAI agents.
Woodward also said international regulation was unlikely to keep pace with AI development and argued that greater attention should be given to identifying and prosecuting people who misuse AI.
