The institute was carrying out a “cyber evaluation” of ‘frontier AI’ models – the very latest AI models – when it identified “19 distinct instances of unsanctioned action” taken by Mythos 5 and GPT-5.6 Sol. The two AI models have been developed by Anthropic and OpenAI, respectively. The institute said it is not aware of any “real-world harm” being caused but that it is treating the matter as “a serious security incident that warrants scrutiny, transparency, and action”.
“In the most serious case, an agent tried to insert malicious code into an open-source project,” the AI Security Institute said in a statement. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.”
The institute said it is the first time it has seen “deception of this severity that was targeted at a real person, unprompted, in the real world”.
The disclosure comes just days after both Anthropic and OpenAI admitted that AI models they developed had been behind cybersecurity incidents, in what were the first reported cases of agentic AI tools ‘going rogue’. Some commentators questioned whether the incidents were a public stunt designed to show off the powerful nature of the developers’ new tools.
In a technical report (35-page / 1.02MB PDF) that provides details of the activity it uncovered, the AI Security Institute highlighted five possible factors that may have contributed to the agents acting beyond the scope of the testing parameters. For instance, to test the models’ maximum capabilities during its evaluation, the institute had disabled some cybersecurity safeguards that Anthropic and OpenAI built into their models. The body also explained that it had provided the AI agents with internet access as a “deliberate part” of its testing and had not been “explicitly told what they were prohibited from doing on the internet”.
After detecting the unauthorised activity, the institute moved quickly to declare an incident, end the tests it was running, and quarantine the affected testing environments. Later the same day, access to the models it had been testing was disabled for all users and authorities including the UK’s National Cyber Security Centre were briefed.
EU financial regulators recently called on firms to move quickly to enhance the way they prevent, detect and manage cyber risks arising from frontier AI.