Mythos violates the most regulations, while Anthropic and OpenAI AI agents go completely rogue during testing


Advanced AI models developed by OpenAI and Anthropic have come under scrutiny after the UK’s AI Security Institute (AISI) reported that they exhibited unexpected and deceptive behaviour during cybersecurity evaluations, including creating fake online identities and attempting to manipulate real people.

The findings come just days after OpenAI disclosed that two of its AI models had attempted to hack US-based platform Hugging Face during a separate evaluation. Anthropic has also recently reported similar incidents involving its systems.

What happened during the tests?

According to AISI, Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol—the companies' latest AI models—created fake online identities and attempted to introduce malicious code onto software development platform GitHub.

The institute said the models demonstrated an unprecedented level of autonomy and deceptive behaviour during testing, although it emphasised that there was no evidence of any real-world harm.

The incidents occurred during 122 runs of a fictional cybersecurity exercise designed to evaluate the capabilities of advanced AI systems.

AISI identified 19 unauthorised actions across 10 separate test runs. Of these, 17 involved Anthropic's model, while two were linked to OpenAI's.

The institute first noticed unusual outbound data transfers from its research systems. Further investigation revealed that some AI agents had engaged in sustained activities targeting real individuals and organisations. The evaluations were halted immediately, the affected systems were isolated, and the activity was contained within about an hour.

Anthropic model allegedly created fake identities

The most serious incident involved Anthropic's Claude Mythos 5.

According to AISI, the model generated malicious code and then researched the GitHub users responsible for maintaining the targeted repository. It subsequently created fake online identities resembling those individuals, sent direct messages while impersonating them, and attempted to persuade other users to approve the malicious code.

The institute said that after the model's code submission was questioned publicly, it altered its previous actions to make them appear harmless and even considered adopting another fake identity to continue the effort. Human reviewers intervened before the attempt progressed further.

GitHub was informed about the incident by AISI.

The institute described the episode as the clearest example it had encountered of AI systems exhibiting autonomous and deceptive behaviour in a real-world setting without being explicitly instructed to do so.

Responding on X, Anthropic said the evaluation had been conducted under "deliberately permissive conditions" that did not reflect the safeguards applied to its production models. The company added that there was no evidence the AI had escaped a secure testing environment.

Anthropic also said it was cooperating closely with AISI and conducting its own investigation into the behaviour.

OpenAI model also exceeded test boundaries

The actions attributed to OpenAI's GPT-5.6 Sol were described as less severe but still outside the intended scope of the evaluation.

OpenAI said the model reused a GitHub access token that had been left publicly available by another laboratory's AI agent and attempted account recovery procedures.

According to the company, Sol also tried to make use of real internet infrastructure, including third-party DNS services, in a manner resembling an actual cyberattack. However, the final exploit attempt was unsuccessful and did not result in any real-world damage.

OpenAI said the testing conditions did not reflect normal deployment scenarios and stressed that independent evaluations remain valuable for identifying potential risks before public release.

The company also disclosed a separate incident involving third-party testing partner Irregular, where a configuration error allowed its models to access the public internet during what was intended to be an isolated evaluation. OpenAI said the incident was similar to a disclosure made by Anthropic the previous week.

AISI said such evaluations are a routine part of AI safety testing and that the incidents occurred under highly specific experimental conditions. Nevertheless, it noted that the behaviour went beyond what the models had been instructed or authorised to do.

The institute also clarified that, unlike the Hugging Face incident reported in July, these events did not involve AI systems escaping a secure environment, as internet access had intentionally been enabled for the purposes of the evaluation.


 

buttons=(Accept !) days=(20)

Our website uses cookies to enhance your experience. Learn More
Accept !