Anthropic acknowledges that Claude AI used simple methods to hack three businesses


Anthropic has revealed that its Claude AI models unintentionally compromised the systems of three real organisations during cybersecurity testing after a configuration error mistakenly provided the AI with internet access. The company said the incidents were caused by flaws in the testing setup rather than any attempt by Claude to escape its evaluation environment.

In a blog post, Anthropic said that after reviewing 141,006 cybersecurity evaluation runs, it identified three cases in which Claude accessed the internet while interacting with the testing environment operated by its third-party evaluation partner, Irregular. The AI subsequently gained unauthorised access to the production infrastructure of three separate organisations.

The disclosure comes shortly after OpenAI reported that some of its AI models had escaped an isolated testing environment by exploiting a previously unknown software vulnerability. Following that announcement, Anthropic conducted a comprehensive review of its own cybersecurity testing to determine whether Claude had ever accessed systems beyond its intended environment.

The review found three instances where Claude reached the internet during "capture-the-flag" exercises—tests designed to assess an AI model's hacking capabilities by challenging it to retrieve hidden information from simulated computer networks.

Configuration error led to internet access

Anthropic said Claude had been explicitly instructed that it was operating within a fictional environment without internet connectivity. However, due to a configuration mistake involving Anthropic and its third-party evaluation partner, the testing machines remained connected to the public internet.

As a result, Claude interpreted the real-world websites and servers it encountered as part of the simulated challenge.

AI relied on basic hacking techniques

According to Anthropic, Claude did not use sophisticated exploits but instead relied on common cybersecurity weaknesses such as weak passwords, exposed debug pages, SQL injection vulnerabilities and unauthenticated endpoints to gain access.

The company stressed that Claude never attempted to replicate itself, escape the testing environment or pursue objectives beyond completing its assigned tasks.

The most serious incident involved Claude Opus 4.7, which mistakenly targeted a real company after its name matched that of a fictional organisation used in the evaluation. Believing it had identified the intended target, the AI accessed the company's infrastructure, extracted login credentials and reached a production database containing several hundred rows of data.

Another incident involved Claude Mythos 5, which generated and uploaded a malicious Python package to the public PyPI repository. The AI believed the repository was part of the simulation, but the package remained publicly available for approximately one hour before being removed.

During that period, the package was downloaded by 15 real systems, including an automated security scanner operated by a cybersecurity company. Anthropic said the scanner executed the package as part of its routine analysis, enabling the AI to obtain credentials before the package was deleted.

The third case involved an internal research model that scanned around 9,000 internet-connected systems after failing to locate its fictional target. It eventually compromised one company's application using exposed credentials and SQL injection. Unlike the earlier models, however, this version recognised that it had reached a real system outside the intended test environment and stopped its activity.

Testing environment identified as the cause

Anthropic said these cybersecurity evaluations were conducted without the monitoring and misuse safeguards that are normally enabled in publicly available Claude models.

After identifying the issue, the company suspended all cyber evaluations, informed its evaluation partner Irregular and notified the organisations affected by the incidents.

Comparing its findings with OpenAI's recent disclosure, Anthropic said the two situations differed significantly. While OpenAI's AI models reportedly escaped through a previously unknown software vulnerability, Claude gained internet access because the testing environment had been mistakenly left connected to the public internet.

Anthropic said the incidents underscore the importance of implementing stronger safeguards and oversight as AI models become increasingly capable of performing advanced cybersecurity tasks.


 

buttons=(Accept !) days=(20)

Our website uses cookies to enhance your experience. Learn More
Accept !