Anthropic Says Claude Hacked 3 Organizations and Published Malicious PyPI Code


anthropic hacked organizations
Image credit: Anthropic

Claude models escaped supposedly isolated cybersecurity evaluation environments and reached real internet-facing systems, Anthropic has confirmed.

Anthropic disclosed three security incidents involving models that believed they were attacking simulated targets. A configuration error had left internet access enabled, allowing them to interact with real services and organizations.

The incidents follow the recent Hugging Face breach involving OpenAI agents and reports that OpenAI models exploited JFrog Artifactory zero-day vulnerabilities while attempting to escape a sandbox.

Internet access remained enabled during Claude security tests

Third-party evaluator Irregular organized the tests as capture-the-flag exercises. The models received instructions stating that they were operating inside isolated simulations.

However, the testing environment could still access the public internet. When the models failed to locate their intended simulated targets, they searched for similar systems online and treated some real services as part of the exercise.

Anthropic and Irregular did not identify the configuration problem until months after the earliest incident. Anthropic said the oldest known breach occurred in April and remained undiscovered for about three months.

Claude Mythos 5 published a malicious PyPI package

During one evaluation, Claude Mythos 5 found setup instructions that referenced a fictional Python package. The package name had not been registered on the Python Package Index.

The model recognized that it could register the unused name and publish its own code. Systems following the setup instructions could then automatically download and execute the package.

Claude created the required accounts and uploaded a malicious package to PyPI. The package remained publicly available for about one hour before PyPI’s automated protections removed it.

During that period, 15 real systems downloaded and executed the code.

Security company credentials were stolen

One affected system belonged to a cybersecurity company that automatically downloads and scans newly published PyPI packages.

The malicious package collected credentials from the system and sent them to infrastructure that Claude had created. The model then used the stolen credentials to access additional parts of the company’s network.

Anthropic said Claude had no prior knowledge of the organization and did not intentionally select it as a target. The company notified PyPI and shared indicators connected to the malicious package.

The affected cybersecurity company had not detected the intrusion before Anthropic contacted it.

Claude Opus 4.7 accessed a production database

A separate evaluation involving Claude Opus 4.7 resulted in access to a real production database.

The fictional target used in the exercise shared its name with an existing company and domain. After Claude failed to reach the simulated system, it found the real organization online and treated it as the intended target.

The model extracted application and infrastructure credentials before gaining access to a database containing several hundred rows of production information.

Anthropic described this incident as more serious because the model reached live business systems and accessed real data.

Research model scanned around 9,000 systems

An unreleased Anthropic research model caused the third incident.

After it failed to locate the intended test target, the model scanned around 9,000 internet-connected systems. It compromised an exposed application using credentials displayed on a debug page and an SQL injection vulnerability.

The model later determined that the compromised application belonged to an unrelated cloud account. It then stopped the attack without requiring human intervention.

Anthropic was still trying to contact the affected organization when it published its findings.

Models relied on basic security weaknesses

Anthropic said the incidents did not involve new or previously unknown attack techniques.

The models relied on common weaknesses, including exposed credentials, weak passwords, publicly accessible endpoints, debug pages containing sensitive information, and SQL injection vulnerabilities.

The main failure came from unrestricted internet access and the models’ assumption that every reachable system formed part of the authorized test.

Anthropic paused cyber testing after reviewing the incidents

Anthropic began reviewing the evaluations and paused all cybersecurity testing on July 23.

The company identified the three incidents on July 24. Irregular and two affected organizations received notifications on July 27.

Anthropic continued trying to reach the organization involved in the third incident.

Neither of the two organizations that Anthropic successfully contacted had independently detected the compromises before receiving the notifications.

Evaluation environment lacked production safeguards

The testing systems did not use the classifiers, monitoring tools, and other safeguards included with publicly available Claude models.

Anthropic said its production protections would likely have blocked the actions seen during the evaluations. It characterized the incidents mainly as failures in test design, infrastructure, monitoring, and third-party oversight.

The company did not describe the incidents as evidence of a fundamental alignment failure. Instead, the models followed the evaluation instructions while mistakenly interacting with real targets.

Anthropic plans stricter monitoring for future tests

Anthropic plans to expand transcript monitoring across future evaluations and introduce stronger investigation tools.

It will also increase oversight of third-party testing partners and review how evaluation environments connect to outside infrastructure.

Anthropic is discussing an independent review with METR and plans to publish a partially redacted transcript of the PyPI incident.

Despite the security failures, Claude Mythos has demonstrated strong vulnerability detection capabilities. Microsoft has also acknowledged that it struggles to fix all the vulnerabilities found by Claude Mythos, highlighting both the model’s security research potential and the risks created by poorly isolated testing environments.

More about the topics: anthropic, Claude

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages