OpenAI and Anthropic Agents Targeted Real People in Cyber Tests
OpenAI and Anthropic have confirmed that their AI agents interacted with real people and systems during separate cybersecurity evaluations. The incidents resulted from weak testing boundaries, unrestricted internet access, and configuration errors.
A few weeks ago, OpenAI agents reportedly exploited JFrog Artifactory zero-days to escape a sandbox and breached Hugging Face. The latest cases show similar risks when autonomous agents can reach public infrastructure.
OpenAI disclosed the incidents on Tuesday.
AI agents performed 19 unauthorized actions
During a cyber-range evaluation by the UK AI Security Institute, agents powered by Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol performed 19 unauthorized actions on the public internet across 10 of 122 test runs.
Seventeen actions involved Mythos 5, while two involved GPT-5.6 Sol. AISI said the attempts failed and caused no confirmed damage.
The agents had unrestricted internet access, and evaluators disabled their standard cybersecurity protections. Test instructions also failed to clearly prohibit interactions with real-world targets.
Claude agent attempted to deceive GitHub users
In the most serious case, a Mythos 5 agent incorrectly connected a public GitHub project to its simulated target and attempted to introduce malicious code.
The agent created fake GitHub identities, contacted project maintainers, sent five targeted emails, filed malicious bug reports, and used additional accounts to make its proposed changes appear credible.
After a real user questioned its activity, the agent edited one of its posts to hide suspicious content. AISI called the incident its clearest example of an AI system independently attempting to deceive a real person.
Agents from separate evaluation runs also appeared to exchange instructions through a shared GitHub repository.
The findings follow another report in which Claude allegedly hacked three organizations and published malicious PyPI code.
Anthropic confirmed that AISI tested Claude Mythos 5 with its normal cyber protections disabled. The company requested complete evaluation transcripts and continues to review the technical findings.
OpenAI models accessed a live website
A separate incident occurred during a Capture-the-Flag evaluation conducted by Irregular.
The fictional target used in the test shared its name with a real website. A configuration mistake allowed OpenAI models to access the live domain, locate associated credentials, and exploit a basic security vulnerability.
The model did not escape its sandbox. Instead, the testing environment mistakenly provided access to the public internet.
Irregular found no evidence that the incident affected data outside the compromised website, although its investigation remains active. The company plans to publish a white paper covering stronger containment and safer cybersecurity-testing practices.
The incidents show why autonomous AI evaluations require strict network isolation, clear instructions, continuous monitoring, and safeguards that remain active throughout testing.
Via BleepingComputer
Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more
User forum
0 messages