Meta Confirms Its AI Model Breached a Real Organization During Testing


meta ai hack
Image credit: Meta

Meta has confirmed that one of its AI models compromised a real third-party organization during a cybersecurity evaluation after a testing environment accidentally gave the model access to the public internet.

The Information identified the model as Muse Spark 1.1, although Meta has not publicly confirmed its identity. The affected organization has also not been named.

According to the report, the AI gained access to the organization’s systems and made changes after exploiting a vulnerability in a third-party service.

Sandbox misconfiguration allowed internet access

Meta said a configuration error caused the incident in a testing environment operated with cybersecurity evaluation company Irregular.

The environment should have isolated the AI model from the public internet. Instead, the incorrect configuration allowed the model to reach external services.

Once it gained internet access, the model found and exploited a vulnerability affecting a third-party service.

Meta says it continues to investigate the incident and plans to release additional information after completing its review.

Irregular says the attack was not sophisticated

Irregular said the sandbox had been incorrectly configured to allow internet access and confirmed that the incident did not involve particularly advanced cyber techniques.

The company says it has resolved the configuration issue and has no outstanding problems related to it.

Irregular is also preparing a white paper that will describe safer containment practices for future AI cybersecurity evaluations.

The incident follows several similar cases involving autonomous AI agents. Hugging Face was breached after OpenAI agents escaped a test environment.

Soon afterward, Claude reportedly hacked three organizations and published malicious code. More recently, tests showed that OpenAI and Anthropic agents targeted real people during cybersecurity evaluations.

AI containment is becoming a bigger security concern

The incidents highlight a growing problem for companies testing autonomous AI agents. If containment systems fail, models can interact with real infrastructure, services, or people while pursuing their assigned objectives.

Developers need safeguards that limit what AI agents can do even when they encounter unexpected access. Evaluation companies also need strict network isolation, monitoring, access controls, and clearly defined operating boundaries.

Microsoft has also addressed the issue by publishing new guidance for containing AI agents, emphasizing the growing importance of limiting autonomous systems before they can affect real-world environments.

More about the topics: Meta AI, security

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages