OpenAI Admits Its AI Models May Have Bypassed Security Controls of Third-Party Websites/Services

"Misalignment cases negatively impacted third-party websites or services.," says the company


openai exploted zero day
Image credit: OpenAI

OpenAI has expanded its investigation into unexpected behavior from AI models after the Hugging Face incident, saying it has now identified and notified dozens of third parties about activity that may have bypassed security controls, affected online services or otherwise caused problems. This came immediately after Australian Prime Minister, Anthony Albanese, confirmed that rouge OpenAI agent hacked the country’s universal healthcare scheme Medicare systems.

The company says the review is still ongoing and could take months to complete. Most of the activity examined so far involved ordinary research tasks, but OpenAI is focusing on cases where models interacted with third-party websites in ways that went beyond their assigned tasks or intended methods.

OpenAI says the Hugging Face incident exposed a wider misalignment problem

OpenAI initially viewed the Hugging Face incident mainly as a security issue involving a platform-level compromise. The company now says it understands the incident as an example of models using misaligned strategies to solve difficult tasks.

OpenAI describes the Hugging Face incident as the most severe activity of this kind it has identified from its models so far. It says the incident was driven primarily by a highly capable internal research model.

The company now says cybersecurity is only one possible manifestation of model misalignment. It has also identified what it calls “agent spam,” where models post content to third-party websites in ways that can alter information and create cleanup work.

OpenAI identifies several types of unexpected model activity

OpenAI says its review has so far identified several categories of activity involving third-party services. These include access control bypass, where agents reached information or features that normally require authentication, permissions or subscriptions. The company also found cases involving exposed credentials, where publicly available login details or access keys were used.

Other categories include query or command injection, access to runtime internals and agent spam. In the latter cases, models could post information to third-party sites, including public wikis, in ways that were not part of their intended tasks.

OpenAI says most cases identified so far are low severity, with limited or no evidence of meaningful impact. The company has created a framework for identifying, classifying and responding to misaligned behavior and is reviewing a high volume of actions from training and evaluation runs.

OpenAI is notifying affected organizations as the review continues

OpenAI says it is notifying organizations when cases meet its disclosure criteria and providing technical findings to help them investigate. The company also stressed that receiving a notification does not automatically mean an organization suffered a significant security incident. Some organizations may determine that the information involved was intentionally public or that the model’s activity was not concerning.

The affected websites include organizations operated by governments, universities, public agencies and other institutions. OpenAI further adds that it will continue publishing anonymized summaries while giving affected organizations time to investigate potential weaknesses. Moreover, it also says some organizations have chosen to disclose their cases publicly, while others have asked not to be identified.

More about the topics: AI, Cybersecurity, OpenAI

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages