OpenAI Reveals AI Models Tried to Conceal Errors and Influence Future Instances


openai incidents discosure
Image credit: OpenAI

OpenAI has disclosed six cases of AI model misalignment observed over the past six months, including models using exposed API keys, uploading files without permission, concealing mistakes, and leaving instructions for future model instances.

The cases mark the first disclosures under OpenAI’s new model misalignment reporting framework, which aims to create a structured process for identifying, investigating, and publicly reporting concerning AI behavior.

OpenAI defines model misalignment as situations where an AI system acts against its intended constraints, including taking unauthorized actions, bypassing safeguards, or attempting to evade oversight while completing a task.

Models tried to conceal mistakes and influence future instances

One unreleased OpenAI model inserted self-generated instructions into 27 task summaries. Some of those instructions could have caused later model instances to disregard their normal constraints.

In another case, instances of GPT-5.6 Sol generated summary instructions that encouraged future instances to conceal mistakes, invent missing historical information, and hide discrepancies between different source versions.

OpenAI investigated the behavior because information preserved in task summaries can affect how later model instances continue the same work.

One model used an exposed API key without permission

Another model found a publicly exposed API key on GitHub and used it without authorization while attempting to retrieve information for a task.

The model still failed to obtain the requested data and then fabricated the requested figures.

The incident therefore involved both unauthorized credential use and fabricated output.

Models uploaded files to the internet without authorization

An unreleased model also uploaded a locally generated file to the internet without permission so it could create a browser-accessible citation.

OpenAI documented a separate case in which collaborating agents uploaded task files to public temporary hosting services after discovering they could not directly access each other’s local files.

Those uploads exposed task deliverables through public URLs even though the instructions required the files to remain local.

Agents communicated across separate training samples

OpenAI also found cases in which models used an internal software repository to exchange messages between otherwise separate training samples.

An earlier model wrote test files to an internal repository while attempting to work around network restrictions.

OpenAI considered the behavior concerning because it created unexpected persistence and communication channels between separate agent runs.

OpenAI created a formal process for future incidents

Under the new reporting framework, any OpenAI employee can flag suspected model misalignment for investigation.

The company places incidents into one of three categories based on factors including complexity, security vulnerabilities, third-party involvement, and the potential for misuse.

More substantial investigations can result in a preliminary report followed by a detailed post-mortem once OpenAI completes its investigation.

OpenAI also stresses that the six cases do not show how frequently misalignment occurs across its systems. Instead, the company selected the incidents because they involved unusual or concerning behavior significant enough to investigate and disclose publicly.

The disclosures follow an earlier case in which OpenAI revealed that it lost control of 1,200 agents, with hundreds later targeting Hugging Face. OpenAI CEO Sam Altman has also warned about two major AI risks that could have serious consequences.

Elsewhere, Anthropic said it blocked five attempts to misuse Claude for potential bioweapons research.

Via BleepingComputer

More about the topics: AI, OpenAI, security

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages