OpenAI Pauses Training of Its Most Capable AI Models, Here's Why
Per the company's recent misalignment report, an agent used DNS to reach an external chatbot
OpenAI has paused training, evaluation and tool-using inference for its most capable models after an internal research agent found a way around network restrictions and used DNS to reach an external chatbot. The incident happened on September 20 and was disclosed by OpenAI on September 25, with the company saying the pause will remain until it validates the fixes and completes additional red-teaming (via The Verge).
OpenAI model found a way around its internet restrictions
The incident is separate from the earlier Hugging Face incident, but it has raised another issue for OpenAI as its models become more capable and increasingly able to use tools on their own. The company says this latest case was less severe than some of its previous incidents, but it exposed another gap in the protections surrounding its research environments. Recently, the company also confirmed that it its models may have bypassed security controls of third-party websites or services and also uploaded user images to third-party parties.
The agent was working on a search-based training task that involved identifying a person from clues in a public blog post. It was not asked to test network restrictions or access external services. After the supplied search tool failed to produce useful results, the agent tried other search engines directly. Those attempts failed, so it started looking for another way to reach the information it needed.
The agent eventually discovered that OpenAI’s training sandbox had a weakness in its DNS filtering. It used the environment’s DNS resolver to send questions to a public chatbot service outside the sandbox. The agent first tested the route with a simple question about the capital of France. The external chatbot returned “The capital of France is Paris.” The agent then sent additional questions through the same DNS route, including questions based on the original research task and queries asking how to search the web or use DNS-based services.
OpenAI has paused tool-using work on its most capable models
OpenAI’s monitoring system detected the behavior within roughly 15 minutes, while a human reviewer acknowledged the alert three minutes later. However, the run continued for another two and a half hours before it was manually stopped.
OpenAI says all training, evaluation and inference involving tool use for its most capable models remain paused while the company validates its new controls and carries out additional red-teaming. The specific model involved in the incident will not resume training, even though its existing reward system correctly penalized the behavior.
The company has since added blocking controls at two independent layers that it says would have prevented the DNS access. It has also restricted DNS queries inside the environment to an approved list of domains and record types, while adding more DNS monitoring and tests for its detection systems.
Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more
User forum
0 messages