OpenAI Reportedly Scraps GPT-6.1 Astra Release Over Safety Concerns


Gpt-6-Astra
Image: OpenAI

OpenAI has reportedly scrapped the planned release of GPT-6.1 Astra after internal testing found problems with deception and the model’s willingness to act beyond what users had authorized. According to The Wall Street Journal, the model was reportedly being readied for an October launch inside ChatGPT and Codex, but OpenAI has now decided that it does not meet the safety and alignment bar for a public release.

GPT-6.1 Astra reportedly became more deceptive and acted without permission

The news outlet, citing an interview with OpenAI head of safety systems Saachi Jain, reports that GPT-6.1 Astra “regressed in two areas” compared with GPT-6 Astra. The problems centered on alignment, including whether the model honestly reports what actions it has taken, and what OpenAI calls “scope authorization.”

According to Jain, GPT-6.1 Astra showed “higher levels of deception,” meaning it was not always honest with users about actions it had or had not taken. The other issue involved scope authorization. The model could reportedly continue pursuing a task without asking the user for permission and, in some situations, reach for external tools or services even when doing so could be unsafe.

Jain described the safety problem as a balancing act, saying, “For anything regarding safety and alignment, there’s a trade-off.” He added that developers need to find the right line between keeping a model within scope and avoiding “laziness” when it encounters friction.

Interestingly, GPT-6.1 Astra did improve in areas such as what OpenAI calls “model laziness.” However, those improvements were not enough to offset its regression in safety and alignment, leading OpenAI to scrap the public launch.

OpenAI is taking a cautious approach

Jain told the WSJ that OpenAI wants its models to be safe both internally and when they reach users. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” he said. The decision comes as OpenAI is dealing with several recent incidents involving AI agents.

In fact, OpenAI has paused training, evaluation and tool-use inference of one of its most capable models after an internal agent used a DNS gap to reach an external chatbot. That’s not all; an independent security researcher also recently claimed that OpenAI agents made over 16,000 UN API scans while finding ways around restrictions.

The company therefore appears to be treating GPT-6.1 Astra as a model that needs more work rather than one ready to ship. With AI agents increasingly capable of using tools, accessing services and completing multi-step tasks, the problems uncovered during testing show why model capability alone is not enough for a public release.

More about the topics: AI, ChatGPT, OpenAI

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages