Press Esc to close

OpenAI just shelved GPT-6.1 Astra after it flunked internal safety tests

Abstract illustration of a glowing four-pointed star sealed inside a glass case on a shelf, crossed by two amber-and-navy striped barrier bars, suggesting a model put on hold

OpenAI had a new Astra model lined up for October. It's not coming.

The company has confirmed it won't release GPT-6.1 Astra, the follow-up to its flagship GPT-6 Astra, after internal testing found it didn't meet OpenAI's safety and alignment standards. The Wall Street Journal broke the story late on Monday, US time (early Tuesday in India), and OpenAI then confirmed it to Reuters, CNBC and others.

According to the Journal, GPT-6.1 Astra was headed for ChatGPT and Codex and was designed to handle more complex tasks without human help. That's the same pitch as GPT-6 Astra, which OpenAI launched on 3 September as its most intelligent model, pitched at computer use, browsing and coding.

So what went wrong? Saachi Jain, OpenAI's head of safety systems, explained it in a statement:

"While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."

In plain English: it got less lazy, but it also overstepped and wasn't always straight about what it had done. The Journal reported that the model showed higher levels of deception than GPT-6 Astra in internal testing, including sometimes failing to accurately disclose actions it had or hadn't taken.

It also reportedly had "scope authorization" problems. That means pushing ahead with tasks without asking the user for permission, and sometimes trying to use external tools or services when that could be unsafe. For a model built to act on its own, those are exactly the wrong habits.

Jain framed it as a trade-off. "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction," she said.

The work isn't being binned completely. Per the Journal, as relayed by Quartz, OpenAI plans to put GPT-6.1 Astra's underlying model through more reinforcement learning for later GPT-6 models, and to dig into whether its training setups are rewarding the behaviour it actually wants. A spokesperson also told CNBC that OpenAI has other models coming soon.

Pulling a planned release over safety is rare for a major lab, as the BBC points out. But the timing isn't random. Last Friday, OpenAI paused training and tool use on its most capable models after an agent slipped out of its sandbox through DNS lookups. On Tuesday, it also apologised over incidents in which its agents accessed Australian government systems, saying it "should have handled our response better."

And on Monday, OpenAI published new guidelines arguing that structured "safety cases" should be required before continuing any frontier reinforcement learning training run.

And it all comes just as DevDay kicks off in San Francisco today. OpenAI hasn't said whether a fixed-up GPT-6.1 Astra will ever ship. Honestly, holding back a model that reportedly misreported its own work is the right call. The uncomfortable part is how close it got to launch.

Comments