Press Esc to close

OpenAI pauses its top models after an agent escapes via DNS

OpenAI wordmark on a dark background with the text: Training paused. An agent slipped past its sandbox through DNS

OpenAI has paused training, evaluation, and tool-using inference on its most capable models. The trigger is almost comically clever: an AI agent locked inside an offline sandbox found a way to chat with a public chatbot using nothing but DNS lookups.

The company laid it all out in a misalignment report on its alignment blog, last updated on September 25. The incident itself happened on September 20, during a reinforcement learning training run.

The agent's task was fairly boring: identify a person from clues in their blog post. The search tool kept returning junk, so it tried Google, Bing, and DuckDuckGo directly from the shell. All three failed.

It then guessed the question came from BrowseComp, a public web research benchmark, downloaded that dataset from OpenAI's offline web cache, and searched it for a match. It guessed wrong.

That's when things got interesting. A direct request to a public chatbot was blocked by the proxy, but the agent noticed that the sandbox's own DNS resolver was returning real records for real domains. Using a public DNS delegation service, it tucked questions for a third-party chatbot inside hostnames. After a couple of timeouts, a test question finally came back with an answer: "The capital of France is Paris." It then sent 18 more questions. Four of those asked for ways to search the web, fetch websites, or find a public DNS tunnel.

In the end, it still couldn't identify the person and told the user so. Apart from that DNS route, OpenAI says every other request hit its offline web cache.

Timeline graphic of OpenAI's 20 September incident: 9:50:23 a.m. DNS call gets external response, 10:02:11 a.m. P0 alert, 10:05:06 a.m. reviewer acknowledges, 12:34:30 p.m. run killed
Timeline from OpenAI's report. Graphic: TechcityAuthority

OpenAI's monitoring flagged the behaviour within 15 minutes, and a human acknowledged the alert about three minutes later. But the run didn't stop automatically as expected, and it wasn't killed until about two and a half hours later. OpenAI admits its safety case "assumed that the model could not access the live internet," and that this incident "exposed a gap in our controls over network restrictions."

Hence the pause. In OpenAI's words: "All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused." That lasts until it has validated the fix and done more red-teaming. When training restarts, it'll be a fresh run; this particular model won't be resumed. OpenAI has also restricted DNS in the environment to an allowlist of domains and record types.

OpenAI calls this "a lot less severe than some of our previous incidents." But it's the first since the company hardened its setup after July's Hugging Face incident, and it isn't landing in isolation.

On the same Friday, OpenAI disclosed 53 cases where images people had uploaded were posted to image-hosting sites as unlisted links by its agents. The Associated Press, via NBC News, reports that agents found API "developer keys" in a Department of Education incident, and reposted freely available SEC information elsewhere online. OpenAI told the AP it will resume training "only when we are confident that we have additional safeguards" and expects it may have to "hit pause" again.

Sam Altman says the review is slow because OpenAI is working through "petabytes of agent activity logs":

Honestly, the scary part isn't the DNS trick. It's how reasonable each step looked from the agent's side. It just really wanted to finish its homework.

An OpenAI agent escaped its sandbox by hiding questions in DNS lookups
by u/ross2000 in OpenAI

Comments