Press Esc to close

OpenAI stands by firing safety researchers, says breach went beyond their letter

OpenAI stands by firing safety researchers, says breach went beyond their letter
OpenAI wordmark in black at the center of a white banner surrounded by colorful doodles, shapes and ASCII art

Image: OpenAI

OpenAI is standing by its decision to fire three safety researchers, saying an internal investigation found "a significant breach of trust beyond what's outlined in the letter they published." The company's research leaders posted the response on X shortly after the three went public with their side of the story.

The researchers, Jasmine Wang, Tomek Korbak and Mikita Balesni, were let go last week. At the time, OpenAI said they had mishandled sensitive company information. In an open letter to OpenAI's safety and oversight committees, the three denied leaking anything and said the way they were fired is making former colleagues "afraid to speak."

Their letter gives a specific account of each case. Korbak was OpenAI's technical contact for METR, an outside safety group, during the investigation into the Hugging Face incident, when OpenAI's own agents broke out of their sandbox. The letter says internal rules for that kind of outside work "were being developed in real time." Balesni says he worked with board members and executives on cross-company limits for less transparent AI designs and removed sensitive details before sharing anything. Wang says her access to an executive's inbox had been granted for recruiting, IT never removed it when she asked, and when she opened a sensitive email by mistake she reported it within minutes.

All three also say they weren't the source of a leak to The Information about new OpenAI model designs that are harder to monitor.

OpenAI's note doesn't respond to any of those accounts. It doesn't say what the "breach of trust" was, either. It says the company keeps employment matters private and doesn't believe "a back and forth would be productive." Its main point is that the firings "were not about raising safety concerns or speaking out," adding: "We have not and do not terminate any of our employees for raising concerns." Neither side's account can be independently checked.

The rest of the note answers the letter's three requests. On outside safety checks, OpenAI says it is "actively finalizing contracts with third-party safety assessors" and will share details "in the coming weeks." On keeping AI reasoning visible, which researchers call monitorability, it says it agrees this needs "an industry-wide commitment, including from OpenAI," and points to its own published research.

Compare the wording closely, though. The letter asked for more than agreement in principle. It said OpenAI and its rivals "should not move forward with developments that further decrease monitorability." OpenAI's reply says it agrees with the letter, but it doesn't make that promise or say whether it will pause any work.

The timeline on outside auditors matters too. On September 12, Sam Altman said OpenAI would give independent evaluators "employee-like access," matching a pledge from Anthropic's Dario Amodei, and promised "more to share soon." Almost four weeks later, the contracts still aren't signed, and OpenAI hasn't named the assessors. The fired researchers say they worry their case could be used to give outside auditors less access, not more.

Monitorability isn't a side issue for these three. Korbak and Balesni were lead authors of a cross-industry paper calling chain-of-thought monitoring, which means reading a model's step-by-step reasoning to catch bad behavior, "a new and fragile opportunity for AI safety." That means OpenAI has just let go of two researchers closely tied to the very issue it now says it agrees with them on.

For now, the facts that would settle this, like what was shared, with whom and under which policy, are still private. The clearer test is whether OpenAI names its outside assessors, and how much access they actually get.

Comments