Safety#AI safety

Fired OpenAI safety researchers publish open letter

Three safety researchers fired by OpenAI published an open letter disputing the misconduct findings and urging the company to embed third-party safety auditors.

Three empty desks in a bright office, an open letter with a red wax seal on the nearest desk

On October 8, the three safety researchers dismissed by OpenAI — Jasmine Wang, Tomek Korbak and Mikita Balesni — published an open letter addressed to the company’s Safety and Security Committee, its Safety Advisory Group and its Mission Advisory Council. The three deny the allegations of mishandling sensitive information point by point, write that terminations “executed and communicated so abruptly” are chilling the open culture OpenAI has prized, and ask for two things: third-party safety auditors embedded in evaluation, and room for internal dissent. An OpenAI spokesperson reiterated to media that an investigation found a “pattern of misconduct.” When we covered the firings on October 1, none of the three had been named.

Key points

  • Three people, three specific scenes: Wang says the executive inbox she accessed was delegated to her for recruiting and she reported an accidental sensitive email within minutes; Korbak acknowledges talking to outside safety evaluators during the Hugging Face sandbox-breach investigation; Balesni says he removed sensitive details and checked his reporting line before sharing.
  • Collective action: this is the first joint public response by fired safety staff, published as a PDF and spread by Balesni on X.
  • Two OpenAI messages at once: an internal memo insists “we do not terminate employees for raising concerns,” while the spokesperson cites an investigation finding — with no specific policy named.

Background

We have followed this thread for three weeks. It starts with the rogue-agent breach of a Hugging Face sandbox in late September, then OpenAI confirming on October 1 that it cut ties with three safety researchers — at that point the reporting, from WSJ and the company, named no one and specified nothing. Two more pieces of context: on October 3 a safety employee resigned calling the culture “broken,” and earlier, on October 2, we covered the departure of safety lead Robinson. The letter stitches these scattered events into a single timeline.

The facts

What the letter and the company’s statements allow us to lay side by side:

PersonTheir accountThe company’s account
Jasmine WangThe executive inbox was delegated for recruiting; she reported the accidental open within minutes — “none of this was hidden”No individual response
Tomek KorbakSpoke with outside safety evaluators during the sandbox-breach investigation, believing it within normsNo individual response
Mikita BalesniRemoved sensitive details before sharing monitorability material and had support from his reporting line and board membersNo individual response

The company speaks in two registers. A spokesperson says the investigation found a “pattern of misconduct” beyond sharing with an outside evaluation group, while declining to name which policies were violated. An internal memo from a research leader praises the three and insists the terminations were unretaliatory — “we do not terminate employees for raising concerns.” The letter’s asks are structural: embed third-party safety auditors and cool down a termination style the authors describe as abrupt in both execution and communication. WSJ’s reading adds a fourth thread: the three also ask OpenAI to preserve visibility into AI reasoning — the subject of the leak they are accused of.

What others say

TechCrunch frames the letter as a direct rebuttal plus a warning about a chilling effect, noting OpenAI did not formally respond to the letter itself. CNN lays out the firing timeline alongside the October 3 resignation. Machine Heart’s Chinese roundup spread widest in community channels, headlining the leaked inner workings. Note the distribution path: Balesni skipped the press and posted the PDF on X, and within hours WSJ, CNN and TechCrunch all cited the same file. Every substantive accusation and denial currently rests on one side’s statement; nothing independently verifiable has been produced by either party.

Our take

The structural implication outweighs the he-said-she-said: a safety team’s effectiveness ultimately depends on the employer’s goodwill, and the power to fire sits with the employer. The third-party audit the three propose is an attempt to convert “independence of safety evaluation” from a cultural promise into an institutional arrangement — if OpenAI accepts, it becomes an industry template; if it declines, the next internal critic faces a starker choice. For other labs the lesson is just as direct: whistleblower channels and exit protection for safety staff are turning from PR topics into governance infrastructure. Honest uncertainty: this is a one-sided letter, and the company’s account is equally one-sided; what the “pattern of misconduct” concretely refers to remains undisclosed by both sides.