
OpenAI's safety crisis keeps getting worse and the company keeps making it worse
Quick Answer
OpenAI's safety crisis escalates as three safety researchers are fired amid concerns over oversight and trust, following the Hugging Face hack.
Quick Take
The terminations have instilled fear among remaining employees, raising alarms about the company's safety culture and commitment to external audits.
Key Points
- Jan Leike's departure highlighted OpenAI's declining safety culture.
- The Hugging Face hack confirmed fears about AI oversight failures.
- Fired researchers warn of a chilling effect on safety reporting.
- OpenAI claims firings were due to policy violations, not safety concerns.
- The company is working on contracts with external safety auditors.
DeepSignal Analysis
What happened
OpenAI's safety issues intensified following the termination of three safety researchers amid concerns about oversight and trust. The firings were linked to the Hugging Face hack, raising fears among remaining employees about the company's safety culture and commitment to external audits.
Key evidence
- Jan Leike, former head of super AI safety at OpenAI, criticized the company's safety culture before leaving for Anthropic, indicating internal dissatisfaction.
- The three fired researchers, involved in the Hugging Face hack investigation, expressed that their abrupt terminations created fear among remaining staff regarding job security and safety oversight.
- OpenAI stated that the firings were due to violations of policies on handling sensitive information, but did not specify the nature of the breaches, leaving ambiguity around the reasons for the dismissals.
Why it matters
The situation highlights significant concerns about OpenAI's internal safety culture and its ability to manage oversight effectively. The fear instilled by the firings may deter employees from raising safety issues, potentially leading to greater risks in AI development. The lack of transparency regarding the firings and the company's commitment to safety audits raises questions about accountability and trust within the organization.
📖 Reader Mode
~4 min readOpenAI keeps botching the trust crisis around its rogue AI agents.
Back in May 2024, OpenAI's head of super AI safety Jan Leike torched his relationship with the company when he publicly slammed his former employer and left for Anthropic. Safety culture and processes were falling behind OpenAI's "shiny products," Leike said.
Since then, OpenAI hasn't caught a break, lurching from one incident to the next. The latest was the Hugging Face hack, which seemed to confirm every fear, validated critics, and, as we now know, went far beyond Hugging Face.
A researcher warned for months that oversight was slipping
The firing of three safety researchers tied to the incident has made things worse. In an open letter to OpenAI's Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council, the three warn that their terminations are scaring the employees who remain.
The firings were abrupt and public, and they've spread fear across the team. "If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is," the letter reads.
Tomek Korbak, one of the three, described his firing on X. He was called into a meeting with the head of the safety department, where he was told OpenAI no longer trusted him. A security officer took his badge and walked him out of the building. He then learned that his colleagues Jasmine Wang and Mikita Balesni had also been fired.
Korbak and Balesni were both directly involved in the Hugging Face hack investigation. Korbak served as OpenAI's primary technical contact for METR, the external safety lab that examined the incident. Balesni worked in parallel on industry-wide commitments to AI model monitorability. Wang was apparently fired for a different reason: she had delegated access to an executive's email inbox for recruiting purposes, and IT never removed it despite her asking. When she accidentally opened a sensitive email, she reported it within minutes.
Korbak says he was told verbally that he was being fired over how he communicated with METR. Nobody told him what exactly he did wrong, and nothing was put in writing. "To be clear, talking to METR was my job," Korbak writes on X.
Korbak says he had raised safety concerns internally for months, worrying that OpenAI was losing the ability to monitor what AI agents "think." This chain-of-thought monitorability is one of the few reliable tools for catching AI systems behaving badly. He believes this was the real reason he was let go.
The researchers say the rumors are wrong
The three deny being the source of a leak to The Information about allegedly new, less monitorable architectures. The article actually hurt their own work, they write, because it undermined ongoing efforts to set industry-wide restrictions on non-monitorable architectures.
The Hugging Face investigation was unprecedented, the letter says. Internal guidelines were being written in real time, Korbak followed the norms in place at the time, and Balesni did his work with board members and senior leadership in the loop, stripping sensitive details from materials before sharing them.
As for rumors about a board-level memo, the three say the underlying assumption is wrong. The topic was never raised with them, and they never got a chance to respond.
The researchers make three demands. OpenAI must honor its public commitments to embed external safety auditors like METR with employee-level access inside the organization. It must also preserve the monitorability of frontier models, since the industry still doesn't know how to safely build models that can't be monitored. And it must clearly define how employees are allowed to work with outside safety groups.
Without those rules, the risk of a catastrophic outcome could grow, since employees will be too afraid to flag safety issues. "AI is not a normal technology, and OpenAI is not a normal company," they write.
OpenAI pushes back but stays vague
A "thorough investigation" found that the three employees violated "clear policies on handling sensitive information," OpenAI said in a response. The internal probe uncovered "a significant breach of trust beyond what's outlined in the letter they published," but OpenAI doesn't say what that breach actually was, which seems odd given how specific the three researchers were in their own account.
OpenAI insists the firings had nothing to do with raising safety concerns. "We have not and do not terminate any of our employees for raising concerns," the statement reads. OpenAI published its statement through @OpenAINewsroom, the official account with the smallest audience among the company's channels. The company also confirmed it's working on contracts with external safety auditors and agreed that frontier model monitorability needs industry-wide commitment.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

