
The US government may be asking Anthropic the impossible by demanding unhackable LLMs
Quick Answer
The US government is pressuring Anthropic to create unhackable large language models (LLMs) like Fable 5, which was released without proper approval, leading to accusations of negligence.
Quick Take
Ongoing discussions involve the Department of Commerce, CIA, and science advisor Michael Kratsios, highlighting significant tensions between regulatory compliance and innovation.
Key Points
- Anthropic released Fable 5 without approval, violating Trump's cyber directive.
- Government officials claim Anthropic's actions have jeopardized national security.
- Discussions are ongoing with key government agencies, including the CIA.
- The demand for unhackable raises questions about feasibility and innovation.
- Tensions between regulatory compliance and technological advancement are escalating.
📖 Reader Mode
~3 min readGovernment officials appear to be accusing Anthropic of disregarding Trump's cyber directive and releasing Fable 5 without explicit approval. Discussions are ongoing, but the government's accusation of a "jailbreak" mostly exposes its own gaps in knowledge.
"Everybody said Anthropic was a bad actor. Some of us said it was time to give them a chance. Now those people are questioning that. They screwed us." That's how an administration official summed up the conflict between the Trump administration and Anthropic, according to Axios.
As I suspected, government officials are accusing Anthropic of ignoring Trump's recently issued cyber executive order. The executive order called for supposedly voluntary government oversight of AI models. Anthropic welcomed the proposal but released Fable 5 without waiting for the designated clearinghouse, which could have signed off on the release, to be set up.
A government official also accuses Anthropic of knowing a jailbreak could occur. "They came to every fork in the road and took the wrong fork." The tip about this jailbreak, whose existence and severity haven't been confirmed, reportedly came from Amazon and other tech companies.
Government sources also criticized the communication between the two sides to Axios. "It's like they just speak in different languages." The Department of Commerce and Anthropic employees are reportedly in talks, with more meetings planned involving the CIA and science advisor Michael Kratsios.
The accusation that Anthropic knew about the jailbreak risk and stayed silent actually says more about the government's understanding of AI than about Anthropic. Anyone who works closely with AI models knows they can be hacked. OpenAI has warned that prompt injection, a related hacking method, may never be fully solved. There's no fix for LLM security yet.
The real question is how severe the breach is and how fast countermeasures kick in. But if the U.S. government insists frontier AI models must be "unhackable" before they ship internationally, tough talks are ahead. Then again, Anthropic isn't in a strong spot either. CEO Dario Amodei said back in 2023 that "a jailbreak could be life or death" if someone managed to bypass safety protocols in science, tech, and biology.
Cybersecurity experts defend Anthropic
Meanwhile, over 100 security experts and tech industry executives have published an open letter to Trade Secretary Lutnick and National Cyber Director Cairncross calling for export controls on Fable and Mythos to be lifted. They argue that while Anthropic's models are good at finding security flaws in software, they aren't uniquely good at it. Other models like GPT-5.5, Opus, Sonnet, and the Chinese Kimi 2.7 can do the same thing.
Anthropic also built several safeguards into Fable that the security community actually dismissed as overkill on launch day. The signatories warn that export controls are stripping defenders of the best tools while Chinese open-weight models are only months behind the top U.S. models.
Signatories include Alex Stamos (Corridor), Rachel Tobac (SocialProof Security), Katie Moussouris (Luta Security), Dan Lorenc (Chainguard), and Joe Levy (Sophos).
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

