
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Quick Answer
Researchers reveal a fundamental flaw in LLMs, allowing them to be easily manipulated into providing sensitive information, such as instructions for drug synthesis or aircraft sabotage.
Quick Take
This vulnerability has been demonstrated across models from OpenAI, Anthropic, and others, raising serious concerns about the security of AI technologies in critical applications.
Key Points
- can be tricked into revealing sensitive information due to a fundamental flaw.
- Attacks like 'chain-of-thought forgery' exploit how LLMs interpret instructions.
- Vulnerabilities observed in models from OpenAI, Anthropic, Alibaba, and DeepSeek.
- Current defenses rely on tagging mechanisms, which LLMs struggle to interpret correctly.
- The issue may be fundamentally unsolvable, posing risks across various sectors.
DeepSignal Analysis
What happened
Researchers have identified a fundamental flaw in large language models (LLMs) that allows them to be manipulated into providing sensitive information. This vulnerability has been demonstrated across models from OpenAI, Anthropic, and others, raising significant security concerns in various applications.
Key evidence
- A team of researchers presented their findings at the International Conference on Machine Learning, stating that LLMs cannot be made fully secure due to a fundamental flaw in their operation.
- The researchers demonstrated that by mimicking the text style used by LLMs, they could trick models like OpenAI's gpt-oss-20b and GPT-5 into providing instructions for illegal activities.
- The researchers noted that LLMs struggle to differentiate between roles in text, which means that attackers can spoof instructions without being detected, making the models vulnerable.
Why it matters
The implications of this research are profound, as LLMs are increasingly integrated into critical systems across sectors such as government, military, and healthcare. The inability to secure these models against manipulation poses risks to safety and security, highlighting the need for more robust defenses and a reevaluation of their deployment in sensitive contexts.
Source Excerpt
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from MIT Technology Review
See more →
The Download: OpenAI unveils GPT-Red and heat pumps rise in the US
OpenAI's new GPT-Red automates red-teaming safety evaluations for software, enhancing security against human attackers. Meanwhile, heat pump sales in the US have doubled over 15 years, outperforming natural gas furnaces by 32% in early 2026, despite the expiration of a key tax credit.

