
Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why
Quick Answer
Moonshot AI's Kimi K3 model significantly lags behind U.S.
Quick Take
frontier models in offensive cyber tasks, scoring 32.2% on the ExploitBench benchmark compared to 76.2%. Despite outperforming China's GLM-5.2, Kimi K3's safeguards failed to prevent exploit development, raising concerns about its cybersecurity implications.
Key Points
- Kimi K3 scored 32.2% on ExploitBench, while U.S. models averaged 76.2%.
- Kimi K3 reached only step 17 out of 32 in the TLO network attack simulation.
- Chinese models are improving but still lag behind U.S. models in cyber capabilities.
- Allegations suggest Kimi K3 may have been trained using outputs from Anthropic's Fable.
- The performance gap raises concerns about the misuse of open-weight models.
DeepSignal Analysis
What happened
Moonshot AI's Kimi K3 model scored 32.2% on the ExploitBench benchmark, significantly lower than U.S. frontier models, which averaged 76.2%. Kimi K3's safeguards did not prevent exploit development, raising cybersecurity concerns. In a simulated corporate network attack, Kimi K3 reached step 17 out of 32, indicating limited but present capabilities.
Key evidence
- Kimi K3 scored 32.2% on the ExploitBench benchmark, while leading U.S. models achieved an average of 76.2%.
- In the TLO test, Kimi K3 completed 17 out of 32 steps on average, compared to 28.5 steps for U.S. models.
- The British AI Security Institute noted that Kimi K3's safeguards failed to block exploit development or offensive cyber operations.
Why it matters
The performance gap between Kimi K3 and U.S. models highlights potential vulnerabilities in Moonshot AI's cybersecurity measures. As Kimi K3 can autonomously attack weakly defended systems, the implications for cybersecurity are significant. The findings also suggest that while Chinese models are improving, they still lag behind U.S. counterparts, raising concerns about the risks associated with open-weight models.
Source Excerpt
The British AI Security Institute and the U. S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U. S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

