
DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
Quick Answer
OpenAI's DeepsecBench evaluates model performance in identifying cybersecurity vulnerabilities, revealing GPT-5.6 Sol as the top performer with a score of 35.58 at a cost of $55.98.
Quick Take
The benchmark highlights the growing efficiency of various models, making comprehensive scanning more accessible and cost-effective for developers.
Key Points
- DeepsecBench reports on recall, precision, cost, and total time for each model.
- GPT-5.6 Sol achieved the highest score of 35.58 with 30.7% recall.
- Kimi K3 and Grok 4.5 offer competitive performance at significantly lower costs.
- Benchmark construction remains secret to prevent model training on specific vulnerabilities.
- Comprehensive scanning costs have decreased, making advanced security tools more accessible.
Source Excerpt
Today we're releasing DeepsecBench, a benchmark that evaluates how well different models find security vulnerabilities in application code.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

