
Microsoft's Decision-1 model enters the fast-growing AI decision model race
Quick Answer
Microsoft's Decision-1 model, based on Qwen3.5-9B, achieves 83.5% accuracy and 85 ms latency, outperforming competitors like Jev and H2O-Lightning-4B.
Quick Take
Available via Microsoft Foundry and OpenRouter, it costs $0.042 per million input tokens, with free output tokens.
Key Points
- Decision-1 is designed for fast, structured decision-making.
- It is 2.5 times faster than H2O-Lightning-4B.
- The model excels across 36 benchmarks with nearly 150,000 questions.
- Input tokens cost $0.042 per million; output tokens are free.
- Jev, the trendsetter in decision models, faces increased competition.
📖 Reader Mode
~1 min readMicrosoft now has its own decision model. Like the ones from OpenAI, Cloudflare, and the "original" from Jev, Decision-1 is built for fast, structured decisions. According to Microsoft, it handles classifications, evaluations, and routing decisions and has "the potential" to "guide and control agents through complex environments."
Decision-1 is based on Qwen3.5-9B. Microsoft says it's the most accurate model tested across 36 benchmarks covering nearly 150,000 questions. On speed, it's 2.5 times faster than the runner-up, H2O-Lightning-4B. Cloudflare's open-source Clef models, which are also based on Qwen, weren't included in the comparison.

Decision-1 is available through Microsoft Foundry and OpenRouter. Input tokens cost $0.042 per million, and output tokens are free. Jev, the startup that kicked off the decision model trend in mid-September, has likely had a rough few weeks. The idea caught on fast, but the tech was quickly adapted and surpassed using open small language models.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

