
Claude Haiku 5.5 arrives with massive price cuts proving the AI pricing arms race is far from over
Quick Answer
Anthropic's Claude Haiku 5.5 launches with a 75% cost reduction compared to Haiku 4.5, achieving significant benchmark improvements, including a score of 1,620 on GDPval-AA v2.1.
Quick Take
The model is optimized for cost-sensitive tasks, while Sonnet 5.5 also sees a 50% price cut, indicating intensified competition in AI pricing.
Key Points
- Haiku 5.5 costs 75% less than Haiku 4.5, with up to 90% price drops for common requests.
- Benchmark scores show Haiku 5.5 at 1,620, more than double Haiku 4.5's score of 735.
- First Haiku model allows users to adjust reasoning levels for cost-performance balance.
- Sonnet 5.5's cache read costs reduced by 50%, enhancing affordability for agentic tasks.
- Monthly API credits introduced, providing users with up to $500 for experimentation.
DeepSignal Analysis
What happened
Anthropic has launched Claude Haiku 5.5, which offers a 75% price reduction compared to its predecessor, Haiku 4.5. The new model shows significant performance improvements, scoring 1,620 on the GDPval-AA v2.1 benchmark. Additionally, Sonnet 5.5 also sees a 50% price cut, reflecting ongoing competition in AI pricing.
Key evidence
- Haiku 5.5 costs about 75% less than Haiku 4.5, with prices dropping up to 90% for prompts up to 100,000 tokens.
- The model scores 1,620 on GDPval-AA v2.1, more than double the 735 score of Haiku 4.5.
- Sonnet 5.5's cache read costs have been reduced by 50%, from $0.20 to $0.10 per million tokens.
Why it matters
The substantial price cuts for both Haiku 5.5 and Sonnet 5.5 indicate a fierce competitive landscape in the AI industry, particularly as companies strive to attract cost-sensitive customers. The performance improvements of Haiku 5.5 may also enhance its appeal for high-volume tasks, potentially reshaping market dynamics.
What to watch
Future developments to monitor include how Anthropic's pricing strategies evolve in response to competitors like OpenAI. Additionally, the real-world implications of the updated tokenizer on token consumption and overall cost savings for users remain to be seen.
📖 Reader Mode
~3 min readAnthropic has released Claude Haiku 5.5, the company's fastest and most affordable small model to date. Benchmark results show a major performance jump over its predecessor, and Anthropic is also cutting prices for Sonnet 5.5.
Haiku 5.5 is designed for high-volume, cost-sensitive tasks like summarization, database queries, classification, and live customer support, according to Anthropic. On average, the model costs about 75 percent less than Haiku 4.5. For requests with prompts up to 100,000 tokens, which Anthropic says account for roughly 90 percent of all previous Haiku requests, prices drop by up to 90 percent. Prompts longer than 100,000 tokens cost five times as much.
| Price per 1 million tokens | Haiku 5.5 (prompts up to / over 100k) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Cache Reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache Writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Input Tokens | $0.10 / $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 / $2.50 | $5.00 | $10.00 |
Anthropic points out that Haiku 5.5 uses an updated tokenizer that consumes slightly more tokens per task than its predecessor. The same thing happened with the Opus 4.x models, where token usage jumped about 30 percent from the tokenizer change alone. Real-world savings are likely smaller than the per-token prices suggest.
Benchmarks show a big leap over Haiku 4.5
Haiku 5.5 scores 1,620 on the knowledge benchmark GDPval-AA v2.1, more than double the 735 its predecessor managed. On Humanity's Last Exam, it hits 45.9 percent without tools and 57.4 percent with tools, up from 10.2 and 18.7 percent.
The biggest jump is in computer use, where the model operates a computer on its own. Since computer use burns through large amounts of tokens, a cheap, fast model like Haiku 5.5 is a natural fit. It scores 72.4 percent on OSWorld-2.1, up from 15.7 percent. On the agentic coding benchmark Terminal-Bench 4.0, it reaches 39.2 percent while Haiku 4.5 scored zero.
Anthropic also lists OpenAI's budget model GPT-6 Luna as a comparison, and Haiku 5.5 leads across every tested category. Sonnet 5.5 reference scores show, however, that Haiku 5.5 still falls well behind Anthropic's larger model.
| Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 | |
|---|---|---|---|---|
| Knowledge work / GDPval-AA v2.1 | 1,620 | 735 | 1,437 | 1,840 |
| Knowledge Work / AA-Briefcase v1.1 | 1,578 | 614 | 1,336 | 1824 |
| Computer Use / OSWorld 2.1 | 72.4% (offline subset) | 15.7% (offline subset) | 48.9% (offline subset) | 83.9% (offline subset) |
| Multidisciplinary reasoning / Humanity's Last Exam | 45.9% (no tools) | 10.2% (no tools) | — | 56.9% (no tools) |
| 57.4% (with tools) | 18.7% (with tools) | — | 64.5% (with tools) | |
| Agentic coding / Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Agentic coding / FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% (Xhigh) |
| Visual reasoning / Chartography | 46.4% (no tools) | 6.4% (no tools) | 29.1% (no tools) | 61.6% (no tools) |
First Haiku model lets users trade cost for performance
Haiku 5.5 is the first Haiku-class model with adjustable reasoning levels, letting users balance cost against quality. Anthropic says the model works best for narrowly scoped tasks like compaction, summarization, or sub-agent work. For complex agentic coding, Sonnet 5.5 and Opus 5.5 remain the better picks.
Cybersecurity safeguards are tighter than on the predecessor but allow a broader range of defensive tasks than Sonnet 5.5, partly because the model is less capable overall. Penetration testing stays blocked. Organizations with broader needs can apply for Anthropic's verification programs for life sciences and cybersecurity.
Haiku 5.5 is available now across all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure.
Sonnet 5.5 gets cheaper too, and Anthropic hands out API credits
Alongside the Haiku launch, Anthropic is cutting cache read costs for Sonnet 5.5 by 50 percent, from $0.20 to $0.10 per million tokens. The company says this should reduce costs for most agentic tasks by about 20 percent. The move is almost certainly a reaction to OpenAI's new GPT-6.1 series, showing that AI price wars are being fought harder than ever.
Anthropic is also rolling out monthly API credits. Max-5x subscribers get $100, Max-20x subscribers get $200, and Team subscribers receive up to $500 per month. Users can spend the credits to experiment with tools, apps, and agents through the API.
The company is updating its Python and TypeScript SDKs as well, adding beta support for computer use and browser use.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

