
Claude Haiku 5.5 降本背后:操作能力暴涨,复杂编程为何仍差一截?
Quick Answer
Anthropic's Claude Haiku 5.5 significantly enhances operational capabilities with a 75% reduction in average running costs and a jump in OSWorld 2.1 scores from 15.7% to 72.4%.
Quick Take
However, it still lags behind Sonnet 5.5 in complex programming tasks, highlighting the need for further improvements in this area.
Key Points
- Haiku 5.5 achieves 72.4% in OSWorld 2.1, up from 15.7%.
- Average running costs for Haiku 5.5 decreased by approximately 75%.
- Complex programming scores show Haiku 5.5 at 39.2% vs Sonnet 5.5's 70.6%.
- Introduces Adaptive Thinking for adjustable reasoning intensity.
- Supports 1 million tokens context but with tiered pricing structure.
DeepSignal Analysis
What happened
Anthropic's Claude Haiku 5.5 has improved operational capabilities, achieving a 72.4% score in OSWorld 2.1, up from 15.7% in the previous version. The model also boasts a 75% reduction in average running costs. However, it still underperforms compared to Sonnet 5.5 in complex programming tasks, indicating room for further enhancement.
Key evidence
- Claude Haiku 5.5 scored 72.4% in OSWorld 2.1, a significant increase from the previous 15.7%.
- The average running costs of Haiku 5.5 have decreased by approximately 75%, according to Anthropic's data.
- In programming tasks, Haiku 5.5 scored 39.2% in Terminal-Bench 4.0, while Sonnet 5.5 achieved 70.6%.
Why it matters
The advancements in Claude Haiku 5.5 suggest a notable leap in operational efficiency and cost-effectiveness for users. However, the persistent gap in complex programming capabilities compared to Sonnet 5.5 raises questions about the model's suitability for more intricate tasks. This discrepancy highlights the need for ongoing development in this area to meet diverse user needs.
What to watch
📖 Reader Mode
~4 min read
作者丨郑佳美
编辑丨岑 峰


01
复杂编程仍有差距






02
自适应推理加入 Haiku



03
百万上下文没有统一低价




04
迁移仍要检查实际任务表现

上车,带你看遍全球 AI 顶会精华
可独家畅览:
专家演讲PPT
大会报告全文
热门论文解读
学术新星访谈

扫描上方二维码
或点击「阅读原文」关注专区。
雷峰网原创文章,未经授权禁止转载。详情见转载须知。
— Originally published at leiphone.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from 雷峰网 AI
See more →
刚刚,GPT 5.6 发布会上,OpenAI 暴露了哪些 Agent 技术路线?
OpenAI's GPT 5.6 integrates ChatGPT and Codex, introducing a for complex task execution, with models Soul, Terra, and Luna for efficient workflow management. The release emphasizes task orchestration, contextual understanding, and robust security measures for enterprise applications.

