
ChatGPT's new health upgrade beats doctor-written answers, OpenAI says
Quick Answer
OpenAI's ChatGPT has been upgraded to GPT-5.5 Instant, outperforming doctor-written answers in healthcare accuracy, clarity, and completeness.
Quick Take
The model's error rate for health-related statements has decreased by 71%, showcasing significant advancements in AI-driven healthcare responses.
Key Points
- GPT-5.5 Instant significantly enhances ChatGPT's healthcare capabilities.
- Model outperforms doctor-written responses in accuracy and clarity.
- Error rate for health-related statements decreased by 71%.
- OpenAI conducted comparative tests to validate performance improvements.
- Implications for AI in healthcare are profound, potentially reshaping patient interactions.
📖 Reader Mode
~1 min readOpenAI has upgraded ChatGPT's healthcare capabilities with the GPT-5.5 Instant model. The updated model matches the performance of the most expensive Thinking models on machine-based health tests like HealthBench and HealthBench Professional, but at a fraction of the cost. GPT-5.5 Instant is available to all free ChatGPT users, though with usage limits.
When compared to doctors, GPT-5.5 Instant's responses scored higher in accuracy, clarity, and completeness. The rate of incorrect health statements has dropped by 71 percent over the past two months.

A network of over 260 doctors from 60 countries is behind these improvements. They've reviewed more than 700,000 model responses. According to OpenAI, more than 230 million people use ChatGPT weekly for health-related questions, things like understanding lab results, prepping for doctor's appointments, or sorting out insurance questions. OpenAI also offers specialized tools for healthcare professionals, including ChatGPT for Clinicians and OpenAI for Healthcare.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

