
AI chatbots reading X-rays can be dangerously confident even when they're wrong
Quick Answer
AI chatbots struggle with X-ray interpretation, scoring significantly lower than human radiologists, with the best model, Google's Gemini 3 Pro, achieving only 758 points out of 2000.
Quick Take
The study emphasizes the dangers of overconfidence in AI, as many models frequently misdiagnose while exhibiting high confidence, undermining their reliability in medical contexts.
Key Points
- Human radiologists scored 988.7 points, outperforming all AI models tested.
- The scoring system penalizes overconfidence, rewarding honesty in diagnoses.
- Meta's Muse Spark 1.1 excelled at recognizing when to defer to human experts.
- Patients increasingly trust AI chatbots for medical image interpretations.
- AI models often misdiagnose with high confidence, raising safety concerns.
DeepSignal Analysis
What happened
A recent study evaluated 16 AI models against human radiologists in interpreting X-rays, revealing that human experts scored significantly higher. The best AI model, Google's Gemini 3 Pro, achieved only 758 points out of 2000, while human radiologists scored 988.7. The study highlights the issue of AI overconfidence, as many models misdiagnose while exhibiting high confidence, raising concerns about their reliability in medical settings.
Key evidence
- The study involved 200 cases and found that human radiologists outperformed all AI models, scoring 988.7 out of 2000 points compared to the best AI model's score of 758.
- Meta's Muse Spark 1.1 was noted for its ability to recognize when to defer cases to human radiologists, showing a significant reduction in its hallucination rate.
- The research team criticized claims that AI systems diagnose better than 99 percent of doctors, stating that such assertions are often based on anecdotes or simulations rather than empirical evidence.
Why it matters
The findings underscore the limitations of current AI models in medical diagnostics, particularly their tendency to misdiagnose with high confidence. This overconfidence can lead to dangerous outcomes in clinical settings, where accurate diagnosis is critical. The study also raises ethical concerns regarding the reliance on AI in healthcare, as patients increasingly trust chatbots for medical advice despite their unreliability.
📖 Reader Mode
~4 min readThe test ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758.

Honest silence beats overconfident guesswork
The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer "I don't know," and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent.
The study tackles a point recently raised by this highly cited paper: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty.
No single model wins across the board
There's no overall winner. Anthropic's Claude Fable 5 performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy.

Meta's Muse Spark 1.1 was the best at knowing when to hand a case off to a human. Meta had recently cut Muse Spark 1.1's hallucination rate nearly in half because the model more often refuses to answer rather than giving a wrong one. Other frontier models trend the opposite way. Grok 4.5, for example, hallucinates significantly more than its predecessor because while it knows more, it's also more convinced of its wrong answers.

According to the research team, several models would have scored much better if they had stayed quiet more often instead of guessing. This was especially obvious among open-weight models and those trained specifically for medical use. They tried to answer nearly every case and were often wrong, usually with high confidence.


The first version of the test painted an even starker picture. Radiologists hit 83 percent accuracy, while the best model managed only about 30 percent. Within three months, Gemini 3 Pro had already surpassed the level of resident radiologists. Accuracy is growing fast, but the models still lack any sense of their own limits.
Patients are already sending their MRIs to chatbots
More and more people are uploading X-rays or MRI scans to chatbots and trusting the responses. A recent study in npj Digital Medicine showed that widely used chatbots frequently give unreliable answers to medical questions.
The research team accuses executives and investors of publicly overstating what AI models can do. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations. As recently as April, a study of 21 models that were then considered state-of-the-art showed they aren't ready for unsupervised clinical use.
RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced.
Two other recent studies on autonomous medical AI agents pointed in a different direction. MIRA, a system for electronic health records, and AMIE were able to keep pace with general practitioners in simulated consultations. Both fueled expectations that AI could soon make diagnoses on its own. The RadLE 2.0 authors push back: before an AI makes decisions independently, it has to know when it's better off not doing so.
Then there's the problem of skill loss. A Polish observational study from 2025 found that doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions without the tool. Detection rates dropped from 28.4 to 22.4 percent. The authors call it the "Google Maps effect": without the navigation aid, users are lost.
Radiology has seen this AI hype before
Radiology already went through one AI hype cycle. In 2016, AI researcher Geoffrey Hinton declared that we should stop training radiologists because deep learning would soon take over the job. Colleagues like Richard Sutton agreed.
Nearly ten years later, radiologists are still overburdened, and Hinton had to walk back his prediction. He had reduced the profession to image analysis and overlooked the complexity of the entire field. The fact that these systems can confidently produce wrong diagnoses means humans remain indispensable.
OpenAI CEO Sam Altman spent years predicting that AI would replace human jobs at a scary pace, then recently walked it back, suggesting AI may have actually created more jobs. So far, research doesn't support either claim.
AI specialists may understand their models, but they routinely overestimate how fast entire professions can be replaced. Those kinds of predictions are back in fashion right now. Much like the AI they build, even people don't always know when they'd be better off staying quiet because they're outside their own expertise.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

