Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning
Quick Answer
This paper shows that This survey evaluates the integration of large language models (LLMs) in healthcare, highlighting a five-level competency scheme for clinical reasoning.
Quick Take
It reveals that specialized medical models outperform general ones in diagnosis tasks, while general models excel in decision support. Key challenges include data limitations and hallucination issues, emphasizing the need for more reliable systems.
Key Points
- Introduces a five-level competency scheme based on Miller's Pyramid for clinical reasoning.
- Specialized medical models excel in diagnosis-centric tasks, outperforming general models.
- General models lead in decision support and dialogue applications.
- Benchmark dataset spans five levels of medical reasoning capability across 18 models.
- Identifies challenges like data limitations and hallucination in current AI systems.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Qi Peng, Jiatong Li, Sirui Huang, Yiyang Jiang, Kaisong Gong, Ronger Ding, Shijie Ye, Changmeng Zheng, Yi Cai, Xiaobo Yang, Jin Huang, Xiao-Yong Wei, Qing Li
Abstract:Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from knowledge recall to dynamic case management. On the computational side, we link deductive, inductive, and abductive reasoning patterns to common medical goals and tasks. We also introduce a benchmark dataset spanning five levels of medical reasoning capability and report results on 18 state-of-the-art models, revealing that medical specialist models excel in diagnosis-centric tasks while general models lead in decision support and dialogue. We conclude by discussing current progress and open challenges, including data limitations, hallucination, and grounding issues, and outline directions toward safer, more reliable, and workflow-ready systems.
| Comments: | Accepted by Machine Intelligence Research |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.07761 [cs.AI] |
| (or arXiv:2607.07761v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.07761 arXiv-issued DOI via DataCite |
Submission history
From: Qi Peng [view email]
[v1]
Wed, 8 Jul 2026 15:19:37 UTC (12,458 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.