GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
Quick Answer
GuideSkill introduces an external reasoning layer for LLMs, enhancing clinical reasoning by executing disease-specific criteria.
Quick Take
GuideSkill-Zero improves accuracy by 13.45% over guideline , while GuideSkill-Evo boosts macro-average accuracy by 18.49% and increases skill coverage from 56.5% to 99.5%. Expert evaluations confirm its clinically sound skills, indicating practical reliability.
Key Points
- GuideSkill-Zero improves macro-average accuracy by 13.45% over guideline RAG.
- GuideSkill-Evo achieves an 18.49% accuracy increase over direct inference.
- Gold-label skill coverage rises from 56.5% to 99.5% with GuideSkill-Evo.
- Expert evaluations indicate GuideSkill's skills are clinically sound and reliable.
- GuideSkill combines guideline-derived procedures with case-derived diagnostic patterns.
DeepSignal Analysis
What happened
GuideSkill introduces an external reasoning layer for LLMs to enhance clinical reasoning by executing disease-specific criteria. GuideSkill-Zero improves accuracy by 13.45% over guideline RAG, while GuideSkill-Evo increases macro-average accuracy by 18.49% and skill coverage from 56.5% to 99.5%. Expert evaluations confirm the clinical reliability of its skills.
Key evidence
- GuideSkill-Zero improves macro-average accuracy by 13.45% over guideline RAG across four benchmarks.
- GuideSkill-Evo achieves an 18.49% improvement in macro-average accuracy over direct inference and increases skill coverage from 56.5% to 99.5%.
- Expert evaluations indicate that GuideSkill produces clinically sound skills, suggesting practical reliability.
Why it matters
The development of GuideSkill represents a significant advancement in integrating clinical practice guidelines into LLMs, potentially enhancing diagnostic accuracy in clinical settings. By executing rules rather than merely retrieving text, it addresses a critical gap in how AI can support healthcare professionals. The improvements in accuracy and skill coverage suggest that such systems could lead to better patient outcomes and more reliable clinical decision-making.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scores. GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case--diagnosis pairs to refine covered skills and add missing diagnoses. At inference, an LLM proposes a differential diagnosis, grounds the features required by each matched skill, and fuses its ranking with the executed skill scores. Across four benchmarks and four backbones, GuideSkill-Zero improves macro-average accuracy over guideline RAG by 13.45% on average. GuideSkill-Evo achieves the highest macro-average for every backbone, improves over direct inference by 18.49% relatively, and increases gold-label skill coverage from 56.5% to 99.5%. On Qwen3.5-9B, it also exceeds the strongest parameter-update baseline by 11.16% without updating the backbone. Expert evaluation further indicates that GuideSkill produces clinically sound and broadly acceptable skills, suggesting that its initialized and evolved rules are reliable and practically meaningful. These results support executable skills as a model-agnostic mechanism for combining guideline-derived procedures with case-derived diagnostic patterns.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.26160 [cs.AI] |
| (or arXiv:2607.26160v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26160 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lang Cao [view email]
[v1]
Tue, 28 Jul 2026 18:10:33 UTC (636 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.