FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning
Quick Answer
FaithMed enhances medical reasoning by integrating clinician-designed rubrics with reinforcement learning, achieving a 9% improvement over agentic-search baselines and a 15.5% increase in evidence-based rubric scores across seven benchmarks.
Quick Take
This framework ensures transparent, evidence-grounded clinical decisions.
Key Points
- FaithMed combines clinician-designed rubrics with reinforcement learning for improved medical reasoning.
- Achieved a 9% average improvement over agentic-search baselines across seven medical benchmarks.
- Increased evidence-based medicine rubric scores by 15.5% compared to agentic-search Qwen3.
- Explicit step-level supervision enhances task success and reasoning faithfulness.
- Code for FaithMed is available on GitHub.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Faithful reasoning is essential in medicine, where clinical decisions require transparent justification grounded in reliable evidence. Current medical LLMs either lack active access to evidence or use retrieved evidence without supervising how it should be appraised and applied during reasoning. To address this, we formalize evidence-based medicine principles as process-level criteria and introduce FaithMed, a framework that combines clinician-designed, automatically refined rubrics with reinforcement learning using step-level process reward assignment and advantage grouping. Across seven medical benchmarks, FaithMed improves over agentic-search baselines (+9% on average) and outcome-only RL (+5.8%), while raising average evidence-based medicine rubric scores over agentic-search Qwen3 baselines (+15.5%). This work demonstrates that explicit step-level supervision can improve both task success and the faithfulness of the reasoning process. Code is available at this https URL.
| Comments: | 15 pages, 5 figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.01440 [cs.CL] |
| (or arXiv:2607.01440v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01440 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhiyun Zhang [view email]
[v1]
Wed, 1 Jul 2026 20:02:55 UTC (947 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.