MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models
Quick Answer
This paper shows that MedBench v5 introduces a dynamic, process-oriented benchmark for clinical multimodal models, enhancing evaluation with 63 tasks across 14 cognitive dimensions and 4 agent environments, while addressing hallucination detection and process stability issues.
Quick Take
Experiments reveal that high task performance does not ensure reliability under stressors like contradiction detection and evidence delay.
Key Points
- Features a dual-dimensional framework with 63 tasks for comprehensive skill evaluation.
- Incorporates three stressors for analyzing model performance degradation.
- Utilizes a dynamic audit protocol to identify model-specific failure fingerprints.
- Monitors hallucination propagation across various stages of reasoning.
- Demonstrates that high overall performance does not guarantee process stability.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Ding Jinru, Jiang Chuchu, Lu Lu, Pang Wenrao, Bian Mouxiao, Gao Zhuangzhi, Chen Jiangyuan, Peng xinwei, Chen Ruiyao, Ren Sijie, Lu Renjie, Han Bin, Liu Meiling, and Xu Jie
Abstract:Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation. MedBench v5 features: (1) a dual-dimensional framework combining Clinical Cognitive Responsiveness (14 sub-dimensions) and Medical Atomic Skills (4 agent environments), covering 63 tasks; (2) three switchable information-flow stressors (omission, contradiction, evidence delay) for factorized degradation analysis; (3) a dynamic process audit protocol with five reasoning nodes that produces model-specific failure fingerprints; (4) hallucination propagation monitoring across initiation, propagation, anchoring, and contradiction interaction-capturing silent hallucination. Experiments on frontier models show that strong overall task performance does not guarantee process stability: stressors mainly disrupt contradiction detection, diagnosis updating, hallucination propagation, and contradiction-based self-correction, while final evidence grounding can remain superficially stable. MedBench v5 provides a unified infrastructure for capability profiling, controllable stress testing, process auditing, and hallucination trajectory analysis in clinical AI evaluation.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.24155 [cs.CL] |
| (or arXiv:2606.24155v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24155 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chuchu Jiang [view email]
[v1]
Tue, 23 Jun 2026 05:23:46 UTC (310 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.