Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
Quick Answer
Med-OPD introduces a novel framework for Medical Vision-Language Models (Med-VLMs) that enhances reasoning by focusing on critical visual evidence.
Quick Take
By employing Medical Evidence Advantage (MEA), it outperforms standard On-Policy Distillation (OPD) and SFT in various medical imaging tasks, demonstrating improved multimodal reasoning capabilities. The approach emphasizes diagnosis-critical tokens, leading to better clinical outcomes.
Key Points
- Med-OPD integrates on-policy distillation with evidence-aware supervision for Med-.
- Medical Evidence Advantage (MEA) focuses teacher scoring on evidence supporting target diagnoses.
- Experiments show Med-OPD outperforms SFT and standard OPD across multiple medical imaging tasks.
- The approach enhances reliance on key visual evidence for improved medical reasoning.
- Source code and data are publicly available for further research.
DeepSignal Analysis
What happened
The Med-OPD framework enhances Medical Vision-Language Models (Med-VLMs) by focusing on critical visual evidence. It introduces Medical Evidence Advantage (MEA) to improve reasoning in medical imaging tasks, outperforming standard On-Policy Distillation (OPD) and Supervised Fine-Tuning (SFT).
Key evidence
- Med-OPD integrates on-policy distillation with medical evidence-aware supervision, marking it as a novel approach for Med-VLMs.
- The framework emphasizes diagnosis-critical tokens, redistributing the distillation signal to enhance reliance on visual evidence.
- Experiments on OmniMedVQA subsets indicate that Med-OPD consistently surpasses SFT and standard OPD across various medical imaging tasks.
Why it matters
Improving reasoning in Med-VLMs is crucial for clinical applications, as existing models often rely on language priors rather than visual evidence. Med-OPD's focus on critical visual evidence could lead to more accurate diagnoses and better patient outcomes, addressing a significant gap in current medical AI capabilities.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions. On-Policy Distillation (OPD) offers dense token-level supervision on student-generated trajectories and provides a privacy-compatible means of capability transfer without requiring the redistribution of raw patient data. However, standard OPD uniformly distills all tokens, causing sparse evidence-dependent tokens to be diluted by abundant clinical narrative tokens. Inspired by the success of OPD in the large language model community, we propose \textbf{Med-OPD}, to our knowledge the first unified post-training framework that integrates on-policy distillation with medical evidence-aware supervision for Med-VLMs. We introduce \textbf{Medical Evidence Advantage} (MEA), a teacher-grounded counterfactual signal that uses an answer-aware hint to focus teacher scoring on evidence supporting the target diagnosis, and measures each token's dependence on medical visual evidence by comparing teacher likelihoods under the original and evidence-degraded imaging modalities. Based on MEA, Med-OPD redistributes the distillation signal at both the token and trajectory levels, emphasizing diagnosis-critical tokens and evidence-reliant rollouts. Experiments on OmniMedVQA subsets show that Med-OPD consistently outperforms SFT and standard OPD across CT, MRI, Disease Diagnosis, and Lesion Grading. These results demonstrate that evidence-aware distillation can better strengthen medical VLMs' reliance on key visual evidence and improve reliable multimodal medical reasoning. The source code and data is publicly available at: this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16303 [cs.CV] |
| (or arXiv:2607.16303v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16303 arXiv-issued DOI via DataCite |
Submission history
From: Yunhang Qian [view email]
[v1]
Tue, 14 Jul 2026 07:08:28 UTC (828 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.