Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention
Quick Answer
The study presents a Concept Bottleneck Model (CBM) that enhances interpretability in medical imaging by allowing concept-level interventions.
Quick Take
Evaluated on Mayo Clinic's ultrasound and CheXpert chest X-ray datasets, the framework enables reliable model diagnosis and can improve predictive performance through guided fine-tuning.
Key Points
- Introduces a plug-and-play framework for concept-based model refinement.
- Aligns a single-modality encoder to BioMedCLIP for enhanced interpretability.
- Evaluated on Mayo Clinic ultrasound and CheXpert datasets.
- Concept interventions help isolate causal concepts and validate insights.
- Guided fine-tuning can maintain or improve predictive performance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.
| Comments: | Accepted at the 5th Workshop on Applications of Medical AI (AMAI), MICCAI 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09031 [cs.CV] |
| (or arXiv:2610.09031v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09031 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Samrajya Thapa [view email]
[v1]
Tue, 6 Oct 2026 19:31:42 UTC (10,646 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.