Joint Medical Image Enhancement and Segmentation with Diffusion-based Symbiotic Information Interaction
Quick Answer
This paper shows that DiSIINet, a novel Diffusion-based Symbiotic Information Interaction Network, enhances and segments medical images simultaneously, significantly outperforming traditional methods.
Quick Take
By integrating enhancement and segmentation through a Symbiotic Information Interaction module, it improves MRI, CT, and ultrasound image quality and accuracy, demonstrating superior performance on multi-modal datasets.
Key Points
- DiSIINet integrates enhancement and segmentation in a unified model.
- Utilizes Denoising Diffusion Implicit Models for high-quality outputs.
- Features a Symbiotic Information Interaction module for dynamic information exchange.
- Demonstrates significant performance improvements over independent methods.
- Code available at https://github.com/Reconsider80/DiSIINet.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Image quality is critical for accurate medical diagnosis. However, MRI, CT, and ultrasound images are often of low resolution and quality due to cost constraints, complicating the visualization of key anatomical structures and lesions. While such limitations are common in practice, traditional methods treat image enhancement as a separate preprocessing step, failing to fully leverage its potential synergy with image segmentation. To address this, we propose DiSIINet (Diffusion-based Symbiotic Information Interaction Network), which is built on the principle that enhancement and segmentation should mutually reinforce each other in a unified model. Based on Denoising Diffusion Implicit Models (DDIM), DiSIINet integrates an enhancement branch and a segmentation branch. These branches interact through a novel Symbiotic Information Interaction (SII) module, which facilitates dynamic, feature-level information exchange via cross-attention during the reverse diffusion process. This design enables both tasks to iteratively improve each other. The DDIM backbone ensures high-quality output and efficient inference through deterministic sampling. Experiments on multi-modal medical datasets (MRI, CT, ultrasound) show that DiSIINet achieves significant performance improvements compared to sequential or independent enhancement and segmentation approaches. The code is available at: this https URL.
| Comments: | Accepted by IJCAI 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.00058 [cs.CV] |
| (or arXiv:2607.00058v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.00058 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Qiankun Li [view email]
[v1]
Tue, 30 Jun 2026 07:50:44 UTC (3,119 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.