Learning What to Trust in Multimodal Learning under Noisy Supervision
Quick Answer
The REFINE framework enhances multimodal learning by effectively detecting label noise using fused and unimodal representations, improving model generalization despite noisy supervision.
Quick Take
Extensive experiments show REFINE outperforms traditional methods in diverse tasks, providing cleaner supervision for multimodal classifiers.
Key Points
- REFINE constructs discriminative eigenvectors for improved noise detection in multimodal settings.
- The framework selects trusted representation spaces for each class to enhance label noise detection.
- REFINE combines subsets from trusted spaces to provide cleaner supervision for training.
- Extensive experiments demonstrate REFINE's superiority over baseline methods in various tasks.
- Source code for REFINE will be publicly available for further research.
DeepSignal Analysis
What happened
The REFINE framework aims to enhance multimodal learning by addressing label noise through the use of both fused and unimodal representations. This approach is designed to improve model generalization in scenarios where high-quality labels are scarce. The framework has been tested across various tasks, demonstrating its effectiveness compared to traditional methods.
Key evidence
- REFINE is a multimodal label-noise detection framework that utilizes both fused and unimodal representations to identify label noise.
- The framework constructs discriminative eigenvectors through analysis of target and background classes, allowing for better noise detection.
- Extensive experiments show that REFINE outperforms baseline methods in diverse tasks, providing cleaner supervision for multimodal classifiers.
Why it matters
The challenge of noisy labels in multimodal learning is significant, as traditional methods often do not leverage multimodal information effectively. REFINE's approach could lead to more robust models that perform better in real-world applications where data quality is inconsistent. By improving noise detection, REFINE may enhance the reliability of multimodal classifiers, which is crucial for tasks in fields like computer vision and machine learning.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11057 [cs.CV] |
| (or arXiv:2610.11057v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11057 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaobo Xia [view email]
[v1]
Thu, 8 Oct 2026 01:14:41 UTC (1,999 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.