A${}^2$BM: Alignment-Aware Bridge Matching for Image-to-Image Translation
Quick Answer
This paper shows that The A²BM method enhances image-to-image translation by incorporating alignment scores to improve fidelity in weakly aligned pairs, outperforming existing GAN and diffusion models in various tasks, including cross-sensor super-resolution.
Key Points
- A²BM leverages alignment scores to distinguish true correspondences from misalignment artifacts.
- The model shows improved translation fidelity in both synthetic and real-world applications.
- Strongly aligned outputs are achieved by prompting with the highest alignment score.
- A²BM consistently outperforms strong baselines in various image translation tasks.
- The approach addresses challenges in real-world applications with weakly aligned image pairs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Paired image-to-image translation underpins a wide range of computer vision tasks, including image editing, sensor translation, and domain adaptation. Bridge matching and flow matching have recently emerged as powerful frameworks, extending diffusion models to arbitrary source and target distributions. However, their standard formulations assume perfectly aligned training pairs, treating all source-target correspondences as equally reliable. In practice, real-world applications often involve weakly aligned pairs due to changes of acquisition conditions, including e.g. asynchronous captures, different illuminations, or misregistration. In this work, we introduce Alignment-Aware Bridge Matching (A${}^2$BM), a bridge matching method that leverages image pairs alignment during training. By incorporating alignment scores, the model learns to disentangle true semantic correspondences from misalignment artifacts. At inference time, we use the alignment score as a control variable over translation fidelity, with strongly aligned outputs obtained when prompting the model with the highest alignment score. We validate A${}^2$BM on both controlled synthetic experiments and on challenging real-world tasks, including cross-sensor super-resolution and pixel-space unsupervised domain adaptation. In all settings, A${}^2$BM consistently improves translation fidelity over strong GAN-, diffusion-, and Schr{ö}dinger bridge-based baselines, establishing alignment conditioning as a principled solution for image translation models with weakly aligned data.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16294 [cs.CV] |
| (or arXiv:2607.16294v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16294 arXiv-issued DOI via DataCite |
Submission history
From: Georges Le Bellier [view email] [via CCSD proxy]
[v1]
Mon, 13 Jul 2026 08:00:58 UTC (7,800 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.