Curvature-Guided Mixing for MLLM Adaptation
Quick Answer
This paper shows that Curvature-Guided Mixing (CGM) enhances MLLM adaptation by merging pre-trained and fine-tuned models using a second-order optimization approach.
Quick Take
Experiments on LLaVA-1.5 and Qwen2.5VL demonstrate improved task specialization and general knowledge retention compared to existing methods. The proposed CGM and its variant CGM† show consistent performance gains across multiple downstream tasks.
Key Points
- CGM uses Hessian approximation for optimal soft mixing ratios in model merging.
- CGM† introduces a robust hard mixing variant with curvature-aware parameter selection.
- Experiments show consistent improvements in task specialization and knowledge retention.
- Code for CGM is available on GitHub for further research and application.
- Applicable to various downstream tasks, enhancing MLLM performance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Fine-tuning Multimodal Large Language Models (MLLMs) on specialized tasks often leads to catastrophic forgetting of their general capabilities. Existing model merging methods to combat this are often heuristic or use sub-optimal objectives. We propose CurvatureGuided Mixing (CGM), a theoretically grounded framework that merges pre-trained and fine-tuned models. CGM formulates a joint optimization objective and uses a second-order (Hessian) approximation of the loss landscapes to analytically derive an optimal, closed-form "soft mixing" ratio. This ratio intelligently blends parameters based on their relative task-specific curvatures. We also introduce CGM$\dagger$, a robust "hard mixing" variant that performs sparse parameter selection guided by a novel, curvature-aware score. Experiments on LLaVA-1.5 and Qwen2.5VL across multiple downstream tasks show that CGM and CGM$\dagger$ consistently improve the trade-off between task specialization and general knowledge retention over existing methods. Code is available at this http URL.
| Comments: | Accepted to ECCV 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.24963 [cs.CV] |
| (or arXiv:2606.24963v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24963 arXiv-issued DOI via DataCite |
Submission history
From: Jinglong Yang [view email]
[v1]
Tue, 23 Jun 2026 09:21:54 UTC (1,662 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.