MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models
Quick Answer
MoR-MLLM introduces a computation-sparse framework for Multimodal Large Language Models, enabling adaptive recursion per token to optimize resource usage.
Quick Take
This model significantly reduces training memory and computational complexity while maintaining high performance on vision-language tasks compared to existing tiny MLLMs.
Key Points
- MoR-MLLM uses adaptive per-token recursion for dynamic computation allocation.
- The model reduces training memory and computation complexity significantly.
- It retains high performance on various vision-language tasks.
- A three-stage MoR-Tuning strategy stabilizes recursive sparsity training.
- Entropy-regularized loss encourages diverse routing distributions.
DeepSignal Analysis
What happened
The MoR-MLLM framework introduces a computation-sparse approach for Multimodal Large Language Models, focusing on adaptive recursion per token. This method aims to optimize resource usage while maintaining performance on vision-language tasks.
Key evidence
- MoR-MLLM employs adaptive per-token recursion, allowing dynamic adjustment of recursive depth based on the complexity of each token.
- The model significantly reduces training memory and computational complexity compared to existing tiny MLLMs, according to extensive experiments.
- A three-stage MoR-Tuning strategy and an entropy-regularized loss are designed to stabilize training and encourage diverse routing distributions.
Why it matters
The development of MoR-MLLM addresses the high computational and memory demands of existing Multimodal Large Language Models, which have limited their real-world application. By optimizing resource usage, this model could facilitate broader deployment in practical scenarios, potentially enhancing the efficiency of AI systems that integrate vision and language.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAuthors:Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen
Abstract:Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.
| Comments: | 13 pages |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08830 [cs.CV] |
| (or arXiv:2610.08830v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08830 arXiv-issued DOI via DataCite |
Submission history
From: Pengcheng Zheng [view email]
[v1]
Sun, 27 Sep 2026 08:33:26 UTC (884 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.