Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
Quick Answer
Omni-Diffusion-Distill introduces a two-stage distillation framework for unified multimodal diffusion large language models, achieving 18.2x speedup in image generation and 21.2x in multimodal understanding.
Quick Take
It reduces decoding steps from 128 to 8 for images and 512 to 64 for understanding, while maintaining high performance on benchmarks like GenEval and DPG-Bench.
Key Points
- Achieves state-of-the-art efficiency for multimodal dLLMs with Omni-Diffusion-Distill.
- Reduces image generation decoding steps from 128 to 8, enhancing speed significantly.
- Improves multimodal understanding decoding from 512 to 64 steps.
- Scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation.
- Surpasses teacher model performance on multimodal understanding tasks.
DeepSignal Analysis
What happened
Omni-Diffusion-Distill presents a two-stage framework for distilling unified multimodal diffusion large language models (dLLMs). This method reduces decoding steps significantly, achieving 18.2x speedup in image generation and 21.2x in multimodal understanding while maintaining performance on key benchmarks.
Key evidence
- The framework reduces image generation decoding steps from 128 to 8 and multimodal understanding from 512 to 64.
- Omni-Diffusion-Distill achieves a score of 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation.
- For multimodal understanding, it reaches GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning, which is double the teacher's score.
Why it matters
This advancement in distillation techniques could significantly enhance the efficiency of multimodal models, which are increasingly important in AI applications. By reducing inference costs while maintaining performance, Omni-Diffusion-Distill may enable broader deployment of these models in real-world scenarios, potentially impacting industries reliant on image and text processing.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10990 [cs.CV] |
| (or arXiv:2610.10990v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10990 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hong Huang [view email]
[v1]
Wed, 7 Oct 2026 23:24:16 UTC (7,611 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.