Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
Quick Answer
Nemotron-Labs-Diffusion is a tri-mode language model that integrates autoregressive, diffusion, and self-speculation decoding, achieving 76.5% more tokens per forward pass than self-speculation.
Quick Take
The model, scaling up to 14B parameters, outperforms existing AR and diffusion models in accuracy and speed, exemplified by its 6x token decoding advantage over Qwen3-8B on SPEED-Bench with SGLang.
Key Points
- Nemotron-Labs-Diffusion unifies AR, diffusion, and self-speculation decoding.
- Achieves 76.5% more tokens per forward pass than self-speculation mode.
- Outperforms Qwen3-8B with 6x more tokens decoded at comparable accuracy.
- Scales to 3B, 8B, and 14B parameters with consistent performance improvements.
- Demonstrates superior efficiency on SPEED-Bench with SGLang on a GB200 GPU.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov
Abstract:We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR-diffusion objective, Nemotron-Labs-Diffusion can switch modes to sustain high throughput across deployment settings and concurrency levels. Our study shows that (1) AR and diffusion objectives are complementary: diffusion improves lookahead planning, while AR provides left-to-right linguistic priors. (2) In self-speculation mode, diffusion drafts while AR verifies, outperforming multi-token prediction (MTP) methods in both acceptance rate and real-device efficiency. (3) A speed-of-light analysis further demonstrates diffusion's long-term potential, with up to 76.5% more tokens per forward pass than self-speculation under an optimal sampler. Scaling to 3B, 8B, and 14B parameters, our Nemotron-Labs-Diffusion family, including base, instruct, and vision-language models, consistently outperforms state-of-the-art open-source AR and diffusion LMs in both accuracy and speed. For example, Nemotron-Labs-Diffusion-8B decodes 6x more tokens per forward than Qwen3-8B with comparable accuracy, translating to 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.05722 [cs.CL] |
| (or arXiv:2607.05722v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05722 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yonggan Fu [view email]
[v1]
Tue, 7 Jul 2026 01:09:54 UTC (2,546 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.