DiffusionGemma Technical Report
Quick Answer
DiffusionGemma is a novel open-weight language model that utilizes discrete diffusion for rapid text generation, achieving around 1,500 tokens per second on an NVIDIA H100 GPU.
Quick Take
By fine-tuning the Gemma 4 model with 3.8B activated parameters, it overcomes the sequential decoding limitations of traditional autoregressive models, generating 20 tokens per forward pass and maintaining multimodal input support.
Key Points
- Generates 1,500 tokens per second on a single NVIDIA H100 GPU.
- Fine-tuned from the Gemma 4 model with 3.8B activated parameters.
- Achieves 20 tokens per forward pass, surpassing traditional AR models.
- Utilizes a two-stage training pipeline for efficiency and quality.
- Supports multimodal inputs and retains autoregressive generation capabilities.
DeepSignal Analysis
What happened
DiffusionGemma is a new open-weight language model that employs discrete diffusion for fast text generation, achieving approximately 1,500 tokens per second on an NVIDIA H100 GPU. It fine-tunes the Gemma 4 model with 3.8 billion activated parameters, allowing it to generate 20 tokens per forward pass.
Key evidence
- DiffusionGemma generates text at around 1,500 tokens per second on a single NVIDIA H100 GPU, significantly outpacing traditional autoregressive models.
- The model fine-tunes the Gemma 4 model, which has 3.8 billion activated parameters and 25.2 billion total parameters, using less than 10% of the original training token budget.
- It maintains multimodal input support and can perform autoregressive generation with only minor performance degradation, indicating potential for hybrid decoding methods.
Why it matters
The introduction of DiffusionGemma represents a significant advancement in language model technology, particularly in overcoming the limitations of sequential decoding found in traditional autoregressive models. Its ability to generate text rapidly while retaining multimodal capabilities could enhance applications in various fields, including natural language processing and AI-driven content generation. This model sets a new benchmark for balancing generation speed and model capability.
Paper Resources
📖 Reader Mode
~2 min readAuthors:DiffusionGemma Team: Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
Abstract:We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2608.00146 [cs.CL] |
| (or arXiv:2608.00146v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00146 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jean Tarbouriech [view email]
[v1]
Fri, 31 Jul 2026 16:11:46 UTC (6,116 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.