Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Quick Answer
This paper shows that The DiffusionGemma-26B model outperforms its autoregressive counterpart Gemma-4-26B in medical visual question answering, achieving faster decoding and superior drafting capabilities.
Quick Take
This diffusion model allows radiologists to infill report fragments bidirectionally, addressing inconsistencies in clinical reports.
Key Points
- DiffusionGemma-26B matches or exceeds AR performance on all tested datasets.
- The finetuned model operates 3.5-4.4x faster than autoregressive models.
- Diffusion models enable any-order infill, enhancing report drafting.
- Medical foundation models remain predominantly autoregressive despite advancements.
- Results are evaluated by a verbosity-robust judge.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competitive with autoregressive (AR) generation. Medical foundation models, however, remain almost entirely autoregressive. We adapt a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge. Diffusion matches or exceeds AR on all of them, and the finetuned model (3.8B active) is competitive with frontier vision-language models; its decoding is also 3.5-4.4x faster. Beyond this parity, the diffusion model offers a drafting capability AR lacks: any-order infill. Because the canvas is denoised bidirectionally, a radiologist can fix report fragments and have the model fill the text between them, an operation inherent to diffusion but not to autoregression, which is subpar at it. This suits real reports, which are often terse or inconsistent across clinicians and institutions.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.2.6; J.3 |
| Cite as: | arXiv:2607.01436 [cs.AI] |
| (or arXiv:2607.01436v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01436 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Max Van Puyvelde [view email]
[v1]
Wed, 1 Jul 2026 19:59:09 UTC (6,181 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.