DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
Quick Answer
DC-Leap introduces a training-free acceleration framework for diffusion large language models (dLLMs), achieving speedups of up to 53.19x on long-sequence generation tasks.
Quick Take
By utilizing a Dynamic Contiguous Verification strategy and draft-guided decoding, it effectively mitigates Joint Probability Dependence Error (JPDE) while maintaining comparable performance. Extensive benchmarks demonstrate significant improvements in inference speed without sacrificing quality.
Key Points
- Achieves up to 53.19x speedup on long-sequence generation tasks.
- Introduces Dynamic Contiguous Verification to address JPDE issues.
- Incorporates draft-guided decoding for improved context retention.
- Maintains comparable performance while enhancing inference speed.
- Code available for implementation and further research.
DeepSignal Analysis
What happened
DC-Leap is a new framework designed to accelerate diffusion large language models (dLLMs) without requiring additional training. It achieves speed improvements of up to 53.19 times for long-sequence generation tasks by addressing inefficiencies in current decoding strategies.
Key evidence
- DC-Leap introduces a Dynamic Contiguous Verification strategy that integrates causal constraints into the parallel decoding process, which helps to reduce redundant iterations.
- The framework mitigates Joint Probability Dependence Error (JPDE), allowing for faster inference speeds while maintaining performance quality.
- Extensive benchmarks indicate that DC-Leap can achieve speedups of up to 105.02 times when combined with KV-Cache, demonstrating significant improvements in inference speed.
Why it matters
The development of DC-Leap is significant as it addresses the limitations of existing decoding methods in dLLMs, particularly in terms of speed and efficiency. By eliminating the need for training, it offers a practical solution for enhancing the performance of language models, which is crucial for real-time applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime. DC-Leap introduces a Dynamic Contiguous Verification strategy that integrates strictly-ordered causal constraints into the parallel decoding process. By progressively validating token dependencies, this mechanism effectively neutralizes the JPDE, enabling reliable acceleration with comparable performance. Furthermore, DC-Leap incorporates the draft-guided decoding mechanism, where the draft helps extend the context by leaping forward across multiple tokens, providing look-ahead context and retaining the structural benefits of bidirectional attention during inference. Extensive experiments on standard benchmarks demonstrate that DC-Leap achieves substantial speedups, up to 53.19x on MBPP for long-sequence generation, and up to 105.02x when combined with KV-Cache with comparable generation quality. Code is available at this https URL .
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.20467 [cs.AI] |
| (or arXiv:2607.20467v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20467 arXiv-issued DOI via DataCite |
Submission history
From: Yulin Li [view email]
[v1]
Tue, 19 May 2026 06:27:58 UTC (1,756 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.