CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models
Quick Answer
CoDR introduces a training-free method to address confidence drift in masked diffusion language models, enhancing accuracy across various tasks with minimal overhead.
Quick Take
By remasking tokens based on their confidence levels, CoDR outperforms existing methods while requiring fewer forward passes, demonstrating significant improvements in model performance.
Key Points
- CoDR improves average accuracy across multiple model-sampler configurations.
- The method requires only k forward passes for confidence drift estimation.
- Targeted remasking leads to performance gains beyond just increased compute.
- CoDR shows improvements across two backbones and four reasoning tasks.
- Code for CoDR is publicly available for further research.
DeepSignal Analysis
What happened
CoDR introduces a method to mitigate confidence drift in masked diffusion language models without requiring additional training. By remasking tokens based on their confidence levels, it enhances model accuracy across various tasks while minimizing computational overhead.
Key evidence
- CoDR operates by estimating confidence drift for committed tokens using k-partition probing, which requires only k forward passes.
- The method shows improvements in average accuracy across multiple configurations, including two model backbones and four reasoning and coding tasks.
- Controlled experiments indicate that the performance gains are attributed to targeted remasking rather than just increased computational resources.
Why it matters
The introduction of CoDR could significantly enhance the performance of masked diffusion language models, which are increasingly used in natural language processing tasks. By addressing confidence drift effectively, it allows models to maintain accuracy even when initial context becomes less relevant, potentially leading to better outcomes in real-world applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Masked diffusion language models (MDLMs) decode by repeatedly committing tokens to masked positions, but these commitments are usually irreversible. A token chosen under sparse, partial context is kept fixed, even when later context no longer supports it. Existing samplers mainly decide when to commit a token, but rarely check whether an already committed token should still be kept, allowing early mistakes to propagate. We trace this issue to confidence drift, where the model's confidence in a committed token drops from its sparse commit-time context to the denser context available later. Based on this signal, we propose CoDR (Confidence Drift Remasking), a training-free and sampler-agnostic refinement pass. CoDR estimates drift for all committed positions in only k forward passes via k-partition probing, then remasks and regenerates only the tokens the model no longer endorses. Across two backbones, four reasoning and coding tasks, and three base samplers, CoDR improves average accuracy across all evaluated model-sampler configurations and improves most individual task settings with modest overhead. Controlled experiments show that the gains come from targeted confidence-drift remasking rather than extra compute alone, and that CoDR uses far fewer forward passes than prior remasking methods. Code is available at this https URL.
| Comments: | 17 pages, 6 figures, and 17 tables |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08833 [cs.CL] |
| (or arXiv:2610.08833v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08833 arXiv-issued DOI via DataCite |
Submission history
From: Jian Huang [view email]
[v1]
Mon, 28 Sep 2026 16:43:09 UTC (864 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.