Plan-and-Patch: Diffusion Language Models for Agentic Planning
Quick Answer
This paper shows that The Plan-and-Patch framework utilizes diffusion language models like DreamReasoner-8B and Qwen3-8B for efficient planning and repair in long-horizon agents.
Quick Take
It achieves a plan repair success rate of 53.7% compared to 27.0% for autoregressive methods, while reducing plan-generation latency by 39-46% after task-specific training on benchmarks like ALFWorld and TextCraft.
Key Points
- Plan-and-Patch generates structured plans through parallel unmasking and targeted repairs.
- Diffusion models outperform autoregressive planners in plan repair success rates.
- After task-specific training, both planners show similar success in plan generation.
- Diffusion reduces mean plan-generation latency significantly compared to autoregressive methods.
- The framework enhances the effectiveness of in dynamic environments.
DeepSignal Analysis
What happened
The Plan-and-Patch framework employs diffusion language models, specifically DreamReasoner-8B and Qwen3-8B, to enhance planning and repair processes for long-horizon agents. It achieves a plan repair success rate of 53.7%, significantly higher than the 27.0% success rate of autoregressive methods, while also reducing plan-generation latency by 39-46% after task-specific training.
Key evidence
- The Plan-and-Patch framework utilizes diffusion language models like DreamReasoner-8B and Qwen3-8B for planning and repair in long-horizon agents.
- Diffusion models achieved a plan repair success rate of 53.7%, compared to 27.0% for autoregressive methods, indicating a substantial performance difference.
- After task-specific training on benchmarks such as ALFWorld and TextCraft, diffusion models reduced mean plan-generation latency by 39-46% relative to autoregressive models.
Why it matters
The ability to efficiently generate and repair plans is crucial for long-horizon agents, as they must adapt to changing environments and unexpected outcomes. The significant improvement in repair success rates and reduced latency offered by the Plan-and-Patch framework suggests a potential shift in how agents can be designed for complex tasks. This could lead to more robust AI systems capable of handling real-world scenarios where flexibility and quick adaptation are essential.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed. We compare DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive (AR) planners. On Natural Plan without task-specific training, diffusion (53.7%) achieves nearly twice the plan repair success rate of AR (27.0%). After task-specific training on agentic benchmarks, ALFWorld and TextCraft, the planners achieve similar observed success in plan generation, while diffusion reduces mean plan-generation latency by 39-46% relative to AR. Our results show that Plan-and-Patch provides a framework for faster plan generation and effective plan repair in long-horizon agents.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10786 [cs.AI] |
| (or arXiv:2610.10786v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10786 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Syamantak Kumar [view email]
[v1]
Wed, 7 Oct 2026 18:45:45 UTC (50 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.