Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Quick Answer
This paper shows that Masked Diffusion Language Models (MDLMs) outperform autoregressive models in text-based world modeling, achieving up to 47% absolute gains in zero-shot transfer across diverse environments.
Quick Take
The study introduces a training framework and provides 239,403 state-action trajectories, demonstrating improved coherence and groundedness without environment-specific fine-tuning.
Key Points
- MDLMs achieve better coherence and groundedness than AR models with 4x the parameters.
- Introduced a GRPO training framework with deterministic state checks for enhanced performance.
- 239,403 grounded state-action trajectories were curated across nine open-source environments.
- Zero-shot transfer ablations showed up to 47% gains without environment-specific fine-tuning.
- Behavioral analysis revealed failure modes and human evaluation on realism and utility.
DeepSignal Analysis
What happened
The study introduces Masked Diffusion Language Models (MDLMs) as a new approach to text-based world modeling, demonstrating significant performance improvements over autoregressive models. It presents a GRPO training framework and a dataset of 239,403 state-action trajectories, achieving up to 47% gains in zero-shot transfer across various environments without specific fine-tuning.
Key evidence
- MDLMs outperform autoregressive models in text-based world modeling, achieving up to 47% absolute gains in zero-shot transfer across diverse environments.
- The study curated 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families.
- MDLMs provide better coherence, groundedness, and rollout diversity than autoregressive models, despite being over four times their parameter size.
Why it matters
This research highlights the limitations of traditional autoregressive models in reinforcement learning and proposes MDLMs as a more effective alternative for creating diverse training environments. The ability to achieve significant performance gains without environment-specific fine-tuning could streamline the development of more robust AI agents, enhancing their adaptability and effectiveness in real-world applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial state, task context, tool schemas, domain rules, and steering directives, and (ii) curate 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families. We compare AR LMs and masked diffusion language models (MDLMs), showing MDLMs, via bidirectional anchor-aware denoising, achieve better coherence, groundedness, and empirically validated rollout diversity than LLMs over 4x their parameter size, at comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three OOD environments (ScienceWorld, ALFWorld, AppWorld) across three 1.2B-7B agent backbones (LFM2.5, Qwen3, Mistral), achieving up to 47% absolute gains over baselines without environment-specific fine-tuning. We further conduct behavioral analysis of failure modes under adversarial scenarios and human evaluation on realism, outcome correctness, and training utility. We open-source our work to encourage research in this direction.
| Comments: | Dataset: this https URL Training code: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.16204 [cs.AI] |
| (or arXiv:2607.16204v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16204 arXiv-issued DOI via DataCite |
Submission history
From: Darshan Deshpande [view email]
[v1]
Thu, 7 May 2026 00:40:32 UTC (211 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.