An Introduction to Causal Reinforcement Learning
Quick Answer
This paper shows that Causal Reinforcement Learning (CRL) merges causal inference with reinforcement learning, enabling agents to optimize policies by leveraging counterfactual reasoning.
Quick Take
This integration allows for a unified framework encompassing various learning modalities, including online, off-policy, and imitation learning, enhancing the understanding of causal relationships in agent-environment interactions.
Key Points
- CRL connects causal inference principles with reinforcement learning methods.
- It enables agents to reason about counterfactual scenarios without existing data.
- The framework includes online, off-policy, and causal calculus learning modalities.
- New learning settings like generalized policy learning and imitation learning are introduced.
- CRL offers a broader perspective for studying causal inference alongside reinforcement learning.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an environment and pursues an exploratory, trial-and-error approach. These two disciplines have evolved independently and with virtually no interaction between them. We note that they operate over different aspects of the same building block, counterfactual relations, which makes them umbilically connected. Based on these observations, novel learning opportunities arise when this connection is explicitly acknowledged and mathematized. To realize this potential, we note that any environment where the RL agent is deployed can be decomposed as a collection of autonomous mechanisms with different causal invariances, parsimoniously modeled as a structural causal model; any standard RL setting implicitly encodes such a model. This formalization allows us to put under a unifying treatment different modes of learning, including online, off-policy, and causal calculus learning, which appear unrelated in the literature. However, these modalities are not exhaustive: we introduce several natural and pervasive classes of learning settings that entail novel dimensions of analysis. Specifically, we introduce and discuss through causal lenses generalized policy learning, where to intervene, imitation learning, and counterfactual learning. These tasks lead to a broader view of counterfactual learning and suggest great potential for studying causal inference and reinforcement learning side by side, which we call causal reinforcement learning (CRL).
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.24160 [cs.AI] |
| (or arXiv:2606.24160v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24160 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Junzhe Zhang [view email]
[v1]
Tue, 23 Jun 2026 05:28:33 UTC (3,015 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.