From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
Quick Answer
This study presents a method to convert deep reinforcement learning policies into executable Prolog logic programs, enhancing explainability.
Quick Take
The approach achieves optimal returns in a two-room key-and-door task and matches neural performance in continuous-control tasks, demonstrating significant improvements in interpretability and policy editing capabilities.
Key Points
- Transforms deep reinforcement learning policies into Prolog programs for better explainability.
- Achieves exact optimal return in a two-room key-and-door task with 16,944 states.
- Surpasses stochastic teacher performance in a budget-capped regime across ten seeds.
- Matches neural teacher performance in Acrobot and recovers 97% return on CartPole.
- Demonstrates exponential cost in observation dimension for oblique decision boundaries.
DeepSignal Analysis
What happened
This study introduces a method for transforming deep reinforcement learning policies into executable Prolog logic programs, enhancing their explainability. The approach demonstrates optimal performance in a specific task and matches neural network performance in continuous-control scenarios, indicating improvements in interpretability and policy modification.
Key evidence
- The method involves a three-stage transformation that converts a proximal policy optimization teacher into a Prolog program, allowing for human readability and machine execution.
- In a two-room key-and-door task with 16,944 states, the Prolog program achieved optimal returns across all trials, outperforming the stochastic teacher in a budget-capped scenario.
- For continuous-control tasks, the Prolog program matched the neural teacher's performance within noise on the Acrobot task and recovered approximately 97% of the return on the CartPole task.
Why it matters
The ability to convert deep reinforcement learning policies into interpretable logic programs could significantly enhance the transparency of AI systems. This transformation allows stakeholders to understand decision-making processes better and facilitates policy adjustments, which is crucial for applications in sensitive areas like healthcare and finance where explainability is paramount.
Paper Resources
📖 Reader Mode
~2 min readAbstract:A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.15459 [cs.AI] |
| (or arXiv:2607.15459v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15459 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Eduardo C. Garrido-Merchán [view email]
[v1]
Thu, 16 Jul 2026 21:10:27 UTC (49 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.