Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model
Quick Answer
The PASE framework introduces a Planning-Aware Semantic self-healing engine that utilizes LLMs for generating recovery plans and a Neural-Symbolic World Model for plan verification, achieving over 40% reduction in recovery time and improved fault detection accuracy in cloud systems.
Key Points
- PASE redefines recovery as a neuro-symbolic program synthesis task.
- Utilizes to generate structured recovery plans from semantic primitives.
- Achieves over 40% reduction in average system recovery time.
- Improves fault detection accuracy in unknown fault scenarios.
- Integrates reasoning, verification, and meta-learning for adaptive recovery.
Paper Resources
📖 Reader Mode
~2 min readAbstract:As the scale and complexity of cloud-based AI systems continue to escalate, ensuring service reliability through rapid fault detection and adaptive recovery has become a critical challenge. While existing approaches integrate Large Language Models (LLMs) for semantic understanding and Deep Reinforcement Learning (DRL) for policy optimization, they often rely on sequential, loosely coupled architectures that underutilize the generative and reasoning capabilities of LLMs. In this paper, we propose a paradigm shift with PASE, a Planning-Aware Semantic self-healing engine, a novel fault self-healing framework that reconceptualizes recovery as a neuro-symbolic program synthesis task. PASE employs an LLM as a core Plan Synthesis Engine to generate structured recovery plans from a library of semantic primitives. A Neural-Symbolic World Model verifies plan feasibility through simulation, while a Meta-Prompt Optimizer, trained via DRL, learns to generate optimal prompts that guide the LLM's planning process. This tight reason-plan-verify-adapt loop enables dynamic, context-aware recovery strategy generation beyond predefined action spaces. Experiments on a real-world cloud fault injection dataset demonstrate that PASE significantly outperforms state-of-the-art methods, reducing average system recovery time by over 40% and improving fault detection accuracy in unknown fault scenarios. Our framework advances autonomous system management by unifying LLM-based reasoning with model-assisted verification and meta-learned guidance.
| Comments: | 13 pages |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.01595 [cs.AI] |
| (or arXiv:2607.01595v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01595 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Junyan Tan [view email]
[v1]
Thu, 2 Jul 2026 01:45:30 UTC (11,561 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.