Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control
Quick Answer
This paper introduces a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints while enabling effective coordination.
Quick Take
It achieves competitive performance with nearly perfect safety rates and theoretical guarantees, making it suitable for safety-critical applications.
Key Points
- Enforces hard safety constraints using a constraint manifold at low levels.
- Achieves competitive performance with nearly perfect safety rates.
- Provides theoretical safety guarantees in settings.
- Enables stable and efficient training with stationary learning dynamics.
- Generalizes effectively to varying numbers of agents and obstacles.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while control-theoretic methods enforce safety but often lead to overly conservative and inefficient behaviors. We propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at low level via a constraint manifold, while enabling effective coordination through high-level policy learning. Our approach provides theoretical safety guarantees in the multi-agent setting and yields stationary learning dynamics, thereby enabling stable and efficient training. Empirically, our method achieves competitive performance while maintaining nearly perfect safety rates, and generalizes effectively to varying numbers of agents and obstacles.
| Comments: | 10 pages |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.24010 [cs.AI] |
| (or arXiv:2606.24010v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24010 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zihao Guo [view email]
[v1]
Mon, 22 Jun 2026 23:32:23 UTC (4,947 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.