MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
Quick Answer
MultiUAV-Plat introduces a lightweight platform for multi-UAV collaborative task planning, featuring 75 mission sessions and 1500 tasks.
Quick Take
The Agent4Drone framework outperforms a ReAct baseline with a 57.9% task pass rate, significantly enhancing -driven UAV autonomy under realistic constraints.
Key Points
- MultiUAV-Plat features RESTful APIs and 2D/3D visualization for realistic task interaction.
- The benchmark includes 9396 validation checks across various UAV scenarios.
- Agent4Drone reduces failed task rates from 32.4% to 12.9% compared to the baseline.
- The platform addresses gaps in existing UAV simulators and LLM-agent benchmarks.
- Agent4Drone achieves a 74.6% average task check pass rate in evaluations.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration remains difficult to evaluate systematically. Existing UAV simulators mainly emphasize dynamics, perception, or low-level control, while existing LLM-agent benchmarks rarely capture aerial-robotics constraints such as partial observability, spatial coverage, UAV assignment, and multi-vehicle coordination. To bridge this gap, we present MultiUAV-Plat, a lightweight, easy-to-use, LLM-agent-oriented simulation platform for multi-UAV collaborative task planning. The platform exposes concise RESTful APIs, agent-facing observations, role-based information access, hidden validation logic, and optional 2D/3D visualization, allowing agents to solve missions through realistic tool interaction rather than privileged simulator access. Built on this platform, the MultiUAV-Plat Benchmark contains 75 mission sessions, 1500 natural-language tasks, and 9396 validation checks across target assignment, area search, and area assignment and patrol scenarios. We further propose Agent4Drone, a task-specific LLM agent framework that structures multi-UAV behavior into memory, observation, task understanding, planning, execution, and verification. In a full paired benchmark comparison, Agent4Drone achieves a 57.9% task pass rate, a 74.6% average task check pass rate, and a 72.0% global check pass rate, substantially outperforming a ReAct baseline at 30.6%, 47.9%, and 43.1%, respectively. Agent4Drone also reduces the total failed task rate from 32.4% to 12.9%. These results demonstrate that MultiUAV-Plat and MultiUAV-Plat Benchmark provide a reproducible foundation for studying LLM-driven multi-UAV autonomy under realistic information and execution constraints.
| Subjects: | Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Robotics (cs.RO) |
| Cite as: | arXiv:2606.31073 [cs.AI] |
| (or arXiv:2606.31073v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31073 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sheng Zhang [view email]
[v1]
Tue, 30 Jun 2026 03:02:12 UTC (1,237 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.