PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
Quick Answer
PlanFlip introduces four prompt injection attacks targeting the planning phase of multi-agent LLM systems, revealing that GPT-5 has the highest attack success rate (ASR = 0.68).
Quick Take
The study shows that homogeneous model pipelines are vulnerable, while reasoning-augmented models like DeepSeek-R1 resist these injections. Key recommendations include enhancing model diversity to improve security against planning-phase attacks.
Key Points
- PlanFlip framework includes GoalSubstitution, PriorityInversion, ContextPollution, and RoleConfusion attacks.
- GPT-4o and Llama-3.3-70B show ASR near 0 but maintain high stealth.
- DeepSeek-R1 achieves StepShift = 0.00, indicating resistance to all attacks.
- Heterogeneous model diversity is crucial for security in .
- GoalAnchorCheck and CrossAgentConsensus detect attacks with rates up to 1.00.
DeepSignal Analysis
What happened
The study introduces PlanFlip, which identifies the planning phase of multi-agent LLM systems as a critical vulnerability. It presents four prompt injection attacks that exploit this phase, with GPT-5 showing the highest attack success rate at 0.68. The research highlights the weaknesses of homogeneous model pipelines and the resilience of reasoning-augmented models.
Key evidence
- PlanFlip introduces four prompt injection attacks targeting the planning phase of multi-agent LLM systems.
- GPT-5 achieves the highest attack success rate (ASR = 0.68), challenging the notion that stronger models are more secure.
- DeepSeek-R1, a reasoning-augmented model, shows resistance to all attacks, achieving StepShift = 0.00.
Why it matters
This research underscores the importance of model diversity in enhancing security against planning-phase attacks in multi-agent systems. The findings suggest that relying on homogeneous models may create blind spots, making systems more susceptible to vulnerabilities. As AI systems become more complex, understanding these weaknesses is crucial for developing robust security measures.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit. We identify the planning phase as a critical attack surface: a single injection into the Planner's context achieves cascade amplification, corrupting all downstream sub-tasks simultaneously. We introduce PlanFlip, a framework comprising four planning-phase prompt injection attacks -- GoalSubstitution (PF-1), PriorityInversion (PF-2), ContextPollution (PF-3), and RoleConfusion (PF-4) -- each disguised as plausible tool outputs to evade keyword filters. Evaluating nine frontier LLMs across 3,479 episodes, we uncover three findings: (1) capability amplifies vulnerability -- GPT-5 achieves the highest attack success rate (ASR = 0.68), contradicting the assumption that stronger models are inherently more secure; (2) homogeneous pipelines exhibit a correlated-agent blind spot -- GPT-4o and Llama-3.3-70B show ASR near 0 yet Stealth = 1.00 and StepShift > 0, with attacks restructuring plans while the same-backbone Critic reports alignment (two independent judges confirm -0.20 to -0.32 semantic deviation, r = 0.943); (3) reasoning-augmented models resist injections -- DeepSeek-R1 achieves StepShift = 0.00 across all attacks. We propose GoalAnchorCheck (D1) and CrossAgentConsensus (D2), achieving detection rates up to 1.00 and outperforming same-backbone baselines in 15 of 16 cells. Our key insight: heterogeneous model diversity is a security prerequisite for multi-agent systems; redundancy within a homogeneous backbone provides no protection against planning-phase attacks.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16199 [cs.AI] |
| (or arXiv:2607.16199v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16199 arXiv-issued DOI via DataCite |
Submission history
From: Yuhang Wang [view email]
[v1]
Thu, 30 Apr 2026 07:22:06 UTC (202 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.