FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
Quick Answer
FluidPD is a novel P/D-disaggregated serving system that enhances SLO attainment by up to 94.6 percentage points over static SGLang.
Quick Take
It employs FluidToken and FluidRole to dynamically manage prefill and decode resource allocation, addressing both transient and sustained imbalances without requiring additional workers. This approach significantly improves service quality in serving environments.
Key Points
- FluidPD improves SLO attainment by up to 94.6 percentage points over static configurations.
- FluidToken offloads prefill tasks to decode workers during transient imbalances.
- FluidRole reassigns worker roles in place, avoiding model reloads and restarts.
- Lightweight pressure indices predict resource pressure before SLO violations occur.
- FluidPD enhances service quality without the need for additional worker provisioning.
DeepSignal Analysis
What happened
FluidPD is a new serving system designed for large language models (LLMs) that enhances service level objective (SLO) attainment by up to 94.6 percentage points compared to static configurations. It addresses the challenges of fluctuating prefill and decode demands without needing additional workers. The system employs two mechanisms, FluidToken and FluidRole, to dynamically manage resource allocation.
Key evidence
- FluidPD improves SLO attainment by up to 94.6 percentage points over static SGLang, indicating a significant enhancement in performance.
- FluidToken offloads prefill computation to decode workers when there is available slack, addressing transient imbalances in resource allocation.
- FluidRole allows for the reassignment of workers between prefill and decode roles without requiring model reloads or engine restarts, thus improving efficiency.
Why it matters
The introduction of FluidPD represents a significant advancement in LLM serving architectures, particularly in managing resource allocation dynamically. By addressing both transient and sustained imbalances, it enhances service quality and operational efficiency. This could lead to better user experiences and more reliable performance in production environments, especially as demand for LLM applications continues to grow.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance.
We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.
| Comments: | 13 pages, 11 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06917 [cs.AI] |
| (or arXiv:2610.06917v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06917 arXiv-issued DOI via DataCite |
Submission history
From: Kaidi Fu [view email]
[v1]
Fri, 2 Oct 2026 18:49:43 UTC (1,043 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.