FluidPD adds in-place P/D elasticity to SGLang, raising SLO attainment by up to 94.6 percentage points on Azure traces without extra GPUs
Synopsis
The authors present FluidPD, a prefill-decode (P/D) disaggregated LLM serving system built on SGLang that offloads a bounded prefill suffix to decode workers with slack via FluidToken and switches running workers between prefill and decode roles in place via FluidRole without reloading models, improving overall SLO attainment over static SGLang by up to 94.6 percentage points on Azure-trace-derived staged and mixed workloads and by up to 88.3 percentage points over Dynamo autoscaling while using fewer GPUs.
Fig. 1 : Time-varying input/output ratios in Azure workloads.
arXivInterpretation
Static P/D partitions lose SLO attainment under both short bursts and sustained demand shifts, even when aggregate capacity is sufficient. The paper attributes SLO loss to P/D imbalance rather than insufficient capacity through controlled experiments: on Azure (staged), 4P4D reaches 86.93% on Azure-conversation but only 3.40% on Azure-code, while 6P2D reaches 99.40% on Azure-code but only 6.64% on Azure-conversation, whereas a per-window oracle that picks a different P/D split performs well. Comparison on Llama-3.1-8B, 8×A100-80GB, constant 40 RPS Poisson arrivals, TTFT/TPOT SLOs of 2000ms/100ms, across 4P4D, 5P3D, and 6P2D static configurations.
FluidToken relieves transient prefill pressure by stealing a bounded suffix of selected prefill requests to their paired decode workers without changing worker roles. The mechanism performs a partial-prefill handoff along the request's existing P/D execution flow, reuses SGLang's async KV-transfer path, and decides using predictive Prefill Pressure Index (PPI) and Decode Pressure Index (DPI) signals before SLO violations occur; the paper notes prior systems adapt capacity only at worker or serving-group granularity and do not exploit token-level stealing for short-lived phase pressure. The design specifies a stealing-aware TTFT constraint, a decode-side stealing budget, and a greedy solver; ablation shows FluidToken contributes the larger individual benefit on Azure (mixed), and full FluidPD performs best on both workloads.
FluidRole rebalances the running P/D worker split through in-place, bidirectional role switching that avoids model reload and engine restart. A role-switching state machine (Active→Drain→Switch→Active) preserves model placement and parallelism configuration, reuses existing model-parallel communication groups, and captures decode CUDA graphs only on the first switch into decode; the paper reports PD transitions of about 1 s and DP transitions of up to 12 s, without dropping admitted requests. On Azure (staged), an observed 4P4D→5P3D transition lets p95 TTFT recover below the SLO after the workload returns to Azure-conversation; ablation shows FluidRole contributes the larger individual benefit on the staged workload.
Under the same GPU allocation, in-place P/D elasticity sustains SLOs better than both static deployments and autoscaling. FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points on Azure (staged) and up to 74.3 percentage points on Azure (mixed); against Dynamo, both start from a 6-GPU 3P3D configuration, FluidPD stays within six GPUs while Dynamo may scale to eight, and at 20 RPS FluidPD reaches 97.6% versus 9.3% on mixed and 92.6% versus 5.2% on staged. Evaluation covers Llama-3.1-8B, Qwen3-14B, and Qwen3-30B-A3B, with baselines SGLang v0.5.4, vLLM v0.26.0, and Dynamo v1.3.0, all using round-robin routing; the paper also reports a slight TPOT attainment decrease for FluidPD at the highest offered load.
Perspective
The work targets LLM serving deployments that use P/D disaggregation and have sufficient aggregate capacity but a poorly matched phase split, covering both short bursts and sustained demand shifts; gains are largest at low to moderate offered loads, where usable headroom remains. The method is implemented on SGLang v0.5.4's disaggregated runtime, preserving its prefill scheduling, decode-side KV preallocation, asynchronous KV transfer, and CUDA graph execution, so the most direct fit is comparable engines on multi-GPU single-server setups with NVLink interconnect. For engineering teams that want to improve TTFT-dominated SLO violations without provisioning additional GPUs and are willing to calibrate latency predictors offline per model and hardware configuration, the mechanisms offer a reproducible starting point; when both phases show persistent pressure, the paper suggests turning to cluster-level admission control or autoscaling.
The predictors currently rely on offline profiling, and the paper lists online recalibration from observed execution times as a natural extension, so behavior under runtime variation and workload drift remains an open question. FluidRole's DP transition time depends on the longest remaining active decode sequence at switch time, reported as up to about 12 s, and how this cost changes under more extreme output-length distributions is worth watching. The paper reports a slight TPOT attainment decrease for FluidPD at the highest offered load because stolen prefill work can interfere with decode iterations, indicating room to tune the interaction between the stealing budget and the saturation point. In addition, evaluation uses two representative workloads constructed from Azure traces rather than the full 168-hour trace, and does not directly compare with TaiChi, TokenScale, or Libra because no public implementations were available at evaluation time; those comparisons remain for follow-up work.
