FluidPD: SLO-aware in-place elasticity for prefill–decode disaggregated LLM serving
This paper introduces FluidPD, a prefill/decode (P/D) disaggregated serving system that implements SLO-aware in-place elasticity via two mechanisms: FluidToken, which offloads a bounded portion of prefill work to decode workers during transient imbalance, and FluidRole, which reassigns running workers between prefill and decode roles without model reloads. Guided by lightweight pressure indices, the authors report up to a 94.6 percentage-point improvement in overall SLO attainment over a static SGLang baseline on production Azure traces, achieved without provisioning additional workers.