RESEARCH · RESEARCH · #1655
FluidPD: SLO-aware in-place elasticity for prefill–decode disaggregated LLM serving
This paper introduces FluidPD, a prefill/decode (P/D) disaggregated serving system that implements SLO-aware in-place elasticity via two mechanisms: FluidToken, which offloads a bounded portion of prefill work to decode workers during transient imbalance, and FluidRole, which reassigns running workers between prefill and decode roles without model reloads. Guided by lightweight pressure indices, the authors report up to a 94.6 percentage-point improvement in overall SLO attainment over a static SGLang baseline on production Azure traces, achieved without provisioning additional workers.
KEY POINTS
- This paper introduces FluidPD, a prefill/decode (P/D) disaggregated serving system that implements SLO-aware in-place elasticity via two mechanisms: FluidToken, which offloads a bounded portion of prefill work to decode workers during transient imbalance, and FluidRole, which reassigns running workers between prefill and decode roles without model reloads.
- Guided by lightweight pressure indices, the authors report up to a 94.6 percentage-point improvement in overall SLO attainment over a static SGLang baseline on production Azure traces, achieved without provisioning additional workers.
- It demonstrates a practical way to reduce latency SLO violations and improve LLM serving quality by adapting role assignment and short-term work offload in-place, without requiring extra GPUs.
WHY IT MATTERS
It demonstrates a practical way to reduce latency SLO violations and improve LLM serving quality by adapting role assignment and short-term work offload in-place, without requiring extra GPUs.