RESEARCH · RESEARCH · #1333
Decode-Latency Feedback Prefill (DLFP): a model-free controller for prefill chunking
The paper (arXiv:2609.38386v1) introduces Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that adaptively resizes prefill chunks overlapping active decodes. On Qwen3-0.6B (BF16) running on a single NVIDIA A100 80 GB, three paired 100-request trials reported mean reductions in P99 inter-token latency of 27.7% (paired 95% CI 21.0%–34.3%) with exact output agreement and unchanged SLO compliance, at the cost of a 34.8% mean increase in P99 time-to-first-token; DLFP did not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel setup, a failure traced to using an asynchronous scheduler-call interval as a proxy for GPU iteration completion.
KEY POINTS
- The paper (arXiv:2609.38386v1) introduces Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that adaptively resizes prefill chunks overlapping active decodes.
- On Qwen3-0.6B (BF16) running on a single NVIDIA A100 80 GB, three paired 100-request trials reported mean reductions in P99 inter-token latency of 27.7% (paired 95% CI 21.0%–34.3%) with exact output agreement and unchanged SLO compliance, at the cost of a 34.8% mean increase in P99 time-to-first-token; DLFP did not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel setup, a failure traced to using an asynchronous scheduler-call interval as a proxy for GPU iteration completion.
- This work offers a practical, reproducible controller that meaningfully lowers tail inter-token latency in concurrent autoregressive inference for some single-GPU setups while explicitly characterizing where the approach fails, guiding serving-scheduler design.
WHY IT MATTERS
This work offers a practical, reproducible controller that meaningfully lowers tail inter-token latency in concurrent autoregressive inference for some single-GPU setups while explicitly characterizing where the approach fails, guiding serving-scheduler design.