Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #5710

Qwen3-32B

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

Decode-Latency Feedback Prefill (DLFP): a model-free controller for prefill chunking

The paper (arXiv:2609.38386v1) introduces Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that adaptively resizes prefill chunks overlapping active decodes. On Qwen3-0.6B (BF16) running on a single NVIDIA A100 80 GB, three paired 100-request trials reported mean reductions in P99 inter-token latency of 27.7% (paired 95% CI 21.0%–34.3%) with exact output agreement and unchanged SLO compliance, at the cost of a 34.8% mean increase in P99 time-to-first-token; DLFP did not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel setup, a failure traced to using an asynchronous scheduler-call interval as a proxy for GPU iteration completion.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

RBS-Attention: radius-adaptive dual-branch sparse prefill for long-context LLMs

arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention. On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.

7.0