RESEARCH · RESEARCH · #729
RBS-Attention: radius-adaptive dual-branch sparse prefill for long-context LLMs
arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention. On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.
KEY POINTS
- arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention.
- On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.
- Faster, training-free sparse prefill that preserves relevant tokens can materially reduce the expensive dense prefill step in long-context LLM inference, enabling longer contexts or lower latency.
WHY IT MATTERS
Faster, training-free sparse prefill that preserves relevant tokens can materially reduce the expensive dense prefill step in long-context LLM inference, enabling longer contexts or lower latency.