Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #729

RBS-Attention: radius-adaptive dual-branch sparse prefill for long-context LLMs

arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention. On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.

KEY POINTS

  1. arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention.
  2. On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.
  3. Faster, training-free sparse prefill that preserves relevant tokens can materially reduce the expensive dense prefill step in long-context LLM inference, enabling longer contexts or lower latency.

WHY IT MATTERS

Faster, training-free sparse prefill that preserves relevant tokens can materially reduce the expensive dense prefill step in long-context LLM inference, enabling longer contexts or lower latency.

SOURCES & TIMELINE

1