Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #728

Attention-Aware Routing (AAR) for MoE models (arXiv:2609.20974v1)

The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters. On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.

KEY POINTS

  1. The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters.
  2. On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.
  3. AAR both raises MoE reasoning performance and reveals a coupled routing–attention circuit plus layer-dependent trade-offs, informing router design and analysis in large models.

WHY IT MATTERS

AAR both raises MoE reasoning performance and reveals a coupled routing–attention circuit plus layer-dependent trade-offs, informing router design and analysis in large models.

SOURCES & TIMELINE

1