RESEARCH · RESEARCH · #728
Attention-Aware Routing (AAR) for MoE models (arXiv:2609.20974v1)
The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters. On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.
KEY POINTS
- The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters.
- On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.
- AAR both raises MoE reasoning performance and reveals a coupled routing–attention circuit plus layer-dependent trade-offs, informing router design and analysis in large models.
WHY IT MATTERS
AAR both raises MoE reasoning performance and reveals a coupled routing–attention circuit plus layer-dependent trade-offs, informing router design and analysis in large models.