Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #5703

transformer

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Attention-Aware Routing (AAR) for MoE models (arXiv:2609.20974v1)

The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters. On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.

7.0