RESEARCH · RESEARCH · #936
Math reasoning in LLMs organized by reasoning approach, not topic
This paper (arXiv:2609.27041v1) introduces a generation-replay protocol to extract activation-importance signatures from reasoning tokens and clusters those signatures across eight open math-capable LLMs and five math-reasoning sources. The authors report that clusters align with reusable reasoning approaches (77–82% approach-level coherence by human judges) rather than benchmark topics, and that controlling the requested reasoning approach shifts cluster assignments in 7 of 8 model conditions while paraphrases largely preserve them.
KEY POINTS
- This paper (arXiv:2609.27041v1) introduces a generation-replay protocol to extract activation-importance signatures from reasoning tokens and clusters those signatures across eight open math-capable LLMs and five math-reasoning sources.
- The authors report that clusters align with reusable reasoning approaches (77–82% approach-level coherence by human judges) rather than benchmark topics, and that controlling the requested reasoning approach shifts cluster assignments in 7 of 8 model conditions while paraphrases largely preserve them.
- If LLMs internally organize math reasoning by reusable approaches rather than task topics, evaluations and training that are balanced by topic can still miss important distributional imbalances over reasoning approaches.
WHY IT MATTERS
If LLMs internally organize math reasoning by reusable approaches rather than task topics, evaluations and training that are balanced by topic can still miss important distributional imbalances over reasoning approaches.