RESEARCH · RESEARCH · #604
Compositional reasoning in LMs shows decomposed-to-composed asymmetry under RL post-training
This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training. Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.
KEY POINTS
- This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training.
- Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.
- The result highlights a fundamental limitation in how RL post-training transfers compositional skills, affecting strategies for training LMs to solve novel, composite tasks.
WHY IT MATTERS
The result highlights a fundamental limitation in how RL post-training transfers compositional skills, affecting strategies for training LMs to solve novel, composite tasks.