Compositional reasoning in LMs shows decomposed-to-composed asymmetry under RL post-training
This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training. Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.