Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #604

Compositional reasoning in LMs shows decomposed-to-composed asymmetry under RL post-training

This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training. Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.

KEY POINTS

  1. This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training.
  2. Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.
  3. The result highlights a fundamental limitation in how RL post-training transfers compositional skills, affecting strategies for training LMs to solve novel, composite tasks.

WHY IT MATTERS

The result highlights a fundamental limitation in how RL post-training transfers compositional skills, affecting strategies for training LMs to solve novel, composite tasks.

SOURCES & TIMELINE

1