Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #393

reinforcement learning (RL)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

Compositional reasoning in LMs shows decomposed-to-composed asymmetry under RL post-training

This arXiv paper proposes a dependency-graph framework defining three compositionality levels and empirically studies compositional generalization of language models after reinforcement-learning post-training. Using deterministic data-structure tasks, the authors find a consistent decomposed-to-composed asymmetry—training on decomposed skills does not reliably transfer to composed tasks, while composed-task training transfers back more readily—provide a theoretical account, test length and structural shifts, and report a pilot on tool-calling benchmarks with preliminary evidence the asymmetry can appear in practical settings.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

This arXiv preprint empirically compares planning methods and reinforcement learning (RL) for multi-asset bearing maintenance using run-to-failure data. The authors find planning enforces reliability as a hard constraint and yields zero-failure policies whose total cost is largely insensitive to failure-penalty magnitude, while RL optimizes expected cost and often trades preventive maintenance for occasional failures (yielding lower costs under low-penalty regimes but persistent non-zero failures even when penalties are high); they also test lightweight constraint mechanisms (reward shaping, action masking) and provide a unified benchmark protocol for comparing decision approaches.

6.0