RLDS: decomposing trajectory rewards by subtask improves RL for language-model agents (arXiv:2609.27035v1)
The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS) and Subtask-Decomposed Advantage Estimation (SDAE), which split trajectory reward into per-subtask shares before policy updates instead of collapsing outcomes to a single scalar as in Group Relative Policy Optimization (GRPO). Evaluated on four benchmarks, RLDS produced substantial gains where subtask heterogeneity is high (ScienceWorld +11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]; FrozenLake +9.8 points, CI [+7.0, +12.8]), showed little effect where diagnostics predicted little recovery (HotpotQA, DeepResearch), and was more compute-efficient on ScienceWorld (-10.9% wall-clock per step).