Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #7779

trajectory reward decomposition

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

RLDS: decomposing trajectory rewards by subtask improves RL for language-model agents (arXiv:2609.27035v1)

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS) and Subtask-Decomposed Advantage Estimation (SDAE), which split trajectory reward into per-subtask shares before policy updates instead of collapsing outcomes to a single scalar as in Group Relative Policy Optimization (GRPO). Evaluated on four benchmarks, RLDS produced substantial gains where subtask heterogeneity is high (ScienceWorld +11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]; FrozenLake +9.8 points, CI [+7.0, +12.8]), showed little effect where diagnostics predicted little recovery (HotpotQA, DeepResearch), and was more compute-efficient on ScienceWorld (-10.9% wall-clock per step).

7.0