Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #939

RLDS: decomposing trajectory rewards by subtask improves RL for language-model agents (arXiv:2609.27035v1)

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS) and Subtask-Decomposed Advantage Estimation (SDAE), which split trajectory reward into per-subtask shares before policy updates instead of collapsing outcomes to a single scalar as in Group Relative Policy Optimization (GRPO). Evaluated on four benchmarks, RLDS produced substantial gains where subtask heterogeneity is high (ScienceWorld +11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]; FrozenLake +9.8 points, CI [+7.0, +12.8]), showed little effect where diagnostics predicted little recovery (HotpotQA, DeepResearch), and was more compute-efficient on ScienceWorld (-10.9% wall-clock per step).

KEY POINTS

  1. The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS) and Subtask-Decomposed Advantage Estimation (SDAE), which split trajectory reward into per-subtask shares before policy updates instead of collapsing outcomes to a single scalar as in Group Relative Policy Optimization (GRPO).
  2. Evaluated on four benchmarks, RLDS produced substantial gains where subtask heterogeneity is high (ScienceWorld +11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]; FrozenLake +9.8 points, CI [+7.0, +12.8]), showed little effect where diagnostics predicted little recovery (HotpotQA, DeepResearch), and was more compute-efficient on ScienceWorld (-10.9% wall-clock per step).
  3. Decomposing rewards by subtask directly targets credit-assignment in long, heterogeneous rollouts for language-model agents, yielding measurable performance and compute gains on complex tasks.

WHY IT MATTERS

Decomposing rewards by subtask directly targets credit-assignment in long, heterogeneous rollouts for language-model agents, yielding measurable performance and compute gains on complex tasks.

SOURCES & TIMELINE

1