Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1005

Estimating step-level advantages via trajectory graphs (GRAFT)

This arXiv v1 paper identifies a systematic bias in group-based RL (e.g., GRPO) when estimating advantages at the step level for multi-turn/agentic tasks, and proposes GRAFT: a Graph-based Faithful sTep-level credit-assignment framework that merges rollouts into a trajectory graph, recovers node state-values via Bellman iteration, and assigns step credit by node-value differences. The authors also introduce Graph GAE (an extension of GAE on the trajectory graph) and report consistent gains over GRPO and recent agentic RL baselines; code is said to be available at the linked GitHub repository.

KEY POINTS

  1. This arXiv v1 paper identifies a systematic bias in group-based RL (e.g., GRPO) when estimating advantages at the step level for multi-turn/agentic tasks, and proposes GRAFT: a Graph-based Faithful sTep-level credit-assignment framework that merges rollouts into a trajectory graph, recovers node state-values via Bellman iteration, and assigns step credit by node-value differences.
  2. The authors also introduce Graph GAE (an extension of GAE on the trajectory graph) and report consistent gains over GRPO and recent agentic RL baselines; code is said to be available at the linked GitHub repository.
  3. More faithful step-level advantage estimates improve credit assignment for training multi-turn/agentic LLMs, which can materially affect downstream policy learning and behavior composition.

WHY IT MATTERS

More faithful step-level advantage estimates improve credit assignment for training multi-turn/agentic LLMs, which can materially affect downstream policy learning and behavior composition.

SOURCES & TIMELINE

1