RESEARCH · RESEARCH · #1005
Estimating step-level advantages via trajectory graphs (GRAFT)
This arXiv v1 paper identifies a systematic bias in group-based RL (e.g., GRPO) when estimating advantages at the step level for multi-turn/agentic tasks, and proposes GRAFT: a Graph-based Faithful sTep-level credit-assignment framework that merges rollouts into a trajectory graph, recovers node state-values via Bellman iteration, and assigns step credit by node-value differences. The authors also introduce Graph GAE (an extension of GAE on the trajectory graph) and report consistent gains over GRPO and recent agentic RL baselines; code is said to be available at the linked GitHub repository.
KEY POINTS
- This arXiv v1 paper identifies a systematic bias in group-based RL (e.g., GRPO) when estimating advantages at the step level for multi-turn/agentic tasks, and proposes GRAFT: a Graph-based Faithful sTep-level credit-assignment framework that merges rollouts into a trajectory graph, recovers node state-values via Bellman iteration, and assigns step credit by node-value differences.
- The authors also introduce Graph GAE (an extension of GAE on the trajectory graph) and report consistent gains over GRPO and recent agentic RL baselines; code is said to be available at the linked GitHub repository.
- More faithful step-level advantage estimates improve credit assignment for training multi-turn/agentic LLMs, which can materially affect downstream policy learning and behavior composition.
WHY IT MATTERS
More faithful step-level advantage estimates improve credit assignment for training multi-turn/agentic LLMs, which can materially affect downstream policy learning and behavior composition.