Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1002

SLCA-GRPO: Segment-Locked Credit Assignment for tool-calling RL (arXiv:2609.29050v1)

This paper introduces SLCA-GRPO, which uses Segment-Locked Credit Assignment (SLCA) and Hierarchical Rewards to prevent cross-segment credit misattribution in tool-calling reinforcement learning; it also presents the Schema-Guided LLM Simulator (SGLS) for scalable training. On a 7B backbone, SLCA-GRPO speeds convergence and outperforms GRPO, ToolPO, and RLTR by +2.53 pp in-domain, +1.36 pp on the BFCL, and +9.15 pp on τ^2-Bench under the same budgets (arXiv:2609.29050v1).

KEY POINTS

  1. This paper introduces SLCA-GRPO, which uses Segment-Locked Credit Assignment (SLCA) and Hierarchical Rewards to prevent cross-segment credit misattribution in tool-calling reinforcement learning; it also presents the Schema-Guided LLM Simulator (SGLS) for scalable training.
  2. On a 7B backbone, SLCA-GRPO speeds convergence and outperforms GRPO, ToolPO, and RLTR by +2.53 pp in-domain, +1.36 pp on the BFCL, and +9.15 pp on τ^2-Bench under the same budgets (arXiv:2609.29050v1).
  3. Fixing advantage contamination between tool-invocation and summary tokens reduces gradient noise and stabilizes optimization, improving training efficiency and accuracy for tool-augmented LLM policies.

WHY IT MATTERS

Fixing advantage contamination between tool-invocation and summary tokens reduces gradient noise and stabilizes optimization, improving training efficiency and accuracy for tool-augmented LLM policies.

SOURCES & TIMELINE

1