SLCA-GRPO: Segment-Locked Credit Assignment for tool-calling RL (arXiv:2609.29050v1)
This paper introduces SLCA-GRPO, which uses Segment-Locked Credit Assignment (SLCA) and Hierarchical Rewards to prevent cross-segment credit misattribution in tool-calling reinforcement learning; it also presents the Schema-Guided LLM Simulator (SGLS) for scalable training. On a 7B backbone, SLCA-GRPO speeds convergence and outperforms GRPO, ToolPO, and RLTR by +2.53 pp in-domain, +1.36 pp on the BFCL, and +9.15 pp on τ^2-Bench under the same budgets (arXiv:2609.29050v1).