Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #6112

Berkeley Function-Calling Leaderboard (BFCL)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

SLCA-GRPO: Segment-Locked Credit Assignment for tool-calling RL (arXiv:2609.29050v1)

This paper introduces SLCA-GRPO, which uses Segment-Locked Credit Assignment (SLCA) and Hierarchical Rewards to prevent cross-segment credit misattribution in tool-calling reinforcement learning; it also presents the Schema-Guided LLM Simulator (SGLS) for scalable training. On a 7B backbone, SLCA-GRPO speeds convergence and outperforms GRPO, ToolPO, and RLTR by +2.53 pp in-domain, +1.36 pp on the BFCL, and +9.15 pp on τ^2-Bench under the same budgets (arXiv:2609.29050v1).

7.0

RESEARCH · 1 SOURCE · NVIDIA Developer

Agent evaluation shifts to executable end-to-end and step-level scoring; NVIDIA Nemotron 3.5 Lightning demonstrates the approach

Agent evaluation is moving from scoring individual function calls to running agents inside executable environments that track state across multi-step tool use, combining step-level (process) scoring to locate failures with end-to-end (outcome) scoring to verify final task completion; metrics roll up through a fixed Benchmark→Trial→Task→Turn→Step hierarchy and favor executable checks over reference- or LLM-judge methods. NVIDIA presents this framework and reports Nemotron 3.5 Lightning achieving 86% accuracy on PinchBench while completing tasks ~30% faster than comparable models, and recommends domain-specific, state-gated production evaluations; reproducibility docs and a demo on build.nvidia.com are provided.

7.0