RESEARCH · RESEARCH · #769
Agent evaluation shifts to executable end-to-end and step-level scoring; NVIDIA Nemotron 3.5 Lightning demonstrates the approach
Agent evaluation is moving from scoring individual function calls to running agents inside executable environments that track state across multi-step tool use, combining step-level (process) scoring to locate failures with end-to-end (outcome) scoring to verify final task completion; metrics roll up through a fixed Benchmark→Trial→Task→Turn→Step hierarchy and favor executable checks over reference- or LLM-judge methods. NVIDIA presents this framework and reports Nemotron 3.5 Lightning achieving 86% accuracy on PinchBench while completing tasks ~30% faster than comparable models, and recommends domain-specific, state-gated production evaluations; reproducibility docs and a demo on build.nvidia.com are provided.
KEY POINTS
- Agent evaluation is moving from scoring individual function calls to running agents inside executable environments that track state across multi-step tool use, combining step-level (process) scoring to locate failures with end-to-end (outcome) scoring to verify final task completion; metrics roll up through a fixed Benchmark→Trial→Task→Turn→Step hierarchy and favor executable checks over reference- or LLM-judge methods.
- NVIDIA presents this framework and reports Nemotron 3.5 Lightning achieving 86% accuracy on PinchBench while completing tasks ~30% faster than comparable models, and recommends domain-specific, state-gated production evaluations; reproducibility docs and a demo on build.nvidia.com are provided.
- Measuring full task completion in executable environments better predicts production reliability and surfaces where multi-step tool chains fail, which single-call scoring misses.
WHY IT MATTERS
Measuring full task completion in executable environments better predicts production reliability and surfaces where multi-step tool chains fail, which single-call scoring misses.