Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

NVIDIA proposes a tool-call-centric agent evaluation framework and uses a SWE-bench trace to show how end-to-end and step-level scoring divide the work

Synopsis

This NVIDIA article traces how agent evaluation moved from scoring a single function call to scoring whole-task completion, proposes a framework built on an execution environment, two scoring layers (step-level process and end-to-end outcome), and a fixed Benchmark to Trial to Task to Turn to Step hierarchy, and uses a real SWE-bench Verified trace for pytest-dev__pytest-5262 (E2E score 1, step-level 3/4, tool-call precision 3/4, argument accuracy 4/4) to show how each metric should be read, alongside a workflow of setting a public floor and building domain evals from your own tickets and APIs.

AI-generated editorial illustration: How to Evaluate AI Agents From Tool Calls to Task Completion

Interpretation

The article argues agent evaluation must move from scoring individual calls to scoring whole tasks, because leaderboards like BFCL that score function selection and argument accuracy cannot show whether underlying checks or updates were skipped. Relative to call-level evaluation, it positions tool calling as the connective tissue of evaluation and states that call accuracy is necessary but not sufficient. A framework argument illustrated by BFCL's scope and an issue_refund example, without controlled experimental data.

It proposes that full agentic evaluation needs an execution environment plus two scoring layers: step-level (process) scoring asking whether a call was valid, relevant, and useful given the state at that point, and end-to-end (outcome) scoring checking only the final state, such as whether a refund posted or a ticket routed correctly. It unifies process and outcome scores onto one object, the trace, an ordered log of a single attempt, and notes that production evals typically gate releases on end-to-end while keeping step-level tracing for debugging. Primarily conceptual definitions and descriptions of production practice, without cross-model statistical comparisons.

It offers an actionable metric set and hierarchy, Benchmark to Trial to Task to Turn to Step, with task success rate, consistency (success-rate range across 3-5 trials), tool-call precision, argument accuracy, steps per success, and cost per success, stressing that metrics must be read in pairs. It makes pairings explicit, such as success rate without consistency being a point estimate on a stochastic system, and tool-call precision without argument accuracy hiding slot-filling failures, and notes parallel tool calling cuts steps and latency but not call count. Presented as formulas and an explanatory table, a methodological recommendation without independent validation experiments.

It demonstrates scoring with a public trace from SWE-bench Verified for pytest-dev__pytest-5262: the end-to-end check passed with an E2E score of 1, step-level 3/4, tool-call precision 3/4, and argument accuracy 4/4, where step 2 viewing the whole file was judged redundant and step 3's grep was judged a recovery. It grounds abstract metrics in one readable real trajectory, showing how outcome and process scores give different information about the same trace. A single public trace example using the OpenHands harness with parallel tool calling off and real filesystem plus git state persisting across turns, illustrative rather than statistical evidence.

Perspective

The article addresses engineering and evaluation teams deciding whether to ship an agent, and applies to settings with executable verification such as database state, passing tests, or a closed ticket; it offers a way to organize evaluation and a metric vocabulary rather than a verdict on any model. It also suggests public scores serve as a signal and a floor, while the real release gate should be a domain eval built from your own tickets, traces, and APIs, judged on environment state rather than an opinion about the final message.

The article notes contamination now extends beyond training-data leaks to live variants, such as web-searching agents retrieving answer keys during evaluation and datasets on Hugging Face being quickly re-scraped into pretraining corpora, and says private domain evals avoid this by being unable to scrape, but it gives no quantified impact. Its claim that Nemotron 3.5 Lightning hits 86% accuracy on PinchBench while finishing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy comes from the vendor's release, so readers should still watch for reproducible configs and independent verification. It also notes that which metric axis varies most across models is benchmark-dependent, with steps-per-turn varying on suites like Terminal-Bench 2.0, so no single axis is universally the most important.

Sources