Skip to main content
Back to timeline
arXivSource publication:

Rethinking Agentic Evaluation: Treating Agents as Configurable Systems, About 54% of Outcome Variance Comes from Repeating the Same Configuration

Synopsis

The work introduces a new benchmark of four scientific tasks in which a coding agent must find and correctly operate a published specialist model, and systematically studies five parts of the agent's configuration—task information, reasoning, self-verification, time budget, and backbone model; it finds that approximately 54% of outcome variance comes from repeating the same configuration, that the information provided to the agent has the largest effect exceeding both time budget and model size while also reducing cost and improving calibration, and that prompting an agent to verify its answer has little effect on its verification behavior whereas providing a dedicated verification tool changes that behavior substantially.

Source-provided article image: Agents Are Systems, Not Models: Rethinking Agentic Evaluation
Figure 1 ·

Figure 1: Architecture of our agentic AI4S benchmark. The agent (center) has access to different specialist models and is evaluated against their published performance (right). By varying the axes of the agent’s configuration within the same task (left), we can systematically study its performance, reliability, and behavior under controlled, repeatable conditions (bottom).

arXiv

Interpretation

Agent performance shows substantial run-to-run variability: approximately 54% of the outcome variance comes from repeating the same configuration rather than changing it. Prior evaluations typically treat the agent itself as fixed and report metrics such as success rate, cost, consistency, and robustness; this work quantifies configuration choices themselves as a source of variance. Based on a new benchmark of four scientific tasks, a systematic comparison across five configuration dimensions, and the release of more than 18,000 agent trajectories.

Among configuration choices, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. It places what the agent is told, how long it runs, and which model it uses within a single comparison framework and gives a relative ordering of importance. The abstract reports cross-configuration comparisons and notes the joint relationship between information, cost, and calibration.

Configuration choices interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. It shows that time budget is not an independent gain; its benefit depends on information and model capability. The abstract reports this finding in the interaction form 'additional time helps only when…'.

A trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. It distinguishes requesting a behavior from implementing that behavior in the system, indicating the latter is more effective. Derived from a trajectory-level taxonomy of agent behavior, accompanied by the release of the benchmark and more than 18,000 trajectories.

Perspective

The work targets scientific task settings in which a coding agent must find and correctly operate a published specialist model, and it is relevant to readers studying how agent configuration affects performance, cost, consistency, and calibration, including benchmark designers and agent system developers. Its conclusions pertain to the four studied tasks and five configuration dimensions, and the release of the benchmark and more than 18,000 trajectories supports later verification and extension under the same setting.

Based only on the abstract, the specific content of the four scientific tasks, the operationalization of the five configuration dimensions, the statistical details of the variance decomposition, and the coding procedure of the trajectory taxonomy are not available; these details affect how far the conclusions extrapolate to other tasks and configurations. In addition, the abstract does not state the specific values of backbone models and time budgets, so readers who wish to reproduce or transfer the results still need the benchmark and experimental setup in the original text.

Sources