Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
Synopsis
This work analyzes large-scale execution trajectories from five agent benchmarks, distills six process signals from twelve observable statistics that are systematically associated with final agent performance, and proposes DualViewEval, which jointly models outcome and process relations to learn an exact-size miniset and predict full-benchmark scores, achieving the best results on all five benchmarks, reaching 24×–40× compression on APEX-Agents and BFCL with only 20 tasks, reducing MAE by 14.5%–28.2% over the strongest competitors, and improving Kendall's τ by up to 7.2% relative to EssenceBench on SWE-bench Verified.
Figure 2 : Association between automatically extracted process measurements and agent performance. (a) Within-benchmark Spearman correlations for the selected six measurements (purple) and six reference statistics. (b) Mean correlation across benchmarks. (c) Distributions of the selected measurements for low, middle, and high scoring BFCL models.
arXivInterpretation
The paper derives six benchmark-agnostic process measurements from large-scale agent trajectories: agent steps, tool failed rate, tool-category entropy, validation-tool rate, required-write execution, and read–write–validate closure. Prior compression methods mainly model redundancy in task–model final-score distributions, whereas this work uses cross-benchmark correlation analysis to select six automatically extractable, semantically non-redundant process signals from twelve candidates, bringing execution behavior into the representation. Based on analysis of large-scale execution traces from five agent benchmarks with open trajectories, reporting systematic associations between these measurements and full-benchmark scores, plus distinct process distributions for low-, middle-, and high-scoring BFCL models.
It proposes DualViewEval, an end-to-end dual-view framework that builds an outcome relation matrix and a process relation matrix per task, fuses them into an agent kernel feeding a Kernel Ridge score predictor, and uses a straight-through hard Top-K gate to keep the miniset exactly sized. Unlike methods such as SparseEval that mainly model relations between columns of the outcome matrix, this framework exploits both agent-level shared structure and the execution process, coupling task selection with the downstream prediction objective in a single feedback loop. The method is reported with means and standard deviations over ten splits across five benchmarks and three miniset sizes, accompanied by outcome/process relation matrix ablations, selector ablations, prediction-architecture ablations, and a budget sweep.
Across five agent benchmarks, DualViewEval achieves the best results on all datasets; with only 20 tasks it reaches 24×–40× compression on APEX-Agents and BFCL, reducing MAE by 14.5%–28.2% over the strongest competitors. Relative to SparseEval, the most directly comparable outcome-driven predictive baseline, it reduces MAE by 30.5%–45.1% averaged across benchmarks and improves mean Kendall's τ by 0.089–0.115; on APEX-Agents at the tightest budget it reduces MAE by 34.3% and raises τ by 0.100. Ten-split mean±standard-deviation results covering BFCL, τ²-Bench, Terminal-Bench 2, SWE-bench Verified, and APEX-Agents, totaling 1,983 tasks, 262 agent configurations, and 132,387 outcome pairs.
The selected minisets reveal capability differences among agents, the process weight γ adapts per benchmark, and training cost is lower than SparseEval. Beyond score prediction, the compressed subsets preserve diagnostic behavioral evidence; the learned γ distributes differently across benchmarks, and a fixed-γ sweep likewise shows intermediate weights generally give the best error–ranking trade-off. Reports aggregate process profile tables for APEX-Agents and SWE-bench Verified, γ distribution figures, a training-data proportion experiment, and roughly 4.2×–8.3× training-time speedups.
Perspective
The result targets evaluation settings that need frequent agent assessment and can access execution trajectories: the method relies on automatically parseable trajectories and tool-call records, and applies to benchmarks such as BFCL, τ²-Bench, Terminal-Bench 2, SWE-bench Verified, and APEX-Agents where outcomes and trajectories can be aligned. It lets researchers and engineering teams estimate full-benchmark scores and rankings under a fixed task budget, and use the selected subsets to observe behavioral differences in tool use, validation, and write-operation closure; for evaluation settings without accessible trajectories or tool-call structure, the applicability of the process view needs separate validation.
Readers may still watch: the extraction stability of process measurements when trajectory formats differ substantially or tool-call parsing is incomplete; the learned process weight γ varies across benchmarks, indicating the optimal outcome–process balance is benchmark-dependent and may need recalibration on new benchmarks; the held-out model-family experiment is conducted only on the Qwen family in APEX-Agents, so generalization to other model families and benchmarks remains an open question; and although this is a full-text read, figures are presented as text, so some graphical details (such as exact values in the γ distribution and efficiency analysis) cannot be fully verified from the text.
