Skip to main content
Back to timeline
arXivSource publication:

Researchers formalize an 'intent-execution' gap and build SSA, reproducing or improving pass@1 on several agentic benchmarks and analyzing 154k trajectories to reveal behavioral differences across models

Related research and updates

Synopsis

The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.

Source-provided article image: Dissecting model behavior through agent trajectories
Figure 1 ·

Figure 1: Reasoning nudges in SSA : Each family is nudged toward tool use in the form, quantitative or qualitative, it responds to.

arXiv

Interpretation

The paper frames AI agent performance as a systems problem rather than purely a modeling problem, and formalizes the 'intent-execution' gap: the mismatch between what the model intends and what the harness executes, and vice versa, which can prevent a model's full capabilities from translating into agent performance. Prior discussion of agent capability tended to focus on the model itself or on harness design elements such as tools and execution loops; this work isolates the consistency between model assumptions and harness behavior as a separately nameable design dimension and argues that minimizing this gap matters as much as tools and execution loops. This is a conceptual formalization and argument; the paper uses the gap as the motivation for designing the SSA harness to illustrate its impact, making it a framing claim rather than a controlled-experiment result.

The authors develop a simple and customizable harness, Simple Strands Agent (SSA), aimed at capturing the bulk of common patterns that generalize across model families (Claude, Gemini, GPT, Grok, Qwen) plus a small number of model-specific preferences. SSA is not tuned to a single model; its design centers on cross-family common patterns with model-specific preferences as a supplement, serving to test how harness-model alignment affects performance. The paper supports this design orientation with SSA's performance on multiple benchmarks; implementation details and ablations are not expanded at the abstract level.

SSA reproduces or improves on the pass@1 performance reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE. This indicates that, with the same models, harness-level alignment can change benchmark scores rather than merely reflecting model capability alone. The evidence is pass@1 comparison on four public agentic benchmarks, reported as reproduction or improvement; the abstract gives no specific values, sample sizes, or statistical tests.

Building on 154k trajectories generated by SSA, the authors represent agent trajectories in code state-spaces, compute solution distance and backtracking, and introduce finer-grained metrics such as edit frequency, testing activity, and phase-transitions, observing behavioral differences across frontier models whose pass@1 is relatively even. The analysis shifts evaluation from a single outcome metric (pass@1) toward process metrics, making each model's effort allocation across problem-solving stages observable and comparable. The evidence comes from large-scale process data of 154k trajectories; metric definitions and computation are only outlined in the abstract, and specific statistical results are not listed there.

Perspective

The work targets agent harness design and evaluation settings: when researchers or engineers use software-engineering and terminal-style agentic benchmarks such as SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE and want comparable performance across model families (Claude, Gemini, GPT, Grok, Qwen), SSA's approach of common patterns plus model-specific preferences can be borrowed directly. The process metrics (solution distance, backtracking, edit frequency, testing activity, phase-transitions) provide operable observation dimensions for analyzing model problem-solving behavior, suited to evaluation and debugging work that needs to go beyond pass@1 and understand how models allocate effort across solving stages.

What remains unexpanded at the abstract level includes: SSA's concrete implementation and configuration, the basis for dividing cross-family common patterns from model-specific preferences, the specific pass@1 values and comparison conditions on the four benchmarks, the collection setup for the 154k trajectories, and the precise definitions and statistical results of metrics such as solution distance, backtracking, edit frequency, testing activity, and phase-transitions. These are the questions readers will keep watching when assessing reproducibility and scope of applicability, and they also indicate that the relationship between process metrics and pass@1 still needs to be confirmed in the paper's details.

Sources