Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Researchers formalize an 'intent-execution' gap and build SSA, reproducing or improving pass@1 on several agentic benchmarks and analyzing 154k trajectories to reveal behavioral differences across models

The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.