Public articles linked to the same research event.
arXiv The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.
The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.
The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.
The work formalizes AI agent performance as an 'intent-execution' gap (the mismatch between what a model intends and what the harness executes), introduces a customizable harness called Simple Strands Agent (SSA) to align different model families, reproduces or improves the pass@1 reported by diverse model-provider families on SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE, and, based on 154k trajectories generated by SSA, represents agent trajectories in code state-spaces to compute solution distance and backtracking, using finer-grained metrics such as edit frequency, testing activity, and phase-transitions to reveal differences in problem-solving behavior and effort allocation across frontier models whose pass@1 numbers are relatively even.