Skip to main content
Back to timeline
arXivSource publication:

Stateless Language Agents (SLA) achieve the best final result on software engineering, kernel optimization, and algorithm design at up to one-billion-token budgets, matching the strongest kernel baseline with 93% fewer tokens

Related research and updates

Synopsis

The work introduces Stateless Language Agents (SLA): the harness owns the research state and reconstructs a fresh, role-specific context for every invocation, with a stateless Advisor assigning concrete experiments to parallel Workers from harness-summarized evidence; evaluated against EvoX, CORAL, and SwarmResearch on software engineering, kernel optimization, and algorithm design at budgets up to one billion cumulative tokens, SLA achieves the best final result on every task, matches the strongest kernel baseline's final performance with 93.1% fewer tokens (84.4% under Claude Code), and ablations show focused contexts and explicit assignments each contribute progress that can compound, while the Advisor consumes 0.24-0.51% of tokens.

Source-provided article image: Stateless Language Agents: Scaling Long-Horizon Automated Research
Figure 1 ·

Figure 1: SLA sustains research progress over long horizons. (a) On Anthropic kernel optimization, SLA finds stronger solutions and reaches the strongest baseline’s mean final performance with 14.5 × \times fewer tokens (a 93.1% reduction); every SLA run finishes ahead of every baseline run. Thick curves show means over three runs per method; thin curves show individual runs. (b) On all four FrontierSWE tasks, SLA trails the strongest per-task baseline at 175M tokens (25% of the budget, open circles) but leads at 700M (full budget, filled circles). All results use Codex with GPT-5.5.

arXiv

Interpretation

SLA achieves the best full-budget result on every evaluated task and reaches strong baseline performance with fewer tokens in most configurations. Prior AutoResearch evaluations mostly budget by iterations or model calls and use benchmarks that saturate early, leaving untested whether a method keeps converting computation into progress over long horizons; this work compares at up to one billion cumulative tokens. On Anthropic kernel optimization all three SLA runs outperform all nine baseline runs (worst SLA 1122 cycles versus best baseline 1191 cycles); relative to the strongest baseline per task, kernel cycles fall by 8.1-12.8%, FrontierSWE scores rise by 4.95 points on average, SOL-ExecBench scores by 5.5% on average across three tasks, and Structured-LWE by 1.5 points; on Anthropic kernel with Codex SLA needs 67.9M tokens to match the target versus 986.3M for SwarmResearch.

Focused contexts and explicit assignments each contribute to progress, and their effects can compound over a run. Ablations resume from identical checkpoints and change only what enters each context or whether Workers receive assignments, making context management and experiment selection separable design variables. From shared checkpoints, removing Advisor context reconstruction, Worker context isolation, or Advisor assignments reduces mean progress in every comparison; removing Worker isolation costs 5-11 cycles over 100M-token continuations but 224 cycles over a full 1B-token run (1336 versus 1112.0 cycles); without assignments, parallel Workers repeatedly implement the same feature.

Explicit coordination is cheap, and separating coordination from implementation lets the Worker model be chosen per task. Unlike SwarmResearch's long-running Shepherd, SLA's Advisor starts each epoch in a new session, so coordination cost and implementation cost can be accounted separately. The Advisor consumes 0.24-0.51% of tokens and 1.20-2.29% of model cost, versus 8.39-10.70% of tokens and 6.14-7.35% of cost for SwarmResearch's Shepherd; at a matched cost of $100, GPT-5.4 mini Workers score 20.88 on GitZig versus 19.68 for GPT-5.5 Workers, while GPT-5.5 Workers lead on Anthropic kernel optimization.

The evaluation horizon changes method rankings and component conclusions. The work reports results at multiple cumulative token budgets rather than a single short budget or an early-saturating benchmark. On FrontierSWE, SLA trails the strongest baseline on all four tasks at 25% of the budget but leads on all four at the full budget; the cost of removing Worker isolation grows by more than an order of magnitude between a 100M-token continuation and a full run; wider Worker pools trail by 337 cycles at 25% of the budget yet finish within 4 cycles of the narrowest at the full budget.

Perspective

The results target long-horizon search problems with executable evaluators and measurable objectives but no known optimum, demonstrated on software engineering, kernel optimization, and algorithm design under two coding-agent configurations, OpenAI Codex and Claude Code. They let researchers keep research state in the harness where it can be inspected, checkpointed, and rolled back, and run controlled ablations from identical checkpoints; they also argue for reporting results at multiple cumulative token budgets. For practitioners who want to split long runs into short invocations for post-training, or to choose Worker count and Worker model per task and time constraint, the work offers a reusable design principle and cost accounting.

Most main-result configurations use a single run; only Anthropic kernel optimization with Codex and the checkpoint ablations are replicated three times, so the stability of task-to-task differences needs more repetition. Cost estimates exclude the compute for running and evaluating experiments. The ablations change what each agent sees, not whether it keeps its conversation across invocations, so the separate contributions of statelessness itself and of context content remain open. All tasks have executable evaluators, leaving vague or slow-to-evaluate goals untested. The Advisor's token share is low, so decision quality and whether to split the Advisor remain open questions, and the best Worker count varies with task, token budget, and time constraint. In addition, some numeric values are missing from the text of the tables, so this summary relies on the numbers that are readable.

Sources