Skip to main content
Back to timeline
arXivSource publication:

TomasuLLM runs LLM agent tool calls out of order speculatively, reaching 1.31x on SWE-bench Verified with zero false accepts across 4,010 commit-validation records

Related research and updates

Synopsis

TomasuLLM is a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state; across three benchmarks spanning sub-second to minutes-long tool calls it reports 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records.

Source-provided article image: TomasuLLM: Out-of-Order Speculative Execution for LLM Agents
Figure 1 ·

Figure 1. Observation-stall slack and the commit boundary. (a) The base agent serializes at observation boundaries. (b) TomasuLLM executes drafted actions in COW overlays and publishes only validated observations in main-agent order. (c) The CPU analogy separates prediction from in-order commit. Three side-by-side panels. Panel a shows the base-agent serial path: LLM reasoning, tool execution T1, LLM reasoning, tool execution T2, and the final answer, where each tool observation must return before the next action issues. Panel b shows the TomasuLLM runtime with a drafting region and an overlay-execution region. The Action Drafter emits predicted tool actions, the Observation Drafter emits a predicted observation, and real tool executions T1 and T2 run in copy-on-write overlays, with T2 finishing first; sandbox observations flow into the Validate and in-order commit stage before publication, and a mismatch discards the speculative suffix. Panel c shows the CPU analogy: branch and value prediction feed fetch, decode, rename, and issue; instructions I1 and I2 execute out of order; validation and the reorder buffer commit in order, with mispredictions squashed.

arXiv

Interpretation

The work presents TomasuLLM, a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness. Relative to a conventional agent loop that executes in trajectory order, it starts predictable future work early to relieve the observation stall caused by long-running tools such as compilers, test suites, and repository commands. The abstract supports this with a mechanism description plus quantitative results on three benchmarks and a correctness statistic of zero false accepts across 4,010 audited commit-validation records.

Its execution flow drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order after validation against committed state. A speculative result becomes visible only after it and every earlier step have been validated, combining out-of-order execution with a sequential visibility interface. The abstract explicitly states sandbox isolation, dependency and effect tracing, in-trajectory-order commit, and validation against committed state.

Speedups scale with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. The three benchmarks span sub-second to minutes-long tool calls, indicating the gains track the tool-latency bottleneck directly. The abstract reports benchmark means and states that the improvement scales with tool latency.

Across 4,010 audited commit-validation records, the system produces zero false accepts. This statistic targets the central risk of speculative execution, namely exposing an unvalidated result early. The abstract gives the record count and the zero-false-accept count.

Perspective

The work targets coding-agent settings where tool calls take from sub-second to minutes, especially workloads dominated by compilers, test suites, and repository commands; its gains are described as scaling with tool latency, so it applies most to trajectories with a high share of long-running tools. For runtime and agent-framework developers seeking to reduce agent idle waiting, it offers a path to start predictable work early while preserving in-trajectory-order visibility.

The abstract reports only benchmark means, without per-task distributions, variance, or item-by-item baseline comparisons, so how stable the gains are across tasks remains an open question. Zero false accepts rests on 4,010 audited records, and the abstract does not say how that conclusion holds at larger or more heterogeneous trajectory sets. In addition, the speculative draft hit rate, sandbox overhead, and fallback cost on validation failure are not quantified in the abstract, all of which affect net benefit in deployment.

Sources