Public articles linked to the same research event.
arXiv TomasuLLM is a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state; across three benchmarks spanning sub-second to minutes-long tool calls it reports 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records.
TomasuLLM is a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state; across three benchmarks spanning sub-second to minutes-long tool calls it reports 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records.
TomasuLLM is a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state; across three benchmarks spanning sub-second to minutes-long tool calls it reports 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records.
TomasuLLM is a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness: it drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state; across three benchmarks spanning sub-second to minutes-long tool calls it reports 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions, with zero false accepts across 4,010 audited commit-validation records.