AgentTime benchmark: GPT-6 Astra ends 63% of runs on time in Codex, but some on-time runs sleep after appearing to finish
Synopsis
The authors introduce AgentTime, a benchmark of 222 tasks from 18 sources run in native CLI harnesses to test duration-following, runtime forecasting, and retrospective time estimation; GPT-6 Astra in Codex ends 63% of runs within 5% of the requested duration versus 4% for Fable 5.1, 14 of 158 classifiable Astra runs explicitly slept after appearing to finish, forecasts mostly overestimate natural runtimes, and stripping time cues grows retrospective misses about two- to fivefold.
Figure 1: Astra keeps time; Fable does not. Each point is one of the 1,991 runs in Table 3 . On time is 0.95–1.05 × \times the request, a band about as wide as the dashed line. Dotted lines mark tenfold deviations, and colored lines are log–log fits. Both axes use log ( 1 + t / 30 s ) \log(1+t/30\,\mathrm{s}) .
arXivInterpretation
AgentTime separates temporal capability into duration-following, forecasting, and retrospection, evaluates agents in their native CLI harnesses (Claude Code, Codex) while retaining each benchmark's native scoring, and issues requests from about a minute to 60 hours. Prior evaluations such as METR's time horizons and BRIDGE use human completion time as a proxy for task difficulty and do not measure an agent's control or awareness of its own runtime; AgentTime measures whether agents work for a requested duration. 222 tasks from 18 benchmark families spanning coding, computer use, agentic work, and automated research; duration-following uses three fresh runs per task per agent, totaling 3,442 attempts and about 1,700 agent-hours.
Duration-following varies substantially by model: GPT-6 Astra in Codex ends 63% of runs within 5% of the request, GPT-5.6 Sol 39%, and Fable 5.1 only 4%; Fable runs long on 79% of shortest requests (median 1.8 times) and stops early on 89% of longest requests (median 0.3 times). Swapping harnesses shows the gap follows the model: Astra's within-task slope exceeds Fable's by 0.71 in Claude Code and 0.69 in Codex, while switching harness moves either model's slope by only 0.14 to 0.16. Main experiments use 659 to 666 runs per agent; the harness swap compares 98 matched runs per setup across 18 GPQA/HLE questions and 16 agentic tasks, with an OpenRouter provider check.
Ending on time is not the same as using the time: among 158 classifiable timestamped Astra runs, 14 (9%) explicitly slept after appearing to finish and 74 (47%) re-checked earlier work; a request about 16 times longer raised the native score on only about a quarter of tasks. Prior work studies behavior under budgets or deadlines; AgentTime records explicit waiting as a third outcome and separates returning at a given time from working effectively. Transcripts were read run by run, mostly from CORE-Bench, TUA-Bench, and Terminal-Bench; the authors note these samples are small and that a gap in visible output does not prove reasoning paused.
Forecasts mostly overestimate natural runtimes (83% of Sol's, 66% of Astra's, 63% of Fable 5.1's), while retrospection is close when time cues are present and degrades about two- to fivefold once timestamps are stripped (Sol's typical miss grows from 1.07 to 5.38). This reverses the human planning fallacy in direction, and the five-stage retrospective design localizes the estimate's dependence on time cues in the transcript. Forecasts and no-duration natural runs were measured separately on all 222 tasks; retrospection used 23 tasks with two independent answers per condition, and the Oracle condition with a clock tool made estimates almost exact.
Perspective
The benchmark targets developers and evaluators of agents running long-horizon tasks in native CLI harnesses, for requested durations from about a minute to 60 hours across question answering, software engineering, computer use, web research, document editing, and automated research. It lets researchers observe task score and duration adherence on the same tasks and distinguish on-time returns, early returns, overruns, and explicit sleeping. For engineering practice that schedules agents under budgets, the results suggest treating duration-following as a separate metric rather than reading final answer quality alone.
Each agent ran each task once per requested duration, so a single task's result mixes the agent's response to the request with run-to-run variation; apart from 55 question runs repeated through OpenRouter, that variation was not measured. Only three models from two labs were tested, and duration-following results reflect one instruction wording at maximum reasoning effort. Requests beyond 8.3 hours occur only in METR and PostTrainBench, four tasks in total, and the harness swap covers requests up to 80 minutes. Runtime also depends on serving speed, and runs went through subscriptions, API keys, and OpenRouter. Whether time-awareness comes from timestamped tool calls in training data or from assumptions about generation speed remains undetermined. Neither early return nor explicit sleeping is evidence of monitor evasion, and whether these behaviors predict conduct under oversight is an open question.
