Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Synopsis
The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
Figure 12. Median absolute ETA error at 50 % and 90 % of the call, as a share of the call’s own duration, pooled over tests, builds and instrumented scripts. Rows are the load condition of the running call. Each history column shows the best method given the history source (idle, mild, high, mixed); the last column is the progress stream. Each call replayed under the four conditions.
arXivInterpretation
The paper shows that no estimate fixed before a call starts can know its duration, and such estimates may not even rank the calls, while the running tool already holds the answer. Relative to prior practice that reads the tool's name, its history, a duration declared before the call, or the engine's own occupancy, the paper places today's timing signals on a ladder and shows with evidence that every rung fixes its estimate before the call runs or never sees inside it. Evidence comes from replays of leaderboard runs, OpenHands evaluation runs, and the authors' own runs: the same command finished an order of magnitude faster on an idle sandbox, adding busy neighbours made most calls several times slower, and the order of long calls recorded under production load has no correlation with their order on an idle machine.
The paper gives the first corpus-scale account of progress in agent tool calls, classifying what a tool holds into strong signals (a fraction of work remaining, or a count out of a known total) and weak signals (only an accurate signal that the end is near), across four sources. No prior systematic count of readable in-flight progress existed; the paper reports the share of tool time in each kind and what it takes to reveal it. The census samples calls from four public corpora, stratified by source, scaffold, command class, and duration, labelled by two frontier language models with a third arbitrating disagreements and a human-labelled subset checking them; with every instrument applied, a strong signal covers 38% of the leaderboard's tool time, 51% of the OpenHands runs', and 48% of the authors' own, and a weak signal raises that to 66%, 59%, and 72%.
The paper builds a harness that turns tool output and the tool's footprint on disk into one stream of progress events without changing what the agent sees. Relative to prior designs that would change the agent's observation or rely on a duration declared before the call, the harness leaves the result path byte-identical and sends progress over a separate side channel, handling trustworthiness issues such as a counter passing its total and folding multi-phase bars into one fraction. On a sample of SWE-bench Verified tasks, the side-channel configuration did not change the agent's score significantly, and steps, tokens, tool time, and wall time were all unchanged as well; together the two layers add less than one percent to the wall clock of a realistic call.
The paper compares the reported progress with four published predictors at the points where a KV cache decision is made, and measures end to end what the signal buys in a production engine. Relative to prior work that measured prediction error on sandboxes of one fixed size, the paper evaluates at the midpoint and near the end of the call, and covers a changing environment, different base models, and lying sessions. At 50% of each call the progress stream's median error is about a fifth of the call's duration against a third for the best history predictor; at 90% it is under 10% against 80%; replayed on a 4×H100 SXM instance, post-tool p90 TTFT falls 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, where the oracle cuts it by a little over a quarter and by 23% respectively.
Perspective
The result is meant for LLM serving systems that must decide, during a tool call, whether a KV cache stays, leaves, or comes back, and applies to the long installs, builds, test suites, and generated scripts of a coding agent; the paper also gives a session-credit mechanism that treats the tool's report as untrusted input, so the worst a liar can do is lose the benefit while honest sessions are unaffected. The paper calls on tools to expose their own state in a form they choose, at which point no switch, dry run, or parser would be needed and the harness would only relay it.
The paper states three boundaries: some tool calls hold no information a harness can read (a remote API that answers only at the end, a command that waits for a person, a download stalled before its first byte); some information cannot be parsed (an output format not yet met, a build system that hides its units, a script that reports in its own words); and drawing the signal out through a switch, a dry run, or an instrumented script is a change to how the tool runs, so every new instrument has to be measured again. In addition, when the load is removed mid-call the progress stream overestimates the remaining time until it has observed the new rate, and a windowed rate shortens that lag; the credit-mechanism result for lying sessions comes from a preliminary simulation.
