Skip to main content
Back to timeline
arXivSource publication:

TokenCast forecasts token consumption during agent runs with composable segment costs, cutting MAE by 14.5% on average across 96 comparisons

Synopsis

TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.

AI-generated editorial illustration: TokenCast: Forecasting Token Consumption During LLM Agent Execution

Interpretation

The paper introduces a segment-cost factorization and an exact composition identity: each execution segment is represented by call count, net input-length change, and a cost residual, and composing adjacent segments explicitly captures the extra input cost incurred when context from earlier segments is re-read by every later call. Prior request-level methods predict a single response length, task-level methods such as Self-Prediction estimate total consumption before execution, and multi-step frameworks rely on program structures or dependency graphs known before execution; TokenCast targets open-ended agents where no complete dependency graph exists and context keeps accumulating, splitting remaining-cost prediction into a prefix segment, a boundary state, and a suffix segment. The paper derives the composition identity and shows the composition is associative and yields the same residual regardless of how a trace is partitioned; ablations show removing composition raises normalized average MAE from 0.69 to 0.78, removing net input-length change raises it to 0.76, and removing the residual raises it to 0.73.

The paper builds a staged prefix–suffix predictor that combines direct and compositional forecasts and uses calibrated quantile models for prediction intervals, refreshing forecasts from observed evidence as execution unfolds with no additional LLM calls. Unlike request-level methods that rely on model internal states or entropy statistics, TokenCast uses only task and execution features visible at the prediction point; unlike Self-Prediction, which spends additional LLM calls to estimate cost, its forecasting incurs zero extra LLM calls. On SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run with about 19.7 forecasts per run; updating every call adds under 0.03% to the 129 s median run wall time; LightGBM inference takes 0.8 ms, versus 3.9 ms for XGBoost and 5.2 ms for CatBoost.

Across 4 benchmarks (SWE-bench Verified, Search-R1, MMLU-Pro, LongBench-v2) and 6 agent models, TokenCast reduces MAE against the strongest comparator in each combination by 14.5% on average over 96 comparisons. The paper reports no losses at In-call Update or Task Update, with losses concentrated at Task Start (15 combinations) and Call Start (9); for example, on SWE-bench Verified with GPT-5.4, In-call Update MAE falls from EGTP's 74.6 to 38.9 tokens and Task Update from TRAIL's 115k to 80k tokens. The study collects 11,712 execution traces from 240 benchmark tasks, 6 agent LLMs, and two harnesses; tasks are split into five folds with three random seeds, and all runs of a task stay in one partition.

In offline budget-control replay, TokenCast matches a fixed-budget policy's trace-completion rate while using 21.3% fewer tokens on average. The replay wires online forecasts into a stopping rule over 288 GPT-5.4 runs from 144 SWE-bench Verified tasks, with seven budgets set at the 0.3 to 0.9 quantiles of recorded consumption and a stop decision at each Task Update checkpoint based on a selected quantile of predicted remaining consumption. The paper states that trace-complete means reaching the recorded terminal state and that the replay does not measure task resolution; prediction processing and time are included in the replay accounting, with TokenCast charged 0.0k prediction tokens versus 166.0k for Self-Prediction.

Perspective

The result targets provider-accounted input and output token consumption in sequential, open-ended agents whose behavior depends on tool feedback and environment state; predictors are learned from recorded traces, and transfer to a new task domain or agent LLM improves with a small set of target tasks. The budget-control finding comes from offline replay, where trace-complete means reaching the recorded terminal state and the paper explicitly notes the replay does not measure task resolution, so the saving should be read as a token-usage comparison at matched completion rather than a task-success claim.

The composition identity and ablations are given in the main text, but some appendix tables (such as the length-extrapolation versus random-split differences and the per-run standard deviations of update frequency) lack numeric values in the loaded text, so those details can only be read as described in the prose. Zero-shot cross-domain transfer at Call Start and Task Update remains above Self-Prediction, so how many target tasks are needed to reliably surpass it, and how the method behaves under different harnesses and reasoning configurations, remain open questions. Budget control is currently an offline replay, and wiring forecasts into a runtime framework that actively manages execution is left as future work.

Sources