AgSpec combines session and workspace corpora with adaptive draft length to lift coding-agent generation throughput to 4.37x over autoregressive decoding
Synopsis
AgSpec is a retrieval-based speculative decoding framework for coding-agent pipelines that organizes reusable text into session, workspace, and global corpora by source, indexes workspace files in the agent's emission format, and controls draft length with offline-profiled per-agent caps plus an online scale adapted from verification feedback; on the two repository-level multi-agent benchmarks SWE-bench Verified and TeamBench it outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37x at batch size 1 and 4.76x at batch size 16, and it remains effective on benchmarks without a repository or a multi-agent pipeline.
Interpretation
AgSpec splits the retrieval corpus by text source into session, workspace, and global corpora, and indexes workspace files in the forms the agent actually emits (such as unified-diff lines prefixed with a space or a minus), making the code, logs, and earlier attempts that agents repeatedly reproduce retrievable. Existing retrieval-based methods index only the current call, a prebuilt datastore, or outputs pooled across requests, so their corpora miss much reusable text, and they store text as the agent read it, which differs from the token sequence the agent generates. Trace-driven simulation on SWE-bench Verified compares corpus configurations by expected accept length: replacing a per-call corpus with the session corpus gives the largest gain, the workspace corpus adds further headroom on its own and on top of the session corpus, and emission-form indexing helps most in the first turn, raising the coder's accept length before its own patches enter the session corpus, with the gain falling from the second turn on.
AgSpec bounds draft length with an offline-profiled per-agent cap and adapts it online with a scale fitted to verification feedback, so draft length follows the differences and drift in accept length across agents and turns. Existing methods either fix the draft length before serving or scale it only by match length or occurrence count, never correcting it from verification outcomes and remaining blind to roles and turns. A controlled comparison on SWE-bench Verified with Devstral and the SAM-Decoding engine shows long drafts are 14% faster than short drafts at batch size 1 but 30% slower at batch size 16; AgSpec is fastest at both batch sizes, 25% and 16% above the short-draft setting, rejecting at most 6 tokens per step versus 11-26 for the long-draft setting, and enabling only the online scale raises throughput by 66% while cutting the rejected-draft share from 78% to 50%.
AgSpec is a framework that leaves the retrieval engine unchanged and can run on top of both SAM-Decoding and SuffixDecoding, improving throughput and accept length for each. The paper positions AgSpec as a policy layer orthogonal to the retrieval engine, supplying the corpus and draft-length policies rather than replacing the matching mechanism. With SuffixDecoding, throughput rises by 16.5% on average and accept length by 18.5%; with SAM-Decoding, throughput rises by 64.6% and accept length nearly doubles; although SAM-Decoding trails SuffixDecoding in all twelve settings, AgSpec-SAM outperforms it in nine.
AgSpec's gains extend beyond repository-level multi-agent settings to LiveCodeBench without a repository and Terminal-Bench with a single agent. These two benchmarks test generality when the workspace corpus is unavailable and when roles are not explicitly separated, with AgSpec falling back to a shared cap and a single scale table. On LiveCodeBench the faster AgSpec variant reaches 4.87-7.94x the generation throughput of autoregressive decoding and exceeds the fastest prior speculative method by 35-145%; on Terminal-Bench throughput is 9.3-27.2% above the fastest prior speculative method and accept length is 16.7-34.3% above its corresponding retrieval-engine baseline.
Perspective
AgSpec targets coding-agent pipelines, especially repository-level, multi-agent, multi-turn sessions; it assumes a retrieval engine that can search multiple corpora and accept a per-step draft length, so it applies to engines such as SAM-Decoding and SuffixDecoding. Its largest gains come when a session repeatedly reproduces code and tool outputs and when workspace files are written back as patches; on LiveCodeBench without a repository and Terminal-Bench with a single agent it falls back to a shared cap and a single scale table and still improves decoding. The global corpus is built per model and benchmark and fixed before serving, while session and workspace corpora are released when the session ends, so the method suits deployments that isolate corpora per session.
The paper reports generation throughput and accept length, not task success rate or output quality, so whether the acceleration changes problem-solving outcomes remains an open question. The global corpus is built per model and benchmark, leaving cross-model or cross-repository corpus reuse unclear; letting the workspace corpus grow across sessions brings no meaningful benefit, and the authors leave the case of many sessions on the same repository for future work. Under the paper's idealization the scale update corresponds to a quantile loss, but in practice sessions are finite and the distribution drifts, so the scale does not converge and only fluctuates around the target. Also, this reading covers the full text, yet some tables and figures present values as ranges or placeholders, so exact per-baseline numbers would still require the original figures.
