Skip to main content
Back to timeline
arXivSource publication:

KVCMAS shares multi-agent KV caches via low-rank anchor pools and chained correction, delivering a 2.0 TTFT speedup over non-shared inference at 32K shared tokens and 8 QPS while cutting peak GPU memory to about 1/3.7 of KVComm

Synopsis

The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.

AI-generated editorial illustration: KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

Interpretation

KVCMAS observes that cross-agent KV cache deviations are concentrated in a low-dimensional feature space, with agent-specific delta corrections showing an effective rank below 32 across workloads and base caches showing a similar structure, so it stores both base caches and delta corrections in the anchor pool in low-rank form while the active KV cache used by attention remains full-dimensional. The prior online correction method KVComm stores full-dimensional base caches together with agent-specific corrections, so memory grows with shared context length, number of anchors, and number of agents; KVCMAS reduces anchor memory from full-dimensional to low-rank, with the paper's example being roughly 50 GiB for full-dimensional anchor states alone with Llama-3.1-8B in BF16, a 4K-token segment pool, and corrections for several agents. Evidence comes from measurements of the effective rank of delta corrections across workloads (reported as below 32) and from single-stream memory comparisons: at 2K, 8K, and 32K shared tokens, KVComm's estimated requirement is 2.5, 6.2, and 14.9 times that of KVCMAS, and KVComm runs out of memory at 8K and 32K shared contexts because of its full-dimensional anchor pool.

KVCMAS uses chained correction: it densely prefills the first agent to obtain an exact cache, and each subsequent workflow edge uses the cache already produced by the source agent as its reference instead of a separately constructed context-free or position-independent reference cache. Existing delta correction methods (GraphFlow, Kamera, KVComm) construct corrections relative to a separately built context-free reference, which for first seen shared context requires an additional prefill outside the workflow, and an approximate correction at the first agent propagates its error to subsequent agents as shared context; the chained structure removes that reference prefill and preserves an exact first-agent cache. In a controlled comparison holding the rest of the configuration fixed at rank 32 and 10 anchors, chained correction improves MMLU by 6.44 points (65.45 versus 59.01) and Video-MME by 1.53 points (52.36 versus 50.83) over a non-chained variant; at 32K shared tokens and 8 QPS, chained correction reduces both p50 and p90 TTFT by 34%.

Across five benchmarks, KVCMAS achieves the highest accuracy among KV cache sharing methods on three of them and ranks second to KVComm on GSM8K and HumanEval within a few points, while providing the lowest end-to-end latency and TTFT among delta correction methods on all workloads. Selective recomputation methods (DroidSpeak, CacheBlend, RelayCaching) rebuild only a selected cache subset, and the paper's Appendix A.1 measurements show at least half of the reuse error remains outside the recomputed subset across workloads, with deviation and attention selection removing only about 20% of the error on MathVista and Video-MME; GraphFlow relies on corrections preconstructed for previously observed context relations and degrades substantially on dynamic requests. Accuracy is averaged over three runs in a framework of three task-specific agents plus a reflection agent, with KVCMAS showing accuracy standard deviations between 0.35 and 0.86 percentage points, comparable to other correction methods; workloads are MMLU, GSM8K, HumanEval, MathVista, and Video-MME with 1,531, 1,319, 161, 1,000, and 1,296 samples respectively.

Under controlled serving traces, KVCMAS achieves the lowest median TTFT in highly concurrent serving, with the advantage growing at longer shared contexts and higher request rates; at 32K shared tokens and 8 QPS it provides a 2.0 TTFT speedup over non-shared inference and 27% lower TTFT than GraphFlow, the next-fastest method. Prior delta correction methods are constrained by reference prefill and full-dimensional anchor states as shared context grows; KVCMAS reduces both costs through low-rank anchors and chained correction, and retains the lowest TTFT and highest throughput as the number of agents and interaction rounds increases. Concurrent experiments are measured on an NVIDIA A100 80 GB at 1 QPS with TTFT averaged across each trajectory, using traces that fix a four-agent schedule and incorporate 128-token agent outputs; single-stream experiments run on an NVIDIA A6000 48 GB, where at 32K shared tokens KVCMAS reaches 3.27 s TTFT and 4,174 tokens/s throughput versus 3.75 s and 3,513 tokens/s for GraphFlow.

Perspective

The result targets multi-agent serving where agents are specialized by distinct role prompts over a shared model, applies to arbitrary agent transitions including sequential, loop, fan-in, and fan-out workflows, and covers both text and vision-language workloads; evaluation uses Llama-3.1-8B-Instruct, Qwen2.5-Coder-7B-Instruct, and LLaVA-OneVision-7B, with concurrent and single-stream settings measured on NVIDIA A100 80 GB and A6000 48 GB respectively. Engineering readers aiming to cut latency and memory in long-context multi-agent serving can directly borrow the low-rank anchor pool and chained correction designs; the paper also notes that works using specialized transfer mechanisms for KV cache sharing require additional training or calibration, while system-level cache optimizations are complementary to cross-agent cache correction.

Reported results are tied to specific models, a specific multi-agent framework, and specific hardware, so the magnitude of the low-rank and chained-correction benefits would need re-verification under other model scales, other agent orchestrations, or different serving stacks; the reliability threshold governs the accuracy-reuse trade-off, and the paper shows that raising it gradually lowers accuracy, so deployments need to reselect the threshold on the target workload; on Video-MME one reuse-eligible transition is consistently routed to dense prefill, saturating the reuse ratio near 0.65, which suggests some transitions may never benefit from correction; and the rank and pool-size ablations show rank 32 with 10 anchors captures most of the gain while larger settings add only marginal accuracy at higher memory, leaving open whether that trade-off holds on other workloads.

Sources