PReCache shares KV caches across multi-LoRA agents via low-rank precomputation and neutral reconstruction, cutting TTFT by up to several times with almost no accuracy loss
Synopsis
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
Interpretation
PreLRShared constructs each agent's low-rank cache by applying every agent's down-projection to the hidden states when each segment of shared context is first processed, so later agents can use their own LR cache without reprocessing the accumulated trajectory. The earlier BaseShared shares the base cache and keeps an agent-specific LR cache, but the current agent still performs full-length backbone processing over the accumulated context to build its LR cache, leaving most repeated prefill unaddressed; PreLRShared replaces that step with lightweight rank-r down-projections. Evaluated on LLaMA-3.1-8B-Instruct and Ministral-8B-Instruct on HotpotQA and ScienceQA, with controlled traces from 1.9k to 66.4k tokens; in single-stream inference at 66.4k tokens PreLRShared reports 27.08 s TTFT versus 73.18 s for NonShared.
ReBaseShared reconstructs the shared base cache from adapter-free hidden states (the neutral base cache), reducing the shared cache's dependence on the previous agent's adapted representation while retaining LR cache precomputation. The paper observes that a shared base cache built from adapter-free hidden states is closer to the current agent's own base cache than one built from the previous agent's adapter-conditioned hidden states, derives the condition under which the neutral base cache has lower relative error, and measures layer-wise error ratios above one for both backbones across the plan-to-action, action-to-reflect, and reflect-to-plan transitions. Table 2 shows ReBaseShared stays closest to NonShared overall across the four backbone-benchmark combinations, with average drops of about 0.57 and 0.42 points for LLaMA-3.1-8B on HotpotQA and ScienceQA and 0.62 and 2.60 points for Ministral-8B; Appendix B.2 reports standard deviations below 1 point across 20 complete benchmark runs.
The paper provides two inference schedules for neutral reconstruction: lazy prefill (LP) performs one contiguous adapter-free prefill after the current turn for single-stream inference, while double batching (DB) folds the adapted and adapter-free paths into the continuous serving batch for concurrent serving. The two schedules produce logically equivalent cache states and differ only in when reconstruction happens, matching edge single-stream and server-side concurrent deployments respectively. Tables 7 and 8 give the numbers: in single-stream at 66.4k tokens LP reports 27.24 s TTFT and 1068 tok/s while DB reports 49.29 s and 889 tok/s; under concurrent serving at 16 QPS DB reports 6.83 s p50 TTFT and 85 tok/s while LP reports 24.79 s and 38 tok/s.
PReCache retains memory efficiency close to BaseShared and keeps its advantage as trajectory length and agent count grow. Both designs store one shared base cache plus compact per-agent LR caches instead of replicating full-dimensional KV caches across agents, whereas selective recomputation retains full-dimensional agent-specific caches for the recomputed layers or tokens. Table 9 shows peak memory of 27.33 GB and 27.34 GB at 66.4k tokens for PreLRShared and ReBaseShared, close to FullShared's 27.13 GB and below NonShared's 39.47 GB; Table 10 shows peak memory rising by only about 0.5 GB when the agent count doubles from 3 to 6, versus about 8 GB for NonShared.
Perspective
The result targets multi-agent systems that share a backbone plus multiple LoRA adapters, especially long-horizon, prefill-dominated trajectories; the default configuration applies LoRA to the query and value projections (QV, rank r), where the key projection has no LR component and is fully shared while the value cache is decomposed into base and LR caches. ReBaseSharedLP suits single-stream edge inference and ReBaseSharedDB suits concurrent server-side serving, producing logically equivalent cache states. Accuracy evaluation uses AutoAct's three-role framework with HotpotQA and ScienceQA, and efficiency evaluation uses controlled traces on a single A6000 or A100.
Several numeric slots in the loaded text are empty, for example the abstract's TTFT speedup factor, per-request throughput gain, and ReBaseShared's average drop are not shown as concrete numbers, and the body similarly reads "up to a TTFT speedup" and "an average drop of only points", so those magnitudes can only be inferred indirectly from appendix tables such as Tables 5, 6, and 9. Accuracy evaluation is also limited to AutoAct's plan-action-reflect workflow with HotpotQA and ScienceQA, and the authors note the analysis across other role structures and task domains is limited; under QKVO adaptation the key cache becomes agent-specific and adds a key LR cache, so the efficiency trend holds but is weaker than under QV. These are scope and open questions rather than refutations of the conclusions.
