Skip to main content
Back to timeline
arXivSource publication:

HeteroFold lets Llama, Qwen, and Ministral share KV caches across families, running about 10.7x faster than native prefill at 32K context

Synopsis

The work proposes HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both sender and receiver frozen and, through token alignment over shared character boundaries, cross-layer and cross-head K/V mapping, and receiver-aware calibration, achieves the best cache-transfer performance on all four long-context benchmarks across six transfer directions among Llama-3.1-8B, Qwen3-4B, and Ministral-3-14B, matches text-based communication on the HiddenBench multi-agent benchmark, and at 32K context makes Llama-3.1-8B to Ministral-3-14B transfer about 10.7x faster than native prefill.

AI-generated editorial illustration: Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

Interpretation

The paper identifies tokenizer mismatch as a key obstacle to prefill-free cross-family KV reuse and introduces Token Alignment (TA), which establishes sender-receiver token correspondence through shared character-end boundaries. Prior prefill-free cache-transfer methods such as Dense Latent and KV Ridge were evaluated mainly within model families with compatible tokenization and did not address cross-family token correspondence. The ablation shows that replacing TA with same-index pairing causes the largest degradation in the Ministral-3-14B to Llama-3.1-8B direction, and Figure 5 shows that same-index position mismatch grows with context length while TA maintains close alignment through shared character boundaries.

The paper develops a cross-family K/V mapping that combines multi-layer sender features, cross-head mixing, and Recolor moment matching to construct receiver-compatible K/V states while keeping both language models frozen. Compared with single-layer and head-local mappings, the method explicitly handles differences in model depth, KV structure, and feature distributions, and uses Recolor to match receiver feature means and covariances rather than only token-wise reconstruction. The ablation shows single-layer and head-local mappings reduce performance; Table D shows three sender layers improve ARC-C, HotpotQA, and QuALITY from 66.89/22.52/51.58 to 77.05/49.44/62.90 over a single layer; Table F shows mapper parameters of 201-252 million across six directions, about 25% fewer than Dense Latent and 62% fewer than KV Ridge.

The paper shows that KV reconstruction alone does not preserve receiver behavior and introduces receiver-aware calibration that directly matches attention patterns and outputs, with learned corrections folded into fixed affine maps so no separate inference module is needed. KV Ridge achieves lower reconstruction error than HeteroFold yet distorts receiver attention, indicating the need for calibration targeting downstream receiver computation rather than cache-value reconstruction alone. Figures 2 and 4(d) show attention-output error persists after Recolor; Table C shows nonzero correction ranks improve HotpotQA over Recolor alone; Appendix D.3 shows prediction divergence does not increase with decoding length on GSM8K autoregressive decoding.

The paper evaluates six transfer directions on four long-context and five short-context benchmarks and reaches performance comparable to text-based communication on HiddenBench multi-round heterogeneous-agent communication while reducing receiver-side transfer latency. Prior prefill-free transfer lacked systematic evaluation across families, across tokenizers, and in multi-round generated-message communication. Long-context benchmarks are Qasper, HotpotQA, LoCoMo, and QuALITY; short-context benchmarks are ARC-Challenge, MMLU, WinoGrande, HellaSwag, and GSM8K; HiddenBench uses all 65 tasks with three or four agents over 15 communication rounds; latency is measured on two H100 80GB GPUs with NVLink at batch size 2, about 10.7x faster than native prefill at 32K and 1.18-1.47x faster than Dense Latent and KV Ridge.

Perspective

The result targets multi-agent settings where the sender has already processed the shared context: sender and receiver reside on two separate H100 80GB GPUs connected by NVLink with batch size 2, and sender prefill, model loading, and offline calibration are excluded from transfer latency; if the sender has not yet processed the context, sender prefill and payload capture must be added (the paper reports 280/1328/3381 ms for Llama and 235/1163/3083 ms for Qwen at 4K/16K/32K). The method is validated across six directions among Llama-3.1-8B, Qwen3-4B, and Ministral-3-14B, with same-family Qwen3 transfers as a complementary evaluation; calibration uses 1,600 training prompts and 400 held-out prompts, excluding gold answers and solutions.

The paper reports performance comparable to text-based communication on HiddenBench but does not give per-item numbers; cross-family transfer is not ahead in every short-context setting, only in most; calibration sensitivity shows correction rank and calibration corpus affect results, for example Open-R1-only calibration gives 12.69 on HotpotQA versus 41.97 for the equal mixture; natural-language reconstruction shows HeteroFold ordered overlap of 83.04 and exact reconstruction of 16/100, indicating cache transfer is not lossless; and cache transfer does not always reduce additional peak receiver memory once incoming payload copies are included.

Sources