Skip to main content
Back to timeline
arXivSource publication:

galahad-kv delivers a 50-million-token memory window on one H100: 100/100 zero-recompute restores, 2.8–4.3× faster and 8.8–12.3× less GPU energy than recompute

Synopsis

The authors release the public package galahad-kv, which writes the KV state a model computes for each 16,000-token block to encrypted local NVMe and later restores any block byte-exact with zero recompute; on a 50-million-token real corpus served through vLLM on one NVIDIA H100, all 100 probed blocks were restored without recompute for both Gemma 4 12B and 31B, restores were 2.8–4.3× faster and used 8.8–12.3× less GPU energy than recompute, VRAM stayed flat, and planted-fact recall was 82/100 and 98/100 with zero fabrication.

AI-generated editorial illustration: Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

Interpretation

The system writes the per-block KV state of each 16,000-token block to local NVMe in MRLNCRY1 format (AES-256-GCM, one key per organisation) and, on restore, grafts that block's KV back so the forward pass continues as if the model had just computed it; one block is resident at a time, giving an O(1) moving window. Prior KV offloading work such as LMCache and CacheBlend stages state in plaintext and re-stages it per request, and some of it is architecture-specific; this store is encrypted, durable and byte-exact, and it covers material far outside the window. On a 50M-token corpus, 100 probed blocks per model (depths 0 to 49,488,000 tokens) were all restored with zero recompute for both 12B and 31B; a store audit found all 3,125 blocks carrying the MRLNCRY1 header, on-disk bytes matching the KV footprint math, and block count equal to the number deposited, with no pruning or eviction.

Reuse is cheaper than recompute: median restore was 0.266 s (12B) and 0.347 s (31B) against 0.760 s and 1.475 s for recompute, with median energy of 59 J and 80 J against 521 J and 976 J. Earlier work often reports latency or throughput gains; this paper compares restore against recompute under the same request shape on both time and GPU energy, and shows restore time does not grow with depth. Energy comes from the GPU's NVML hardware counter, idle-subtracted; restore TTFT p90 was 0.284 s (12B) and 0.579 s (31B), and low-third versus high-third depth TTFT was 0.193/0.274 s and 0.343/0.375 s, indicating depth independence. The authors note a per-probe restore of about 0.3 s approaches the NVML counter resolution, so the ratio is reliable while absolute per-probe joules carry a wider band.

Answer quality is set by the reading model, not by the memory: on identical stored blocks the 12B recalled the exact code 82/100 and the 31B 98/100, neither fabricated a code absent from the block, and both refused in the negative control (0/20 fabricated). The result separates the engine property (zero-recompute restore) from the model property (text recall) and shows one durable memory can be served by a cheaper model and escalated to a stronger one. Two metrics are recorded per probe and kept separate; the answer key is held out of the engine; every miss was a wrong line genuinely present in the block, either a same-format decoy or another in-block number, rather than an invented answer.

The authors define a cheat-resistant evaluation protocol and a single-GPU reproduction harness: a runtime mutator replaces names, dates, amounts and e-mails in the real source with random nonces, ingestion is blind (all 50M tokens deposited before any question), one high-entropy needle per block is paired with same-format decoys, and the physical store is audited. The protocol pairs a countermeasure with each of five ways a long-context benchmark can be gamed: pretraining contamination, query lookahead, context dropping, lexical short-circuit and telemetry spoofing; earlier related work did not apply this combination. The four-stage harness (build_hardened.py, deposit_blind.py, audit_storage.py, probe_watertight.py) is a thin public HTTP client (transformers, datasets, urllib) with no private component, and every reported figure has a raw artifact and a SHA-256 manifest.

Perspective

The result applies to KV reuse for a model served through vLLM on a single GPU: one block is resident in VRAM and deeper content is restored from disk on request, so it suits deployments that repeatedly revisit the same already-processed long text and can provide local NVMe with the required capacity. The authors state the method runs on any single 24 GB NVIDIA GPU, with a smaller disk simply holding fewer blocks; the 50M-token store is 1.86 TB at 12B and 6.25 TB at 31B. Because restore time does not grow with depth and the larger model gains more (4.25×/12.3× for 31B against 2.8×/8.8× for 12B), the benefit of reuse scales with model size. The authors also provide a public harness so a third party can swap in any datasets, model and needle design as long as they follow the anti-cheat protocol.

The authors state the system does not widen single-pass attention and does not do learned retrieval (Blaise is deliberately left out and combining it with this window at 50M scale is future work); what is restored is a block that was processed and held in VRAM. A per-probe restore takes about 0.3 s, approaching the NVML counter resolution, so absolute per-probe joules carry a wider band while the restore-versus-recompute ratio is considered reliable. The store must be on local NVMe, since a network volume voids the latency and energy figures. The paper also does not re-establish byte-exact logit equality of grafted KV, which is the subject of the companion Taliesin work; readers who need that property should consult that companion paper.

Sources