Galahad turns LLM document reading into a one-time cost with byte-exact KV memory: 100 of 100 on a 97,000-token corpus at 0.59–0.64 s per question
Synopsis
Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
Interpretation
Taliesin writes the KV state of a block of text to storage under a fingerprint of the exact input bytes, the model and the tenant, and loads it when the same bytes appear again instead of recomputing; restored state is bit-identical to a fresh prefill (all 262,144 output logits matched after process restart, rehydration from disk and hot-load). Previously vLLM's PagedAttention and SGLang's RadixAttention shared KV state for common prefixes only while it stayed in GPU memory, LMCache and Mooncake offloaded KV to CPU memory or disk, and CacheBlend and RAGCache accepted approximation when combining retrieved chunks; Galahad extends reuse beyond GPU memory and process lifetime and makes bit-identical output the acceptance test, treating any load that fails a check as a miss. The correctness table reports 262,144 of 262,144 logits bit-identical on save and restore, 20 of 20 byte-equal deterministic prefills, 64 of 64 greedy tokens identical when state moved across GPU types, 5 of 5 token-exact on a hybrid architecture, 276 of 276 tokens bit-identical in snapshots, and an encryption round trip of 9.77 GiB with zero failures plus 4,167 NIST CAVP vectors with zero failures; each of four safety mechanisms was attacked three times (defence on, off, restored) and all four met the criterion.
On a recall test with 100 facts in a 96,726-token corpus of 11 blocks, Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.01 s and 572 J per question, against 10 of 100, 9.25 s and 2,754 J for the no-Galahad baseline limited to the last 12,000 tokens; each question sent 53,219 prompt tokens on average, 99.5% of them loaded from Taliesin. This result involves no retrieval and no tuning on the test questions, so it isolates the effect of KV persistence: the model processed 5.5 times as many prompt tokens as the truncated baseline yet finished 3.1 times faster with 79% less GPU energy. Reported separately for llama.cpp, vLLM and SGLang, where Taliesin alone answered 98–100 of 100; a repeat on llama.cpp gave 98 correct at 3.00 s and 580 J; with vLLM's own cache kept between questions, Taliesin alone answered 99 of 100 at 1.09 s and 329 J.
Blaise keeps each document as byte-exact text divided into sections, selects one section on the CPU for a question, and passes only that section's text to the model; in the recall test the model read about 668 tokens per question and answered 100 of 100 on every runtime, at 0.59 s and 200 J per question on vLLM, 13.7 times faster and 92% lower GPU energy than the no-Galahad baseline. Taliesin removes repeated computation and Blaise removes unnecessary reading, and each can run without the other; on llama.cpp Taliesin alone was already 3.1 times faster than the baseline, and Blaise, by giving the model 668 tokens instead of 96,726, took that to 14.5. Blaise was developed on this recall test but still answered 91–92% (319–321 of 349) on seven datasets it had never seen and was 3–10 times faster per question than Taliesin alone; as an external baseline, a tuned RAGFlow v0.27.2 pipeline answered 77 of 100.
Across 349 questions from seven real-world datasets (help-desk tickets, customer-support logs, PDFs with extraction errors, SWE-bench Lite, The Stack, AgentBench and WebArena), 98.69% of prompt tokens were loaded from Taliesin rather than recomputed (llama.cpp 99.49%, vLLM 99.37%, SGLang 97.21%); Taliesin alone answered 329–333 (94–95%) and with Blaise 319–321 (91–92%). This carries reuse from a controlled recall test to untuned real workloads and separates the contribution of each part; it also shows selection decides accuracy, since in 9 of 50 messy-PDF questions the answer was lost during text extraction, and on the 41 answerable questions Taliesin alone answered 37 while Taliesin plus Blaise answered 30. One run per runtime with seed 20260927, Gemma 4 31B deployed on an A6000, an RTX 6000 Ada and an L40S; every retry and every loaded block counts toward time, energy and token totals.
Perspective
The work targets serving settings that return to the same documents, such as support assistants, coding agents, contract review and agent tasks; Taliesin and Blaise can be used separately or together. It explicitly does not try to raise the per-token computation bound described by Sikka and Sikka, but moves repeated computation, search and verification away from the model: search runs on the CPU, and verification is decided by byte comparison and cryptographic hashes. On availability, Galahad ships as a Linux x86-64 shared library, enters public beta on 1 October 2026, is free for non-commercial use on one GPU for 12 months, and commercial pilots are available on request. The paper describes itself as a system description and does not disclose the storage format, the internal structure of Blaise or the security design in implementation detail, which are covered by pending patent applications; Blaise's second mode, in which the model reads the corpus index once and Taliesin keeps that reading, is left for later benchmarking.
Blaise's selection method is not described, so the conclusion that selection decides accuracy can currently only be observed from the results: on seven unseen datasets it answered 91–92%, below the 94–95% of Taliesin alone, indicating that reading less text helps only when the selected section contains the answer. Blaise was developed on the recall test, so its 100 of 100 there and its 91–92% on seven datasets should be read separately. Some results (throughput, the storage-medium comparison, the 284B model) are single runs, and the author marks the SGLang Taliesin-only time as using an older backend and therefore an upper bound. The 9 of 50 messy-PDF misses happened during text extraction, and no memory layer can restore text an upstream parser dropped. In addition, although this evidence bundle is full text, the paper states that it omits the storage format and Blaise's internal structure, so readers interested in implementation detail will need later material.
