Draft-KV makes a receiver actually use the sharer's draft KV states: a 0.5B receiver rises from 37.45% to 78.04% on MMLU-Redux, and falls back to 36.40% when the message comes from an unrelated question
Synopsis
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
Interpretation
The paper names and quantifies an information-use gap: accuracy gains of existing latent interfaces need not depend on the correctly paired message content. Prior work largely treats end-task accuracy gains for the receiver as evidence that communication works; this work introduces Matched, Deranged, and Receiver-only conditions that separate system gain from pairing gain. Across five method-dataset pairs covering C2C and LatentMAS, pairing gain never exceeds 0.60 points while system gain reaches 15.44 points; for C2C fusers retrained under the authors' recipe, 13 of 20 per-question paired bootstrap intervals on the five Public benchmarks contain zero, and the seven exceptions stay within a few points of zero.
Draft-KV makes the message the key-value states formed while the sharer drafts an answer, rather than a static prompt cache. Existing interfaces mostly transmit the sharer's prompt cache and fuse it into the receiver's cache; Draft-KV transmits draft-position KV states, projects them with two bias-free linear maps into a separate side memory, and reads them through a gated attention branch that reuses the receiver's frozen query and output projections, with gates initialized to zero. An ablation replacing draft-position KV with prompt-position KV lowers Matched accuracy by 33.70 and 21.48 points on ARC-Challenge and MMLU-Redux and nearly eliminates pairing gain, which falls to 0.17 and 0.09 points, reproducing the information-use gap inside the authors' own architecture.
Progressive training (message reconstruction, then answer alignment, then answer-text training with a one-sided guard) makes the receiver genuinely depend on the paired message. Readability and use are trained separately, and the third stage adds a one-sided hinge guard that acts only when a mismatched packet makes the gold answer more than a threshold costlier per token than sending nothing. Deleting either reconstruction pretraining or answer alignment lowers both Matched accuracy and pairing gain on both benchmarks, with markedly larger losses on MMLU-Redux, which no stage trains on; removing the guard raises mismatch harm and roughly doubles the mean hinge while Matched accuracy falls only slightly.
Draft-KV's gains grow with sharer capability and transfer to held-out tasks and to settings where each model holds different evidence. In existing interfaces a stronger sharer does not consistently improve the system; Draft-KV, at fixed interface size, raises mean system gain by 32.6 points from a 0.6B to an 8B sharer, against 30.2 for T2T and 1.4 for C2C. Under the Public protocol it beats the receiver alone in all 35 settings, is most accurate in 30, and exceeds T2T in 33, with an average gain of 5.85 points; under the Private protocol it beats T2T in all 14 model-benchmark settings by 6.41 points on average, and in seven settings it also exceeds both models answering independently.
Perspective
The result targets heterogeneous model collaboration: both models stay frozen, an interface is trained per model pair, and the setting suits tasks where the sharer can first draft an answer and the receiver needs to use that draft. Under the Public protocol both models see the same question; under the Private protocol each holds half of the annotated evidence, so the message is the receiver's only route to the other half. The reported latency and arithmetic decomposition shows Draft-KV's added cost comes mainly from one compute-bound forward pass over prompt and draft, while the interface itself is 0.885 GFLOPs on MMLU-Redux and 0.323 GFLOPs on HotpotQA, two orders of magnitude below C2C's 118.3 GFLOPs. For systems that want to reuse a stronger partner's capability, this means gains can be obtained by raising sharer capability at fixed interface size.
The message replacement tests cross-question mismatches only and does not speak to robustness when the draft itself is wrong for the same question; the adapter-only control isolates a receiver-conditioned component of the interface, which the paper notes is not a causal decomposition of source information. In the sharer-scaling sweep the 0.6B point is the one case where draft length and capability are hard to separate, since its drafts are short and highly variable. DLC has no released implementation and is not reproduced, so its numbers are taken from the published paper and the authors make no claim about its behaviour under message intervention. Per-question paired intervals cover only four C2C fuser configurations the authors trained; the released checkpoint and the LatentMAS runs have point estimates only. In addition, some numeric values appear blank in the abstract and body text of the loaded markdown, so the concrete figures should be read from the experiments section as listed there.
