Skip to main content
Back to timeline
arXivSource publication:

WhiteMatter lets every Transformer layer read past-token representations from any depth, beating matched baselines at half KV cache

Synopsis

The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.

AI-generated editorial illustration: WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

Interpretation

WhiteMatter uses a content-dependent router to mix past-token states from any depth into specialized KV channels for each target layer, removing the standard Transformer constraint that each layer can only use same-depth KV. Earlier deep-to-shallow feedback methods compress available sources into a single layer-width representation, such as Feedback Transformer pooling all layer states into one weighted sum and LCKV building shared KV from the deepest state; WhiteMatter instead gives different target layers different source mixtures and can share channels across layers. The paper provides the method derivation plus router parameterization and initialization details, and ablations support distinct mixtures, dynamic routing, and deep-to-shallow feedback: replacing distinct mixtures with one shared mixture performs worse, and replacing the dynamic router with static learnable weights raises perplexity.

Under matched training-token budgets, half-cache WhiteMatter lowers held-out perplexity relative to matched standard Transformers and improves zero-shot downstream average scores at two model scales up to about 1.3B parameters; the full-cache version performs comparably to a standard Transformer with more layers at the smaller scale. At the smaller scale the paper reports that both half-cache and full-cache WhiteMatter exceed all baselines on average, including the deeper standard Transformer, and outperform equal-cache LCKV and FusedKV. Evidence comes from Qwen3-based decoders trained from scratch on FineWeb-Edu, with matched token budgets and effective batch sizes at two scales, evaluated on final checkpoints and full evaluation splits; the paper also reports that autoregressive perplexity for WhiteMatter and LCKV differs only slightly from fixed-point iteration results.

Cyclic Gauss–Seidel iteration divides token positions into groups processed in order with tokens within each group evaluated in parallel, letting updates propagate farther within a single full-sequence evaluation and speeding up training and prefill convergence. LCKV's Jacobi iteration processes all tokens in parallel but can only read the preceding pass's KV, so updates move forward only between passes; cyclic grouping lets later groups read updates from earlier groups in the same pass. On a reference model trained with exact autoregressive execution, cyclic iteration reaches the target perplexity in four passes at a millisecond-scale per-sequence time while Jacobi needs more passes, which the paper reports as a 12.5x speedup; on a larger model cyclic iteration likewise reaches the target in fewer passes.

At about 1.3B parameters, WhiteMatter matches standard Transformer decoding throughput with lower peak device memory, while iterative training and prefill remain more expensive than the standard Transformer. The paper reports quality benefits alongside execution costs, including decode, prefill, and training FLOP comparisons plus prefill throughput and peak memory measurements. Runtime measurements use compiled BF16 execution at a fixed batch size on an RTX A6000, reporting median throughput over repeated runs and peak memory; FLOPs are measured with PyTorch at fixed batch size and sequence length.

Perspective

The result targets pretraining and inference for autoregressive language models trained from scratch: under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores at two scales up to about 1.3B parameters, and the full-cache version matches a deeper standard Transformer at the smaller scale; decoding throughput matches the standard Transformer with lower peak device memory. It suits research and engineering settings that want to reuse past-token representations without adding layers and are willing to accept an iterative training and prefill procedure; the ablations also indicate that distinct mixtures, dynamic routing, and deep-to-shallow feedback drive the gains.

The paper itself states that iterative training and prefill remain more expensive than standard Transformers, that the main quality results cover up to about 1.3B parameters trained for a fixed token budget, that effectiveness on larger models trained with more data requires further investigation, and that full-cache and alternative KV-sharing baselines are evaluated only at the smaller scale. In addition, the homepage evidence bundle's external story leaves several numbers blank (such as parameter counts, perplexity reductions, and throughput ratios), so this summary relies on the tables and prose explicitly present in the paper text; readers needing exact values should check the corresponding tables and appendices in the original.

Sources