Skip to main content
Back to timeline
arXivSource publication:

CanonicalMerge turns multi-agent KV-cache merging into a convergent replicated state via content-determined ordering, matching the best BagMerge ordering on a partitioned-reasoning benchmark

Synopsis

The work models KV-cache exchange in multi-agent latent reasoning as a convergent replicated state: CanonicalMerge fixes the layout by a content-determined ordering (mean K-norm at a middle layer), making the merged cache byte-identical under any input permutation (verified on synthetic tensors with N<=5 and bit-for-bit on real Qwen3-1.7B and Qwen3-4B KV state), and on a partitioned-reasoning benchmark it matches the best BagMerge ordering without knowing which it is (within 4 points in all 12 cells at 1.7B, same picture at 4B, and tracking the best ordering on Llama-3.1-8B where the gap grows to 19 points), while being the best cache-level method on HotpotQA (n=200) and MuSiQue.

Source-provided article image: Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
Figure 2 ·

Figure 2: Slot asymmetry measured at the mechanism level. For a fixed pair of thinker caches, both BagMerge orders are rendered and the judger’s prompt is forwarded over each merged cache; the curves show the per-token attention mass on a fragment block when the same fragment occupies the prefix slot (blue) versus the tail slot (red). Mean over 10 10 problems (two per family) and both fragments, ± 1 \pm 1 standard error (bands are narrower than the line width); Qwen3-1.7B, query-blind, 40 40 latent steps. The same bytes receive 8.5 % 8.5\% less per-token attention in the prefix slot (pooled ratio 0.915 0.915 ; the final query position alone gives 0.920 0.920 ). The tail block is closer to the judger’s positions and draws more raw attention, while the prefix block keeps its original rotation geometry ( Section 4.2 ); the net effect on accuracy is regime-dependent ( Section 5 ), which is why no fixed order wins consistently.

arXiv

Interpretation

It introduces CanonicalMerge, which fixes the merged layout by a content-determined ordering (mean K-norm at a middle layer), making the merged cache byte-identical under any input permutation. Prior work (Agent Primitives), which the authors name BagMerge, concatenates per-agent caches with RoPE re-encoding; that construction is non-commutative and the best input ordering is not predictable a priori, shifting with deployment regime, latent-step budget, model scale, and model family. Verified on synthetic tensors (N<=5) and bit-for-bit on real Qwen3-1.7B and Qwen3-4B KV state.

It separates state from layout: the durable object is a set of content-addressed latent fragments merged by set union, a state-based CvRDT, and CanonicalMerge is its deterministic render. As a result every accuracy number is inherited and re-delivered duplicates are absorbed, making cache exchange a convergent replicated state. Constructive argument from the set-union semantics of a state-based CvRDT and its deterministic render, building on the bit-for-bit consistency verification above.

On a partitioned-reasoning benchmark, CanonicalMerge matches the best BagMerge ordering without knowing which ordering that is. This indicates content-determined ordering can replace a priori selection of input order, which otherwise drifts with deployment conditions. Within 4 points in all 12 cells at 1.7B, the same picture at 4B, and it again tracks the best ordering on Llama-3.1-8B where the ordering gap grows to 19 points.

On HotpotQA (n=200) and MuSiQue, CanonicalMerge is the best cache-level method. Relative to shipping the text it costs 7 points of F1 on HotpotQA and is at parity on MuSiQue, while the output-fusion baseline PackLLM trails by 45 points. HotpotQA uses n=200 and MuSiQue is at parity; comparisons include shipping the text and the PackLLM output-fusion baseline.

Perspective

The result targets deployment settings where the KV caches of several agents must be composed into one context, especially when input ordering cannot be determined a priori and the merged result should be insensitive to permutation. It makes cache exchange a convergent replicated state: re-delivered fragments are absorbed, accuracy numbers are inherited, and CanonicalMerge as a deterministic render needs no knowledge of which ordering is best. Applicability is bounded by the models and benchmarks verified in the text, including Qwen3-1.7B, Qwen3-4B, and Llama-3.1-8B, and by the partitioned-reasoning benchmark, HotpotQA (n=200), and MuSiQue.

The authors explicitly delimit that at k>2 cache merge transports latent traces but does not by itself compose them, so composing multiple fragments remains an open question. In addition, bit-for-bit consistency verification covers synthetic tensors with N<=5 and Qwen3-1.7B and Qwen3-4B KV state, so behavior at other scales and model families needs separate confirmation; on HotpotQA there remains a 7-point F1 cost against shipping the text while MuSiQue is at parity, and the text does not further specify the conditions under which this difference is stable.

Sources