LatentPort hands a 4B hybrid model's recurrent memory straight to a 9B receiver, cutting continuation loss to just 0.076 nats/token above native 9B without rereading the prefix
Synopsis
In one architecture-matched Qwen3.5 4B→9B hybrid (Gated DeltaNet plus full attention) sibling pair, the work installs the source model's persistent inference state after reading a prefix — attention KV plus GDN recurrent matrices and convolution history — directly into the larger receiver with no target prefix replay; holding translated KV fixed, adding the GDN persistent-state package lowers teacher-forced NLL by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]) and improves all 64 PG19 documents, direct recurrent and convolution reuse beats the tested learned GDN maps, and a rank-4 correction with 434,176 trainable parameters brings continuation loss to only 0.076 nats/token above native 9B, JS divergence 0.022, native context recovery 0.
Interpretation
In a hybrid language model, the transferable inference state is more than attention KV: with translated KV fixed, additionally installing the GDN persistent-state package (recurrent matrices, convolution history, and initialization semantics) lowers NLL from 3.115 to 2.368, a 0.7473 nats/token reduction (95% CI [0.6921, 0.8047]), improving all 64 documents and removing about 80% of KV-only excess NLL (mean 79.7%, median 80.2%). Cross-model KV translation, hidden-state messaging, external memory reuse, and same-model recurrent cache reuse are established directions, but per the authors no prior work has demonstrated cross-model transfer of built-in persistent recurrent inference state between differently sized attention–recurrent hybrid LLMs without target prefix replay. Frozen teacher-forced NLL comparison over 64 PG19 documents with 4,096-token prefixes and 64 scored targets each, with the document as the statistical unit and 10,000 paired document bootstrap resamples; same-model restoration checks reach top-1 agreement 1.0 and maximum absolute logit difference 0 on 16 contexts per model across four lengths.
Directly reusing the source's recurrent and convolution state outperforms the tested learned GDN translation: with translated KV unchanged, direct copying lowers NLL from 2.368 to 2.291 (a 0.077-nat gain); in the 32-document component factorial, translating recurrent state instead raises NLL by 0.0280 (CI [0.0104, 0.0475]), and the selected TDD path reaches excess NLL 0.105 on the 64 held-out documents versus 0.134 for TTT. The learned recurrent mapper has lower validation normalized state reconstruction error (0.5037 versus 0.6732 for direct copying) yet performs worse behaviorally, indicating that tensor reconstruction error is insufficient to predict cross-model continuation quality, consistent in direction with CacheBridge's attention-sensitive weighting observation. Two independent experiments: Experiment 1 fits on 128 FineWeb-Edu documents, validates on 32, and tests on 64 PG19 books; Experiment 2 runs an eight-cell component factorial on 32 fresh documents and a primary test on 64 further fresh documents, both excluding Experiment 1 split identities and text hashes.
A rank-4 correction with only 434,176 trainable parameters (identity-anchored, both language models frozen) lowers handoff continuation loss from 2.018 to 1.989 and excess NLL from 0.105 to 0.076, a paired improvement of 0.0290 nats/token (CI [0.0228, 0.0354]) that removes 27.5% of the remaining native gap (CI [22.4%, 33.9%]), with JS divergence 0.022, NCR 0.918, top-1 agreement 0.864, and a significant advantage over continued 4B (2.042). This supplies an intuitive behavioral endpoint: after receiving the imported state, the receiver predicts the observed continuation better than the source that read the prefix, while processing zero historical prefix tokens; the correction's layer-relative residual norm has median 0.0264 and maximum 0.0440. Frozen 64 held-out documents, teacher-forced continuation, paired document bootstrap intervals; the result passes the protocol's full-state gate (FULL_STATE_HANDOFF) but fails the stronger near-native gate (excess NLL 0.076 > 0.05, NCR 0.918 < 0.95, top-1 0.864 < 0.90), so the conditional 16K branch was not run.
Wrong-donor controls show the receiver does use document-specific transferred state: rotating complete donor states removes the advantage, with correct full state beating shuffled state by 0.948 nats/token in Experiment 1 (CI [0.878, 1.022]), and in Experiment 2 corrected minus shuffled NLL is positive while shuffled NLL 3.178 is worse than empty 9B's 2.842. Together with the fixed-KV primary contrast, these controls support useful transfer beyond attention KV, though because the shuffle changes the complete donor state it does not independently localize that context specificity to recurrent matrices or identify particular retained facts. Rotations between equal-length documents with no self-donors, executed independently in both experiments alongside a frozen continuation schedule and state-tracking diagnostics.
Perspective
The result applies to one direction (4B→9B), one geometry-matched Qwen3.5 Base-model pair, 4K prefixes with 64 teacher-forced targets per document, and a receiver processing zero historical prefix tokens; it makes worth testing the design direction that model families could be deliberately trained to preserve compatible persistent-state geometry and functional coordinates across sizes, exposing a stable switching interface, and it lets a reader see that a hybrid model's inference state is more than KV. The authors name unconstrained free generation and replication on an independent model pair as the most decisive next tests, with mismatched geometry as a subsequent, harder test.
Open questions the authors raise include whether small state errors compound under free generation, whether compatibility is peculiar to these Qwen siblings, whether mismatched geometry or different learned update dynamics would break it, and whether shared pretraining or a shared training recipe contributes. The original primary contrast does not separate recurrent matrices from convolution history and initialization semantics, and the frozen path rejects activating only one GDN component, so isolating R from C requires a separately validated intervention. Bootstrap uncertainty is conditional on the frozen models, fitted maps, selected correction, and corpus procedure, and does not cover training-seed or model-pair variation; post-verdict component removals reuse test documents and are exploratory. The timing table omits the translation pass, serialized I/O, model loading, batching, and scheduling, so it is not an end-to-end production benchmark. This summary is based on the loaded full-text evidence bundle; no model evaluation or refitting was performed.
