LOCI conditions a linear memory on camera geometry alongside historical KV, reaching 14.36 dB revisit PSNR on the MIND memory benchmark and streaming 300 seconds at a constant 23.6 GiB
Synopsis
LOCI introduces a hybrid spatial-memory architecture in which 15 of 30 transformer blocks keep a key-value cache of past observations while the other 15 are restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, with recurrent readouts flowing into subsequent cache-backed blocks to supply their queries with accumulated scene context; on the public MIND memory benchmark and on held-out recorded trajectories it reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model, lowers peak memory at equal length by about 30% with full history, and streams long videos at constant memory with a bounded bank of retained observations.
Interpretation
LOCI keeps both memory representations in one network: half the blocks retain a KV cache of past observations to preserve observation-level detail, while the other half compress the entire history into a fixed-size recurrent linear-attention state whose reads and writes are conditioned on projective camera geometry via PRoPE, so viewpoint enters both memory addressing and stored content. Earlier explicit-memory world models select past observations by field-of-view overlap, a surfel index, or query-key similarity, while local-global architectures consolidate history into a fixed state and lose scene-specific detail; LOCI adds the recurrent readout to the token features from which later historical-attention blocks form their queries, letting compressed history direct access to retained observations without a separate state-to-query adapter or router. The paper trains on a 5B Wan2.2-TI2V-5B backbone for 5,000 updates, with the recurrent branch adding 90.7M parameters (1.7% of 5.38B); the controlled ablation shows the hybrid raises PSNR by 0.89 dB over a same-recipe full-softmax model on MIND under an identical bounded KV budget and is better in 44 of 50 segments.
On the public MIND memory benchmark, LOCI evaluated zero-shot attains lower MSE and higher PSNR and SSIM than the values reported for GIM-World, which is trained on MIND, and is best on all four metrics among the streaming world models scored on all segments. The paper reports LOCI at MSE 0.0455, PSNR 14.36, SSIM 0.464 and LPIPS 0.643, against reported values of 0.0614, 13.40, 0.414 and 0.630 for GIM-World; among the streaming models the authors run, LOCI is 1.82 dB PSNR above the strongest (HY-WorldPlay), and several of those external models are larger (8-28B). Evaluation follows the official MIND protocol over the entire prediction segment, with 95% bootstrap intervals over clips (10,000 resamples); the authors note that external models differ in size, training data, sampling steps and how they receive the memory segment, so this comparison is a reference under a unified protocol while the controlled comparison comes from the same-recipe ablation.
In bounded sparse mode, LOCI streams 300 seconds of video at constant memory using a fixed-capacity bank of retained observations plus the recurrent state, and remains more faithful than full softmax under the same budget. Sparse mode retains a conditioning-frame sink, the 8 most recent frames and a panorama bank of at most 20 frames, so softmax attention sees at most 34 latent frames; discarded observations are not archived, but the recurrent state continues to carry their information, so neither explicit storage nor the recurrent state grows with video length. The paper reports 300-second runs at a constant 23.6 GiB in bounded sparse mode, 15% less memory than full softmax under the same access (27.6 GiB) at the same speed (5.4 s per video second for both on one H200); on recorded trajectories, bounded sparse reaches revisit PSNR 11.21 dB and LPIPS 0.547, slightly higher than the full-history setting.
The recurrent state does carry camera-decodable scene information and shapes how downstream historical attention is allocated: camera yaw is linearly decodable from the accumulated state, and LOCI's historical attention concentrates more on co-visible content. The paper uses PRoPE to write camera geometry into the recurrent keys and values and tests whether yaw relative to the first frame remains linearly decodable from the accumulated state; it also replays the same ground-truth history through both models and compares attention at the same query positions to measure the share of attention going to co-visible tokens. Reading the state at every chunk boundary from the ninth chunk on (32 designed trajectories, mean over the 15 hybrid layers, cross-validation and bootstrap intervals clustered by trajectory) yields held-out R-squared 0.951 and a median error of 3.6 degrees; co-visible tokens make up only about 0.1% of the history, yet LOCI's 15 softmax layers place 11.1% of their history attention on them versus 10.45% for full softmax, consistent in 8 of 8 clips; on each model's own rollouts LOCI also reconstructs the co-visible region more closely (masked LPIPS 0.597 vs. 0.654, lower at 131 of 154 instants).
Perspective
The result targets research and applications in camera-controllable, long-horizon streaming video world models that require revisit consistency, and it applies to revisit-fidelity evaluation on rendered or recorded trajectories and to continued generation under bounded explicit storage. It lets follow-up work extend generation duration without history storage growing with video length and makes camera geometry an explicit cue for memory addressing; for teams using similar hybrid-attention backbones that need to preserve observation-level detail, this combination of a recurrent state with retained KV offers a directly reusable inter-layer interaction. The held-out trajectories are rendered in Unreal Engine and real-world captures are not evaluated, so the conclusions apply to the evaluated rendered and benchmark settings.
Absolute fidelity is low for all models (benchmark-average PSNR below 15 dB), so the gains should be read as relative; the study covers one 5B backbone with a short fine-tuning budget of 5,000 updates. Both models keep an explicit-history camera-attention branch in all 30 blocks, so the recurrent path halves main-attention KV rather than all stored history. The recurrent state is built from noised chunk features in training but committed from generated chunks at inference, and how that difference affects long-horizon behavior is worth watching. The per-token versus per-chunk retention comparison uses an extended metric defined after the short-range comparison and is exploratory. Speed parity with full softmax in bounded sparse mode relies on the optimized implementation described in the paper. In addition, some tables and interval values are missing in the text read here, so specific confidence intervals and per-item numbers should be checked against the original.
