Skip to main content
Back to timeline
arXivSource publication:

Memorizon trains world models on arbitrarily long spans at fixed sequence length, leading five of eight memory columns

Synopsis

Memorizon proposes a training recipe in which each scored chunk retrieves its own top-k latents by camera co-visibility and their union forms a bounded shared memory bank, so a training sample can cover a span of any length while the attention sequence stays bounded; revisit consistency improves on every split over a sliding-window baseline, and a span long enough to reach the first visit adds further gains at some cost in image quality.

AI-generated editorial illustration: Memorizon: Training World Models Beyond Their Context Window

Interpretation

Memorizon decouples the span a training sample covers from the sequence the transformer attends over: a sample may cover a span of any length but is scored only on its last chunks, the unretrieved history before them is never tokenized, and each scored chunk retrieves its own top-k latents by camera co-visibility whose union forms a shared bank bounded in size, so the sequence does not grow with the span. Prior routes either lengthen the training window, which scales attention quadratically and activation memory linearly, or use sparse attention and history compression to cheapen processing long sequences, but neither changes which pairs of visits a sample can relate; a single retrieval for the whole clip, as in Context-as-Memory, gives every chunk the same frames, whereas here each scored chunk retrieves its own set and shares the union. The paper proves in Proposition 1 that the shared bank is lossless, since a chunk re-selecting inside it recovers exactly the set it would have chosen from the whole history, and gives a bound on bank size; the training-cost table reports that quadrupling the maximum span from 10 s to 40 s leaves the sequence at its bound and peak memory unchanged, adding only a small amount to step time.

Retrieval raises revisit consistency on every split, and training on a span long enough to reach the first visit raises it further, while beyond that point more length no longer helps. The paper treats the span as a sampled variable rather than a fixed clip length and separates the jobs of the two parts: retrieval lets a model read frames far older than any it was trained on, while the span determines whether what is retrieved contains the first visit at all. The ablation adds retrieval, the bank and a longer span one at a time across seen scenes, unseen scenes and web photographs; the paper reports that the 40 s span improves on 10 s and that the 80 s span falls back on every split, and notes that spans of 20, 40 and 80 s reach the first visit of a certain fraction of returning latents in the corpus.

The model reads what the bank holds rather than treating it as filler: changing only the bank contents changes revisit performance substantially. The paper directly tests whether retrieval is used, holding checkpoint, trajectory and seed fixed and varying only the contents of the bank's slots, which speaks to the standing concern that a retrieval mechanism may go unused. The paper reports mid-path Revisit of 0.440 with the retrieved bank, falling to 0.207 with random frames from the same history, 0.105 with an empty bank and 0.073 with frames from another episode, and Gain falling from 0.319 to 0.109, 0.019 and 0.016 respectively.

Against open world models, Memorizon and CaR are the only two systems with a clear memory, with Memorizon highest in five of the eight memory columns and ahead on returns met mid-path. The paper uses the Gain columns to subtract the correlation any two frames of one video already share at the same time gap, separating revisit consistency from ordinary frame-to-frame agreement. The paper reports that for the other five baselines Gain at the starting pose is at most 0.075 and 0.006 for LingBot-World 2.0, against 0.230 for Memorizon and 0.220 for CaR; it also states plainly that Memorizon does not lead the VBench columns and reports a camera-following check to rule out the objection that the margin measures camera control rather than memory.

Perspective

The recipe targets streaming world models dominated by camera motion in static scenes: the paper states it mainly focuses on static scenes where nothing moves but the camera, so a return is always to a place that should look the same. It applies to training pipelines that use a chunked causal diffusion transformer, can obtain camera poses and intrinsics, and want to supervise long-span revisits at bounded sequence length. The retrieval criterion needs camera poses, and both the pose-distance term and the depth range of frustum overlap carry a scene scale; the paper notes its corpus shares one metric unit scaling and reports a scale-sensitivity measurement. The span is positioned as a coverage setting: once it covers the first visits of returns, more length only adds candidates no chunk chooses.

The paper reports that image quality falls as the span grows and traces it to exposure bias, where training conditions on clean ground truth while inference conditions on the model's own output; how much the distillation closes that gap remains to be measured. Gains from spans beyond about a minute cannot be separated by the paper's scored window, so it instead rolls web photographs out longer and bins returns by interval, noting that returns more than four minutes apart are few per seed and that the last bin is the least certain. Numbering the bank in temporal order rather than giving it one shared index is higher on pixel revisit but not consistently better on feature revisit and is lower on image quality on every split, leaving which index is preferable an open question. The retrieval criterion does not account for occlusion or for motion along the viewing axis, both stated in the appendix. In addition, this evidence bundle is full text without the figures themselves, so details tied to specific figure values are taken from the prose.

Sources