Honeycomb stores scene memory in six fixed-size feature planes, lifting revisit consistency and novel-view synthesis on WorldScore and RealEstate10K
Synopsis
The work introduces Honeycomb, a video world model built on HexMemory, a low-rank representation that stores scene features in a fixed-size memory of three spatial plus three spatiotemporal planes; a feed-forward writer maps each generated chunk into new plane features, warps the previous planes while preserving their dimensions, and fuses them via confidence-weighted pooling plus a learned residual correction, while a reader retrieves latents to condition subsequent generation; on WorldScore and RealEstate10K it reports stronger generation quality and revisit consistency with feature storage held constant throughout generation.
Interpretation
HexMemory represents persistent scene memory as six fixed-size 2D feature planes, three spatial and three spatiotemporal, with each pair of orthogonal axes covering all four coordinates exactly once, plus confidence maps recording accumulated interpolation weights and a bounding box over world space and write time. Existing spatial memories such as Spatia's updatable point cloud of RGB observations and LSM-World's diffusion latents tied to 3D points grow as observations accumulate; this replaces them with a low-rank plane factorization whose dimensions stay fixed throughout generation. The paper gives the plane sampling formulation (bilinear interpolation, spatial planes sharing information across write times, products with spatiotemporal planes capturing time-varying content) and notes that HexPlane originally optimizes planes per scene whereas HexMemory uses a writer shared across scenes without per-scene optimization.
A feed-forward writer maps each new chunk's latent observations into plane features; the recurrent update expands the bounds as needed, warps previous planes into the new bounds by bilinear interpolation while preserving grid dimensions, and fuses them with the new features through confidence-weighted pooling and a zero-initialized learned residual correction. Compared with per-scene direct optimization (the original HexPlane approach) and a replacement writer that rebuilds planes from all stored points each chunk, recurrent writing processes only the new chunk and avoids reprocessing the full history. Ablations show recurrent writing holds write time at about 13 ms for chunks 2, 5, and 9, while the replacement writer grows from 11.2 ms to 40.1 ms and direct optimization takes about 3217 ms (described as roughly 240x slower); WorldScore closed-loop PSNR is 17.22, 17.16, and 17.44 dB respectively.
On WorldScore closed-loop revisit evaluation, Honeycomb achieves the best results across all metrics among the compared methods, improving PSNR by 1.23 dB over the next-best method and reducing flow error from 6.64 to 3.00 pixels relative to Spatia; on RealEstate10K novel-view synthesis it reaches 18.45 dB PSNR versus 15.58 dB for Spatia and 17.46 dB for LSM-World. These results are obtained with feature memory fixed in size and without excluding dynamic objects or sky from memory writes, whereas the compared methods rely on spatial caches that grow with observations. WorldScore contains 3,000 image-to-video samples, and the closed-loop evaluation starts from input images of 100 scenes and generates along trajectories that leave and return to the initial viewpoint; RealEstate10K evaluation uses 100 test-set videos and reports PSNR, SSIM, and LPIPS.
Ablations on memory resolution and write scheme quantify the trade-off among storage, write cost, and quality: lowering plane resolution from 512 to 256 cuts HexMemory from 73.9 MB to 19.8 MB (described as 73% less) with only a 0.12 dB drop in WorldScore closed-loop PSNR; an appendix retention probe on a 129-frame sequence shows chunk-1 reconstruction falling from 16.71 dB to 16.30 dB across successive writes, a total decline of 0.41 dB over three additional writes and still 4.01 dB above the 12.29 dB memory-disabled baseline. This provides a tunable capacity knob for fixed-size memory and quantitative evidence that earlier content remains recoverable as new observations are incorporated. Ablations run on WorldScore closed-loop samples with resolution settings of 512/384/256/128; the retention probe fixes the input frame, camera trajectory, and prompt, and excludes the fixed input frame when computing accuracy.
Perspective
The result targets camera-trajectory-driven long-horizon video world model generation, in settings where the camera leaves and returns to the same viewpoint and layout and appearance must stay consistent; memory capacity is tunable through plane resolution, and at 256 resolution 19.8 MB of storage sustains revisit quality close to the default configuration, which is directly relevant to storage-constrained deployment. The method relies on estimated depth and camera poses for world-space backprojection and writing, and builds on a pretrained camera-controllable video diffusion backbone, which defines its current applicable setting.
The text presents the abstract, method, experiments, and appendix, with figures described in prose, so the specific details of plane visualizations and qualitative comparisons still need to be confirmed against the original figures. The retention probe is reported on a single 129-frame sequence, so how early content decays over longer rollouts and more writes remains an open question. The resolution ablation shows that further reductions cause more noticeable degradation, leaving the relationship between the capacity floor and scene complexity to be characterized. Evaluation is concentrated on WorldScore and RealEstate10K, and how the design choice of writing dynamic objects and sky into memory behaves in more complex dynamic scenes, as well as when combined with other backbones, are directions worth continued observation.
