Public articles linked to the same research event.
arXiv The work proposes Spatial Memory Intelligence (SMI), described as the first framework to systematically employ an understanding model (a multimodal large language model) for spatial-memory management in long-video world models, through four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering; experiments across multiple baselines, benchmarks, and world-model backbones report comprehensive improvements in memory sparsity, generation stability, and spatial consistency, supporting effectiveness and generalizability.
The work proposes Spatial Memory Intelligence (SMI), described as the first framework to systematically employ an understanding model (a multimodal large language model) for spatial-memory management in long-video world models, through four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering; experiments across multiple baselines, benchmarks, and world-model backbones report comprehensive improvements in memory sparsity, generation stability, and spatial consistency, supporting effectiveness and generalizability.
The work proposes Spatial Memory Intelligence (SMI), described as the first framework to systematically employ an understanding model (a multimodal large language model) for spatial-memory management in long-video world models, through four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering; experiments across multiple baselines, benchmarks, and world-model backbones report comprehensive improvements in memory sparsity, generation stability, and spatial consistency, supporting effectiveness and generalizability.
The work proposes Spatial Memory Intelligence (SMI), described as the first framework to systematically employ an understanding model (a multimodal large language model) for spatial-memory management in long-video world models, through four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering; experiments across multiple baselines, benchmarks, and world-model backbones report comprehensive improvements in memory sparsity, generation stability, and spatial consistency, supporting effectiveness and generalizability.