Skip to main content
Back to timeline
arXivSource publication:

Latent-Foresight jointly trains a tokenizer and a flow-matching predictor, consistently beating two-stage baselines on future scene understanding in Cityscapes and nuScenes

Synopsis

The work proposes Latent-Foresight, an end-to-end framework that jointly learns a VFM feature tokenizer and a flow-based generative dynamics model, using stop-gradient, latent normalization, a denoised-latent reconstruction loss and a noise schedule to prevent latent collapse, and it consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons on Cityscapes, nuScenes and Kubric while eliminating separate training stages, including during high-resolution adaptation.

AI-generated editorial illustration: Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

Interpretation

The authors propose Latent-Foresight, which optimizes a VFM feature tokenizer and a flow-matching predictor under a single end-to-end objective, so the latent space is shaped directly by the requirements of generative temporal prediction rather than compressed first and then frozen. Prior VFM world models followed a two-stage route: compress features with PCA or an independently trained autoencoder, then train a predictor on the frozen latent space; the authors describe this as the first approach to enable end-to-end training of both the tokenizer and a generative latent dynamics model in VFM-based world modeling. The paper gives a full method description and ablations: DINOv2-Reg ViT-B/14 intermediate-layer concatenated features by default, bottleneck dimension 256, a 12-layer predictor, evaluated on Cityscapes, nuScenes and Kubric under established protocols.

The authors identify that naive joint optimization collapses latent representations to trivial constants, and provide the key designs that make training stable: stopping gradients through the velocity target and the noisy input, normalizing latents with BatchNorm without affine parameters, adding an auxiliary reconstruction loss on denoised latents, and selecting the noise level distribution. These designs turn coupling representation learning with generative dynamics from infeasible into feasible and clarify each component's role: removing latent normalization or the stop-gradient on the target latent causes training collapse, while removing the stop-gradient on the noisy latent keeps training stable but degrades forecasting. The training recipe ablation table reports reconstruction cosine similarity and short/mid-term segmentation metrics item by item, with collapse cases explicitly marked as Training Collapse; the noise distribution ablation shows shifting the logit-normal toward higher noise levels beats uniform sampling and the JiT default.

End-to-end training consistently outperforms two-stage baselines across datasets and prediction horizons, with gains more pronounced at longer horizons and on movable-object classes. Compared with the two-stage variant that trains the autoencoder independently, joint training improves forecasting while preserving reconstruction fidelity; on Cityscapes, Latent-Foresight reaches mid-term segmentation ALL 61.7 and MO 59.9, above DINO-Foresight's 59.8 and 57.6 and above DeltaTok's 60.0; on nuScenes the gains extend to 27-frame, 2.25-second long-horizon autoregressive prediction. The comparison spans Cityscapes, nuScenes and Kubric, includes an Oracle upper bound, Copy Last, VISTA, DINO-Foresight, DeltaTok, VFMF and a two-stage autoencoder baseline, and reports zero-shot transfer results.

The end-to-end formulation simplifies the training pipeline, including high-resolution adaptation, and shows scalability. Two-stage pipelines require separately fine-tuning the autoencoder and then the predictor, whereas this method needs a single training stage; at high resolution the two-stage autoencoder reaches mid-term segmentation ALL 60.18 and MO 58.05 versus 61.74 and 59.92 for end-to-end; Latent-Foresight+, trained on more data with a longer schedule, further improves Cityscapes mid-term to ALL 63.3 and MO 61.9. The high-resolution comparison is given on Cityscapes and nuScenes, Latent-Foresight+ is trained on combined Cityscapes, nuScenes and CoVLA data, and on Kubric a single generation exceeds VFMF's 32-generation setting.

Perspective

The result targets latent world models that predict in VFM feature space, suited to settings such as autonomous driving and robotics that must anticipate scene evolution, with evaluation on Cityscapes, nuScenes and the synthetic Kubric benchmark, including zero-shot transfer from Cityscapes to nuScenes. It enables joint optimization of tokenizer and predictor in a single training stage and removes the need to fine-tune the two components separately for high-resolution adaptation; the authors note the autoencoder architecture is kept lightweight, point to more expressive tokenizer architectures and strategies for reducing spatial tokens, and suggest extending the current next-frame prediction to generating multiple frames in a single forward pass.

The authors state in the appendix that the autoencoder architecture remains lightweight, leaving stronger tokenizers and spatial-token reduction to future work; the current framework focuses on next-frame prediction, with multi-frame joint modeling not yet explored. The method builds on pretrained VFMs, and the authors caution that biases embedded in those models may propagate into predicted semantic representations, requiring evaluation before deployment. Ablations show a trade-off between reconstruction fidelity and temporal predictability in bottleneck dimension and noise distribution, so the best setting depends on the task and data. A reader seeing only the abstract may not be able to judge each design component's individual contribution and would need the training-recipe ablation table.

Sources