Skip to main content
Back to timeline
arXivSource publication:

LIFT controls video generation with a last-frame layout plus camera trajectory, raising mIoU from 0.41 to 0.51

Synopsis

LIFT introduces a unified image-to-video framework that adds Layout-In-FuTure control on top of camera control, letting users specify what should appear in a future view and where; it transfers a dense per-frame layout teacher to a last-frame-layout student via dual-mode on-policy self-distillation (OPSD) and curates the LIFT-Vista dataset for large viewpoint changes, reporting improvements in video quality, future-layout controllability, and camera controllability over compared methods.

AI-generated editorial illustration: LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Interpretation

LIFT uses the last-frame layout as an explicit control signal jointly conditioned with the camera trajectory, so the model can control the content and spatial arrangement of newly revealed regions under large viewpoint changes. Existing camera-control methods only specify how the viewpoint moves and text prompts give only coarse semantic guidance, while existing video layout methods typically rely on dense per-frame boxes, masks, or trajectories and mainly control objects already visible in the first frame; LIFT needs only a last-frame layout to specify the semantics and composition of future views. The paper presents a unified framework with two inference modes (a single-condition mode using only the camera trajectory and a dual-condition mode adding the last-frame layout) and compares against camera-control, object-motion-control, and joint-control baselines.

The paper proposes dual-mode on-policy self-distillation (OPSD), using dense spatiotemporal layouts as privileged information to train a shared student across both the single-condition and dual-condition modes. The authors report that directly learning last-frame-only layouts with standard supervised flow matching performs poorly, and that progressively reducing layout density through SFT incurs substantial training cost with limited gains; OPSD instead supervises the student on its own rollout states with a dense-layout teacher and concentrates supervision on the first 10 high-noise states of each rollout. Ablations show OPSD surpasses SFT baselines in 5 of 8 metrics using about 136K sample updates (steps times batch size) versus 256K for SFT; removing the SFT anchor loss raises FVD from 99.35 to 129.14, indicating the anchor term helps maintain generation quality.

The paper curates LIFT-Vista, a dataset specifically for future-view layout control under large viewpoint changes, with camera trajectories and temporally consistent layout annotations. Existing camera-annotated video datasets generally lack object-level layout labels, while datasets with boxes or layouts typically lack camera trajectories and focus on first-frame objects; LIFT-Vista provides both and emphasizes clips where camera motion reveals regions outside the initial view. Data is built from RealEstate10K, Sekai, and SpatialVID and filtered by FoV expansion ratio, accumulated translation, and a DINOv2-based content change ratio; the result is 120,898 Stage 1 training samples, 58,272 Stage 2/3 samples, and 600 test samples, with layout annotations averaging 5.5 objects per clip.

In the paper's comparisons, LIFT supports both camera control and future-layout control and achieves the best layout-control results. Compared with MagicMotion, which is additionally provided with dense per-frame bounding-box trajectories, LIFT uses only a last-frame layout and improves mIoU from 0.41 to 0.51; the paper reports that LIFT (1.3B) is competitive in video quality with 14B Uni3C and 7B GEN3C and achieves the best camera-control accuracy among the compared methods. Quantitative results use FVD, FID, LPIPS, RotErr, TransErr, mIoU, SRe, and CLIPlocal; a randomized user study with 7 participants, 30 comparisons each and 210 total responses gives LIFT a 65.96% preference rate.

Perspective

The work targets the setting of large viewpoint changes in image-to-video generation: given a first-frame image, a text caption, a target camera trajectory, and (in dual-condition mode) a last-frame object layout, it generates 81-frame clips at 16 FPS. It suits creative scenarios where camera motion reveals regions outside the initial view and the user cares about the final-frame composition, for example turning from one room toward another while specifying where newly appearing furniture should be. The paper also shows LIFT supports different layout sparsity: providing more layout frames at inference generally improves generation quality, spatial alignment, and semantic consistency, so users can trade annotation effort for finer spatial control while retaining the practical last-frame-only interface.

The paper states that it currently uses 2D bounding boxes with local text prompts to specify the desired future-view composition, which provides only coarse spatial constraints and does not explicitly capture depth, orientation, or occlusion relationships, so finer-grained geometric or instance-level control remains an open direction. In addition, the benefit of OPSD is presented mainly through comparison with SFT baselines, and its behavior on other base models, resolutions, or longer videos remains to be examined; the user study involves 7 participants and 210 responses, which is small-sample preference evidence. In the data-filtering ablation, both the filtered and random-sampling sets contain 30K clips, and the filtered set is better on FVD, FID, RotErr, and TransErr, though the margins are limited.

Sources