Skip to main content
Back to timeline
arXivSource publication:

First survey of post-training and alignment for video generation unifies four method families and reports dimension-wise cross-benchmark evidence

Synopsis

This TMLR survey presents the first comprehensive review of post-training and alignment for video generation models, distinguishing implicit from explicit alignment by how alignment signals are enforced and organizing methods into supervised fine-tuning, self-training and distillation, preference- and reward-based methods, and inference-time methods, while reviewing datasets, benchmarks, and evaluation practices and discussing open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation.

AI-generated editorial illustration: Video Generation Models: A Survey of Post-Training and Alignment

Interpretation

The survey proposes a unifying framework organized around how alignment signals are enforced, defining post-training as any optimization, adaptation, or control procedure applied after large-scale pretraining that modifies model behavior without retraining from scratch, and distinguishing implicit from explicit alignment. Earlier surveys organize the field by architecture or conditioning signal, such as broad overviews of video diffusion models, controllable video generation, and human-centric video generation; this survey instead organizes by how alignment is enforced, which the authors describe as the first comprehensive review of post-training and alignment for video generation. This is a conceptual framing contribution grounded in literature classification and definitions rather than experimental measurement; the text explicitly states that alignment exists on a spectrum rather than forming a strict binary.

The survey groups methods into four categories: supervised fine-tuning, self-training and distillation, preference- and reward-based methods, and inference-time methods, and notes that these families are not mutually exclusive, with modern systems often composing several stages sequentially. It documents recurring cross-family composition patterns, such as supervised fine-tuning followed by preference optimization, distillation followed by reward-based fine-tuning, and a larger pipeline combining supervised fine-tuning, video-specific reinforcement learning, and multi-stage distillation. The classification rests on the survey's organization of a large body of representative methods, supported by summary tables listing each method's sub-category, number of stages, base model, GPU scale, venue, and year.

The survey systematically reviews post-training datasets, benchmarks organized by alignment dimension, and evaluation protocols, and summarizes reported results on VBench, VBench2, VideoPhy, and VideoPhy2. The authors stress that these numbers are benchmark-wise evidence under heterogeneous protocols rather than a unified leaderboard, noting that methods reporting under the same benchmark name may use different prompt sets, submetrics, or evaluation subsets. The tables cover reported scores across dimensions such as temporal consistency, visual quality, subject consistency, controllability, and physical plausibility for supervised fine-tuning, self-training and distillation, preference- and reward-based, and inference-time methods, with missing entries indicating scores not reported in the corresponding papers.

The survey observes that post-training often optimizes specific behavioral dimensions, so gains in one aspect of alignment may expose or even amplify weaknesses in another, for example high temporal smoothness does not necessarily imply strong motion generation. This diagnostic perspective links benchmark decomposition to post-training objectives, explaining why targeted benchmarks are more informative than broad ones for judging whether a method achieves its intended alignment goal. The observation draws on the dimension decomposition in VBench and VBench2, and on the design of VideoPhy and VideoPhy2, which separate semantic adherence from physics consistency.

Perspective

The survey targets researchers and practitioners who want a systematic understanding of post-training and alignment for video generation, and it applies to settings that adapt pretrained video generation models without retraining from scratch, including controllable generation, personalization, domain specialization, efficiency optimization, and safety-aware generation. It also maintains a companion repository for continuously updated papers and resources, serving as a starting point for entering the field and a reference for developing new methods. The aggregated benchmark results are intended for dimension-wise diagnostic comparison rather than for building a unified cross-method leaderboard.

As a survey, its conclusions rest on literature classification and reported results, and the authors themselves caution that cross-benchmark numbers are not directly comparable because methods differ in task formulation, base model, model size, sampling budget, resolution, video length, and optimization objective. Readers should still watch how open problems such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation develop in subsequent work; in addition, the text read here is truncated at the reference list, so some entries and table details are not fully visible, and specific methods or numbers should be verified against the original paper and companion repository.

Sources