Skip to main content
Back to timeline
arXivSource publication:

DuoMatching adds joint-marginal distribution matching to few-step video generation, with over 80% overall human preference against every baseline

Synopsis

DuoMatching introduces a unified joint-marginal distribution matching framework: beyond existing joint DMD, it applies direct frame-level marginal supervision from an image generator, uses LatentBridge to resolve the video-image latent mismatch, and uses Latent Variation Sampling to spread frame-level supervision across temporal segments; in causal and bidirectional few-step video generation it improves visual quality, composition and semantic alignment while largely preserving motion dynamics, with human overall preference above 80% against all evaluated baselines.

AI-generated editorial illustration: DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

Interpretation

The paper extends distribution matching for few-step video generation from joint matching alone to a joint-marginal formulation: the video teacher keeps constraining cross-frame dependencies while an image teacher directly supervises the marginal distribution of sampled frames, transferring fine-grained appearance and semantic priors to the video student. Existing DMD uses the video teacher as a proxy for the real video distribution and matches the joint distribution of all frames, with no dedicated objective for individual frames; the paper reports that this joint matching substantially mitigates drift over the rollout but yields only limited improvement in initial-frame quality, illustrating frame-level deficiencies such as insufficient fur detail and a missing feather. Appendices C and D provide idealized analyses: under matched marginal capacity, marginal-only fitting achieves no greater forward-KL marginal error than joint fitting, and there exists a positive marginal weight for which the ideal joint-marginal optimum is closer to the real distribution in forward KL; main experiments report results on VBench and a curated set of 400 challenging prompts.

LatentBridge is a lightweight differentiable module that maps temporally compressed video latents to frame-specific representations in the image teacher's latent space, enabling marginal DMD on the video student without decode-and-re-encode. Applying marginal DMD directly to temporally compressed video latent slices improves semantic alignment and visual quality but suppresses the dynamic information encoded in the latents, while the decode-encode route runs out of memory under the training configuration. Ablation shows LatentBridge yields further gains in visual quality and semantic alignment while largely retaining video dynamics, with only a small additional memory overhead over the direct baseline; Appendix E describes eight FiLM residual blocks with 256 hidden channels and 10.75M trainable parameters, trained on 29,400 videos at 480p for 3,000 iterations in about 15 minutes on eight 80-GB GPUs.

Latent Variation Sampling computes the mean squared difference between adjacent latent slices, selects the positions with the largest values, splits the sequence into contiguous temporal segments, and samples one slice per segment for marginal DMD, distributing the limited frame-level supervision budget across temporal regions. Uniform sampling can concentrate on slowly changing regions, producing redundant supervision and potentially suppressing motion dynamics by overemphasizing nearly static content. Ablation shows that increasing the number of sampled latent slices from a smaller to a moderate value improves the Total score for both uniform sampling and LVS, while further increasing it reduces Dynamic and Total scores; LVS outperforms uniform sampling across all tested numbers of slices and temporally stratified sampling at the same setting.

Across causal and bidirectional few-step generation settings, DuoMatching achieves higher VBench total scores within matched generation modes and NFE budgets, without additional inference computation. The causal setting covers one-, two- and four-step chunk-wise and frame-wise generation, and the bidirectional setting uses the bidirectional DMD variant of CausVid as the primary baseline; gains concentrate in visual quality and semantic alignment while motion quality is largely preserved. Table 1 reports Semantic, Aesthetic, Imaging, Dynamic, Smoothness and Total scores per mode and NFE; human evaluation follows a two-alternative forced-choice protocol with 23 participants each evaluating 40 prompt-matched video pairs, showing overall preference rates above 80% against every baseline, higher preference for visual quality and semantic alignment, and temporal and motion preference near or above parity.

Perspective

The framework targets few-step video generation, especially autoregressive/causal streaming generation, and also extends to bidirectional generation; its gains come from frame-level priors supplied by an image teacher, so it suits teams that can obtain a strong image generation model and accept a one-time LatentBridge pretraining cost. The default setting uses Qwen-Image as the image teacher, produces videos at a specific frame count and resolution, and trains on eight 80-GB GPUs; the marginal weight and the number of sampled latent slices have ranges that best balance visual quality and dynamics.

The joint-marginal improvement analysis rests on idealized assumptions, including matched marginal capacity, positive densities and differentiability, so behavior when training departs from these assumptions remains to be observed. Human evaluation is a 2AFC result from 23 participants each judging 40 video pairs, and confidence intervals for the preference rates are not given in the main text. LatentBridge reconstruction quality and loss-transfer analysis appear in the appendices, and how well they generalize across different video VAE and image teacher combinations remains an open question. In addition, a too-large marginal weight markedly lowers the Dynamic score, indicating that the quality-motion balance is sensitive to hyperparameters.

Sources