Skip to main content
Back to timeline
arXivSource publication:

SGF+ separates context-writing and denoising parameters in autoregressive video generation, improving long-horizon consistency on VBench and enabling 24-hour continuous generation

Synopsis

The authors find that context writing and denoising in autoregressive video diffusion exhibit systematically conflicting gradients when they share parameters, and introduce SGF+ to assign separate parameters to the two roles while jointly optimizing them under the original generation objective; across VBench framewise and chunkwise generation at 5s, 60s, and 240s, SGF+ improves most subject consistency, background consistency, flickering, motion smoothness, aesthetics, and imaging metrics over SF and SGF, and supports up to 24 hours of continuous generation from 5s training rollouts.

AI-generated editorial illustration: SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Interpretation

Isolating the gradient contributions of context writing and denoising within SGF's second differentiable reconstruction pass, the authors find strongly misaligned directions: angular-distance t-SNE shows separated gradient distributions in both Attention and FFN, mean gradient directions differ by more than 90 degrees in both module families, and all 512 measured paired cosine similarities are negative. SGF had already closed the historical context-gradient gap via differentiable KV reconstruction, but context writing and denoising still shared one set of DiT parameters; this work identifies parameter sharing itself as an optimization bottleneck and quantifies role-level gradient conflict. Based on gradient decomposition and cosine-similarity measurement within the same Pass-2 computation, covering Attention and FFN across multiple prompts and denoising timesteps, with layer-wise distributions in Appendix E.

SGF+ assigns independent parameters to context writing and denoising (a context writer and a denoiser) that remain forward-coupled through causal attention and are jointly trained with the original generation objective alone, with context writing supervised through its contribution to future predictions, requiring no auxiliary losses, additional video training data, or long-video fine-tuning. Unlike SGF, SGF+ does not alter history construction or context management; it intervenes on the parameter-sharing design dimension, making it mechanistically complementary to history-construction and context-management approaches. Training retains SGF's self-generated rollouts, differentiable replay, and short training window; ablations show that separating parameters only in Attention or only in FFN underperforms the full model, indicating conflict in both module families.

At 60s and 240s generation horizons in both framewise and chunkwise generation on VBench, SGF+ outperforms SF and SGF on most of subject consistency, background consistency, temporal flickering, motion smoothness, aesthetic quality, and imaging quality; Dynamic Degree is the main exception, which the authors explain by noting that scene jumps, object deformation, subject disappearance, or additional people appearing can produce large but incoherent apparent motion that inflates that metric. Prior long-video methods often rely on longer training horizons, long-video fine-tuning, or inference-time context management; SGF+ obtains long-horizon extrapolation gains within a 5s training window. Comparisons share the same sink-plus-FIFO context policy, with 60s evaluated on VBench-Long prompts and 240s on 128 MovieGen prompts, scores multiplied by 100; 5s results appear in Appendix A.

On efficiency, SGF+ doubles generator parameters from 1.4B to 2.8B while preserving SGF's denoising and context-writing call schedule, and each token uses only its role-specific parameters within a forward pass, so computation does not double: inference time for 81 frames is 4.969 s versus 4.962 s for SGF, inference memory rises from 24.85 GB to 27.96 GB, and training time per step increases by about 8.2%. This shows the cost of role-specific parameterization is mainly parameter storage and modest training overhead rather than inference latency. Reports measured training peak and stable memory, training time per step, inference memory, and inference time for 81 frames.

Perspective

The results target streaming, long-horizon generation with autoregressive video diffusion: training on 5s video windows, evaluation at 5s, 60s, and 240s generation horizons in framewise and chunkwise modes, and demonstrations of up to 24 hours of continuous generation. The method applies to pipelines that retain SGF's self-generated rollouts and differentiable context reconstruction, and the authors position it as mechanistically complementary to history construction, context management, spectral correction, positional adaptation, and longer-horizon training. The authors also propose extending role-specific parameterization to teacher-forced world action models and to asymmetric architectures for KV compression with a lightweight context writer, which remain future work.

The gradient-conflict statistics rest on 512 measured pairs across the tested prompts and denoising timesteps, and the local optimization analysis in Appendix G indicates that minibatch comparisons require negative alignment of aggregated gradients, so negative alignment of individual samples or selected modules does not establish that condition. The interpretation that lower Dynamic Degree reflects scene jumps and deformation inflating the metric relies on the authors' qualitative comparisons, which readers can weigh alongside the reported examples. The 24-hour continuous generation is presented through examples, and its stability across different prompts and scenes would benefit from further observation. In addition, this material is a full-text parse, and some appendix figures and tables are not fully expanded in the text.

Sources