Skip to main content
Back to timeline
arXivSource publication:

SVEET trains only on a bidirectional video diffusion model and transfers editing zero-shot to a streaming backbone at 15 FPS on a single H100

Synopsis

The work proposes SVEET, a framework that trains a control branch only on a pretrained bidirectional video diffusion model and, through a temporally independent 2D spatial-attention control branch plus Orthogonal Decoupled Training (ODT) that bridges the bidirectional-to-autoregressive feature gap, enables high-quality streaming video editing without retraining or distilling the streaming backbone, reaching 15 FPS on a single H100 GPU and outperforming several baselines on style transfer, video inpainting, and depth-to-video generation.

AI-generated editorial illustration: Streaming Video Editing with Easy Adaptation

Interpretation

The authors systematically revisit existing video-to-video diffusion architectures and identify two principles for streaming adaptation: backbone feature disentanglement, where the control mechanism stays decoupled from the base model to preserve pretrained knowledge, and conditional frame independence, where source-video encoding is per-frame without temporal dependency to remain causally compatible. DiT-based streaming video editing was largely unexplored, and existing editing methods rely on bidirectional full-sequence processing that conflicts with streaming constraints; this work turns the transfer question from engineering trial and error into two testable design principles. Comparative experiments on three representative architectures: Wan-Fun, which fine-tunes backbone parameters; VACE with full spatiotemporal attention; and VACE with temporally independent 2D attention. Wan-Fun degrades severely after transfer, full spatiotemporal VACE retains only partial editing capability with unstable temporal behavior, while the 2D-attention variant transfers substantially more reliably.

SVEET replaces the original full spatiotemporal self-attention in the VACE control branch with temporally independent 2D spatial attention, so each frame's conditioning representation depends only on that frame while temporal coherence is modeled by the causal streaming backbone through its autoregressive history, meaning the control branch needs no extra KV cache at streaming inference. Compared with the Full-Attn baseline that keeps full spatiotemporal attention, this change makes the conditioning pathway compatible with causal chunk-wise inference and avoids adding cache-growth overhead on top of the autoregressive backbone. Ablations show 2D attention improves source preservation and streaming stability over 3D attention (VLM editing accuracy 7.1898 vs 7.0653; motion smoothness 0.9830 vs 0.9809), and combining it with ODT gives the best overall trade-off (7.4322, 0.9898).

The authors propose Orthogonal Decoupled Training (ODT): ridge regression estimates a per-block linear map between bidirectional and causal hidden states, a thin SVD of the residual transformation yields the top right singular vectors spanning the discrepancy subspace, and the control branch's learnable update (LoRA) is projected onto its orthogonal complement, with the rank chosen adaptively per layer by cumulative spectral energy. The scheme explicitly decouples the optimization directions of video controllability and model causality to mitigate transfer degradation from misaligned bidirectional and autoregressive feature spaces, and it comes with theoretical bounds on feature-level decoupling and output additivity. Appendix A reports that discrepancy energy concentrates in a compact set of singular directions, with a small fraction of directions explaining most spectral energy; ablations show 3D with ODT beats 3D without ODT (7.3315 vs 7.0653), adaptive rank beats fixed rank, and two alternative transfer strategies, inference-stage projection and two-stage teacher forcing, perform worse.

Across style transfer, video inpainting, and depth-to-video generation, SVEET achieves the best average performance on most metrics, leads all baselines in VLM editing accuracy, and runs at 15 FPS on a single H100 GPU without auxiliary acceleration techniques. Open-source real-time video editing models remain scarce, and recent work reports streaming distillation can take about 128 H100 GPU days with thousands of synthesized ODE pairs; this work demonstrates a lightweight path that avoids retraining the streaming backbone. Evaluation uses GPT-4o 1-10 VLM assessment, CLIP-T text alignment, and six VBench metrics, on test sets of 120 video pairs for style transfer and 80 samples each for inpainting and depth-to-video; a user study with 10 participants scoring 15 generated videos preferred SVEET on editing correctness 9.0800, structural preservation 9.2067, and overall smoothness 8.1133.

Perspective

The result targets settings that need video editing under causal streaming, such as live style transfer and online inpainting, for researchers and engineering teams who want to reuse an existing bidirectional editing model without retraining the streaming backbone. The method assumes a Wan2.1-1.3B-VACE bidirectional backbone and chunk-wise Causal Forcing as the default streaming backbone, trains the control branch with rank-128 LoRA while all pretrained backbone parameters stay frozen, trains each task for 10 epochs on a single A100 80GB with batch size 1 on 81-frame clips at a given spatial resolution, and trains separate task-specific control branches for the three tasks. The authors note the method still relies on the capability of the underlying bidirectional editing model, and extending the transfer paradigm to broader editing tasks and more heterogeneous backbones is left as future work.

The authors state the method still relies on the capability of the underlying bidirectional editing model, and transfer to broader editing tasks and more heterogeneous backbones remains to be verified. ODT depends on the assumption that the bidirectional-to-causal discrepancy concentrates in a compact subspace, supported in Appendix A by spectral energy concentration, but whether that holds for other backbones or data distributions is an open question. Evaluation covers three tasks with limited test-set sizes, and the user study involves 10 participants and 15 videos, so longer-horizon generation and more complex editing instructions need more evidence. In addition, this is a full-text read, but some figures are referenced by number without being expanded in the text, so specific visual details require checking the original figures.

Sources