FlashForward reuses in-flight KV cache with sparse clean anchors to speed 20s+ video generation by 1.16–1.69x on four GPUs while reaching 0.838 VBench Total
Synopsis
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
Interpretation
The paper reframes cross-chunk memory as a state-availability problem: a cross-chunk state contract specifies what representation each chunk publishes, at what noise level, and when later chunks may consume it, and on that basis FlashForward lets every ordinary renderer forward publish same-stage KV to later chunks while advancing the current chunk's output. Previously Self-Forcing needed an extra cache-update-only forward after a chunk completed to build clean KV, and HiAR re-encoded each consumed predecessor at every denoising stage; FlashForward is, in the paper's Tab. 1, the only memory both emitted by an output-advancing forward and available early enough to pipeline chunks across stage workers. The paper gives a formal definition and the mechanism comparison in Tab. 1, and measures latency across 1.3B and 14B backbones at 480p and 720p with up to four GPUs, reporting 1.16–1.69x over HiAR and 1.42–2.92x over Self-Forcing from 20 seconds onward.
Because same-stage history comes from unfinished noisy chunks and used alone causes appearance and motion drift, the paper builds clean anchor KV from sparse auxiliary anchor latents generated in advance, giving two-sided conditioning: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance while dense stage-matched history preserves fine, recent evolution. The paper repurposes the established sparse-planning and two-sided-conditioning idea as a coarse-timescale state complementary to fine-timescale stage-matched renderer history; ablations show that stage-matched history alone flickers so heavily that VBench cannot score its temporal flickering dimension, while clean history and future alone degenerates into a high-latency Self-Forcing variant, and combining them yields 0.838 Total. Tab. 3 bottom compares three conditioning modules under the same setup, with the final design scoring higher than both single-module variants on Total 0.838, Quality 0.859, and Semantic 0.753.
The paper realizes planner and renderer roles with a shared generator backbone plus role embeddings and role-specific LoRA adapters, trained in two phases: packed supervised fine-tuning establishes the planner–renderer graph from real videos, and self-rollout distillation distills guidance and denoising stages and adapts few-step generation to generated context. The supervised phase learns directly from real videos rather than teacher trajectories, so it can use clips of arbitrary duration; the paper uses 20-second clips containing nine anchors and exactly three anchor blocks to explicitly train the planner's autoregressive dependence, while the distillation phase uses full-clip self-rollout with a distribution-matching objective to mitigate exposure bias. Tab. 3 top ablations show time-rebased supervision gives a better initialization for distillation, and role-specific embedding and LoRA raise the distillation-phase Total from 0.8184 to 0.8380 while reducing seams between anchor and non-anchor positions.
The paper reports that the renderer regenerates anchor positions rather than splicing planner anchors into the final video, because splice designs produce periodic seams at anchor positions; using planner anchors only as conditioning and letting the renderer generate every output position substantially reduces the seam. Panels (a), (b), and (c) separate the effects of the anchor-KV source and the final anchor source, and show that the seam remains even when generated planner anchors are replaced with anchors encoded from real videos, indicating an output-interface mismatch rather than planner generation error alone. At 50 steps, switching the final anchor from planner to renderer cuts the anchor seam from 3.93 to 1.83 and the periodic total from 7.66 to 3.54; in the four-step deployment, the anchor seam falls from 6.20 to 2.43 and the periodic total from 9.55 to 5.61.
Perspective
The paper targets long-sequence, few-step generation on multiple GPUs: with up to four GPUs, 16 FPS, 20 seconds or longer, 1.3B and 14B backbones, and 480p and 720p, FlashForward is the fastest evaluated autoregressive schedule and holds stable quality at 20s, 35s, and 65s for the 1.3B model at 480p. It naturally supports streaming generation by dedicating one GPU to the planner and the remaining GPUs to the four-stage renderer wavefront, so rendering can begin as soon as the first anchor block is published; in the five-GPU configuration the 1.3B model gains further speed from 20 to 65 seconds, and for 14B across two GB300 nodes most anchor generation stays hidden behind rendering. The paper also provides a forward-count model and two break-even conditions for when the forward reduction becomes a wall-clock speedup, and notes that for small forwards the planner should run on one rank rather than with context parallelism.
The 14B rows are systems-only timing measurements that instantiate the same tensor shapes, masks, cache paths, and execution schedules with the 14B backbone, so the quality conclusions come mainly from the 1.3B 480p VBench evaluation. In the quality comparison, competing methods' results are taken from HiAR, while VBench-Long is run by the paper for all four methods under the same protocol. The paper uses a fixed anchor stride, conditioning window, and a vanilla training recipe except for the time-rebased augmentation, and it lists relaxing the current clean anchor-state assumption, adapting anchor selection, combining with KV selection and compression, RoPE modifications, and anti-drift techniques, and parallel generation of anchor-delimited intervals as future work. In addition, training data consists only of real-world videos, which may bias the model toward photorealistic content, leaving performance on animation and other stylized domains an open question.
