Skip to main content
Back to timeline
Amazon ScienceSource publication:

A kernel-centric path to real-time video generation on Trainium

Synopsis

Using Rolling Forcing as a proxy for autoregressive diffusion video generation, Reactor and the Amazon Neuron Science team adopted a kernel-centric, bottom-up approach—writing hardware-level kernels via the Neuron Kernel Interface, applying hybrid sequence and tensor parallelism sharding, and rewriting model code to batch shared components of the diffusion and cache-update phases—successfully generating a correct video on the first end-to-end run on Trainium and delivering real-time generation above 16 fps.

AI-generated editorial illustration: A kernel-centric path to real-time video generation on Trainium

Interpretation

The team used NKI to write compute kernels running directly on NeuronCore hardware, replacing bottleneck operations that generic compilers struggle to optimize: the 3D-RoPE kernel went from five seconds to 1.8 milliseconds, cache copies from 23 milliseconds to 1.9 milliseconds per layer, and attention transposes were eliminated entirely by fusing them into the attention kernel. Relative to relying on generic compilers, this pushes the three recurring challenges of real-time video generation—dynamic shapes, unusual memory access patterns, and heavy cache management—down to hardware-level kernels, with quantified operator-level gains. The text reports specific operator-level latency figures (five seconds to 1.8 milliseconds, 23 milliseconds to 1.9 milliseconds) and a memory result (the pipeline used 11 GB of high-bandwidth memory while the standard eager-mode path ran out of memory), as an engineering measurement report.

For self-attention (23,400 query tokens attending to 32,760 context tokens, about 70% of compute time), the team used a hybrid sharding strategy of sequence parallelism and tensor parallelism: splitting attention heads across 4 cores and sequences across 2, addressing both the divisibility issue between 12 attention heads and 8 Neuron cores per chip and the frame-granularity constraint of 3D video tokens. Tensor parallelism alone would require padding and waste computation, while sequence parallelism alone can break the 3D structure; the hybrid approach keeps both the math and the data layout correct, and the VAE decoder used spatial W-axis sharding to achieve a super-linear 8.25 times speedup. The text provides specific numbers for attention scale, compute share, sharding dimensions, and the 8.25 times speedup, as engineering measurements for this model structure.

The team optimized the model code so that components shared by the diffusion phase and the cache-update phase were batched together, splitting again only when components differed slightly, to mitigate the low hardware utilization of running the less compute-heavy cache-update phase separately. Relative to executing the two phases as two separate runs, this batching targets Trainium's topology of 16 chips per instance and eight cores per chip, avoiding partitioning computation across too many cores so that each call has too little partitioned computation. The text grounds this in engineers' descriptions of the two phases' compute volume and hardware utilization, without giving a separate speedup figure for this rewrite.

After these optimizations, the team successfully generated a correct video on the first end-to-end run on Trainium, and emphasized that these techniques target all autoregressive diffusion models with real-time streaming requirements rather than optimizing a single model. Relative to tuning a single model, the team used Rolling Forcing as a proxy, arguing it contains the full system components—encoder, DiT, VAE, decoder—so the resulting methods can transfer to more complex systems. This is the team's own account of the first end-to-end success and its generalizability judgment; the text provides no comparison data against other hardware or baselines.

Perspective

This work targets developers and platform providers who need to deploy real-time streaming autoregressive diffusion video models on Trainium; the applicable setting is interactive video generation where generation must stay ahead of the playback timeline and frame rates need to exceed 16 fps. The team validated the path using Rolling Forcing as a proxy model and argues its methods can transfer to more complex systems containing encoder, DiT, VAE, and decoder.

The text provides no controlled comparisons against other hardware or baselines, nor does it describe how these optimizations perform on larger or differently architected autoregressive diffusion models; the specific measurement conditions behind the operator-level latencies and the 8.25 times speedup, as well as the video quality and stability corresponding to the first end-to-end success, still warrant confirmation through subsequent public materials.

Sources