Why Video Diffusion Models Break Physics: Researchers Trace It to RoPE Spatial Anchoring and Fix It by Rescaling RoPE Frequency
Synopsis
The work presents an interpretability study of the "motion planning" process in text-to-video diffusion models, finding that excessive spatial attention decay induced by RoPE in self-attention makes early candidate regions lock prematurely onto physically implausible positions, and proposes a lightweight architectural change that merely rescales RoPE frequency across denoising steps, improving physical commonsense in both training-free and training-based experiments.
Interpretation
The study offers the first interpretability study of the "motion planning" process in text-to-video diffusion models, extending the empirical "first shape, then details" finding to a mechanistic level: using Wan2.1-T2V-1.3B and a "basketball falling freely and bouncing" case, cross-attention maps converge from multiple candidate regions to a deterministic shape at around step 5 of 50, quantified by two proposed metrics, attention entropy and support quality. Prior work established that early denoising steps finalize spatial layouts but left the underlying formation mechanism unexplored; this work examines it from the model-architecture view, resolving the phenomenon down to attention heads and denoising steps. Based on layer-by-layer and head-by-head visualization of a single main case plus two custom metrics, with the authors reporting the conclusions hold in the absolute majority of other cases and random seeds.
Using convergence speed and causal intervention (attribution patching, mean ablation, zero ablation), the study separates cross-attention heads: only a small subset with visible trajectory patterns truly affects motion planning, and a clear trajectory pattern is not a sufficient condition for a head to be responsible for motion planning; ablating the third type collapses the object trajectory, while ablating the first two types mainly affects static content such as object appearance and background. It moves the question of which attention heads govern motion from observational correlation to a causal-intervention-based classification, and shows that a large contribution score does not imply a large real effect on the trajectory. Evidence comes from a patching metric on velocity prediction and from generated-video quality after zero ablation, with the classification reproduced on other cases and random seeds.
The study identifies excessive spatial attention decay from RoPE in self-attention, a "spatial anchoring effect": a region that stabilizes early in one frame and is physically implausible raises the confidence of spatially adjacent regions in other frames, suppressing more reasonably positioned candidates that are farther away and triggering failure modes such as a basketball stopping in mid-air. It attributes physically implausible generation to an architectural factor, RoPE-induced spatial decay in self-attention, rather than to missing external priors or data. Evidence includes the strengthening correlation between mutual consistency and anchor distance across denoising steps, and visualizations of confidence changes for regions K1 and K3 in the seed-20 failure case.
The study proposes a lightweight RoPE modification that applies different scalings to RoPE frequency across denoising steps to reduce spatial decay and encourage exploration of more candidate regions during motion planning; on VideoPhy's 343 test cases, using human-evaluated Semantic Adherence and Physical Commonsense, the training-based method generally outperforms training-free and other baselines, with advantages mainly on the solid-* subset, and it can be combined with prompt refinement. Unlike approaches that add physics simulators, specialized data, or external foundation models, this method changes only a scaling factor, preserving native scalability and combining with prompt refinement for further gains. Both training-free and training-based settings (LoRA fine-tuning with a custom timestep sampler) with blinded human evaluation; the authors report that automated evaluation shows significant bias, so they use human evaluation.
Perspective
The results are aimed at researchers and engineering teams studying the internal mechanisms and physical commonsense of text-to-video diffusion models, and apply to flow-matching video generation models with a DiT backbone and 3D RoPE, such as Wan2.1-T2V-1.3B/14B, mainly for solid dynamics where trajectories are easy to track. At inference the RoPE modification is applied only during the first 5 denoising steps, and training combines a custom timestep sampler with LoRA fine-tuning; the authors also state the idea can extend to broader settings such as image-to-video and few-step autoregressive diffusion models, which they leave for future work.
The mechanistic conclusions rest mainly on a single main case (basketball free fall) and visualizations from a limited set of random seeds; the authors report the classification holds in the absolute majority of scenarios, but replication on larger-scale cases and different models remains to be seen. Effectiveness is evaluated on VideoPhy's 343 cases with blinded human evaluation, and the authors note the judge model in automated evaluation shows severe hallucination and low correlation with human results, so comparability of numbers across evaluation protocols is still open. The training-free setting requires manual tuning of the scaling factor, which the authors consider hard to scale; attempts at a learnable scaling factor did not yield clear improvement, suggesting current architectures and flow-matching losses still capture object motion trajectories insufficiently. In addition, this is a full-text read, but figures in the text appear as placeholders, so specific visual details require consulting the original paper and appendices.
