Reversing frame order often leaves VideoLLM answers unchanged; researchers locate a mid-layer temporal peak and inject it back into later layers without training
Synopsis
The work defines the temporal divergence vector τ_l, the layer-wise representational difference induced by reversing frame order, and finds across Qwen2.5-VL-7B, Qwen3-VL-8B, and InternVL2.5-8B that its normalized magnitude peaks at intermediate-to-late layers and then declines toward the output; question-conditioned analysis and attention knockout confirm the peak is specific to temporal reasoning and functionally critical, motivating Temporal Activation Injection (TAI), a training-free method that extracts τ_l at the peak and reinjects it into subsequent layers following the measured decay, consistently improving temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks.
Interpretation
The paper introduces the temporal divergence vector τ_l and its layer-wise profile: pairing each video with its frame-reversed counterpart, which shares all spatial content and text conditioning, isolates the contribution of temporal order, and across three models the normalized magnitude follows a shared shape that rises through intermediate layers, peaks in the mid-to-late layers, and then declines toward the output. Prior work such as T3 and ArrowRL localized the bottleneck to the language model or noted output-level insensitivity to reversal, but did not characterize how the signal evolves inside the layers; this work turns where temporal information is strongest and how it decays into a measurable per-layer curve. Based on 50 reversed video pairs with binary temporal questions from TempCompass, averaged over Qwen2.5-VL-7B, Qwen3-VL-8B, and InternVL2.5-8B, which differ in vision encoder, LLM backbone, and positional encoding; the peak-then-decay shape emerges with as few as 5 pairs, and the Pearson correlation with the 50-pair profile exceeds 0.95.
The peak is shown to be specific to temporal reasoning and functionally critical: holding the video pair constant and varying only the question, non-temporal conditions (a spatial question and a video-irrelevant question) substantially attenuate the peak; attention knockout on the last token produces the largest drop in ground-truth answer probability when blocking layers around the peak. This separates every representational difference caused by reversal from the subset that actually serves temporal reasoning, upgrading the profile from a descriptive curve into a diagnostic that locates an intervention point. The question-conditioned analysis compares three question types on the same video pairs; the functional test sweeps attention knockout across all layers and measures changes in ground-truth answer probability on temporal subtasks, with the affected region aligning with the profile peak.
This motivates TAI: it extracts τ_l at the profile peak and reinjects it into the last token at subsequent layers using the measured decay as a per-layer schedule, with strength adapted per input by the magnitude of τ_l, requiring no training, no annotations, and no model modification. Unlike contrastive decoding methods (TCD, VTD, SEASON) that intervene on output distributions, or methods that learn steering directions from labeled data, TAI intervenes at the intermediate representation level, with both direction and schedule grounded in quantities measured from the model, leaving a single global scaling coefficient as the only free hyperparameter. Ablations show the profile-guided schedule beats uniform and reversed schedules (TempCompass average 75.6 versus 75.4 and 75.1, baseline 73.4); extraction at the peak layer yields the largest gain, earlier layers give near-zero improvement, and single-layer extraction is best while widening the window degrades only gradually.
Across three VideoLLMs and four benchmarks, TAI consistently improves temporal reasoning with negligible impact on non-temporal ability: TempCompass averages rise from 73.4/77.1/70.9 to 75.6/78.0/73.6, TVBench from 44.6/49.9/57.2 to 46.6/52.6/58.3, and AoTBench from 54.5/56.6/54.6 to 57.7/59.9/56.6, while on MVBench the temporal-relevant group improves and the temporal-irrelevant and hybrid groups fluctuate only slightly. Among training-free methods TAI achieves the highest average and approaches ArrowRL on Qwen2.5-VL-7B, which requires reinforcement learning with curated temporal rewards; applying TAI on top of ArrowRL-trained weights still adds gains, indicating activation-level steering and weight-level optimization are complementary. All numbers come from the authors' own evaluation under the same hardware and 16-frame protocol; an Anti-TAI sign-inversion control selectively degrades reversal-sensitive categories (Attribute Change down 36.3%, Order 27.4%, Direction 12.6%) while reversal-invariant categories (Action, Speed) stay within 1% of baseline; gains remain statistically significant under frame-sampling and question-phrasing perturbations via the McNemar test.
Perspective
The result targets temporal reasoning tasks that depend on frame order, such as event ordering, directional motion, and attribute change over time, and applies to decoder-only VideoLLMs that expose intermediate representations; it can be layered onto existing weights at inference time, and the profile is estimated once per model and reused without per-domain recalibration. The authors note the work focuses on temporal information arising from frame ordering, while spatial reasoning, object interactions, and audio-visual correspondence fall outside its scope, though the framework could in principle be extended to other axes of variation by designing contrastive pairs beyond frame-order reversal.
The profile and injection schedule are estimated from 50 reversed TempCompass video pairs; although re-estimating from other benchmarks leaves accuracy nearly unchanged, whether this holds for longer, more event-dense, or more domain-shifted videos is not settled in the text. TAI's gains are clear on temporal categories and near zero on reversal-invariant ones, so tasks whose answers do not depend on frame order should expect limited benefit; the Direction category is internally heterogeneous, which the authors use to explain its smaller average τ_l. Free-form generation is examined only qualitatively (a sunrise video and its reversed counterpart) without quantitative evaluation. In addition, this summary is based on the paper's text and appendix rather than the original figures and tables, so judgments about curve shapes and visual examples follow the written descriptions.
