Skip to main content
Back to timeline
arXivSource publication:

ThinkV2V makes MLLMs think before editing: a 5B model tops both complex and standard video-editing benchmarks

Synopsis

The work proposes ThinkV2V, a reasoning-driven video-editing framework that explicitly activates multimodal large language model thinking before visual generation: it uses Qwen3-VL-Thinking-8B to perform chain-of-thought reasoning over the source video and instruction and output a refined prompt, injects the hidden states into a Wan-2.1-5B DiT through a learnable-query connector, pairs this with Progressive Curriculum Training and Inference-Time Thinking Scaling, and builds the ThinkV2V-150K dataset and ThinkV2V-Bench; under two judge models, the 5B-scale model achieves the best overall scores on both complex and standard editing scenarios and surpasses several 10B-scale baselines.

AI-generated editorial illustration: ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Interpretation

ThinkV2V turns the MLLM from a semantic encoder into an explicit reasoner: the MLLM first performs chain-of-thought reasoning over the source video and original instruction, produces a refined prompt, and extracts the last-layer hidden states of answer tokens after the </think> tag as high-level semantic features. Prior instruction-guided video editing methods mainly use MLLMs as stronger semantic encoders that jointly process prompt and video, whereas this work makes the thinking process explicit and converts it into editing conditions. Ablations show that treating the MLLM as a static encoder ("w/o Thinking") yields an overall score of 2.51, enabling thinking with all features raises it to 2.63, and using answer features raises it further to 2.68; qualitative comparisons also show edits better aligned with intent when thinking is enabled.

Architecturally it adopts an MLLM-to-DiT design: a learnable-query connector compresses MLLM hidden states into fixed-length conditioning features, while original-instruction text embeddings and source-video VAE features are also fed into the DiT to mitigate the information bottleneck. Compared with directly exposing the DiT to a long and potentially unstable token sequence, learnable queries perform feature compression and cross-module alignment; the multi-condition design avoids over-reliance on compressed MLLM features and preserves fine-grained content details. Implementation details report a learnable-query token length of 512, zero-initialization of the connector's final layer to stabilize early training, DiT initialization from Lucy-Edit, and a frozen MLLM during training.

The training-and-inference recipe combines Progressive Curriculum Training and Inference-Time Thinking Scaling: the curriculum advances in three stages along resolution and instruction complexity, and at inference time candidate prompts are serially refined and then selected by the MLLM via best-of-N. The work designs "how reasoning is learned" and "how reasoning is used at test time" as separate components rather than relying on a single forward pass. In ablations, "Simple-to-Complex" reaches an overall score of 2.68, above "Complex Only" 2.10, "Complex-to-Simple" 2.41, "Mixed Simple+Complex" 2.47, and "Simple Only" 2.52; serial refinement alone gives 2.67, and adding selection raises it to 2.72.

On data and evaluation, it builds ThinkV2V-150K and ThinkV2V-Bench: the former retains five editing categories from OpenVE-3M, applies quality and suitability filtering, and uses Gemini-2.5-Pro to rewrite instructions into reasoning-oriented ones, while the latter rewrites OpenVE-Bench's direct instructions into reasoning-oriented ones, yielding 308 video-instruction pairs. Existing datasets and benchmarks mainly emphasize basic editing ability and lack training and evaluation support for implicit intent and causal reasoning. The data pipeline includes Gemini-2.5-Flash rescoring, average inter-frame CLIP similarity and Temporal Flickering filtering, retaining the top 50% highest-quality samples to form OpenVE-HQ-1M, then selecting 20K to 40K samples per category to synthesize 150K video pairs.

Perspective

The result targets instruction-guided video editing that requires implicit intent understanding and causal reasoning, covering five editing categories: Global Style, Background Change, Local Remove, Local Add, and Local Change; both training and evaluation are designed around this scope. For creators and content production workflows seeking to reduce trial and error in complex editing, it offers a reusable "reason first, then generate" path; for researchers, ThinkV2V-150K and ThinkV2V-Bench provide a data and evaluation foundation that can be further extended.

The authors note in the limitations that explicit thinking inevitably introduces additional computation and latency, increasing both training and inference overhead; because rigorous quantitative evaluation protocols for semantic reasoning traces are unavailable, the correctness of the thinking content itself is not independently evaluated; and the current coverage of ThinkV2V-150K and ThinkV2V-Bench is still limited for more open-domain, longer-horizon, and rarer complex editing scenarios. In addition, the benchmark has 308 video-instruction pairs, deliberately focused rather than large-scale, so extrapolation of its conclusions to broader editing types still needs further validation.

Sources