Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

A recoverable motion signal in the intermediate states of frozen video diffusion model Wan2.1 is decoded into 3D human motion generation

This work asks whether a frozen large-scale text-to-video diffusion model implicitly contains explicit 3D human motion knowledge, probes Wan2.1 and finds a recoverable motion signal present in its intermediate states across the entire denoising schedule, and introduces parasitic co-denoising together with its instantiation, the Parasitic Motion Decoder (PMD), an efficient flow-matching decoder that shares the host's noise schedule and reads its intermediate features through sigma-adaptive multi-layer fusion, leaving the host unmodified, leading dedicated motion generators on text-motion alignment at a small fraction of their trainable parameters while producing paired video and motion in a single pass.