4Director controls video world models with complete meshes and per-frame rigid transforms, beating four baselines on camera and object control
Synopsis
4Director introduces a video world model conditioned on an explicit 4D scene representation in which each object is reconstructed once from the input image as a complete canonical mesh and moved by one prescribed rigid transformation per frame; the controlled scene is rendered as a depth video and a Motion Adapter injects it into a pretrained video generator to supply appearance, illumination and non-rigid dynamics, with the authors building RealCOD-Rigid, a dataset of 20,774 clips annotated by an automatic pipeline on RealCOD-25K, and introducing Identity-Gated IoU (IG-IoU), reporting that 4Director outperforms MotionCtrl, Perception-as-Control, SymphoMotion and VerseCrafter on visual quality, camera control and object control.
Interpretation
The paper introduces an explicit 4D scene representation in which every object is a complete canonical mesh reconstructed once from the input image and moved by one rigid transformation per frame, with meshes, a background point cloud and the camera sharing one coordinate frame. Compared with image-plane cues (dragged paths, bounding boxes, masks, point tracks) that are ambiguous in depth and rotation, and with 3D trajectories, 3D boxes and lifted proxies (tracked 3D points, spheres, one 3D Gaussian per object) whose geometry is incomplete, this representation fixes position, orientation and the revealed surface before generation. Supported by the method description and a comparison table covering camera trajectory, 3D object position, object orientation, complete geometry and new object insertion, plus an ablation that places 2D box, 3D box, 3D mesh and full rigid 3D geometry side by side.
The paper presents the Motion Adapter, a trainable branch that injects the depth video rendered from the rigid scene into a pretrained video generator, Wan2.1-VACE-14B, so that geometry fixes camera and object motion while the generator supplies appearance, illumination and non-rigid dynamics. The depth video carries only the viewpoint, each object's rigid transformation and occlusion, and lacks appearance, illumination, non-rigid dynamics and the background beyond what the image shows; the adapter is trained to supply exactly what the rendering omits. The adapter has 3.0B parameters and is trained on the 20,774 pairs with a flow-matching objective for three epochs on 24 GPUs with a global batch of 24; a rank-16 LoRA variant (15.3M parameters) is also reported.
The paper builds RealCOD-Rigid, a dataset of 20,774 clips annotated with complete object meshes, per-frame rigid transformations and camera trajectories by an automatic pipeline, and introduces Identity-Gated IoU (IG-IoU), which accumulates mask IoU only over frames in which the object's identity is preserved. The authors state that no dataset provides clips paired with their rigid 3D scenes and that such annotation cannot be done by hand at scale, and that mask IoU alone rewards correct placement regardless of identity. The pipeline starts from the 25,318 source clips of RealCOD-25K and runs four stages (camera and depth estimation, first frame to 3D, rigid body tracking, depth rendering); the 20,774 clips that pass every stage and are not used for evaluation form the training set, and the IG-IoU identity gate comes from a Qwen3-VL vision-language judge with masks from SAM3.
Experiments report that 4Director outperforms the four baselines on visual quality, camera control and object control, and keeps the same object when it leaves the camera view and later returns. IG-IoU rises from 54.8 for VerseCrafter, the next best, to 60.4 at nearly the same recognition rate (94.9 against 94.8), which the authors attribute to more accurate placement rather than more frequent recognizability; in the ablation, full rigid 3D geometry scores 60.4, dropping to 50.3 without the surface and 53.0 without rotation. Evaluated on the same 100 clips of RealCOD-25K excluded from training, using FID, FVD, CLIP-SIM, rotation and relative translation error, IG-IoU and eight VBench-I2V dimensions, plus a user study in which 20 participants rated nine cases from 1 to 5 on visual quality, control accuracy and consistency.
Perspective
The representation controls motion only at the level of a rigid body: an articulated or deforming object moves as one whole, so finer motion such as running, jumping or limb movement cannot be prescribed and is decided by the generator, which the paper's failure case shows may even keep a mostly articulated object in its first-frame pose. The authors list extending the rigid 3D scene with articulated parts as a natural next step. The method targets settings that start from a single image and let a user author camera and object trajectories in a 3D viewer; a new object can be inserted as a mesh reconstructed from a separate image and given a trajectory in the same coordinate frame without additional training. The paper reports 61 authored cases but presents them only qualitatively, since they have no source video under the authored trajectories and the training data contain no insertion pairs.
The quantitative comparison centers on 100 evaluation clips and the user study on 20 participants over nine cases, so the sample sizes are limited and the reach of the conclusions still needs observation across more scenes and longer clips. The IG-IoU identity gate depends on a vision-language judge, and the paper notes that mask IoU does not penalize a turn of a front-back symmetric object, which is shown only qualitatively. Camera control error is obtained by re-estimating trajectories with the same MegaSaM pipeline used for annotation, so evaluation and annotation share one estimation pipeline. In addition, the rigid representation cannot prescribe articulated motion, and the paper's failure case shows the generator may keep a mostly articulated object in its first-frame pose; how control capability changes once articulated parts enter the scene representation remains an open question.
