EditWorld moves video world models from navigation to streaming editing, scoring 73.8 overall and 80.0 on editing in WBench-Editing
Synopsis
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
Interpretation
EditWorld extends world-model interaction from navigation to precise modification of existing world content, supporting continuous streaming editing instructions during autoregressive generation and flexible injection of reference-image content. The paper states that existing video world models mainly focus on navigation and text-driven event generation, and that XGEN-JING and ABot-World support identity conditioning from initial references but do not let users flexibly inject content from different reference images over time; EditWorld makes both editing and reference conditioning time-varying streaming conditions. The abstract and introduction give the capability positioning, and the methodology gives concrete designs for Gated Causal Attention and unidirectional gated self-attention over reference-image tokens, constituting system- and architecture-level evidence.
To support streaming editing, EditWorld introduces Gated Causal Attention for temporally varying editing conditions and reference images, and uses a Sparse Context mechanism that keeps historical context within a fixed budget for long-horizon inference. The paper explains that the editing instruction prompt is activated only when the corresponding chunk is labeled as during, and that each chunk also attends to the prompts of the two preceding chunks to reduce abrupt transitions from prompt switching; reference-image tokens attend only within the same reference image, a video chunk's visibility of reference tokens is gated by its textual context, and reference tokens receive a negative temporal RoPE margin to discourage direct copy-and-paste. The methodology specifies the attention visibility rules, gating conditions, and RoPE treatment, supported by ablations: removing unidirectional self-attention loses reference detail, and removing short-term textual context produces a pronounced scene shift at editing-prompt switches.
For training, the work adopts joint autoregressive and bidirectional objectives, annealed self-resampling, and two-stage few-step distillation, so the model gains causal generation while retaining responsiveness to input conditions. The paper reports that optimizing only the causal objective biases the model toward video continuation and weakens responsiveness to diverse input conditions, so it jointly optimizes both branches following LingBot-World 2.0; the ablation shows joint AR+BI raises Editing Overall from 69.7 to 80.0 and Detail Accuracy from 34.9 to 53.2. The ablation table gives paired numbers for AR Only and AR+BI across Editing Overall, Editing Presence, Semantic Alignment, Editing Completion, Detail Accuracy, and Edit Cleanliness.
The paper builds a data synthesis and annotation pipeline for world editing and introduces WBench-Editing to systematically evaluate streaming world editing, with EditWorld achieving the best overall performance on that benchmark. On the data side it designs separate synthesis pipelines for global and local editing, with annotations covering scene description, editing instruction, editing state, and camera poses, and re-estimates camera intrinsics and extrinsics with ViPE; on the evaluation side it redesigns roughly 150 cases on the WBench framework, each spanning 240–480 frames with one to three editing instructions and a subset including reference images. The paper reports EditWorld at 73.8 overall and 80.0 on editing in WBench-Editing, 6.8 points above the second-best YUME 1.5 at 67.0 overall and 25.2 points above the second-best editing score of 54.8; on the reference-image subset it reports 74.0 overall and 74.6 on editing.
Perspective
This work targets interactive settings where users continuously modify content in a generated world, such as multi-turn editing, injecting reference images at chosen turns, and long-horizon autoregressive generation; the paper describes a data pipeline that separates global and local editing, and evaluates on roughly 150 WBench-Editing cases of 240–480 frames each with one to three editing instructions and a subset with reference images. For researchers building on it, the reusable pieces include Gated Causal Attention, Sparse Context, joint autoregressive and bidirectional training, annealed self-resampling, two-stage few-step distillation, and WBench-Editing as an evaluation entry point.
The reported results are on WBench-Editing and WBench, with editing-related metrics evaluated through VLM-based question answering; behavior under different bases, data conditions, or interaction turns still needs more independent verification. The data pipeline relies on several off-the-shelf models for synthesis and annotation, and the effect of annotation quality on final capability is not separately quantified. In addition, because most existing models do not support interactive reference inputs, the comparison uses text-only streaming editing instructions, so cross-model comparability under reference conditioning remains something to watch.
