Skip to main content
Back to timeline
arXivSource publication:

StoryEngine constrains multi-shot video generation with an explicit story world state, lifting anchor persistence to 0.9389 on a 60-story benchmark

Synopsis

The work proposes StoryEngine, a state-grounded agentic framework that separates an authoritative semantic plan from fallible generated pixels, using story-state propagation, grounded render planning, and a bounded evaluation-guided repair loop to produce long-form multi-shot video; on a self-built 60-story benchmark with two video backbones, Veo 3.1 and Wan2.2-TI2V-5B, it outperforms ViMax, MovieAgent, and the planner-free Direct I2V baseline on all metrics of narrative realization, cross-shot coherence, and visual consistency.

AI-generated editorial illustration: StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Interpretation

StoryEngine makes the story world state explicit: a structured representation records recurring entities' placements and story-relevant attributes, converts narrative events into explicit state transitions, and uses them to define each shot's intended start and end states. Earlier agentic pipelines mostly coordinate shots through loosely structured textual shot descriptions or by conditioning later shots on previously generated frames or clips, lacking an explicit mechanism for propagating event consequences; here a shot's opening state follows the planned state trajectory rather than being inferred from potentially erroneous generated pixels. The paper formalizes state reduction and validation, noting that events are processed in narrative order, that events without state-changing effects leave the state unchanged, and that entity references, placement relations, and attribute domains are validated after each update; in ablations, removing this stage (w/o SSP) lowers Avg by 0.0142, SPS by 0.0307, and APR by 0.0278.

Grounded render planning builds canonical references for recurring entities and environments (character identity sheets, state-specific prop references, environment panoramas), selects action-relevant camera views, and compiles state, camera, and visual constraints into executable render plans. The work translates the semantic contract of what must be true into generator-ready shot-level contracts with phase-tagged criteria (start, motion, end, always), so visual conditioning is selected from the planned state trajectory instead of being regenerated from each shot's wording. The paper describes asset screening against identity or layout criteria and absence of unrequested text or people, six canonical panorama probes, and criterion normalization; in ablations, removing this stage (w/o GRP) causes the largest drop, with Avg falling from 0.7690 to 0.7119 and GPC and LCS decreasing by 0.1074 and 0.1411.

Continuity-aware generation uses a one-way evidence gate to decide whether the previous shot's tail frame is reused, partially referenced, or discarded, and a bounded evaluation-guided repair loop corrects local state inconsistencies within a fixed retry budget. Existing systems repeatedly reuse generated pixels as context, so a misplaced object or corrupted detail may be read as a new narrative fact; here preceding pixels can condition the next shot only when compatible with its intended opening state, and can never modify the propagated state trajectory or its compiled criteria. The paper specifies a deterministic mode resolver with reuse, reference, and fresh modes and a monotone fallback order, and states that UNKNOWN is not converted to PASS; in ablations, removing this stage (w/o SCVC) lowers Avg to 0.7251 and yields the lowest ECS, APR, and MIR among the three component ablations.

The authors build a 60-story benchmark with three diagnostic suites (N20 narrative realization, T20 cross-shot coherence, C20 visual consistency) and eight metrics, and StoryEngine achieves the best score on all seven video-based metrics with both backbones. Evaluation targets are fixed before generation so failures can be localized to individual shots and cuts, with annotation sheets frozen under a content hash and hidden from all methods; the metrics cover plan event coverage, event completion, anchor persistence, state progression, location adherence, geometric place consistency, minimum identity retention, and lighting coherence. The paper reports StoryEngine's Avg of 0.7690 with Veo 3.1 and 0.7372 with Wan2.2-TI2V-5B, exceeding the strongest baseline ViMax by 0.1416 and 0.1304 respectively; APR rises to 0.9389 and 0.8444, LAR to 0.9800 and 0.9749, and LCS to 0.5691 and 0.5372, while every baseline's LCS lies between 0.21 and 0.27.

Perspective

The framework targets narrative long-form video generation with well-defined entities, locations, and shot plans, fitting production settings that require cross-shot causal progression and visual identity, such as storyboard-driven short films or advertising-style content; its gains are larger with a weaker video backbone, suggesting that explicit render criteria and the repair loop matter more when the backend is less reliable. For researchers, it offers a reusable evaluation approach: fixing evaluation targets before generation, constraining judgments with frozen annotations and distractor events, and reporting plan-text and video metrics separately.

Open questions the paper itself notes include: once the authoritative plan is wrong, later visual evidence cannot correct it autonomously, and schema checks catch structural contradictions but not every commonsense or cultural error; canonical references and a panorama do not constitute a true 3D scene, so large viewpoint changes, severe occlusion, mirrors, transparent objects, deformable props, crowds, and fine hand-object contact remain difficult; and a compatible endpoint does not prove the intervening motion is physically valid. On evaluation, the benchmark is English, fictional, and without audio, with 10 shots, two locations, and one or two named characters per story, so it does not cover long narratives with many locations, crowded multi-person interaction, dialogue or audio continuity, lip synchronization, on-screen text, or interactive editing, and lighting contrasts are annotated only between locations. In addition, SPS remains below 0.67 for all methods, and irreversible state progression is listed by the authors as an open challenge; on metric interpretation, PEC, ECS, APR, SPS, LAR, and the location gate of GPC rely on a VLM judge, and that VLM is the same model as StoryEngine's in-loop evaluator, so although the inputs differ, shared preferences could still affect those metrics, while MIR and LCS involve no VLM. This reading covered the full text, but the per-item numbers in the tables are not fully expanded in the text, so specific per-metric comparisons rest on the prose descriptions.

Sources