PROWBench tests video models on 170 replayable programmatic worlds: proxy-conditioned models place entities better, but 30-second horizons and multi-view identity binding still break down
Synopsis
The work introduces PROWBench, comprising 170 programmatically constructed episodes, 200 selected camera views and 600 synchronized proxy videos, using replayable world states and timestamped engine events as ground truth to evaluate whether video models faithfully render program-specified scene structure, rules and interactions, and proposes two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate; results show proxy-conditioned models clearly outperform camera-conditioned world models on entity control, while long-horizon memory, multi-view identity and appearance binding remain open problems.
Interpretation
PROWBench separates executable world dynamics from visual generation for evaluation: its data engine logs entity states and timestamped events, including events outside the camera's field of view, as replayable world records, then renders synchronized views and proxy representations so generated videos can be checked against the observable consequences of program execution. Existing benchmarks assess visual quality, controllability, instruction or physical adherence, but lack replayable records of entity states and timestamped events, so a generated video can be judged for plausibility but not against what the program executed; PROWBench adds that comparison layer. The paper reports 170 programmatically constructed episodes, 200 selected camera views and 600 proxy videos, yielding 37,500 view-specific camera-frame observations and 112,500 proxy video frames, totaling 62.5 minutes of aggregate proxy-video duration; each selected view provides three synchronized representations (coarse 3D, CWM proxy, Colored-OBB), and a subset of episodes provides synchronized multi-view observations.
The paper proposes two VLM-based metrics: Logic-Render Alignment measures whether each segment of the engine-derived action timeline is visibly depicted within its time window, while Interaction Success Rate counts an event successful only when both the prescribed action and the expected end state are judged present. Global similarity cannot determine whether a prescribed interaction is successfully executed; these metrics move evaluation from visual plausibility to timeline and event realization, with IoU-weighted variants gISR and gLRA that jointly reflect entity placement. For ISR, Qwen3.6-27B receives eight frames sampled around each event window, including full-frame context and an engine-localized crop; for LRA, the judge receives twelve frames from the corresponding temporal window plus neighboring descriptions as context; table entries average equally across samples so samples with more annotated events do not receive greater weight.
Main-track results show proxy-conditioned models clearly outperform camera-conditioned world models on entity control: in Verified-FF, LynnReal-Omni, MiniMax-H3 and CWM reach BBox IoU of 0.518, 0.494 and 0.454, whereas the four camera-conditioned world models reach only 0.303-0.351; in Unverified-FF, LynnReal-Omni and MiniMax-H3 reach 0.457 and 0.409 versus 0.204-0.255. The paper derives paired proxy representations from the same recorded world-state trajectory, isolating representation and viewpoint effects under fixed scene evolution, and reports class-level paired comparisons rather than confounded cross-model, cross-dataset comparisons. Class means are higher in 38 of 40 Verified-FF scenes and 87 of 89 Unverified-FF scenes, with paired sign tests reported; the camera-control gap is much smaller, with most methods' rotation error comparable to that of the pose estimator on the engine renders.
The challenge track shows long-horizon and multi-view settings remain open: over 30 seconds BBox IoU is only 0.175-0.280 and Reappearance IoU for returning entities is only 0.053-0.131; in the multi-view setting, methods receiving identity-colored boxes or a first frame reach 82.2-94.0 Appearance Compliance and 81.0-89.7 Cross-view Consistency, whereas C2R, whose grey coarse 3D proxy draws all players alike and which receives no first frame, reaches only 32.2 and 25.5. The paper treats long-horizon memory and cross-view identity consistency as separate evaluation dimensions and shows geometric agreement across independently generated views does not ensure consistent participant identity or appearance, so geometric metrics alone are not a reliable proxy. The long-horizon setting contains ten 30-second recordings with one camera each; the multi-view setting contains ten four-player scenes with four synchronized third-person follow views generated independently; the paper also notes state persistence covers only 17-21 events per method from an uncalibrated judge, so these numbers indicate tendencies rather than a stable ranking.
Perspective
The benchmark targets rendering fidelity of programmable world models: it applies where programmatic world records exist and one needs to check whether generated videos realize engine events along the prescribed timeline, and its main users are researchers building game-engine-style world models and generative renderers. Its data engine is extensible, allowing new scene layouts, actions and visual representations to be added and recorded episodes to be reused for aligned multi-view and multi-representation observations without rerunning behavior controllers. The main track covers five-second clips, while the challenge track covers 30-second long-horizon and four-player, four-view synchronized scenes; the paper also notes Verified-FF models the case where an agent already has a concept image or character design, and Unverified-FF the case with only world state, text and control signals.
The paper states several scope limits: ISR, LRA and the appearance metrics rely on a VLM judge not yet calibrated against human annotation; camera metrics inherit the error of the pose estimator, which reaches a comparable error level even on the engine renders; the Unverified-FF set was curated with MiniMax-H3 outputs and all first frames are generated by MiniMax-H3, which may favor it; and the challenge-track sets are small, with ten recordings or scenes each. Long-horizon state persistence covers only 17-21 events per method, and window-chaining protocols differ (for example, Seedance 2.5 additionally receives the last frame of the previous window), so that section measures end-to-end long-horizon preservation under each method's actual protocol rather than an isolated internal memory. The proxy design itself also remains an open problem: outputs sometimes follow the input geometry too closely and inherit the coarse shapes of low-poly proxies, box proxies leave open which described appearance belongs to which entity, and facing direction can be rendered the opposite way.
