Skip to main content
Back to timeline
arXivSource publication:

Physis-Lang rewrites video world models with self-evolving physical language: open-source Cosmos3-Nano surpasses Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench

Synopsis

Physis-Lang treats physical language as a shared, optimizable representation across data curation, model training and video generation, using a physics-aware critic and PhysCapBench to drive an agentic loop that refines physical captioning instructions and language-guided retrieval to cover missing physical processes, yielding consistent gains on four physical video benchmarks with Wan and Cosmos backbones, where models built on the open-source Cosmos3-Nano surpass Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench.

AI-generated editorial illustration: Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model

Interpretation

The paper proposes treating language itself as a physical world representation, rather than using language only as a semantic condition and adding visual, latent, numerical or planning signals on top. Prior work largely assumes natural language is insufficient for the physical knowledge needed for reliable generation and therefore adds geometry, motion or physical-correctness signals; this work revisits that assumption and makes language explicitly describe physical entities, causes, interactions, governing principles, temporal evolution and effects. The claim is supported by method design and ablation: the physical reasoning prompt raises the pretrained model from 61.67 to 63.33 without fine-tuning, showing that structuring the language side alone yields a zero-shot gain.

It builds PhysCapBench and a physics-aware critic that decomposes physical processes into atomic assertions, evaluates captions with precision and recall, and drives an agentic loop that refines the captioning instruction. PhysCapBench contains 246 physics-rich videos and 3,794 human-verified physical assertions, averaging 15.4 per video; the critic splits a generated caption into independently verifiable claims and labels them correct, incorrect or uncertain, while recall measures coverage of the human assertions. Iterating on a fixed 20-video development set, PhysCapBench F1 rises from 78.64 at iteration 1 to 87.82 at iteration 9; changing only the inference captions with the backbone fixed lifts PhyGenBench from 64.17 to 65.63, 65.83 and 67.29.

A language-guided data engine summarizes model failures into a category-level deficiency profile, then uses physics-tag matching to retrieve visually diverse videos covering missing physical processes and re-captions them with self-evolving physical captions for training. Unlike retrieval based solely on visual similarity, this matching selects videos by physical content rather than appearance, tilting the training distribution toward the generator's observed physical weaknesses. Adding retrieved videos improves the four benchmarks by 3.01 points on average; broken down by physical category on VideoPhy-2, gains are positive across frequently represented categories such as chemical processes (+8.00), fracture mechanics (+7.45) and cloth deformation (+7.19).

After fine-tuning Wan2.1-14B and the Cosmos3 family (Edge-4B, Nano-16B, Super-64B), physical plausibility improves consistently, and models built on the open-source Cosmos3-Nano surpass Veo 3.1 on three benchmarks. The paper reports average gains of 7.05 points for Wan2.1-14B and 6.22 for Cosmos3-Nano-16B across model families, and 3.24, 6.22 and 5.02 points across Cosmos scales; on the full VideoPhy-2 set it trails Veo 3.1 by 0.85 points (68.02 vs. 68.87) but leads by 3.93 points on the hard split (62.36 vs. 58.43). The four benchmarks are PhyGenBench and VideoPhy-2 for T2V and Physics-IQ Verified and PhyGround for I2V, evaluated under the same scoring protocol with the offline VLM evaluators of VideoPhy-2 and PhyGenBench replaced by GPT-5.5; on VBench-I2V, Wan2.1-14B moves from 86.86 to 87.51 and Cosmos3-Nano from 88.32 to 88.69, indicating general generation quality is preserved.

Perspective

The work targets developers and deployers of video world models that need physical plausibility, applies to text-to-video and image-to-video generation as well as driving and robotics demos, and its gains come from combining physical captions, language-guided retrieval and inference-time positive and negative prompts; it can be layered onto Wan2.1 and the Cosmos3 family without changing backbone architectures or training objectives, and PhysThinker-C and PhysThinker-U let local 4B models replace commercial ones, cutting estimated external annotation cost from $24.12K to $0.12K and to zero.

PhysCapBench's 246 videos and 3,794 assertions, and the development set's 20 videos and 273 assertions, are limited in scale, so whether they cover the broader space of physical phenomena remains an open question; the evaluators for VideoPhy-2 and PhyGenBench were replaced with GPT-5.5, so comparability with the original leaderboards and the influence of evaluator bias deserve continued watching; after distillation, PhysThinker's average gain drops from 7.05 to 4.76 points, leaving the accuracy-cost trade-off under different deployment budgets to be validated in more settings; and whether the physical language representation extends to longer horizons, multi-object coupling and interactive simulation is still to be explored.

Sources