PhysVista tests 16 VLMs with a seeing–reasoning–assessment loop: perception holds up, physical reasoning and plausibility scoring lag
Synopsis
PhysVista introduces a benchmark that links physical state perception, physical dynamics reasoning, and physical plausibility assessment into a closed loop, includes both real-world and AI-generated videos, and evaluates 16 VLMs, finding that models do relatively better on observable state perception but fall short on physical dynamics prediction, counterfactual reasoning, and fine-grained plausibility scoring and ranking, with performance generally better on real-world than generated videos.
Interpretation
PhysVista organizes physical intelligence evaluation as a seeing–reasoning–assessment loop that jointly measures physical state perception, physical dynamics reasoning, and physical plausibility assessment. The paper argues existing benchmarks cover only isolated stages of this loop: PhysBench focuses on perception, most physical reasoning benchmarks cover limited tasks and mainly event-level reasoning, and PAI-Bench covers perception and reasoning but lacks the final assessment component. PhysVista places all three in one framework and distinguishes event-level from scale-level reasoning. The paper uses a task-type comparison in Tab. 1 to show the coverage difference and defines each stage's subtasks and data construction; the appendix lists per-task sample counts such as Spatial 200, Localization 200, Camera 200, Scale 60, Uncertainty 180, Order 150, Mechanism 192, Violation 224, Prediction 135, Counterfactual 119, Critique 100, Score 200, and Comparison 100.
The benchmark includes both real-world and AI-generated videos and finds a systematic generalization gap on generated content. The paper notes early benchmarks were mostly simulation-based and recent work began adding real-world videos, while evaluation of AI-generated content remains "largely unexplored." PhysVista builds tasks from WISA-80K real videos and VideoPhy-2 generated videos, and in the assessment stage assigns real-world videos the highest plausibility score of 5. The paper reports higher accuracy on real-world than generated videos across most tasks, with the gap widening notably for Physical Uncertainty Awareness and Physical Dynamics Prediction; an exception is Quantitative Scale Estimation, where generated videos slightly exceed real-world videos, which the authors suggest may be because generated videos present cleaner geometric layouts or more controlled object configurations.
Evaluation across 16 VLMs shows a cognitive gap, with inconsistent capability rankings across the perception, reasoning, and assessment stages. Rather than reporting a single overall score, the paper uses cross-stage ranking shifts to show that strong perception does not imply strong causal reasoning, and strong mechanism reasoning does not imply calibrated plausibility judgment. The paper states GPT-5.2 and Kimi K2.5 rank third and fourth on State Perception but drop substantially in relative ranking on Dynamics Reasoning; Gemini 3.1 Pro and Doubao Seed 2.0 Pro rank third and fourth on Physical Dynamics Reasoning but decline on Physical Plausibility Assessment. Tab. 2 and Tab. 3 give per-task and average accuracies, for example GPT-6 Sol at 68.59% perception average and 62.68% reasoning average, and Claude Opus 5.5 at 67.89% reasoning average.
The plausibility assessment tasks expose two behavioral patterns: score collapse in scoring and over-confidence in pairwise comparison. The paper splits assessment into single-video 1–5 scoring and video-pair ranking, and analyzes score distributions and difficulty differences across pair types rather than reporting only one accuracy number. The paper reports that some models (e.g., Qwen3-VL-8B and Kimi K2.5) concentrate a disproportionately large fraction of predictions on the highest score level, while others (e.g., Gemini 3.1 Pro and Grok 4) shift the default score to the lowest level; models with genuine scoring ability show distinct distributions between generated and real-world videos. In comparison, real vs generated is easiest, generated vs generated is substantially harder, and real vs real yields low accuracy, indicating difficulty recognizing equivalence.
Perspective
The benchmark targets researchers and model developers who need to judge the physical consistency of dynamic scenes, especially those evaluating the physical authenticity of generated video. It organizes evaluation into perception, reasoning, and assessment stages and distinguishes event-level from scale-level reasoning, so it is suited to locating which part of the loop a model is weak in rather than producing a single overall score. Placing real and generated videos side by side lets results directly compare model behavior across the two data sources.
The paper reports results on a specific benchmark and a specific set of models, with model versions and access dates noted as 2026 in the references, so readers should note that conclusions are tied to those model versions. Phenomena such as score collapse are described as model-dependent, and the explanation attributing them to differing priors toward certain score tokens remains an open question. The slightly better Quantitative Scale Estimation on generated videos is offered as a possible explanation rather than a settled finding. The main text also notes that more detailed benchmark statistics are in the appendix, so reading only the main text may miss full data distribution details.
