PhysBrain 1.5: A Physical Foundation Model Unifying Language, Action, and Future World States in One Autoregressive Vocabulary
Synopsis
PhysBrain 1.5 builds on a pretrained Qwen3-VL backbone extended with dedicated action and visual-state tokens, formulating language responses, structured spatial outputs, end-effector trajectories, and future world states as discrete tokens jointly learned under a single next-token prediction objective, with embodied pretraining supervision drawn entirely from human interaction videos (egocentric, synchronized ego-exocentric, and panoramic recordings structured into task-centered episodes) and supervised fine-tuning combining human demonstrations, real-robot trajectories, and simulated experience; across 28 embodied spatial-intelligence and planning benchmarks the 8B version reaches an overall average of 72.
Interpretation
Language, action, and visual states are unified as discrete tokens in a single autoregressive stream, with no task-specific heads. Where earlier work often designed separate prediction heads for action or world modeling, here semantic, motion, and future-state supervision all update the same parameters through one next-token prediction objective. The paper presents this as an architectural design with a unified vocabulary, built on pretrained Qwen3-VL with dedicated action and visual-state tokens; the evidence form is model design plus overall benchmark performance rather than ablation studies.
ActionPiece tokens encode end-effector trajectories in a unified action codebook shared across control configurations and arm setups. This allows one generalist checkpoint to drive varied robotic platforms and manipulation settings instead of training a separate policy per platform. The paper describes the unified action codebook and vocabulary mechanism and pairs it with qualitative results of action trajectory prediction; evidence is mainly qualitative plus overall benchmark scores.
Future world states are predicted inside the same vocabulary as spatially aligned RGB images, depth maps, and robot masks. The world model is folded into the same autoregressive vocabulary as language and action rather than being a separate module. Supported by the description of a multimodal world model predicting one step ahead and by qualitative future visual prediction results across diverse robot embodiments.
On 28 embodied spatial-intelligence and planning benchmarks, the 8B version reaches an overall average of 72.5, first among open-source models and close to leading proprietary systems. The paper states all models are independently re-evaluated with one canonical metric per benchmark for direct comparability, and reports 14 first-place and 10 second-place results among open-source models. Evidence is a summary score table across 28 benchmarks and 13 models, with the overall average being the unweighted mean across the 28 benchmarks; closed-source models and the 2B version are shown for reference and excluded from ranking.
Perspective
The results target embodied spatial-intelligence and planning tasks: given a task instruction, current observations, and optional recent action history, the model predicts the next end-effector action chunk and predicts one step of future RGB, depth, and robot masks. The intended audience is researchers and engineering teams who want a single generalist checkpoint to drive varied robotic platforms and manipulation settings and to reuse a standard model interface for inference and post-training. Embodied pretraining supervision comes from human interaction videos, and fine-tuning combines human demonstrations, real-robot trajectories, and simulated experience, so the positioning is a physical foundation model under this data and task setting.
The text loaded here is an incomplete scope, lacking the full technical report body, figure and table details, and ablation information, so details about pretraining data scale, training recipe, how the action codebook is constructed, and per-benchmark differences remain open questions. Readers may continue to watch: whether language, action, and visual states influence one another within the unified vocabulary; how the unified action codebook performs across more control configurations and arm setups; and how future-state prediction accuracy relates to downstream manipulation success. In addition, the overall average is an unweighted mean across 28 benchmarks, so how the difficulty distribution across benchmarks affects that summary value is also worth noting when reading the full report.
