Skip to main content
Back to timeline
arXivSource publication:

ACT3 tri-stream transformer makes the action expert the sole integrator of semantics and dynamics, lifting RoboCasa success from 59.4% to 71.2%

Synopsis

The authors introduce ACT3, an action-centric tri-stream transformer in which a semantic stream (VLM) and a dynamics stream (world model) compute independently while only the action expert reads both through layerwise attention, with all streams jointly optimized under action supervision. It reaches 71.2% success on 24 RoboCasa kitchen tasks versus 59.4% for the baseline, 98.6% average on LIBERO versus 96.9%, and 82.4% versus 68.4% on LIBERO-Plus distribution shifts.

Source-provided article image: Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
Figure 2 ·

Figure 2: Overview of the proposed ACT 3 . The architecture comprises three specialized transformer streams for semantic understanding, action generation, and dynamics prediction. The semantic and dynamics streams maintain independent forward computation, while the action expert attends to semantic, action, and dynamics representations at every action transformer block, making the action stream the sole integration point for control.

arXiv

Interpretation

ACT3 makes the action stream the sole integration interface for semantics and dynamics: each action layer reads VLM and WM keys and values through layerwise attention, while neither context stream reads the other's representations, preserving independent forward computation. VLA models built on VLM backbones carry limited physical dynamics priors, while replacing the backbone with a video world model can sacrifice task-level semantics; existing multi-stream methods add semantic-dynamics or dynamics-action coupling beyond the action pathway. ACT3 removes that extra coupling so the action expert is the first point where semantics and dynamics meet. 71.2% versus 59.4% on 24 RoboCasa kitchen tasks under the same downstream demonstrations and camera views; 98.6% versus 96.9% average across four LIBERO suites; 82.4% versus 68.4% on the same 10,030 perturbed LIBERO-Plus instances without additional fine-tuning.

Ablations over cross-stream attention topologies show success dropping as extra coupling is added (64.6%, 64.0%, 59.1%), while ACT3's topology, in which only action queries read both contexts, is highest at 71.2%. The ordering links avoiding additional cross-stream dependencies during context construction to performance, indicating that independently constructed semantics and dynamics can be weighted and combined by action queries as the action representation evolves. Controlled topology comparison within one architecture; all variants let action queries read both semantic and dynamics information, differing only in how context is constructed, with masks applied to the last 18 Cosmos blocks interleaved with semantic and action layers.

Layerwise access outperforms final-layer reuse: reading corresponding VLM and WM layers gives 71.2%, versus 65.8%, 63.5%, and 64.6% when final-layer K/V reuse is applied to the VLM only, the WM only, or both. This suggests intermediate semantic and dynamics features may expose fine-grained motion cues not present in final-layer representations, letting action refinement draw on local control detail together with task-level guidance. All variants reuse existing K/V projections and differ only in which context layers action queries read, making the comparison controlled.

Generated-future context and dynamics-prediction supervision each independently improve control: 61.3% to 65.8% without prediction supervision and 66.0% to 71.2% with it, while prediction supervision also helps current-frame-only control. The study separates the dynamics stream's two roles, context for action generation and an auxiliary prediction objective, and shows that learning demonstrated transitions shapes control-relevant representations even when no generated future reaches the action stream. Four factorial configurations share the same Cosmos backbone and action-context interface with an unchanged context-token budget; point estimates suggest approximately additive benefits.

Perspective

The design targets robotic manipulation: it is validated on 24 RoboCasa kitchen tasks, four LIBERO suites and LIBERO-Plus perturbed instances, and three real-world tasks (Bowl Stacking, Block Collection, Flower Arrangement) on the AgileX CobotMagic bimanual platform. The method directly reuses a pretrained VLA and the Cosmos-Predict2.5-2B checkpoint, with semantic and action streams initialized from the pretrained base model and the dynamics stream from a video-generation checkpoint, so it suits teams that already hold MoT-style VLA and video world-model weights. At inference, one future latent is sampled per policy query and layerwise contexts are cached, the action stream integrates predicted velocity with 10 Euler steps, and both context caches are refreshed for the next query; the authors state that future work will extend this action-centric design to other embodied tasks such as navigation.

Real-world evaluation uses 30 trials per task with a separately trained policy per task, so sample sizes are limited; on LIBERO-Plus the camera-viewpoint category is the lowest-scoring for both policies (47.1% versus 58.5%), indicating observation-geometry changes remain demanding. Inference latency is 423/418 ms on LIBERO/RoboCasa versus 161/160 ms for the baseline, and the authors note the Cosmos Policy figures of 610 ms and 160 ms are published timings rather than a controlled comparison. The attention analysis covers 300 rollouts and 21,307 policy queries, with event-related changes reported in percentage points and framed as possible interpretations. Both privileged self-distillation settings (65.0%, 63.3%) fall below the 71.2% reference, and the authors offer possible explanations rather than a settled account. Supplementary target-timing experiments favor consecutive near-future targets (71.2%) over later consecutive (65.4%) and approximately uniform (62.5%) selections.

Sources