Skip to main content
Back to timeline
arXivSource publication:

ExStereo lifts 2D VLA models to 3D with explicit stereo representations, raising simulated success from 51.2% to 76.4% and real-world success from 16/60 to 50/60

Related research and updates

Synopsis

The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.

Source-provided article image: ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations
Fig. 1 ·

Fig. 1: ExStereo is a stereo module that enables 3D spatial perception in pre-trained, off-the-shelf Vision-Language-Action models. Given a stereo image pair, it constructs four orthogonal views of the scene with metric 3D coordinates. This explicit 3D representation allows the models to selectively retrieve visual features relevant to the action tokens through our proposed action-stereo cross-attention mechanism, improving robotic manipulation performance.

arXiv

Interpretation

ExStereo lifts a pre-trained 2D VLA into a policy with metric 3D perception without re-pretraining the base model. StereoVLA previously required training its VLA from scratch and StereoPolicy gave only modest gains to pre-trained VLAs; ExStereo attaches as a plug-in module to off-the-shelf VLAs. Average success on ten RoboTwin tasks rises from 51.2% to 76.4%, and SmolVLA from 18.7% to 42.0%; real-world aggregate is 50/60 versus 16/60 for the baseline.

Using explicit stereo representations directly in the action expert outperforms concatenating implicit stereo features into the VLM. The paper characterizes StereoVLA and StereoPolicy as implicit representations and systematically compares explicit multi-view tokens fed to the action expert. In ablation, explicit features reach 66.7 on Lift Pot and 72.0 on Open Laptop versus 32.0 and 20.0 for implicit features; on real hardware StereoVLA reaches 11/60, below the RGB baseline.

A mid-training stage with a masked autoencoding objective lets the policy learn robust 3D representations from incomplete observations. An intermediate stage is inserted between pre-training and task post-training, trained on large-scale stereo data generated by StereoEngine, with the stereo module frozen during post-training to prevent overfitting. Mid-training with MAE reaches 61.3 on Lift Pot and 92.0 on Open Laptop, above 57.3 and 90.7 without MAE; removing mid-training drops real-world performance to 36/60.

Perspective

The results target bimanual manipulation settings with a calibrated stereo camera and a mostly tabletop workspace, and suit researchers and engineering teams who want fast metric 3D perception without retraining the base VLA. The method depends on a frozen Fast-FoundationStereo and on large-scale stereo data generated by StereoEngine; mid-training tasks do not overlap with evaluation tasks in either task or object, and post-training uses only 100 clean episodes, indicating the pipeline adapts with limited task demonstrations.

The stereo foundation model may degrade on distant, textureless, and thin structures, and those errors propagate into the multi-view observations; the point cloud comes from a single stereo viewpoint, so occluded regions leave holes across all four rendered views; multi-view rendering assumes a planar surface such as a tabletop to define the world frame, which may not hold in cluttered or non-tabletop scenes; mid-training data comes from a single benchmark and a single bimanual embodiment, so scaling to larger stereo robot datasets and more diverse embodiments remains open. In addition, mid-training degrades performance on some tasks, which the authors attribute to the choice of mid-training tasks and hypothesize could be mitigated by tasks more similar to post-training ones, a point still to be verified.

Sources