Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ExStereo lifts 2D VLA models to 3D with explicit stereo representations, raising simulated success from 51.2% to 76.4% and real-world success from 16/60 to 50/60

The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.