Public articles linked to the same research event.
arXiv The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.
The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.
The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.
The authors introduce ExStereo, a stereo module that uses a frozen stereo foundation model to reconstruct a metric point cloud from stereo image pairs, renders it into four orthographic views (Top, Front, Left, Right) as an explicit stereo representation, encodes them with a ViT, and lets action tokens selectively read them through action-stereo cross-attention; after mid-training on large-scale stereo data with a masked autoencoding objective and freezing the module during post-training, it raises average success on ten RoboTwin simulation tasks from 51.2% to 76.4% and real-world success on three bimanual PiPER tasks from 16/60 to 50/60, outperforming point-cloud and implicit-stereo baselines.