GAM swaps language and video pretraining for a frozen 3D grounding backbone, lifting RoboTwin 2.0 randomized-scene success from 30.4% to 47.6%
Synopsis
The work proposes Grounded Action Model (GAM), which builds a robot foundation model on a pretrained promptable 3D grounding model (WildDet3D), maps language, 2D point, or 2D box prompts into a shared object-centric representation (image tokens keeping target visual features plus detection tokens encoding point clouds, metric positions, and extents), and uses a multi-stream transformer (MM-DiT) with robot state history to predict action chunks; it reaches 55.3% average success across 50 RoboTwin 2.0 tasks (vs. 52.0% for Spatial Forcing), 47.6% under scene randomization (vs. 30.4% for Abot-M0), a state-of-the-art 61% average across 16 LIBERO-PRO perturbation settings (vs. 53%), 17/20 successes under visual shift on a bimanual YAM (vs. 4/20), and 64.7% in-distribution and 49.
Interpretation
It proposes 3D grounding as a pretraining foundation for robot action learning and instantiates it as GAM: the policy is built on a pretrained promptable 3D grounding model, WildDet3D, rather than a language-generation or video-generation backbone, and the grounding backbone stays frozen during policy training while only the action head is optimized. Dominant VLA and WAM lines inherit language or video pretraining, whose objectives do not directly require explicit metric grounding, so object localization must be learned implicitly from robot demonstrations; GAM makes 'which objects matter, where they are, and what geometry they occupy' an explicit backbone output. The paper gives a full method description: language is parsed by a span-tagging head over a frozen Flan-T5, points and boxes specify objects directly, and WildDet3D outputs 2D boxes, metric 3D boxes, and a dense depth map; the action head is a 12-block MM-DiT trained with flow matching, optimizing only the action head. This is a method-level argument, with effectiveness evidence in the benchmarks and real-robot experiments.
A unified task-specification interface: language, 2D points, and 2D boxes are all mapped to the same object-centric representation, so GAM can run autonomously or serve as a low-level controller driven by a high-level planner (such as Molmo2) through the point interface, supporting long-horizon and memory-dependent manipulation. Prior work often conditions on language alone or handles geometry and semantics separately; GAM links semantic reasoning and spatial control through one grounded observation, with the planner updating target selections at sub-task transitions while the action policy stays unchanged across sub-tasks. On a Franka composed with a Molmo2 planner through the point interface, four tasks (two long-horizon, two memory-dependent) reach 64.7% in-distribution and 49.8% out-of-distribution average step completion, versus 42.4%/24.0% for and 35.3%/17.1% for MolmoAct2, over 20 in-distribution and 10 out-of-distribution trials per task.
Clear advantages on two generalization axes, randomized scenes and relocated or newly designated targets: 47.6% on the RoboTwin 2.0 randomized setting (vs. 30.4% for Abot-M0) with the action policy trained only on clean-scene demonstrations, and 0.61 average across 16 LIBERO-PRO perturbations (vs. 0.53 for ), with the largest gains on the Pos and Task perturbations that change layout or goal. Baselines stay strong under Obj and Sem perturbations that leave layout and goal intact but collapse below 0.11 under Pos and Task, indicating their success tracks whether the memorized training-layout trajectory remains valid; GAM's observation is defined by what is grounded, so a change of target is reflected in the input. RoboTwin 2.0 uses the official single-task protocol with 50 tasks and 100 rollouts per task; LIBERO-PRO covers four suites and 16 settings, with baseline numbers from the official leaderboard and Flex- evaluated by the authors. GAM trails under Obj perturbations, which the authors attribute to object-size changes shifting the appropriate grasp locations.
Ablations show the gains come from object-centric filtering and complementary streams: removing image masking or point cropping drops average success from 46.8% to 22.3% and 12.8%, removing both drops it to 10.3%, and keeping only detection tokens or only image tokens yields 16.0% and 20.3%. This separates 'a stronger perception backbone' from 'how the backbone's outputs are selected and represented': all variants share the same frozen backbone, so differences come from observation construction, and metric geometry and appearance are shown to be complementary. Evaluated on 10 RoboTwin 2.0 tasks with 20 rollouts per task in both Easy and Hard, holding the backbone, action head, training data, and schedule fixed.
Perspective
The work targets robot manipulation policies that must keep acting on designated objects when scene appearance changes or targets are relocated or newly designated, and it is evaluated on simulation benchmarks (RoboTwin 2.0, LIBERO-PRO) and two real robots (bimanual pick-and-place on a YAM, long-horizon and memory-dependent tasks on a single-arm Franka). It lets follow-up work swap the grounding backbone under the same interface, plug in different high-level planners, or replace 3D boxes with finer-grained geometric representations; for researchers and engineering teams wanting to connect semantic planning with spatial control, the paper provides a reproducible observation-construction and training recipe (frozen backbone, action-head-only training, flow matching).
The paper states that GAM depends on the quality and completeness of its grounded observations: grounding errors propagate to actions without recovery, and object-centric filtering can omit relevant context such as unselected obstacles. The authors call for more diverse 3D grounding datasets and models that generalize across objects and scenes, and plan to explore finer-grained representations such as dense 3D instance segmentation alongside grounding-aware failure recovery and adaptive context selection. A reader would also watch how far the authors' explanation holds that GAM trails some baselines under Obj perturbations (object appearance and size) because the policy may struggle to adjust grasping motions to sizes not seen in training, and what the cost is of adapting the grounding backbone alone to a new domain.
