Skip to main content
Back to timeline
arXivSource publication:

BeyondCSe lets a robot grasp "the thing I just used": 77% success on occluded targets, 95% under heavy occlusion with mean views cut from 3.35 to 2.20

Synopsis

The work presents BeyondCSe, a zero-shot robotic grasping system that uses off-the-shelf multimodal large language models and a point tracker to recover a 3D event prior for an event-referenced target and, when the target is occluded, combines that prior with current scene geometry to actively select camera viewpoints; in real-robot experiments with a single wrist-mounted RGB-D camera it reaches 76% and 77% grasp success for initially visible and initially occluded targets (versus 40% and 55% for the strongest baseline), and on four heavily occluded scenes it raises success from 75% to 95% while reducing mean views from 3.35 to 2.20.

AI-generated editorial illustration: Beyond the Current Scene: Event-Referential Grasping with Active View Selection

Interpretation

The system turns the "event-referential" instruction type into an executable grasping pipeline: the instruction specifies a target by the role it played in a past event (for example the object moved first or last, a flower stem, a cup handle) rather than by name or appearance, and video reasoning first uses an MLLM to produce a target appearance description and an event cue, then locates the target in the initial image and lifts it to a 3D action point. Prior language-guided grasping methods (LERF-TOGO, GraspSplats, Point2Act, GraspMolmo) resolve targets within the current scene by name, category, appearance, or part; a target defined by its role in a past event cannot be identified from the current scene alone, so this work makes the event history the source of both identity and spatial evidence. Real-robot experiments use a ROBOTIS OMY-F3M with a wrist-mounted Intel RealSense D435i; the visible condition has 10 scene-query pairs and the occluded condition 10 scenes with two queries each, five trials per pair, 150 trials total; localization succeeds on 43/50 visible and 88/100 occluded targets, exceeding the strongest baselines by 32 and 20 percentage points.

When the target is not visible from the initial view, the system recovers a 3D track of the target from the event history and fuses it with current scene geometry into a volumetric probability density over target location, then scores candidate viewpoints with transmittance-aware visibility and performs Bayesian belief updates after each observation. Existing active-perception methods derive view scores from current-scene geometric information gain, predicted grasp affordance, or object-occluder relations; this work instead builds an instance-specific prior from the target's own past observations and attenuates confidence along sight lines through unobserved space rather than treating that space as free. The method builds a TSDF map, constructs a Gaussian prior centered at the last observation and shaped by the terminal path, normalized over the admissible workspace, with a mixture weight that keeps nonzero probability so search can recover from an inaccurate historical estimate; ablations show removing transmittance attenuation raises mean views to 3.10 (S3 needs 5.2 views), and removing the event prior leaves success unchanged but raises mean views from 2.20 to 4.00.

On four heavily occluded scenes where every queried target is initially out of sight (concealed inside a basket or behind other objects and invisible from a fixed bird's-eye view), the system achieves 95% grasp success with 2.20 mean views, versus 75% success and 3.35 views for an active-perception baseline that is given the target's ground-truth 3D bounding box. The comparison indicates that even with a ground-truth box removing localization ambiguity, the box neither reveals occluded target geometry nor guarantees grasp success, whereas the event prior provides spatial guidance that reduces the number of search views. All methods use the same scenes and queries across S1-S4 and stop at the first feasible target grasp; ablation variants keep candidate views, feasibility checks, target confirmation, and the grasp pipeline fixed, changing only belief weights, transmittance, or the event prior.

The same pipeline transfers to external egocentric video without dataset-specific prompt changes: on clips from EgoDex and EPIC-KITCHENS it selects targets referred to by the order of plate placement or washing and produces 2D points on the requested objects in the final frames. This illustrates use beyond the authors' own capture setup, while the authors note that in the EPIC-KITCHENS example, which includes large camera viewpoint changes, region estimation in the final frame can remain imprecise. This is a qualitative demonstration; the authors explicitly frame it as illustrating use beyond their capture setup and state that region estimation can remain imprecise.

Perspective

The result targets tabletop settings in which a person first interacts with objects and the robot later executes an event-referential instruction, with a single wrist-mounted RGB-D camera and perception models running on one RTX 4090 workstation; the method uses off-the-shelf MLLMs and a point tracker without task-specific training, so it can plug into existing grasp generation and planning stacks such as AnyGrasp. For readers building home or laboratory assistants, this means a robot can still grasp when the target has left the initial view, following instructions like "the thing I just used," and can search with fewer views under heavy occlusion.

The system assumes the target remains stationary during search, and the authors list interactions in which objects continue to move as future work; validation on external egocentric video is qualitative, and the authors note final-frame region estimation can remain imprecise. In addition, this reading covers the paper's main text and abstract; the specific image content of Figures 3, 5, 6, 7, and 8 is not expanded in the text, so per-scene visual detail and failure modes still need confirmation from the original figures.

Sources