Skip to main content
Back to timeline
arXivSource publication:

SpatialCORE folds both box accuracy and box confidence into the reward, letting an 8B model top same-scale open-source and specialized spatial-reasoning models on OmniSpatial and SpatiaLab

Synopsis

SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.

AI-generated editorial illustration: SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Interpretation

The paper introduces SpatialCORE, which turns the model's own confidence in its generated grounding into a learning signal for spatial reasoning, centered on a self-regulating spatial reward that multiplies each predicted bounding box's matching quality by a confidence weight derived from coordinate-token uncertainty (normalized entropy). Earlier grounded spatial-reasoning methods mainly optimize final-answer correctness, or additionally reward whether localization is correct, but none makes how certain the model is while producing a box part of the reward; the paper notes that when the answer is correct, a final-answer reward reinforces the entire trajectory, including its uncertain grounding. The method is specified with full formulas: the normalized coordinate-token entropy definition, a confidence floor, confidence-weighted matching quality, a pseudo-GT-validity-weighted soft recall term, and an F-beta harmonic mean combining them; Appendix B uses a GRPO gradient analysis to show that under answer-only supervision coordinate tokens from trajectories with the same answer receive the same advantage, and that confidence weighting breaks this indifference.

The paper couples the spatial reward to final-answer correctness through an answer gate: the full spatial reward applies when the answer is correct, while positive spatial rewards are reduced by a coefficient when the answer is incorrect, so incorrect-answer trajectories retain partial credit for useful grounding. This differs from rewarding only correct answers and from supervising grounding separately from the reasoning trace; the paper states the aim is to let grounding improve before the model consistently predicts the correct final answer. In the ablation, removing the answer gate lowers accuracy from 56.33% to 54.67%, which the paper reports as a 1.66-point contribution; the reward configuration sets the spatial answer gate coefficient to 0.3.

On the held-out OmniSpatial test set (1,533 samples, 4 reasoning dimensions, 50 subcategories), SpatialCORE-8B achieves the highest weighted-average accuracy among open-source and specialized spatial-reasoning models, surpassing its GRPO-trained backbone, with the largest gains in traffic analysis and localization, tasks that require comparing the positions of multiple task-relevant objects at once. The paper reports particularly large margins over the specialized spatial-reasoning models VST-RL-7B and SpaceThinkerQwen2.5VL-3B, which it reads as evidence that confidence-aware grounding provides a stronger training signal than grounding supervision alone; SpatialCORE-4B also remains competitive on average accuracy. Results are given in a table of per-subcategory and weighted-average accuracy across methods on OmniSpatial; the paper also reports comparisons against the backbone and a standard GRPO variant, and notes that allocentric and hypothetical reasoning show smaller gains because they require reasoning across multiple viewpoints, which predicted boxes from a single egocentric view cannot address.

On the unseen SpatiaLab (1,400 visual question-answer pairs from realistic scenes, 6 spatial-reasoning categories), SpatialCORE-8B leads open-source and specialized models zero-shot, with the strongest gains in relational positioning; grounding analyses show the lowest-uncertainty quartile attains the highest mean matched IoU, the entropy-IoU correlation shifts from -0.32 to -0.61, answer accuracy rises from 43.90% to 48.46%, and bounding-box coordinate uncertainty falls from 0.42 to 0.31. The paper links zero-shot transfer to improved alignment between confidence and localization quality, indicating that the self-regulating spatial reward instills a grounding behavior beyond the training distribution; it also notes smaller gains in size and scale estimation, where scene-level scale estimation remains a distinct challenge. Zero-shot evaluation uses different task designs and visual distributions from training; the alignment analysis reports matched IoU by uncertainty quartile together with the entropy-IoU correlation and the accuracy and uncertainty values; the pseudo-GT quality audit found 251 of 262 boxes correct (95.8%), with all observed errors in the two lower Grounding DINO confidence ranges.

Perspective

This work targets post-training for spatial reasoning where grounding is generated as bounding boxes, in question-answering settings where task-relevant objects can be boxed and reliable pseudo-GT boxes can be constructed; the paper instantiates it with Qwen3-VL-8B-Thinking and Qwen3-VL-4B-Thinking backbones, post-trains on the OmniSpatial training split, and evaluates on the held-out OmniSpatial test set and SpatiaLab. For readers who want to bring confidence signals into multimodal reinforcement learning, it provides a directly reusable reward construction: coordinate-token entropy, pseudo-GT validity weighting, Hungarian matching, and an answer gate. The paper suggests future work could add depth, multi-view context, or geometry-enhanced encoders to address 3D and non-boxable spatial concepts.

The gains concentrate on tasks that require comparing the positions of multiple task-relevant objects, while allocentric and hypothetical reasoning and size and scale estimation improve less; whether these directions are limited by the single egocentric view input remains an open question. Pseudo-GT boxes are generated offline by Grounding DINO and manually audited, so the training signal depends on that reference quality; the 40% corruption experiment still beats the unadapted backbone, but accuracy falls 1.76 points relative to original pseudo-GT, so the relationship between reference quality and final gains deserves continued observation. Evaluation is also concentrated on two benchmarks, OmniSpatial and SpatiaLab, leaving performance under other task designs and in closed-loop real-robot settings to be verified.

Sources