SpatialSpeak's QA-style reconstruction pretraining lifts the spatial chain-of-thought gain from 2.6 to 6.9 points, reaching 62.8 on ReVSI
Synopsis
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
Interpretation
QA-RP makes spatial chain-of-thought supervision more effective: on ReVSI, CoT-VC yields only a 2.6-point gain without QA-RP, rising to 6.9 points with QA-RP. Prior work typically injects features from pretrained reconstruction models into VLMs or adds geometric training objectives; here reconstruction itself is recast as text-based question answering so geometric estimation and later reasoning share the same autoregressive output interface. ReVSI ablations: full method 62.8, without CoT-VC 55.9, without QA-RP 55.0, without both 52.4; the authors read this as QA-RP amplifying the effect of CoT-VC.
Local and global reconstruction supervision are complementary: removing global object-center queries drops the score to 59.9, removing local marked-point queries to 59.3, and removing both to 55.0. Local queries ground image points in 3D, while global queries enumerate visible object instances and predict their semantic categories and 3D centers, both expressed in the first-frame camera coordinate system, covering object-distribution information a single task does not. Ablation comparisons on ReVSI, where all three settings fall below the full method's 62.8, supporting joint learning of local geometry and global scene context.
Visual compensation (VC) provides a fallback when geometric derivations are unreliable: adding spatial CoT without VC scores 58.5, and adding VC further raises it to 62.8, an additional 4.3 points. The CoT target explicitly includes a reliability token and a visual-inspection fallback; when a 3D estimate is labeled Low, the model is trained to acknowledge the limitation and revise the answer from visual evidence rather than only emitting the geometric derivation. ReVSI ablations: 55.9 without CoT-VC, 58.5 with CoT but no VC, 62.8 full; the reliability threshold of 0.1 gives 62.8 versus 59.6 at 0.05 and 59.5 at 0.2.
With a 4B backbone it reaches the best results among compared methods on three spatial reasoning benchmarks: ReVSI 62.8 (8.7 points above the highest-scoring compared method, SpatialStack-4B at 54.1), VSI-Bench 73.0 (versus GeoThinker-8B at 72.6), and SPAR-Bench 76.0 (above SpatialStack-4B at 72.0 and GeoThinker-8B at 68.2). It obtains the best reported result on five of the seven ReVSI task categories, including object counting (64.9 vs. 36.3), room size (66.0 vs. 54.4), and relative distance (71.9 vs. 62.9). Main result tables span proprietary, open-source, and spatial-enhanced model families; VSI-Bench is reported under normal and scaled training settings, with SPAR-Bench breakdowns in the appendix.
Perspective
The result targets multi-view indoor spatial question answering: training data come from ScanNet 3D annotations and existing subsets such as LLaVA-Hound and SPAR-7M, evaluation centers on ReVSI, VSI-Bench, and SPAR-Bench, and reconstruction accuracy is assessed on ScanNet. The method suits settings that need quantitative spatial answers (distance, size, count, closest object) and relies on the first-frame camera coordinate system as a shared reference. For question types not covered by the geometric CoT templates (for example relative direction), the authors retain direct answer supervision to preserve coverage of the full task distribution. Extending reconstruction-augmented reasoning to broader scenes and task types is a natural next direction.
After GT Sim(3) alignment the reconstruction gap narrows (SpatialSpeak 5.0 cm versus CUT3R 4.7/4.6 cm), so the advantage under the no-alignment protocol and the aligned behavior are not fully aligned, and which downstream conditions make that difference matter remains an open question. The reliability threshold is compared only at 0.05, 0.1, and 0.2, with finer threshold behavior not expanded in the main text. In addition, SPAR-Bench breakdowns and several qualitative examples sit in the appendix while the main text reports summary scores, so readers wanting per-category judgment need the appendix.
