Vantage lets VLMs seek a question-relevant viewpoint before reasoning, lifting six models by 6.7% on average across five spatial benchmarks without fine-tuning
Synopsis
The work formulates a Seek-and-View paradigm and instantiates it as Vantage, a training-free, model-agnostic framework in which a VLM analyzes the question and plans a camera action relative to a reference view, a 3D foundation model reconstructs the scene and reprojects it to synthesize that view, and the synthesized view plus reasoning context augment the final answer; across six VLMs and five spatial reasoning benchmarks it yields consistent gains without fine-tuning, with a mean accuracy gain of 6.7%.
Figure 1: (a) Existing approaches reason over fixed views, suffering from fragile cross-view alignment and the geometry-to-language bottleneck. (b) Our novel Seek-and-View reasoning paradigm seeks a question-relevant view (see bottom right) to make the spatial evidence directly observable.
arXivInterpretation
It formulates Seek-and-View: instead of reasoning only over fixed sparse input views, the model first seeks a question-relevant viewpoint that consolidates spatial evidence distributed across views into one observation. Relative to prior View-and-Reason approaches that verbalize geometry or inject geometric priors, this paradigm makes where-to-look part of the reasoning, reducing reliance on language-based cross-view alignment. The paper supports the paradigm with experiments: consistent improvements across five benchmarks and six VLMs, with a mean accuracy gain of 6.7%.
Vantage realizes the paradigm in two stages: viewpoint-grounded reasoning for question analysis and view planning, then geometry-grounded evidence augmentation where a 3D foundation model synthesizes the new view and augments the final VQA with reasoning context. It decouples semantic camera intent from executable geometry: the VLM only selects among 17 predefined camera actions (rotation, translation, hybrid, each with three coarse magnitude levels), while the reconstruction backend determines the precise pose. Ablations show structured action selection outperforms direct 6-DoF pose prediction, random actions, and view interpolation; removing question analysis drops Qwen3-VL-4B from 38.9% to 30.0% on MindCube-tiny.
Synthesized view, question analysis, and reasoning guidance play complementary roles: the view alone improves average accuracy, adding textual context generally helps further, and the full combination performs best. This indicates the gains come from seeking question-relevant viewpoints rather than simply adding denser or more continuous observations. For Qwen3-VL-4B the full combination improves the average by 7.9 points over baseline (relative +24.7%); for Gemma-4-31B by 4.1 points (relative +8.7%).
An audit of 600 incorrect predictions finds residual errors concentrate in the viewpoint-grounded reasoning stage, namely incorrect analysis and incorrect view plan, rather than in synthesis quality. It locates the bottleneck of current multi-view spatial understanding in deciding what should be seen and how to reach it, not only in geometric reasoning capability. Fifty incorrect predictions were sampled per model across four models and three multi-view benchmarks, weighted by empirical error rate and macro-averaged; unreliable synthesis is rare and mostly occurs under reflections, transparency, severe occlusion, close-up views, darkness, blur, or low resolution.
Perspective
The results target multi-view spatial VQA in static real-world scenes, where answers are determined by geometric relations within the scene; evaluation covers MindCube-tiny, the Positional Relationship and Attribute subsets of MMSI-Bench, and the Multi-view Reasoning subset of BLINK, with additional evaluation on the Perspective Taking subset of OmniSpatial and the Dynamic Rotation and Dynamic Translation subsets of SPINBench. The method is training-free and model-agnostic, so it can be layered onto existing VLMs, suited to researchers and engineering teams who want to improve cross-view spatial question answering without task-specific fine-tuning; the synthesized view only rearranges pixels observed in the input views, leaving surfaces not covered by any input view empty, so it supplements rather than replaces the original observations.
Gains do not increase monotonically with model size, and on Qwen3-VL-8B view interpolation slightly surpasses the method in some settings, indicating that effectiveness is coupled to the capabilities of the underlying VLM and 3D foundation model. The paper attributes residual errors mainly to question analysis and view planning, but the failure categorization relies on manual annotation of 600 incorrect samples with error-rate weighting, so readers may want to judge the category boundaries and estimate stability against the protocol in the appendix. In addition, action magnitudes have only small, medium, and large levels, translation distances are defined in scene units, and absolute distances are not directly transferable across scenes given the non-metric scale of the reconstructed point cloud; the field-of-view choice also affects how much spatial context the synthesized view can capture. As a training-free framework, rollout quality varies with the underlying models, and end-to-end optimization is left as future work.
