SLR supervises continuous latent reasoning with a fingertip ray state plus parity-pooled visual targets, raising EgoPoint-Ground mIoU over same-backbone SFT by up to 21.1 points
Synopsis
The work proposes Spatial Latent Reasoning (SLR), which organizes pointing-gesture geometry and target-region appearance into an ordered sequence of continuous latent states: one spatial ray state supervised by fingertip position and pointing direction, followed by four states aligned with target-region features whose visual targets are built by parity pooling, which groups ROI tokens by row and column parity and averages each group; on EgoPoint-Ground it improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, with gains on both hard subsets, and reaches 77.6% precision at IoU 0.5 on YouRefIt.
Figure 1: Comparison of reasoning approaches on an EgoPoint-Ground example ( Li et al., 2026c ) . (a) Tool-assisted reasoning crops and re-encodes image regions during text generation ( Zheng et al., 2025 ) . (b) LVR ( Li et al., 2026a ) supervises continuous states by reconstructing annotated region features. (c) SLR combines one spatial ray state with four states supervised by parity-pooled target features. Image crops and ray overlays illustrate the reasoning operations and supervision targets; generic latent boundary markers are omitted for clarity.
arXivInterpretation
SLR factors pointing grounding into an ordered sequence of continuous latent states: one spatial ray state supervised by fingertip position and pointing direction, followed by four states aligned with target-region features; all states are generated recurrently at training and inference, and auxiliary annotations enter only the losses. Prior region-grounded latent reasoning (e.g., LVR supervising visual-feature reconstruction with teacher-forced region tokens, RIS combining box and semantic supervision) does not directly supervise hand geometry; SLR brings the relation of how the observed hand indicates the referent into intermediate supervision and orders geometric and visual evidence. The method specifies the ray origin and orientation, a lightweight readout head with LayerNorm and two linear maps, a spatial loss combining SmoothL1 over fingertip coordinates with directional cosine distance, and validity masking that skips only the spatial term for invalid finger geometry; ablations show the joint spatial-plus-visual configuration attains the best mIoU on the standard and Hard-Similar sets.
Parity pooling converts a variable-size target region into four phase-specific feature targets by grouping ROI tokens by global row and column parity and averaging each group, with phase indices anchored to the full image grid and fixed as the ROI changes. Against quadrant grouping (contiguous regions), random grouping (changed group membership), and Max (changed aggregation on parity supports), parity pooling exceeds them in standard-set mIoU by 3.6, 2.9, and 1.8 points and leads at all three thresholds, so the advantage is not explained by a larger latent-step budget. The appendix gives structural properties: each ROI token contributes to exactly one descriptor, empty phases are omitted without duplicating tokens, and phase means recover the overall ROI mean when weighted by counts; it also notes phase-wise averaging is lossy and does not reconstruct the feature grid or imply translation invariance.
On EgoPoint-Ground, SLR achieves higher mIoU than same-backbone supervised fine-tuning in all nine backbone-evaluation combinations, with standard-set gains of 2.8 to 21.1 percentage points; on Qwen2.5-VL-7B and Qwen3-VL-8B it exceeds the strongest reported reasoning baseline in standard-set mIoU by 4.3 points in each case. The gains appear over both the direct-prediction baseline and explicit-reasoning baselines, and on Qwen3.5-4B P@0.7 rises from 0.761 to 0.811, indicating the improvement includes boxes meeting a stricter overlap criterion. Results are grouped by backbone to distinguish framework contributions from underlying-model differences, with human performance reported separately as a reference; the authors note aggregate IoU metrics do not separate referent-selection errors from box-placement errors.
On YouRefIt, SLR gains over the same-backbone zero-shot model are similar across IoU thresholds 0.25/0.5/0.75 (8.4/8.0/8.1 points), reaching 77.6% precision at IoU 0.5. The gain persists at the strictest threshold, supporting improved precise grounding on this benchmark; the authors state this comparison measures the trained framework against the pretrained model and does not isolate the contribution beyond task-specific SFT, and literature scores retain their original protocols. On this benchmark SLR's 77.6% precision is a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols.
Perspective
The result targets the single-image, two-dimensional pointing-grounding setting: input is an image plus a language query, training may use finger-base and fingertip annotations and the target box, and inference generates five recurrent states and greedily decodes the box without any ray-object intersection. Teams using comparable multimodal backbones that can supply hand keypoints and target boxes can reuse this supervision directly; the authors list depth and temporal supervision, weaker annotation requirements, and scene-adaptive states as future directions. The method builds on LVR's continuous latent-state feedback and retains Coconut's generic boundary markers.
Readers should still watch: aggregate IoU does not separate referent-selection errors from box-placement errors; on Hard-Complex, Target-only leads the full model (0.597 versus 0.551), showing the benefit depends on scene conditions; on Qwen3-VL-8B Hard-Similar, Text CoT has higher mIoU (0.353 versus 0.315), so the advantage over explicit reasoning is setting-dependent; attention maps illustrate complementary state roles but do not establish a causal reasoning mechanism; the authors also note configuration comparisons do not isolate the causal contribution of a specific geometric mechanism and that repeated runs and broader evaluations would assess robustness. In addition, reported times include device transfer, hashing, and scoring overhead rather than pure decoding latency, so latent-state count and inference speed cannot be equated directly.
