Where-OPD gives the teacher textual spatial hints over procedurally generated scenes, raising real-image averages by 3.23, 1.07 and 1.29 points across three MLLMs
Synopsis
The work introduces Where-OPD, an on-policy self-distillation scheme in which the teacher additionally receives textual spatial hints naming question-relevant visual elements and their coordinates while the student sees only the image and question; training data is generated procedurally from geometric counting scenes with no human annotation or external teacher, and across Qwen3.5-4B, Qwen3.5-9B and Qwen3-VL-4B it improves counting, document and chart benchmarks and lifts the average over CVBench, V*, ZoomBench, BLINK, HR-Bench and MME-RealWorld by 3.23, 1.07 and 1.29 points despite training only on synthetic scenes.
Interpretation
A new form of privileged information is proposed: the teacher receives textual spatial hints identifying question-relevant visual elements and their locations, rather than a cropped or zoomed view of the image. Prior multimodal on-policy self-distillation (Vision-OPD, Imagine-OPD, OPD-V, RP-OPSD, S2VOPD, VCSD) mainly creates the teacher-student gap through better visual observation; this work keeps the image unchanged and places the asymmetry in textual spatial guidance. The paper compares against Vision-OPD, OPD-V, Imagine-OPD and S2VOPD on Qwen3.5-4B, Qwen3.5-9B and Qwen3-VL-4B, and reports ablations: without privileged information OPSD stays at the base-model average, answer-only reaches 71.27, coordinates without the total reach 72.35, and coordinates plus count reach 72.94 on a seven-benchmark average.
The training data is generated entirely procedurally: scenes are built from colored geometric shapes whose identities, attributes and coordinates are known at generation time, so spatial hints can be constructed automatically without human annotation, region proposals or an external large model. Existing methods rely on human-annotated grounding data or external teacher models to locate relevant regions; here the hint is derived directly from the generator state, making post-training scalable and annotation-free. The main post-training set contains 3,000 image-question pairs; scenes hold 12 to 40 objects with radii of 14 to 64 pixels across seven shapes and twelve colors, and hints follow a fixed template, for example "Scanning for green crosses: found one near (377, 400), found one near (213, 805). Counting: 2 total."
Gains from synthetic-scene training transfer to real images and to visual capabilities beyond the training task. Prior zoom-based methods concentrate gains on tasks that benefit from localized visual evidence; this work reports improvements across counting, document, chart and general perception benchmarks. The paper reports a 3.23-point average gain over CVBench, V*, ZoomBench, BLINK, HR-Bench and MME-RealWorld, with 3.23, 1.07 and 1.29 points for the three models; on Qwen3.5-4B, ChartQA and EvoChart improve by 7.20 and 10.11 points and CountQA and OCRBench by 2.53 and 1.93 points; category-level results show a 24.70-point gain on the MME-RealWorld autonomous-driving subset.
Design choices are examined systematically: a frozen teacher outperforms EMA updates, 3,000 examples give the best average among tested sizes, and the default resolution and object count are best. The question of what the privileged information should be is decomposed into measurable ablations rather than reported only end to end. A teacher update rate of 0 (frozen) reaches 72.94 on Qwen3.5-4B and the average declines as the rate grows; data-size ablations show 3,000 examples are highest among tested sizes for both Qwen3.5-4B and Qwen3-VL-4B; scene-generation ablations show the default resolution and 12-40 objects per scene give the highest average.
Perspective
The result targets researchers and engineering teams who want low-cost post-training that improves visual localization and evidence integration in multimodal models; it applies where training environments can be generated procedurally with known object identities and coordinates. The paper uses counting questions as the supervision carrier because answering them requires identifying all matching instances, naturally covering multiple question-relevant image locations. No hint is used at inference, so deployment cost matches the base model. Reported training cost is about 1.4 hours for Qwen3.5-4B, 1.2 hours for Qwen3-VL-4B and 5.8 hours for Qwen3.5-9B on 4 H100 GPUs.
A careful reader would still ask how much the benefit of spatial hints depends on counting, a task that requires enumerating multiple locations, and whether the hint template carries over to tasks needing long reasoning chains or external knowledge. The paper notes that CountQA benefits from denser scenes, suggesting a trade-off between scene parameters and specific benchmarks, so the best configuration may shift with the target capability. Increasing the teacher update rate lowers the average, which the authors attribute to a possible domain shift on synthetic images, an explanation that invites further verification. In addition, only the Where-OPD results for Qwen3.5-4B and Qwen3.5-9B are averaged over three independent runs, while other post-training results, ablations and analyses come from a single run per setting, so the stability of individual differences would need more repetitions to confirm.
