EviRover teaches a 4B model to look beyond a glance: 30-point average gain on EviLens and 15 points on BrowseComp-VL
Synopsis
The work defines perception under insufficient evidence, reformulates perception tasks such as grounding, segmentation and counting as an agentic evidence-seeking process, builds two data generation pipelines yielding EviRover-SFT-5K and EviRover-RL-12K plus the human-verified 688-instance EviLens benchmark spanning five perception categories, and shows that a 4B model trained with supervised fine-tuning followed by agentic reinforcement learning improves over its Qwen3-VL-4B-Instruct backbone by 30 points on average on EviLens, reaches performance comparable to advanced proprietary models, and transfers to WebEyes, ReasonSeg, RefCOCOg, MMMU, MMMU-Pro, MathVerse and BrowseComp-VL with a 15-point gain.
Interpretation
The paper identifies a neglected class of perception cases in which the evidence needed to resolve a query is not available from a single glance, which it terms perception under insufficient evidence. Grounding, segmentation and counting have largely been formulated as one-shot predictions from an image-query pair, assuming image content and parametric knowledge suffice; this work makes the perceptual output itself the objective of evidence seeking rather than treating search as an aid to textual answering. The problem is argued and formalized as a task setting, supported by the observation that existing prompt-based workflows rely on manually designed inference-time strategies and that prior search-based segmentation agents are limited to segmentation and assume missing evidence is always external knowledge.
To support the setting, the authors design two data generation pipelines that yield the EviRover-SFT-5K and EviRover-RL-12K training sets, and construct the human-verified EviLens benchmark. One pipeline targets evidence present in the image but not resolvable at a glance, collecting high-resolution images whose targets often occupy well under 0.1% of the image area plus I-spy and spot-the-difference scenes; the other targets evidence beyond the image, verifying identities in real group photographs or synthesizing group photographs from a single reference with GPT-Image-2 and filtering with Seed-2.0-Pro, then rewriting queries into multi-hop form through iterative entity replacement. EviLens contains 688 instances covering 140 localization, 182 recognition, 15 spot-the-difference, 195 segmentation and 156 counting instances, with 79 annotated differences in the spot-the-difference subset, and every instance is manually verified; training set sizes are supported by the stated 5K and 12K designations.
EviRover is described as the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. At each step the model decides whether further evidence is needed and which action can provide it before producing a location, a boundary or a count; the RL stage uses EMA-GRPO to normalize rewards across heterogeneous tasks and a sequence-mean-token-mean reduction to avoid long rollouts dominating the policy update. The method description lists ten tools (text search, text-to-image search, image search, web browsing, crop, verify, verify_part, SAM3 mask generation, left-right comparison and a Python interpreter), with training on a single node of 8 NVIDIA A800-80GB GPUs using FSDP and evaluation capped at 35 tool calls per trajectory.
Experiments show EviRover improves over its backbone by 30 points on average on EviLens and transfers to several external benchmarks. The 4B model achieves the best localization and recognition results among evaluated models, with localization IoU of 0.444 versus 0.373 for Gemini-3.5-Flash; on WebEyes it improves grounding by 14.4 IoU and segmentation by 24.7 gIoU; it beats its backbone on ReasonSeg, RefCOCOg, MMMU, MMMU-Pro and MathVerse, and raises BrowseComp-VL accuracy from 9.5 to 24.8. Results are reported in tables spanning proprietary models, open-source segmentation and grounding models, and open-source general models; the authors also report that EviLens is broadly difficult, with segmentation specialists reaching at most 0.343 gIoU and grounding specialists near-zero IoU on localization.
Perspective
The result targets perceptual queries that require evidence beyond a single glance, covering localization, recognition, spot-the-difference, segmentation and counting, and is evaluated on EviLens, WebEyes, ReasonSeg, RefCOCOg, MMMU, MMMU-Pro, MathVerse and BrowseComp-VL. For practitioners this means deciding when and what to search can be trained rather than hand-prompted at inference time, and the released code, models and data provide a base for further extension; the paper notes that accurate perception underpins downstream applications such as embodied manipulation and image editing.
The spot-the-difference category of EviLens contains only 15 instances with 79 annotated differences in total, so conclusions on counting and spot-the-difference rest on relatively limited samples; synthesized group photographs of real individuals contain exactly one verified identity, and identity coverage varies across sources; training and evaluation cap each trajectory at 35 tool calls, leaving behavior under longer horizons open; and although this reading covered the full text, some numeric details in tables and appendices remain as given in the original paper.
