Skip to main content
Back to timeline
arXivSource publication:

ReImaGin uses image generation as a multimodal reasoning tool, beating text-only reasoning and specialist vision-tool baselines by up to 25% across six visual reasoning tasks

Synopsis

The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.

AI-generated editorial illustration: Reasoning with Image Generation

Interpretation

ReImaGin replaces a fixed suite of specialist vision tools with a single instruction-following image generation model that serves as a free-form visual operation interface inside the multimodal LLM's reasoning loop. Prior tool-augmented methods such as Visual Sketchpad rely on pre-implemented narrow operations like cropping, depth estimation, or object detection and cannot perform open-ended generation or transformation; ReImaGin's generate_image accepts natural-language prompts and can remove an occlusion or generate a floorplan from multiple disjoint views of a room, without task-specific training. The paper uses the same image generation model (primarily Nano-Banana-Pro) across all six tasks and provides a full reasoning trajectory on collision prediction: the agent calls generate_image to draw a forward vector from the nose of the aeroplane and circle the first object it hits, then reads off the answer.

Across six visual reasoning tasks and three multimodal LLM backbones, ReImaGin consistently outperforms text-only reasoning and the specialist vision-tool baseline. The paper reports gains of up to 25% on path tracing and about 40% relative on occlusion counting; on puzzle completion Visual Sketchpad even falls below the no-tools baseline (16.7% vs. 29.5% with GPT-5), indicating that fixed specialist tools struggle on tasks requiring generative capability. Results cover Gemini-3.1-Pro, GPT-5, and the open-weights Qwen-3.5-27B, with multiple-choice accuracy (sMAPE for occlusion counting) averaged over 3 seeds with standard errors; an ablation replacing generate_image with a no-op that returns the input image unchanged, holding prompts, tool definitions, and the code environment fixed, degrades all six tasks, with the largest drops on path tracing and puzzle completion.

The faithfulness of generated images is strongly associated with final answer correctness and can be improved by test-time scaling. The authors manually audited generated images across all six tasks and found that most generations are faithful, with correctness markedly higher when the image is faithful than when it is not; on spatial reasoning, sampling multiple generated images and having a multimodal LLM select the most faithful one raises Gemini-3.1-Pro from 51.0% to 59.0% and Qwen-3.5-27B from 42.0% to 54.3%, while the textual reasoning process is unchanged. The authors explicitly label this relationship as correlational rather than causal; the test-time scaling experiment provides complementary evidence because the extra compute is spent only on image sampling. The other five tasks show little variance across generated samples, so test-time scaling is not applied there.

Effective visual reasoning strategies can be discovered automatically and can transfer across multimodal LLMs. Previously the strategy was specified through handcrafted in-context examples, requiring per-task human effort; the paper formulates strategy discovery as an iterative search over agent prompts, where a proposal model generates candidate prompts, evaluates them on a development set, and refines based on observed successes and failures. Discovered strategies often resemble handcrafted ones (e.g., generating a depth map for depth reasoning, drawing an arrow along the predicted trajectory for collision prediction) and recover most of the gains of handcrafted strategies. With Gemini-3.1-Pro and Nano-Banana-Pro, the automatic strategy beats the no-tools baseline on all five tasks and consistently beats the no-strategy setting; with Qwen-3.5-27B and Nano-Banana-2 it beats no-tools on four of five tasks (tying on MMSI). Transferring Gemini-discovered prompts unchanged to Qwen still beats Visual Sketchpad and no-strategy on collision, spatial reasoning, and path tracing, though gains shrink on occlusion counting and depth.

Perspective

The framework targets multimodal visual reasoning tasks that demand spatial and physical intuition, and is meant for researchers and engineering teams who want to add open-ended visual operation capability to a multimodal LLM without task-specific training; its gains scale with the underlying image generation model, and the paper shows open-weights generators are already competitive on local editing-style transformations. Automated strategy discovery applies to tasks with a small training and development split where the strategy can be expressed through a prompt, with a reported one-time discovery cost of roughly $214 per task.

Most numeric cells in the paper's main tables are blank in the loaded text, so beyond the numbers stated explicitly in the abstract and body (up to 25% on path tracing, about 40% relative on occlusion counting, 51.0% to 59.0% and 42.0% to 54.3% on spatial reasoning, 16.7% vs. 29.5% on puzzle completion, and 94.6% vs. 91.7% vs. 99.2% on depth), the full per-task per-model comparison cannot be restated here. The relationship between generation faithfulness and answer correctness is labeled correlational by the authors and still needs stricter causal verification. The automatically discovered strategy slightly underperforms the handcrafted one on path tracing because of extra coloring, and the composite MMSI strategy is not recovered under the current search setting, so the search space and generator reliability remain open questions. Evaluation task sizes are also limited (e.g., 124 BLINK validation instances, 100 randomly sampled CAPTURe instances, 100 MMSI instances, 100 programmatically generated path-tracing instances), so generalization to larger and more realistic settings remains to be seen.

Sources