Skip to main content
Back to timeline
arXivSource publication:

SolveEdit's 2,728 scene-transformation tasks leave the best model at 57.0% SolveScore, while a two-stage planner lifts GPT-Image-2 to 71.6%

Synopsis

The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.

AI-generated editorial illustration: SolveEdit: Benchmarking Visual Problem Solving in Generative Models

Interpretation

The paper formulates visual problem solving as scene transformation: given an image and a goal, a model must infer a valid transformation and realize it in the same scene without disturbing unrelated content, and builds the SolveEdit benchmark of 2,728 cases spanning 10 application domains and 54 subdomains. Existing evaluations tend to isolate perception, abstract reasoning, generation, or the execution of explicitly specified transformations, without isolating where the required transformation comes from; SolveEdit makes that information source an explicit evaluation axis with three non-overlapping regimes, IS, SD, and RD. Cases follow a fixed annotation procedure and quality review requiring the intended change to be semantically checkable, its evidence visible in the image, correctness statable without exact reference matching, and the change separable from protected content; image provenance includes 1,375 AI-generated inputs, 646 real photographs, 377 animation and game captures, and 330 film and cinematic frames.

The paper introduces atomic transition contracts and SolveScore: each case records required and protected conditions, and scoring measures completion while applying a bounded deduction for unintended changes, reporting six diagnostic scores across required and protected semantic (SA), relational (RA), and visual-quality (VQ) conditions. Reference-based scoring can penalize valid solutions that differ from the single authored output; contract-based scoring admits multiple valid renderings of the same transition and reports failure to accomplish separately from collateral damage. The canonical contract inventory contains 29,460 atomic criteria (14,831 required and 14,629 protected; 14,322 SA, 10,520 RA, 4,618 VQ), with a mean of 10.8 and median of 10 criteria per case and at least one protected criterion per case; evaluation uses a VLM plus fixed preassigned SAM and YOLO checkers, and abstentions do not redistribute weight.

Across 11 systems, the strongest model, GPT-Image-2, reaches only 57.0% SolveScore, with Seedream 5.0 Pro at 56.6% and Gemini 3.1 Flash Image at 55.5%; performance declines with information dependence, with RD trailing IS by 17.2 to 23.7 points and SD by 7.0 to 8.8 points, and after adjusting for domain composition the gap concentrates in required semantic (-25.0 points) and required relational (-13.0 points) conditions versus -5.6 points for required visual quality. This diagnosis separates solution-recovery errors (choosing an invalid transition) from visual-execution errors (a valid transition rendered poorly) and indicates that current failures are largely the former rather than rendering quality alone. Same 2,728-case manifest, same hidden contracts, and a deterministic scorer; GPT-Image-2 has the highest completion but more collateral damage than Seedream (23.9 versus 17.6), showing similar aggregate scores can hide different completion-preservation profiles.

The two-stage planner SolveEdit-Plan instantiates the transition before generation: Inspect finds unresolved variables and gathers targeted visual evidence, and Resolve compares candidate transitions and compiles one instruction with completion conditions and preservation obligations for a single generation call; under a matched single-generation budget it lifts GPT-Image-2 from 57.0% to 71.6%, 6.3 points above the budget-matched Generic Vision Rewrite, and transfers to Qwen-Image-Edit-2509 (20.8% to 26.9%) and HunyuanVideo-1.5 (7.4% to 14.0%). The intervention directly tests the RQ2 diagnosis: if failures are largely solution-recovery errors, explicitly determining the transition before generation should help; the planner has no learned parameters and leaves the editor unchanged, acting only through the instruction interface. The matched comparison holds the case manifest, output count, scoring contract, and final generation call count constant; relative to Generic Vision Rewrite, required completion rises from 72.8 to 75.8 and collateral damage falls from 16.3 to 8.6; planning adds two LLM calls per case, averaging 3,207 input and 635 output tokens.

Perspective

The benchmark targets generative visual problem solving that answers by changing the scene, applying to image editing and to video evaluated at the terminal frame; for video, the same contracts are applied to the final frame to test cross-modal transfer of the transition contracts, and temporal generation is outside scope. The planner targets a setting with one final generation call and an unchanged editor, and its gains hold under a matched call budget; for developers and evaluators who want to improve scene- and rule-dependent completion and reduce collateral changes without retraining a model, the contracts and planning interface are directly reusable.

The planning gains come with added test-time computation (two LLM calls per case), and equal call counts do not imply equal token use, so cost and latency accounting depends on implementation details; absolute scores for the open-source image editor and the video generator remain low (26.9% and 14.0%), and whether residual errors originate in planning or generation is not separated; evaluation is mostly VLM-based, with some criteria supported by SAM and YOLO evidence and one additional VLM call to reconcile disagreements, and abstentions reduce evidence coverage; additionally, this reading is of the full text, but per-category, per-regime, and sensitivity results in figures and appendices are described mainly in prose, so specific numeric details still warrant consulting the original and the release artifact.

Sources