LEGO-Anything has coding agents rebuild 3D scenes from a single image as Blender code, with GPT-6-astra scoring 53.4% indoors and 39.6% outdoors and LEGO-Plugin lifting all six models by up to 62.7%
Synopsis
The work presents LEGO-Anything, an Image-to-Code framework in which a general-purpose coding agent reconstructs a 3D scene from a single image by iteratively writing, executing, and revising Blender programs, yielding an executable, inspectable, editable, and queryable scene program, together with LEGO-Bench, a simulator-grounded benchmark of 208 images rendered from 104 indoor and outdoor scenes that separately scores artifact validity, visible-surface geometry, and rendered appearance; among evaluated agents GPT-6-astra achieves the strongest overall results (53.4% indoor, 39.
Interpretation
Single-image 3D reconstruction is reformulated as Image-to-Code: instead of predicting a scene in one shot, the agent alternates between editing code, executing Blender, inspecting the evolving scene and its renderings, and revising the program, finally submitting an executable scene program with its export and a rendered view. Compared with meshes, point maps, or object sets as fixed 3D outputs, and with modular pipelines that compose perception, reconstruction, asset retrieval, and scene assembly, this representation makes objects, geometry, layout, and camera explicit so the result can be run, inspected, edited, and queried like code. The paper formalizes the setting and describes a construction trajectory of intermediate programs, scenes, and observations in Blender; Figure 2 illustrates the workflow as abstracted from a GPT-6-astra trajectory.
LEGO-Bench evaluates end-to-end reconstruction on 208 images rendered from 104 scenes across 8 environments, 17 themes, and 443 registered assets, scoring artifact validity, visible-surface geometry, and rendered appearance separately. Existing benchmarks focus on object- or part-level modeling, specify worlds through text rather than images, or rely on calibrated multi-view indoor observations; LEGO-Bench targets single-image, scene-level, executable reconstruction with indoor and outdoor coverage and controlled difficulty. The benchmark is built in LychSim with Fab assets, with candidate scenes passing collision and stability checks and human review; evaluation compares visible surfaces in camera coordinates without alignment or rescaling, appearance is computed from an evaluator re-render of the submitted scene as the fraction of pixels within a channel-error threshold, and a human study reports 86.6% overall compatibility agreement and 83.7% human-metric compatibility.
Evaluated coding agents reliably deliver valid artifacts but lag clearly in geometric and visual fidelity; GPT-6-astra achieves the strongest overall results at 53.4% indoor and 39.6% outdoor, and increasing scene complexity degrades fidelity rather than executability. The results separate delivering an evaluable artifact from faithfully recovering a scene: all six GPT configurations reach near-saturated validity while overall scores range widely from GPT-6-astra down to the strongest GPT-5.6 configurations, and among baselines Gen3DSR attains higher reconstruction than GPT-6-astra but produces no evaluable appearance. The main table reports means and standard deviations over three runs; as difficulty moves from Easy to Hard, validity stays stable while overall score declines, with the steepest drop in outdoor reconstruction.
Trajectory diagnostics identify weak initialization, regressive edits, and unreliable self-evaluation, motivating the training-free LEGO-Plugin, whose Enhanced Initialization, Version Control, and Grounded Refinement modules improve all six evaluated models on the 42-case Office subset with relative gains of up to 62.7%. The plugin plugs into existing harnesses as MCP tools, workflow skills, and runtime hooks rather than a separate planner, and uses only the reference image, never private benchmark ground truth; gains are inversely related to base performance, with the largest improvements for weaker agents. Diagnostics show 29.6% of GPT-5.6-sol edits decrease the score and its final submission trails its best intermediate scene by 3.2 points; across 36 builder-judge pairs, self-judgments agree with the deterministic metric direction at near or below chance on geometry.
Perspective
The work targets scene reconstruction from a single RGB image with an executable Blender program as output, aimed at researchers and system builders who want inspectable, editable, queryable 3D artifacts; LEGO-Bench's simulator-grounded pipeline lets users convert their own scenes and assets into new evaluation cases with the same ground truth and scoring protocol, and LEGO-Plugin attaches to existing harnesses as MCP tools, skills, and runtime hooks without changing the division of labor in which the agent plans and Blender MCP edits the scene. LEGO-World's readouts address downstream vision tasks that treat a frozen reconstructed scene as a representation of a natural image, and its conclusions are scoped to scenes reconstructed by GPT-6-astra without the plugin.
Worth watching: LEGO-Plugin is evaluated only on the 42-case Office subset, so its up-to-62.7% relative gain comes from that subset and full indoor-outdoor benchmark behavior remains to be seen; LEGO-World detection and segmentation use an equal-confidence protocol because scene exports provide no calibrated confidences, so AP should be read as a compatibility diagnostic under a fixed ordering; the depth comparison is relative-depth, and LEGO-Anything has 99 quality-valid outputs out of 100 images, a coverage difference from the direct baseline. In addition, the main table pools 90 ground-level outdoor views with 10 bird's-eye views, and the bird's-eye portion is reported separately in the appendix as a stress test rather than in the paired difficulty analysis.
