EditHero benchmarks long-horizon 3D editing on 457 part-level chains: LLM agents preserve untouched parts better but need minutes per edit
Related research and updatesSynopsis
This work builds EditHero, a benchmark of part-level 3D edit chains in which a data engine assembles library parts onto segmented host objects and a JSON log rebuilds the exact target state after every turn, covering 457 chains, 2755 edits and 252 objects with 3 to 30 turns per chain; under a self-rollout protocol it evaluates 4 non-agentic methods (PartFlow, Nano3D, 3DEditFormer, VoxHammer) and 6 LLM/VLM agents, finding that non-agentic methods often miss the requested change from the first turn and accumulate errors along a chain, while agents preserve untouched regions better (CC 0.95–0.98 versus at most 0.90) and most follow instructions better (IF up to 0.62 versus at most 0.36), but take about 1.5–6 minutes per edit.
Interpretation
EditHero provides a part-level edit-chain benchmark with exact per-turn ground truth: a data engine assembles library parts onto segmented hosts with four operations (add, remove, replace, retexture) and records every operation in a JSON file that, together with the part library, rebuilds any turn's state exactly. Existing paired datasets provide only isolated single edits whose targets come from learned editing models or from pose and semantic-part changes, so they cannot tell whether a method preserves earlier changes across a sequence of instructions; EditHero is the first benchmark to evaluate different natural-language part edits in sequence on one asset with an exact 3D reference after every turn. After human review it contains 457 chains, 2755 edits and 252 hosts, with chains of 3 to 30 turns (median 6) and a median instruction of 17 words; each chain is stored as a JSON log of accepted operations with their library part, pose and texture, and rebuilds any state deterministically.
Under self-rollout, non-agentic methods often miss the requested change as early as the first turn, and errors accumulate along a chain while untouched parts keep drifting. Prior evaluations score a single edit applied to a clean source object and cannot expose the degradation that comes from inheriting one's own output turn after turn; this work makes each turn start from the method's own previous output and scores the edit region and the unchanged region separately. IF is only 0.12–0.32 over all turns and already 0.24–0.36 at turn 1; from turn 1 to turn 7 CC0 falls from 0.60–0.78 to 0.33–0.67; when every turn instead starts from the ground-truth previous state, IoU to the target at turn 5 rises from 0.52 to 0.68 for Nano3D and from 0.36 to 0.46 for 3DEditFormer.
LLM/VLM agents work bottom up and change only the named part, so they preserve untouched content better than every non-agentic method, and most also follow instructions better. Non-agentic methods work top down and regenerate the whole object from a learned 3D representation, changing nearly every corner each turn; an agent inspects the current mesh and writes code that changes only the named part, leaving the remainder untouched by construction. On a shared comparison subset of 55 chains and 416 turns, the 6 LLMs reach CC 0.95–0.98 against at most 0.90 for non-agentic methods; all except GLM 5.3 Flash also follow instructions better, with Opus 5.5 reaching IF 0.62 against at most 0.36; removal is easiest (IF 0.88–0.99) while additions and replacements are harder (IF 0.10–0.49 and 0.17–0.52).
Agentic methods remain too slow for interactive use, and building new geometry from code stays hard. The work reports editing quality together with the per-edit cost, so the fidelity advantage of the agentic route and its time price can be read side by side. An LLM needs about 1.5–6 minutes and 6–16 model calls per edit (medians), with Opus 5.5 at 149 s and 10.9 model calls per edit, while a non-agentic method takes 20–50 seconds on a local GPU.
Perspective
The benchmark targets multi-turn part-level editing that starts from a textured, part-segmented host object and is conditioned on a text instruction plus a fixed-view render of the target state, with operations limited to add, remove, replace and retexture and chains of 3 to 30 turns. Evaluation runs under self-rollout, where each turn takes the method's own previous output as input and ground-truth states are used only for scoring; all states of a chain share one coordinate frame normalized by the bounding box of its ground-truth states and 4 fixed views (1 conditioning view and 3 held-out views), with image metrics averaged over the 4 views. The authors state they will release the data synthesis engine and the edit chains for further 3D editing research.
The per-turn curves shrink to 80 chains after turn 7, and the authors draw no conclusions beyond that point; how much of the self-rollout decline comes from later instructions being harder versus from inherited errors is separated by a ground-truth-input diagnostic, but that diagnostic covers only 200 chains for Nano3D and 3DEditFormer. Material edits are not registered by geometry metrics and are instead measured with a changed-pixel diagnostic whose region is narrower than the part-based definition; Nano3D has no material mode, so its material turns run in replace mode and its material behavior cannot be read as material accuracy. Agent evaluation uses the tools provided by each runtime under differing budgets, with cost reported as median time, tokens and model calls per edit. In addition, the readable material here is the main text and appendices without the figures themselves, so numbers tied to specific figure numbers rest on the prose descriptions.
