StepCAD reconstructs meshes into executable CAD programs with a state-conditioned policy plus IoU-guided tree search, improving relative IoU by up to 87.2%
Related research and updatesSynopsis
The work introduces StepCAD, a generative optimization approach that combines a state-conditioned CAD policy with geometry-guided search: given an input mesh, the policy predicts construction actions conditioned on both target and intermediate geometry, and an IoU-guided tree search refines the resulting program through local edits; it also releases ARCADE-1.5M, a dataset of 1.5M executable CAD programs with sequences up to 150+ counted operations and 12.5M intermediate state-action transitions; across multiple CAD reconstruction benchmarks StepCAD achieves state-of-the-art geometric reconstruction accuracy with consistently high validity, yielding up to 87.2% relative IoU improvement over the strongest evaluated baseline, with particularly large gains on complex shapes.
Figure 1: ARCADE-1.5M: large-scale synthetic CAD program generation. Our pipeline synthe- sizes diverse long-horizon CAD programs through stochastic strategy-guided composition of sketch primitives and CAD operations. Each program is executed and validated, yielding 1.5M executable programs and 12.5M intermediate construction states with broad operation diversity.
arXiv · Page 4Interpretation
StepCAD formulates mesh-to-CAD-program reconstruction as generative optimization: a state-conditioned policy predicts construction actions and an IoU-guided tree search refines the program through local edits. Unlike many learning-based methods that predict complete programs in a single pass and rely predominantly on sketch-extrude representations, this framework keeps opportunities to correct geometric errors during reconstruction and broadens operation diversity. The abstract states that the policy is conditioned on both target and intermediate geometry and that search is IoU-guided with local edits; search width, edit operators, and policy architecture are not detailed in the provided text.
It releases ARCADE-1.5M, a large-scale dataset of 1.5M executable CAD programs, sequences with a maximum length of 150+ counted operations, and 12.5M intermediate state-action transitions. The dataset spans diverse operations and explicitly records intermediate state-action transitions, providing resources for state-conditioned policies and search-based refinement. Scale and composition are given directly in the abstract; data provenance, licensing, and splits are not described in the provided text.
Across multiple CAD reconstruction benchmarks, StepCAD achieves state-of-the-art geometric reconstruction accuracy with consistently high validity, yielding up to 87.2% relative IoU improvement over the strongest evaluated baseline, with particularly large gains on complex shapes. Relative to prior single-pass prediction methods, the result reports accuracy gains together with maintained validity and indicates that gains grow with shape complexity. Evidence consists of multi-benchmark experiments and a relative IoU improvement figure; benchmark names, the baseline set, validity metric definitions, and per-category results are not listed in the provided text.
Perspective
The work targets recovering executable CAD programs from 3D meshes, suited to modeling pipelines that need operation diversity and error correction during reconstruction; its gains are especially large on complex shapes, indicating a focus on objects with many construction steps and rich geometric detail. The release of ARCADE-1.5M lets state-conditioned policies and search-based refinement be trained and compared on shared data, providing a basis for follow-up work on longer operation sequences and more operation types. For readers interested in reverse engineering, design reuse, or generative CAD modeling, the result offers a path where accuracy and validity improve together.
The provided text is abstract-level and does not list specific benchmark names, the evaluated baseline set, validity metric definitions, or per-shape-category results, so the comparison conditions behind the 87.2% relative IoU improvement still need confirmation in the full text. ARCADE-1.5M's data provenance, licensing, splits, and deduplication strategy are not described, which affects judgment of its coverage bias. Search computational cost, the set of local edit operators, and policy architecture details are likewise absent; these are open questions for assessing reproducibility and extensibility.
