Skip to main content
Back to timeline
arXivSource publication:

CADFather coordinates learned, geometric, and optimization tools through a vision-language assistant, reaching zero invalid outputs across six CAD reconstruction benchmarks

Synopsis

CADFather is an autonomous, training-free agent in which a vision-language assistant decides which candidate CAD program to extend, which tool to invoke, and how many proposals to sample, unifying learned proposal generation, geometry-based algorithmic proposals, and parameter optimization in one CadQuery representation, and achieving zero invalidity with higher mean IoU and GMS and lower median Chamfer distance on the full DeepCAD, Fusion360, and MCB test sets.

Source-provided article image: CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
Figure 1 ·

Figure 1: CADFather overview. Given a target mesh and feedback from existing CAD candidates, the visual assistant selects candidate programs to continue, tools to apply, and the number of proposals to generate. Learned and algorithmic proposals create alternative continuations, while parameter optimization refines existing candidates. The resulting programs are executed and evaluated, and the assistant visually accepts or rejects them before pool admission. Their renders and scores inform subsequent decisions. * In the figures, “stepwise” denotes learned proposals, and “deterministic” or “det” denotes algorithmic proposals.

arXiv

Interpretation

The system places three complementary tools in one candidate pool and one representation: learned proposals come from a pretrained CADENA-RL operation generator, algorithmic proposals come from cross-section extrusion fitting against the target mesh, and parameter optimization changes only sketch coordinates, radii, and extrusion depths without altering the operation sequence. Prior methods typically commit to a single source of operations or repeatedly edit one program; CADFather lets the visual assistant decide both which candidate to continue and which tool to apply at each step, while retaining earlier candidates for revisiting. The method is described in full, with explicit tool interfaces, candidate-table fields, budgets, and termination rules; both the generator and the assistant are frozen and require no additional training.

On the full DeepCAD, Fusion360, and MCB test sets, mean IoU reaches 0.987, 0.976, and 0.913, above the 0.966, 0.952, and 0.895 of CADENA-RL sampling; the invalidity ratio is 0.0000, so every part ends as a closed solid. The learned tool uses the same CADENA-RL checkpoint, so the gain does not come from a stronger generator but from tool coordination and candidate search. The three test sets contain 8046, 1725, and 5000 parts; the IoU gain is of similar size across all three (0.018–0.024), and mean GMS and median Chamfer distance at both sampling densities are consistently better.

Ablations show algorithmic proposals carry the largest contribution: removing them raises MCB invalidity from zero to 0.124 and lowers the mean score about as much as removing both tool classes, whereas removing the optimizer lowers the mean score by only 0.004–0.006 and leaves invalidity at zero. This separates the role of supplying a closed solid to build on from that of refining existing parameters, indicating that on mechanical parts starting from a valid base matters more than parameter precision. Ablations run on fixed 1000-part subsets of each test set under identical budgets, limits, and model servers, and report paired per-part score differences with 95% bootstrap confidence intervals.

On CADENA-Bench the invalidity ratio is zero with mean IoU 0.909 and GMS 0.721, above the published CADENA-RL greedy row (0.876 and 0.670); on BenchCAD it reaches voxel IoU 0.968 with no invalid predictions, against 0.910 for the other mesh-conditioned entry, CADENA-RL. The same configuration extends to a mechanical-part-family benchmark and an industrial-part-family benchmark, with contextual comparisons to image-conditioned systems and tool-using frontier models. CADENA-Bench contains 3396 parts in six families; BenchCAD contains 17900 CadQuery programs from 106 industrial families in three difficulty tiers, with IoU and GMS decreasing monotonically from easy to hard.

Perspective

The result targets settings where a single part mesh is the input and an executable, editable CAD program is the goal, such as components available only as scans or without construction history; the system ends with zero invalidity on the full DeepCAD, Fusion360, and MCB test sets, also zero on CADENA-Bench, and only 12 of 18000 CADBench parts end without a valid prediction. Readers who want to reuse the pipeline can use the public code repository and adopt its candidate pool, protected best result, and budget mechanism; algorithmic proposals matter most on parts where no sampled first operation yields a closed solid, so they are worth prioritizing in mechanical-part workflows. The evaluation covers geometric reconstruction and program validity, not recovery of original design history, manufacturing intent, or engineering constraints, and does not cover assemblies, tolerances, or simulation-validated designs.

The assistant decides from incomplete visual and numerical evidence, so its choices can depend on the prompt, model behavior, and which candidates are available for inspection, and the best reconstruction is not guaranteed. Parameter optimization cannot correct an unsuitable operation sequence, and retaining earlier candidates does not help if a useful alternative was never generated. Step and time limits may terminate a trajectory that would have succeeded. The ablations use fixed budgets but are limited to 1000-part subsets, do not match consumed compute, and do not isolate the contribution of the assistant's selection policy; direct comparison with other CAD agents remains future work. Mean GMS on MCB is 0.766, well below the IoU, suggesting that matching volume does not imply matching surface structure, a gap worth watching on mechanical parts. On CADBench the valid shape rate is below CADFit because 576 parts exceed the evaluator's 30 s limit, and the practical weight of that cost boundary is left to the reader.

Sources