Skip to main content
Back to timeline
arXivSource publication:

EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts

Synopsis

The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.

AI-generated editorial illustration: EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

Interpretation

EngiWorld shifts evaluation from comparing final interface states or reference solutions to inspecting the submitted engineering artifacts themselves, using a unified domain-verifier suite that checks geometric dimensions and topology, simulation quantities, schematic connectivity, building-model entities and relationships, and 3D scene structure, with multi-software tasks additionally checking intermediate artifacts and cross-stage consistency. Prior benchmarks such as OSWorld, Spider2-V, and ScienceBoard judge success by final states or reference solutions, while FEABench, BIM-Edit, and CADTestBench validate artifact quality but each is confined to a single software platform and a single task format; EngiWorld extends these single-domain precedents into a unified verifier suite spanning six domains. The paper reports that verifiers were tested with reference artifacts, alternative valid solutions, and controlled defective submissions; all 96 reference artifacts in the completed audit group pass, 40 Abaqus/ANSYS alternative-solver cases per interface and 5 FreeCAD lightweight-design cases are accepted, while 21 empty deliveries and 21 metrics-only submissions are rejected.

The benchmark covers the complete design loop: 1,301 tasks form six disjoint task types, including single-software (931), multi-software (60), software-selection (140), quantitative design (40), image-based modeling (120), and open-ended (10), with both GUI and CLI interfaces (691 CLI and 610 GUI tasks). Existing work is often limited to single-turn generation or short step budgets, for example Text2CAD is dominated by single-turn generation and CADWorld caps its step budget at one hundred steps; EngiWorld requires agents to interact with real professional software over multiple turns while maintaining parametric logic across long operation sequences. Tasks were built through an expert-calibrated, LLM-assisted pipeline, and every final task was executed and reviewed in its designated software environment; software coverage is reported as 1,723 task–software associations.

Evaluation of seven frontier models reveals a substantial capability gap: the strongest model reaches an EngiScore of only 44.3, 128 of the 300 tasks receive a zero score from all seven models, and only 6 of 168 multi-software attempts succeed (3.6%). The results separate general computer-use ability from professional engineering execution: roughly half of single-software tasks and about two thirds of image-based modeling tasks are completed, but performance drops sharply once intermediate artifacts must be transferred across applications and dependencies preserved. Seven models were evaluated zero-shot on a stratified subset of 300 tasks (152 CLI and 148 GUI) covering 73 software-and-task-category strata; in the failure analysis, DONE with a zero score and decision-turn exhaustion together account for 94.1% of main-evaluation failures.

Ablations show that observation design and interaction history affect models differently: placing the reference drawing directly in the task message improves GUI performance for all three models (GPT rises from 15.0 to 31.7), native resolution yields the highest score for all three, and the optimal history length is model dependent (Gemini benefits from extending five to fifteen turns, GPT largely plateaus after ten, and Kimi performs best at ten). These results indicate that more structured observations do not automatically translate into better engineering decisions: adding an accessibility tree reduces decision rounds for all three models but increases API cost by 50.7%–147.9% and lowers Gemini's score. Ablations were run on GPT-5.6 Sol, Gemini 3.7 Flash, and Kimi K3 with paired task-level gain and regression analysis, for example GPT's GUI gain combines 15 failures becoming successes and five successes becoming failures.

Perspective

The benchmark targets evaluation of agents operating end to end in real professional engineering software, covering six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization), 26 software platforms and workbenches, and both GUI and CLI interaction. It can be used to compare models and agent frameworks on single-software construction, tool selection, quantitative design, cross-toolchain delivery, and open-ended tasks, and to diagnose which stage of the design loop a failure occurs in. The verifier suite and artifact-checking methodology can be extended and reused by later work when adding new software or new engineering domains.

The main evaluation runs on a stratified subset of 300 tasks rather than all 1,301, so model rankings and absolute scores on the full benchmark remain to be seen. Results for quantitative design and open-ended tasks are partly reported in the appendix, with the main text giving only aggregates. Ablations show that observation design and history length affect models differently, with the same model moving in opposite directions on different tasks, indicating that these settings interact with task composition in ways still to be clarified. Verifier audits cover reference artifacts, alternative solutions, and controlled defective submissions, but tolerance-boundary tests are outside the completed audit populations. In addition, DONE with a zero score and decision-turn exhaustion dominate the failure analysis, suggesting that the gap between declaring completion and passing artifact verification is a direction later work will need to keep watching.

Sources