Skip to main content
Back to timeline
arXivSource publication:

MOOSEnger raises executable success for MOOSE inputs from 5% to 89.5% with a simulation-aware agent framework

Synopsis

MOOSEnger is a modeling-and-simulation AI agent framework for the Multiphysics Object-Oriented Simulation Environment (MOOSE) ecosystem, whose simulation-aware harness combines an interchangeable reasoning model with grounded domain knowledge retrieval, HIT-aware parsing, syntax metadata, language-server diagnostics, revision-controlled authoring, and local or MCP-backed validation and execution in a generate-check-repair-run workflow; across 200 prompts spanning eight simulation families it raises executable success from 10/200 (5%) to 179/200 (89.5%) with GPT 5.2 API and from 0/200 to 153/200 (76.

Source-provided article image: MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem
Figure 1 ·

Figure 1 : Core-plus-domain architecture. The reasoning model is replaceable, while the MOOSE plugin and core runtime retain the domain evidence, artifact lifecycle, validation policy, and execution gates. Model substitution therefore changes how a candidate is produced, not the artifact-level conditions used to accept it.

arXiv

Interpretation

The work introduces MOOSEnger, a modeling-and-simulation AI agent framework for the MOOSE ecosystem, centered on a simulation-aware harness that combines an interchangeable reasoning model with grounded domain knowledge, revised simulation artifacts, MOOSE-specific validation, and executable solver feedback. Relative to one-shot large language model generation, the framework places generation inside a complete system that includes MOOSE knowledge retrieval, HIT-aware parsing, syntax metadata, language-server diagnostics, revision-controlled authoring, and local or MCP-backed validation and execution. Evidence comes from the described framework composition and from evaluation results on 200 prompts spanning eight simulation families and a ten-case Method of Manufactured Solutions benchmark.

Across 200 prompts spanning eight simulation families, the MOOSEnger harness increases executable success from 10/200 (5%) to 179/200 (89.5%) with GPT 5.2 API and from 0/200 to 153/200 (76.5%) with Gemma 4 31B. The result attributes executable reliability to the complete agent system rather than to the reasoning model alone, since the same harness produces large gains with two different reasoning models. Evidence is the executable-success counts for 200 prompts on each of two models, spanning eight simulation families.

On a ten-case Method of Manufactured Solutions benchmark, all ten generated inputs satisfy the semantic-alignment criterion, and eight execute successfully while meeting the prescribed single-mesh numerical-accuracy criterion. This benchmark moves evaluation beyond executability toward semantic alignment and numerical accuracy, pointing toward physics-informed verification. Evidence is the pass counts on the semantic-alignment criterion and the single-mesh numerical-accuracy criterion across ten cases.

The framework's generate-check-repair-run workflow binds evidence to each input revision and guides bounded repair before acceptance. Relative to one-shot generation, this workflow targets the problem that small syntax, schema, reference, or solver-configuration errors can prevent a plausible input from executing, while noting that successful execution alone does not establish scientific correctness. Evidence comes from the described workflow and validation mechanisms together with the reported executable-success rates and Method of Manufactured Solutions results.

Perspective

The work targets modeling-and-simulation input generation for the MOOSE ecosystem and applies to settings where an interchangeable reasoning model is connected to a simulation-aware harness that includes MOOSE knowledge retrieval, HIT-aware parsing, syntax metadata, language-server diagnostics, revision-controlled authoring, and local or MCP-backed validation and execution. Evaluation covers 200 prompts spanning eight simulation families and a ten-case Method of Manufactured Solutions benchmark, so its conclusions apply to executable reliability and to the semantic-alignment and single-mesh numerical-accuracy criteria within that scope. The framework offers a path toward physics-informed verification and future full application-level and engineering verification and validation implementation.

This reading is at summary scope and does not include the full paper's figures and details, so the concrete implementation of each harness component, the boundaries of the repair strategy, and the specific reasons the two cases in the Method of Manufactured Solutions benchmark did not meet the numerical-accuracy criterion remain open questions to check in the original. In addition, the distinction that successful execution alone does not establish scientific correctness suggests caution when extrapolating executable-success rates to scientific correctness.

Sources