NeutronGym grades neutron instrument design with McStas physics and trains an LLM agent, lifting Qwen3-8B from 11% to 77%
Related research and updatesSynopsis
The work introduces NeutronGym, an executable environment for neutron instrument design in which agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge; on the 16 published-instrument tasks of McStasBench seven models reproduce at most 7, none retrieves a reference and none meets an improvement target, while reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B.
Figure 2: Pass rates on held-out instances ( n = 300 n=300 ; 150 150 for the no-model probe) with Wilson 95% intervals, covering instance sampling for a fixed policy, not training-seed variation (8 points) or grading-seed fragility (§ 5.3 ). (a) The trained 8B passes the untrained 32B on all four gated families and no fixed-answer or lookup baseline comes close; ∗∗∗ : p < 10 − 3 p<10^{-3} , paired exact McNemar against the untrained 32B (discordant 208 vs. 12, 89 vs. 43, 116 vs. 12, 125 vs. 29). Guide match is replicated at a second seed; the other three are single runs. (b) Guide match under a tightening tolerance: the trained policy’s designs sit inside the tolerance, the untrained one’s on its edge.
arXivInterpretation
It introduces NeutronGym, described as the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. It turns instrument design from text question answering into an executable, automatically graded physics task, where the grade rests on ray-tracing output rather than model self-assessment or human judgment. The environment consists of validating tools, McStas ray-tracing and a four-level grading ladder, and the authors emphasize grading with 'no LLM judge'; procedural families supply unlimited instances of a fixed layout with held-out parameter regimes.
It builds McStasBench, a curated slice of 16 tasks from published instruments behind memorization probes and a sandbox; seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. Memorization probes and a sandbox separate reproduction from memorization or retrieval, yielding a quantified baseline for current models on this task. The 16 tasks come from published instruments, evaluation covers seven models, and the reported outcomes are the reproduction ceiling, retrieval, and the improvement target.
The environment also trains: reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. It shows the environment's reward signal can itself drive capability gains, not merely serve as an evaluation harness. Concrete figures are reported for 11% to 77% and 69% at a second seed, plus replication on three further families at one seed each.
The analysis characterizes the gain: without the ladder's partial credit it collapses by 60 points; from reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. It decomposes 'training works' into partial credit and closed-form physics information, and sets the result against a classical optimizer and frontier models. It reports the 60-point collapse, the 77% versus 81% comparison with the statement that the gap does not separate at this size, and the 98-99% frontier-model contrast; the authors also note failing four task designs that no-model baselines could solve and releasing the probes that found them.
Perspective
The environment targets the specific setting of neutron instrument design, where agents set design parameters within a fixed layout of procedural families and are evaluated on held-out parameter regimes; the 16 McStasBench tasks come from published instruments and sit behind memorization probes and a sandbox. The training results concern Qwen3-8B and several gated families, with a second seed and one seed each on three further families. The authors note that reaching a classical optimizer from reward alone requires being handed the closed-form physics, and that this gap does not separate at this size, while frontier models still solve 98-99%. These choices define the scope of the result: simulatable, automatically gradable, parameterized instrument design tasks.
A careful reader may still watch: whether instrument designs beyond the procedural families and McStasBench follow the same grading ladder; how stable the training results are across more seeds and families; and how the condition that reaching a classical optimizer from reward alone requires closed-form physics behaves in more complex layouts. The authors mention failing four task designs that no-model baselines could solve and releasing the probes that found them, which are informative for judging the difficulty distribution of tasks. The loaded text is abstract-level information without figures or full experimental detail, so the exact experimental settings behind these numbers remain to be checked in the original.
