Public articles linked to the same research event.
arXiv The work introduces NeutronGym, an executable environment for neutron instrument design in which agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge; on the 16 published-instrument tasks of McStasBench seven models reproduce at most 7, none retrieves a reference and none meets an improvement target, while reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B.
The work introduces NeutronGym, an executable environment for neutron instrument design in which agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge; on the 16 published-instrument tasks of McStasBench seven models reproduce at most 7, none retrieves a reference and none meets an improvement target, while reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B.
The work introduces NeutronGym, an executable environment for neutron instrument design in which agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge; on the 16 published-instrument tasks of McStasBench seven models reproduce at most 7, none retrieves a reference and none meets an improvement target, while reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B.
The work introduces NeutronGym, an executable environment for neutron instrument design in which agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge; on the 16 published-instrument tasks of McStasBench seven models reproduce at most 7, none retrieves a reference and none meets an improvement target, while reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B.