OMatGRPO reinforces a crystal generator's discrete composition channel directly, raising mSUN from 13.4% to 45.5% and showing single-element repacking can inflate the benchmark
Synopsis
Using group-relative policy optimization (GRPO), this work reinforces the discrete atom-type channel of the crystal generator OMatG, whose closed-form likelihood ratio makes element choices policy actions; with a reward combining stability, novelty and diversity plus sparse-hull and single-element guards, the yield of metastable, unique and novel structures (mSUN) on LeMat-GenBench rises from 13.4% for the pretrained model to 45.5%, while the paper records how the policy exploited earlier rewards and shows that single-element repacking is counted as discovery by existing mSUN implementations.
Figure 1: Reinforcement learning on a crystal generator can satisfy the benchmark without real discoveries, but guards in the reward close the exploit. (a) The OMatG generator turns noise with masked element identities into crystals and is scored by mSUN under LeMat-GenBench (Section 4.1 ). Without the single-element guard, RL inflates the pretrained model’s 13.4% to 37.6%, but over half of that mSUN set is single-element repackings of known elements, such as face-centered-cubic Mg counted as a novel material. With the full reward (right, Section 3.3 ), which routes sparse-hull structures to abstention and single-element structures to the penalty reward, mSUN reaches 45.5% with 0.4% single-element structures and DFT-confirmed compounds on well-referenced hulls. (b) The composition channel. Under masked discrete flow matching, each atom’s element identity is a discrete commit event with an exact log-probability, while positions and lattice evolve as a continuous SDE (Section 3.2 ).
arXivInterpretation
It is, to the authors' knowledge, the first application of GRPO to the discrete flow-matching composition channel of a crystal generator, where each atom type is committed in a single categorical draw and the likelihood ratio has a closed form, so the policy gradient reaches the composition channel exactly. Prior reinforcement learning for diffusion and flow generators either froze atom types, relaxed them to continuous variables, or decoded them from a latent space, leaving composition only indirectly controlled; here the elements are native discrete actions of the policy. Appendix A derives the composition-channel ratio and KL terms, showing the model enters only through the unmask rate and that schedule factors are identical under current and old policy and cancel in the ratio.
The proposed reward raises mSUN on the community benchmark LeMat-GenBench from 13.4% (336 of 2,500) for the pretrained model to 45.5% (1,138 of 2,500), and also improves the published Chemeleon2 model. Best-of-N selection reaches only 7.7% and reinforcement learning with the composition channel frozen only 9.8%, indicating most of the gain comes from reinforcement learning on the discrete composition channel; the same reward raises Chemeleon2 from 36.6% to 43.7% inside its own pipeline. The benchmark scores stability with a three-MLIP ensemble (Orb, MACE, UMA) requiring agreement of at least two, and uniqueness and novelty against roughly 5.3 million LeMat-Bulk structures; DFT confirms 63 of 67 trusted-tier metastability claims with a mean absolute deviation of 0.016 eV/atom between UMA and DFT hull distances.
Each term and guard of the reward was added after the policy had exploited a previous version, and the final single-element guard brings single-element structures to 0.4% of the mSUN set. In the discovery run without the single-element guard, 53.5% of the mSUN set is single-element, and these repacked cells are counted as metastable, unique and novel by LeMat-GenBench, MatterGen's evaluation code and Chemeleon2's definition. Appendix C catalogs five reward designs and five distinct exploits with per-run numbers, including drift into Li-O compositions under an unbounded stability term, drift into sparse above-hull O/F binaries, and single-element repacking; the guard holds across three seeds.
A stability claim is only as good as its reference hull, so every result is reported split by the number of reference phases behind each structure. In the pretrained prior 21.7% of the mSUN set sits on hulls with fewer than 12 reference phases, the discovery run has 63.4% there, and OMatGRPO holds the share at 23.5%; on the sparse tier DFT tracks UMA energies but cannot adjudicate stability because the DFT hull is missing the same competing phases. A second machine-learned potential reproduces the stability energies to a median of 0.015 eV/atom after removing a composition-linear offset, and 19 of 21 sparse-tier spot-checks agree on energy.
Perspective
The result is aimed at crystal-generation research targeting thermodynamic stability, novelty and diversity, and applies to inverse-design pipelines scored with machine-learned potentials and reference hulls. The reward swap shows the same reward behaves differently across decoder pipelines, so conclusions should be read per host pipeline. The authors recommend reading sparse-tier mSUN as candidates rather than discoveries, and note that a promising sparse candidate can be settled by computing its missing competing phases with DFT. Code and run configurations, trained models and generated structure sets are released, supporting reproduction and extension under the same benchmark conventions.
The reward and the benchmark share the novelty and uniqueness components, and the authors note that about 80% of the gain comes from a reward that optimizes neither, so the gain should be read with that circularity in mind. The four reward-swap cells are single-seed, and Chemeleon2's sampler is unseeded and not bit-reproducible, with the authors proposing a seed replicate to bound sensitivity. The 0.015 eV/atom median difference for the second potential was computed on a development run and not repeated for OMatGRPO. Two Gd-bearing DFT calculations were excluded for electronic non-convergence or ionic divergence and enter no statistic. Sparse-tier DFT tracks UMA energies but cannot adjudicate stability because the DFT hull is missing the same competing phases.
