Skip to main content
Back to timeline
arXivSource publication:

A flow-matching generative model reconstructs 3D cosmic baryon fields from galaxy surveys and yields per-sightline FRB dispersion-measure posteriors

Synopsis

The work introduces a flow-matching generative model that reconstructs the posterior distribution of 3D gas density fields from redshift-space galaxy density fields, conditioned on cosmological parameters, a minimum stellar mass threshold, and the baryon suppression power spectrum, and validates on CAMELS50 (IllustrisTNG) test boxes that the reconstructed dispersion measures agree with the truth both sightline by sightline and in overall distribution, with generalization to TNG300-1 at roughly 43 times the training volume.

Source-provided article image: Feedback-Conditional 3D Reconstruction of Cosmic Baryons with Flow Matching
Figure 1 ·

Figure 1: Overview of our method for reconstructing the 3D distribution of cosmic baryons. The flow matching model learns to reconstruct the gas density field conditioned on galaxy density field δ g ​ ( 𝐱 ) \delta_{g}(\mathbf{x}) , minimum stellar mass threshold log 10 ⁡ M ⋆ , min \log_{10}M_{\star,\mathrm{min}} of the galaxy sample, the cosmological parameters 𝛀 \mathbf{\Omega} , and power spectrum suppression s ​ P ​ ( k ) sP(k) . The output is a gas density field that can be integrated to predict FRB dispersion measures. The model is probabilistic, allowing sampling from a reconstruction posterior.

arXiv

Interpretation

The model learns a conditional posterior over gas density fields rather than a single best-fit field, so the dispersion measure, a linear functional of the gas field, inherits a posterior directly with no additional modeling. Prior galaxy-to-gas reconstructions (e.g., Kvasiuk et al., Krywonos et al.) return a single best-fit field; this work samples the posterior directly with an implicit-likelihood generative model and explicitly conditions on the scale-dependent baryon suppression power spectrum as a feedback measure. Validated on 10 CAMELS50 test boxes: sightline-by-sightline residual scatter of about 0.1 dex, best-fit slope near unity, and mean and width matching the true distribution; Spearman rank correlation about 0.9; 100 posterior samples drawn per test box.

The reconstruction responds physically to the feedback condition: changing the input suppression spectrum systematically changes gas morphology and the width of the dispersion-measure distribution. Feedback previously entered as CAMELS astrophysical parameters rather than as a scale-resolved quantity; conditioning on the suppression power spectrum makes the feedback effect directly tunable and interpretable across feedback implementations. Across all 10 test boxes the low-feedback reconstruction yields a broader dispersion-measure distribution than the high-feedback one, by about 0.1 dex (about 10%) on average, with a paired comparison p about 0.002 and a sign test p about 0.002, all ten differences sharing the same sign.

The fully convolutional architecture lets the model be applied to inputs far larger than the training volume with no change to the network or its trained weights, reaching the scales needed for cosmological FRB sightlines. Training boxes are 50 Mpc on a side, whereas TNG300-1 is 205 Mpc on a side (about 43 times the training volume); the model processes the whole volume in a single pass with no tiling or stitching. The reconstructed dispersion-measure histogram on TNG300-1 closely matches the true distribution; TNG300-1 shares the IllustrisTNG galaxy formation model with the training set, so this tests generalization in box size and resolution rather than in underlying physics; the reconstruction underpredicts dispersion measure by about 0.1 dex on average, calibratable through the known cosmic baryon density.

The galaxy selection threshold smoothly affects reconstruction fidelity, and the model widens its posterior in the right direction when information is scarce. The minimum stellar mass threshold is passed as a conditioning parameter and marginalized over at inference, drawn log-uniformly during training; this work further quantifies degradation when the threshold lies outside the training range. Raising the threshold from 10^9 to 10^10.5 solar masses drops the average galaxy count per box from 2558 to 314; the scale at which the correlation coefficient falls to 0.5 moves from about 1 Mpc to about 4 Mpc; rank correlation degrades smoothly from about 0.83 to about 0.63; posterior width grows by 21% while reconstruction error grows by 39%, with the posterior about 15% too narrow at the lowest threshold and about 30% too narrow at the highest.

Perspective

The result is aimed at researchers who need a 3D distribution of diffuse gas with quantified uncertainty on cosmological scales, especially those working with fast radio burst dispersion measures, kSZ/tSZ, weak lensing, and X-ray stacking analyses. The method is established on simulated data, with training and validation assuming uniform galaxy sampling and redshift-independent stellar mass thresholds, and evaluated at redshift z=0; its intended setting is therefore a simulation-calibrated probabilistic reconstruction framework for predicting dispersion-measure distributions given an observed foreground galaxy field and for comparing feedback scenarios. The fully convolutional property allows direct application to larger volumes, reaching the scales spanned by cosmological FRB sightlines; the absolute dispersion-measure scale can be anchored through the known cosmic baryon density, calibrating out the overall normalization offset.

The model inherits the assumptions and systematics of the CAMELS IllustrisTNG simulations, and inferred parameters remain limited to the simulation prior range; training and evaluation are at redshift z=0, whereas real FRB sightlines span a range of redshifts, requiring treatment of gas redshift evolution. The galaxy-position input carries no information about the gas distribution below about 1 Mpc, forcing reliance on the prior, so the current algorithm is positioned as a cosmic-web reconstruction rather than a halo-gas reconstruction; applications requiring halo gas would need a hybrid halo-model/SBI approach or higher-resolution training data. The minimum stellar mass threshold has not been calibrated against a real survey selection function, and training and validation data assume uniform sampling, so model mismatch may degrade reconstruction quality. In addition, several equations, figure captions, and table values are missing from the parsed text (for example, parts of the equations and the specific content of some figures), so quantitative restatement here is limited to the numbers explicitly given in the body text.

Sources