Skip to main content
Back to timeline
medRxivSource publication:

Simulating Community Health Behavior with LLM Synthetic Populations: A Cross-State Evaluation of LLMPopSim

Synopsis

The study introduces LLMPopSim, a generative population simulation framework that integrates U.S. Census and CDC data to construct synthetic individuals and uses an LLM to simulate individual health behaviors whose aggregate outcomes can be evaluated at the community level; developed with historical Hawaiʻi data and evaluated for temporal and geographic generalizability using held-out 2022 cohorts from Hawaiʻi and New York State with colorectal cancer screening and mammography as proof-of-concept behaviors, the four state-outcome evaluations showed mean absolute error of 3.5 to 15.0 percentage points and correlations between simulated and observed ZCTA-level prevalence of 0.26 to 0.69.

Source-provided article image: Evaluating agentic simulation for local public health estimation
Figure 1 ·

Figure 1. LLMPopSim framework for geographically grounded generative population simulation. Population-

medRxiv · Page 3

Interpretation

It proposes and implements LLMPopSim, a generative population simulation framework that integrates U.S. Census ACS data with CDC PLACES data to construct synthetic individuals, uses an LLM to simulate individual health behaviors, and allows aggregate outcomes to be evaluated at the community level. Prior LLM-based generative agents largely reproduced aspects of individual behavior; this work extends that idea to geographically grounded populations and provides an aggregate evaluation path comparable at the community level. The framework was developed with historical Hawaiʻi data and evaluated for temporal and geographic generalizability on held-out 2022 cohorts from Hawaiʻi and New York State, making it a framework study with held-out validation.

Across the four state-outcome evaluations, correlations between simulated and observed ZCTA-level prevalence ranged from 0.26 to 0.69 and mean absolute error ranged from 3.5 to 15.0 percentage points, indicating that individually represented LLM synthetic agents can aggregate into population-level patterns that retain measurable features of real-world health behavior. This moves the question of whether LLM agents can reproduce real-world health behavior under temporal and geographic transfer from conceptual discussion to empirical evaluation with quantitative error and correlation metrics. Evidence comes from held-out evaluation across four state-outcome combinations, reporting both mean absolute error and ZCTA-level correlation.

Behaviors differed across dimensions of population fidelity: colorectal cancer screening predictions preserved geographic ranking more strongly but systematically overestimated prevalence and compressed geographic variation, whereas mammography achieved lower absolute error but weaker geographic correlation and inconsistent preservation of between-community variability. This finding indicates that accuracy is not a single dimension, since absolute error, geographic ranking, and distributional variation can trade off in different directions, offering a multi-dimensional evaluation lens for simulation design. Based on the comparative description of the two behaviors within the same set of four state-outcome evaluations.

Prediction error was greatest in communities with lower observed screening prevalence and varied across community characteristics without a uniform socioeconomic gradient. This suggests error distribution relates to community characteristics rather than aligning along a single socioeconomic axis, pointing to directions for subsequent calibration and subgroup performance work. From descriptive analysis of the relationship between error and community characteristics across the four evaluations.

Perspective

The framework targets settings where synthetic populations are built from public aggregate data and aggregate health behaviors are evaluated at the community level; current validation is limited to two proof-of-concept behaviors, colorectal cancer screening and mammography, and to held-out 2022 cohorts from Hawaiʻi and New York State, with the stated intent of providing an empirical foundation for developing synthetic populations that may ultimately enable simulation of heterogeneous population responses to public health interventions.

Readers would still watch how calibration and distributional fidelity are specifically implemented, what drives subgroup performance differences, and whether the framework maintains similar error and correlation for behaviors and regions beyond the two behaviors and two states; the available text here is the abstract and declarations, missing figures and full methodological detail, so judgments on these points remain open.

Sources