Microsoft Research intern builds a machine learning pipeline that gives 30-60 minute space-weather risk warnings for 66,935 U.S. substations, detecting nearly 80% of major events
Synopsis
Developed during a Microsoft Research summer internship, this end-to-end machine learning pipeline uses solar-wind observations from the L1 Lagrange point, forecasts of the AE and Dst indices, physics-informed constraints, local geological conductivity, and grid-infrastructure data to produce location-specific geomagnetically induced current risk estimates 30-60 minutes ahead for 66,935 substations in the continental United States, detecting nearly 80% of major space-weather events over the 2020-2026 evaluation period, with an AE forecast RMSE of 410.2 nT and a Dst forecast RMSE of 7.2 nT, outperforming the Burton equation on 62.2% of high-activity hours.
Interpretation
Built a three-stage end-to-end pipeline: first using L1 solar-wind measurements to forecast the AE and Dst indices while assembling geological conductivity and location features for each substation, then applying a gradient-boosting model to estimate dB/dt, the rate of magnetic-field change associated with GIC risk, and finally converting predictions into location-specific risk estimates aggregated into a continental risk assessment. Prior space-weather warnings largely stayed at global or regional storm-intensity alerts; this work carries the forecast all the way to per-substation, location-specific risk and explicitly incorporates local latitude and geology. The text describes the three-stage structure (Figure 1) and the public data sources used (NASA OMNI, NASA-aggregated Kyoto World Data Center data, INTERMAGNET and U.S. Geological Survey magnetometer observations, and GridSFM-derived grid data), and notes that 50 AI agents helped explore features, validation strategies, and model configurations.
Over the 2020-2026 evaluation period the AE predictor produced forecasts spanning nearly the full observed range of AE activity with an RMSE of 410.2 nT, lower than empirical and solar-wind-only baselines; the Dst predictor reached 7.2 nT RMSE, outperformed the Burton equation on 62.2% of individual hours during the most active periods, and produced a substantially wider prediction range, improving severe-event detection in the end-to-end system by 1.2 percentage points when combined with AE forecasts. Relative to Burton-style empirical formulas, the model wins on a majority of peak-activity hours and yields a wider prediction distribution, indicating greater capacity to represent extreme events. Results are presented as RMSE, hourly win rate, and percentage-point improvement, with a two-row results table for AE and Dst; the evaluation period is 2020-2026.
Because no widely deployed operational system offers a direct industry benchmark for the GIC risk calculation, that stage was compared with simple linear regression and achieved detection rates of 76.5% for major events (≥10 nT/min), 81.2% for severe events (≥20 nT/min), and 64.1% for extreme events (≥50 nT/min), with false-alarm rates rising with storm severity and the highest detection rates at northern, higher-latitude stations. It translates forecast metrics into operational event-level detection rates and a false-alarm trade-off, and reports performance by latitude, bringing the risk output closer to what grid operations care about. Detection rates are reported across three thresholds (Figure 2), and the text explicitly states that false-alarm rates increase with severity and that performance varies by latitude, while also noting the absence of an operational benchmark.
In measured inference the pipeline produced estimates for all 66,935 substations in approximately 333 milliseconds, allowing many scenarios to be evaluated quickly; Figure 3 illustrates the output for a representative major-storm scenario, and the text explicitly states the map is a demonstration of the model's continental-scale output rather than a record of a live operational event. It combines continental-scale, per-substation risk mapping with sub-second inference speed, supporting a workflow that moves from broad warnings toward targeted analysis. Speed is given as measured inference time covering all 66,935 substations; the demonstrative nature of Figure 3 is explicitly qualified in the text.
Perspective
The result is aimed at grid operations and planning: it provides 30-60 minute-ahead, location-specific risk estimates for 66,935 substations in the continental United States, serving grid operators and planners who need to turn a broad space-weather warning into a concrete judgment about which locations warrant closer analysis, and to prioritize engineering review or consider adjusting reactive-power reserves or temporarily reconfiguring parts of the network. Methodologically it applies to regions with public solar-wind and magnetometer data plus substation location and geological conductivity information; the authors also propose extending it to other regions, while accounting for different geological conditions and grid topologies.
Several points remain worth watching: the GIC risk stage has no directly comparable deployed operational system and was compared only with simple linear regression, so its frame of reference for relative advantage is limited; false-alarm rates rise with storm severity, reflecting a trade-off between missed events and cautious alerts; and performance varies by latitude, with the highest detection rates at northern stations. Figure 3 is explicitly labeled a demonstration of model output rather than a record of a live operational event, so it cannot be used to judge real operational performance. The authors also note that further validation with utilities and operational data would be needed before the system could be used in grid operations; future directions include temporal transformers to extend forecast horizons, international scaling, integration with grid operations, and moving from substation-level to transformer-level risk. In addition, although the available text is the full text, it does not include the graphical content of Figures 1, 2, and 3, so judgments that depend on figure details remain open questions.
