Retraining AIFS on satellite precipitation observations yields Laxmi, which lifts global probabilistic accuracy by 19% and gives the most accurate 150 mm event-total forecast in 7 of 10 Indian tropical storms
Synopsis
The authors retrained AIFS, ECMWF's open-source operational 0.25-degree probabilistic graph-transformer weather model, on satellite-based precipitation observations to produce Laxmi, which improves global probabilistic accuracy by 19%, cuts drizzle overprediction by 33% for amounts below 3 mm per day, raises the global 95th percentile Brier skill score by 57%, and delivers the most accurate 150 mm event-total precipitation forecast in 7 of 10 Indian tropical storms (versus 1 for AIFS and 2 for the leading physical model IFS).
Figure 1: Global precipitation intensity distributions compared to IMERG observations. 99 th percentile precipitation for ( a ) ERA5, ( b ) IMERG, and ( c ) the difference computed between 2001 and 2023. Fraction of 6-hour time windows which experience between 0.1 and 1 mm 6 hr − 1 1\;\mathrm{mm\;6hr^{-1}} for ( d ) ERA5, ( e ) IMERG, and ( f ) the difference, computed between 2001 and 2023. Mean precipitation rates over the period are shown in Supplementary Fig. 1, and regional ERA5-to-IMERG frequency ratios are shown in Supplementary Fig.2̃. . ( g ) Precipitation intensity ratios, or the model frequency divided by observed IMERG frequency for a given precipitation intensity, are evaluated at 0.25 ∘ spatial resolution for IFS, Laxmi, AIFS-MSE, and AIFS-CRPS. To compute these ratios, we aggregate 24-hour accumulated precipitation for 94 forecasts for each model for all lead times up to 15 days. IMERG frequencies are computed over the global domain between 60 ∘ S and 60 ∘ N between January 2023 to September 2025. A ratio of 1, shown with a horizontal black line, indicates that a model perfectly emulates the observed precipitation frequency distribution of IMERG. Solid lines distinguish models trained with IMERG precipitation observations, whereas dashed lines denote the physics-based IFS and AIWP models trained on ERA5 precipitation. We exclude observations that are less than 0.1 mm day − 1 0.1\mathrm{\,mm\;day{{}^{-1}}} and truncate at the high end above 908 mm day − 1 908\mathrm{\,mm\;day^{-1}} where the baseline IMERG probability density drops below 10 − 10 10^{-10} . The probability densities before a ratio is taken are shown in Supplementary Fig. 3.
arXivInterpretation
Training an AI weather model directly on satellite precipitation observations, rather than on reanalysis data, can substantially improve global probabilistic precipitation forecast accuracy. Most AI weather models are trained and evaluated against the ERA5 reanalysis product, which has well-known biases; this work retrains AIFS directly on satellite precipitation observations to produce Laxmi, improving global probabilistic accuracy by 19%. Presented as an overall model-level comparison through the 19% global probabilistic accuracy gain, supported by the distributional and case-study results below.
Laxmi mitigates systematic distributional biases inherited from ERA5, reducing both drizzle overprediction and extreme rainfall underprediction. Targeting ERA5's known biases, Laxmi reduces drizzle overprediction by 33% for amounts less than 3 mm per day and improves the global 95th percentile Brier skill score by 57%, showing the gain is not confined to mean error. Quantified by two specific bias metrics (33% and 57%) covering the light-precipitation and extreme-precipitation ends respectively.
In Indian tropical storm cases, Laxmi forecasts heavy event-total precipitation better than AIFS and the leading physical model IFS. Across a case study of 10 Indian tropical storms, Laxmi delivered the most accurate forecast of 150 mm event-total precipitation in 7 events, compared with 1 for AIFS and 2 for IFS. A case study of 10 storms, compared by counts on the actionable metric of event-total precipitation, constituting small-sample case-study evidence.
Perspective
The result addresses global probabilistic precipitation forecasting, particularly the two problems of light-precipitation bias and extreme-precipitation underprediction, and applies to forecasting pipelines built on AIFS that wish to replace reanalysis precipitation with observational precipitation in training; for sectors such as agriculture and flood management, its value lies in translating improvements into actionable metrics such as event-total precipitation. The case-study portion focuses on Indian tropical storms, indicating the method was also examined in strongly convective precipitation settings.
Only abstract-level information is available here: the specific source and period of the training data, the observational benchmark used for evaluation, and the statistical uncertainty around the 19%, 33%, and 57% figures are not elaborated in the text. The case-study result rests on 10 Indian tropical storms, so whether the conclusions generalize to other regions and other types of precipitation systems remains an open question to watch.
