Skip to main content
Back to timeline
NVIDIA ResearchSource publication:

University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

Synopsis

Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.

AI-generated editorial illustration: University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

Interpretation

The NVIDIA Earth-2 generative weather framework was transferred to pollution fields, with the CorrDiff generative downscaling model working on the first attempt. The Earth-2 family had mainly addressed weather forecasting; this work extends it to air quality, using generative downscaling in place of the heavy compute of traditional chemistry-based models. Grounded in Topping's statement that 'The model worked on the first attempt' and the described training workflow, which is project-side narration without error metrics or controlled comparisons.

A UK-wide pollution model at 2-3 square kilometer resolution was trained from a year of hourly UK pollution simulation data, taking just two days on a single eight-GPU node of Isambard-AI. Traditional chemistry-based models are expensive enough to limit detail and update frequency; this shows a path to national-scale, high-resolution pollution fields with relatively low GPU hours. The text gives concrete data scale, resolution, and hardware (a single Isambard-AI node, 5,448 GH200 chips, 21 exaflops), plus a comment from the Bristol supercomputing director on relatively low GPU hours and power draw.

With Earth-2 StormCast added, the model enables time-dependent forecasts that directly use air quality observations, and the same workflow runs on the DGX Spark desktop system for inference and smaller training runs. Pollution modeling moves from supercomputers toward desktop-class devices, changing who can do this science and how quickly. Based on Hao Zhang's account of training StormCast, Topping retraining models on a DGX Spark in his office, and NVIDIA's Niall Robinson's comment, which is demonstration-level evidence.

The team plans to release open source training data and workflows for the pollution models so similar models can be trained for other countries and regions. The approach is framed as a replicable open workflow rather than a single UK case, pointing toward localized pollution modeling worldwide. Stated as a team plan ('The team plans to release open source training data and workflows'), with no release date or licensing details given.

Perspective

The result targets UK-wide pollution fields at 2-3 square kilometer resolution, trained on a year of hourly chemistry-climate simulation data; the team plans to raise resolution toward street scale by incorporating additional open data. It enables countries or cities with a small burst of supercomputer AI time to train their own pollution models with local data, and moves inference and smaller training runs down to desktop systems such as the DGX Spark. The health alerts for asthma patients, policy-change scenario exploration, and pairing with edge AI devices for real-time decisions such as wildfires are described as envisioned or being explored.

A careful reader would still watch how closely the model reproduces or improves on traditional chemistry-based pollution fields, how usable 2-3 square kilometer resolution is for street-scale applications, how StormCast performs once it directly uses observations, and the exact scope, licensing, and timing of the open source training data and workflows. The text gives no error metrics, controlled comparisons, or independent evaluation, which are key open questions for judging the method's usability in health alerts and policy scenario exploration.

Sources