Google search data plus machine learning forecast weekly influenza counts across Australian states, with best-model correlation from -0.353 to 0.977
Synopsis
Using weekly Google Trends search volumes for 2018 and 2019 compared against weekly influenza notifications from Australia's National Notifiable Disease Surveillance System (NNDSS), the study fitted four supervised regression models (elastic net, support vector regression, random forest, and feedforward neural network) independently for each state and territory except the Australian Capital Territory, for nowcast and one- and two-week-ahead predictions, finding that search volumes correlate with reported influenza rates over time, that random forest and elastic net generally performed better than the other models, that every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notificatio
Interpretation
The study examined the feasibility of forecasting influenza rates in Australian states and territories from Google Trends data and found that search volumes correlate with reported influenza rates over time. Earlier search-data influenza forecasting work was largely national or single-city; this work moves the unit of analysis down to Australian states and territories and models each independently. Based on comparison of weekly Google Trends search volumes with weekly NNDSS influenza notifications across 2018 and 2019, with every state and territory except the Australian Capital Territory included.
Among the four supervised regression models, random forest and elastic net generally showed better predictive performance than support vector regression and the feedforward neural network. The study compares four model families side by side on the same data and prediction task rather than reporting a single model. The four models were elastic net, support vector regression, random forest, and feedforward neural network, fitted independently per state and territory and evaluated for nowcast and one- and two-week-ahead predictions.
Predictive performance varied markedly by location: every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notifications, and best-model accuracy ranged from -0.353 to 0.977. The result ties the strength of search-data predictive utility to specific jurisdictions and notes lower accuracy in jurisdictions with smaller populations. Accuracy was measured by the Pearson correlation coefficient, reported per jurisdiction, and was lower for one- and two-week-ahead forecasts than for nowcasts.
Perspective
The work targets public health surveillance settings in Australian states and territories, applicable to weekly nowcasting and one- and two-week-ahead forecasting of influenza rates; the data window is 2018 to 2019 and the Australian Capital Territory was not modelled. The authors propose future work on location-specific keywords, disease prediction at a more geographically specific level, and considerations for smaller populations and robustness over time.
A careful reader would still watch which specific search queries were selected in each jurisdiction, how models were tuned and validated, and whether the same predictive performance holds in seasons beyond 2018 and 2019; the authors themselves flag accuracy in smaller-population jurisdictions and robustness over time as open questions. Because the reading scope is incomplete, figures and result tables are unavailable, so per-jurisdiction numbers cannot be checked further.
