Skip to main content
Back to timeline
International Journal of Data Science and AnalyticsSource publication:

Calendar-only LSTM and TCN detected 7 of 8 vineyard mildew risk events in 2023, while environmental-only models were substantially weaker

Synopsis

The study reformulates vineyard mildew risk prediction as an event-onset warning task—after a minimum disease-free gap, will a new treatment-associated risk event begin within the following 3–7 days?—and, under a chronological 2020–2021/2022/2023 train-validation-test split, compares calendar-only, environmental-only, and combined representations with logistic regression, XGBoost, LSTM, and TCN, finding that calendar-only models were highly competitive (calendar-only LSTM and TCN detected 7 of 8 test events in every run, while a monthly climatological baseline detected 6), that environmental-only models were substantially weaker, that gains from adding environmental variables varied across model families, and that event-onset prediction and conventional daily-status classification show dif

Source-provided article image: Event-based early warning of vineyard disease risk from environmental time series
Fig. 1

Fig. 1

Interpretation

The paper recasts vineyard mildew risk prediction from daily-status classification into new-event onset warning: a new event is the first positive day following at least g consecutive observed disease-free days, and models estimate whether such an event falls within a future 3–7-day interval, with reference settings L=30, g=5, and horizon [3,7]. The same dataset (Arvanitis et al.) had been framed as daily classification of grapevine disease labels from IoT environmental observations; this study argues that positive annotations persist over consecutive days, so favorable point-wise scores may merely reproduce an already active state, and therefore targets transitions instead. The task definition, event indicator, and sample eligibility rules are given in the methodology; the split is strictly chronological by year (2020–2021 training, 2022 validation and threshold selection, 2023 test), windows and target intervals never cross a year boundary, and the current-day annotation is used only for eligibility, never as an input feature.

Calendar-only representations were highly competitive: calendar-only LSTM and TCN detected 7 of 8 test events in every run with F1 of 0.638±0.021 and 0.608±0.022, and calendar-only TCN obtained the highest AUROC (0.913±0.007) and PR-AUC (0.537±0.025) among all evaluated configurations, while a monthly climatological baseline using no environmental observations and no fitted model detected 6/8 events with F1 0.507, AUROC 0.830, and PR-AUC 0.401. The study places explicit seasonal baselines (monthly climatology and cyclic calendar encodings) alongside learned models, quantifying how much predictive performance comes from recurring seasonal structure in treatment-derived annotations rather than from short-term meteorological information. Results are from the 2023 test year with thresholds selected only on the 2022 validation year and applied unchanged; stochastic models are reported as mean ± standard deviation over five random seeds; the authors state that the 7/8 versus 6/8 difference is interpreted descriptively and give approximate 95% Wilson intervals of [0.529,0.978] and [0.409,0.929] that overlap substantially.

Environmental-only representations were substantially weaker and less stable: AUROC for environmental-only XGBoost, LSTM, and TCN was approximately 0.64–0.65 with PR-AUC below 0.21, and environmental-only LSTM detected (4.0±3.0)/8 events while generating 84.6±54.7 alerts; adding environmental variables to calendar information changed results in a model-dependent way, raising XGBoost F1 from 0.495±0.016 to 0.555±0.028 while reducing AUROC, PR-AUC, and mean event coverage, and leaving calendar-only input stronger for LSTM and TCN on the principal discrimination and event-coverage measures. Rather than treating added meteorological variables as a default improvement, the study reports the direction of change per model family and notes that no family was uniformly superior: calendar-only LSTM had the highest F1, calendar-only TCN the highest AUROC and PR-AUC, and combined XGBoost the fewest strict false-alert episodes. Three input representations (C with four cyclic calendar features, E with 22 environmental features, E+C with 26 features) are compared under the same chronological protocol and the same threshold-selection procedure, with stochastic models reported across five seeds.

Event-onset prediction and conventional daily-status classification are different tasks with different trade-offs: under the common event-oriented protocol, daily-status logistic regression detected 7/8 events (versus 6/8 for event-onset logistic regression) but generated 52 warning alerts (versus 46), and event-onset XGBoost detected one more event than daily-status XGBoost while generating roughly four additional alert days, with the same mean number of strict false-alert episodes (3.2±0.4). The study performs a matched comparison of the two target formulations using the same model families, the same chronological split, and the same event-oriented protocol, and states explicitly that the two formulations define different positive labels and eligible samples, so their sample-level metrics are interpreted within each task. The comparison uses the combined E+C representation with L=30; for daily-status models only positive outputs issued while the current annotation was disease-free were retained post hoc as candidate warnings, and the paper notes this does not convert the daily-status classifier into a prospectively trained forecasting model.

Perspective

The framework targets a single-site, year-forward deployment setting: train on historical seasons and warn for an unseen future season, with reference settings L=30, g=5, and a 3–7-day horizon. It directly serves crop-protection decision support that uses treatment-record risk proxies, and its event-oriented protocol (event coverage, warning lead time, alert burden, episode formation, and near-miss versus strict false alerts) can be reused by other proxy-labelled agricultural time-series warning studies. Sensitivity analyses indicate that the 3–7-day horizon gave the most balanced trade-off among sample-level discrimination, event coverage, and alert burden, while no historical window length was uniformly optimal, so these parameters are best treated as adjustable operational choices.

The 2023 test year contains only eight reference events, so the outcome of one event changes event recall by 0.125, and the paper reports approximate 95% Wilson intervals for 7/8 and 6/8 that overlap substantially, meaning one-event differences do not establish statistical superiority. Labels come from fungicide-treatment records rather than pathological confirmation, and the downy- and powdery-mildew annotations agree on 98.9% of observed days and are identical in the validation and test years, so the unified any-mildew target is an operational management-risk proxy. The SHAP analysis is based on one representative XGBoost run, does not quantify variability across seeds, and characterizes model usage conditional on included and potentially correlated predictors rather than biological causality or independent predictive contribution; the high attribution of the 14–29-day lags may reflect longer-range seasonal context, redundancy among lagged descriptors, or recurring management-related timing. Changing the historical window or prediction horizon alters the number and composition of positive samples, and changing the disease-free gap changes the event ground truth itself, so the reference configuration is best viewed as an operational formulation that performed reasonably in this case study.

Sources