Unified Benchmark Plus RL Feature Selection: Learned Models Consistently Beat Fixed-Coefficient USGS Model for Post-Wildfire Debris Flow Prediction, with Drivers Varying Sharply by Region
Synopsis
The work builds a unified benchmark for data-driven post-wildfire debris flow (PFDF) prediction, uses an LLM to convert operational weather forecast narratives into structured meteorological indicators and augments features with terrain and soil-moisture context, and proposes a reinforcement-learning feature-selection framework that identifies key drivers by perturbing features until positive and negative events become indistinguishable; learned models consistently outperform the fixed-coefficient USGS logistic-regression baseline, augmented features further improve most model families, and storm variables dominate globally while feature importance varies substantially across regions.
Figure 1. Benchmark Framework of Data-driven PFDF Prediction, including data curation, evaluation metrics, baseline, and feature selection. Heterogeneous environmental inputs support standardized model comparisons, while reinforcement learning identifies influential features. Benchmark Framework of Data-driven PFDF Prediction, including data curation, evaluation metrics, baseline, and feature selection. Heterogeneous environmental inputs support standardized model comparisons, while reinforcement learning identifies influential features.
arXivInterpretation
Establishes a unified PFDF data-driven prediction benchmark comparing the fixed USGS operational model with trainable baselines (LR, CART, SVM, NN, RF, HGB) and their augmented-feature variants under identical data splits, feature tables, and evaluation protocols. Prior studies were fragmented across feature spaces, model architectures, and evaluation protocols, making rigorous comparison difficult; the benchmark uses zone-stratified repeated splits and Threat Score as the primary metric, enabling fair cross-method comparison and scientific insight derivation. Built on the Staley et al. post-fire debris-flow database with 1,243 storm-watershed events from 33 fires across 10 geographic zones, including 316 positive cases (25.4%); the complete-case subset has 1,131 events with 292 positives (25.8%). Four zone-stratified train/validation/test ratios (50/25/25, 60/20/20, 70/15/15, 80/10/10) and five split seeds are used, reporting mean test Threat Score and standard deviation across five seeds.
Uses LLM-based text processing to convert operational weather forecasts (Area Forecast Discussions, Hazardous Weather Outlooks) into structured PFDF weather indicators, and combines USGS 3DEP terrain attributes with NLDAS-2 soil-moisture summaries and time since fire into an augmented feature table. Most existing approaches rely on rainfall-specific statistics while overlooking the richer meteorological context and storm dynamics in operational weather forecast reports; this work turns narrative forecasts into usable predictive signals. Augmented-feature variants further improve most model families, particularly SVM, random forest, and neural network models; the original feature table already provides strong performance for tree-based ensembles.
Proposes a reinforcement-learning feature-selection framework that learns a binary feature-removal mask, measures predictive separation between positive and negative samples via a standardized mean gap, and optimizes masks to maximally collapse that separation relative to the unmasked input, thereby identifying factors important for PFDF discrimination. Provides interpretable global and regional driver analysis rather than prediction accuracy alone; masked features are replaced by training-set reference values, and a sparsity term avoids suppressing excessive features. Optimizes the discrete masking objective with a REINFORCE-style policy gradient; results show storm-related variables dominate globally, soil erodibility, burn severity, and terrain slope also consistently emerge as critical, and the geographical zone itself appears as a highly important feature highlighting strong region-specific characteristics.
Regional diagnostics show PFDF mechanisms are not uniform: OSD and SGSBSJ are dominated by storm/rainfall factors, while SWCO and WMT are dominated by terrain and geographical features; the fixed USGS formula misses 9 of 25 PFDF events in OSD, 8 of 10 in SWCO, and only 1 of 18 in MT. Indicates that globally fixed coefficients are insufficient across diverse geographic zones, whereas adaptive learning models can capture region-specific dynamics while remaining physically interpretable. In SWCO, observed PFDF events have a smaller mean rainfall forcing than non-PFDF events, violating the monotonic relationship assumed by the USGS equation and causing the fixed model to rank many PFDF events too low; in MT, USGS variables already align well with observed outcomes, so there is less capacity for a flexible neural model to improve, and additional learned interactions can slightly reduce performance under the repeated split setting.
Perspective
The benchmark targets binary PFDF occurrence prediction at the storm-watershed-event level for monitored burned watersheds in the western United States, applicable where storm forcing, watershed and burn attributes, and binary labels are available; augmented features depend on the availability of USGS 3DEP terrain, NLDAS-2 soil moisture, and operational weather forecast text. The zone-stratified evaluation and Threat Score design serve imbalanced hazard detection and enable comparison of the fixed operational model with trainable models under identical protocols. Code and benchmark are publicly released, supporting reproduction and extension under the same data and splits.
The loaded text is a fast parse: Table 2 and Figures 2 and 3 do not include their numeric values, so the Threat Score of each model under each split ratio, the magnitude of improvement from augmented features, and the full top-10 feature rankings per zone cannot be quantified here; these details require the original figures and tables. Regional mechanism conclusions are based on a specific database and zone partition, and generalization to other geographic regions or datasets remains to be verified. Future work proposes developing physics-informed equations for regional PFDF prediction and investigating the trustworthiness of debris-flow prediction systems, including vulnerability to knowledge poisoning and data extraction attacks; these directions are yet to be pursued.
