PyroAdapt lifts daily California wildfire average precision from 21.62% to 24.35–24.57% and captures 344 extra positive cell-days under a fixed 34-cell budget
Synopsis
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
Interpretation
PyroAdapt recasts wildfire occurrence prediction as a batch-transductive pretrain–retrieve–rank adaptation pipeline: historical focal pretraining yields a risk representation, unlabeled target-year covariates retrieve similar historical examples, and the retrieved historical labels form positive–negative pairs for ranking fine-tuning, with no target-year labels used. Compared with source-only training or continued focal fine-tuning, it separates representation learning, target-conditioned sample selection, and decision-aligned adaptation, and it makes retrieval a sample-selection step rather than label propagation. The paper gives the full pipeline and equations and runs it in two settings: Yosemite (9 cells, 3,285 target grid-days in 2021 with 363 positives) and California (666 cells, 351 fire dates and 14,760 positives in 2022); California retrieval queries all 243,090 unlabeled target covariates and yields 363,643 unique historical examples.
A unified score-gap formulation compares direct ranking, residual pairwise DPO (RDPO), and selective ranking, showing how each allocates gradients across pretrained-correct and pretrained-wrong pairs, and yields a finite-sample bound on avoidable misses under a fixed daily budget for nonnegative targets. DPO-style objectives have mostly been used for language-model preference optimization; here they are connected to pairwise learning-to-rank and cost-sensitive classification, with a derived reduction and conditional minimizer for binary-label RDPO. Propositions 1, 2, 5, and 6 give the gradient allocation, target-sensitivity bound, and budget certificate, with proofs in Appendix A; the paper also notes these are soft objectives whose task-level correction and retention still require empirical evaluation.
Across the full California domain, all three ranking objectives beat continued focal fine-tuning without focal supervision: daily average precision of 24.35%/24.47%/24.57% versus 21.62%, and Top5% recall of 22.25%/22.11%/22.79% versus 18.70%; at a fixed daily budget of 34 cells, selective ranking captures a mean of 3,210.3 positive cell-days versus 2,866.0 for continued focal, 344 more. Rather than treating wildfire prediction as pointwise classification or global ranking, it restricts training pairs to the same day so the ranking objective aligns with the decision of which cells to inspect first under a limited daily monitoring budget. Results come from three model seeds (11, 22, 33) evaluated on the complete 666-cell daily grid with reported means and standard deviations; the paper notes these are neighborhood-smoothed positive cell-days, not independent incidents.
In strata defined by dry matter consumption, selective ranking raises recall by 39.70/28.18/20.50 percentage points for the top 5%/10%/20% strata, and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift. This extends evaluation from overall ranking to sensitivity for high-activity events and tests across rolling years whether adaptation degrades as conditions change. The California strata use 2,193 raw-DM-positive test examples with each policy's historical F1 threshold; the Yosemite rolling evaluation covers 2019–2021 across nine year–seed cases, where direct ranking reaches rolling AUROC 74.19% and AUPRC 24.30% versus 73.12% and 22.85% for the pretrained model.
Perspective
The work targets the decision of which cells to prioritize under a fixed daily monitoring budget: the California experiment covers 666 0.25° cells for a single target year (2022) and only locations present historically, while the Yosemite experiment tests year-to-year adaptation on a fixed 9-cell grid. It fits as a downstream adaptation layer on a historically pretrained risk model, for research and operational teams that need to rank candidate locations within a limited daily alarm count, and it pairs well with region-stratified auditing and region-level performance reporting. The paper also lays out untested designs for larger-area deployment, including partial pooling of regional residuals, support-aware retrieval distances, within-region and cross-region preference mixtures, and calibration shrinkage for sparse regions.
The California evaluation covers a single target year and uses covariates from the entire target year during adaptation, a batch-transductive setting; the paper states that issue-time adaptation using only then-available covariates remains unevaluated. The three ranking objectives share a stopping rule but differ in realized update counts, so small differences between them should not be attributed to the loss form alone. The high-DM strata are defined retrospectively and each policy uses its own historical F1 threshold with unmatched alarm rates, so that diagnostic speaks to sensitivity for high observed activity rather than severity forecasting or equal-budget severe-event discrimination. Retrieval depends on useful overlap between historical and target covariates, and the paper notes that retrieval cannot reconstruct a target regime absent from the archive. Labels are neighborhood-smoothed occurrence indicators rather than verified ignitions, and positive cell-days may come from the same fire, so they should not be read as counts of independent incidents. The loaded text is the full paper, but the homepage bundle's community discussion and citation entries are empty, so external replication or follow-up citation activity would need to be checked separately.
