FairRSFM evaluates three remote sensing foundation models by biome group and finds aggregate metrics mask ecological performance gaps
Related research and updatesSynopsis
The authors introduce FairRSFM, which maps georeferenced samples from four remote sensing downstream datasets (m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K) to 14 terrestrial biome classes consolidated into six macro-groups, and evaluates Prithvi-EO-2.0, SatMAE, and DOFA under a unified frozen-backbone protocol across three random seeds, showing that aggregate metrics consistently mask biome-dependent disparities, for example Prithvi-EO-2.0 reaching high overall macro-F1 on m-EuroSAT while its worst-group score is markedly lower, and compares BOLP, DBR, and GroupDRO as mitigation baselines whose effectiveness is model- and task-dependent.
Interpretation
The paper builds FairRSFM, a biome-aware benchmark and debiasing framework for remote sensing foundation model robustness, covering classification and semantic segmentation across four datasets. Existing remote sensing benchmarks such as GEO-Bench mainly summarize performance at the dataset level, whereas FairRSFM turns standard transfer evaluation into a group-robustness problem that asks whether models perform consistently across biomes. The benchmark includes m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K, all providing public geocoordinates for reproducible biome assignment; MMEarth20K is a 20,000-sample subset built from MMEarth with Dynamic World label maps, with macro-group quotas derived by inverse-frequency weighting and allocated uniformly across the 14 biomes.
The paper proposes a two-level biome-labeling pipeline that first assigns georeferenced samples to 14 terrestrial biome classes and then consolidates them into six macro-groups reflecting spectral, phenological, hydrological, cryospheric, and surface-property regimes. Using all 14 raw classes as primary evaluation groups can be unstable for datasets with limited or uneven biome coverage, so the macro-groups pool under-represented biomes to reduce sparsity while the 14 labels are retained for fine-grained analysis. The mapping is deterministic and fixed across all datasets, models, and mitigation methods; unknown, water-only, or unmatched samples are excluded from primary group-based summaries; the supplementary material reports that centroid and majority-pixel spatial joins agreed on a high fraction of samples across all four datasets.
Under a unified frozen-backbone protocol, all three models across the four datasets show that aggregate performance masks biome-dependent disparities. The result establishes ecological group robustness as an evaluation dimension independent of average accuracy, indicating that high aggregate scores do not imply consistent reliability across ecological regions. Three backbones (Prithvi-EO-2.0, SatMAE, DOFA) and three random seeds are used, reporting overall metrics alongside worst-group score, normalized failure range, ECE, EOdd, and DPM; for example, Prithvi-EO-2.0 reaches high overall macro-F1 on m-EuroSAT while its worst-group score is markedly lower, and m-SA-Crop-Type shows much lower mIoU on the Xeric and Mineralogical group than overall.
The paper compares BOLP, DBR, and GroupDRO as mitigation baselines and finds that worst-group robustness gains are model- and task-dependent and can involve trade-offs with aggregate performance. BOLP removes dominant biome-associated directions from frozen embeddings through a closed-form projection without updating the backbone or adding trainable parameters, offering a reusable representation-level approach to ecological debiasing. On Prithvi-EO-2.0 with m-BigEarthNet, BOLP improves both overall and worst-group F1@opt and reduces NFR without updating the backbone; on dense segmentation such as MMEarth20K, BOLP reduces both overall and worst-group performance while GroupDRO improves the worst-group score, indicating no single strategy dominates across architectures and tasks.
Perspective
The work targets remote sensing foundation model researchers and practitioners who evaluate downstream transfer with frozen backbones, and applies to classification and dense segmentation datasets that carry geocoordinates and support biome assignment; its diagnostic and mitigation findings hold under the four datasets, three backbones, and unified frozen-backbone protocol studied, and mitigation effects should be read per model and task.
Biome labels are assigned at the patch level from sample coordinates and can be noisy near biome boundaries or in multi-biome tiles, which the paper treats as a consistent grouping rule rather than a pixel-perfect ecological map; datasets cover different numbers of macro-groups, so worst-group and disparity metrics are computed only over matched groups in each split, and small groups (for example 14-25 test samples) can increase sampling variability; the study covers three backbones and four datasets, leaving broader sensor families, additional backbones, and sensitivity analysis over the number of macro-groups as open directions; ecological group-disparity metrics are less standardized for segmentation than for classification; and biome grouping captures only one spatially meaningful aspect of geographic bias, not sensor and source shifts, geographic imbalance, spatial-resolution differences, temporal acquisition effects, or uneven imagery coverage.
