A satellite image holds far fewer independent samples than pixels: correlation range sets the effective count, and random holdout overstates accuracy
Lead
For a remote sensing image whose spatial correlation persists over about r pixels, a classifier has roughly n/r² independent samples rather than n pixels, a rate locked by matching upper and lower bounds that explains why random holdout overstates accuracy.
Story
Neighboring pixels in a remote sensing image carry redundant information, so an image with n labeled pixels contains only about n/r² independent samples, where r is the distance over which spatial correlation persists, the correlation range. Earlier learning-theoretic guarantees assumed independent and identically distributed samples and computed error bounds directly from the pixel count, which fails on spatially autocorrelated imagery. On synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1), the measured effective sample size falls as a power law in block size with a slope consistent with the theoretical prediction, and effective sample sizes are typically one to three orders of magnitude smaller than pixel counts.
The method partitions the image into small blocks separated by gaps comparable to the correlation range, so the blocks are nearly independent and standard concentration inequalities apply to them. Applying independent-sample concentration directly to correlated pixels underestimates error, while gaps that are too wide waste pixels and reduce the number of usable blocks. The proof gives a finite-sample upper bound on a two-dimensional mixing random field, and a matching minimax lower bound shows the rate is tight up to logarithmic factors, so no algorithm can do better.
What to watch
A next step is to carry the blocking argument to space-time lattices, where a spatial and a temporal correlation range both enter the effective sample size and quantify the marginal value of each revisit. Teams doing land cover mapping and accuracy assessment can compute the effective sample size per stratum, set cross-validation fold separation from the correlation range, and replace the nominal sample size when reporting confidence intervals.
The bounds are worst-case and conservative, with unoptimized constants and polylogarithmic factors, so the value lies in the scaling and the cross-validation guidance rather than a sharp constant. The correlation range must be estimated from an empirical variogram, and the bounds are driven by the mixing rate of the loss field while the measured variograms are computed on raw features; the two need not coincide, and when intra-class variability is small the feature range overestimates the loss-field range, making the effective sample size conservative. Whether the assumption that label noise is conditionally independent given features still holds for tasks with spatial label structure, such as segmentation, remains to be tested.
