Texture-driven visual learning suffers low-frequency shortcuts: pruning low-frequency components raises in-distribution accuracy by up to 10% and out-of-distribution accuracy by up to 40%
Related research and updatesSynopsis
This work analyzes shortcut learning in texture-driven visual domains and compares it with a shape-driven standard benchmark, finding that texture-driven domains base most decisions on a few low-frequency components (LFCs) with skewed spectral behavior even though higher-frequency components (HFCs) have higher predictive power; pruning LFCs from training and test sets mitigates the shortcut and yields more balanced spectral behavior, improving in-distribution accuracy by up to 10% and out-of-distribution accuracy by up to 40% under algorithmic and real-world domain shifts, while general-purpose and domain-specific foundation models can also suffer from low-frequency shortcuts.
Figure 1 : Spectral behavior of models trained on texture-driven domains show that models make majority of their decisions based on a few low-frequency components (first three figures above, from left). Once we prune low-frequency components (LFCs) from training and test sets, spectral behavior shifts towards higher-frequencies (forth figure above), and we obtain a significantly higher ID accuracy, up to 10% (Section 5 ). These results show that texture-driven domains suffer from low-frequency shortcuts, where models learn from LFCs despite that higher-frequency components have a higher predictive power. We show that low-frequency shortcuts persist across different model sizes, architectures, and hyperparameters (Section 5 ), under various OOD conditions (Section 6 & 7 ), for pretrained foundation models (Section 8 ), and different resolutions, transformation block sizes, and color spaces (Section 9 ).
arXivInterpretation
Texture-driven domains exhibit low-frequency shortcuts: models make the majority of their decisions based on a few low-frequency components with skewed spectral behavior, despite higher-frequency components having higher predictive power. Existing shortcut-learning studies are all based on a few shape-driven standard benchmarks; this work extends the analysis to texture-driven domains and compares them with a standard benchmark, identifying a low-frequency shortcut specific to that setting. Based on a comparative analysis of texture-driven domains versus a standard benchmark, characterized through spectral behavior and predictive-power comparisons; the abstract does not name specific datasets or sample sizes.
Pruning low-frequency components from training and test sets mitigates the shortcut and produces more balanced spectral behavior, improving in-distribution accuracy by up to 10% and out-of-distribution accuracy by up to 40% under algorithmic and real-world domain shifts. Proposes LFC pruning as a mitigation and reports both ID and OOD accuracy gains, covering algorithmic domain shifts and real-world domain shifts. Evidence comes from accuracy comparisons before and after LFC pruning, with gains reported as 'up to' values, i.e., observations under the paper's experimental settings.
General-purpose and domain-specific foundation models can also suffer from low-frequency shortcuts; while large models can mitigate the shortcuts, they incur high computational cost and may achieve significantly lower accuracy than shortcut-pruned from-scratch trained small models. Extends the low-frequency shortcut diagnosis from conventionally trained models to foundation models and offers a cost-versus-accuracy comparison between large models and pruned small models. Based on observations of general-purpose and domain-specific foundation models and a comparison between large models and shortcut-pruned small models; the abstract does not give specific model names or numbers.
Reduced image resolutions amplify the degree of shortcuts; large frequency-transformation block sizes capture low-frequency shortcuts better than small block sizes; and low-frequency shortcuts persist across different color spaces. Provides three conditional observations about what modulates low-frequency shortcuts (resolution, frequency-transformation block size, color space), offering actionable cues for practitioners diagnosing and designing in new domains. Derived from conditional comparisons across resolution, block size, and color space; the abstract reports no specific values or statistical tests.
Perspective
This work targets texture-driven visual learning domains and is meant for practitioners building or evaluating models in new, understudied domains, especially where out-of-distribution generalization and domain shift matter. Its mitigation (pruning low-frequency components from training and test sets) can be embedded directly into existing training and evaluation pipelines and combined with from-scratch trained small models; for foundation-model users, the analysis suggests that pruned small models can be a competitive option when compute is constrained. The scope of the conclusions is set by texture-driven domains, and the comparison with a shape-driven standard benchmark serves to highlight differences rather than to replace that benchmark.
The abstract reports accuracy gains as 'up to' values and omits specific datasets, sample sizes, model architectures, and statistical tests, so the stability and applicability boundaries of the gains still need to be judged against the paper's experimental settings. The observation that low-frequency shortcuts persist across color spaces is not expanded in the abstract regarding which color spaces or transformations were used. The accuracy and cost comparison between large models and shortcut-pruned small models lacks specific models and numbers in the abstract, so readers interested in that trade-off should consult the full text. In addition, whether pruning LFCs on the test set applies to online settings where the test distribution is not known in advance is not addressed in the abstract and is an open question worth watching.
