Learning Robust Global Representations by Penalizing Local Predictive Power
Haohan WangSongwei GeEric P. XingZachary C. Lipton
Proposes a training method that forces convolutional networks to prioritize global shapes over local textures by penalizing early-layer predictive power, significantly boosting out-of-domain generalization and introducing the ImageNet-Sketch benchmark for cross-domain evaluation.
Modern computer vision models often achieve high accuracy under standard conditions but suffer sharp performance drops when deployed in new environments. This instability occurs because conventional convolutional neural networks tend to rely heavily on superficial, local visual signals—such as background colors, textures, and small image patches—rather than grasping the overall shape and structure of objects. When real-world test conditions change, these superficial correlations break down, creating severe operational and safety risks for machine learning systems.
The article demonstrates a novel training strategy, termed Patch-wise Adversarial Regularization, designed to force image classifiers to discard predictive local cues and instead learn robust global object concepts. It evaluates whether penalizing the predictive utility of early-layer image representations improves generalization across unexpected domain shifts without requiring advance knowledge or data from the target deployment domain.
The authors implemented a training mechanism that adds secondary, patch-level classifiers to early neural network layers. Using an adversarial reverse-gradient approach, the network is trained to maximize classification accuracy at the final output layer while simultaneously preventing early layers from making accurate predictions based on isolated, local image patches alone. The authors evaluated this method across synthetic benchmarks, established domain shift testbeds, and a newly constructed large-scale evaluation set called ImageNet-Sketch, which comprises 50,000 black-and-white sketch images across 1,000 categories designed to match the scale of standard image benchmarks.
The experiments show that Patch-wise Adversarial Regularization consistently improves model generalization across altered environments. On perturbed datasets testing robustness against altered color and texture, the approach outperformed baseline architectures and domain adaptation techniques, raising average accuracy on perturbed CIFAR-10 tasks from 63.9% to 66.5%. On the multi-domain PACS benchmark, the method achieved state-of-the-art results among approaches that do not use domain labels, showing its largest advantage in the colorless sketch domain where local color cues are absent. Furthermore, on the new ImageNet-Sketch benchmark, regularized networks successfully classified complex shapes where standard models were misled by local textures—such as confusing a furry dog for a mop or a tricycle frame for a safety pin.
These findings indicate that actively suppressing reliance on local visual shortcuts is an effective, practical way to build resilient visual classifiers. By encouraging models to base decisions on holistic object geometry rather than background or surface texture, organizations can deploy vision systems that generalize better to real-world variations without needing prior access to target-domain data or complex multi-domain training labels. This reduces the risk of silent out-of-distribution failures in critical applications.
For practitioners seeking to improve computer vision robustness, adopting patch-wise regularization offers a straightforward enhancement that can be applied directly by fine-tuning existing pretrained models. When implementing the method, engineering teams should tune the regularization penalty carefully, as excessively high regularization strengths can destabilize early training. Future work should focus on establishing structured guidelines for choosing between architectural variants and exploring broader applications across diverse computer vision pipelines.
- Paper: ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness, Robert Geirhos et al. (2018). This paper establishes that standard CNNs rely disproportionately on local textures rather than global shapes, directly motivating the source paper's objective of penalizing local predictive power to learn robust global representations.
- Paper: Deeper, Broader and Artier Domain Generalization, Da Li et al. (2017). This foundational work introduces multi-domain visual evaluation across sketches and photos (PACS), providing critical context for out-of-domain generalization and sketch-based evaluation.
- Paper: Deep Domain Confusion: Maximizing for Domain Invariance, Eric Tzeng et al. (2014). This study introduces early deep domain-confusion methods to align representations across disparate visual domains, foundational to representation learning for domain adaptation.
- Paper: Improved Regularization of Convolutional Neural Networks with Cutout, Terrance Devries et al. (2017). This work demonstrates how masking localized image regions forces networks to rely on broader visual context, providing conceptual groundwork for discouraging local feature reliance.
- Paper: Learning to Generalize: Meta-Learning for Domain Generalization, Da Li et al. (2017). This paper formulates domain generalization via meta-learning over shifted source distributions, establishing core benchmark setups used to test out-of-domain generalization.
- Paper: The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization, Dan Hendrycks et al. (2021). This study systematically evaluates diverse distribution shifts and out-of-domain generalization techniques across large-scale vision benchmarks, directly building on the evaluation paradigms and datasets like ImageNet-Sketch introduced in the source.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This comprehensive survey contextualizes representation-learning and feature-alignment methods for domain generalization, providing a broad framework that synthesizes techniques like local-signal suppression.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). This survey analyzes domain generalization taxonomy and benchmarks, extending the discussion on how learning domain-invariant, global representations enhances robustness to unseen domains.
