Learning Robust Global Representations by Penalizing Local Predictive Power

Haohan WangSongwei GeEric P. XingZachary C. Lipton

article2019NeurIPS1,396 citationsBest Paper Honorable Mention

Proposes a training method that forces convolutional networks to prioritize global shapes over local textures by penalizing early-layer predictive power, significantly boosting out-of-domain generalization and introducing the ImageNet-Sketch benchmark for cross-domain evaluation.

Listen

Modern computer vision models often achieve high accuracy under standard conditions but suffer sharp performance drops when deployed in new environments. This instability occurs because conventional convolutional neural networks tend to rely heavily on superficial, local visual signals—such as background colors, textures, and small image patches—rather than grasping the overall shape and structure of objects. When real-world test conditions change, these superficial correlations break down, creating severe operational and safety risks for machine learning systems.

The article demonstrates a novel training strategy, termed Patch-wise Adversarial Regularization, designed to force image classifiers to discard predictive local cues and instead learn robust global object concepts. It evaluates whether penalizing the predictive utility of early-layer image representations improves generalization across unexpected domain shifts without requiring advance knowledge or data from the target deployment domain.

The authors implemented a training mechanism that adds secondary, patch-level classifiers to early neural network layers. Using an adversarial reverse-gradient approach, the network is trained to maximize classification accuracy at the final output layer while simultaneously preventing early layers from making accurate predictions based on isolated, local image patches alone. The authors evaluated this method across synthetic benchmarks, established domain shift testbeds, and a newly constructed large-scale evaluation set called ImageNet-Sketch, which comprises 50,000 black-and-white sketch images across 1,000 categories designed to match the scale of standard image benchmarks.

The experiments show that Patch-wise Adversarial Regularization consistently improves model generalization across altered environments. On perturbed datasets testing robustness against altered color and texture, the approach outperformed baseline architectures and domain adaptation techniques, raising average accuracy on perturbed CIFAR-10 tasks from 63.9% to 66.5%. On the multi-domain PACS benchmark, the method achieved state-of-the-art results among approaches that do not use domain labels, showing its largest advantage in the colorless sketch domain where local color cues are absent. Furthermore, on the new ImageNet-Sketch benchmark, regularized networks successfully classified complex shapes where standard models were misled by local textures—such as confusing a furry dog for a mop or a tricycle frame for a safety pin.

These findings indicate that actively suppressing reliance on local visual shortcuts is an effective, practical way to build resilient visual classifiers. By encouraging models to base decisions on holistic object geometry rather than background or surface texture, organizations can deploy vision systems that generalize better to real-world variations without needing prior access to target-domain data or complex multi-domain training labels. This reduces the risk of silent out-of-distribution failures in critical applications.

For practitioners seeking to improve computer vision robustness, adopting patch-wise regularization offers a straightforward enhancement that can be applied directly by fine-tuning existing pretrained models. When implementing the method, engineering teams should tune the regularization penalty carefully, as excessively high regularization strengths can destabilize early training. Future work should focus on establishing structured guidelines for choosing between architectural variants and exploring broader applications across diverse computer vision pipelines.

arXiv: 1905.13549
  • Paper: The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization, Dan Hendrycks et al. (2021). This study systematically evaluates diverse distribution shifts and out-of-domain generalization techniques across large-scale vision benchmarks, directly building on the evaluation paradigms and datasets like ImageNet-Sketch introduced in the source.
  • Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This comprehensive survey contextualizes representation-learning and feature-alignment methods for domain generalization, providing a broad framework that synthesizes techniques like local-signal suppression.
  • Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). This survey analyzes domain generalization taxonomy and benchmarks, extending the discussion on how learning domain-invariant, global representations enhances robustness to unseen domains.
Cover for Learning Robust Global Representations by Penalizing Local Predictive Power

Abstract

Despite their renowned predictive power on i.i.d. data, convolutional neural networks are known to rely more on high-frequency patterns that humans deem superficial than on low-frequency patterns that agree better with intuitions about what constitutes category membership. This paper proposes a method for training robust convolutional networks by penalizing the predictive power of the local representations learned by earlier layers. Intuitively, our networks are forced to discard predictive signals such as color and texture that can be gleaned from local receptive fields and to rely instead on the global structures of the image. Across a battery of synthetic and benchmark domain adaptation tasks, our method confers improved generalization out of the domain. Also, to evaluate cross-domain transfer, we introduce ImageNet-Sketch, a new dataset consisting of sketch-like images, that matches the ImageNet classification validation set in categories and scale.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Patch-wise Adversarial Regularization
  • 3.2 Other Extensions and Training Heuristics
  • 4 Experiments
  • 4.1 MNIST with Perturbation
  • 4.2 CIFAR with Perturbation
  • 4.3 PACS
  • 4.4 ImageNet-Sketch
  • 4.4.1 The ImageNet-Sketch Data
  • 4.4.2 Experiment Results
  • 5 Conclusion
  • References
  • A Other Hyperparameter Choices for MNIST experiment
  • B Cifar10 discussion
  • C More results of ImageNet-Sketch

Knowls

  1. Knowl 1 — Patch-wise Adversarial Regularization (PAR) Framework and Objective

    model/method

    Patch-wise Adversarial Regularization (PAR) is a learning framework designed to force convolutional neural networks (CNNs) to rely on global object structures rather than local, superficial predictive patterns (such as texture or local color cues). Let ⟨X,y⟩\langle X, y \rangle denote an input image and its ground-truth label, and let the neural network be decomposed as f(g(X;δ);θ)f(g(X; \delta); \theta), where g(⋅;δ)g(\cdot; \delta) denotes the bottom convolutional layer(s) parametrized by δ\delta producing an intermediate activation tensor of dimension c×m′×n′c \times m' \times n', and f(⋅;θ)f(\cdot; \theta) represents the remaining layers parametrized by θ\theta.

    A patch-wise side classifier h(⋅;ϕ)h(\cdot; \phi) takes an individual cc-dimensional feature vector at spatial location (i,j)(i, j) in g(X;δ)g(X; \delta) and predicts the label yy. The joint optimization problem is formulated as:

    min⁡δ,θE(X,y)[l(f(g(X;δ);θ),y)−λm′n′∑i=1m′∑j=1n′l(h(g(X;δ)i,j;ϕ),y)]\min_{\delta, \theta} \mathbb{E}_{(X, y)} \left[ l(f(g(X; \delta); \theta), y) - \frac{\lambda}{m' n'} \sum_{i=1}^{m'} \sum_{j=1}^{n'} l(h(g(X; \delta)_{i,j}; \phi), y) \right]

    min⁡ϕE(X,y)[λm′n′∑i=1m′∑j=1n′l(h(g(X;δ)i,j;ϕ),y)]\min_{\phi} \mathbb{E}_{(X, y)} \left[ \frac{\lambda}{m' n'} \sum_{i=1}^{m'} \sum_{j=1}^{n'} l(h(g(X; \delta)_{i,j}; \phi), y) \right]

    where l(⋅,⋅)l(\cdot, \cdot) is the cross-entropy classification loss and λ>0\lambda > 0 is a regularization hyperparameter. Via gradient reversal between g(⋅;δ)g(\cdot; \delta) and h(⋅;ϕ)h(\cdot; \phi), the lower representation layers are penalized whenever any individual spatial receptive field contains sufficient predictive power for yy, compelling the network to base its final predictions on representations formed by aggregating patterns across multiple receptive fields.

  2. Knowl 2 — Efficient Implementation and Architectural Variants of PAR

    model/method

    In the standard PAR setup, the patch classifier h(⋅;ϕ)h(\cdot; \phi) is implemented as a 1×11 \times 1 convolutional operation with cc input channels and kk output channels (where kk is the number of classes), parametrized by a c×kc \times k weight matrix and a kk-length bias vector. Applying a single set of parameters ϕ\phi across all m′×n′m' \times n' spatial locations reduces computational overhead and accounts for the fact that predictive local patterns may appear at arbitrary spatial coordinates across images.

    Three extensions to the base formulation explore classifier capacity, receptive field size, and depth:

    • PARM\text{PAR}_M (More Powerful Pattern Classifier): Replaces the single linear layer in h(⋅;ϕ)h(\cdot; \phi) with a 3-layer multilayer perceptron (MLP) with ReLU activations.
    • PARB\text{PAR}_B (Broader Local Pattern): Replaces the 1×11 \times 1 convolution with a 3×33 \times 3 convolution, expanding the receptive field of what is regularized as local.
    • PARH\text{PAR}_H (Higher-Level Local Concept): Applies the adversarial regularization penalty to the activations produced at the output of the second convolutional layer rather than the first layer.

    Additionally, a two-stage training heuristic is employed: the CNN backbone is first pretrained under standard classification loss until convergence (or for a set number of epochs), after which all parameters are fine-tuned jointly with the adversarial regularization objective.

  3. Knowl 3 — ImageNet-Sketch Benchmark Dataset

    definition

    ImageNet-Sketch is an out-of-domain evaluation benchmark designed to measure the generalization capability of vision models trained on standard ImageNet (ILSVRC 2012) when tested on out-of-domain sketch images matching the exact ImageNet taxonomy and scale.

    • Taxonomy and Scale: Comprises 50,000 black-and-white sketch images spanning all 1,000 ImageNet categories, with exactly 50 images per class.
    • Data Collection: Images were retrieved independently of the original ImageNet dataset by querying Google Images with the string "sketch of <class_name>" within a black-and-white color filter.
    • Cleaning and Augmentation: An initial set of 100 images per category was retrieved and manually curated to remove irrelevant images or confusing category overlaps. For categories with fewer than 50 cleaned images, data augmentation (horizontal flipping and rotation) was applied to achieve 50 images per class.

    Because sketch images eliminate realistic colors and photographic surface textures while retaining canonical shapes, the dataset evaluates whether vision models utilize global geometric structure rather than surface textures.

  4. Knowl 4 — Out-of-Domain Classification Accuracy on ImageNet-Sketch

    data/table

    Evaluating AlexNet-based models on the 50,000-image ImageNet-Sketch benchmark shows that penalizing local predictive representations via PAR improves zero-shot transfer performance compared to standard training and baseline domain adaptation/regularization methods.

    Method AlexNet InfoDrop HEX PAR PARB\text{PAR}_B PARM\text{PAR}_M PARH\text{PAR}_H
    Top-1 Accuracy 0.1204 0.1224 0.1292 0.1306 0.1273 0.1287 0.1266
    Top-5 Accuracy 0.2480 0.2560 0.2654 0.2627 0.2575 0.2603 0.2544
    Method DANN†\text{DANN}^\dagger Jigen∗\text{Jigen}^* PAR∗\text{PAR}^* PARB∗\text{PAR}_B^* PARM∗\text{PAR}_M^* PARH∗\text{PAR}_H^*
    Top-1 Accuracy 0.1360 0.1469 0.1494 0.1494 0.1501 0.1499
    Top-5 Accuracy 0.2712 0.2898 0.2949 0.2945 0.2957 0.2954

    In the table, †^\dagger denotes methods with access to unlabeled target domain images during training, and ∗^* denotes methods utilizing extra data augmentation techniques (e.g., converting image patches to grayscale). Fine-tuning AlexNet with vanilla PAR for 5 epochs on original ImageNet images achieves a top-1 accuracy of 0.1306 (vs. 0.1204 for AlexNet and 0.1292 for HEX). When combined with data augmentation, PARM∗\text{PAR}_M^* achieves 0.1501 top-1 accuracy.

  5. Knowl 5 — Domain Generalization Performance on the PACS Benchmark

    data/table

    The PACS benchmark evaluates domain generalization across four visual domains: Art painting, Cartoon, Photo, and Sketch. Models are trained on three source domains without target domain identifiers or target domain samples and tested on the remaining held-out domain using an AlexNet backbone.

    Method Art Cartoon Photo Sketch Average
    AlexNet Baseline 63.3 63.1 87.7 54.0 67.03
    DSN 61.1 66.5 83.2 58.5 67.33
    L-CNN 62.8 66.9 89.5 57.5 69.18
    MLDG 63.6 63.4 87.8 54.9 67.43
    Fusion 64.1 66.8 90.2 60.1 70.30
    MetaReg 69.8 70.4 91.1 59.2 72.63
    HEX 66.8 69.7 87.9 56.3 70.18
    PAR 66.9 67.1 88.6 62.6 71.30
    PARB\text{PAR}_B 66.3 67.8 87.2 61.8 70.78
    PARM\text{PAR}_M 65.7 68.1 88.9 61.7 71.10
    PARH\text{PAR}_H 66.3 68.3 89.6 64.1 72.08
    Jigen∗\text{Jigen}^* 67.6 71.7 89.0 65.1 73.38
    PAR∗\text{PAR}^* 68.0 71.6 90.8 61.8 73.05
    PARB∗\text{PAR}_B^* 67.6 70.7 90.1 62.0 72.59
    PARM∗\text{PAR}_M^* 68.7 71.5 90.5 62.6 73.33
    PARH∗\text{PAR}_H^* 68.7 70.5 90.4 64.6 73.54

    Methods marked with ∗^* follow the training setup of Carlucci et al. (2019) with grayscale patch data augmentation and random splits. Among methods without data augmentation, PARH\text{PAR}_H achieves the highest accuracy on the Sketch domain (64.1%64.1\%, outperforming the AlexNet baseline of 54.0%54.0\% and Fusion's 60.1%60.1\%). This indicates PAR's effectiveness when generalizing to colorless domains where local color cues are removed.

  6. Knowl 6 — Robustness on CIFAR-10 with Color and Texture Perturbations

    data/table

    To evaluate robustness against superficial statistical shifts in color and texture, a ResNet-50 model trained on standard CIFAR-10 (achieving ~92% clean accuracy) was tested against four perturbed test variations: Greyscale, Negative Color, Random Fourier Kernel filtering, and Radial Fourier Kernel filtering.

    Dataset Shift ResNet-50 DANN InfoDrop HEX PAR PARB\text{PAR}_B PARM\text{PAR}_M PARH\text{PAR}_H
    Greyscale 87.7% 87.3% 86.4% 87.6% 88.1% 87.9% 87.8% 86.9%
    Negative Color 62.8% 64.3% 57.6% 62.4% 66.2% 65.3% 67.6% 62.7%
    RandKernel 43.0% 33.4% 41.3% 42.5% 47.0% 40.5% 47.5% 40.8%
    RadialKernel 62.4% 63.3% 60.3% 61.9% 63.8% 63.2% 63.2% 61.4%
    Average 63.9% 62.0% 61.4% 63.6% 66.3% 64.2% 66.5% 62.9%

    PAR models were pre-trained for 250 epochs followed by 150 epochs of adversarial regularization training, while competing methods were trained for 400 epochs. PAR and PARM\text{PAR}_M outperform standard ResNet-50 and unsupervised domain adaptation (DANN, despite DANN having access to unlabelled test data), achieving up to 66.5%66.5\% average perturbed test accuracy compared to ResNet-50's 63.9%63.9\%.

  7. Knowl 7 — Domain Generalization on MNIST with Superficial Texture Perturbations

    empirical result

    In domain generalization experiments on MNIST perturbed with superficial texture patterns (original digit, radial Fourier kernel mask, random Fourier kernel mask), training sets contained two pattern types and testing sets contained the held-out third pattern type under two settings:

    1. Independent setting: Perturbation patterns were assigned randomly and independently of the digit label.
    2. Dependent setting: Digit labels 0–40\text{--}4 received one pattern and digits 5–95\text{--}9 received another, creating a strong spurious correlation between superficial texture and target label during training.

    PAR and its variants (PARM\text{PAR}_M, PARB\text{PAR}_B, PARH\text{PAR}_H) consistently outperform vanilla CNN baselines, DANN, InfoDrop, and HEX. In the dependent setting, PARM\text{PAR}_M achieves the highest accuracy when the test domain is "original" or "radial", whereas its performance decreases when tested on the "random" pattern, indicating that random kernel patterns are more easily identified and stripped during training by the multi-layer pattern classifier.

  8. Knowl 8 — Local Pattern Reliance Analysis on ImageNet-Sketch Misclassifications

    empirical result

    An analysis of conflicting predictions on ImageNet-Sketch between standard AlexNet and AlexNet trained with Patch-wise Adversarial Regularization (AlexNet-PAR) reveals distinct error mechanisms:

    • Standard AlexNet exhibits a strong inductive bias toward matching local visual primitives without contextual integration across the image. For example, AlexNet classifies a sketch of a stethoscope as a hook (confidence 0.39) due to the local tube curvature, a tricycle as a safety pin (confidence 0.51) due to the local frame structure, an Afghan hound as a swab (mop) (confidence 0.74) due to hair texture, and red wine as a goblet (confidence 0.74) due to the local presence of a glass.
    • AlexNet-PAR correctly predicts these instances (stethoscope with 0.66 confidence, tricycle with 0.92, Afghan hound with 0.89, red wine with 0.59) by forcing the network to integrate global compositional features across spatial locations.
    • Conversely, when local components are ambiguous or degraded, PAR can misclassify objects with deceptive global geometries (such as predicting a keyboard with missing keys as a crossword puzzle).
  9. Knowl 9 — Regularization Strength and Layer Depth Sensitivity in PAR

    limitation

    The effectiveness of Patch-wise Adversarial Regularization depends critically on the regularization weight λ\lambda and the layer level chosen for penalization:

    • Regularization Magnitude (λ\lambda): A value of λ=1.0\lambda = 1.0 can be overly aggressive on large-scale datasets such as ImageNet-Sketch, causing training loss instability and deteriorating early-epoch performance unless accompanied by a reduced learning rate. Empirical sweeps indicate that smaller regularization weights (e.g., λ∈[0.01,0.2]\lambda \in [0.01, 0.2]) yield more stable optimization.
    • Layer Selection: Applying PAR to higher convolutional layers or across multiple layers with geometric decay weights (1,λ,λ2,λ31, \lambda, \lambda^2, \lambda^3) can over-penalize semantic representations needed for classification, causing unstable training dynamics without improving out-of-distribution performance over first-layer penalization.
    • In-Domain Accuracy Trade-off: When local features (such as texture) are genuine, non-spurious predictors of category membership in the target deployment distribution, penalizing local representations may offer no gain or slightly reduce in-domain classification accuracy.

Coverage note — No substantial contributed material was omitted; minor qualitative examples in Appendix C and specific Fourier filter formulas were condensed into the experimental knowls.

References

  1. 1.A. Achille and S. Soatto. Information dropout: Learning optimal representations through noisy computation. IEEE transactions on pattern analysis and machine intelligence, 40(12):2897–2905, 2018.
  2. 2.Y. Balaji, S. Sankaranarayanan, and R. Chellappa. Metareg: Towards domain generalization using meta-regularization. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 998–1008. Curran Associates, Inc., 2018.
  3. 3.S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan. A theory of learning from different domains. Machine learning, 79(1):151–175, 2010a.
  4. 4.S. Ben-David, T. Lu, T. Luu, and D. Pál. Impossibility theorems for domain adaptation. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2010b.
  5. 5.K. Bousmalis, G. Trigeorgis, N. Silberman, D. Krishnan, and D. Erhan. Domain separation networks. In Advances in Neural Information Processing Systems, pages 343–351, 2016.
  6. 6.K. Bousmalis, N. Silberman, D. Dohan, D. Erhan, and D. Krishnan. Unsupervised pixel-level domain adaptation with generative adversarial networks. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul 2017. doi: 10.1109/cvpr.2017.18.
  7. 7.J. S. Bridle and S. J. Cox. Recnorm: Simultaneous normalisation and classification applied to speech recognition. In Advances in Neural Information Processing Systems, pages 234–240, 1991.
  8. 8.F. M. Carlucci, P. Russo, T. Tommasi, and B. Caputo. Agnostic domain generalization. arXiv preprint arXiv:1808.01102, 2018.
  9. 9.F. M. Carlucci, A. D’Innocente, S. Bucci, B. Caputo, and T. Tommasi. Domain generalization by solving jigsaw puzzles. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  10. 10.G. Csurka. Domain adaptation for visual applications: A comprehensive survey. arXiv preprint arXiv:1702.05374, 2017.
  11. 11.J. Deng, W. Dong, R. Socher, L. jia Li, K. Li, and L. Fei-fei. Imagenet: A large-scale hierarchical image database. In In CVPR, 2009.
  12. 12.Z. Ding and Y. Fu. Deep domain generalization with structured low-rank constraint. IEEE Transactions on Image Processing, 27(1):304–313, 2018.
  13. 13.S. Erfani, M. Baktashmotlagh, M. Moshtaghi, V. Nguyen, C. Leckie, J. Bailey, and R. Kotagiri. Robust domain generalisation by enforcing distribution invariance. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, pages 1455–1461. AAAI Press/International Joint Conferences on Artificial Intelligence, 2016.
  14. 14.Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky. Domain-adversarial training of neural networks. The Journal of Machine Learning Research, 17(1):2096–2030, 2016.
  15. 15.T. Gebru, J. Hoffman, and L. Fei-Fei. Fine-grained recognition in the wild: A multi-task domain adaptation approach. 2017 IEEE International Conference on Computer Vision (ICCV), Oct 2017. doi: 10.1109/iccv.2017.151.
  16. 16.R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019.
  17. 17.M. Ghifary, W. Bastiaan Kleijn, M. Zhang, and D. Balduzzi. Domain generalization for object recognition with multi-task autoencoders. In Proceedings of the IEEE international conference on computer vision, pages 2551–2559, 2015.
  18. 18.A. Gretton, A. J. Smola, J. Huang, M. Schmittfull, K. M. Borgwardt, and B. Schölkopf. Covariate shift by kernel mean matching. Journal of Machine Learning Research, 2009.
  19. 19.J. J. Heckman. Sample selection bias as a specification error (with an application to the estimation of labor supply functions), 1977.
  20. 20.D. Hendrycks and T. Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019.
  21. 21.J. Hoffman, E. Tzeng, T. Darrell, and K. Saenko. Simultaneous deep transfer across domains and tasks. Advances in Computer Vision and Pattern Recognition, page 173–187, 2017. ISSN 2191-6594. doi: 10.1007/978-3-319-58347-1_9.
  22. 22.J. Hoffman, E. Tzeng, T. Park, J.-Y. Zhu, P. Isola, K. Saenko, A. Efros, and T. Darrell. CyCADA: Cycle-consistent adversarial domain adaptation. In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 1989–1998, Stockholmsmässan, Stockholm Sweden, 10–15 Jul 2018. PMLR.
  23. 23.W. Hu, G. Nio, I. Sato, and M. Sugiyama. Does distributionally robust supervised learning give robust classifiers? arXiv preprint arXiv:1611.02041, 2016.
  24. 24.J. Jo and Y. Bengio. Measuring the tendency of cnns to learn surface statistical regularities. arXiv preprint arXiv:1711.11561, 2017.
  25. 25.F. D. Johansson, R. Ranganath, and D. Sontag. Support and invertibility in domain-invariant representations. arXiv preprint arXiv:1903.03448, 2019.
  26. 26.A. Kumagai and T. Iwata. Zero-shot domain adaptation without domain semantic descriptors. arXiv preprint arXiv:1807.02927, 2018.
  27. 27.A. Kumar, P. Sattigeri, K. Wadhawan, L. Karlinsky, R. Feris, B. Freeman, and G. Wornell. Co-regularized alignment for unsupervised domain adaptation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9345–9356. Curran Associates, Inc., 2018.
  28. 28.K. Lee, K. Lee, H. Lee, and J. Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 1019–1030. Curran Associates, Inc., 2018.
  29. 29.D. Li, Y. Yang, Y.-Z. Song, and T. M. Hospedales. Deeper, broader and artier domain generalization. In Computer Vision (ICCV), 2017 IEEE International Conference on, pages 5543–5551. IEEE, 2017a.
  30. 30.D. Li, Y. Yang, Y.-Z. Song, and T. M. Hospedales. Learning to generalize: Meta-learning for domain generalization. arXiv preprint arXiv:1710.03463, 2017b.
  31. 31.H. Li, S. J. Pan, S. Wang, and A. C. Kot. Domain generalization with adversarial feature learning. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(CVPR), 2018a.
  32. 32.W. Li, Z. Xu, D. Xu, D. Dai, and L. Van Gool. Domain generalization and adaptation using low rank exemplar svms. IEEE transactions on pattern analysis and machine intelligence, 2017c.
  33. 33.Y. Li, m. Murias, g. Dawson, and D. E. Carlson. Extracting relationships by multi-domain matching. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 6798–6809. Curran Associates, Inc., 2018b.
  34. 34.Z. C. Lipton, Y.-X. Wang, and A. Smola. Detecting and correcting for label shift with black box predictors. In International Conference on Machine Learning (ICML), 2018.
  35. 35.M. Long, H. Zhu, J. Wang, and M. I. Jordan. Unsupervised domain adaptation with residual transfer networks. In Advances in Neural Information Processing Systems, pages 136–144, 2016.
  36. 36.M. Long, Z. CAO, J. Wang, and M. I. Jordan. Conditional adversarial domain adaptation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 1640–1650. Curran Associates, Inc., 2018.
  37. 37.M. Mancini, S. R. Bulò, B. Caputo, and E. Ricci. Best sources forward: domain generalization through source-specific nets. arXiv preprint arXiv:1806.05810, 2018.
  38. 38.C. F. Manski and S. R. Lerman. The estimation of choice probabilities from choice based samples. Econometrica: Journal of the Econometric Society, 1977.
  39. 39.Y. Mansour, M. Mohri, and A. Rostamizadeh. Domain adaptation: Learning bounds and algorithms. arXiv preprint arXiv:0902.3430, 2009.
  40. 40.S. Motiian, M. Piccirilli, D. A. Adjeroh, and G. Doretto. Unified deep supervised domain adaptation and generalization. 2017 IEEE International Conference on Computer Vision (ICCV), Oct 2017a. doi: 10.1109/iccv.2017.609.
  41. 41.S. Motiian, M. Piccirilli, D. A. Adjeroh, and G. Doretto. Unified deep supervised domain adaptation and generalization. In The IEEE International Conference on Computer Vision (ICCV), volume 2, page 3, 2017b.
  42. 42.K. Muandet, D. Balduzzi, and B. Schölkopf. Domain generalization via invariant feature representation. In International Conference on Machine Learning, pages 10–18, 2013.
  43. 43.L. Niu, W. Li, and D. Xu. Multi-view domain generalization for visual recognition. In Proceedings of the IEEE International Conference on Computer Vision, pages 4193–4201, 2015.
  44. 44.QuickDraw. Quick draw! the data, 2018. URL https://quickdraw.withgoogle.com/data.
  45. 45.B. Recht, R. Roelofs, L. Schmidt, and V. Shankar. Do imagenet classifiers generalize to imagenet?, 2019.
  46. 46.P. Sangkloy, N. Burnell, C. Ham, and J. Hays. The sketchy database: Learning to retrieve badly drawn bunnies. ACM Transactions on Graphics (proceedings of SIGGRAPH), 2016.
  47. 47.A. Schoenauer-Sebag, L. Heinrich, M. Schoenauer, M. Sebag, L. Wu, and S. Altschuler. Multi-domain adversarial learning. In International Conference on Learning Representations, 2019.
  48. 48.B. Schölkopf, D. Janzing, J. Peters, E. Sgouritsa, K. Zhang, and J. Mooij. On causal and anticausal learning. In International Coference on International Conference on Machine Learning (ICML-12), pages 459–466. Omnipress, 2012.
  49. 49.H. Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of statistical planning and inference, 2000.
  50. 50.A. Storkey. When training and test sets are different: characterizing learning transfer. Dataset shift in machine learning, 2009.
  51. 51.E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell. Adversarial discriminative domain adaptation. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul 2017. doi: 10.1109/cvpr.2017.316.
  52. 52.R. Volpi, H. Namkoong, O. Sener, J. C. Duchi, V. Murino, and S. Savarese. Generalizing to unseen domains via adversarial data augmentation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, pages 5334–5344. Curran Associates, Inc., 2018.
  53. 53.H. Wang, A. Meghawat, L.-P. Morency, and E. P. Xing. Select-additive learning: Improving generalization in multimodal sentiment analysis. arXiv preprint arXiv:1609.05244, 2016.
  54. 54.H. Wang, Z. He, Z. C. Lipton, and E. P. Xing. Learning robust representations by projecting superficial statistics out. In International Conference on Learning Representations, 2019.
  55. 55.M. Wang and W. Deng. Deep visual domain adaptation: A survey. Neurocomputing, 312:135–153, Oct 2018. ISSN 0925-2312. doi: 10.1016/j.neucom.2018.05.083.
  56. 56.K. Weiss, T. M. Khoshgoftaar, and D. Wang. A survey of transfer learning. Journal of Big Data, 3 (1):9, 2016.
  57. 57.Y. Wu, E. Winston, D. Kaushik, and Z. Lipton. Domain adaptation with asymmetrically-relaxed distribution alignment. International Conference on Machine Learning, 2019.
  58. 58.S. Xie, Z. Zheng, L. Chen, and C. Chen. Learning semantic representations for unsupervised domain adaptation. In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 5423–5432, Stockholmsmässan, Stockholm Sweden, 10–15 Jul 2018. PMLR.
  59. 59.K. Zhang, B. Schölkopf, K. Muandet, and Z. Wang. Domain adaptation under target and conditional shift. In International Conference on Machine Learning, 2013.
  60. 60.A. Zhao, M. Ding, J. Guan, Z. Lu, T. Xiang, and J.-R. Wen. Domain-invariant projection learning for zero-shot recognition. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 1019–1030. Curran Associates, Inc., 2018a.
  61. 61.H. Zhao, S. Zhang, G. Wu, J. M. F. Moura, J. P. Costeira, and G. J. Gordon. Adversarial multiple source domain adaptation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 8559–8570. Curran Associates, Inc., 2018b.

Citation

MLA
Wang, H., et al. “Learning Robust Global Representations by Penalizing Local Predictive Power”. NeurIPS 2019, 2019, http://arxiv.org/abs/1905.13549v2.
APA
Wang, H., Ge, S., Xing, E. P., & Lipton, Z. C. (2019). Learning Robust Global Representations by Penalizing Local Predictive Power. NeurIPS 2019. http://arxiv.org/abs/1905.13549v2
Chicago
Wang, H., S. Ge, E. P. Xing, and Z. C. Lipton. 2019. “Learning Robust Global Representations by Penalizing Local Predictive Power”. NeurIPS 2019. http://arxiv.org/abs/1905.13549v2.
Harvard
Wang, H. et al. (2019) “Learning Robust Global Representations by Penalizing Local Predictive Power”, NeurIPS 2019 [Preprint]. Available at: http://arxiv.org/abs/1905.13549v2.
Vancouver
1. Wang H, Ge S, Xing EP, Lipton ZC (2019) Learning Robust Global Representations by Penalizing Local Predictive Power. NeurIPS 2019

BibTeX

@article{wang2019learning,
  title = {Learning Robust Global Representations by Penalizing Local Predictive Power},
  author = {Wang, Haohan and Ge, Songwei and Xing, Eric P. and Lipton, Zachary C.},
  year = {2019},
  journal = {NeurIPS 2019},
  url = {http://arxiv.org/abs/1905.13549v2},
  eprint = {1905.13549}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors