WILDS: A Benchmark of in-the-Wild Distribution Shifts

Pang Wei KohShiori SagawaHenrik MarklundSang Michael XieMarvin ZhangAkshay BalsubramaniWeihua HuMichihiro YasunagaRichard Lanas PhillipsIrena Gao

article2020ICML2,007 citations

Introduces WILDS, a benchmark of ten real-world datasets across diverse applications, demonstrating that standard algorithms fail under naturally occurring distribution shifts and providing standardized evaluations to develop models with better out-of-distribution generalization.

arXiv: 2012.07421
  • Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Introduces Invariant Risk Minimization (IRM), a foundational multi-environment distribution shift method whose real-world efficacy WILDS explicitly evaluates and challenges.
  • Paper: Fairness Without Demographics in Repeated Loss Minimization, Tatsunori B. Hashimoto et al. (2018). Establishes distributionally robust optimization (DRO) principles for mitigating subpopulation shifts that form the baseline methodologies analyzed across WILDS datasets.
  • Paper: Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, Dan Hendrycks et al. (2019). Pioneers standardized evaluation protocols and synthetic corruptions for machine learning robustness benchmarks, setting the stage for WILDS to focus on naturally occurring in-the-wild shifts.
  • Paper: Unbiased look at dataset bias, A. Torralba et al. (2011). Provides the classic foundational analysis of cross-dataset generalization failures and dataset bias that motivates benchmarking real-world distribution shifts.
  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). Develops the core theoretical learning bounds for generalization across distinct domains that underpin empirical out-of-distribution evaluation suites.
  • Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). Introduces domain-adversarial neural networks (DANN), one of the standard baseline domain adaptation algorithms evaluated on the WILDS benchmark.
  • Paper: Deep CORAL: Correlation Alignment for Deep Domain Adaptation, Baochen Sun et al. (2016). Proposes correlation alignment (CORAL) for deep domain adaptation, providing an essential baseline algorithm benchmarked throughout the WILDS package.
  • Paper: Do ImageNet Classifiers Generalize to ImageNet?, Benjamin Recht et al. (2019). Demonstrates empirical generalization drops on replicated test sets, motivating the need for comprehensive benchmarks that test models beyond standard i.i.d. splits.
  • Paper: Natural Adversarial Examples, Dan Hendrycks et al. (2019). Highlights the vulnerability of neural networks to natural out-of-distribution examples, serving as a direct precursor to curated real-world shift benchmarks.
  • Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). Presents the Open Graph Benchmark methodology of application-driven, non-random splits and unified toolkits that strongly influenced WILDS' design philosophy.
Cover for WILDS: A Benchmark of in-the-Wild Distribution Shifts

Abstract

Distribution shifts -- where the training distribution differs from the test distribution -- can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets widely used in the ML community today. To address this gap, we present WILDS, a curated benchmark of 10 datasets reflecting a diverse range of distribution shifts that naturally arise in real-world applications, such as shifts across hospitals for tumor identification; across camera traps for wildlife monitoring; and across time and location in satellite imaging and poverty mapping. On each dataset, we show that standard training yields substantially lower out-of-distribution than in-distribution performance. This gap remains even with models trained by existing methods for tackling distribution shifts, underscoring the need for new methods for training models that are more robust to the types of distribution shifts that arise in practice. To facilitate method development, we provide an open-source package that automates dataset loading, contains default model architectures and hyperparameters, and standardizes evaluations. Code and leaderboards are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Existing ML benchmarks for distribution shifts
  • 3 Problem settings
  • 3.1 Domain generalization (Figure -Top)
  • 3.2 Subpopulation shift (Figure -Bottom)
  • 3.3 Hybrid settings
  • 4 Wilds datasets
  • 4.1 Domain generalization datasets
  • 4.1.1 iWildCam2020-wilds: Species classification across different camera traps
  • 4.1.2 Camelyon17-wilds: Tumor identification across different hospitals
  • 4.1.3 RxRx1-wilds: Genetic perturbation classification across experimental batches
  • 4.1.4 OGB-MolPCBA: Molecular property prediction across different scaffolds
  • 4.1.5 GlobalWheat-wilds: Wheat head detection across regions of the world
  • 4.2 Subpopulation shift datasets
  • 4.2.1 CivilComments-wilds: Toxicity classification across demographic identities
  • 4.3 Hybrid datasets
  • 4.3.1 FMoW-wilds: Land use classification across different regions and years
  • 4.3.2 PovertyMap-wilds: Poverty mapping across different countries
  • 4.3.3 Amazon-wilds: Sentiment classification across different users
  • 4.3.4 Py150-wilds: Code completion across different codebases
  • 5 Performance drops from distribution shifts
  • 5.1 In-distribution performance should be measured on P𝗍𝖾𝗌𝗍P^{\mathsf{test}}, not P𝗍𝗋𝖺𝗂𝗇P^{\mathsf{train}}
  • 5.2 Types of in-distribution settings
  • 5.3 Model selection
  • 5.4 Results
  • 6 Baseline algorithms for distribution shifts
  • 6.1 Domain generalization baselines
  • 6.2 Subpopulation shift baselines
  • 6.3 Setup
  • 6.4 Results
  • 7 Empirical trends
  • 7.1 Underspecification
  • 7.2 Model selection with in-distribution versus out-of-distribution validation sets
  • 7.3 The compounding effects of multiple distribution shifts
  • 8 Distribution shifts in other application areas
  • 8.1 Algorithmic fairness
  • 8.2 Medicine and healthcare
  • 8.3 Genomics
  • 8.4 Natural language and speech processing
  • 8.5 Education
  • 8.6 Robotics
  • 8.7 Feedback loops
  • 9 Guidelines for method developers
  • 9.1 General-purpose and specialized training algorithms
  • 9.2 Methods beyond training algorithms
  • 9.3 Avoiding overfitting to the test distribution
  • 9.4 Reporting both ID and OOD performance
  • 9.5 Extensions to other problem settings
  • 10 Using the Wilds package
  • References
  • A Dataset realism
  • B Prior work on ML benchmarks for distribution shifts
  • C Potential extensions to other problem settings
  • C.1 Problem settings in domain shifts
  • C.2 Unsupervised domain adaptation
  • C.3 Test-time adaptation
  • C.4 Selective prediction
  • D Additional experimental details
  • D.1 Model hyperparameters
  • D.2 Replicates
  • D.3 Baseline algorithms
  • E Additional dataset details and results
  • E.1 iWildCam2020-wilds
  • E.1.1 Setup
  • E.1.2 Baseline results
  • E.1.3 Broader context
  • E.1.4 Additional details
  • E.2 Camelyon17-wilds
  • E.2.1 Setup
  • E.2.2 Baseline results
  • E.2.3 Broader context
  • E.2.4 Additional details
  • E.3 RxRx1-wilds
  • E.3.1 Setup
  • E.3.2 Baseline results
  • E.3.3 Broader context
  • E.3.4 Additional details
  • E.4 OGB-MolPCBA
  • E.4.1 Setup
  • E.4.2 Baseline results
  • E.4.3 Broader context
  • E.4.4 Additional details
  • E.5 GlobalWheat-wilds
  • E.5.1 Setup
  • E.5.2 Baseline results
  • E.5.3 Broader context
  • E.5.4 Additional details
  • E.6 CivilComments-wilds
  • E.6.1 Setup
  • E.6.2 Baseline results
  • E.6.3 Broader context
  • E.6.4 Additional details
  • E.7 FMoW-wilds
  • E.7.1 Setup
  • E.7.2 Baseline results
  • E.7.3 Broader context
  • E.7.4 Additional details
  • E.8 PovertyMap-wilds
  • E.8.1 Setup
  • E.8.2 Baseline results
  • E.8.3 Broader context
  • E.8.4 Additional details
  • E.9 Amazon-wilds
  • E.9.1 Setup
  • E.9.2 Baseline results
  • E.9.3 Broader context
  • E.9.4 Additional details
  • E.10 Py150-wilds
  • E.10.1 Setup
  • E.10.2 Baseline results
  • E.10.3 Broader context
  • E.10.4 Additional details
  • F Datasets with distribution shifts that do not cause performance drops
  • F.1 SQF: Criminal possession of weapons across race and locations
  • F.1.1 Setup
  • F.1.2 Baseline results
  • F.1.3 Additional details
  • F.2 ENCODE: Transcription factor binding across different cell types
  • F.2.1 Setup
  • F.2.2 Baseline results
  • F.2.3 Additional details
  • F.3 BDD100K: Object recognition in autonomous driving across locations
  • F.3.1 Setup
  • F.3.2 Time of day shift
  • F.3.3 Location shift
  • F.4 Amazon: Sentiment classification across different categories and time
  • F.4.1 Setup
  • F.4.2 Time shifts
  • F.4.3 Category shifts
  • F.5 Yelp: Sentiment classification across different users and time
  • F.5.1 Setup
  • F.5.2 Time shifts
  • F.5.3 User shift

Knowls

  1. Knowl 1 — Mathematical Formulation of Domain Generalization, Subpopulation Shift, and Hybrid Distribution Shifts

    definition

    In the WILDS distribution shift framework, a dataset represents an underlying mixture over a discrete set of DD domains D={1,,D}\mathcal{D} = \{1, \dots, D\}. Each domain dDd \in \mathcal{D} corresponds to a fixed data distribution PdP_d over triplets (x,y,d)(x, y, d), where xXx \in \mathcal{X} is the input, yYy \in \mathcal{Y} is the prediction target, and every point sampled from PdP_d is associated with domain label dd.

    The training and test distributions are defined as domain mixtures with domain-specific categorical weights:

    Ptrain=dDqdtrainPd,Ptest=dDqdtestPdP^{\text{train}} = \sum_{d \in \mathcal{D}} q^{\text{train}}_d P_d, \quad P^{\text{test}} = \sum_{d \in \mathcal{D}} q^{\text{test}}_d P_d

    where qdtrain0q^{\text{train}}_d \ge 0 and qdtest0q^{\text{test}}_d \ge 0 satisfy dDqdtrain=1\sum_{d \in \mathcal{D}} q^{\text{train}}_d = 1 and dDqdtest=1\sum_{d \in \mathcal{D}} q^{\text{test}}_d = 1. The sets of active training and test domains are denoted by Dtrain={dDqdtrain>0}\mathcal{D}^{\text{train}} = \{d \in \mathcal{D} \mid q^{\text{train}}_d > 0\} and Dtest={dDqdtest>0}\mathcal{D}^{\text{test}} = \{d \in \mathcal{D} \mid q^{\text{test}}_d > 0\}, respectively.

    Under this formulation, distribution shifts are categorized into three distinct regimes:

    1. Domain Generalization: The training and test domain sets are mutually disjoint, DtrainDtest=\mathcal{D}^{\text{train}} \cap \mathcal{D}^{\text{test}} = \emptyset. The model is evaluated on unseen domains drawn from PtestP^{\text{test}}, aiming to minimize the expected test loss E(x,y)Ptest[(f(x),y)]\mathbb{E}_{(x, y) \sim P^{\text{test}}}[\ell(f(x), y)].
    2. Subpopulation Shift: The test domains are a subset of the training domains, DtestDtrain\mathcal{D}^{\text{test}} \subseteq \mathcal{D}^{\text{train}}, but their relative mixture proportions differ (qtestqtrainq^{\text{test}} \neq q^{\text{train}}). The objective is typically to minimize the worst-case error across all test subpopulations: maxdDtestE(x,y)Pd[(f(x),y)]\max_{d \in \mathcal{D}^{\text{test}}} \mathbb{E}_{(x, y) \sim P_d}[\ell(f(x), y)].
    3. Hybrid Shift: The task simultaneously exhibits domain generalization along one domain attribute (e.g., disjoint time periods or geographical locations) and subpopulation shift along another attribute (e.g., evaluating worst-region accuracy over regions present in both train and test).
  2. Knowl 2 — The WILDS Benchmark Dataset Suite

    experimental setup

    The WILDS benchmark comprises 10 curated datasets spanning diverse modalities, shift types, and domain structures:

    • IWILDCAM2020-WILDS: Domain generalization task classifying animal species (y{1,,182}y \in \{1, \dots, 182\}) from camera trap photos (xx). The domain dd is the camera trap identifier (323 total domains across 203,029 examples). Evaluated by Macro F1 score.
    • CAMELYON17-WILDS: Domain generalization task identifying tumor tissue (y{0,1}y \in \{0, 1\}) from 96×9696 \times 96 lymph node whole-slide histopathology patches (xx). The domain dd is the source hospital (5 hospitals across 455,954 examples). Evaluated by average classification accuracy.
    • RXRX1-WILDS: Domain generalization task predicting genetic siRNA treatments (y{1,,1139}y \in \{1, \dots, 1139\}) from 3-channel fluorescent microscopy cell images (xx). The domain dd is the experimental batch (51 batches across 125,510 examples). Evaluated by average classification accuracy.
    • OGB-MOLPCBA: Domain generalization multi-label classification task predicting 128 bioassay outcomes (y{0,1}128y \in \{0, 1\}^{128}) from molecular graphs (xx). The domain dd is the molecular scaffold cluster (120,084 scaffolds across 437,929 examples). Evaluated by mean Average Precision (AP) across assays.
    • GLOBALWHEAT-WILDS: Domain generalization object detection task predicting bounding boxes for wheat heads (yy) from overhead crop field images (xx). The domain dd is the acquisition session defined by location, date, and sensor (47 sessions across 6,515 images). Evaluated by average domain accuracy at a fixed Intersection-over-Union threshold.
    • CIVILCOMMENTS-WILDS: Subpopulation shift binary classification task predicting online comment toxicity (y{0,1}y \in \{0, 1\}) from raw text (xx). The domain dd is an 8-dimensional binary indicator covering demographic identity mentions (16 groups across 448,000 examples). Evaluated by worst-group accuracy over all demographic identities.
    • FMOW-WILDS: Hybrid shift task predicting 62 land use and building categories (yy) from satellite images (xx). The domain dd represents acquisition year ×\times geographical region (16 years ×\times 5 regions across 523,846 images). Evaluated by worst-region accuracy on post-2016 test years.
    • POVERTYMAP-WILDS: Hybrid shift task predicting continuous asset wealth index (yRy \in \mathbb{R}) from multispectral satellite images (xx). The domain dd represents country ×\times urban/rural status (23 countries ×\times 2 settlement types across 19,669 images). Evaluated by worst Pearson correlation (RR) across urban and rural subpopulations.
    • AMAZON-WILDS: Hybrid shift task predicting 1-to-5 star ratings (y{1,,5}y \in \{1, \dots, 5\}) from user product reviews (xx). The domain dd is the user reviewer ID (2,586 users across 539,502 reviews). Evaluated by the 10th percentile of per-user classification accuracies.
    • PY150-WILDS: Hybrid shift language modeling task predicting the next code token (yy) given prior token context (xx). The domain dd is the source GitHub repository (8,421 repositories across 150,000 examples). Evaluated by accuracy on class and method name token subpopulations.
  3. Knowl 3 — Evaluation Protocols for In-Distribution Baseline Comparisons

    definition

    To isolate and quantify performance degradation driven specifically by distribution shift, WILDS establishes four standard In-Distribution (ID) comparison protocols to benchmark against Out-of-Distribution (OOD) test performance:

    1. Fixed-train: The training set remains fixed. Performance is evaluated on a distinct, held-out ID test set sampled from the identical domain mixture Dtrain\mathcal{D}^{\text{train}} as the training set. This is used when training and test domains are interchangeable random draws from the same distribution.
    2. Fixed-test: The OOD test set is held constant, but the training distribution is altered by mixing in labeled samples from the test distribution (e.g., swapping 50% of the training data with otherwise unused data from the target domain distribution) while matching the original training dataset size. This controls for whether lower OOD performance is caused by intrinsic difficulty of the test domains.
    3. Randomized: All available instances across all domains are randomly shuffled into standard i.i.d. train, validation, and test partitions. This is employed when the number of examples per domain is too small to construct separate held-out domain splits.
    4. Average: For subpopulation shift benchmarks where the OOD metric is worst-case subpopulation accuracy, the ID baseline is defined as the unweighted mean performance evaluated across the entire test set.
  4. Knowl 4 — In-Distribution vs. Out-of-Distribution Performance Drops under Empirical Risk Minimization

    data/table

    When standard deep learning models are trained via Empirical Risk Minimization (ERM), out-of-distribution (OOD) performance drops substantially across every WILDS dataset compared to in-distribution (ID) baselines:

    Dataset Metric ID Comparison Type In-distribution Out-of-distribution
    IWILDCAM2020-WILDS Macro F1 Fixed-train 47.0 (1.4) 31.0 (1.3)
    CAMELYON17-WILDS Average accuracy Fixed-train 93.2 (5.2) 70.3 (6.4)
    RXRX1-WILDS Average accuracy Fixed-test 39.8 (0.2) 29.9 (0.4)
    OGB-MOLPCBA Average AP Randomized 34.4 (0.9) 27.2 (0.3)
    GLOBALWHEAT-WILDS Average domain accuracy Fixed-test 64.8 (0.4) 48.4 (1.8)
    CIVILCOMMENTS-WILDS Worst-group accuracy Average 92.2 (0.1) 56.0 (3.6)
    FMOW-WILDS Worst-region accuracy Fixed-test 48.6 (0.9) 32.3 (1.3)
    POVERTYMAP-WILDS Worst-U/R Pearson R Fixed-test 0.60 (0.06) 0.45 (0.06)
    AMAZON-WILDS 10th percentile accuracy Average 71.9 (0.1) 53.8 (0.8)
    PY150-WILDS Method/class accuracy Fixed-train 75.4 (0.4) 67.9 (0.1)

    Numbers in parentheses denote standard deviations across 3 or more replicates. On datasets allowing Fixed-test comparisons, models trained on an equal-sized mixture of ID and OOD distributions achieve high performance simultaneously on both distributions, proving that the OOD performance degradation stems from domain shift rather than inherent sample difficulty.

  5. Knowl 5 — Empirical Comparison of Baseline Robustness Algorithms (CORAL, IRM, Group DRO) against ERM on WILDS

    data/table

    Existing algorithms designed for domain generalization, unsupervised domain adaptation, and subpopulation shifts fail to consistently outperform standard Empirical Risk Minimization (ERM) on real-world distribution shifts:

    Dataset Setting ERM CORAL IRM Group DRO
    IWILDCAM2020-WILDS Domain gen. 31.0 (1.3) 32.8 (0.1) 15.1 (4.9) 23.9 (2.1)
    CAMELYON17-WILDS Domain gen. 70.3 (6.4) 59.5 (7.7) 64.2 (8.1) 68.4 (7.3)
    RXRX1-WILDS Domain gen. 29.9 (0.4) 28.4 (0.3) 8.2 (1.1) 23.0 (0.3)
    OGB-MOLPCBA Domain gen. 27.2 (0.3) 17.9 (0.5) 15.6 (0.3) 22.4 (0.6)
    GLOBALWHEAT-WILDS Domain gen. 49.2 (1.5) 46.1 (1.6)
    CIVILCOMMENTS-WILDS Subpop. shift 56.0 (3.6) 65.6 (1.3) 66.3 (2.1) 70.0 (2.0)
    FMOW-WILDS Hybrid 32.3 (1.3) 31.7 (1.2) 30.0 (1.4) 30.8 (0.8)
    POVERTYMAP-WILDS Hybrid 0.45 (0.06) 0.44 (0.06) 0.43 (0.07) 0.39 (0.06)
    AMAZON-WILDS Hybrid 53.8 (0.8) 52.9 (0.8) 52.4 (0.8) 53.3 (0.0)
    PY150-WILDS Hybrid 67.9 (0.1) 65.9 (0.1) 64.3 (0.2) 65.9 (0.1)

    Parentheses report standard deviation across 3+ experimental replicates. Key empirical takeaways include:

    1. On 9 out of 10 datasets, CORAL, IRM, and Group DRO match or underperform standard ERM.
    2. The exception is CIVILCOMMENTS-WILDS, where Group DRO improves worst-group accuracy from 56.0% to 70.0% by upweighting minority demographic groups during training (though this remains far below the 92.2% average accuracy of ERM).
    3. IRM severely degrades performance on several domain generalization tasks (e.g., dropping to 8.2% on RXRX1 and 15.1% on IWILDCAM2020).
  6. Knowl 6 — Deep Correlation Alignment (CORAL) Objective for Domain Invariance

    model/method

    Deep Correlation Alignment (CORAL) encourages domain invariance by aligning the second-order statistics (covariance matrices) of neural network feature activations across different domains. Given training domains dDtraind \in \mathcal{D}^{\text{train}}, the total objective is:

    LCORAL(θ)=LERM(θ)+λi,jDtrain,i<jCiCjF2\mathcal{L}_{\text{CORAL}}(\theta) = \mathcal{L}_{\text{ERM}}(\theta) + \lambda \sum_{i, j \in \mathcal{D}^{\text{train}}, i < j} \| C_i - C_j \|_F^2

    where LERM(θ)\mathcal{L}_{\text{ERM}}(\theta) is the standard empirical classification loss, λ>0\lambda > 0 is a regularization hyperparameter, F\|\cdot\|_F denotes the Frobenius norm, and CiC_i is the sample covariance matrix of the last-layer feature representations computed over samples from domain ii in a batch:

    Ci=1Ni1k=1Ni(ϕ(xi,k)ϕˉi)(ϕ(xi,k)ϕˉi)TC_i = \frac{1}{N_i - 1} \sum_{k=1}^{N_i} (\phi(x_{i, k}) - \bar{\phi}_i)(\phi(x_{i, k}) - \bar{\phi}_i)^T

    where ϕ(xi,k)\phi(x_{i, k}) is the feature representation for the kk-th sample of domain ii, NiN_i is the number of samples from domain ii, and ϕˉi\bar{\phi}_i is the mean feature vector for domain ii.

  7. Knowl 7 — Invariant Risk Minimization (IRM) Objective in the WILDS Benchmark

    model/method

    Invariant Risk Minimization (IRM) aims to learn a feature representation Φ:XZ\Phi: \mathcal{X} \to \mathcal{Z} such that the optimal linear classifier ww operating on top of Φ\Phi is invariant across all training environments/domains dDtraind \in \mathcal{D}^{\text{train}}. In its first-order continuous surrogate form (IRMv1), the objective is formulated as:

    minΦdDtrainRd(wΦ)+λdDtrainww=1.0Rd(wΦ)2\min_{\Phi} \sum_{d \in \mathcal{D}^{\text{train}}} R_d(w \circ \Phi) + \lambda \sum_{d \in \mathcal{D}^{\text{train}}} \left\| \nabla_{w \mid w = 1.0} R_d(w \cdot \Phi) \right\|^2

    where Rd(f)=E(x,y)Pd[(f(x),y)]R_d(f) = \mathbb{E}_{(x, y) \sim P_d}[\ell(f(x), y)] represents the empirical risk on domain dd, ww is a fixed scalar dummy classifier initialized to 1.01.0, and λ>0\lambda > 0 is a penalty parameter balancing empirical risk minimization against the gradient penalty enforcing cross-domain classifier optimality.

  8. Knowl 8 — Group Distributionally Robust Optimization (Group DRO) Formulation

    model/method

    Group Distributionally Robust Optimization (Group DRO) protects against performance degradation on minority subpopulations by optimizing the empirical risk of the worst-performing group during training. The optimization problem is defined as:

    minθmaxdDtrainE(x,y)Pd[(fθ(x),y)]\min_{\theta} \max_{d \in \mathcal{D}^{\text{train}}} \mathbb{E}_{(x, y) \sim P_d} [\ell(f_\theta(x), y)]

    During training, Group DRO tracks the group-specific losses and dynamically updates a probability vector over groups via exponentiated gradient ascent, upweighting groups with higher instantaneous training loss.

    In WILDS, Group DRO is evaluated both in its intended subpopulation shift setting (e.g., optimizing worst-case demographic accuracy in CIVILCOMMENTS-WILDS) and in domain generalization settings (e.g., treating each distinct training hospital in CAMELYON17-WILDS or each camera trap in IWILDCAM2020-WILDS as an individual group to enforce uniform cross-domain loss).

  9. Knowl 9 — Curational Selection Criteria for Real-World Distribution Shift Benchmarks

    definition

    To ensure benchmarks reflect meaningful real-world challenges, WILDS adopts three formal selection criteria for including datasets:

    1. Distribution shifts with performance drops: The train/test splits must produce a substantial drop in out-of-distribution performance relative to in-distribution baselines when evaluated using standard ERM.
    2. Real-world relevance: Training and test partitions, as well as evaluation metrics, must directly model realistic deployment obstacles identified in collaboration with application-domain experts (e.g., testing on new hospitals, new camera locations, unseen molecular scaffolds, or across time periods).
    3. Potential leverage: Datasets must not present unconstrained arbitrary shifts; they must supply training data spanning multiple distinct domains Dtrain\mathcal{D}^{\text{train}} accompanied by explicit domain annotations dd and contextual metadata. This enables models to leverage multi-domain structure to learn invariant features or balance subpopulation risks.
  10. Knowl 10 — OOD Validation Guidelines to Prevent Test Distribution Overfitting

    limitation

    In real-world distribution shift datasets, practical collection constraints often limit the number of available test domains (e.g., CAMELYON17-WILDS contains test data from only a single unseen hospital). Because inter-domain variance can be large, model performance can exhibit high stochastic variability across individual test domains.

    To prevent developers from inadvertently overfitting to specific test domains, model selection, hyperparameter tuning (such as penalty weights for IRM and CORAL), and early stopping must be conducted exclusively on designated out-of-distribution (OOD) validation splits consisting of validation domains separate from both training and test domains. Final out-of-distribution test sets must be reserved solely for final evaluation.

Coverage note — None was omitted; all primary formalisms, benchmark dataset specifications, comparison protocols, tabular empirical results, baseline algorithms, and methodological guidelines from the paper's contribution are fully covered.

References

  1. 1.Abelson, B., Varshney, K. R., and Sun, J. Targeting direct cash transfers to the extremely poor. In International Conference on Knowledge Discovery and Data Mining (KDD), 2014.
  2. 2.Adragna, R., Creager, E., Madras, D., and Zemel, R. Fairness and robustness in invariant learning: A case study in toxicity classification. arXiv preprint arXiv:2011.06485, 2020.
  3. 3.Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A. Don’t just assume; look and answer: Overcoming priors for visual question answering. In Computer Vision and Pattern Recognition (CVPR), pp. 4971–4980, 2018.
  4. 4.Ahadi, A., Lister, R., Haapala, H., and Vihavainen, A. Exploring machine learning methods to automatically identify students in need of assistance. In Proceedings of the Eleventh Annual International Conference on International Computing Education Research, pp. 121–130, 2015.
  5. 5.Ahumada, J. A., Fegraus, E., Birch, T., Flores, N., Kays, R., O’Brien, T. G., Palmer, J., Schuttler, S., Zhao, J. Y., Jetz, W., Kinnaird, M., Kulkarni, S., Lyet, A., Thau, D., Duong, M., Oliver, R., and Dancer, A. Wildlife insights: A platform to maximize the potential of camera trap and other passive sensor wildlife data for the planet. Environmental Conservation, 47(1):1–6, 2020.
  6. 6.Aich, S., Josuttes, A., Ovsyannikov, I., Strueby, K., Ahmed, I., Duddu, H. S., Pozniak, C., Shirtliffe, S., and Stavness, I. Deepwheat: Estimating phenotypic traits from crop images with deep learning. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 323–332. IEEE, 2018.
  7. 7.AlBadawy, E., Saha, A., and Mazurowski, M. Deep learning for segmentation of brain tumors: Impact of cross-institutional training and testing. Med Phys., 45, 2018.
  8. 8.Alexandari, A., Kundaje, A., and Shrikumar, A. Maximum likelihood with bias-corrected calibration is hard-to-beat at label shift adaptation. In International Conference on Machine Learning (ICML), pp. 222–232, 2020.
  9. 9.Allamanis, M. and Brockschmidt, M. Smartpaste: Learning to adapt source code. arXiv preprint arXiv:1705.07867, 2017.
  10. 10.Allamanis, M., Barr, E. T., Bird, C., and Sutton, C. Suggesting accurate method and class names. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, pp. 38–49, 2015.
  11. 11.Amorim, E., Cançado, M., and Veloso, A. Automated essay scoring in the presence of biased ratings. In Association for Computational Linguistics (ACL), pp. 229–237, 2018.
  12. 12.Ando, D. M., McLean, C. Y., and Berndl, M. Improving phenotypic measurements in high-content imaging screens. BioRxiv, pp. 161422, 2017.
  13. 13.Ardila, R., Branson, M., Davis, K., Kohler, M., Meyer, J., Henretty, M., Morais, R., Saunders, L., Tyers, F., and Weber, G. Common voice: A massively-multilingual speech corpus. In Language Resources and Evaluation Conference (LREC), pp. 4218–4222, 2020.
  14. 14.Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019.
  15. 15.Asuncion, A. and Newman, D. UCI Machine Learning Repository, 2007.
  16. 16.Attene-Ramos, M. S., Miller, N., Huang, R., Michael, S., Itkin, M., Kavlock, R. J., Austin, C. P., Shinn, P., Simeonov, A., Tice, R. R., et al. The tox21 robotic platform for the assessment of environmental chemicals–from vision to reality. Drug Discovery Today, 18(15):716–723, 2013.
  17. 17.Atwood, J., Halpern, Y., Baljekar, P., Breck, E., Sculley, D., Ostyakov, P., Nikolenko, S. I., Ivanov, I., Solovyev, R., Wang, W., et al. The Inclusive Images competition. In Advances in Neural Information Processing Systems (NeurIPS), pp. 155–186, 2020.
  18. 18.Aviv, R., Teichmann, S. A., Lander, E. S., Ido, A., Christophe, B., Ewan, B., Bernd, B., Campbell, P., Piero, C., Menna, C., et al. The human cell atlas. eLife, 6, 2017.
  19. 19.Avsec, Ž., Weilert, M., Shrikumar, A., Alexandari, A., Krueger, S., Dalal, K., Fropf, R., McAnany, C., Gagneur, J., Kundaje, A., and Zeitlinger, J. Deep learning at base-resolution reveals motif syntax of the cis-regulatory code. bioRxiv, 2019.
  20. 20.Ayalew, T. W., Ubbens, J. R., and Stavness, I. Unsupervised domain adaptation for plant organ counting. In European Conference on Computer Vision, pp. 330–346. Springer, 2020.
  21. 21.Azizzadenesheli, K., Liu, A., Yang, F., and Anandkumar, A. Regularized learning for domain adaptation under label shifts. In International Conference on Learning Representations (ICLR), 2019.
  22. 22.Badgeley, M. A., Zech, J. R., Oakden-Rayner, L., Glicksberg, B. S., Liu, M., Gale, W., McConnell, M. V., Percha, B., Snyder, T. M., and Dudley, J. T. Deep learning predicts hip fracture using confounding patient and healthcare variables. npj Digital Medicine, 2, 2019.
  23. 23.Balaji, Y., Sankaranarayanan, S., and Chellappa, R. Metareg: Towards domain generalization using meta-regularization. In Advances in Neural Information Processing Systems (NeurIPS), pp. 998–1008, 2018.
  24. 24.Bandi, P., Geessink, O., Manson, Q., Dijk, M. V., Balkenhol, M., Hermsen, M., Bejnordi, B. E., Lee, B., Paeng, K., Zhong, A., et al. From detection of individual metastases to classification of lymph node status at the patient level: the CAMELYON17 challenge. IEEE Transactions on Medical Imaging, 38(2):550–560, 2018.
  25. 25.Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems (NeurIPS), pp. 9453–9463, 2019.
  26. 26.Bartlett, P. L. and Wegkamp, M. H. Classification with a reject option using a hinge loss. Journal of Machine Learning Research (JMLR), 9(0):1823–1840, 2008.
  27. 27.Baumann, T., Köhn, A., and Hennig, F. The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening. Language Resources and Evaluation, 53(2):303–329, 2019.
  28. 28.BBC. A-levels and GCSEs: How did the exam algorithm work? The British Broadcasting Corporation, 2020. URL https://www.bbc.com/news/explainers-53807730.
  29. 29.Beck, A. H., Sangoi, A. R., Leung, S., Marinelli, R. J., Nielsen, T. O., Vijver, M. J. V. D., West, R. B., Rijn, M. V. D., and Koller, D. Systematic analysis of breast cancer morphology uncovers stromal features associated with survival. Science, 3(108), 2011.
  30. 30.Becke, A. D. Perspective: Fifty years of density-functional theory in chemical physics. The Journal of Chemical Physics, 140(18):18A301, 2014.
  31. 31.Beede, E., Baylor, E., Hersch, F., Iurchenko, A., Wilcox, L., Ruamviboonsuk, P., and Vardoulakis, L. M. A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. In Conference on Human Factors in Computing Systems (CHI), pp. 1–12, 2020.
  32. 32.Beery, S., Horn, G. V., and Perona, P. Recognition in terra incognita. In European Conference on Computer Vision (ECCV), pp. 456–473, 2018.
  33. 33.Beery, S., Morris, D., and Yang, S. Efficient pipeline for camera trap image review. arXiv preprint arXiv:1907.06772, 2019.
  34. 34.Beery, S., Cole, E., and Gjoka, A. The iWildCam 2020 competition dataset. arXiv preprint arXiv:2004.10340, 2020a.
  35. 35.Beery, S., Wu, G., Rathod, V., Votel, R., and Huang, J. Context r-cnn: Long term temporal context for per-camera object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13075–13085, 2020b.
  36. 36.Bejnordi, B. E., Veta, M., Diest, P. J. V., Ginneken, B. V., Karssemeijer, N., Litjens, G., Laak, J. A. V. D., Hermsen, M., Manson, Q. F., Balkenhol, M., et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. JAMA, 318(22):2199–2210, 2017.
  37. 37.Bellamy, D., Celi, L., and Beam, A. L. Evaluating progress on machine learning for longitudinal electronic healthcare data. arXiv preprint arXiv:2010.01149, 2020.
  38. 38.Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. Autonomous navigation of stratospheric balloons using reinforcement learning. Nature, 588, 2020.
  39. 39.Ben-David, S., Blitzer, J., Crammer, K., and Pereira, F. Analysis of representations for domain adaptation. In Advances in Neural Information Processing Systems (NeurIPS), pp. 137–144, 2006.
  40. 40.Bender, E. M. and Friedman, B. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics (TACL), 6:587–604, 2018.
  41. 41.BenTaieb, A. and Hamarneh, G. Adversarial stain transfer for histopathology image analysis. IEEE Transactions on Medical Imaging, 37(3):792–802, 2017.
  42. 42.Berman, G., de la Rosa, S., and Accone, T. Ethical considerations when using geospatial technologies for evidence generation. Innocenti Discussion Paper, UNICEF Office of Research, 2018.
  43. 43.Beyene, A. A., Welemariam, T., Persson, M., and Lavesson, N. Improved concept drift handling in surgery prediction and other applications. Knowledge and Information Systems, 44(1):177–196, 2015.
  44. 44.Blanchard, G., Lee, G., and Scott, C. Generalizing from several related classification tasks to a new unlabeled sample. In Advances in Neural Information Processing Systems (NeurIPS), pp. 2178–2186, 2011.
  45. 45.Blitzer, J., Dredze, M., and Pereira, F. Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification. In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, pp. 440–447, 2007.
  46. 46.Blodgett, S. L. and O’Connor, B. Racial disparity in natural language processing: A case study of social media African-American English. arXiv preprint arXiv:1707.00061, 2017.
  47. 47.Blodgett, S. L., Green, L., and O’Connor, B. Demographic dialectal variation in social media: A case study of African-American English. In Empirical Methods in Natural Language Processing (EMNLP), pp. 1119–1130, 2016.
  48. 48.Blumenstock, J., Cadamuro, G., and On, R. Predicting poverty and wealth from mobile phone metadata. Science, 350, 2015.
  49. 49.Bohacek, R. S., McMartin, C., and Guida, W. C. The art and practice of structure-based drug design: a molecular modeling perspective. Medicinal Research Reviews, 16(1):3–50, 1996.
  50. 50.Borkan, D., Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L. Limitations of pinned AUC for measuring unintended bias. arXiv preprint arXiv:1903.02088, 2019a.
  51. 51.Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L. Nuanced metrics for measuring unintended bias with real data for text classification. In WWW, pp. 491–500, 2019b.
  52. 52.Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E. Counterfactual reasoning and learning systems: The example of computational advertising. Journal of Machine Learning Research (JMLR), 14:3207–3260, 2013.
  53. 53.Boutros, M., Heigwer, F., and Laufer, C. Microscopy-based high-content screening. Cell, 163(6):1314–1325, 2015.
  54. 54.Bray, M.-A., Singh, S., Han, H., Davis, C. T., Borgeson, B., Hartland, C., Kost-Alimova, M., Gustafsdottir, S. M., Gibson, C. C., and Carpenter, A. E. Cell painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes. Nature protocols, 11(9):1757, 2016.
  55. 55.Broach, J. R., Thorner, J., et al. High-throughput screening for drug discovery. Nature, 384(6604):14–16, 1996.
  56. 56.Broussard, M. When algorithms give real students imaginary grades. The New York Times, 2020. URL https://www.nytimes.com/2020/09/08/opinion/international-baccalaureate-algorithm-grades.html.
  57. 57.Bruch, M., Monperrus, M., and Mezini, M. Learning from examples to improve code completion systems. In European software engineering conference and the ACM SIGSOFT symposium on the foundations of software engineering, 2009.
  58. 58.Bruzzone, L. and Marconcini, M. Domain adaptation problems: A DASVM classification technique and a circular validation strategy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(5):770–787, 2009.
  59. 59.Bug, D., Schneider, S., Grote, A., Oswald, E., Feuerhake, F., Schüler, J., and Merhof, D. Context-based normalization of histological stains using deep convolutional features. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, pp. 135–142, 2017.
  60. 60.Bunel, R., Hausknecht, M., Devlin, J., Singh, R., and Kohli, P. Leveraging grammar and reinforcement learning for neural program synthesis. In International Conference on Learning Representations (ICLR), 2018.
  61. 61.Buolamwini, J. and Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on Fairness, Accountability and Transparency, pp. 77–91, 2018.
  62. 62.Burke, M., Heft-Neal, S., and Bendavid, E. Sources of variation in under-5 mortality across sub-Saharan Africa: a spatial analysis. Lancet Global Health, 4, 2016.
  63. 63.Byrd, J. and Lipton, Z. What is the effect of importance weighting in deep learning? In International Conference on Machine Learning (ICML), pp. 872–881, 2019.
  64. 64.Caicedo, J. C., Cooper, S., Heigwer, F., Warchal, S., Qiu, P., Molnar, C., Vasilevich, A. S., Barry, J. D., Bansal, H. S., Kraus, O., et al. Data-analysis strategies for image-based cell profiling. Nature methods, 14(9):849–863, 2017.
  65. 65.Caicedo, J. C., McQuin, C., Goodman, A., Singh, S., and Carpenter, A. E. Weakly supervised learning of single-cell feature embeddings. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 9309–9318, 2018.
  66. 66.Caldas, S., Wu, P., Li, T., Konecnˇ y, J., McMahan, H. B., ` Smith, V., and Talwalkar, A. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018.
  67. 67.Campanella, G., Hanna, M. G., Geneslaw, L., Miraflor, A., Silva, V. W. K., Busam, K. J., Brogi, E., Reuter, V. E., Klimstra, D. S., and Fuchs, T. J. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature Medicine, 25(8): 1301–1309, 2019.
  68. 68.Cao, K., Wei, C., Gaidon, A., Arechiga, N., and Ma, T. Learning imbalanced datasets with label-distribution-aware margin loss. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  69. 69.Cao, K., Chen, Y., Lu, J., Arechiga, N., Gaidon, A., and Ma, T. Heteroskedastic and imbalanced deep learning with adaptive regularization. arXiv preprint arXiv:2006.15766, 2020.
  70. 70.Carlucci, F. M., D’Innocente, A., Bucci, S., Caputo, B., and Tommasi, T. Domain generalization by solving jigsaw puzzles. In Computer Vision and Pattern Recognition (CVPR), pp. 2229–2238, 2019.
  71. 71.Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. The Open Catalyst 2020 (oc20) dataset and community challenges. arXiv preprint arXiv:2010.09990, 2020.
  72. 72.Chen, I., Johansson, F. D., and Sontag, D. Why is my classifier discriminatory? In Advances in Neural Information Processing Systems (NeurIPS), pp. 3539–3550, 2018.
  73. 73.Chen, I. Y., Szolovits, P., and Ghassemi, M. Can AI help reduce disparities in general medical and mental health care? AMA Journal of Ethics, 21(2):167–179, 2019a.
  74. 74.Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., and Ghassemi, M. Ethical machine learning in health care. arXiv preprint arXiv:2009.10576, 2020.
  75. 75.Chen, V., Wu, S., Ratner, A. J., Weng, J., and Ré, C. Slice-based learning: A programming model for residual learning in critical data slices. In Advances in Neural Information Processing Systems (NeurIPS), pp. 9397–9407, 2019b.
  76. 76.Ching, T., Himmelstein, D. S., Beaulieu-Jones, B. K., Kalinin, A. A., Do, B. T., Way, G. P., Ferrero, E., Agapow, P., Zietz, M., Hoffman, M. M., et al. Opportunities and obstacles for deep learning in biology and medicine. Journal of The Royal Society Interface, 15(141), 2018.
  77. 77.Christie, G., Fendley, N., Wilson, J., and Mukherjee, R. Functional map of the world. In Computer Vision and Pattern Recognition (CVPR), 2018.
  78. 78.Chung, J. S., Nagrani, A., and Zisserman, A. Voxceleb2: Deep speaker recognition. Proc. Interspeech, pp. 1086–1090, 2018.
  79. 79.Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J. Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages. arXiv preprint arXiv:2003.05002, 2020.
  80. 80.Codella, N., Rotemberg, V., Tschandl, P., Celebi, M. E., Dusza, S., Gutman, D., Helba, B., Kalloo, A., Liopyris, K., Marchetti, M., et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (ISIC). arXiv preprint arXiv:1902.03368, 2019.
  81. 81.Conneau, A. and Lample, G. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems (NeurIPS), pp. 7059–7069, 2019.
  82. 82.Conneau, A., Rinott, R., Lample, G., Williams, A., Bowman, S., Schwenk, H., and Stoyanov, V. Xnli: Evaluating cross-lingual sentence representations. In Empirical Methods in Natural Language Processing (EMNLP), pp. 2475–2485, 2018.
  83. 83.Consortium, E. P. et al. An integrated encyclopedia of DNA elements in the human genome. Nature, 489(7414):57–74, 2012.
  84. 84.Consortium, G. et al. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science, 369(6509):1318–1330, 2020.
  85. 85.Consortium, H. et al. The human body at cellular resolution: the NIH human biomolecular atlas program. Nature, 574(7777), 2019.
  86. 86.Corbett-Davies, S. and Goel, S. The measure and mismeasure of fairness: A critical review of fair machine learning. arXiv preprint arXiv:1808.00023, 2018.
  87. 87.Corbett-Davies, S., Pierson, E., Feller, A., and Goel, S. A computer program used for bail and sentencing decisions was labeled biased against blacks. It’s actually not that clear. Washington Post, 2016. ISSN 0190-8286. URL https://www.washingtonpost.com/news/monkey-cage/wp/2016/10/17/can-an-algorithm-be-racist-our-analysis-is-more-cautious-than-propublicas/.
  88. 88.Corbett-Davies, S., Pierson, E., Feller, A., Goel, S., and Huq, A. Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 797–806, 2017.
  89. 89.Cordella, L. P., Stefano, C. D., Tortorella, F., and Vento, M. A method for improving classification reliability of multilayer perceptrons. IEEE Transactions on Neural Networks, 6(5):1140–1147, 1995.
  90. 90.Courtiol, P., Maussion, C., Moarii, M., Pronier, E., Pilcer, S., Sefta, M., Manceron, P., Toldo, S., Zaslavskiy, M., Stang, N. L., et al. Deep learning-based classification of mesothelioma improves prediction of patient outcome. Nature Medicine, 25(10):1519–1525, 2019.
  91. 91.Croce, F., Andriushchenko, M., Sehwag, V., Flammarion, N., Chiang, M., Mittal, P., and Hein, M. Robustbench: a standardized adversarial robustness benchmark. arXiv preprint arXiv:2010.09670, 2020.
  92. 92.Crunchant, A.-S., Borchers, D., Kühl, H., and Piel, A. Listening and watching: Do camera traps or acoustic sensors more efficiently detect wild chimpanzees in an open habitat? Methods in Ecology and Evolution, 11(4):542–552, 2020.
  93. 93.Cuccarese, M. F., Earnshaw, B. A., Heiser, K., Fogelson, B., Davis, C. T., McLean, P. F., Gordon, H. B., Skelly, K., Weathersby, F. L., Rodic, V., et al. Functional immune mapping with deep-learning enabled phenomics applied to immunomodulatory and COVID-19 drug discovery. bioRxiv, 2020.
  94. 94.Cui, Y., Jia, M., Lin, T., Song, Y., and Belongie, S. Class-balanced loss based on effective number of samples. In Computer Vision and Pattern Recognition (CVPR), pp. 9268–9277, 2019.
  95. 95.Dai, D. and Van Gool, L. Dark model adaptation: Semantic image segmentation from daytime to nighttime. In International Conference on Intelligent Transportation Systems (ITSC), 2018.
  96. 96.D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al. Underspecification presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395, 2020a.
  97. 97.D’Amour, A., Srinivasan, H., Atwood, J., Baljekar, P., Sculley, D., and Halpern, Y. Fairness is not static: deeper understanding of long term fairness via simulation studies. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp. 525–534, 2020b.
  98. 98.David, E., Madec, S., Sadeghi-Tehran, P., Aasen, H., Zheng, B., Liu, S., Kirchgessner, N., Ishikawa, G., Nagasawa, K., Badhon, M. A., Pozniak, C., de Solan, B., Hund, A., Chapman, S. C., Baret, F., Stavness, I., and Guo, W. Global wheat head detection (gwhd) dataset: a large and diverse dataset of high-resolution rgb-labelled images to develop and benchmark wheat head detection methods. Plant Phenomics, 2020, 2020.
  99. 99.David, E., Serouart, M., Smith, D., Madec, S., Velumani, K., Liu, S., Wang, X., Espinosa, F. P., Shafiee, S., Tahir, I. S. A., Tsujimoto, H., Nasuda, S., Zheng, B., Kichgessner, N., Aasen, H., Hund, A., Sadhegi-Tehran, P., Nagasawa, K., Ishikawa, G., Dandrifosse, S., Carlier, A., Mercatoris, B., Kuroki, K., Wang, H., Ishii, M., Badhon, M. A., Pozniak, C., LeBauer, D. S., Lilimo, M., Poland, J., Chapman, S., de Solan, B., Baret, F., Stavness, I., and Guo, W. Global wheat head dataset 2021: an update to improve the benchmarking wheat head localization with more diversity, 2021.
  100. 100.Davis, S. E., Lasko, T. A., Chen, G., Siew, E. D., and Matheny, M. E. Calibration drift in regression and machine learning models for acute kidney injury. Journal of the American Medical Informatics Association, 24(6):1052–1061, 2017.
  101. 101.DeGrave, A. J., Janizek, J. D., and Lee, S. AI for radiographic COVID-19 detection selects shortcuts over signal. medRxiv, 2020.
  102. 102.Desmarais, M. C. and Baker, R. A review of recent advances in learner and skill modeling in intelligent learning environments. User Modeling and User-Adapted Interaction, 22(1):9–38, 2012.
  103. 103.DigitalGlobe, N. and Works, C. Spacenet. https://aws.amazon.com/publicdatasets/spacenet/, 2016.
  104. 104.Dill, K. A. and MacCallum, J. L. The protein-folding problem, 50 years on. Science, 338(6110):1042–1046, 2012.
  105. 105.Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L. Measuring and mitigating unintended bias in text classification. In Association for the Advancement of Artificial Intelligence (AAAI), pp. 67–73, 2018.
  106. 106.Djolonga, J., Yung, J., Tschannen, M., Romijnders, R., Beyer, L., Kolesnikov, A., Puigcerver, J., Minderer, M., D’Amour, A., Moldovan, D., et al. On robustness and transferability of convolutional neural networks. arXiv preprint arXiv:2007.08558, 2020.
  107. 107.Dodge, S. and Karam, L. A study and comparison of human and deep learning recognition performance under visual distortions. In 26th International Conference on Computer Communication and Networks (ICCCN), pp. 1–7. IEEE, 2017.
  108. 108.Dou, Q., Castro, D., Kamnitsas, K., and Glocker, B. Domain generalization via model-agnostic learning of semantic features. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  109. 109.Dreccer, M. F., Molero, G., Rivera-Amado, C., John-Bejai, C., and Wilson, Z. Yielding to the image: how phenotyping reproductive growth can assist crop improvement and production. Plant science, 282:73–82, 2019.
  110. 110.Dressel, J. and Farid, H. The accuracy, fairness, and limits of predicting recidivism. Science Advances, 4(1), 2018.
  111. 111.Duchi, J. and Namkoong, H. Learning models with uniform performance via distributionally robust optimization. Annals of Statistics, 2021.
  112. 112.Duchi, J., Hashimoto, T., and Namkoong, H. Distributionally robust losses for latent covariate mixtures. arXiv preprint arXiv:2007.13982, 2020.
  113. 113.Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R. Fairness through awareness. In Innovations in Theoretical Computer Science (ITCS), pp. 214–226, 2012.
  114. 114.Echeverri, C. J. and Perrimon, N. High-throughput rnai screening in cultured cells: a user’s guide. Nature Reviews Genetics, 7(5):373, 2006.
  115. 115.Elvidge, C. D., Sutton, P. C., Ghosh, T., Tuttle, B. T., Baugh, K. E., Bhaduri, B., and Bright, E. A global poverty map derived from satellite data. Computers and Geosciences, 35, 2009.
  116. 116.Eraslan, G., Avsec, Ž., Gagneur, J., and Theis, F. J. Deep learning: new computational modelling techniques for genomics. Nature Reviews Genetics, 20(7):389–403, 2019.
  117. 117.Espey, J., Swanson, E., Badiee, S., Chistensen, Z., Fischer, A., Levy, M., Yetman, G., de Sherbinin, A., Chen, R., Qiu, Y., Greenwell, G., Klein, T., , Jutting, J., Jerven, M., Cameron, G., Rivera, A. M. A., Arias, V. C., , Mills, S. L., and Motivans, A. Data for development: A needs assessment for SDG monitoring and statistical capacity development. Sustainable Development Solutions Network, 2015.
  118. 118.Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., and Thrun, S. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017.
  119. 119.et al, O. Solving Rubik’s cube with a robot hand. arXiv preprint arXiv:1910.07113, 2019.
  120. 120.Fan, Z., Lu, J., Gong, M., Xie, H., and Goodman, E. D. Automatic tobacco plant detection in uav images via deep neural networks. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 11(3): 876–887, 2018.
  121. 121.Fang, C., Xu, Y., and Rockmore, D. N. Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias. In International Conference on Computer Vision (ICCV), pp. 1657–1664, 2013.
  122. 122.Feng, J., Sondhi, A., Perry, J., and Simon, N. Selective prediction-set models with coverage guarantees. arXiv preprint arXiv:1906.05473, 2019.
  123. 123.Filmer, D. and Scott, K. Assessing asset indices. Demography, 49, 2011.
  124. 124.Franks, C., Tu, Z., Devanbu, P., and Hellendoorn, V. Cacheca: A cache language model based code suggestion tool. In International Conference on Software Engineering (ICSE), 2015.
  125. 125.Fuentes, A., Yoon, S., Kim, S. C., and Park, D. S. A robust deep-learning-based detector for real-time tomato plant diseases and pests recognition. Sensors, 17(9):2022, 2017.
  126. 126.Futoma, J., Simons, M., Panch, T., Doshi-Velez, F., and Celi, L. A. The myth of generalisability in clinical research and machine learning in health care. The Lancet Digital Health, 2(9):e489–e492, 2020.
  127. 127.Gal, Y. and Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning (ICML), 2016.
  128. 128.Ganin, Y. and Lempitsky, V. Unsupervised domain adaptation by backpropagation. In International Conference on Machine Learning (ICML), pp. 1180–1189, 2015.
  129. 129.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., March, M., and Lempitsky, V. Domain-adversarial training of neural networks. Journal of Machine Learning Research (JMLR), 17, 2016.
  130. 130.Garg, S., Wu, Y., Balakrishnan, S., and Lipton, Z. C. A unified view of label shift estimation. arXiv preprint arXiv:2003.07554, 2020.
  131. 131.Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Ill, H. D., and Crawford, K. Datasheets for datasets. arXiv preprint arXiv:1803.09010, 2018.
  132. 132.Geifman, Y. and El-Yaniv, R. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  133. 133.Geifman, Y. and El-Yaniv, R. Selectivenet: A deep neural network with an integrated reject option. In International Conference on Machine Learning (ICML), 2019.
  134. 134.Geifman, Y., Uziel, G., and El-Yaniv, R. Bias-reduced uncertainty estimation for deep neural classifiers. In International Conference on Learning Representations (ICLR), 2018.
  135. 135.Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv preprint arXiv:1811.12231, 2018a.
  136. 136.Geirhos, R., Temme, C. R., Rauber, J., Schütt, H. H., Bethge, M., and Wichmann, F. A. Generalisation in humans and deep neural networks. Advances in Neural Information Processing Systems, 31:7538–7550, 2018b.
  137. 137.Geirhos, R., Jacobsen, J., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. Shortcut learning in deep neural networks. arXiv preprint arXiv:2004.07780, 2020.
  138. 138.Gelman, A., Fagan, J., and Kiss, A. An Analysis of the New York City Police Department’s “Stop-and-Frisk” Policy in the Context of Claims of Racial Bias. Journal of the American Statistical Association, 102(479):813–823, Sep 2007. ISSN 0162-1459. doi: 10.1198/016214506000001040. URL https://amstat.tandfonline.com/doi/abs/10.1198/016214506000001040. Publisher: Taylor & Francis.
  139. 139.Geva, M., Goldberg, Y., and Berant, J. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  140. 140.Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E. Neural message passing for quantum chemistry. In International Conference on Machine Learning (ICML), pp. 1273–1272, 2017.
  141. 141.Godinez, W. J., Hossain, I., and Zhang, X. Unsupervised phenotypic analysis of cellular images with multi-scale convolutional neural networks. BioRxiv, pp. 361410, 2018.
  142. 142.Goel, K., Gu, A., Li, Y., and Ré, C. Model patching: Closing the subgroup performance gap with data augmentation. arXiv preprint arXiv:2008.06775, 2020.
  143. 143.Goel, S., Rao, J. M., and Shroff, R. Precinct or prejudice? Understanding racial disparities in New York City’s stop-and-frisk policy. The Annals of Applied Statistics, 10(1): 365–394, March 2016. ISSN 1932-6157. doi: 10.1214/15-AOAS897. URL http://projecteuclid.org/euclid.aoas/1458909920.
  144. 144.Gogoll, D., Lottes, P., Weyler, J., Petrinic, N., and Stachniss, C. Unsupervised domain adaptation for transferring plant classification systems to new field environments, crops, and robots. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 2636–2642. IEEE, 2020.
  145. 145.Goh, W. W. B., Wang, W., and Wong, L. Why batch effects matter in omics data, and how to avoid them. Trends in biotechnology, 35(6):498–507, 2017.
  146. 146.Goodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR), 2015.
  147. 147.Graetz, N., Friedman, J., Osgood-Zimmerman, A., Burstein, R., Biehl, M. H., Shields, C., Mosser, J. F., Casey, D. C., Deshpande, A., Earl, L., Reiner, R. C., Ray, S. E., Fullman, N., Levine, A. J., Stubbs, R. W., Mayala, B. K., Longbottom, J., Browne, A. J., Bhatt, S., Weiss, D. J., Gething, P. W., Mokdad, A. H., Lim, S. S., Murray, C. J. L., Gakidou, E., and Hay, S. I. Mapping local variation in educational attainment across Africa. Nature, 555, 2018.
  148. 148.Grooten, M., Peterson, T., and Almond, R. Living Planet Report 2020 - Bending the curve of biodiversity loss. WWF, Gland, Switzerland, 2020.
  149. 149.Gu, S., Holly, E., Lillicrap, T., and Levine, S. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In International Conference on Robotics and Automation (ICRA), 2017.
  150. 150.Gulrajani, I. and Lopez-Paz, D. In search of lost domain generalization. arXiv preprint arXiv:2007.01434, 2020.
  151. 151.Guo, J., Shah, D. J., and Barzilay, R. Multi-source domain adaptation with mixture of experts. arXiv preprint arXiv:1809.02256, 2018.
  152. 152.Gupta, A., Murali, A., Gandhi, D., and Pinto, L. Robot learning in homes: Improving generalization and reducing dataset bias. In Advances in Neural Information Processing Systems (NIPS), 2018.
  153. 153.Gurcan, M. N., Boucheron, L. E., Can, A., Madabhushi, A., Rajpoot, N. M., and Yener, B. Histopathological image analysis: A review. IEEE reviews in biomedical engineering, 2:147–171, 2009.
  154. 154.Han, X. and Tsvetkov, Y. Fortifying toxic speech detectors against veiled toxicity. arXiv preprint arXiv:2010.03154, 2020.
  155. 155.Hand, D. J. Classifier technology and the illusion of progress. Statistical science, pp. 1–14, 2006.
  156. 156.Hansen, M. C., Potapov, P. V., Moore, R., Hancher, M., Turubanova, S. A., Tyukavina, A., Thau, D., Stehman, S. V., Goetz, S. J., Loveland, T. R., Kommareddy, A., Egorov, A., Chini, L., Justice, C. O., and Townshend, J. R. G. High-resolution global maps of 21st-century forest cover change. Science, 342, 2013.
  157. 157.Harrill, J., Shah, I., Setzer, R. W., Haggard, D., Auerbach, S., Judson, R., and Thomas, R. S. Considerations for strategic use of high-throughput transcriptomics chemical screening data in regulatory decisions. Current Opinion in Toxicology, 15:64–75, 2019. ISSN 2468-2020. doi: https://doi.org/10.1016/j.cotox.2019.05.004. URL https://www.sciencedirect.com/science/article/pii/S2468202019300129. Risk Assessment in Toxicology.
  158. 158.Hashimoto, T. B., Srivastava, M., Namkoong, H., and Liang, P. Fairness without demographics in repeated loss minimization. In International Conference on Machine Learning (ICML), 2018.
  159. 159.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Computer Vision and Pattern Recognition (CVPR), 2016.
  160. 160.He, Y., Shen, Z., and Cui, P. Towards non-IID image classification: A dataset and baselines. Pattern Recognition, 110, 2020.
  161. 161.Heinze-Deml, C. and Meinshausen, N. Conditional variance penalties and domain shift robustness. arXiv preprint arXiv:1710.11469, 2017.
  162. 162.Hellendoorn, V. J., Proksch, S., Gall, H. C., and Bacchelli, A. When code completion fails: A case study on real-world completions. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), pp. 960–970. IEEE, 2019.
  163. 163.Henderson, B. E., Lee, N. H., Seewaldt, V., and Shen, H. The influence of race and ethnicity on the biology of cancer. Nature Reviews Cancer, 12(9):648–653, 2012.
  164. 164.Hendrycks, D. and Dietterich, T. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations (ICLR), 2019.
  165. 165.Hendrycks, D. and Gimpel, K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In International Conference on Learning Representations (ICLR), 2017.
  166. 166.Hendrycks, D., Basart, S., Mazeika, M., Mostajabi, M., Steinhardt, J., and Song, D. Scaling out-of-distribution detection for real-world settings. arXiv preprint arXiv:1911.11132, 2020a.
  167. 167.Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., and Gilmer, J. The many faces of robustness: A critical analysis of out-of-distribution generalization. arXiv preprint arXiv:2006.16241, 2020b.
  168. 168.Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D. Pretrained transformers improve out-of-distribution robustness. arXiv preprint arXiv:2004.06100, 2020c.
  169. 169.Ho, J. W., Jung, Y. L., Liu, T., Alver, B. H., Lee, S., Ikegami, K., Sohn, K., Minoda, A., Tolstorukov, M. Y., Appert, A., et al. Comparative analysis of metazoan chromatin organization. Nature, 512(7515):449–452, 2014.
  170. 170.Hoffman, J., Tzeng, E., Park, T., Zhu, J., Isola, P., Saenko, K., Efros, A. A., and Darrell, T. Cycada: Cycle consistent adversarial domain adaptation. In International Conference on Machine Learning (ICML), 2018.
  171. 171.Hovy, D. and Spruit, S. L. The social impact of natural language processing. In Association for Computational Linguistics (ACL), pp. 591–598, 2016.
  172. 172.Hu, J., Ruder, S., Siddhant, A., Neubig, G., Firat, O., and Johnson, M. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. arXiv preprint arXiv:2003.11080, 2020a.
  173. 173.Hu, W., Niu, G., Sato, I., and Sugiyama, M. Does distributionally robust supervised learning give robust classifiers? In International Conference on Machine Learning (ICML), 2018.
  174. 174.Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J. Open graph benchmark: Datasets for machine learning on graphs. arXiv preprint arXiv:2005.00687, 2020b.
  175. 175.Huang, G., Liu, Z., Maaten, L. V. D., and Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 4700–4708, 2017.
  176. 176.Hughes, J. P., Rees, S., Kalindjian, S. B., and Philpott, K. L. Principles of early drug discovery. British journal of pharmacology, 162(6):1239–1249, 2011.
  177. 177.Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436, 2019.
  178. 178.Jaganathan, K., Panagiotopoulou, S. K., McRae, J. F., Darbandi, S. F., Knowles, D., Li, Y. I., Kosmicki, J. A., Arbelaez, J., Cui, W., Schwartz, G. B., et al. Predicting splicing from primary sequence with deep learning. Cell, 176(3):535–548, 2019.
  179. 179.Jean, N., Burke, M., Xie, M., Davis, W. M., Lobell, D. B., and Ermon, S. Combining satellite imagery and machine learning to predict poverty. Science, 353, 2016.
  180. 180.Jean, N., Xie, S. M., and Ermon, S. Semi-supervised deep kernel learning: Regression with unlabeled data by minimizing predictive variance. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  181. 181.Jin, W., Barzilay, R., and Jaakkola, T. Enforcing predictive invariance across structured biomedical domains. arXiv preprint arXiv:2006.03908, 2020.
  182. 182.Johnson, A. E., Pollard, T. J., Shen, L., Li-Wei, H. L., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L. A., and Mark, R. G. Mimic-iii, a freely accessible critical care database. Scientific Data, 3(1):1–9, 2016.
  183. 183.Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Computer Vision and Pattern Recognition (CVPR), 2017.
  184. 184.Jones, E., Sagawa, S., Koh, P. W., Kumar, A., and Liang, P. Selective classification can magnify disparities across groups. In International Conference on Learning Representations (ICLR), 2021.
  185. 185.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Tunyasuvunakool, K., Ronneberger, O., Bates, R., Žídek, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Potapenko, A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Steinegger, M., Pacholska, M., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D. High accuracy protein structure prediction using deep learning. Fourteenth Critical Assessment of Techniques for Protein Structure Prediction, 2020.
  186. 186.Jung, J., Goel, S., Skeem, J., et al. The limits of human predictions of recidivism. Science Advances, 6(7), 2020.
  187. 187.Jørgensen, A. K., Hovy, D., and Søgaard, A. Challenges of studying and processing dialects in social media. In ACL Workshop on Noisy User-generated Text, pp. 9–18, 2015.
  188. 188.Kahn, G., Abbeel, P., and Levine, S. BADGR: An autonomous self-supervised learning-based navigation system. arXiv preprint arXiv:2002.05700, 2020.
  189. 189.Kallus, N. and Zhou, A. Residual Unfairness in Fair Machine Learning from Prejudiced Data. arXiv:1806.02887 [cs, stat], June 2018. URL http://arxiv.org/abs/1806.02887. arXiv: 1806.02887.
  190. 190.Kamath, A., Jia, R., and Liang, P. Selective question answering under domain shift. In Association for Computational Linguistics (ACL), 2020.
  191. 191.Katona, Z., Painter, M., Patatoukas, P. N., and Zeng, J. On the capital market consequences of alternative data: Evidence from outer space. Miami Behavioral Finance Conference, 2018.
  192. 192.Kaushik, D., Hovy, E., and Lipton, Z. Learning the difference that makes a difference with counterfactually-augmented data. In International Conference on Learning Representations (ICLR), 2019.
  193. 193.Kearns, M., Neel, S., Roth, A., and Wu, Z. S. Preventing fairness gerrymandering: Auditing and learning for subgroup fairness. In International Conference on Machine Learning (ICML), pp. 2564–2572, 2018.
  194. 194.Keilwagen, J., Posch, S., and Grau, J. Accurate prediction of cell type-specific transcription factor binding. Genome Biology, 20(1), 2019.
  195. 195.Kelley, D. R., Snoek, J., and Rinn, J. L. Basset: learning the regulatory code of the accessible genome with deep convolutional neural networks. Genome Research, 26(7): 990–999, 2016.
  196. 196.Kim, J. H., Xie, M., Jean, N., and Ermon, S. Incorporating spatial context and fine-grained detail from satellite imagery to predict poverty. Stanford University, 2016a.
  197. 197.Kim, N. and Linzen, T. Cogs: A compositional generalization challenge based on semantic interpretation. arXiv preprint arXiv:2010.05465, 2020.
  198. 198.Kim, S., Thiessen, P. A., Bolton, E. E., Chen, J., Fu, G., Gindulyte, A., Han, L., He, J., He, S., Shoemaker, B. A., Wang, J., Yu, B., Zhang, J., and Bryant, S. H. Pubchem substance and compound databases. Nucleic Acids Research, 44(D1):D1202–D1213, 2016b.
  199. 199.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  200. 200.Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., and Goel, S. Racial disparities in automated speech recognition. Science, 117(14):7684–7689, 2020.
  201. 201.Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P. Concept bottleneck models. In International Conference on Machine Learning (ICML), 2020.
  202. 202.Kompa, B., Snoek, J., and Beam, A. Empirical frequentist coverage of deep learning uncertainty quantification procedures. arXiv preprint arXiv:2010.03039, 2020.
  203. 203.Komura, D. and Ishikawa, S. Machine learning methods for histopathological image analysis. Computational and Structural Biotechnology Journal, 16:34–42, 2018.
  204. 204.Kulal, S., Pasupat, P., Chandra, K., Lee, M., Padon, O., Aiken, A., and Liang, P. S. Spoc: Search-based pseudocode to code. In Advances in Neural Information Processing Systems, pp. 11906–11917, 2019.
  205. 205.Kulkarni, C., Koh, P. W., Huy, H., Chia, D., Papadopoulos, K., Cheng, J., Koller, D., and Klemmer, S. R. Peer and self assessment in massive online classes. Design Thinking Research, pp. 131–168, 2015.
  206. 206.Kulkarni, C. E., Socher, R., Bernstein, M. S., and Klemmer, S. R. Scaling short-answer grading by combining peer assessment with algorithmic scoring. In Proceedings of the first ACM conference on Learning@Scale conference, pp. 99–108, 2014.
  207. 207.Kumar, A., Ma, T., and Liang, P. Understanding self-training for gradual domain adaptation. In International Conference on Machine Learning (ICML), 2020.
  208. 208.Kundaje, A., Meuleman, W., Ernst, J., Bilenky, M., Yen, A., Heravi-Moussavi, A., Kheradpour, P., Zhang, Z., Wang, J., Ziller, M. J., et al. Integrative analysis of 111 reference human epigenomes. Nature, 518(7539):317–330, 2015.
  209. 209.Kuznichov, D., Zvirin, A., Honen, Y., and Kimmel, R. Data augmentation for leaf segmentation and counting tasks in rosette plants. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 0–0, 2019.
  210. 210.Lake, B. and Baroni, M. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International Conference on Machine Learning (ICML), 2018.
  211. 211.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  212. 212.Landrum, G. et al. Rdkit: Open-source cheminformatics, 2006.
  213. 213.Larrazabal, A. J., Nieto, N., Peterson, V., Milone, D. H., and Ferrante, E. Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proceedings of the National Academy of Sciences, 117(23):12592–12594, 2020.
  214. 214.Larson, J., Mattu, S., Kirchner, L., and Angwin, J. How we analyzed the compas recidivism algorithm. ProPublica, 9(1), 2016.
  215. 215.Latessa, E. J., Lemke, R., Makarios, M., and Smith, P. The Creation and Validation of the Ohio Risk Assessment System (ORAS). Federal Probation, 74:16, 2010. URL https://heinonline.org/HOL/Page?handle=hein.journals/fedpro74&id=16&div=&collection=.
  216. 216.Lau, R. Y., Li, C., and Liao, S. S. Social analytics: Learning fuzzy product ontologies for aspect-oriented sentiment analysis. Decision Support Systems, 65:80–94, 2014.
  217. 217.LeCun, Y., Cortes, C., and Burges, C. J. The MNIST database of handwritten digits. http://yann.lecun.com/exdb/mnist/, 1998.
  218. 218.Leek, J. T., Scharpf, R. B., Bravo, H. C., Simcha, D., Langmead, B., Johnson, W. E., Geman, D., Baggerly, K., and Irizarry, R. A. Tackling the widespread and critical impact of batch effects in high-throughput data. Nature Reviews Genetics, 11(10), 2010.
  219. 219.Li, D., Yang, Y., Song, Y., and Hospedales, T. M. Deeper, broader and artier domain generalization. In Proceedings of the IEEE International Conference on Computer Vision, pp. 5542–5550, 2017a.
  220. 220.Li, D., Yang, Y., Song, Y.-Z., and Hospedales, T. Learning to generalize: Meta-learning for domain generalization. In Association for the Advancement of Artificial Intelligence (AAAI), 2018a.
  221. 221.Li, H. and Guan, Y. Leopard: fast decoding cell type-specific transcription factor binding landscape at singlenucleotide resolution. bioRxiv, 2019.
  222. 222.Li, H., Pan, S. J., Wang, S., and Kot, A. C. Domain generalization with adversarial feature learning. In Computer Vision and Pattern Recognition (CVPR), pp. 5400–5409, 2018b.
  223. 223.Li, H., Quang, D., and Guan, Y. Anchor: trans-cell type prediction of transcription factor binding sites. Genome Research, 29(2):281–292, 2019a.
  224. 224.Li, J., Miller, A. H., Chopra, S., Ranzato, M., and Weston, J. Dialogue learning with human-in-the-loop. In International Conference on Learning Representations (ICLR), 2017b.
  225. 225.Li, T., Sanjabi, M., Beirami, A., and Smith, V. Fair resource allocation in federated learning. arXiv preprint arXiv:1905.10497, 2019b.
  226. 226.Li, Y., Wang, N., Shi, J., Liu, J., and Hou, X. Revisiting batch normalization for practical domain adaptation. In International Conference on Learning Representations Workshop (ICLRW), 2017c.
  227. 227.Li, Y., Tian, X., Gong, M., Liu, Y., Liu, T., Zhang, K., and Tao, D. Deep domain generalization via conditional invariant adversarial networks. In European Conference on Computer Vision (ECCV), pp. 624–639, 2018c.
  228. 228.Liang, S., Li, Y., and Srikant, R. Enhancing the reliability of out-of-distribution image detection in neural networks. In International Conference on Learning Representations (ICLR), 2018.
  229. 229.Libbrecht, M. W. and Noble, W. S. Machine learning applications in genetics and genomics. Nature Reviews Genetics, 16(6):321–332, 2015.
  230. 230.Lipton, Z., Wang, Y., and Smola, A. Detecting and correcting for label shift with black box predictors. In International Conference on Machine Learning (ICML), 2018.
  231. 231.Liu, L. T., Dean, S., Rolf, E., Simchowitz, M., and Hardt, M. Delayed impact of fair machine learning. In International Conference on Machine Learning (ICML), 2018.
  232. 232.Liu, Y., Gadepalli, K., Norouzi, M., Dahl, G. E., Kohlberger, T., Boyko, A., Venugopalan, S., Timofeev, A., Nelson, P. Q., Corrado, G. S., et al. Detecting cancer metastases on gigapixel pathology images. arXiv preprint arXiv:1703.02442, 2017.
  233. 233.Ljosa, V., Sokolnicki, K. L., and Carpenter, A. E. Annotated high-throughput microscopy image sets for validation. Nature methods, 9(7):637–637, 2012.
  234. 234.Long, M., Cao, Y., Wang, J., and Jordan, M. Learning transferable features with deep adaptation networks. In International Conference on Machine Learning, pp. 97–105, 2015.
  235. 235.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  236. 236.Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664, 2021.
  237. 237.Lum, K. and Isaac, W. To predict and serve? Significance, 13(5):14–19, 2016.
  238. 238.Lum, K. and Shah, T. Measures of fairness for New York City’s Supervised Release Risk Assessment Tool. Human Rights Data Analytics Group, pp. 21, 2019.
  239. 239.Lyu, J., Wang, S., Balius, T. E., Singh, I., Levit, A., Moroz, Y. S., O’Meara, M. J., Che, T., Algaa, E., Tolmachova, K., et al. Ultra-large library docking for discovering new chemotypes. Nature, 566(7743):224–229, 2019.
  240. 240.Macarron, R., Banks, M. N., Bojanic, D., Burns, D. J., Cirovic, D. A., Garyantes, T., Green, D. V., Hertzberg, R. P., Janzen, W. P., Paslay, J. W., et al. Impact of high-throughput screening in biomedical research. Nature reviews Drug discovery, 10(3):188, 2011.
  241. 241.Macenko, M., Niethammer, M., Marron, J. S., Borland, D., Woosley, J. T., Guan, X., Schmitt, C., and Thomas, N. E. A method for normalizing histology slides for quantitative analysis. In 2009 IEEE International Symposium on Biomedical Imaging: From Nano to Macro, pp. 1107–1110, 2009.
  242. 242.Madec, S., Jin, X., Lu, H., De Solan, B., Liu, S., Duyme, F., Heritier, E., and Baret, F. Ear density estimation from high resolution rgb imagery using deep learning technique. Agricultural and forest meteorology, 264:225–234, 2019.
  243. 243.Malloy, B. A. and Power, J. F. Quantifying the transition from python 2 to 3: an empirical study of python applications. In 2017 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), pp. 314–323. IEEE, 2017.
  244. 244.Mansour, Y., Mohri, M., and Rostamizadeh, A. Domain adaptation with multiple sources. In Advances in Neural Information Processing Systems (NeurIPS), pp. 1041–1048, 2009.
  245. 245.Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B. Building a large annotated corpus of English: the Penn Treebank. Computational Linguistics, 19:313–330, 1993.
  246. 246.McCloskey, K., Sigel, E. A., Kearns, S., Xue, L., Tian, X., Moccia, D., Gikunju, D., Bazzaz, S., Chan, B., Clark, M. A., et al. Machine learning on DNA-encoded libraries: A new paradigm for hit finding. Journal of Medicinal Chemistry, 2020.
  247. 247.McCoy, R. T., Min, J., and Linzen, T. Berts of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv preprint arXiv:1911.02969, 2019a.
  248. 248.McCoy, R. T., Pavlick, E., and Linzen, T. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Association for Computational Linguistics (ACL), 2019b.
  249. 249.McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., Back, T., Chesus, M., Corrado, G. C., Darzi, A., et al. International evaluation of an AI system for breast cancer screening. Nature, 577(7788):89–94, 2020.
  250. 250.Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. A survey on bias and fairness in machine learning. arXiv preprint arXiv:1908.09635, 2019.
  251. 251.Meinshausen, N. and Bühlmann, P. Maximin effects in inhomogeneous large-scale data. Annals of Statistics, 43, 2015.
  252. 252.Miller, J., Krauth, K., Recht, B., and Schmidt, L. The effect of natural distribution shift on question answering models. arXiv preprint arXiv:2004.14444, 2020.
  253. 253.Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, C., Kumaran, D., and Hadsell, R. Learning to navigate in complex environments. In International Conference on Learning Representations (ICLR), 2017.
  254. 254.Moore, J. E., Purcaro, M. J., Pratt, H. E., Epstein, C. B., Shoresh, N., Adrian, J., Kawli, T., Davis, C. A., Dobin, A., Kaul, R., et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature, 583(7818):699–710, 2020.
  255. 255.Moult, J., Pedersen, J. T., Judson, R., and Fidelis, K. A large-scale experiment to assess protein structure prediction methods. Proteins: Structure, Function, and Bioinformatics, 23(3):ii–iv, 1995.
  256. 256.Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Kolawole, T., Fagbohungbe, T., Akinola, S. O., Muhammad, S. H., Kabongo, S., Osei, S., Freshia, S., Niyongabo, R. A., Macharm, R., Ogayo, P., Ahia, O., Meressa, M., Adeyemi, M., Mokgesi-Selinga, M., Okegbemi, L., Martinus, L. J., Tajudeen, K., Degila, K., Ogueji, K., Siminyu, K., Kreutzer, J., Webster, J., Ali, J. T., Abbott, J., Orife, I., Ezeani, I., Dangana, I. A., Kamper, H., Elsahar, H., Duru, G., Kioko, G., Murhabazi, E., van Biljon, E., Whitenack, D., Onyefuluchi, C., Emezue, C., Dossou, B., Sibanda, B., Bassey, B. I., Olabiyi, A., Ramkilowan, A., Öktem, A., Akinfaderin, A., and Bashir, A. Participatory research for low-resourced machine translation: A case study in African languages. In Findings of Empirical Methods in Natural Language Processing (Findings of EMNLP), 2020.
  257. 257.Nestor, B., McDermott, M., Boag, W., Berner, G., Naumann, T., Hughes, M. C., Goldenberg, A., and Ghassemi, M. Feature robustness in non-stationary health records: caveats to deployable model performance in common clinical machine learning tasks. arXiv preprint arXiv:1908.00690, 2019.
  258. 258.Nguyen, A. T. and Nguyen, T. N. Graph-based statistical language model for code. In International Conference on Software Engineering (ICSE), 2015.
  259. 259.Ni, J., Li, J., and McAuley, J. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Empirical Methods in Natural Language Processing (EMNLP), pp. 188–197, 2019.
  260. 260.Nita, M. and Notkin, D. Using twinning to adapt programs to alternative apis. In 2010 ACM/IEEE 32nd International Conference on Software Engineering, volume 1, pp. 205–214. IEEE, 2010.
  261. 261.Noor, A., Alegana, V., Gething, P., Tatem, A., and Snow, R. Using remotely sensed night-time light as a proxy for poverty in Africa. Population Health Metrics, 6, 2008.
  262. 262.Norouzzadeh, M. S., Morris, D., Beery, S., Joshi, N., Jojic, N., and Clune, J. A deep active learning system for species identification and counting in camera trap images. arXiv preprint arXiv:1910.09716, 2019.
  263. 263.Nygaard, V., Rødland, E. A., and Hovig, E. Methods that remove batch effects while retaining group differences may lead to exaggerated confidence in downstream analyses. Biostatistics, 17(1):29–39, 2016.
  264. 264.NYTimes. The Times is partnering with Jigsaw to expand comment capabilities. The New York Times, 2016. URL https://www.nytco.com/press/the-times-is-partnering-with-jigsaw-to-expand-comment-capabilities/.
  265. 265.Obermeyer, Z., Powers, B., Vogeli, C., and Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, 2019.
  266. 266.Oren, Y., Sagawa, S., Hashimoto, T., and Liang, P. Distributionally robust language modeling. In Empirical Methods in Natural Language Processing (EMNLP), 2019.
  267. 267.Osgood-Zimmerman, A., Millear, A. I., Stubbs, R. W., Shields, C., Pickering, B. V., Earl, L., Graetz, N., Kinyoki, D. K., Ray, S. E., Bhatt, S., Browne, A. J., Burstein, R., Cameron, E., Casey, D. C., Deshpande, A., Fullman, N., Gething, P. W., Gibson, H. S., Henry, N. J., Herrero, M., Krause, L. K., Letourneau, I. D., Levine, A. J., Liu, P. Y., Longbottom, J., Mayala, B. K., Mosser, J. F., Noor, A. M., Pigott, D. M., Piwoz, E. G., Rao, P., Rawat, R., Reiner, R. C., Smith, D. L., Weiss, D. J., Wiens, K. E., Mokdad, A. H., Lim, S. S., Murray, C. J. L., Kassebaum, N. J., and Hay, S. I. Mapping child growth failure in Africa between 2000 and 2015. Nature, 555, 2018.
  268. 268.Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  269. 269.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an ASR corpus based on public domain audio books. In International Conference on Acoustics, Speech, and Signal Processing (ICASSP), pp. 5206–5210, 2015.
  270. 270.Parham, J., Crall, J., Stewart, C., Berger-Wolf, T., and Rubenstein, D. I. Animal population censusing at scale with citizen science and photographic identification. In AAAI Spring Symposium-Technical Report, 2017.
  271. 271.Park, J. H., Shin, J., and Fung, P. Reducing gender bias in abusive language detection. In Empirical Methods in Natural Language Processing (EMNLP), pp. 2799–2804, 2018.
  272. 272.Parker, H. S. and Leek, J. T. The practical effect of batch on genomic prediction. Statistical applications in genetics and molecular biology, 11(3), 2012.
  273. 273.Patro, G. K., Biswas, A., Ganguly, N., Gummadi, K. P., and Chakraborty, A. Fairrec: Two-sided fairness for personalized recommendations in two-sided platforms. In Proceedings of The Web Conference 2020, pp. 1194–1204, 2020.
  274. 274.Peng, X., Usman, B., Kaushik, N., Wang, D., Hoffman, J., and Saenko, K. VisDA: A synthetic-to-real benchmark for visual domain adaptation. In Computer Vision and Pattern Recognition (CVPR), pp. 2021–2026, 2018.
  275. 275.Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B. Moment matching for multi-source domain adaptation. In International Conference on Computer Vision (ICCV), 2019.
  276. 276.Peng, X., Coumans, E., Zhang, T., Lee, T., Tan, J., and Levine, S. Learning agile robotic locomotion skills by imitating animals. In Robotics: Science and Systems (RSS), 2020.
  277. 277.Perelman, L. When “the state of the art” is counting words. Assessing Writing, 21:104–111, 2014.
  278. 278.Peters, J., Bühlmann, P., and Meinshausen, N. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society. Series B (Methodological), 78, 2016.
  279. 279.Phillips, N. A., Rajpurkar, P., Sabini, M., Krishnan, R., Zhou, S., Pareek, A., Phu, N. M., Wang, C., Ng, A. Y., and Lungren, M. P. Chexphoto: 10,000+ smartphone photos and synthetic photographic transformations of chest x-rays for benchmarking deep learning robustness. arXiv preprint arXiv:2007.06199, 2020.
  280. 280.Piech, C., Huang, J., Chen, Z., Do, C., Ng, A., and Koller, D. Tuned models of peer assessment in moocs. Educational Data Mining, 2013.
  281. 281.Pierson, E., Corbett-Davies, S., and Goel, S. Fast Threshold Tests for Detecting Discrimination. arXiv:1702.08536 [cs, stat], March 2018. URL http://arxiv.org/abs/1702.08536. arXiv: 1702.08536.
  282. 282.Pimentel, M. A., Clifton, D. A., Clifton, L., and Tarassenko, L. A review of novelty detection. Signal Processing, 99: 215–249, 2014.
  283. 283.Pipal, K. A., Notch, J. J., Hayes, S. A., and Adams, P. B. Estimating escapement for a low-abundance steelhead population using dual-frequency identification sonar (didson). North American Journal of Fisheries Management, 32(5):880–893, 2012.
  284. 284.Price, W. N. and Cohen, I. G. Privacy in the age of medical big data. Nature Medicine, 25(1):37–43, 2019.
  285. 285.Proksch, S., Lerch, J., and Mezini, M. Intelligent code completion with bayesian networks. ACM Transactions on Software Engineering and Methodology (TOSEM), 2015.
  286. 286.Proksch, S., Amann, S., Nadi, S., and Mezini, M. Evaluating the evaluations of code recommender systems: A reality check. In 2016 31st IEEE/ACM International Conference on Automated Software Engineering (ASE), 2016.
  287. 287.Quang, D. and Xie, X. Factornet: a deep learning framework for predicting cell type specific transcription factor binding from nucleotide-resolution sequential data. Methods, 166:40–47, 2019.
  288. 288.Quiñonero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D. Dataset shift in machine learning. The MIT Press, 2009.
  289. 289.Raychev, V., Vechev, M., and Yahav, E. Code completion with statistical language models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation, pp. 419–428, 2014.
  290. 290.Raychev, V., Bielik, P., and Vechev, M. Probabilistic model for code with decision trees. ACM SIGPLAN Notices, 2016.
  291. 291.Ré, C., Niu, F., Gudipati, P., and Srisuwananukorn, C. Overton: A data system for monitoring and improving machine-learned products. arXiv preprint arXiv:1909.05372, 2019.
  292. 292.Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do ImageNet classifiers generalize to ImageNet? In International Conference on Machine Learning (ICML), 2019.
  293. 293.Reiner, R. C., Graetz, N., Casey, D. C., Troeger, C., Garcia, G. M., Mosser, J. F., Deshpande, A., Swartz, S. J., Ray, S. E., Blacker, B. F., Rao, P. C., Osgood-Zimmerman, A., Burstein, R., Pigott, D. M., Davis, I. M., Letourneau, I. D., Earl, L., Ross, J. M., Khalil, I. A., Farag, T. H., Brady, O. J., Kraemer, M. U., Smith, D. L., Bhatt, S., Weiss, D. J., Gething, P. W., Kassebaum, N. J., Mokdad, A. H., Murray, C. J., and Hay, S. I. Variation in childhood diarrheal morbidity and mortality in Africa, 2000–2015. New England Journal of Medicine, 379, 2018.
  294. 294.Reker, D. Practical considerations for active machine learning in drug discovery. Drug Discovery Today: Technologies, 2020.
  295. 295.Ren, S., He, K., Girshick, R., and Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497, 2015.
  296. 296.Reynolds, M., Chapman, S., Crespo-Herrera, L., Molero, G., Mondal, S., Pequeno, D. N., Pinto, F., Pinera-Chavez, F. J., Poland, J., Rivera-Amado, C., et al. Breeder friendly phenotyping. Plant Science, pp. 110396, 2020.
  297. 297.Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Association for Computational Linguistics (ACL), pp. 4902–4912, 2020.
  298. 298.Richter, S. R., Vineet, V., Roth, S., and Koltun, V. Playing for data: Ground truth from computer games. In European Conference on Computer Vision, pp. 102–118, 2016.
  299. 299.Rigaki, M. and Garcia, S. Bringing a GAN to a knife-fight: Adapting malware communication to avoid detection. In 2018 IEEE Security and Privacy Workshops (SPW), pp. 70–75, 2018.
  300. 300.Robbes, R. and Lanza, M. How program history can improve code completion. In International Conference on Automated Software Engineering, 2008.
  301. 301.Rolf, E., Jordan, M. I., and Recht, B. Post-estimation smoothing: A simple baseline for learning with side information. In Artificial Intelligence and Statistics (AISTATS), 2020.
  302. 302.Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A. M. The SYNTHIA dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3234–3243, 2016.
  303. 303.Rosenfeld, A., Zemel, R., and Tsotsos, J. K. The elephant in the room. arXiv preprint arXiv:1808.03305, 2018.
  304. 304.Sadeghi, F. and Levine, S. CAD2RL: Real single-image flight without a single real image. In Robotics: Science and Systems (RSS), 2017.
  305. 305.Sadeghi-Tehran, P., Virlet, N., Sabermanesh, K., and Hawkesford, M. J. Multi-feature machine learning model for automatic segmentation of green fractional vegetation cover for high-throughput field phenotyping. Plant methods, 13(1):1–16, 2017.
  306. 306.Saenko, K., Kulis, B., Fritz, M., and Darrell, T. Adapting visual category models to new domains. In European Conference on Computer Vision, pp. 213–226, 2010.
  307. 307.Saerens, M., Latinne, P., and Decaestecker, C. Adjusting the outputs of a classifier to new a priori probabilities: a simple procedure. Neural Computation, 14(1):21–41, 2002.
  308. 308.Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. In International Conference on Learning Representations (ICLR), 2020a.
  309. 309.Sagawa, S., Raghunathan, A., Koh, P. W., and Liang, P. An investigation of why overparameterization exacerbates spurious correlations. In International Conference on Machine Learning (ICML), 2020b.
  310. 310.Sahn, D. E. and Stifel, D. Exploring alternative measures of welfare in the absence of expenditure data. The Review of Income and Wealth, 49, 2003.
  311. 311.Sanh, V., Debut, L., Chaumond, J., and Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  312. 312.Santurkar, S., Tsipras, D., and Madry, A. Breeds: Benchmarks for subpopulation shift. arXiv, 2020.
  313. 313.Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A. The risk of racial bias in hate speech detection. In Association for Computational Linguistics (ACL), 2019.
  314. 314.Schneider, S. and Zhuang, A. Counting fish and dolphins in sonar images using deep learning. arXiv preprint arXiv:2007.12808, 2020.
  315. 315.Seyyed-Kalantari, L., Liu, G., McDermott, M., and Ghassemi, M. Chexclusion: Fairness gaps in deep chest X-ray classifiers. arXiv preprint arXiv:2003.00827, 2020.
  316. 316.Shakoor, N., Lee, S., and Mockler, T. C. High throughput phenotyping to accelerate crop breeding and monitoring of diseases in the field. Current opinion in plant biology, 38:184–192, 2017.
  317. 317.Shankar, S., Halpern, Y., Breck, E., Atwood, J., Wilson, J., and Sculley, D. No classification without representation: Assessing geodiversity issues in open data sets for the developing world. Advances in Neural Information Processing Systems (NeurIPS) Workshop on Machine Learning for the Developing World, 2017.
  318. 318.Shankar, V., Dave, A., Roelofs, R., Ramanan, D., Recht, B., and Schmidt, L. Do image classifiers generalize across time? arXiv preprint arXiv:1906.02168, 2019.
  319. 319.Shapiro, A., Dentcheva, D., and Ruszczynski, A. ´ Lectures on stochastic programming: modeling and theory. SIAM, 2014.
  320. 320.Shen, J., Qu, Y., Zhang, W., and Yu, Y. Wasserstein distance guided representation learning for domain adaptation. In Association for the Advancement of Artificial Intelligence (AAAI), 2018.
  321. 321.Shermis, M. D. State-of-the-art automated essay scoring: Competition, results, and future directions from a united states demonstration. Assessing Writing, 20:53–76, 2014.
  322. 322.Shetty, R., Schiele, B., and Fritz, M. Not using the car to see the sidewalk–quantifying and controlling the effects of context in classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8218–8226, 2019.
  323. 323.Shi, Y., Thomasson, J. A., Murray, S. C., Pugh, N. A., Rooney, W. L., Shafian, S., Rajan, N., Rouze, G., Morgan, C. L. S., Neely, H. L., Rana, A., Bagavathiannan, M. V., Henrickson, J., Bowden, E., Valasek, J., Olsenholler, J., Bishop, M. P., Sheridan, R., Putman, E. B., Popescu, S., Burks, T., Cope, D., Ibrahim, A., McCutchen, B. F., Baltensperger, D. D., Avant, Jr, R. V., Vidrine, M., and Yang, C. Unmanned aerial vehicles for high-throughput phenotyping and agronomic research. PLOS ONE, 11(7):1–26, 07 2016. doi: 10.1371/journal.pone.0159781. URL https://doi.org/10.1371/journal.pone.0159781.
  324. 324.Shimodaira, H. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90:227–244, 2000.
  325. 325.Shin, R., Kant, N., Gupta, K., Bender, C., Trabucco, B., Singh, R., and Song, D. Synthetic datasets for neural program synthesis. In International Conference on Learning Representations (ICLR), 2019.
  326. 326.Shiu, Y., Palmer, K., Roch, M. A., Fleishman, E., Liu, X., Nosal, E.-M., Helble, T., Cholewiak, D., Gillespie, D., and Klinck, H. Deep neural networks for automated detection of marine mammal species. Scientific Reports, 10(1):1–12, 2020.
  327. 327.Shoichet, B. K. Virtual screening of chemical libraries. Nature, 432(7019):862–865, 2004.
  328. 328.Slack, D., Friedler, S., and Givental, E. Fairness Warnings and Fair-MAML: Learning Fairly with Minimal Data. arXiv:1908.09092 [cs, stat], December 2019. URL http://arxiv.org/abs/1908.09092. arXiv: 1908.09092.
  329. 329.Sohoni, N., Dunnmon, J., Angus, G., Gu, A., and Ré, C. No subclass left behind: Fine-grained robustness in coarse-grained classification problems. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  330. 330.Soneson, C., Gerster, S., and Delorenzi, M. Batch effect confounding leads to strong bias in performance estimates obtained by cross-validation. PloS one, 9(6):e100335, 2014.
  331. 331.Srivastava, D. and Mahony, S. Sequence and chromatin determinants of transcription factor binding and the establishment of cell type-specific binding patterns. Biochimica et Biophysica Acta (BBA)-Gene Regulatory Mechanisms, 1863(6), 2020.
  332. 332.Srivastava, M., Hashimoto, T., and Liang, P. Robustness to Spurious Correlations via Human Annotations. In International Conference on Machine Learning, pp. 9109–9119. PMLR, November 2020. URL http://proceedings.mlr.press/v119/srivastava20a.html. ISSN: 2640-3498.
  333. 333.Sterling, T. and Irwin, J. J. Zinc 15 – ligand discovery for everyone. Journal of Chemical Information and Modeling, 55(11):2324–2337, 2015. doi: 10.1021/acs.jcim.5b00559. PMID: 26479676.
  334. 334.Stowell, D., Wood, M. D., Pamuła, H., Stylianou, Y., and Glotin, H. Automatic acoustic detection of birds through deep learning: the first bird audio detection challenge. Methods in Ecology and Evolution, 10(3):368–380, 2019.
  335. 335.Subbaswamy, A., Adams, R., and Saria, S. Evaluating model robustness to dataset shift. arXiv preprint arXiv:2010.15100, 2020.
  336. 336.Sun, B. and Saenko, K. Deep CORAL: Correlation alignment for deep domain adaptation. In European conference on computer vision, pp. 443–450, 2016.
  337. 337.Sun, B., Feng, J., and Saenko, K. Return of frustratingly easy domain adaptation. In Association for the Advancement of Artificial Intelligence (AAAI), 2016.
  338. 338.Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., Vasudevan, V., Han, W., Ngiam, J., Zhao, H., Timofeev, A., Ettinger, S., Krivokon, M., Gao, A., Joshi, A., Zhao, S., Cheng, S., Zhang, Y., Shlens, J., Chen, Z., and Anguelov, D. Scalability in perception for autonomous driving: Waymo open dataset. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020a.
  339. 339.Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M. Test-time training with self-supervision for generalization under distribution shifts. In International Conference on Machine Learning (ICML), 2020b.
  340. 340.Svyatkovskiy, A., Zhao, Y., Fu, S., and Sundaresan, N. Pythia: ai-assisted code completion system. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 2727–2735, 2019.
  341. 341.Swinney, D. C. and Anthony, J. How were new medicines discovered? Nature reviews Drug discovery, 10(7):507, 2011.
  342. 342.Tabak, G., Fan, M., Yang, S., Hoyer, S., and Davis, G. Correcting nuisance variation using wasserstein distance. PeerJ, 8:e8594, 2020.
  343. 343.Tabak, M. A., Norouzzadeh, M. S., Wolfson, D. W., Sweeney, S. J., VerCauteren, K. C., Snow, N. P., Halseth, J. M., Di Salvo, P. A., Lewis, J. S., White, M. D., et al. Machine learning to classify animal species in camera trap images: Applications in ecology. Methods in Ecology and Evolution, 10(4):585–590, 2019.
  344. 344.Taghipour, K. and Ng, H. T. A neural approach to automated essay scoring. In Proceedings of the 2016 conference on empirical methods in natural language processing, pp. 1882–1891, 2016.
  345. 345.Taori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L. Measuring robustness to natural distribution shifts in image classification. arXiv preprint arXiv:2007.00644, 2020.
  346. 346.Tatman, R. Gender and dialect bias in YouTube’s automatic captions. In Workshop on Ethics in Natural Langauge Processing, volume 1, pp. 53–59, 2017.
  347. 347.Taylor, J., Earnshaw, B., Mabey, B., Victors, M., and Yosinski, J. Rxrx1: An image set for cellular morphological variation across many experimental batches. In International Conference on Learning Representations (ICLR), 2019.
  348. 348.Taylor, M. J., Lukowski, J. K., and Anderton, C. R. Spatially resolved mass spectrometry at the single cell: Recent innovations in proteomics and metabolomics. Journal of the American Society for Mass Spectrometry, 32(4): 872–894, 2021.
  349. 349.Tellez, D., Balkenhol, M., Otte-Höller, I., van de Loo, R., Vogels, R., Bult, P., Wauters, C., Vreuls, W., Mol, S., Karssemeijer, N., et al. Whole-slide mitosis detection in h&e breast histology using phh3 as a reference to train distilled stain-invariant convolutional networks. IEEE Transactions on Medical Imaging, 37(9):2126–2136, 2018.
  350. 350.Tellez, D., Litjens, G., Bándi, P., Bulten, W., Bokhorst, J., Ciompi, F., and van der Laak, J. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology. Medical Image Analysis, 58, 2019.
  351. 351.Temel, D., Lee, J., and AlRegib, G. Cure-or: Challenging unreal and real environments for object recognition. In 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 137–144. IEEE, 2018.
  352. 352.Thorp, K. R., Thompson, A. L., Harders, S. J., French, A. N., and Ward, R. W. High-throughput phenotyping of crop water use efficiency via multispectral drone imagery and a daily soil water balance model. Remote Sensing, 10 (11):1682, 2018.
  353. 353.Tiecke, T. G., Liu, X., Zhang, A., Gros, A., Li, N., Yetman, G., Kilic, T., Murray, S., Blankespoor, B., Prydz, E. B., and Dang, H. H. Mapping the world population one building at a time. arXiv, 2017.
  354. 354.Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In International Conference on Intelligent Robots and Systems (IROS), 2017.
  355. 355.Toda, Y. and Okura, F. How convolutional neural networks diagnose plant disease. Plant Phenomics, 2019, 2019.
  356. 356.Torralba, A. and Efros, A. A. Unbiased look at dataset bias. In Computer Vision and Pattern Recognition (CVPR), pp. 1521–1528, 2011.
  357. 357.Tuschl, T. Rna interference and small interfering rnas. Chembiochem, 2(4):239–245, 2001.
  358. 358.Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., and Darrell, T. Deep domain confusion: Maximizing for domain invariance. arXiv preprint arXiv:1412.3474, 2014.
  359. 359.Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T. Adversarial discriminative domain adaptation. In Computer Vision and Pattern Recognition (CVPR), 2017.
  360. 360.Ubbens, J. R., Ayalew, T. W., Shirtliffe, S., Josuttes, A., Pozniak, C., and Stavness, I. Autocount: Unsupervised segmentation and counting of organs in field images. In European Conference on Computer Vision, pp. 391–399. Springer, 2020.
  361. 361.Uzkent, B. and Ermon, S. Learning when and where to zoom with deep reinforcement learning. In Computer Vision and Pattern Recognition (CVPR), 2020.
  362. 362.Vasic, M., Kanade, A., Maniatis, P., Bieber, D., and Singh, R. Neural program repair by jointly learning to localize and repair. In International Conference on Learning Representations (ICLR), 2019.
  363. 363.Vatnehol, S., Peña, H., and Handegard, N. O. A method to automatically detect fish aggregations using horizontally scanning sonar. ICES Journal of Marine Science, 75(5): 1803–1812, 2018.
  364. 364.Veeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M. Rotation equivariant cnns for digital pathology. In International Conference on Medical Image Computing and Computer-assisted Intervention, pp. 210–218, 2018.
  365. 365.Venkateswara, H., Eusebio, J., Chakraborty, S., and Panchanathan, S. Deep hashing network for unsupervised domain adaptation. In Computer Vision and Pattern Recognition (CVPR), pp. 5018–5027, 2017.
  366. 366.Veta, M., Diest, P. J. V., Jiwa, M., Al-Janabi, S., and Pluim, J. P. Mitosis counting in breast cancer: Object-level interobserver agreement and comparison to an automatic method. PloS one, 11(8), 2016.
  367. 367.Veta, M., Heng, Y. J., Stathonikos, N., Bejnordi, B. E., Beca, F., Wollmann, T., Rohr, K., Shah, M. A., Wang, D., Rousson, M., et al. Predicting breast tumor proliferation from whole-slide images: the tupac16 challenge. Medical image analysis, 54:111–121, 2019.
  368. 368.Volpi, R., Namkoong, H., Sener, O., Duchi, J., Murino, V., and Savarese, S. Generalizing to unseen domains via adversarial data augmentation. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  369. 369.Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems (NeurIPS), 2019a.
  370. 370.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations (ICLR), 2019b.
  371. 371.Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T. Fully test-time adaptation by entropy minimization. arXiv preprint arXiv:2006.10726, 2020a.
  372. 372.Wang, H., Ge, S., Lipton, Z., and Xing, E. P. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems (NeurIPS), 2019c.
  373. 373.Wang, S., Bai, M., Mattyus, G., Chu, H., Luo, W., Yang, B., Liang, J., Cheverie, J., Fidler, S., and Urtasun, R. Torontocity: Seeing the world with a million eyes. In International Conference on Computer Vision (ICCV), 2017.
  374. 374.Wang, S., Chen, W., Xie, S. M., Azzari, G., and Lobell, D. B. Weakly supervised deep learning for segmentation of remote sensing imagery. Remote Sensing, 12, 2020b.
  375. 375.Ward, D. and Moghadam, P. Scalable learning for bridging the species gap in image-based plant phenotyping. Computer Vision and Image Understanding, 197:103009, 2020.
  376. 376.Wearn, O. and Glover-Kapfer, P. Camera-trapping for conservation: a guide to best-practices. WWF conservation technology series, 1(1):2019–04, 2017.
  377. 377.Weinberger, S. Speech accent archive. George Mason University, 2015.
  378. 378.Weinstein, B. G. A computer vision for animal ecology. Journal of Animal Ecology, 87(3):533–545, 2018.
  379. 379.Weinstein, J. N., Collisson, E. A., Mills, G. B., Shaw, K. R. M., Ozenberger, B. A., Ellrott, K., Shmulevich, I., Sander, C., Stuart, J. M., Network, C. G. A. R., et al. The cancer genome atlas pan-cancer analysis project. Nature genetics, 45(10), 2013.
  380. 380.West, R., Paskov, H. S., Leskovec, J., and Potts, C. Exploiting social network structure for person-to-person sentiment analysis. Transactions of the Association for Computational Linguistics (TACL), 2:297–310, 2014.
  381. 381.Widmer, G. and Kubat, M. Learning in the presence of concept drift and hidden contexts. Machine learning, 23(1):69–101, 1996.
  382. 382.Williams, J. J., Kim, J., Rafferty, A., Maldonado, S., Gajos, K. Z., Lasecki, W. S., and Heffernan, N. Axis: Generating explanations at scale with learnersourcing and machine learning. In Proceedings of the Third (2016) ACM Conference on Learning@Scale, pp. 379–388, 2016.
  383. 383.Wilson, B., Hoffman, J., and Morgenstern, J. Predictive inequity in object detection. arXiv preprint arXiv:1902.11097, 2019.
  384. 384.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J. HuggingFace’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  385. 385.Wong, H. Y. F., Lam, H. Y. S., Fong, A. H.-T., Leung, S. T., Chin, T. W.-Y., Lo, C. S. Y., Lui, M. M.-S., Lee, J. C. Y., Chiu, K. W.-H., Chung, T., et al. Frequency and distribution of chest radiographic findings in covid-19 positive patients. Radiology, pp. 201160, 2020.
  386. 386.Worrall, D. E., Garbin, S. J., Turmukhambetov, D., and Brostow, G. J. Harmonic networks: Deep translation and rotation equivariance. In Computer Vision and Pattern Recognition (CVPR), pp. 5028–5037, 2017.
  387. 387.Wu, M., Mosse, M., Goodman, N., and Piech, C. Zero shot learning for code education: Rubric sampling with deep learning inference. In Association for the Advancement of Artificial Intelligence (AAAI), volume 33, pp. 782–790, 2019a.
  388. 388.Wu, M., Davis, R. L., Domingue, B. W., Piech, C., and Goodman, N. Variational item response theory: Fast, accurate, and expressive. International Conference on Educational Data Mining, 2020.
  389. 389.Wu, Y., Winston, E., Kaushik, D., and Lipton, Z. Domain adaptation with asymmetrically-relaxed distribution alignment. In International Conference on Machine Learning (ICML), pp. 6872–6881, 2019b.
  390. 390.Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., and Pande, V. Moleculenet: a benchmark for molecular machine learning. Chemical Science, 9(2):513–530, 2018.
  391. 391.Wulfmeier, M., Bewley, A., and Posner, I. Incremental adversarial domain adaptation for continually changing environments. In International Conference on Robotics and Automation (ICRA), 2018.
  392. 392.Xiao, K., Engstrom, L., Ilyas, A., and Madry, A. Noise or signal: The role of image backgrounds in object recognition. arXiv preprint arXiv:2006.09994, 2020.
  393. 393.Xie, M., Jean, N., Burke, M., Lobell, D., and Ermon, S. Transfer learning from deep features for remote sensing and poverty mapping. In Association for the Advancement of Artificial Intelligence (AAAI), 2016.
  394. 394.Xie, S. M., Kumar, A., Jones, R., Khani, F., Ma, T., and Liang, P. In-N-Out: Pre-training and self-training using auxiliary information for out-of-distribution robustness. arXiv, 2020.
  395. 395.Xiong, H., Cao, Z., Lu, H., Madec, S., Liu, L., and Shen, C. Tasselnetv2: in-field counting of wheat spikes with context-augmented local regression networks. Plant Methods, 15(1):1–14, 2019.
  396. 396.Xu, K., Hu, W., Leskovec, J., and Jegelka, S. How powerful are graph neural networks? In International Conference on Learning Representations (ICLR), 2018.
  397. 397.Yang, Y. and Newsam, S. Bag-of-visual-words and spatial extensions for land-use classification. Geographic Information Systems, 2010.
  398. 398.Yang, Y., Caluwaerts, K., Iscen, A., Zhang, T., Tan, J., and Sindhwani, V. Data efficient reinforcement learning for legged robots. In Conference on Robot Learning (CoRL), 2019.
  399. 399.Yasunaga, M. and Liang, P. Graph-based, self-supervised program repair from diagnostic feedback. In International Conference on Machine Learning (ICML), 2020.
  400. 400.Yeh, C., Perez, A., Driscoll, A., Azzari, G., Tang, Z., Lobell, D., Ermon, S., and Burke, M. Using publicly available satellite imagery and deep learning to understand economic well-being in Africa. Nature Communications, 11, 2020.
  401. 401.You, J., Li, X., Low, M., Lobell, D., and Ermon, S. Deep gaussian process for crop yield prediction based on remote sensing data. In Association for the Advancement of Artificial Intelligence (AAAI), 2017.
  402. 402.Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. BDD100K: A diverse driving dataset for heterogeneous multitask learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  403. 403.Yuval, N., Tao, W., Adam, C., Alessandro, B., Bo, W., and Y, N. A. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011.
  404. 404.Zafar, M. B., Valera, I., Rodriguez, M. G., Gummadi, K. P., and Weller, A. From Parity to Preference-based Notions of Fairness in Classification. arXiv:1707.00010 [cs, stat], Nov 2017. URL http://arxiv.org/abs/1707.00010. arXiv: 1707.00010.
  405. 405.Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., and Oermann, E. K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. In PLOS Medicine, 2018.
  406. 406.Zhang, K., Schölkopf, B., Muandet, K., and Wang, Z. Domain adaptation under target and conditional shift. In International Conference on Machine Learning (ICML), pp. 819–827, 2013.
  407. 407.Zhang, M., Marklund, H., Dhawan, N., Gupta, A., Levine, S., and Finn, C. Adaptive risk minimization: A meta-learning approach for tackling group shift. arXiv preprint arXiv:2007.02931, 2020.
  408. 408.Zhang, Y., Baldridge, J., and He, L. Paws: Paraphrase adversaries from word scrambling. In North American Association for Computational Linguistics (NAACL), 2019.
  409. 409.Zhao, J., Wang, T., Yatskar, M., Ordoñez, V., and Chang, K. Gender bias in coreference resolution: Evaluation and debiasing methods. In North American Association for Computational Linguistics (NAACL), 2018.
  410. 410.Zhou, J. and Troyanskaya, O. G. Predicting effects of noncoding variants with deep learning–based sequence model. Nature Methods, 12(10):931–934, 2015.
  411. 411.Zhou, X., Nie, Y., Tan, H., and Bansal, M. The curse of performance instability in analysis datasets: Consequences, source, and suggestions. arXiv preprint arXiv:2004.13606, 2020.
  412. 412.Zhou, Y., Zhu, S., Cai, C., Yuan, P., Li, C., Huang, Y., and Wei, W. High-throughput screening of a crispr/cas9 library for functional genomics in human cells. Nature, 509(7501):487, 2014.
  413. 413.Zitnick, C. L., Chanussot, L., Das, A., Goyal, S., Heras-Domingo, J., Ho, C., Hu, W., Lavril, T., Palizhati, A., Riviere, M., Shuaibi, M., Sriram, A., Tran, K., Wood, B., Yoon, J., Parikh, D., and Ulissi, Z. An introduction to electrocatalyst design using machine learning for renewable energy storage. arXiv preprint arXiv:2010.09435, 2020.

Citation

MLA
Koh, P. W., et al. “WILDS: A Benchmark of in-the-Wild Distribution Shifts”. arXiv, 2020, http://arxiv.org/abs/2012.07421v3.
APA
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., Lee, T., David, E., Stavness, I., Guo, W., Earnshaw, B. A., Haque, I. S., Beery, S., Leskovec, J., Kundaje, A., … Liang, P. (2020). WILDS: A Benchmark of in-the-Wild Distribution Shifts. arXiv. http://arxiv.org/abs/2012.07421v3
Chicago
Koh, P. W., S. Sagawa, H. Marklund, et al. 2020. “WILDS: A Benchmark of in-the-Wild Distribution Shifts”. arXiv. http://arxiv.org/abs/2012.07421v3.
Harvard
Koh, P.W. et al. (2020) “WILDS: A Benchmark of in-the-Wild Distribution Shifts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2012.07421v3.
Vancouver
1. Koh PW, Sagawa S, Marklund H, et al (2020) WILDS: A Benchmark of in-the-Wild Distribution Shifts. arXiv

BibTeX

@article{koh2020wilds,
  title = {WILDS: A Benchmark of in-the-Wild Distribution Shifts},
  author = {Koh, Pang Wei and Sagawa, Shiori and Marklund, Henrik and Xie, Sang Michael and Zhang, Marvin and Balsubramani, Akshay and Hu, Weihua and Yasunaga, Michihiro and Phillips, Richard Lanas and Gao, Irena and Lee, Tony and David, Etienne and Stavness, Ian and Guo, Wei and Earnshaw, Berton A. and Haque, Imran S. and Beery, Sara and Leskovec, Jure and Kundaje, Anshul and Pierson, Emma and Levine, Sergey and Finn, Chelsea and Liang, Percy},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2012.07421v3},
  eprint = {2012.07421}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/