Underspecification Presents Challenges for Credibility in Modern Machine Learning

Alexander D'AmourKatherine A. HellerDan MoldovanBen AdlamBabak AlipanahiAlex BeutelChristina ChenJonathan DeatonJacob EisensteinMatthew D. Hoffman

article2022JMLR830 citations

Demonstrates how standard machine learning pipelines routinely suffer from underspecification, causing models with identical in-distribution accuracy to exhibit widely divergent failure modes in real-world deployment across computer vision, NLP, and clinical domains.

Listen

Modern machine learning models frequently suffer from unexpected failures, performance degradation, and brittle behavior when deployed in real-world settings, even after achieving state-of-the-art accuracy during standard testing. In conventional workflows, models are validated on held-out test data that follow the exact same statistical distribution as the training data. This evaluation paradigm assumes that identical validation scores imply equivalent real-world utility, obscuring critical behavioral flaws and undermining the credibility of artificial intelligence systems in high-stakes domains such as healthcare and automated reasoning.

The article demonstrates that underspecification in machine learning training pipelines is a pervasive root cause of these deployment failures. A pipeline is underspecified when its design and validation criteria can be satisfied equally well by many distinct predictive models that nevertheless behave radically differently when faced with real-world distribution shifts or stress tests.

To establish how underspecification operates, the article analyzes simple epidemiological, genomic, and theoretical models, and executes an empirical stress-testing protocol across several deployable deep learning domains. The experimental approach trains ensembles of models using identical architectures, hyperparameters, and datasets, varying only arbitrary operational choices such as the initialization random seed. The resulting predictors—which achieve near-identical performance on standard in-distribution validation sets—are then subjected to targeted stress tests, including stratified demographic evaluations, domain distribution shifts, and contrastive input perturbations across computer vision, ophthalmology, dermatology, natural language processing, and electronic health record analysis.

The findings show that underspecification is widespread across model architectures and application areas. In computer vision, models with identical standard accuracy showed variation an order of magnitude larger when tested on corrupted image benchmarks, and scaled-up transfer learning models exhibited five times greater dispersion on natural distribution shifts than on standard validation data. In medical imaging, identically trained models exhibited statistically significant disparities in diagnostic calibration across unobserved camera hardware and variable accuracy across different patient skin types. In natural language processing, varying the random seed in pretraining or fine-tuning caused large fluctuations in reliance on societal stereotypes and syntactic shortcuts, with gender-correlation scores swinging widely from 0.3 to 0.7 despite near-identical benchmark accuracy. In clinical risk prediction from electronic health records, models exhibited unstable sensitivity to operational artifacts—such as the time of day a lab test was ordered—leading to conflicting, flipped clinical alert decisions across different random initializations.

These results demonstrate that substantive real-world behavior, fairness, and safety are routinely dictated by arbitrary training choices rather than deliberate engineering. This ambiguity exposes organizations to serious operational, clinical, and reputational risks, especially when assuming that high validation scores guarantee robustness. Furthermore, the findings show that standard in-distribution evaluations are largely uncorrelated with stress-test performance, confirming that underspecification is distinct from structural domain mismatch and cannot be resolved merely by selecting the model with the highest standard validation score.

To address this challenge, organizations deploying machine learning systems must implement application-specific behavioral stress tests and domain-tailored operational checks before clearance for production. Practitioners should move away from evaluating single model checkpoints and instead evaluate multi-run ensembles to detect underspecified variation. When training systems, teams should integrate domain constraints, causal knowledge, and targeted invariances into the optimization process to restrict the set of acceptable predictors without sacrificing predictive accuracy.

The study's primary limitation lies in its conservative exploration: by focusing primarily on perturbations to random initialization seeds, the analysis likely underestimates the full magnitude of underspecification present in commercial pipelines where architectures, optimizers, and data preprocessing pipelines also vary. Decision-makers can place high confidence in the finding that standard validation metrics do not guarantee reliable deployment, and should exercise strict caution when deploying models that have not undergone explicit, domain-specific stress testing.

arXiv: 2011.03395
  • Paper: Shortcut learning in deep neural networks, Robert Geirhos et al. (2020). This paper establishes the foundational concept of shortcut learning and unintended decision rules in neural networks, which directly underpins how underspecified training pipelines latch onto arbitrary statistical artifacts.
  • Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). This work introduces standard benchmarks and protocols for evaluating subpopulation and domain shifts in high-stakes domains, providing the empirical foundation used to test underspecification across real-world data.
  • Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). By demonstrating that standard empirical risk minimization matches specialized robust algorithms under controlled benchmarks, this paper highlights the baseline model selection ambiguities explored in underspecification analysis.
  • Paper: Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al. (2019). This study analyzes how overparameterized neural networks fail under group distribution shifts despite high in-distribution accuracy, providing essential context for why identical training metrics conceal severe subgroup disparities.
  • Paper: Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles, Balaji Lakshminarayanan et al. (2017). This work introduces deep ensembles across random initializations as a scalable uncertainty estimator, providing the core multi-run methodology recommended by the source to detect underspecified predictive variation.
  • Paper: Do ImageNet Classifiers Generalize to ImageNet?, Benjamin Recht et al. (2019). This paper provides foundational evidence that machine learning models drop in performance under subtle distribution shifts even when standard test benchmarks suggest strong generalization.
  • Paper: The ML test score: A rubric for ML production readiness and technical debt reduction, Eric Breck et al. (2017). This rubric outlines production readiness tests and hidden technical debt in machine learning pipelines, motivating the need for the rigorous behavioral stress-testing framework formalized by the source.
  • Paper: Deep Reinforcement Learning that Matters, Peter Henderson et al. (2018). This study demonstrates how non-deterministic factors like random seeds cause large behavioral variance in empirical models, offering early empirical evidence for the randomness-driven disparities analyzed in underspecification.
Cover for Underspecification Presents Challenges for Credibility in Modern Machine Learning

Abstract

Machine learning (ML) systems often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification in ML pipelines as a key reason for these failures. An ML pipeline is the full procedure followed to train and validate a predictor. Such a pipeline is underspecified when it can return many distinct predictors with equivalently strong test performance. Underspecification is common in modern ML pipelines that primarily validate predictors on held-out data that follow the same distribution as the training data. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment domains. This ambiguity can lead to instability and poor model behavior in practice, and is a distinct failure mode from previously identified issues arising from structural mismatch between training and deployment domains. We provide evidence that underspecification has substantive implications for practical ML pipelines, using examples from computer vision, medical imaging, natural language processing, clinical risk prediction based on electronic health records, and medical genomics. Our results show the need to explicitly account for underspecification in modeling pipelines that are intended for real-world deployment in any domain.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Related Work
  • 2.1 Machine Learning Pipelines
  • 2.2 Underspecification
  • 2.3 Structural and Underspecified Failure Modes
  • 2.4 Stress Tests and Credibility
  • 3. Warm-Up: Underspecification in Simple Models
  • 3.1 Underspecification in a Simple Epidemiological Model
  • 3.2 Underspecification in a Linear Polygenic Risk Score Model
  • 3.3 Theoretical Analysis of Underspecification in a Random Feature Model
  • 4. Empirical Strategy for Probing Underspecification in Deep Learning Pipelines
  • 5. Case Studies in Computer Vision
  • 5.1 ImageNet-C
  • 5.2 ObjectNet
  • 5.3 Conclusions
  • 6. Case Studies in Medical Imaging
  • 6.1 Ophthalmological Imaging
  • 6.2 Dermatological Imaging
  • 6.3 Conclusions
  • 7. Case Study in Natural Language Processing
  • 7.1 Gendered Correlations in Downstream Tasks
  • 7.1.1 Semantic textual similarity (STS)
  • 7.1.2 Pronoun resolution
  • 7.1.3 Gender correlations and underspecification
  • 7.2 Stereotypical Associations in Pretrained Language Models
  • 7.3 Spurious Correlations in Natural Language Inference
  • 7.4 Conclusions
  • 8. Case Study in Clinical Predictions from Electronic Health Records
  • 8.1 Data, Predictor Ensemble, and Metrics
  • 8.2 Reliance on Operational Signals
  • 8.3 Conclusions
  • 9. Discussion: Implications for ML Practice
  • Acknowledgements
  • References
  • Appendix A. Computer Vision: Marginalization versus Model Selection
  • Appendix B. Natural Language Processing
  • B.1 Analysis of Static Embeddings
  • B.2 Exploratory Analysis of Gendered Correlations in STS Task
  • Appendix C. Clinical Prediction with EHR: Additional Details and Supplementary Ablation Experiment
  • C.1 Lab Order Patterns and Time of Day
  • C.2 Details of Predictor Performance on Intervened Data
  • C.3 Preliminary Ablation Experiment
  • Appendix D. Genomics: Full Experimental Details
  • D.1 Background
  • D.2 Methods
  • Appendix E. Random Feature Model: Complete Theoretical Analysis
  • E.1 General Definitions
  • E.2 Random Featurization Maps
  • E.3 Random features model: Risk
  • E.4 Random features model: Sensitivity to random featurization
  • E.5 Random features model: Distribution shift
  • E.6 Random features model: Derivation Eq. (29)

Knowls

  1. Knowl 1 — Underspecification in Machine Learning Pipelines

    definition

    In supervised machine learning, an ML pipeline comprises a training dataset D\mathcal{D} drawn from distribution PP, a hypothesis class F\mathcal{F}, an optimization algorithm, and an evaluation procedure (typically risk minimization evaluated on an independent and identically distributed holdout dataset D′∼P\mathcal{D}' \sim P).

    An ML pipeline is defined as underspecified when its specification admits a non-trivial set of candidate predictors F∗⊆F\mathcal{F}^* \subseteq \mathcal{F} (often termed the Rashomon set) that achieve equivalent validation performance on the training distribution PP, yet differ systematically in how they process inputs to produce outputs. Because standard validation criteria are agnostic to the specific internal mechanisms or signals utilized by a model, arbitrary or opaque design choices—such as random initialization seeds, data ordering, hardware details, or regularization hyperparameters—determine which specific predictor f∈F∗f \in \mathcal{F}^* the pipeline returns. When deployed to target domains with distribution P′≠PP' \neq P or evaluated on behavioral stress tests, predictors within F∗\mathcal{F}^* can exhibit widely divergent generalization performance, fairness characteristics, and robustness.

  2. Knowl 2 — Empirical Protocol for Probing Pipeline Underspecification

    model/method

    To empirically detect and measure underspecification in complex machine learning pipelines where analytical characterization of the validation-equivalent set F∗\mathcal{F}^* is intractable, a perturbation-and-stress-testing protocol is used:

    1. Instantiate an Ensemble via Pipeline Perturbations: Construct an ensemble of models {f1,f2,…,fK}⊂F\{f_1, f_2, \dots, f_K\} \subset \mathcal{F} by executing the identical pipeline specification multiple times, varying only arbitrary degrees of freedom that do not alter the training objective or expected in-distribution (i.i.d.) performance (e.g., random initialization seeds, data shuffling seeds, or pretraining checkpoints).
    2. Verify In-Distribution Equivalence: Confirm that all models in the ensemble achieve approximately identical validation performance on the standard i.i.d. holdout set D′∼P\mathcal{D}' \sim P, establishing that the ensemble samples from the empirical validation-equivalent set F∗\mathcal{F}^*.
    3. Evaluate on Domain-Specific Stress Tests: Probe model behaviors along deployment-critical axes not guaranteed by i.i.d. validation:
      • Shifted performance evaluations: Measure performance on out-of-distribution or corrupted datasets P′≠PP' \neq P.
      • Stratified performance evaluations: Measure performance across specific subgroups Da′={(xi,yi):Ai=a}\mathcal{D}'_a = \{(x_i, y_i) : A_i = a\} defined by auxiliary or demographic attributes AA.
      • Contrastive evaluations: Measure consistency or invariance over matched sets (x,T(x))(x, T(x)) generated by label-preserving transformations TT.
    4. Quantify Underspecification Signatures: Measure the variance in stress test performance across ensemble members and compute correlation coefficients between i.i.d. validation performance and stress test metrics. High between-model variance on stress tests alongside near-zero correlation with i.i.d. metrics confirms that deployment behavior is underspecified by the pipeline.
  3. Knowl 3 — Asymptotic Sensitivity and Adversarial Shift Vulnerability in Overparameterized Random Feature Regression

    theoretical result

    Consider a random features regression model fW(x)=θ^(W)Tσ(Wx)f_W(x) = \hat{\theta}(W)^T \sigma(W x) mapping inputs x∈Rdx \in \mathbb{R}^d to responses y∈Ry \in \mathbb{R}, where W∈RN×dW \in \mathbb{R}^{N \times d} has independent rows wi∼Unif(Sd−1(1))w_i \sim \text{Unif}(S^{d-1}(1)), input covariates satisfy xi∼Unif(Sd−1(d))x_i \sim \text{Unif}(S^{d-1}(\sqrt{d})), and the target response is linear yi=β0Txiy_i = \beta_0^T x_i with ∥β0∥2=r\|\beta_0\|_2 = r. Second-layer weights θ^(W)∈RN\hat{\theta}(W) \in \mathbb{R}^N are obtained via minimum ℓ2\ell_2-norm interpolation (minimizing ∥θ∥2\|\theta\|_2 subject to fW(xi)=yif_W(x_i) = y_i for all i=1,…,ni=1,\dots,n).

    In the proportional asymptotic regime N,n,d→∞N, n, d \to \infty with N/d→ψ1N/d \to \psi_1 and n/d→ψ2n/d \to \psi_2 (with ψ1>ψ2≥1\psi_1 > \psi_2 \ge 1):

    1. Model Orthogonality in-Distribution: Two independently drawn feature matrices W1,W2W_1, W_2 yield predictors with identical in-distribution mean squared error R(W1,P)=R(W2,P)=R(W,P)R(W_1, P) = R(W_2, P) = R(W, P). However, the normalized model sensitivity satisfies: S(W1,W2;P)R(W,P)=EX∼P[(fW1(X)−fW2(X))2]R(W,P)→2,\frac{S(W_1, W_2; P)}{R(W, P)} = \frac{\mathbb{E}_{X \sim P}\left[(f_{W_1}(X) - f_{W_2}(X))^2\right]}{R(W, P)} \to 2, meaning that the residual functions (fW1−f∗)(f_{W_1} - f_*) and (fW2−f∗)(f_{W_2} - f_*) are asymptotically orthogonal in L2(P)L^2(P).
    2. Vulnerability to Adversarial Mean Shifts: For a specific weight draw W0W_0, define a shift x0∈Rdx_0 \in \mathbb{R}^d orthogonal to β0\beta_0 with norm ∥x0∥2=Δ≪∥x∥2\|x_0\|_2 = \Delta \ll \|x\|_2: x0=−ΔPβ0⊥W0Tθ^(W0)∥Pβ0⊥W0Tθ^(W0)∥2,x_0 = -\Delta \frac{P_{\beta_0}^\perp W_0^T \hat{\theta}(W_0)}{\|P_{\beta_0}^\perp W_0^T \hat{\theta}(W_0)\|_2}, where Pβ0⊥=I−β0β0T∥β0∥22P_{\beta_0}^\perp = I - \frac{\beta_0 \beta_0^T}{\|\beta_0\|_2^2}. Under the shifted distribution PW0,ΔP_{W_0, \Delta} (where xtest=x+x0x_{\text{test}} = x + x_0), the prediction risk decomposes as: R(W,PW0,Δ)=R(W,P)+Δ2μ12T(W,W0)+oP(1),R(W, P_{W_0, \Delta}) = R(W, P) + \Delta^2 \mu_1^2 T(W, W_0) + o_P(1), where μ1=E[Gσ(G)]\mu_1 = \mathbb{E}[G \sigma(G)] for standard Gaussian GG, and T(W,W0)=⟨Pβ0⊥WTθ^(W),Pβ0⊥W0Tθ^(W0)⟩∥Pβ0⊥W0Tθ^(W0)∥22.T(W, W_0) = \frac{\langle P_{\beta_0}^\perp W^T \hat{\theta}(W), P_{\beta_0}^\perp W_0^T \hat{\theta}(W_0) \rangle}{\|P_{\beta_0}^\perp W_0^T \hat{\theta}(W_0)\|_2^2}. For W=W0W = W_0, E[T(W0,W0)]\mathbb{E}[T(W_0, W_0)] remains strictly positive, inducing a multi-fold risk increase. For an independent random draw W≠W0W \neq W_0, E[T(W,W0)]→0\mathbb{E}[T(W, W_0)] \to 0, leaving R(W,PW0,Δ)≈R(W,P)R(W, P_{W_0, \Delta}) \approx R(W, P).
  4. Knowl 4 — Underspecification in Image Classification Under Synthetic Corruptions and Natural Shifts

    empirical result

    Deep convolutional image classification pipelines trained on ImageNet exhibit substantial underspecification when evaluated on corrupted (ImageNet-C) and natural distribution-shifted (ObjectNet) test sets, despite achieving virtually identical top-1 accuracy on the standard ImageNet validation set.

    Evaluating an ensemble of 50 ResNet-50 models trained from scratch with different random seeds, and an ensemble of 30 Big Transfer (BiT) ResNet-101x3 models pre-trained on JFT-300M and fine-tuned on ImageNet across different seeds:

    • Variance on Corruptions: Accuracy variation across random seeds on high-severity ImageNet-C corruptions (severity level 5) is up to an order of magnitude larger than the variation observed on the standard i.i.d. validation set. ResNet-50 models show validation accuracy of 0.759±0.0010.759 \pm 0.001, but pixelation accuracy varies at 0.197±0.0240.197 \pm 0.024 and contrast accuracy at 0.091±0.0080.091 \pm 0.008. BiT models show validation accuracy of 0.862±0.0010.862 \pm 0.001, while contrast accuracy is 0.462±0.0190.462 \pm 0.019.
    • Unpredictability from Validation Accuracy: Pearson and Spearman rank correlations between i.i.d. validation accuracy and corruption stress test accuracy contain zero within 95% confidence intervals, demonstrating that standard validation accuracy does not predict out-of-distribution robustness.
    • Disagreement on Natural Shifts: On ObjectNet (which features 113 overlapping classes in varied backgrounds and poses), pairwise prediction disagreement reaches 50.9%50.9\% for ResNet-50 and 25.3%25.3\% for BiT (compared to 16.0%16.0\% and 6.4%6.4\% on standard ImageNet). Permutation tests show that predictor accuracy variance on ObjectNet is statistically non-random (p=0.002p = 0.002 for ResNet-50, p<0.001p < 0.001 for BiT).
    Dataset / Stress Test ImageNet (i.i.d.) Pixelate (sev. 5) Contrast (sev. 5) Motion Blur (sev. 5) Brightness (sev. 5) ObjectNet
    ResNet-50 Top-1 Accuracy 0.759 (0.001) 0.197 (0.024) 0.091 (0.008) 0.100 (0.007) 0.607 (0.003) 0.259 (0.002)
    BiT ResNet-101x3 Top-1 Accuracy 0.862 (0.001) 0.555 (0.008) 0.462 (0.019) 0.515 (0.008) 0.723 (0.002) 0.520 (0.005)

    Values reported are the ensemble mean (and standard deviation across random seeds) of top-1 accuracy proportions.

  5. Knowl 5 — Underspecification and Shortcut Learning Across Random Seeds in Pretrained and Fine-Tuned NLP Models

    empirical result

    Transformer language models (BERT-Large, 340M parameters) exhibit large variance in their reliance on spurious shortcuts and societal stereotypes across different random pretraining and fine-tuning seeds, despite invariant accuracy on standard NLP benchmarks.

    Evaluating an ensemble of 5 distinct pretraining runs of BERT-Large and 20 fine-tuning runs per checkpoint (100 models total):

    1. Semantic Textual Similarity (STS-B) and Gender Bias: Fine-tuned models (iid correlation 0.870.87--0.900.90) vary widely in their alignment with U.S. Bureau of Labor Statistics (BLS) occupational gender proportions (Spearman ρ\rho with BLS data ranges from 0.300.30 to 0.700.70). Pretraining seed significantly affects gender bias variance (F=9.66,p=1×10−6F = 9.66, p = 1\times 10^{-6}), while fine-tuning task accuracy does not predict gender correlation (Spearman ρ=0.21\rho = 0.21, 95% CI: [0.00,0.40][0.00, 0.40]).
    2. Pronoun Resolution (OntoNotes): Model accuracy is tightly constrained (0.9600.960--0.9650.965), yet gender correlation spans 0.260.26 to 0.510.51, strongly driven by the pretraining seed (F=7.91,p=2×10−5F = 7.91, p = 2\times 10^{-5}).
    3. StereoSet Pretrained Bias: Evaluating the 5 pretrained checkpoints on StereoSet reveals an Idealized Context Association Test (ICAT) score range of 3.353.35 points across random seeds alone—larger than the performance gap across different model architectures on the public benchmark leaderboard.
    4. Natural Language Inference (MNLI) Stress Tests: Across 100 fine-tuned models on MNLI (matched validation accuracy tightly constrained between 83.4%83.4\% and 84.4%84.4\%), performance on heuristic stress tests (such as HANS lexical overlap and Naik et al. stress suites) shows large spreads (e.g., spelling error accuracy ranges from ∼76%\sim 76\% to 81%81\%, F=25.11,p=3×10−14F = 25.11, p = 3\times 10^{-14}). Stress test performance exhibits weak rank correlation with MNLI matched validation accuracy (Spearman ρ∈[−0.07,0.43]\rho \in [-0.07, 0.43] across tests).
  6. Knowl 6 — Underspecification in Clinical Risk Prediction from Electronic Health Records

    empirical result

    In clinical risk prediction from longitudinal Electronic Health Record (EHR) data, recurrent neural networks (RNNs) predicting acute kidney injury (AKI) within a 48-hour window rely heavily on operational artifacts (time of day of lab orders and test panel composition), with the degree of reliance underspecified by the training pipeline.

    Using de-identified EHR data from 703,782 patients from the US Department of Veterans Affairs, an ensemble of 15 RNN models (5 random seeds across SRU, LSTM, and UGRNN architectures) was trained using 6-hour aggregated time buckets:

    • Baseline Equivalence: All models achieved tightly constrained in-distribution normalized Precision-Recall Area Under the Curve (normalized PRAUC) between 34.59%34.59\% and 36.61%36.61\% (baseline test set AKI prevalence: 2.269%2.269\%).
    • Operational Perturbations: When subjected to a time-of-day shift (shifting 6-hour buckets by a constant offset) and lab-panel subsetting (restricting orders strictly to the basic metabolic panel CHEM-7), normalized PRAUC dropped to 32.09%32.09\%--36.08%36.08\% on creatinine-sampled timepoints and exhibited wide between-model dispersion.
    • Decision Flipping Across Seeds: Two LSTM models differing only by random initialization seed flipped clinical risk decisions (crossing the calibrated decision threshold) on over 1,500 patient-timepoints under operational shifts, but the specific patient-timepoints whose decisions flipped were largely disjoint across seeds (only 26% to 35% intersection in flipped decisions).
    • Feature Ablation: An LSTM trained with the timestamp feature completely ablated achieved an identical in-distribution normalized PRAUC of 0.3680.368, confirming that high i.i.d. performance does not require learning operational timing shortcuts.
  7. Knowl 7 — Underspecification in Medical Imaging Across Hardware Variations and Demographic Strata

    empirical result

    Deep convolutional neural networks (Inception-V4) trained for clinical diagnosis using standard i.i.d. validation pipelines exhibit significant underspecification when evaluated across unseen imaging devices and patient demographic subgroups:

    1. Ophthalmology (Diabetic Retinopathy Detection):

      • Models trained on retinal fundus photographs from EyePACS and Indian clinics achieve consistent area under the ROC curve (AUC) on camera types present during training.
      • When evaluated on a camera type held out from training (n=287n = 287), models differing only by the random initialization seed at the fine-tuning stage exhibit substantial AUC variance. A two-sample zz-test comparing AUC standard deviations between the held-out camera test set and the standard validation set (n=3,712n = 3,712) yields z=2.47z = 2.47 (p=0.007p = 0.007).
      • Calibration curves across seeds remain identical for training camera types 1–4, but diverge markedly for the held-out camera type 5.
    2. Dermatology (Skin Condition Classification):

      • An ensemble of 10 Inception-V4 models fine-tuned from ImageNet on clinical skin images showed stable overall top-1 accuracy.
      • When stratified across Fitzpatrick skin types (I to VI), model accuracy variance was significantly elevated in specific subgroups, notably Fitzpatrick skin type IV (n=798,19.6%n = 798, 19.6\% of test set), where permutation tests confirmed systematic, non-sampling-noise variance across random initialization seeds (p=0.03p = 0.03).
  8. Knowl 8 — Underspecification via Feature Collinearity in Polygenic Risk Scores

    empirical result

    In medical genomics, Polygenic Risk Score (PRS) pipelines exhibit severe underspecification due to linkage disequilibrium (LD)—the strong collinearity between nearby single-nucleotide polymorphisms (SNPs).

    When constructing linear regression models to predict intraocular pressure (IOP) using demographic covariates and 129 clumped genomic feature clusters (derived from a genome-wide association study of 4,054 SNPs on the UK Biobank):

    • In-Distribution Equivalence: An ensemble of 1,000 PRS models—each constructed by selecting one representative SNP per cluster uniformly at random, including the index SNP chosen by the standard PLINK heuristic—achieves nearly identical normalized mean squared error (NMSE) on the British training set (n=82,309n = 82,309) and the i.i.d. British validation set (n=9,662n = 9,662).
    • Divergence Across Populations: When evaluated on a genetically shifted non-British evaluation cohort (n=14,898n = 14,898), NMSE displays wide dispersion across ensemble members.
    • Lack of Generalization Correlation: Model performance on the British validation set is weakly correlated with performance on the non-British set (Spearman rank correlation ρ=0.135\rho = 0.135, 95% CI: [0.070,0.200][0.070, 0.200]). The standard PLINK heuristic does not outperform random cluster representatives on the shifted population.
  9. Knowl 9 — Parameter Underspecification in Epidemiological SIR Dynamical Models

    empirical result

    In dynamical epidemic modeling using the Susceptible-Infected-Recovered (SIR) differential equation system: dSdt=−β(IN)S,dIdt=−ID+β(IN)S,dRdt=ID,\frac{dS}{dt} = -\beta \left(\frac{I}{N}\right) S, \quad \frac{dI}{dt} = -\frac{I}{D} + \beta \left(\frac{I}{N}\right) S, \quad \frac{dR}{dt} = \frac{I}{D}, where β\beta is the transmission rate, DD is the average infectious duration, and NN is total population size, estimating parameters from early-stage epidemic data is fundamentally underspecified.

    During early epidemic phases (t≤Tobst \le T_{\text{obs}}), S(t)≈NS(t) \approx N, causing infections to grow exponentially according to dIdt≈(β−1/D)I\frac{dI}{dt} \approx (\beta - 1/D)I. Consequently, training by minimizing squared-error loss on observed early infections identifies only the net growth rate g=β−1/Dg = \beta - 1/D, admitting an infinite continuum of parameter pairs (β,D)(\beta, D) that fit the training data equally well.

    When these fitted models are used to forecast epidemic trajectories beyond TobsT_{\text{obs}}, the choice of initial parameter value D0D_0 in gradient-based optimization determines the resulting (β,D)(\beta, D), producing peak infection forecasts that differ by orders of magnitude despite identical in-sample fit.

  10. Knowl 10 — Model Selection Versus Marginalization Under Pipeline Underspecification

    empirical result

    While ensembling (marginalizing predictions over a set of validation-equivalent models F∗\mathcal{F}^*) consistently improves in-distribution (i.i.d.) accuracy by reducing model variance, ensembling does not reliably outperform individual model selection on out-of-distribution stress tests when behavioral variance across F∗\mathcal{F}^* is high.

    Evaluating ensembles of 50 ResNet-50 models on ImageNet image classification demonstrates that:

    1. On the i.i.d. ImageNet test set, averaging predictions across a small subset of models quickly surpasses the accuracy of the best individual model in the ensemble.
    2. On stress tests with moderate performance variance across seeds (e.g., ObjectNet and ImageNet-C contrast), a larger ensemble is required to surpass the performance of the best single model.
    3. On stress tests with extreme performance variance across seeds (e.g., ImageNet-C pixelation corruption), the average of all 50 models never reaches the performance of the best single member of F∗\mathcal{F}^*.

    Because F∗\mathcal{F}^* contains models that rely on brittle or erroneous shortcuts, naive marginalization incorporates these failure modes rather than eliminating them, showing that explicit model selection or constrained training is required to resolve underspecification.

  11. Knowl 11 — Underspecification in Word Embedding Association Tests Across Random Seeds

    empirical result

    Static word embedding algorithms (word2vec 500-dimensional skip-gram/CBOW models trained on news and Wikipedia corpora) exhibit underspecification with respect to demographic associations measured by the Word Embedding Association Test (WEAT).

    For target word sets X,YX, Y and attribute word sets A,BA, B, the normalized WEAT effect size measures relative association via: s(X,Y,A,B)=∑x∈Xs(x,A,B)−∑y∈Ys(y,A,B),s(X, Y, A, B) = \sum_{x \in X} s(x, A, B) - \sum_{y \in Y} s(y, A, B), where s(w,A,B)=meana∈Acos⁡(w,a)−meanb∈Bcos⁡(w,b)s(w, A, B) = \text{mean}_{a \in A} \cos(w, a) - \text{mean}_{b \in B} \cos(w, b).

    In an ensemble of 20 word2vec models trained with identical hyperparameters and varying only the initialization random seed:

    • All models achieved consistent performance on standard word analogy tasks (76.2%76.2\%--76.7%76.7\% accuracy).
    • On WEAT evaluations of gender associations, all random seeds consistently showed strong, statistically significant associations (p<0.01p < 0.01).
    • On WEAT evaluations of racial associations, scores varied substantially across random seeds, with the statistical significance (p<0.01p < 0.01 via permutation tests) of racial bias appearing or disappearing depending entirely on the random training seed.

Coverage note — Intermediate algebraic derivations of the Stieltjes transform asymptotics for random feature models in Appendix E and exhaustive per-task numerical breakdowns for all individual sub-benchmarks in Naik et al. and ImageNet-C were omitted in favor of the core analytical results and representative summary statistics.

References

  1. 1.Adewole S Adamson and Avery Smith. Machine learning and health care disparities in dermatology. JAMA dermatology, 154(11):1247–1248, 2018.
  2. 2.Ademide Adelekun, Ginikanwa Onyekaba, and Jules B Lipoff. Skin color in dermatology textbooks: An updated evaluation and analysis. Journal of the American Academy of Dermatology, 2020.
  3. 3.R Ambrosino, B G Buchanan, G F Cooper, and M J Fine. The use of misclassification costs to learn rule-based decision support models for cost-effective hospital admission strategies. Proceedings. Symposium on Computer Applications in Medical Care, pages 304–8, 1995. ISSN 0195-4210. URL http://www.ncbi.nlm.nih.gov/pubmed/8563290http://www.pubmedcentral.nih.gov/articlerender.fcgi?artid=PMC2579104.
  4. 4.Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019.
  5. 5.Susan Athey. Beyond prediction: Using big data for policy problems. Science, 355(6324):483–485, 2017.
  6. 6.Marzieh Babaeianjelodar, Stephen Lorenz, Josh Gordon, Jeanna Matthews, and Evan Freitag. Quantifying gender bias in different corpora. In Companion Proceedings of the Web Conference 2020, pages 752–759, 2020.
  7. 7.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems, pages 9448–9458, 2019.
  8. 8.Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis. A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1–12, 2020.
  9. 9.Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine learning and the bias-variance trade-off. arXiv preprint arXiv:1812.11118, 2018.
  10. 10.Emily M. Bender and Alexander Koller. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.463. URL https://www.aclweb.org/anthology/2020.acl-main.463.
  11. 11.Yoshua Bengio. The consciousness prior. arXiv preprint arXiv:1709.08568, 2017.
  12. 12.Jeremy J Berg, Arbel Harpak, Nasa Sinnott-Armstrong, Anja Moltke Joergensen, Hakhamanesh Mostafavi, Yair Field, Evan August Boyle, Xinjun Zhang, Fernando Racimo, Jonathan K Pritchard, and Graham Coop. Reduced signal for polygenic adaptation of height in UK biobank. Elife, 8, March 2019.
  13. 13.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in neural information processing systems, pages 4349–4357, 2016.
  14. 14.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1075. URL https://www.aclweb.org/anthology/D15-1075.
  15. 15.Kendrick Boyd, Vítor Santos Costa, Jesse Davis, and C. David Page. Unachievable region in precision-recall space and its effect on empirical evaluation. In Proceedings of the 29th International Conference on Machine Learning, ICML 2012, volume 1, pages 639–646, 2012. ISBN 9781450312851.
  16. 16.Leo Breiman. Statistical modeling: The two cultures (with comments and a rejoinder by the author). Statistical science, 16(3):199–231, 2001.
  17. 17.Theodora S. Brisimi, Tingting Xu, Taiyao Wang, Wuyang Dai, and Ioannis Ch Paschalidis. Predicting diabetes-related hospitalizations based on electronic health records. Statistical Methods in Medical Research, 28(12):3667–3682, dec 2019. ISSN 14770334. doi: 10.1177/0962280218810911.
  18. 18.Joy Buolamwini and Timnit Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency, pages 77–91, 2018.
  19. 19.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186, 2017.
  20. 20.CARDIoGRAMplusC4D Consortium, Panos Deloukas, Stavroula Kanoni, Christina Willenborg, Martin Farrall, Themistocles L Assimes, John R Thompson, Erik Ingelsson, Danish Saleheen, Jeanette Erdmann, Benjamin A Goldstein, Kathleen Stirrups, Inke R König, Jean-Baptiste Cazier, Asa Johansson, Alistair S Hall, Jong-Young Lee, Cristen J Willer, John C Chambers, Tõnu Esko, Lasse Folkersen, Anuj Goel, Elin Grundberg, Aki S Havulinna, Weang K Ho, Jemma C Hopewell, Niclas Eriksson, Marcus E Kleber, Kati Kristiansson, Per Lundmark, Leo-Pekka Lyytikäinen, Suzanne Rafelt, Dmitry Shungin, Rona J Strawbridge, Gudmar Thorleifsson, Emmi Tikkanen, Natalie Van Zuydam, Benjamin F Voight, Lindsay L Waite, Weihua Zhang, Andreas Ziegler, Devin Absher, David Altshuler, Anthony J Balmforth, Inãs Barroso, Peter S Braund, Christof Burgdorf, Simone Claudi-Boehm, David Cox, Maria Dimitriou, Ron Do, DIAGRAM Consortium, CARDIOGENICS Consortium, Alex S F Doney, Noureddine El Mokhtari, Per Eriksson, Krista Fischer, Pierre Fontanillas, Anders Franco-Cereceda, Bruna Gigante, Leif Groop, Stefan Gustafsson, Jörg Hager, Göran Hallmans, Bok-Ghee Han, Sarah E Hunt, Hyun M Kang, Thomas Illig, Thorsten Kessler, Joshua W Knowles, Genovefa Kolovou, Johanna Kuusisto, Claudia Langenberg, Cordelia Langford, Karin Leander, Marja-Liisa Lokki, Anders Lundmark, Mark I McCarthy, Christa Meisinger, Olle Melander, Evelin Mihailov, Seraya Maouche, Andrew D Morris, Martina Müller-Nurasyid, MuTHER Consortium, Kjell Nikus, John F Peden, N William Rayner, Asif Rasheed, Silke Rosinger, Diana Rubin, Moritz P Rumpf, Arne Schäfer, Mohan Sivananthan, Ci Song, Alexandre F R Stewart, Sian-Tsung Tan, Gudmundur Thorgeirsson, C Ellen van der Schoot, Peter J Wagner, Wellcome Trust Case Control Consortium, George A Wells, Philipp S Wild, Tsun-Po Yang, Philippe Amouyel, Dominique Arveiler, Hanneke Basart, Michael Boehnke, Eric Boerwinkle, Paolo Brambilla, Francois Cambien, Adrienne L Cupples, Ulf de Faire, Abbas Dehghan, Patrick Diemert, Stephen E Epstein, Alun Evans, Marco M Ferrario, Jean Ferrières, Dominique Gauguier, Alan S Go, Alison H Goodall, Villi Gudnason, Stanley L Hazen, Hilma Holm, Carlos Iribarren, Yangsoo Jang, Mika Kähönen, Frank Kee, Hyo-Soo Kim, Norman Klopp, Wolfgang Koenig, Wolfgang Kratzer, Kari Kuulasmaa, Markku Laakso, Reijo Laaksonen, Ji-Young Lee, Lars Lind, Willem H Ouwehand, Sarah Parish, Jeong E Park, Nancy L Pedersen, Annette Peters, Thomas Quertermous, Daniel J Rader, Veikko Salomaa, Eric Schadt, Svati H Shah, Juha Sinisalo, Klaus Stark, Kari Stefansson, David-Alexandre Trégouët, Jarmo Virtamo, Lars Wallentin, Nicholas Wareham, Martina E Zimmermann, Markku S Nieminen, Christian Hengstenberg, Manjinder S Sandhu, Tomi Pastinen, Ann-Christine Syvänen, G Kees Hovingh, George Dedoussis, Paul W Franks, Terho Lehtimäki, Andres Metspalu, Pierre A Zalloua, Agneta Siegbahn, Stefan Schreiber, Samuli Ripatti, Stefan S Blankenberg, Markus Perola, Robert Clarke, Bernhard O Boehm, Christopher O’Donnell, Muredach P Reilly, Winfried März, Rory Collins, Sekar Kathiresan, Anders Hamsten, Jaspal S Kooner, Unnur Thorsteinsdottir, John Danesh, Colin N A Palmer, Robert Roberts, Hugh Watkins, Heribert Schunkert, and Nilesh J Samani. Large-scale association analysis identifies new risk loci for coronary artery disease. Nat. Genet., 45(1):25–33, January 2013.
  21. 21.Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. Intelligible Models for HealthCare. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD ’15, pages 1721–1730, 2015. ISBN 9781450336642. doi: 10.1145/2783258.2788613. URL http://dx.doi.org/10.1145/2783258.2788613http://dl.acm.org/citation.cfm?doid=2783258.2788613.
  22. 22.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/S17-2001. URL https://www.aclweb.org/anthology/S17-2001.
  23. 23.Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. Entropy-sgd: Biasing gradient descent into wide valleys. Journal of Statistical Mechanics: Theory and Experiment, 2019(12):124018, 2019.
  24. 24.Lenaic Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. 2019.
  25. 25.Edward Choi, Mohammad Taha Bahadori, Le Song, Walter F Stewart, and Jimeng Sun. GRAM: Graph-based Attention Model for Healthcare Representation Learning. 2017. doi: 10.1145/3097983.3098126. URL http://dx.doi.org/10.1145/3097983.3098126.
  26. 26.Gary S Collins, Johannes B Reitsma, Douglas G Altman, and Karel GM Moons. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod): the tripod statement. British Journal of Surgery, 102(3):148–158, 2015.
  27. 27.Jasmine Collins, Jascha Sohl-Dickstein, and David Sussillo. Capacity and trainability in recurrent neural networks. In 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings, 2017.
  28. 28.Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 120–128, 2019.
  29. 29.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009.
  30. 30.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, 2019.
  31. 31.Josip Djolonga, Jessica Yung, Michael Tschannen, Rob Romijnders, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Matthias Minderer, Alexander D’Amour, Dan Moldovan, et al. On robustness and transferability of convolutional neural networks. arXiv preprint arXiv:2007.08558, 2020.
  32. 32.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305, 2020.
  33. 33.L Duncan, H Shen, B Gelaye, J Meijsen, K Ressler, M Feldman, R Peterson, and B Domingue. Analysis of polygenic risk score usage and performance in diverse human populations. Nat. Commun., 10(1):3328, July 2019.
  34. 34.Michael W Dusenberry, Dustin Tran, Edward Choi, Jonas Kemp, Jeremy Nixon, Ghassen Jerfel, Katherine Heller, and Andrew M Dai. Analyzing the role of model uncertainty for electronic health records. In Proceedings of the ACM Conference on Health, Inference, and Learning, pages 204–213, 2020.
  35. 35.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. Journal of Machine Learning Research, 20(55):1–21, 2019. URL http://jmlr.org/papers/v20/18-598.html.
  36. 36.Andre Esteva, Brett Kuprel, Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017.
  37. 37.Chenchen Feng, David Le, and Allison B. McCoy. Using Electronic Health Records to Identify Adverse Drug Events in Ambulatory Care: A Systematic Review. Applied Clinical Informatics, 10(1):123–128, 2019. ISSN 18690327. doi: 10.1055/s-0039-1677738.
  38. 38.Aaron Fisher, Cynthia Rudin, and Francesca Dominici. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177):1–81, 2019.
  39. 39.TB Fitzpatrick. Sun and skin. Journal de Medecine Esthetique, 2:33–34, 1975.
  40. 40.Seth Flaxman, Swapnil Mishra, Axel Gandy, H Juliette T Unwin, Thomas A Mellan, Helen Coupland, Charles Whittaker, Harrison Zhu, Tresnia Berah, Jeffrey W Eaton, et al. Estimating the effects of non-pharmaceutical interventions on covid-19 in europe. Nature, 584(7820):257–261, 2020.
  41. 41.Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep ensembles: A loss landscape perspective. arXiv preprint arXiv:1912.02757, 2019.
  42. 42.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin. Linear mode connectivity and the lottery ticket hypothesis. In Proceedings of the 37th International Conference on Machine Learning, 2020.
  43. 43.Joseph Futoma, Morgan Simons, Trishan Panch, Finale Doshi-Velez, and Leo Anthony Celi. The myth of generalisability in clinical research and machine learning in health care. The Lancet Digital Health, 2(9):e489 – e492, 2020. ISSN 2589-7500. doi: https://doi.org/10.1016/S2589-7500(20)30186-2. URL http://www.sciencedirect.com/science/article/pii/S2589750020301862.
  44. 44.Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. Counterfactual fairness in text classification through robustness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 219–226, 2019.
  45. 45.Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. In Advances in Neural Information Processing Systems, pages 8789–8798, 2018.
  46. 46.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bygh9j09KX.
  47. 47.Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. arXiv preprint arXiv:2004.07780, 2020.
  48. 48.Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. Deep learning, volume 1. MIT press Cambridge, 2016.
  49. 49.Varun Gulshan, Lily Peng, Marc Coram, Martin C Stumpe, Derek Wu, Arunachalam Narayanaswamy, Subhashini Venugopalan, Kasumi Widner, Tom Madams, Jorge Cuadros, et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. Jama, 316(22):2402–2410, 2016.
  50. 50.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  51. 51.Christina Heinze-Deml, Jonas Peters, and Nicolai Meinshausen. Invariant causal prediction for nonlinear models. Journal of Causal Inference, 6(2), 2018.
  52. 52.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HJz6tiCqYm.
  53. 53.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. arXiv preprint arXiv:1907.07174, 2019.
  54. 54.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. arXiv preprint arXiv:2006.16241, 2020.
  55. 55.Sepp Hochreiter and Jürgen Schmidhuber. Long Short-Term Memory. Neural Computation, 9(8):1735–1780, 1997. URL http://www7.informatik.tu-muenchen.de/{~}hochreithttp://www.idsia.ch/{~}juergen.
  56. 56.Wolfgang Hoffmann, Ute Latza, Sebastian E Baumeister, Martin Brünger, Nina Buttmann-Schweiger, Juliane Hardt, Verena Hoffmann, André Karch, Adrian Richter, Carsten Oliver Schmidt, et al. Guidelines and recommendations for ensuring good epidemiological practice (gep): a guideline developed by the german society for epidemiology. European journal of epidemiology, 34(3):301–317, 2019.
  57. 57.Sara Hooker. The hardware lottery. arXiv preprint arXiv:2009.06489, 2020.
  58. 58.Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. OntoNotes: The 90% solution. In Proceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers, pages 57–60, New York City, USA, June 2006. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/N06-2015.
  59. 59.Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, 2018.
  60. 60.Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. Adversarial examples are not bugs, they are features. In Advances in Neural Information Processing Systems, pages 125–136, 2019.
  61. 61.International Schizophrenia Consortium, Shaun M Purcell, Naomi R Wray, Jennifer L Stone, Peter M Visscher, Michael C O’Donovan, Patrick F Sullivan, and Pamela Sklar. Common polygenic variation contributes to risk of schizophrenia and bipolar disorder. Nature, 460(7256):748–752, August 2009.
  62. 62.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407, 2018.
  63. 63.Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai. arXiv preprint arXiv:2010.07487, 2020.
  64. 64.Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. Learning the difference that makes a difference with counterfactually-augmented data. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=Sklgs0NFvr.
  65. 65.John A Kellum and Azra Bihorac. Artificial intelligence to predict aki: is it a breakthrough? Nature Reviews Nephrology, pages 1–2, 2019.
  66. 66.Christopher J Kelly, Alan Karthikesalingam, Mustafa Suleyman, Greg Corrado, and Dominic King. Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17(1):195, 2019.
  67. 67.Amit V Khera, Mark Chaffin, Krishna G Aragam, Mary E Haas, Carolina Roselli, Seung Hoan Choi, Pradeep Natarajan, Eric S Lander, Steven A Lubitz, Patrick T Ellinor, and Sekar Kathiresan. Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations. Nat. Genet., 50(9):1219–1224, September 2018.
  68. 68.Arif Khwaja. KDIGO clinical practice guidelines for acute kidney injury. Nephron - Clinical Practice, 120(4), oct 2012. ISSN 16602110. doi: 10.1159/000339789.
  69. 69.Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan, and Ziad Obermeyer. Prediction policy problems. American Economic Review, 105(5):491–95, 2015.
  70. 70.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Large scale learning of general visual representations for transfer. arXiv preprint arXiv:1912.11370, 2019.
  71. 71.Jonathan Krause, Varun Gulshan, Ehsan Rahimy, Peter Karth, Kasumi Widner, Greg S Corrado, Lily Peng, and Dale R Webster. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology, 125(8):1264–1272, 2018.
  72. 72.Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. Counterfactual fairness. In Advances in neural information processing systems, pages 4066–4076, 2017.
  73. 73.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 6402–6413. Curran Associates, Inc., 2017. URL http://papers.nips.cc/paper/7219-simple-and-scalable-predictive-uncertainty-estimation-using-deep-ensembles.pdf.
  74. 74.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.
  75. 75.Olivier Ledoit and Sandrine Péché. Eigenvectors of some large sample covariance matrix ensembles. Probability Theory and Related Fields, 151(1-2):233–264, 2011.
  76. 76.Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, and Yoav Artzi. Simple recurrent units for highly parallelizable recurrence. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMzhou 2018, pages 4470–4481. Association for Computational Linguistics, sep 2018. ISBN 9781948087841. doi: 10.18653/v1/d18-1477. URL http://arxiv.org/abs/1709.02755.
  77. 77.Tal Linzen. How can we accelerate progress towards human-like linguistic generalization? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5210–5217, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.465. URL https://www.aclweb.org/anthology/2020.acl-main.465.
  78. 78.Xiaoxuan Liu, Samantha Cruz Rivera, David Moher, Melanie J Calvert, and Alastair K Denniston. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the consort-ai extension. bmj, 370, 2020a.
  79. 79.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  80. 80.Yuan Liu, Ayush Jain, Clara Eng, David H Way, Kang Lee, Peggy Bui, Kimberly Kanada, Guilherme de Oliveira Marinho, Jessica Gallegos, Sara Gabriele, et al. A deep learning system for differential diagnosis of skin diseases. Nature Medicine, pages 1–9, 2020b.
  81. 81.Sara Magliacane, Thijs van Ommen, Tom Claassen, Stephan Bongers, Philip Versteeg, and Joris M Mooij. Domain adaptation by using causal inference to predict invariant conditional distributions. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31 (NeurIPS2018), pages 10869–10879. Curran Associates, Inc., 2018. URL http://papers.nips.cc/paper/8282-domain-adaptation-by-using-causal-inference-to-predict-invariant-conditional-distributions.pdf.
  82. 82.Maggie Makar, Ben Packer, Dan Moldovan, Davis Blalock, Yoni Halpern, and Alexander D’Amour. Causally-motivated shortcut removal using auxiliary labels. arXiv preprint arXiv:2105.06422, 2021.
  83. 83.Alicia R Martin, Christopher R Gignoux, Raymond K Walters, Genevieve L Wojcik, Benjamin M Neale, Simon Gravel, Mark J Daly, Carlos D Bustamante, and Eimear E Kenny. Human demographic history impacts genetic risk prediction across diverse populations. Am. J. Hum. Genet., 100(4):635–649, April 2017.
  84. 84.Alicia R Martin, Masahiro Kanai, Yoichiro Kamatani, Yukinori Okada, Benjamin M Neale, and Mark J Daly. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet., 51(4):584–591, April 2019.
  85. 85.Charles T Marx, Flavio du Pin Calmon, and Berk Ustun. Predictive multiplicity in classification. arXiv preprint arXiv:1909.06677, 2019.
  86. 86.R Thomas McCoy, Junghyun Min, and Tal Linzen. Berts of a feather do not generalize together: Large variability in generalization across models with similar test set performance. arXiv preprint arXiv:1911.02969, 2019a.
  87. 87.R Thomas McCoy, Ellie Pavlick, and Tal Linzen. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. arXiv preprint arXiv:1902.01007, 2019b.
  88. 88.Song Mei and Andrea Montanari. The generalization error of random features regression: Precise asymptotics and double descent curve. arXiv:1908.05355, 2019.
  89. 89.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  90. 90.Joannella Morales, Danielle Welter, Emily H Bowler, Maria Cerezo, Laura W Harris, Aoife C McMahon, Peggy Hall, Heather A Junkins, Annalisa Milano, Emma Hastings, Cinzia Malangone, Annalisa Buniello, Tony Burdett, Paul Flicek, Helen Parkinson, Fiona Cunningham, Lucia A Hindorff, and Jacqueline A L MacArthur. A standardized framework for representation of ancestry data in genomics studies, with application to the NHGRI-EBI GWAS catalog. Genome Biol., 19(1):21, February 2018.
  91. 91.Sendhil Mullainathan and Jann Spiess. Machine learning: an applied econometric approach. Journal of Economic Perspectives, 31(2):87–106, 2017.
  92. 92.Moin Nadeem, Anna Bethke, and Siva Reddy. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456, 2020.
  93. 93.Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. Stress test evaluation for natural language inference. arXiv preprint arXiv:1806.00692, 2018.
  94. 94.Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=B1g5sA4twr.
  95. 95.National Institute for Health and Care Excellence (NICE). Acute kidney injury: prevention, detection and management. NICE Guideline NG148, 2019.
  96. 96.Radford M Neal. Priors for infinite networks. In Bayesian Learning for Neural Networks, pages 29–53. Springer, 1996.
  97. 97.Anna C Need and David B Goldstein. Next generation disparities in human genomics: concerns and remedies. Trends Genet., 25(11):489–494, November 2009.
  98. 98.Bret Nestor, Matthew B. A. McDermott, Willie Boag, Gabriela Berner, Tristan Naumann, Michael C Hughes, Anna Goldenberg, and Marzyeh Ghassemi. Feature Robustness in Non-stationary Health Records: Caveats to Deployable Model Performance in Common Clinical Machine Learning Tasks. Proceedings of Machine Learning Research, 106:1–23, 2019. URL https://mimic.physionet.org/mimicdata/carevue/http://arxiv.org/abs/1908.00690.
  99. 99.Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In Proceedings of the ACM Conference on Health, Inference, and Learning, pages 151–159, 2020.
  100. 100.Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, oct 2019. ISSN 10959203. doi: 10.1126/science.aax2342.
  101. 101.Cecilia Panigutti, Alan Perotti, and Dino Pedreschi. Doctor XAI An ontology-based approach to black-box sequential data classification explanations. In FAT 2020 - Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency*, pages 629–639, 2020. ISBN 9781450369367. doi: 10.1145/3351095.3372855. URL https://doi.org/10.1145/3351095.3372855.
  102. 102.Jonas Peters, Peter Bñhlmann, and Nicolai Meinshausen. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78(5):947–1012, 2016. doi: 10.1111/rssb.12167. URL https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/rssb.12167.
  103. 103.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, 2018.
  104. 104.Alice B Popejoy and Stephanie M Fullerton. Genomics is failing on diversity. Nature, 538(7624):161–164, October 2016.
  105. 105.Mihail Popescu and Mohammad Khalilia. Improving disease prediction using ICD-9 ontological features. In IEEE International Conference on Fuzzy Systems, pages 1805–1809, 2011. ISBN 9781424473175. doi: 10.1109/FUZZY.2011.6007410.
  106. 106.Alkes L Price, Nick J Patterson, Robert M Plenge, Michael E Weinblatt, Nancy A Shadick, and David Reich. Principal components analysis corrects for stratification in genome-wide association studies. Nat. Genet., 38(8):904–909, August 2006.
  107. 107.Shaun Purcell, Benjamin Neale, Kathe Todd-Brown, Lori Thomas, Manuel A R Ferreira, David Bender, Julian Maller, Pamela Sklar, Paul I W de Bakker, Mark J Daly, and Pak C Sham. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet., 81(3):559–575, September 2007.
  108. 108.Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang. Understanding and mitigating the tradeoff between robustness and accuracy. arXiv preprint arXiv:2002.10716, 2020.
  109. 109.Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. In Advances in neural information processing systems, pages 1177–1184, 2008.
  110. 110.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.442. URL https://www.aclweb.org/anthology/2020.acl-main.442.
  111. 111.Samantha Cruz Rivera, Xiaoxuan Liu, An-Wen Chan, Alastair K Denniston, and Melanie J Calvert. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the spirit-ai extension. bmj, 370, 2020.
  112. 112.Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. Right for the right reasons: training differentiable models by constraining their explanations. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, pages 2662–2670. AAAI Press, 2017.
  113. 113.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 8–14, 2018.
  114. 114.Bernhard Schölkopf. Causality for machine learning. arXiv preprint arXiv:1911.10500, 2019.
  115. 115.Montgomery Slatkin. Linkage disequilibrium — understanding the evolutionary past and mapping the medical future. Nature Reviews Genetics, 9:477–485, 2008.
  116. 116.Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D Sculley, Joshua Dillon, Jie Ren, and Zachary Nado. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems, pages 13969–13980, 2019.
  117. 117.Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, Paul Downey, Paul Elliott, Jane Green, Martin Landray, Bette Liu, Paul Matthews, Giok Ong, Jill Pell, Alan Silman, Alan Young, Tim Sprosen, Tim Peakman, and Rory Collins. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med., 12(3):e1001779, March 2015.
  118. 118.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision, pages 843–852, 2017.
  119. 119.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-first AAAI conference on artificial intelligence, 2017.
  120. 120.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. arXiv preprint arXiv:2007.00644, 2020.
  121. 121.Daniel Shu Wei Ting, Carol Yim-Lui Cheung, Gilbert Lim, Gavin Siew Wei Tan, Nguyen D Quang, Alfred Gan, Haslina Hamzah, Renata Garcia-Franco, Ian Yew San Yeo, Shu Yen Lee, et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. Jama, 318(22):2211–2223, 2017.
  122. 122.Nenad Tomašev, Xavier Glorot, Jack W Rae, Michal Zielinski, Harry Askham, Andre Saraiva, Anne Mottram, Clemens Meyer, Suman Ravuri, Ivan Protsyuk, Alistair Connell, Cían O Hughes, Alan Karthikesalingam, Julien Cornebise, Hugh Montgomery, Geraint Rees, Chris Laing, Clifton R Baker, Kelly Peterson, Ruth Reeves, Demis Hassabis, Dominic King, Mustafa Suleyman, Trevor Back, Christopher Nielson, Joseph R Ledsam, and Shakir Mohamed. A clinically applicable approach to continuous prediction of future acute kidney injury. Nature, 572(7767):116–119, aug 2019a. ISSN 0028-0836. doi: 10.1038/s41586-019-1390-1.
  123. 123.Nenad Tomašev, Xavier Glorot, Jack W. Rae, Michal Zielinski, Harry Askham, Andre Saraiva, Anne Mottram, Clemens Meyer, Suman Ravuri, Ivan Protsyuk, Alistair Connell, Cian O. Hugues, Alan Kathikesalingam, Julien Cornebise, Hugh Montgomery, Geraint Rees, Chris Laing, Clifton R. Baker, Kelly Peterson, Ruth Reeves, Demis Hassabis, Dominic King, Mustafa Suleyman, Trevor Back, Christopher Nielson, Joseph R. Ledsam, and Shakir Mohamed. Developing Deep Learning Continuous Risk Models for Early Adverse Event Prediction in Electronic Health Records: an AKI Case Study. PROTOCOL available at Protocol Exchange, version 1, jul 2019b. doi: 10.21203/RS.2.10083/V1.
  124. 124.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  125. 125.Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. Counterfactual invariance to spurious correlations: Why and how to pass stress tests. arXiv preprint arXiv:2106.00545, 2021.
  126. 126.Bjarni J Vilhjálmsson, Jian Yang, Hilary K Finucane, Alexander Gusev, Sara Lindström, Stephan Ripke, Giulio Genovese, Po-Ru Loh, Gaurav Bhatia, Ron Do, Tristan Hayeck, Hong-Hee Won, Schizophrenia Working Group of the Psychiatric Genomics Consortium, Discovery, Biology, and Risk of Inherited Variants in Breast Cancer (DRIVE) study, Sekar Kathiresan, Michele Pato, Carlos Pato, Rulla Tamimi, Eli Stahl, Noah Zaitlen, Bogdan Pasaniuc, Gillian Belbin, Eimear E Kenny, Mikkel H Schierup, Philip De Jager, Nikolaos A Patsopoulos, Steve McCarroll, Mark Daly, Shaun Purcell, Daniel Chasman, Benjamin Neale, Michael Goddard, Peter M Visscher, Peter Kraft, Nick Patterson, and Alkes L Price. Modeling linkage disequilibrium increases accuracy of polygenic risk scores. Am. J. Hum. Genet., 97(4):576–592, October 2015.
  127. 127.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, 2018.
  128. 128.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems, pages 10506–10518, 2019.
  129. 129.Haohan Wang, Xindi Wu, Zeyi Huang, and Eric P Xing. High-frequency component helps explain the generalization of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8684–8694, 2020.
  130. 130.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032, 2020.
  131. 131.Florian Wenzel, Jasper Snoek, Dustin Tran, and Rodolphe Jenatton. Hyperparameter ensembles for robustness and uncertainty quantification. arXiv preprint arXiv:2006.13570, 2020.
  132. 132.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, 2018.
  133. 133.Andrew Gordon Wilson and Pavel Izmailov. Bayesian deep learning and a probabilistic perspective of generalization. arXiv preprint arXiv:2002.08791, 2020.
  134. 134.Julia K Winkler, Christine Fink, Ferdinand Toberer, Alexander Enk, Teresa Deinlein, Rainer Hofmann-Wellenhof, Luc Thomas, Aimilios Lallas, Andreas Blum, Wilhelm Stolz, et al. Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition. JAMA dermatology, 155(10):1135–1141, 2019.
  135. 135.Naomi R Wray, Michael E Goddard, and Peter M Visscher. Prediction of individual genetic risk to disease from genome-wide association studies. Genome Res., 17(10):1520–1528, October 2007.
  136. 136.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems, pages 5753–5763, 2019.
  137. 137.Dong Yin, Raphael Gontijo Lopes, Jon Shlens, Ekin Dogus Cubuk, and Justin Gilmer. A fourier perspective on model robustness in computer vision. In Advances in Neural Information Processing Systems, pages 13255–13265, 2019.
  138. 138.Bin Yu et al. Stability. Bernoulli, 19(4):1484–1500, 2013.
  139. 139.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. Swag: A large-scale adversarial dataset for grounded commonsense inference. arXiv preprint arXiv:1808.05326, 2018.
  140. 140.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.
  141. 141.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, 2018.
  142. 142.Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. The curse of performance instability in analysis datasets: Consequences, source, and suggestions. arXiv preprint arXiv:2004.13606, 2020.

Citation

MLA
D'Amour, A., et al. “Underspecification Presents Challenges for Credibility in Modern Machine Learning”. arXiv, 2020, http://arxiv.org/abs/2011.03395v2.
APA
D'Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., Hormozdiari, F., Houlsby, N., Hou, S., Jerfel, G., Karthikesalingam, A., Lucic, M., Ma, Y., McLean, C., Mincu, D., … Sculley, D. (2020). Underspecification Presents Challenges for Credibility in Modern Machine Learning. arXiv. http://arxiv.org/abs/2011.03395v2
Chicago
D'Amour, A., K. Heller, D. Moldovan, et al. 2020. “Underspecification Presents Challenges for Credibility in Modern Machine Learning”. arXiv. http://arxiv.org/abs/2011.03395v2.
Harvard
D'Amour, A. et al. (2020) “Underspecification Presents Challenges for Credibility in Modern Machine Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2011.03395v2.
Vancouver
1. D'Amour A, Heller K, Moldovan D, et al (2020) Underspecification Presents Challenges for Credibility in Modern Machine Learning. arXiv

BibTeX

@article{damour2020underspecification,
  title = {Underspecification Presents Challenges for Credibility in Modern Machine Learning},
  author = {D'Amour, Alexander and Heller, Katherine and Moldovan, Dan and Adlam, Ben and Alipanahi, Babak and Beutel, Alex and Chen, Christina and Deaton, Jonathan and Eisenstein, Jacob and Hoffman, Matthew D. and Hormozdiari, Farhad and Houlsby, Neil and Hou, Shaobo and Jerfel, Ghassen and Karthikesalingam, Alan and Lucic, Mario and Ma, Yian and McLean, Cory and Mincu, Diana and Mitani, Akinori and Montanari, Andrea and Nado, Zachary and Natarajan, Vivek and Nielson, Christopher and Osborne, Thomas F. and Raman, Rajiv and Ramasamy, Kim and Sayres, Rory and Schrouff, Jessica and Seneviratne, Martin and Sequeira, Shannon and Suresh, Harini and Veitch, Victor and Vladymyrov, Max and Wang, Xuezhi and Webster, Kellie and Yadlowsky, Steve and Yun, Taedong and Zhai, Xiaohua and Sculley, D.},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2011.03395v2},
  eprint = {2011.03395}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/