Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR

Sandra WachterBrent MittelstadtChris Russell

article2017Harvard Journal of Law & Technology3,341 citations

Proposes counterfactual explanations as a practical method to satisfy GDPR transparency requirements, showing affected individuals the minimal input changes needed to reverse automated decisions without disclosing proprietary model mechanics.

Listen

The paper examines how individuals can receive meaningful explanations for automated decisions under the EU General Data Protection Regulation, despite the absence of a legally binding right to explanation and the technical difficulty of revealing the internal logic of complex machine-learning models. Current GDPR provisions require only high-level information about intended processing before decisions occur and offer limited practical support for understanding specific outcomes, contesting them, or changing behavior to achieve better results in the future.

The authors set out to demonstrate that counterfactual explanations—simple statements of the form “if variable V had value X instead of Y, the decision would have been different”—can meet these three aims without requiring disclosure of model internals or trade secrets. They ground the approach in philosophical accounts of knowledge and justification, then show how such explanations can be generated efficiently for standard classifiers by optimizing a distance metric (weighted L1 norm) that favors sparse, human-interpretable changes while holding other variables fixed.

Examples computed on the LSAT admissions dataset and the Pima diabetes database illustrate that a small number of minimal changes to input variables suffice to produce a different outcome. These statements can be rendered directly in plain language and remain valid even when the underlying model contains millions of interdependent parameters. Legal analysis of Articles 12–15 and Recital 71 confirms that counterfactuals satisfy the GDPR’s transparency requirements while avoiding the narrow applicability conditions and potential conflicts with other rights that hinder a full right to explanation.

The approach therefore offers data subjects actionable information at lower regulatory cost to controllers. It enables verification of data accuracy, identification of grounds for contest, and limited guidance on future changes without exposing proprietary algorithms or infringing the privacy of others. Because the method works on existing differentiable models and requires only modest additional computation, it can be deployed immediately.

Limitations remain. The examples assume variable independence and do not incorporate causal structure; distance metrics must still be tuned to domain-specific notions of mutability and relevance. Counterfactuals also supply evidence about individual decisions rather than the statistical patterns needed to audit systemic bias. Nonetheless, the concrete demonstrations and alignment with GDPR text provide strong support for treating unconditional counterfactual explanations as a practical first step toward greater accountability in automated decision-making.

arXiv: 1711.00399
  • Paper: A Survey of Methods for Explaining Black Box Models, Riccardo Guidotti et al. (2018). Reading this comprehensive taxonomy of explainability methods first provides the essential classification framework needed to understand where unconditional counterfactual explanations fit within the broader landscape of black-box remediation.
  • Paper: Towards A Rigorous Science of Interpretable Machine Learning, Finale Doshi-Velez et al. (2017). Understanding this foundational work on the rigorous science of interpretable machine learning provides the necessary evaluation criteria for judging when and why recourse explanations are genuinely useful to humans.
  • Paper: The Mythos of Model Interpretability, Zachary C. Lipton (2016). This critical examination of model interpretability clarifies the ambiguous terminology surrounding black-box models, setting a precise conceptual baseline before exploring actionable recourse mechanisms.
Cover for Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR

Abstract

There has been much discussion of the right to explanation in the EU General Data Protection Regulation, and its existence, merits, and disadvantages. Implementing a right to explanation that opens the black box of algorithmic decision-making faces major legal and technical barriers. Explaining the functionality of complex algorithmic decision-making systems and their rationale in specific cases is a technically challenging problem. Some explanations may offer little meaningful information to data subjects, raising questions around their value. Explanations of automated decisions need not hinge on the general public understanding how algorithmic systems function. Even though such interpretability is of great importance and should be pursued, explanations can, in principle, be offered without opening the black box. Looking at explanations as a means to help a data subject act rather than merely understand, one could gauge the scope and content of explanations according to the specific goal or action they are intended to support. From the perspective of individuals affected by automated decision-making, we propose three aims for explanations: (1) to inform and help the individual understand why a particular decision was reached, (2) to provide grounds to contest the decision if the outcome is undesired, and (3) to understand what would need to change in order to receive a desired result in the future, based on the current decision-making model. We assess how each of these goals finds support in the GDPR. We suggest data controllers should offer a particular type of explanation, unconditional counterfactual explanations, to support these three aims. These counterfactual explanations describe the smallest change to the world that can be made to obtain a desirable outcome, or to arrive at the closest possible world, without needing to explain the internal logic of the system.

Table of Contents

  • I. INTRODUCTION
  • II. COUNTERFACTUALS
  • A. HISTORIC CONTEXT AND THE PROBLEM OF KNOWLEDGE
  • C. ADVERSARIAL PERTURBATIONS AND COUNTERFACTUAL EXPLANATIONS
  • D. CAUSALITY AND FAIRNESS
  • III. GENERATING COUNTERFACTUALS
  • Equation 3
  • Equation 4
  • A. LSAT DATASET
  • Equation 5
  • B. PIMA DIABETES DATABASE
  • *C. CAUSAL ASSUMPTIONS AND COUNTERFACTUAL EXPLANATIONS*
  • IV. ADVANTAGES OF COUNTERFACTUAL EXPLANATIONS
  • V. COUNTERFACTUAL EXPLANATIONS AND THE GDPR
  • A. EXPLANATIONS TO UNDERSTAND DECISIONS
  • 2. Understanding through counterfactuals
  • *B. EXPLANATIONS TO CONTEST DECISIONS*
  • 1. Contesting through counterfactuals
  • C. EXPLANATIONS TO ALTER FUTURE DECISIONS
  • CONCLUSION
  • APPENDIX 1: SIMPLE LOCAL MODELS AS EXPLANATIONS
  • APPENDIX 2: EXAMPLE TRANSPARENCY INFOGRAPHIC

Knowls

  1. Knowl 1 — Definition and Conceptual Structure of Counterfactual Explanations

    definition

    A counterfactual explanation of an automated algorithmic decision is defined as a statement of the form:

    "Score pp was returned because variables VV had values (v1,v2,… )(v_1, v_2, \dots). If VV instead had values (v1′,v2′,… )(v_1', v_2', \dots), and all other variables had remained constant, score p′p' would have been returned."

    Unlike traditional interpretability methods that attempt to convey the internal parameters, feature weights, or logic of a complex "black box" model, a counterfactual explanation describes the model's output as a dependency on external input facts. It identifies the smallest required alteration to an individual's input attributes to obtain an alternative, desired outcome p′p'.

    Conceptually grounded in the philosophical analysis of knowledge (specifically Nozick and Sosa's modal sensitivity condition: if qq were false, subject SS would not believe proposition pp), counterfactual explanations identify a "close possible world" in feature space that alters the outcome without requiring causal knowledge of the world or disclosure of the internal model architecture.

  2. Knowl 2 — Objective Function for Counterfactual Explanation Generation

    equation

    Let fw:Rd→Rf_w: \mathbb{R}^d \to \mathbb{R} denote a differentiable predictive model (such as a neural network or regression classifier) parameterized by fixed weights ww. For an original input vector xi∈Rdx_i \in \mathbb{R}^d and a target classification or prediction outcome y′∈Ry' \in \mathbb{R}, a counterfactual data point x′∈Rdx' \in \mathbb{R}^d is determined by solving the optimization problem:

    arg⁡min⁡x′max⁡λ[λ(fw(x′)−y′)2+d(xi,x′)]\arg\min_{x'} \max_{\lambda} \left[ \lambda \left(f_w(x') - y'\right)^2 + d(x_i, x') \right]

    where d(xi,x′)d(x_i, x') is a distance metric quantifying the cost or feature-space distance between the original observation xix_i and the candidate counterfactual x′x', and λ>0\lambda > 0 is an adaptive penalty multiplier balancing proximity to the target prediction y′y' against proximity to the factual instance xix_i.

  3. Knowl 3 — Median Absolute Deviation Weighted Manhattan Distance Metric

    equation

    To evaluate feature distance d(xi,x′)d(x_i, x') between an original data vector xix_i and a candidate counterfactual x′x' across a set of features FF, the L1L_1 norm (Manhattan distance) weighted by the inverse Median Absolute Deviation (MADk\text{MAD}_k) of each feature k∈Fk \in F is employed:

    d(xi,x′)=∑k∈F∣xi,k−xk′∣MADkd(x_i, x') = \sum_{k \in F} \frac{|x_{i,k} - x_k'|}{\text{MAD}_k}

    where MADk\text{MAD}_k is computed over the reference dataset PP of sample points as:

    MADk=medianj∈P(∣Xj,k−medianl∈P(Xl,k)∣)\text{MAD}_k = \text{median}_{j \in P} \left( \left| X_{j,k} - \text{median}_{l \in P}(X_{l,k}) \right| \right)

    Here, Xj,kX_{j,k} denotes the value of feature kk for the jj-th data sample in PP. The L1L_1 norm induces sparsity in the solution vector, ensuring that only a minimal subset of input variables are modified. Weighting by MADk−1\text{MAD}_k^{-1} scales feature perturbations according to their empirical variability while maintaining robustness against outliers.

  4. Knowl 4 — Gradient-Based Optimization for Sparse Counterfactual Discovery

    algorithm

    Counterfactual generation optimizes candidate point x′x' via gradient descent while adaptively scaling the target constraint penalty λ\lambda, handling discrete features by clamping, and bounding continuous variables within the valid data range.

    Input: Factual data vector xix_i, differentiable model fwf_w, target prediction y′y', feature scale vector MAD\text{MAD}, learning rate α\alpha, initial penalty λ0\lambda_0, maximum penalty λmax⁡\lambda_{\max}, growth factor γ>1\gamma > 1, iteration limit TT, tolerance ϵ\epsilon, bounds [l,u][\mathbf{l}, \mathbf{u}], discrete feature indices DD.
    Output: Counterfactual instance x∗x^*.
    Initialize x′←xi+δx' \leftarrow x_i + \delta with random noise δ\delta
    λ←λ0\lambda \leftarrow \lambda_0
    while λ≤λmax⁡\lambda \leq \lambda_{\max} and ∣fw(x′)−y′∣>ϵ|f_w(x') - y'| > \epsilon do
        for t=1t = 1 to TT do
            L←λ(fw(x′)−y′)2+∑k∣xi,k−xk′∣MADk\mathcal{L} \leftarrow \lambda (f_w(x') - y')^2 + \sum_{k} \frac{|x_{i,k} - x'_k|}{\text{MAD}_k}
            g←∇x′Lg \leftarrow \nabla_{x'} \mathcal{L}
            x′←x′−α⋅ADAM(g)x' \leftarrow x' - \alpha \cdot \text{ADAM}(g)
            for each continuous feature index k∉Dk \notin D do
                xk′←max⁡(lk,min⁡(uk,xk′))x'_k \leftarrow \max(l_k, \min(u_k, x'_k))
            end for
        end for
        λ←λ⋅γ\lambda \leftarrow \lambda \cdot \gamma
    end while
    for each discrete feature index k∈Dk \in D do
        Evaluate separate optimization runs clamping xk′x'_k to each admissible discrete level and retain the candidate that minimizes total distance d(xi,x′)d(x_i, x')
    end for
    return x∗x^*

    To discover diverse alternative counterfactuals ("close possible worlds"), the algorithm is executed from multiple distinct random initializations, returning the local minima corresponding to varied viable feature modifications.

  5. Knowl 5 — Tripartite Functional Aims of Counterfactual Explanations

    model/method

    Counterfactual explanations are designed to serve three core objectives for affected data subjects without exposing internal algorithmic mechanisms:

    1. Understanding: Provides individual-specific insight into the external facts and decisive features that led to an automated decision.
    2. Contesting: Identifies the specific input variables that drove an adverse outcome, enabling the data subject to audit those specific data points for factual errors or to challenge decisions influenced by protected attributes (e.g., race or gender).
    3. Altering Future Decisions (Actionable Recourse): Offers explicit guidance on how an individual can modify their circumstances (e.g., increasing income by a specific amount) to obtain a favorable outcome upon reapplication under the current decision model.
  6. Knowl 6 — Unconditional Counterfactual Explanations Under Data Protection Law

    theoretical result

    Under the European Union General Data Protection Regulation (GDPR, Articles 12, 13–15, 22, and Recital 71), data subjects have no legally binding right to post-hoc individual explanations of complex black-box logic. Articles 13–15 mandate only ex-ante general information about intended system functionality and processing logic, while Article 22 safeguards apply strictly to solely automated decisions with legal or similarly significant effects.

    Providing unconditional counterfactual explanations upon request—regardless of whether a decision is solely automated, positive, or negative—fulfills the transparency and contestability goals of data protection law. This approach avoids key legal barriers to algorithmic transparency: it does not reveal proprietary source code or trade secrets (Article 15(4) and Recital 63), does not compromise the privacy of other individuals in the training data, and provides actionable recourse to lay subjects without requiring comprehension of high-dimensional parameter spaces.

  7. Knowl 7 — Empirical Comparison of Distance Metrics and Discrete Clamping on LSAT Admissions Data

    empirical result

    A 3-layer neural network (two hidden layers of 20 neurons, 941 weights) trained on the Law School Admission Council (LSAT) dataset was evaluated to find counterfactuals achieving a target average predicted first-year grade (y′=0y'=0). The empirical comparison demonstrated three key behaviors:

    1. Distance Metric Normalization: Unnormalized Euclidean distance (L2L_2) disproportionately altered grade point average (GPA) because GPA has a smaller numerical variance than LSAT scores. Normalizing L2L_2 by feature standard deviation, or using MAD-weighted L1L_1, yielded sparse counterfactuals where GPA remained constant and only LSAT was perturbed.
    2. Discrete Feature Clamping: Unconstrained gradient optimization assigned nonsensical fractional or negative values to binary race indicators (Race ∈{0,1}\in \{0, 1\}). Clamping race to binary values during separate runs and selecting the minimal distance produced valid discrete outputs.
    3. Detection of Bias: For Black applicants (Race = 1), counterfactual optimization consistently shifted race to White (Race = 0) to achieve the target score, revealing direct model dependence on a protected demographic attribute.
    Original Data Continuous Normalized L1L_1 Hybrid Normalized L1L_1
    Sample GPA LSAT Race GPA LSAT Race GPA LSAT Race
    Person 1 3.1 39.0 0 (White) 3.1 35.0 0.1 3.1 34.0 0 (White)
    Person 2 3.7 48.0 0 (White) 3.7 33.5 0.0 3.7 32.4 0 (White)
    Person 3 3.3 28.0 1 (Black) 3.3 34.4 0.1 3.3 33.5 0 (White)
    Person 4 2.4 28.5 1 (Black) 2.4 39.3 0.2 2.4 35.8 0 (White)
    Person 5 2.7 18.3 0 (White) 2.7 35.8 0.1 2.7 34.9 0 (White)

    The hybrid normalized L1L_1 results translate directly into human-interpretable text: for Person 3, "If your LSAT was 33.5, and you were 'White', you would have an average predicted score (0)."

  8. Knowl 8 — Sparse Counterfactual Generation for Pima Indian Diabetes Risk Prediction

    empirical result

    On the Pima Indian Diabetes dataset, a 3-layer neural network with two 20-neuron hidden layers evaluated diabetes onset risk in [0,1][0, 1] from 8 clinical features (including number of pregnancies, age, BMI, 2-hour serum insulin, and plasma glucose concentration). Features were clipped to valid ranges observed in the training distribution.

    Applying MAD-weighted L1L_1 counterfactual optimization to reach a target risk score of y′=0.51y'=0.51 produced sparse modifications altering only 1 or 2 clinical variables while holding all remaining variables constant. For instance:

    • Person 1: "If your 2-Hour serum insulin level was 154.3, you would have a score of 0.51."
    • Person 2: "If your 2-Hour serum insulin level was 169.5, you would have a score of 0.51."
    • Person 3: "If your Plasma glucose concentration was 158.3 and your 2-Hour serum insulin level was 160.5, you would have a score of 0.51."
  9. Knowl 9 — Instability and Scale-Dependence of Local Linear Approximations vs. Counterfactuals

    model/method

    Local surrogate methods (such as LIME) approximate black-box models by fitting simple linear models within a localized sampling window around an input point. These methods suffer from severe scale instability: varying the sampling bandwidth or input domain scale across a non-linear decision boundary causes the direction and magnitude of the estimated linear coefficients to fluctuate wildly or reverse sign.

    Furthermore, linear surrogate models extrapolate poorly and cannot accurately guide recourse. When querying how to reduce an output score below a threshold on a non-linear function, linear approximations can predict that negative scores occur at points that are actually local maxima or regions of high positive values. In contrast, counterfactual explanation optimization searches directly along the non-linear decision surface fw(x′)=y′f_w(x') = y', identifying exact, valid input coordinates that achieve the target score.

  10. Knowl 10 — Assumptions and Limitations of Optimization-Based Counterfactual Explanations

    limitation

    Counterfactual explanations generated via loss minimization exhibit three main theoretical and practical limitations:

    1. Lack of Structural Causal Modeling: The method assumes features can be perturbed independently. In real-world domains with complex causal structures, changing an intervention variable (such as career) inherently induces shifts in dependent variables (such as income). Without an underlying causal graph, the generated counterfactuals may propose physically or temporally implausible combinations.
    2. Asymmetry in Discrimination and Fairness Auditing: If a generated counterfactual alters a protected attribute (e.g., race or gender), it establishes that the individual decision is dependent on that attribute. However, the converse does not hold: a counterfactual that leaves a protected attribute unchanged does not prove that the model is fair or that the attribute was irrelevant across the population.
    3. Manifold Divergence in High-Dimensional Spaces: In continuous, high-dimensional spaces with complex manifolds (e.g., natural images), gradient-based perturbations can generate adversarial solutions that satisfy fw(x′)=y′f_w(x') = y' but lie slightly off the valid data manifold, yielding counterfactuals that are visually imperceptible to humans rather than realistic possible worlds.

Coverage note — No substantial contributed material was omitted. The knowls cover the formal definition, loss formulation, MAD-weighted L1 metric, optimization algorithm, tripartite functional aims, GDPR legal analysis, empirical experiments on LSAT and Pima datasets, surrogate model comparisons from Appendix 1, and stated limitations.

Citation

MLA
Wachter, S., et al. “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”. Harvard Journal of Law & Technology, 2018, 2017, http://arxiv.org/abs/1711.00399v3.
APA
Wachter, S., Mittelstadt, B., & Russell, C. (2017). Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR. Harvard Journal of Law & Technology, 2018. http://arxiv.org/abs/1711.00399v3
Chicago
Wachter, S., B. Mittelstadt, and C. Russell. 2017. “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR”. Harvard Journal of Law & Technology, 2018. http://arxiv.org/abs/1711.00399v3.
Harvard
Wachter, S., Mittelstadt, B. and Russell, C. (2017) “Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR”, Harvard Journal of Law & Technology, 2018 [Preprint]. Available at: http://arxiv.org/abs/1711.00399v3.
Vancouver
1. Wachter S, Mittelstadt B, Russell C (2017) Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR. Harvard Journal of Law & Technology, 2018

BibTeX

@article{wachter2017counterfactual,
  title = {Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR},
  author = {Wachter, Sandra and Mittelstadt, Brent and Russell, Chris},
  year = {2017},
  journal = {Harvard Journal of Law & Technology, 2018},
  url = {http://arxiv.org/abs/1711.00399v3},
  eprint = {1711.00399}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors