Linear Adversarial Concept Erasure

Shauli RavfogelMichael TwitonYoav GoldbergRyan Cotterell

article2022ICML112 citations

Develops a constrained minimax framework and convex relaxation method, R-LACE, to identify and remove linear concept subspaces from neural representations, providing closed-form and tractable solutions for post-hoc bias mitigation without sacrificing downstream performance.

Listen

Modern artificial intelligence models rely heavily on pre-trained text and image representations that learn patterns without direct supervision. As these representations are integrated into critical real-world systems, an urgent operational and ethical challenge has emerged: pre-trained models frequently encode sensitive concepts, such as demographic attributes and social biases, that decision-makers cannot easily control or remove. Post-hoc debiasing methods attempt to neutralize this unwanted information from fixed, pre-trained vectors, but existing approaches often damage model utility by removing too many dimensions or fail to eliminate the target concept completely.

The article aims to formulate a principled mathematical framework for concept erasure that prevents linear predictors from recovering unwanted information while preserving the original representation's useful content as much as possible. It demonstrates the theoretical guarantees of this approach and evaluates its practical effectiveness in removing binary gender bias and visual attributes.

To achieve this, the article frames concept erasure as a linear minimax game between a predictor attempting to recover the concept and an adversary that projects the data onto an orthogonal subspace to block prediction. The authors derive closed-form mathematical solutions for linear regression and Rayleigh quotient objectives. For classification problems, they introduce Relaxed Linear Adversarial Concept Erasure (R-LACE), a convex relaxation that optimizes the adversarial projection matrix. They evaluated the framework across static word embeddings (GloVe), deep contextual language models (BERT fine-tuned on hundreds of thousands of online biography profiles), and image datasets (CelebA face images).

The analysis produced several key findings. First, R-LACE completely neutralized linear gender information in word embeddings using only a single-dimensional subspace projection, dropping linear classifier accuracy from 100% to roughly 50% (random chance), whereas previous iterative methods failed to reach majority accuracy even after removing 20 dimensions. Second, in deep language model evaluations, removing a single dimension via R-LACE reduced gender prediction accuracy from nearly 97% to approximately 55% while maintaining core profession classification accuracy (85.09% versus 85.12% unmitigated), avoiding the substantial performance drops observed with prior approaches. Third, the method substantially mitigated biased word associations in standard benchmark tests and reduced gender clustering without degrading overall semantic word similarity. Finally, adversarial training baselines failed to prevent post-hoc extraction of gender information, highlighting the superior reliability of explicit post-hoc linear projection.

These findings indicate that organizations can mitigate sensitive attributes from pre-trained representations with minimal computational overhead and near-zero degradation in primary task performance. By restricting the intervention to linear projections, the approach preserves interpretability and transparency, enabling practitioners to audit the specific subspaces being removed. This provides a low-risk, practical mechanism to enhance fairness and regulatory compliance in production systems without retraining entire foundational models.

Decision-makers should consider adopting linear adversarial concept erasure as a lightweight post-processing step for linear decision heads and embedding representations where linear leakage of protected attributes is a compliance concern. However, because the method specifically neutralizes linear predictors, technical teams must ensure that erased representations are only fed directly into linear classifiers (such as the final classification layer of a network) and not into subsequent deep, non-linear layers that can still extract non-linear signals. Furthermore, since the relationship between internal concept removal and broader downstream fairness metrics remains complex, organizations should conduct context-specific evaluations across varied fairness metrics rather than treating linear concept erasure as a complete bias solution.

The findings are supported by rigorous theoretical proofs and consistent empirical evaluations across multiple domains. Nevertheless, users should remain aware of key limitations: the method is designed strictly against linear adversaries, does not guarantee protection against deep non-linear probing, and was primarily tested on binary representations of attributes like gender. Further research and validation are required before extending these techniques to complex, non-linear concept removal settings.

Cover for Linear Adversarial Concept Erasure

Abstract

Machine learning models can leak sensitive information about their training data. Concept erasure aims to remove undesirable concepts—such as protected attributes—from representation spaces while preserving performance downstream. We study concept erasure through the lens of adversarial robustness, introducing LEACE, a closed-form method achieving linear-concept erasure provably...

Table of Contents

  • 1. Introduction
  • 2. Linear Minimax Games
  • 2.1. Notation and Generalized Linear Modeling
  • 2.2. The Linear Bias Subspace Hypothesis
  • 2.3. Linear Minimax Games
  • 3. Solving the Linear Minimax Game
  • 3.1. Linear Regression
  • 3.2. Rayleigh Quotient Maximization
  • 3.3. Classification
  • 3.4. R-LACE: A Convex Relaxation
  • 4. Relation to INLP
  • 4.1. Linear Regression
  • 4.2. Rayleigh quotient losses
  • 4.3. Classification
  • 5. Experiments
  • 5.1. Static Word Vectors
  • 5.2. Deep Classification
  • 5.3. Erasing Concepts in Image Data
  • 6. Related Work
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Pseudocode
  • B. Appendices
  • B.1. Ethical Considerations
  • B.2. Rayleigh-quotient
  • B.3. Linear Regression
  • B.4. Optimizing the Relaxed Objective
  • B.4.1. ALTERNATE OPTIMIZATION WITH PROJECTED GRADIENT DESCENT
  • B.5. The INLP Algorithm
  • B.6. Experimental Setting: Static Word Vectors
  • B.7. Experimental Setting: Deep Classifiers
  • B.8. V-Measure
  • B.9. Influence on Neighbors in Embedding Space
  • B.10. Additional results on the CelebsA dataset
  • B.11. Relaxation Quality

Knowls

  1. Knowl 1 — Linear minimax formulation of concept erasure

    model/method

    Given representations xn∈RDx_n\in R^D and scalar labels yny_n for n=1,…,Nn=1,\ldots,N, concept erasure is formulated as a game between a linear predictor and an adversary that projects representations. For a chosen erasure dimension kk, the adversary selects an orthogonal projection P=I−WTWP=I-W^TW, where W∈Rk×DW\in R^{k\times D} and WWT=IkWW^T=I_k; the erased subspace is the row span of WW, and each representation becomes PxnPx_n. The objective is

    min⁡θ∈Θmax⁡P∑n=1Nℓ(yn,g−1(θTPxn)),\min_{\theta\in\Theta}\max_{P}\sum_{n=1}^N \ell\bigl(y_n,g^{-1}(\theta^TPx_n)\bigr),

    where θ\theta parameterizes the predictor, Θ\Theta is its parameter set, g−1g^{-1} is the inverse link function, and ℓ\ell is the selected prediction loss. The adversary seeks a projection that makes prediction difficult under that chosen linear model and loss, while constraining the intervention to a fixed-dimensional orthogonal erasure.

  2. Knowl 2 — R-LACE relaxes the projection game to a convex set

    algorithm

    Relaxed Linear Adversarial Concept Erasure (R-LACE) replaces the nonconvex set of orthogonal projections with the Fantope Fk={A:A=AT, 0⪯A⪯I, tr⁡(A)=k}\mathcal{F}_k=\{A: A=A^T,\ 0\preceq A\preceq I,\ \operatorname{tr}(A)=k\}. It alternates gradient descent on the predictor parameters and gradient ascent on the adversarial matrix. After each ascent update, the matrix is symmetrized and projected onto the Fantope. For a symmetric matrix with eigenvalues λi\lambda_i and eigenvectors viv_i, this projection replaces each eigenvalue with min⁡(max⁡(λi−γ,0),1)\min(\max(\lambda_i-\gamma,0),1), where γ\gamma is chosen so the projected eigenvalues sum to kk; the paper uses bisection to find γ\gamma. At the end, a spectral decomposition converts the relaxed solution into an orthogonal projection that erases kk dimensions.

    Input: representations X, labels y, loss L, erasure dimension k, outer iterations T, inner iterations M
    Initialize predictor θ and adversarial matrix P randomly
    For each outer iteration from 1 to T:
        Repeat M times:
            Update θ by stochastic gradient descent to reduce L(θ, X, y, P)
        Repeat M times:
            Update P by stochastic gradient ascent to increase L(θ, X, y, P)
            Replace P with (P + P^T) / 2
            Project P onto the Fantope using eigenvalue clipping and bisection
    Convert the relaxed solution to an orthogonal projection erasing k dimensions by spectral decomposition
    Return the projection

    The static-word-vector experiment used 50,000 iterations, one predictor/adversary update per iteration, SGD learning rate 0.005, and batch size 128. The deep-classifier experiment used learning rate 0.005, weight decay 10−410^{-4}, and batch size 256. In both settings, the selected adversarial projection was the checkpoint yielding the highest development-set prediction loss after retraining the predictor.

  3. Knowl 3 — Closed-form erasure direction for linear regression

    theoretical result

    For scalar-response linear regression with squared-error loss, let X∈RN×DX\in R^{N\times D} contain the representations as rows and let y∈RNy\in R^N contain the responses. The minimax solution for erasing one dimension projects out the representation–response covariance direction v=XTyv=X^Ty. When v≠0v\ne 0, the optimal projection is

    P=I−vvTvTv.P=I-\frac{vv^T}{v^Tv}.

    This removes the direction used by the best linear predictor to explain the response. The paper reports that at this solution the objective equals the variance of yy; thus the predictor cannot recover additional response information through the projected representations under the stated linear-regression objective.

  4. Knowl 4 — Rayleigh-quotient erasure removes leading eigen-directions

    theoretical result

    For a symmetric matrix A∈RD×DA\in R^{D\times D} with ordered eigenvalues λ1≥λ2≥⋯≥λD\lambda_1\ge\lambda_2\ge\cdots\ge\lambda_D and corresponding orthonormal eigenvectors v1,…,vDv_1,\ldots,v_D, the paper's Rayleigh-quotient game with kk erased dimensions has an optimum that removes the span of the first kk eigenvectors. The retained maximizing direction is θ∗=vk+1\theta^*=v_{k+1}, and the game value is λk+1\lambda_{k+1}. This result supplies a closed-form solution for Rayleigh-quotient objectives, including the paper's partial-least-squares formulation, for which A=XTyyTXA=X^Tyy^TX.

  5. Knowl 5 — INLP is not generally minimal for regression but is optimal for Rayleigh losses

    theoretical result

    The paper compares iterative nullspace projection (INLP), which repeatedly trains a linear predictor and projects away its direction, with the minimax erasure solution. For linear regression, INLP's first fitted direction is the least-squares coefficient (XTX)−1XTy(X^TX)^{-1}X^Ty, whereas the minimax rank-one erasure direction is XTyX^Ty. Since these directions generally differ, INLP is not guaranteed to find a minimal-dimensional erasure for maximizing regression error and can remove unnecessary dimensions. For Rayleigh-quotient losses, by contrast, the paper states that INLP's iterative direction-finding and projection steps recover an optimal set of directions.

  6. Knowl 6 — Rank-one R-LACE nearly eliminates linear gender prediction from GloVe

    empirical result

    On 300-dimensional uncased GloVe vectors labeled for binary gender association, R-LACE reduced held-out linear gender-prediction accuracy to approximately majority accuracy (about 50%) even when only one dimension was erased; this near-chance result held across the tested erasure dimensions k=1,…,20k=1,\ldots,20. The experiment used the word-vector train, development, and test splits of 7,350, 3,150, and 4,500 examples, respectively, and averaged five runs. INLP remained above majority accuracy even after erasing 20 dimensions, while the tested PCA-based gender subspace did not significantly reduce gender accuracy for k=1,…,10k=1,\ldots,10. The result supports the paper's finding that, for this dataset and linear prediction criterion, gender information can be neutralized by a one-dimensional projection.

  7. Knowl 7 — GloVe association tests decrease while similarity quality is retained

    empirical result

    On three Word Embedding Association Tests (WEAT), rank-one R-LACE reduced measured gender associations more than the tested PCA and INLP projections. The table reports WEAT effect size dd (lower indicates weaker association) and p-value after projection; projected-method entries are mean ±\pm standard deviation across runs.

    Test Method WEAT d p-value
    Math-art Original 1.57 0.000
    Math-art PCA 1.37 ±\pm 0.00 0.002 ±\pm 0.000
    Math-art R-LACE 0.80 ±\pm 0.01 0.062 ±\pm 0.002
    Math-art INLP 1.10 ±\pm 0.10 0.016 ±\pm 0.009
    Professions-family Original 1.69 0.000
    Professions-family PCA 1.24 ±\pm 0.00 0.005 ±\pm 0.000
    Professions-family R-LACE 0.78 ±\pm 0.01 0.072 ±\pm 0.003
    Professions-family INLP 1.15 ±\pm 0.07 0.007 ±\pm 0.003
    Science-art Original 1.63 0.000
    Science-art PCA 1.16 ±\pm 0.00 0.003 ±\pm 0.000
    Science-art R-LACE 0.77 ±\pm 0.01 0.073 ±\pm 0.003
    Science-art INLP 1.03 ±\pm 0.11 0.022 ±\pm 0.016

    The R-LACE p-values are near or above the conventional 0.05 threshold in all three tests. On SimLex-999, the Pearson correlation between embedding cosine similarities and human similarity judgments changed from 0.399 for original vectors to 0.392 after rank-one R-LACE (and 0.395 after one INLP iteration); the paper reports no significant effect on this semantic-similarity measure.

  8. Knowl 8 — R-LACE reduces gender predictability in BERT while largely preserving profession accuracy

    empirical result

    The deep-classification experiment used biographies annotated with gender and profession. The authors extracted BERT [CLS] representations, reduced them to 300 dimensions with PCA, applied erasure to frozen or profession-finetuned representations, and evaluated profession prediction and fairness. The table gives gender and profession accuracy in percent, the root-mean-square true-positive-rate gap across professions (lower is better), and the correlation between profession-level TPR gaps and percentage of women (lower magnitude is reported as better). Finetuned-model entries are mean ±\pm standard deviation across five runs.

    Setting Gender accuracy Profession accuracy TPR-gap RMS Gap/BERT-frozen 99.32 79.14 0.145 0.813
    BERT-frozen + R-LACE (rank 1) 52.48 78.86 0.109 0.680
    BERT-frozen + R-LACE (rank 100) 52.77 77.28 0.102 0.615
    BERT-frozen + INLP (rank 1) 98.98 79.09 0.137 0.816
    BERT-frozen + INLP (rank 100) 53.21 71.94 0.099 0.604
    BERT-finetuned 96.89 ±\pm 1.01 85.12 ±\pm 0.08 0.123 ±\pm 0.011 0.810 ±\pm 0.023
    BERT-finetuned + R-LACE (rank 1) 54.59 ±\pm 0.66 85.09 ±\pm 0.07 0.117 ±\pm 0.011 0.794 ±\pm 0.025
    BERT-finetuned + R-LACE (rank 100) 54.33 ±\pm 0.36 85.04 ±\pm 0.09 0.115 ±\pm 0.014 0.792 ±\pm 0.025
    BERT-finetuned + INLP (rank 1) 93.52 ±\pm 1.42 85.12 ±\pm 0.08 0.122 ±\pm 0.011 0.808 ±\pm 0.024
    BERT-finetuned + INLP (rank 100) 53.04 ±\pm 0.97 84.98 ±\pm 0.06 0.113 ±\pm 0.009 0.797 ±\pm 0.027
    BERT-finetuned-adv (MLP adversary) 99.57 ±\pm 0.05 84.87 ±\pm 0.11 0.128 ±\pm 0.004 0.840 ±\pm 0.015
    BERT-finetuned-adv (linear adversary) 99.23 ±\pm 0.09 84.92 ±\pm 0.12 0.124 ±\pm 0.005 0.827 ±\pm 0.012
    Majority 53.52 30.0 - -

    Rank-one R-LACE brought gender accuracy close to majority while changing finetuned profession accuracy only from 85.12% to 85.09%. Rank-one INLP did not comparably erase gender predictability; rank-100 INLP did so but reduced frozen-BERT profession accuracy to 71.94%. The adversarially finetuned baselines retained nearly perfect gender predictability at test time and showed no reduction in the reported TPR-gap RMS relative to the unmodified finetuned model. The fairness measure is the difference in true-positive rates between protected groups, conditioned on the true profession; the paper cautions that its relationship to gender predictability is not clear-cut.

  9. Knowl 9 — Rank-one projections suppress linearly classifiable image attributes

    empirical result

    To inspect the effect of erasure directly, the authors applied R-LACE to raw pixel vectors from CelebA face images. Images were converted to 50-by-50 grayscale pixels and flattened into 2,500-dimensional vectors; the target attributes were glasses, smile, mustache, beard, baldness, and hat. For each attribute, a rank-one projection reduced classification accuracy to less than one percentage point above majority accuracy. The resulting images changed in attribute-related ways—for example, the glasses intervention could add pseudo-sunglasses, while the smile intervention blurred mouth-area features. Because the operation is a projection, the paper notes that it is more limited in expressivity and can remove features more readily than it can add them.

  10. Knowl 10 — Erasure guarantees are limited to linear predictability

    limitation

    R-LACE is designed to hinder predictors in the selected linear model family; it does not guarantee that a concept is absent from the representation or inaccessible to nonlinear predictors. In the GloVe experiment, an RBF-SVM and a one-hidden-layer ReLU MLP still predicted gender with accuracy above 90% after linear erasure. The paper therefore cautions against interpreting the method as a general solution to bias, especially because its experiments use limited datasets and a binary-gender operationalization. More broadly, alternating optimization is not guaranteed to reach the minimax equilibrium and can exhibit rotational behavior; the experiments instead selected projections using development-set prediction loss.

Coverage note — Supplementary clustering and nearest-neighbor checks, projection-eigenvalue diagnostics, and additional image examples are omitted because they provide supporting illustrations rather than separate load-bearing contributions.

References

  1. 1.Mathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig, Stephan Hoyer, Felipe Llinares-Lopez, Fabian Pedregosa, and Jean-Philippe Vert. 2021. Efficient and modular implicit differentiation.
  2. 2.Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai. 2016. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. Advances in Neural Information Processing Systems, 29:4349–4357.
  3. 3.Stephen P. Boyd and Lieven Vandenberghe. 2014. Convex Optimization. Cambridge University Press.
  4. 4.Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  5. 5.Hande Celikkanat, Sami Virpioja, Jörg Tiedemann, and Marianna Apidianaki. 2020. Controlling the imprint of passivization and negation in contextualized representations. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 136–148, Online. Association for Computational Linguistics.
  6. 6.Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger. 2018. Adversarial deep averaging networks for cross-lingual sentiment classification. Transactions of the Association for Computational Linguistics, 6:557–570.
  7. 7.Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, page 120–128, New York, NY, USA. Association for Computing Machinery.
  8. 8.Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2021. OSCaR: Orthogonal subspace correction and rectification of biases in word embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5034–5050, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  9. 9.Sunipa Dev and Jeff Phillips. 2019. Attenuating bias in word vectors. In Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume 89 of Proceedings of Machine Learning Research, pages 879–887. PMLR.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Harrison Edwards and Amos Storkey. 2016. Censoring representations with an adversary. In International Conference in Learning Representations, pages 1–14.
  12. 12.Yanai Elazar and Yoav Goldberg. 2018. Adversarial removal of demographic attributes from text data. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 11–21, Brussels, Belgium. Association for Computational Linguistics.
  13. 13.Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175.
  14. 14.Yaroslav Ganin and Victor Lempitsky. 2015. Unsupervised domain adaptation by backpropagation. In Proceedings of the 32nd International Conference on International Conference on Machine Learning, volume 37, page 1180–1189.
  15. 15.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1926–1940, Online. Association for Computational Linguistics.
  16. 16.Hila Gonen and Yoav Goldberg. 2019. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 609–614, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Hila Gonen, Shauli Ravfogel, Yanai Elazar, and Yoav Goldberg. 2020. It’s not Greek to mBERT: Inducing word-level translations from multilingual BERT. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 45–56, Online. Association for Computational Linguistics.
  18. 18.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems, volume 27.
  19. 19.Pantea Haghighatkhah, Wouter Meulemans, Bettina Speckmann, Jérôme Urhausen, and Kevin Verbeek. 2021. Obstructing classification via projection. In 46th International Symposium on Mathematical Foundations of Computer Science, Leibniz International Proceedings in Informatics, LIPIcs. Schloss Dagstuhl - Leibniz-Zentrum für Informatik.
  20. 20.Moritz Hardt, Eric Price, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  21. 21.Yuzi He, Keith Burghardt, and Kristina Lerman. 2020. A Geometric Solution to Fair Representations, page 279–285. Association for Computing Machinery, New York, NY, USA.
  22. 22.Evan Hernandez and Jacob Andreas. 2021. The low-dimensional linear geometry of contextualized word representations. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 82–93, Online. Association for Computational Linguistics.
  23. 23.Felix Hill, Roi Reichart, and Anna Korhonen. 2015. SimLex-999: Evaluating semantic models with (genuine) similarity estimation. Computational Linguistics, 41(4):665–695.
  24. 24.Roger A. Horn and Charles R. Johnson. 2012. Matrix Analysis. Cambridge University Press.
  25. 25.Harold Hotelling and Margaret Richards Pabst. 1936. Rank correlation and tests of significance involving no assumption of normality. The Annals of Mathematical Statistics, 7(1):29–43.
  26. 26.Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, Melbourne, Australia. Association for Computational Linguistics.
  27. 27.Masahiro Kaneko and Danushka Bollegala. 2021. Debiasing pre-trained contextualised embeddings. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1256–1266, Online. Association for Computational Linguistics.
  28. 28.Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory sayres. 2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2668–2677. PMLR.
  29. 29.Hellmuth Kneser. 1952. Sur un theoreme fondamental de la theorie des jeux. Comptes rendus de l’Academie des Sciences Paris, 234:2418–2420.
  30. 30.Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016. Context2vec: Learning generic context embedding with bidirectional LSTM. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 51–61, Berlin, Germany. Association for Computational Linguistics.
  31. 31.J. A. Nelder and Robert W. M. Wedderburn. 1972. Generalized linear models. Journal of the Royal Statistical Society: Series A (General), 135(3):370–384.
  32. 32.Maher Nouiehed, Maziar Sanjabi, Tianjian Huang, Jason D Lee, and Meisam Razaviyayn. 2019. Solving a class of non-convex min-max games using iterative first order methods. In Advances in Neural Information Processing Systems, volume 32.
  33. 33.Hadas Orgad, Seraphina Goldfarb-Tarrant, and Yonatan Belinkov. 2022. How gender debiasing affects internal model representations, and why it matters. arXiv preprint arXiv:2204.06827.
  34. 34.Jong-Shi Pang and Meisam Razaviyayn. 2016. A unified distributed algorithm for non-cooperative games. In Shuguang Cui, Alfred O. Hero, III, Zhi-Quan Luo, and Jose M. F. Moura, editors, Big Data over Networks, chapter 4, page 101–134. Cambridge University Press.
  35. 35.Lance Parsons, Ehtesham Haque, and Huan Liu. 2004. Subspace clustering for high dimensional data: A review. ACM SIGKDD Explorations Newsletter, 6(1):90–105.
  36. 36.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  37. 37.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  38. 38.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  39. 39.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256, Online. Association for Computational Linguistics.
  40. 40.Shauli Ravfogel, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021. Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 194–209, Online. Association for Computational Linguistics.
  41. 41.Andrew Rosenberg and Julia Hirschberg. 2007. V-measure: A conditional entropy-based external cluster evaluation measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pages 410–420, Prague, Czech Republic. Association for Computational Linguistics.
  42. 42.Bashir Sadeghi and Vishnu Boddeti. 2021. On the fundamental trade-offs in learning invariant representations. arXiv preprint arXiv:2109.03386.
  43. 43.Bashir Sadeghi, Runyi Yu, and Vishnu Boddeti. 2019. On the global optima of kernelized adversarial representation learning. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7970–7978. IEEE.
  44. 44.Shun Shao, Yftah Ziser, and Shay B. Cohen. 2022. Gold doesn’t always glitter: Spectral removal of linear and nonlinear guarded attribute information. arXiv preprint arXiv:2203.07893.
  45. 45.Hoang Tuy. 2004. Minimax theorems revisited. Acta Mathematica Vietnamica, 29(3):217–229.
  46. 46.Francisco Vargas and Ryan Cotterell. 2020. Exploring the linear subspace hypothesis in gender bias mitigation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2902–2913, Online. Association for Computational Linguistics.
  47. 47.John von Neumann and Oskar Morgenstern. 1947. Theory of games and economic behavior. Princeton University Press.
  48. 48.Vincent Q. Vu, Juhee Cho, Jing Lei, and Karl Rohe. 2013. Fantope projection and selection: A near-optimal convex relaxation of sparse PCA. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pages 2670–2678.
  49. 49.Liwen Wang, Yuanmeng Yan, Keqing He, Yanan Wu, and Weiran Xu. 2021. Dynamically disentangling social bias from task-oriented representations with adversarial attack. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3740–3750, Online. Association for Computational Linguistics.
  50. 50.Yuanhao Wang and Jian Li. 2020. Improved algorithms for convex-concave minimax optimization. In Advances in Neural Information Processing Systems, volume 33, pages 4800–4810.
  51. 51.Herman Wold. 1966. Estimation of principal components and related models by iterative least squares. Multivariate Analysis, pages 391–420.
  52. 52.Herman Wold. 1973. Nonlinear iterative partial least squares (NIPALS) modelling: Some current developments. In Paruchuri R. Krishnaiah, editor, Multivariate Analysis–III, pages 383–407. Academic Press.
  53. 53.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  54. 54.Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig. 2017. Controllable invariance through adversarial feature learning. In Advances in Neural Information Processing Systems, volume 30, pages 585–596.
  55. 55.Shuo Yang, Ping Luo, Chen-Change Loy, and Xiaoou Tang. 2015. From facial parts responses to face detection: A deep learning approach. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 3676–3684.
  56. 56.Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018. Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, page 335–340, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Ravfogel, S., et al. “Linear Adversarial Concept Erasure”. International Conference on Machine Learning, vol. 162, 2022, pp. 18400–21, https://proceedings.mlr.press/v162/ravfogel22a.html.
APA
Ravfogel, S., Twiton, M., Goldberg, Y., & Cotterell, R. D. (2022). Linear Adversarial Concept Erasure. International Conference on Machine Learning, 162, 18400–18421. https://proceedings.mlr.press/v162/ravfogel22a.html
Chicago
Ravfogel, S., M. Twiton, Y. Goldberg, and R. D. Cotterell. 2022. “Linear Adversarial Concept Erasure”. International Conference on Machine Learning 162: 18400–18421. https://proceedings.mlr.press/v162/ravfogel22a.html.
Harvard
Ravfogel, S. et al. (2022) “Linear Adversarial Concept Erasure”, International Conference on Machine Learning. PMLR, pp. 18400–18421. Available at: https://proceedings.mlr.press/v162/ravfogel22a.html.
Vancouver
1. Ravfogel S, Twiton M, Goldberg Y, Cotterell RD (2022) Linear Adversarial Concept Erasure. In: International Conference on Machine Learning. PMLR, pp 18400–18421

BibTeX

@InProceedings{pmlr-v162-ravfogel22a,
  title = 	 {Linear Adversarial Concept Erasure},
  author =       {Ravfogel, Shauli and Twiton, Michael and Goldberg, Yoav and Cotterell, Ryan D},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {18400--18421},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/ravfogel22a/ravfogel22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/ravfogel22a.html},
  abstract = 	 {Modern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in real-world applications, the inability to control their content becomes an increasingly important problem. In this work, we formulate the problem of identifying a linear subspace that corresponds to a given concept, and removing it from the representation. We formulate this problem as a constrained, linear minimax game, and show that existing solutions are generally not optimal for this task. We derive a closed-form solution for certain objectives, and propose a convex relaxation that works well for others. When evaluated in the context of binary gender removal, the method recovers a low-dimensional subspace whose removal mitigates bias by intrinsic and extrinsic evaluation. Surprisingly, we show that the method—despite being linear—is highly expressive, effectively mitigating bias in the output layers of deep, nonlinear classifiers while maintaining tractability and interpretability.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/