Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

Fan ZhouYuzhou MaoLiu YuYi YangTing Zhong

article2023ACL76 citations

Introduces a causal invariant learning framework that unifies language model debiasing with downstream fine-tuning, preventing stereotypical biases from resurfacing during task adaptation without hurting model utility.

Listen

Pretrained language models often absorb social stereotypes and demographic biases from large text corpora. While existing techniques attempt to remove these biases during pretraining or representation stages, the article demonstrates that standard downstream fine-tuning causes unwanted biases to resurface or amplify in deployed applications. The article introduces and evaluates Causal-Debias, a unified framework designed to prevent bias resurgence by integrating debiasing directly into the downstream fine-tuning process.

The framework addresses bias propagation using a structural causal model that separates task-essential causal factors from non-causal demographic attributes. Causal-Debias generates counterfactual sentences and retrieves semantically similar examples from external corpora to establish interventional demographic environments. It then applies an invariant risk minimization loss during fine-tuning. The approach was evaluated across three widely used language models (BERT, ALBERT, and RoBERTa) and three downstream language processing tasks (SST-2, CoLA, and QNLI) to mitigate gender and racial biases.

The findings show that Causal-Debias consistently achieves lower bias scores than existing standalone debiasing methods while preserving application performance. Under the Sentence Encoder Association Test, Causal-Debias outperformed existing methods on downstream tasks and minimized bias deviations on benchmark evaluation pairs. In contrast, previously debiased baseline models exhibited significant bias resurgence once fine-tuned on task data, often suffering performance degradation as well. In racial debiasing experiments, Causal-Debias effectively handled word ambiguity issues that degraded the performance of prior methods.

These results indicate that treating bias mitigation as an isolated pre-deployment step introduces compliance, fairness, and reputational risks for real-world artificial intelligence deployments. To ensure fair and accountable language applications, organizations should incorporate causal invariant debiasing directly into downstream fine-tuning workflows rather than relying solely on off-the-shelf debiased models.

Nevertheless, decision-makers should recognize current limitations: the framework relies on predefined gender and race word pairs, treats demographic categories as binary, operates only on English-language text, and relies on bias metrics that primarily capture North American contexts. Future initiatives should pilot causal invariant methods across multilingual settings, explore intersectional and domain-specific biases, and develop broader evaluation standards.

Zhou et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

Abstract

Pretrained Language Models (PLMs) have achieved remarkable success in various NLP tasks, but they are also known to encode social biases from the training data, which can lead to harmful consequences when applied to downstream tasks. Existing debiasing methods for PLMs and fine-tuning are often studied separately, and the debiasing effect is not well understood from a causal perspective. In this paper, we propose a novel framework, Causal-Debias, which unifies debiasing in PLMs and fine-tuning via causal invariant learning. We first formulate the debiasing problem from a causal perspective, and then propose a causal invariant learning approach to learn debiased representations that are invariant across different environments. Specifically, we introduce a causal intervention module to disentangle the biased and unbiased components in the representation, and a causal invariant learning module to enforce the unbiased component to be invariant across different environments. We conduct extensive experiments on three real-world datasets, and the results demonstrate the effectiveness of our proposed framework.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Works
  • 3 Methodology
  • 3.1 Biases from a Causal View
  • 3.2 Causal-Debias
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Results on Mitigating Gender Bias
  • SEAT Tests and Downstream Tasks Evaluation.
  • 4.3 Ablation Study
  • 4.4 Results on Mitigating Racial Bias
  • 5 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Instantiated Causal Graphs
  • B Bias Words List
  • C External Corpora
  • D SEAT Details

Knowls

  1. Knowl 1 — Causal model of task-relevant and bias-relevant sentence factors

    model/method

    Causal-Debias represents a supervised NLP example using an observed sentence XX, its label YY, and latent factors CC and NN. The sentence is generated from both the task-relevant causal factor CC and the bias-relevant, non-causal factor NN; the assumed causal link to the label is C→YC\to Y, not N→YN\to Y. The factors may nevertheless be statistically dependent, yielding a spurious association between NN and YY along the path N←C→YN\leftarrow C\to Y. Which features belong to CC or NN depends on the task: for sentiment classification, sentiment-bearing adjectives may be causal while nouns or pronouns are non-causal; for coreference, pronouns may instead be causal. The learning assumption is that the label is conditionally independent of the non-causal factor given the causal factor, Y⊥N∣CY\perp N\mid C, so the relation between CC and YY remains stable when NN is intervened upon.

  2. Knowl 2 — Counterfactual and corpus-based demographic interventions

    algorithm

    Causal-Debias constructs intervention data by pairing downstream examples with counterfactual and semantically related sentences. Let XoX_o be downstream sentences containing an attribute word or target word, where attribute words WaW_a denote demographic groups and target words WtW_t are concepts whose associations are being examined. First form Xd=Xo∪XcX_d=X_o\cup X_c, where XcX_c is created by replacing an attribute word with its paired alternative, such as a masculine term with its feminine counterpart; the counterfactual retains the original label because the intended sentence meaning is unchanged. Next, use cosine semantic similarity to retrieve the top-kk sentences from external corpus EE that are most similar to the bias-related data, expanding the original and counterfactual subsets into X~o\widetilde X_o and X~c\widetilde X_c. The intervention set is described as X~d=Xd∪TopK⁡(sim⁡(Xd,E))\widetilde X_d=X_d\cup\operatorname{TopK}(\operatorname{sim}(X_d,E)). The expanded intervention data is combined with the remaining bias-unrelated downstream examples for training.

  3. Knowl 3 — Invariant-risk objective for debiasing during fine-tuning

    equation

    For each demographic intervention value nn, let RnR_n be the prediction risk of the language model under the interventional distribution do(N=n)do(N=n), where NN is the non-causal demographic factor. Causal-Debias seeks low average risk while penalizing variation in risk across interventions: Linvariant=En[Rn]+Var⁡n(Rn)L_{\mathrm{invariant}}=\mathbb{E}_{n}[R_n]+\operatorname{Var}_{n}(R_n). In practice, the model is encouraged to produce agreeing predictions for original and counterfactual sentences with equivalent intended meanings but different attribute words; the paper uses Wasserstein distance to measure agreement between the corresponding predictions. Fine-tuning jointly optimizes the task prediction loss and the invariant loss: Ltotal=Lprediction+τLinvariantL_{\mathrm{total}}=L_{\mathrm{prediction}}+\tau L_{\mathrm{invariant}}, where LpredictionL_{\mathrm{prediction}} is the downstream task loss and τ\tau controls the trade-off between task performance and invariance.

  4. Knowl 4 — Evaluation protocol and experimental settings

    experimental setup

    The experiments evaluate gender and racial debiasing on SST-2 sentiment classification, CoLA grammatical acceptability, and QNLI question-answer inference. Gender experiments use BERT-base-uncased, ALBERT-large-v2, and RoBERTa-base; racial experiments use BERT-base-uncased and ALBERT-base-v2. Causal-Debias is trained for 5 epochs with learning rate 2×10−52\times10^{-5}, and reported downstream results average 5 runs. Its external intervention corpora contain 183,060 sentences drawn from WikiText-2, the Stanford Sentiment Treebank, Reddit, MELD, and POM. Gender bias is assessed with SEAT tests 6, 6b, 7, 7b, 8, and 8b; racial bias is assessed with SEAT tests 3, 3b, 4, 5, and 5b. SEAT effect sizes closer to zero indicate lower measured bias. Gender results also use CrowS-Pairs, for which a score closer to 50% indicates less stereotypical preference. Downstream performance is reported as accuracy for SST-2 and QNLI and Matthews correlation coefficient for CoLA. Comparisons include non-task-specific methods such as CDA, Dropout, Context-Debias, Auto-Debias, and MABEL, and task-specific methods such as Sent-Debias and FairFil.

  5. Knowl 5 — Gender debiasing results across downstream tasks

    empirical result

    After fine-tuning, Causal-Debias obtains the following average SEAT effect sizes and downstream scores. For BERT, SST-2 is 0.11 SEAT and 92.9 accuracy, CoLA is 0.11 SEAT and 58.1 Matthews correlation coefficient, and QNLI is 0.15 SEAT and 91.6 accuracy. For ALBERT, the corresponding results are 0.06 and 92.9, 0.16 and 57.1, and 0.09 and 91.6. For RoBERTa, they are 0.09 and 93.9, 0.17 and 54.1, and 0.06 and 92.9. Lower SEAT effect sizes indicate less measured gender association. The paper reports that Causal-Debias has the lowest average SEAT scores among the compared methods for each downstream task. On BERT, its SEAT scores are lower than FairFil's by 0.07, 0.01, and 0.07 for SST-2, CoLA, and QNLI, respectively, and lower than Auto-Debias's by 0.27, 0.21, and 0.09. The reported task scores show that this bias reduction is achieved alongside maintained downstream performance.

  6. Knowl 6 — Fine-tuning can restore bias in separately debiased models

    empirical result

    The experiments show that bias mitigation performed separately from downstream fine-tuning often does not persist after fine-tuning: SEAT scores increased for almost all non-task-specific debiasing methods in the reported settings. For example, with BERT plus Auto-Debias, the reported SST-2 SEAT score rises from 0.14 before fine-tuning to 0.38 after; the post-fine-tuning scores are 0.32 on CoLA and 0.24 on QNLI, with increases of 0.18 and 0.10 reported for those tasks. In contrast, the original, non-debiased BERT often becomes less biased after fine-tuning, including a decrease from 0.35 to 0.29 on SST-2 and from 0.35 to 0.18 on CoLA. These results support the paper's finding that standalone debiasing and downstream fine-tuning can interact in ways that reintroduce measured bias, motivating joint treatment of the two stages.

  7. Knowl 7 — CrowS-Pairs evaluation of gender bias

    empirical result

    On BERT fine-tuned for SST-2, Causal-Debias obtains a CrowS-Pairs score of 48.94%, an absolute deviation of 1.06 percentage points from the 50% less-stereotypical reference. This is the smallest deviation among the reported fine-tuned models. The post-fine-tuning scores for the other methods are BERT 53.18%, CDA 58.42%, Dropout 44.56%, Context-Debias 58.89%, Auto-Debias 44.96%, MABEL 46.75%, and Sent-Debias 55.04%. The score's distance from 50%, rather than whether it is above or below 50%, is used to indicate measured stereotypical preference; the results therefore provide a second evaluation, beyond SEAT, in which Causal-Debias is closest to the stated reference.

  8. Knowl 8 — Racial debiasing results after fine-tuning

    empirical result

    For racial-bias evaluation, Causal-Debias reports the following post-fine-tuning SEAT scores and task results. With BERT, SST-2 achieves 0.11 SEAT and 92.9 accuracy, CoLA 0.06 SEAT and 57.1 Matthews correlation coefficient, and QNLI 0.11 SEAT and 91.6 accuracy. With ALBERT, the corresponding results are 0.13 and 91.9, 0.16 and 59.6, and 0.01 and 92.5. For comparison, Auto-Debias obtains post-fine-tuning SEAT scores of 0.31, 0.20, and 0.24 on BERT and 0.39, 0.18, and 0.36 on ALBERT for SST-2, CoLA, and QNLI, respectively. The paper reports that Causal-Debias reduces measured racial bias while maintaining comparable task performance, whereas Auto-Debias exhibits bias recurrence after fine-tuning.

  9. Knowl 9 — Ablation findings on corpus expansion and invariant learning

    empirical result

    The ablation compares the full Causal-Debias model with variants that remove external-corpus expansion, remove both expansion and invariant loss while retaining downstream counterfactual augmentation, or remove invariant loss and retain only task prediction loss. On SST-2, the variant without external-corpus expansion performs worst on debiasing, indicating that corpus-based expansion contributes substantially to reducing measured bias. The comparisons also show a task-utility cost from augmented examples when invariant learning is absent: the paper attributes degraded accuracy to additional semantics that can obscure the task signal. The full model's invariant objective mitigates this adverse effect, producing a marked accuracy improvement over the prediction-only variant while retaining low SEAT bias.

  10. Knowl 10 — Scope and limitations of the debiasing evidence

    limitation

    Causal-Debias relies on human-curated gender and racial attribute lists, which the authors acknowledge do not cover all bias-related demographic groups, and on external corpora to create intervention distributions. The experiments focus primarily on gender and racial bias considered separately, use binary demographic contrasts, and are based on an English-language system; they do not establish performance for intersectional biases, other demographic axes, or other languages. The authors also caution that favorable SEAT and CrowS-Pairs results do not demonstrate complete bias removal: these measures primarily reflect selected North American social biases and positive predictive evidence of bias, not its absence. The paper identifies the lack of a universal, reliable, and broadly agreed debiasing evaluation framework as a further limitation.

Coverage note — No substantial contributed method or result was omitted; detailed per-corpus sentence examples and individual SEAT word-list contents were left out because they describe evaluation resources rather than distinct contributions.

References

  1. 1.Ahmed Abbasi, David Dobolyi, John P Lalor, Richard G Netemeyer, Kendall Smith, and Yi Yang. 2021. Constructing a psychometric testbed for fair natural language processing. In EMNLP, pages 3748–3758.
  2. 2.Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant risk minimization. arXiv :1907.02893.
  3. 3.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  4. 4.Shiyu Chang, Yang Zhang, Mo Yu, and Tommi Jaakkola. 2020. Invariant rationalization. In ICML, pages 1448–1458. PMLR.
  5. 5.Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, and Lawrence Carin. 2021. Fairfil: Contrastive neural debiasing method for pretrained text encoders. In ICLR.
  6. 6.Chengyu Chuang and Yi Yang. 2022. Buy tesla, sell ford: Assessing implicit stock market preference in pre-trained language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 100–105.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186.
  8. 8.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In ACL, pages 1926–1940.
  9. 9.Hila Gonen and Yoav Goldberg. 2019. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. arXiv preprint arXiv:1903.03862.
  10. 10.Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2013. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211.
  11. 11.Yue Guo, Yi Yang, and Ahmed Abbasi. 2022. Auto-debias: Debiasing masked language models with automated biased prompts. In ACL, pages 1012–1023.
  12. 12.Jacqueline He, Mengzhou Xia, Christiane Fellbaum, and Danqi Chen. 2022. Mabel: Attenuating gender bias using textual entailment data. arXiv:2210.14975.
  13. 13.Masahiro Kaneko and Danushka Bollegala. 2021. Debiasing pre-trained contextualised embeddings. In EACL.
  14. 14.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526.
  15. 15.John P Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, and Ahmed Abbasi. 2022. Benchmarking intersectional biases in nlp. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3598–3609.
  16. 16.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. Albert: A lite bert for self-supervised learning of language representations. In ICLR.
  17. 17.Shaobo Li, Xiaoguang Li, Lifeng Shang, Zhenhua Dong, Cheng-Jie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang, and Qun Liu. 2022. How pre-trained language models capture factual knowledge? a causal-inspired analysis. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1720–1732.
  18. 18.Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Towards debiasing sentence representations. In ACL, pages 5502–5515.
  19. 19.Paul Pu Liang, Yao Chong Lim, Yao-Hung Hubert Tsai, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2019. Strong and simple baselines for multimodal utterance embeddings. In NAACL, pages 2599–2609.
  20. 20.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692.
  21. 21.Fangrui Lv, Jian Liang, Shuang Li, Bin Zang, Chi Harold Liu, Ziteng Wang, and Di Liu. 2022. Causality inspired representation learning for domain generalization. In CVPR, pages 8046–8056.
  22. 22.Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019. Black is to criminal as caucasian is to police: Detecting and removing multi-class bias in word embeddings. In NAACL, pages 615–621.
  23. 23.Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In NAACL, pages 622–628.
  24. 24.Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022. An empirical survey of the effectiveness of debiasing techniques for pre-trained language models. In ACL, pages 1878–1898.
  25. 25.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In ICLR.
  26. 26.Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf. 2013. Domain generalization via invariant feature representation. In ICML, pages 10–18. PMLR.
  27. 27.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020. Crows-pairs: A challenge dataset for measuring social biases in masked language models. In EMNLP, pages 1953–1967.
  28. 28.Sunghyun Park, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, and Louis-Philippe Morency. 2014. Computational analysis of persuasiveness in social multimedia: A novel dataset and multimodal prediction approach. In ICMI, pages 50–57.
  29. 29.Judea Pearl, Madelyn Glymour, and Nicholas P Jewell. 2016. Causal Inference in Statistics: A Primer. John Wiley & Sons.
  30. 30.Judea Pearl et al. 2000. Models, reasoning and inference. Cambridge, UK: CambridgeUniversityPress, 19(2).
  31. 31.Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. 2016. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78(5):947–1012.
  32. 32.Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. 2017. Elements of causal inference: foundations and learning algorithms. The MIT Press.
  33. 33.Maxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal, Barun Patra, Dean Carignan, Emre Kiciman, and Robert West. 2021. Invariant language modeling. arXiv:2110.08413.
  34. 34.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. In ACL, pages 527–536.
  35. 35.Chen Qian, Fuli Feng, Lijie Wen, Chunping Ma, and Pengjun Xie. 2021. Counterfactual inference for text classification debiasing. In ACL, pages 5434–5445.
  36. 36.Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. 2022. Perturbation augmentation for fairer nlp. arXiv preprint arXiv:2205.12586.
  37. 37.Aaditya Ramdas, Nicolás García Trillos, and Marco Cuturi. 2017. On wasserstein two-sample testing and related families of nonparametric tests. Entropy, 19(2):47.
  38. 38.Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris M Mooij. 2012. On causal and anticausal learning. In ICML.
  39. 39.Seonguk Seo, Joon-Young Lee, and Bohyung Han. 2022. Unsupervised learning of debiased representations with pseudo-attributes. In CVPR, pages 16742–16751.
  40. 40.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642.
  41. 41.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  42. 42.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. In EMNLP, pages 353–355.
  43. 43.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv:2010.06032.
  44. 44.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In EMNLP, pages 38–45.
  45. 45.Ying-Xin Wu, Xiang Wang, An Zhang, Xiangnan He, and Tat-Seng Chua. 2022. Discovering invariant rationales for graph neural networks. arXiv:2201.12872.
  46. 46.Xiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang, Renzhe Xu, Peng Cui, and Tong Zhang. 2022. Model agnostic sample reweighting for out-of-distribution learning. In ICML, pages 27203–27221. PMLR.
  47. 47.Ran Zmigrod, Sabrina J Mielke, Hanna Wallach, and Ryan Cotterell. 2019. Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. In ACL, pages 1651–1661.

Citation

MLA
Zhou, F., et al. “Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 4227–41, https://doi.org/10.18653/v1/2023.acl-long.232.
APA
Zhou, F., Mao, Y., Yu, L., Yang, Y., & Zhong, T. (2023). Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4227–4241. https://doi.org/10.18653/v1/2023.acl-long.232
Chicago
Zhou, F., Y. Mao, L. Yu, Y. Yang, and T. Zhong. 2023. “Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4227–41. https://doi.org/10.18653/v1/2023.acl-long.232.
Harvard
Zhou, F. et al. (2023) “Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4227–4241. Available at: https://doi.org/10.18653/v1/2023.acl-long.232.
Vancouver
1. Zhou F, Mao Y, Yu L, Yang Y, Zhong T (2023) Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4227–4241

BibTeX

@inproceedings{zhou-etal-2023-causal,
    title = "Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning",
    author = "Zhou, Fan  and
      Mao, Yuzhou  and
      Yu, Liu  and
      Yang, Yi  and
      Zhong, Ting",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.232/",
    doi = "10.18653/v1/2023.acl-long.232",
    pages = "4227--4241"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/