Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Causal-Debias

Causal-Debias is a machine learning debiasing framework designed to mitigate social biases across both pretrained language models and downstream fine-tuning stages using causal invariant learning. The approach addresses bias from a causal perspective by disentangling biased attributes from core, task-relevant data representations. Through causal interventions and invariant learning mechanisms, it encourages models to generate representations that remain invariant across diverse demographic and contextual environments, effectively preventing language models from relying on spurious or harmful social correlations while preserving their predictive performance on downstream tasks.

1 item

Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

Fan Zhou, Yuzhou Mao, Liu Yu, Yi Yang, Ting Zhong

Why you should read this

Introduces a causal invariant learning framework that unifies language model debiasing with downstream fine-tuning, preventing stereotypical biases from resurfacing during task adaptation without hurting model utility.

Pretrained Language Models (PLMs) have achieved remarkable success in various NLP tasks, but they are also known to encode social biases from the training data, which can lead to harmful consequences when applied to downstream tasks. Existing debiasing methods for PLMs and fine-tuning are often studied separately, and the debiasing effect is not well understood from a causal perspective. In this paper, we propose a novel framework, Causal-Debias, which unifies debiasing in PLMs and fine-tuning via causal invariant learning. We first formulate the debiasing problem from a causal perspective, and then propose a causal invariant learning approach to learn debiased representations that are invariant across different environments. Specifically, we introduce a causal intervention module to disentangle the biased and unbiased components in the representation, and a causal invariant learning module to enforce the unbiased component to be invariant across different environments. We conduct extensive experiments on three real-world datasets, and the results demonstrate the effectiveness of our proposed framework.

Added

2026-10-03