Mitigating Label Biases for In-context Learning
Yu FeiYifan HouZeming ChenAntoine Bosselut
Proposes a domain-context calibration method that estimates and removes pre-existing task corpus biases from language models, improving in-context text classification performance by up to 37% Macro-F1.
Large language models are increasingly used for automated text classification via in-context learning, where models make predictions based on a small context prompt of example demonstration pairs without updating underlying parameters. However, in-context learning is notoriously brittle and vulnerable to systemic label biases—unwanted preferences toward predicting certain target labels regardless of input meaning. When unaddressed, these biases can lead to severely skewed outputs and unpredictable failures in production environments.
The article systematically categorizes label biases in in-context learning and introduces Domain-context Calibration, an efficient, tuning-free method to evaluate and eliminate these distortions. Specifically, the article formalizes three distinct sources of bias: vanilla-label bias (inherent preferences from word frequencies during pre-training), context-label bias (distortions introduced by prompt formatting and example order), and domain-label bias (a newly identified mechanism where vocabulary specific to a target task domain triggers pre-trained associations that skew predictions).
To address all three bias types, the proposed method estimates a model's holistic bias by passing random, grammatically meaningless word sequences sampled from unlabeled domain data into the prompt, effectively simulating average input length. The model's baseline probability distribution over labels is estimated across twenty sampled text variations and then used to calibrate and normalize final prediction probabilities. The authors validated this approach across 24 standard text classification datasets spanning sentiment analysis, natural language inference, and abusive content detection, using models including GPT-J (6 billion parameters), GPT-3 variants (up to 175 billion parameters), and RoBERTa.
Key findings show that Domain-context Calibration significantly outperforms existing methods. First, the method improved overall classification performance by an average of 20% on GPT-J and 18% on GPT-3 in Macro-F1 scores compared to uncalibrated baselines. Second, on tasks with severe domain-label bias—such as hate speech and content moderation, where uncalibrated models collapsed to random chance—the method delivered massive performance gains of up to 37% for GPT-J and 35% for GPT-3. Third, conventional scaling and standard calibration techniques failed to resolve domain-level bias: expanding model scale to 175 billion parameters, increasing demonstration examples from zero to 16, or adding task instructions did not mitigate severe domain bias, whereas Domain-context Calibration consistently produced stable, superior accuracy. Finally, the calibration proved effective on smaller architectures, improving zero-shot performance in RoBERTa-large by 26% across the evaluation suite.
These results demonstrate that standard in-context learning risks catastrophic classification failures when deployed in domain-specific tasks due to learned word associations, even when using state-of-the-art models. Domain-context Calibration provides a cost-effective, inference-time safeguard that prevents expensive fine-tuning or manual prompt engineering while maintaining robust decision boundaries. The findings also caution against masking performance issues behind single aggregate benchmark scores, as domain bias varies substantially across use cases.
Organizations deploying large language models should immediately integrate domain-based calibration into few-shot and zero-shot classification pipelines. Practitioners should utilize task-indicative label names calibrated with domain-specific text rather than relying on arbitrary placeholder labels or uncalibrated instructions. Calibration requires minimal data overhead, performing effectively with as few as 50 unlabeled task examples. Decision-makers should note that the article’s empirical validation is focused on English text classification tasks and models up to the GPT-3 family; high confidence is warranted for similar text classification workloads, but additional validation is recommended before applying the method to multilingual settings or open-ended text generation.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Its contextual calibration method is a key prior approach to measuring and correcting prompt-induced label bias that this paper evaluates and extends.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Its analysis of how demonstration labels and label spaces affect in-context learning provides useful groundwork for the paper’s distinctions among label-bias types.
- Paper: Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, Yao Lu et al. (2021). Its study of prompt-order sensitivity establishes how demonstration arrangements can destabilize few-shot predictions, motivating the source’s focus on bias calibration.
No sufficiently relevant recommendations were found.
