Mitigating Label Biases for In-context Learning

Yu FeiYifan HouZeming ChenAntoine Bosselut

article2023ACL120 citations

Proposes a domain-context calibration method that estimates and removes pre-existing task corpus biases from language models, improving in-context text classification performance by up to 37% Macro-F1.

Listen

Large language models are increasingly used for automated text classification via in-context learning, where models make predictions based on a small context prompt of example demonstration pairs without updating underlying parameters. However, in-context learning is notoriously brittle and vulnerable to systemic label biases—unwanted preferences toward predicting certain target labels regardless of input meaning. When unaddressed, these biases can lead to severely skewed outputs and unpredictable failures in production environments.

The article systematically categorizes label biases in in-context learning and introduces Domain-context Calibration, an efficient, tuning-free method to evaluate and eliminate these distortions. Specifically, the article formalizes three distinct sources of bias: vanilla-label bias (inherent preferences from word frequencies during pre-training), context-label bias (distortions introduced by prompt formatting and example order), and domain-label bias (a newly identified mechanism where vocabulary specific to a target task domain triggers pre-trained associations that skew predictions).

To address all three bias types, the proposed method estimates a model's holistic bias by passing random, grammatically meaningless word sequences sampled from unlabeled domain data into the prompt, effectively simulating average input length. The model's baseline probability distribution over labels is estimated across twenty sampled text variations and then used to calibrate and normalize final prediction probabilities. The authors validated this approach across 24 standard text classification datasets spanning sentiment analysis, natural language inference, and abusive content detection, using models including GPT-J (6 billion parameters), GPT-3 variants (up to 175 billion parameters), and RoBERTa.

Key findings show that Domain-context Calibration significantly outperforms existing methods. First, the method improved overall classification performance by an average of 20% on GPT-J and 18% on GPT-3 in Macro-F1 scores compared to uncalibrated baselines. Second, on tasks with severe domain-label bias—such as hate speech and content moderation, where uncalibrated models collapsed to random chance—the method delivered massive performance gains of up to 37% for GPT-J and 35% for GPT-3. Third, conventional scaling and standard calibration techniques failed to resolve domain-level bias: expanding model scale to 175 billion parameters, increasing demonstration examples from zero to 16, or adding task instructions did not mitigate severe domain bias, whereas Domain-context Calibration consistently produced stable, superior accuracy. Finally, the calibration proved effective on smaller architectures, improving zero-shot performance in RoBERTa-large by 26% across the evaluation suite.

These results demonstrate that standard in-context learning risks catastrophic classification failures when deployed in domain-specific tasks due to learned word associations, even when using state-of-the-art models. Domain-context Calibration provides a cost-effective, inference-time safeguard that prevents expensive fine-tuning or manual prompt engineering while maintaining robust decision boundaries. The findings also caution against masking performance issues behind single aggregate benchmark scores, as domain bias varies substantially across use cases.

Organizations deploying large language models should immediately integrate domain-based calibration into few-shot and zero-shot classification pipelines. Practitioners should utilize task-indicative label names calibrated with domain-specific text rather than relying on arbitrary placeholder labels or uncalibrated instructions. Calibration requires minimal data overhead, performing effectively with as few as 50 unlabeled task examples. Decision-makers should note that the article’s empirical validation is focused on English text classification tasks and models up to the GPT-3 family; high confidence is warranted for similar text classification workloads, but additional validation is recommended before applying the method to multilingual settings or open-ended text generation.

arXiv: 2305.19148fywalter/label-bias

No sufficiently relevant recommendations were found.

Cover for Mitigating Label Biases for In-context Learning

Abstract

Various design settings for in-context learning (ICL), such as the choice and order of the in-context examples, can bias the model’s predictions. While many studies discuss these design choices, there have been few systematic investigations into categorizing them and mitigating their impact. In this work, we define a typology for three types of label biases in ICL for text classification: vanilla-label bias, context-label bias, and domain-label bias (which we conceptualize and detect for the first time).

Our analysis demonstrates that prior label bias calibration methods fall short of addressing all three types of biases. Specifically, domain-label bias restricts LLMs to random-level performance on many tasks regardless of the choice of in-context examples. To mitigate the effect of these biases, we propose a simple bias calibration method that estimates a language model’s label bias using random in-domain words from the task corpus. After controlling for this estimated bias when making predictions, our novel domain-context calibration significantly improves the ICL performance of GPT-J and GPT-3 on a wide range of tasks. The gain is substantial on tasks with large domain-label bias (up to 37% in Macro-F1). Furthermore, our results generalize to models with different scales, pretraining methods, and manually-designed task instructions, showing the prevalence of label biases in ICL.

Table of Contents

  • 1 Introduction
  • 2 Categorizing Label Biases in ICL
  • 2.1 A Typology of Label Biases
  • 3 Domain Label Bias
  • 4 Domain-context Calibration
  • 5 Experimental Setup
  • 6 Experimental Results
  • 6.1 Generalizability
  • 6.2 Analysis
  • 6.3 Zero-shot Prompting
  • 7 Discussion
  • 8 Related Work
  • 9 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Domain-label Bias of All Datasets
  • B Correlation of Domain-label Bias Estimated with Different LLMs
  • C Full Dataset Information
  • D Full Few-shot Results
  • E Templates and Task Instructions
  • F Sampling Analysis
  • G Zero-shot Prompting Experiment
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — Domain-context calibration estimates and removes label priors

    model/method

    Domain-context calibration (DC) adjusts in-context classification probabilities using a model’s estimated label prior on random words sampled from the task corpus. Given unlabeled task texts, DC collects their words into a bag of words, samples random sequences of length LL (the average task-text length), and evaluates the model on MM such sequences, optionally prepended with the same in-context prompt CC used for prediction. For each label yy in label set L\mathcal{L}, it averages the model probabilities to estimate the bias prior:

    PˉM(y∣C)=1M∑j=1MPM(y∣rj,C),\bar P_M(y\mid C)=\frac{1}{M}\sum_{j=1}^{M}P_M(y\mid r_j,C),

    where rjr_j is the jjth random in-domain sequence and PMP_M is the model’s label probability. For an input xix_i, DC predicts

    y^i=arg⁡max⁡y∈LPM(y∣xi,C)PˉM(y∣C).\hat y_i=\arg\max_{y\in\mathcal{L}}\frac{P_M(y\mid x_i,C)}{\bar P_M(y\mid C)}.

    The random sequences are intended to be nearly content-free because they are not grammatical, while retaining the task corpus’s word distribution to capture domain effects. The paper’s main experiments use M=20M=20.

  2. Knowl 2 — Three sources of label bias in in-context classification

    definition

    For a classification task with text input xx, label name yy, and in-context prompt CC, the paper distinguishes three sources of a model’s preference for some labels over others. Vanilla-label bias is a preference for label names independent of the task context, potentially reflecting label-name frequencies in pretraining. Context-label bias arises from the prompt, including the selected or ordered demonstrations and the task template. Domain-label bias arises from the task corpus: learned associations between words in that domain and label names can affect predictions beyond the text-label examples in the prompt. Domain-context calibration is designed to estimate and correct for all three sources, whereas contextual calibration addresses vanilla- and context-label bias but not domain-label bias.

  3. Knowl 3 — A prior-distance measure of domain-label bias

    definition

    The paper measures domain-label bias for a model and classification task by comparing its label distributions on random English words and random in-domain words. Let L\mathcal{L} be the task’s label-name set, and let xEngx_{\mathrm{Eng}} and xi.d.x_{\mathrm{i.d.}} be random English and random in-domain word sequences, respectively, each containing LL words, where LL is the dataset’s average text length. The measure is

    bias=12∑y∈L∣PM(y∣xEng)−PM(y∣xi.d.)∣.\mathrm{bias}=\frac{1}{2}\sum_{y\in\mathcal{L}}\left|P_M(y\mid x_{\mathrm{Eng}})-P_M(y\mid x_{\mathrm{i.d.}})\right|.

    It is the total-variation distance between the two label distributions. In the paper’s hate-detection example, GPT-J’s label prior on random in-domain words was 0.95 for “hate” and 0.05 for “neutral,” while random English words did not produce the same class preference. The measure depends on the model and task formulation, including label names; nevertheless, the authors report high correlations between estimates from different language models.

  4. Knowl 4 — DC improves few-shot performance, especially on high-bias tasks

    empirical result

    The main evaluation covered 24 English text-classification datasets spanning sentiment and topic classification, natural-language inference, and detection. GPT-J (6B) was evaluated on all 24 datasets and GPT-3 (175B) on a budget-limited subset. Unless otherwise specified, prompts used eight randomly sampled training examples, simple unified templates without task instructions, and five random seeds; GPT-3 test sets larger than 500 examples were subsampled to 500. DC used unlabeled test texts to construct its in-domain word pool and averaged 20 random-text priors. Performance was measured with Macro-F1 and compared with chance, uncalibrated predictions, and contextual calibration.

    Across the evaluated datasets, DC raised average Macro-F1 performance by a reported 20% for GPT-J and 18% for GPT-3. The gains were larger as domain-label bias increased. On the highest-bias tasks, DC was the only evaluated calibration method to significantly outperform chance, with reported improvements of up to 37% for GPT-J and 35% for GPT-3 over the other baselines. Contextual calibration could fail to help, and sometimes reduced performance, on these tasks.

  5. Knowl 5 — More demonstrations do not remove bias on tweet classification

    data/table

    These results average Macro-F1 across TweetEval hate, irony, and offensive classification. The comparison tests whether adding demonstrations resolves domain-label bias; DC has the best score at every shot count, while the original and contextual-calibration (CC) scores do not improve consistently as demonstrations are added. Values are reported exactly as in the paper.

    Could not parse LaTeX table
  6. Knowl 6 — DC generalizes across model scales and instruction settings

    empirical result

    On three tweet-classification datasets with large domain-label bias, GPT-3 models from 350M to 175B parameters did not show reliably better original or contextual-calibration performance simply from scaling. DC improved performance at all tested scales and reduced variation across five choices of in-context examples. In a separate comparison, the table reports average Macro-F1 across the three tweet datasets for GPT-3 and instruction-tuned GPT-3 (text-davinci-002), with and without task instructions:

    Could not parse LaTeX table

    DC outperformed the original and CC in all four settings. Task instructions raised the GPT-3 DC score from 61.8 to 68.6; for text-davinci-002, the instruction condition scored 68.1 with DC, compared with 71.9 without instructions.

  7. Knowl 7 — Random-word source and length determine calibration quality

    empirical result

    The paper’s analysis identifies three reasons DC can outperform calibration with a fixed content-free token such as “N/A.” First, a fixed token can itself favor a label: on sentiment tasks, calibrating GPT-J with “N/A” introduced a positive-label preference, which became stronger when repeated; a random English word did not show the same effect. Second, using more random words improved calibration, and the preferred calibration-sequence length tracked task-input length: longer inputs benefited from longer random sequences, and shorter inputs from shorter ones. Third, random in-domain words improved over random English words on tasks with domain-label bias because they better represent the task corpus’s word distribution. On TweetEval hate, increasing the number of in-domain calibration words progressively reduced the model’s tendency to predict “hate”; increasing random English words did not comparably remove that preference.

  8. Knowl 8 — Calibration also improves zero-shot RoBERTa prompting

    empirical result

    The authors applied DC to zero-shot cloze prompting with RoBERTa-large on the same 24 classification datasets. They report a significant 26% performance gain across the suite, showing that the calibration approach is not limited to few-shot autoregressive GPT models. On individual detection datasets, for example, TweetEval hate Macro-F1 was 32.9 without calibration, 30.5 with CC, and 59.0 with DC; TweetEval offensive was 28.4, 32.8, and 56.3, respectively. The results support the paper’s finding that label priors can also impair masked-language-model prompting and can be mitigated with DC.

  9. Knowl 9 — Task-indicative label names work best after calibration

    empirical result

    On TweetEval hate, the authors compared the task-relevant labels “neutral” and “hate,” the less task-relevant antonyms “favor” and “against,” and meaningless placeholders “X” and “Y.” The latter two pairs showed little difference between calibration with random English and in-domain words, indicating little domain-label bias for those label choices. However, after applying DC with in-domain words, “neutral” and “hate” produced the best performance. The manually chosen “favor” and “against” labels performed worse than the meaningless placeholders after DC. The authors therefore find that task-indicative labels can be preferable when their induced bias is calibrated, and that intuitive human label choices do not necessarily work best for a language model.

  10. Knowl 10 — Scope and limitations of the evaluation

    limitation

    The analysis and DC experiments focus on English classification tasks and emphasize word-level domain effects; the paper does not establish how the method behaves for multilingual tasks, open-ended generation, or other domain-associated factors such as topics and gender. The large-model experiments focus on GPT models, while RoBERTa-large is the only smaller model examined. Because of evaluation-budget limits, GPT-3 was tested on only a subset of the 24 datasets. The authors identify broader task types, domain effects, and model families as open areas for evaluation.

Coverage note — The paper’s full per-dataset score tables and detailed dataset/template listings are omitted because the aggregate results and representative per-task evidence capture the main contribution without reproducing the appendix inventory.

References

  1. 1.Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1644–1650, Online. Association for Computational Linguistics.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  4. 4.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  5. 5.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Machine learning challenges workshop, pages 177–190. Springer.
  6. 6.Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018. Hate speech dataset from a white supremacy forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2), pages 11–20, Brussels, Belgium. Association for Computational Linguistics.
  7. 7.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124.
  8. 8.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.
  9. 9.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. arXiv preprint arXiv:2104.08315.
  10. 10.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 168–177.
  11. 11.Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning.
  12. 12.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al. 2021. Datasets: A community library for natural language processing. arXiv preprint arXiv:2109.02846.
  13. 13.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804.
  14. 14.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  15. 15.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  16. 16.Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala. 2014. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65(4):782–796.
  17. 17.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. A sick cure for the evaluation of compositional distributional semantic models. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 216–223.
  18. 18.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  19. 19.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021. Reframing instructional prompts to gptk’s language. arXiv preprint arXiv:2109.07830.
  20. 20.Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. 2022. Ethos: a multi-label hate speech detection dataset. Complex & Intelligent Systems, pages 1–16.
  21. 21.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  22. 22.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058.
  23. 23.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. arXiv preprint cs/0506075.
  24. 24.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206.
  25. 25.Timo Schick and Hinrich Schütze. 2020. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676.
  26. 26.Emily Sheng and David Uthus. 2020. Investigating societal biases in a poetry composition system. In Proceedings of the Second Workshop on Gender Bias in Natural Language Processing, pages 93–106, Barcelona, Spain (Online). Association for Computational Linguistics.
  27. 27.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  28. 28.Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate. 2022. An information-theoretic approach to prompt engineering without ground truth labels. arXiv preprint arXiv:2203.11364.
  29. 29.Ellen M Voorhees and Dawn M Tice. 2000. Building a question answering test collection. In Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, pages 200–207.
  30. 30.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  31. 31.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  32. 32.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846.
  33. 33.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  34. 34.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28.
  35. 35.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.

Citation

MLA
Fei, Y., et al. “Mitigating Label Biases for In-context Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 14014–31, https://doi.org/10.18653/v1/2023.acl-long.783.
APA
Fei, Y., Hou, Y., Chen, Z., & Bosselut, A. (2023). Mitigating Label Biases for In-context Learning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14014–14031. https://doi.org/10.18653/v1/2023.acl-long.783
Chicago
Fei, Y., Y. Hou, Z. Chen, and A. Bosselut. 2023. “Mitigating Label Biases for In-context Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14014–31. https://doi.org/10.18653/v1/2023.acl-long.783.
Harvard
Fei, Y. et al. (2023) “Mitigating Label Biases for In-context Learning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14014–14031. Available at: https://doi.org/10.18653/v1/2023.acl-long.783.
Vancouver
1. Fei Y, Hou Y, Chen Z, Bosselut A (2023) Mitigating Label Biases for In-context Learning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14014–14031

BibTeX

@inproceedings{fei-etal-2023-mitigating,
    title = "Mitigating Label Biases for In-context Learning",
    author = "Fei, Yu  and
      Hou, Yifan  and
      Chen, Zeming  and
      Bosselut, Antoine",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.783/",
    doi = "10.18653/v1/2023.acl-long.783",
    pages = "14014--14031"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/