Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Yuhang ZhouPaiheng XuXiaoyu LiuBang AnWei AiFurong Huang

article2024ACL67 citations

Reveals how language models learn shortcut predictions from broader concept-level biases in both fine-tuning and in-context learning, and presents a counterfactual data rebalancing method to eliminate these errors while preserving classification accuracy.

Listen

Modern language models frequently rely on misleading statistical shortcuts rather than genuine comprehension when classifying text. While prior research focused on superficial shortcuts tied to specific words or sentence syntax, the article investigates broader semantic shortcuts, known as concept-level spurious correlations. These occur when high-level concepts—such as food, service, or style—consistently coincide with particular sentiment ratings or outcomes in training data, causing models to misclassify new, unseen inputs that mention those concepts.

The article demonstrates that language models systematically adopt concept-level shortcuts across both model fine-tuning and prompt-based in-context learning. It also introduces and evaluates a practical data-rebalancing framework to mitigate these biases using counterfactual text generation.

To evaluate this issue, the authors analyzed five standard benchmark datasets covering review sentiment and question answering. They used automated large language model prompts to identify and label target concepts across thousands of examples, achieving annotation accuracy exceeding 90% when compared against human baselines. The researchers tested multiple model architectures—including DistilBERT, LLAMA 2 (7B), and GPT-3.5—across both naturally occurring distributions and synthetically skewed environments, while introducing a formal metric (Bias@C) to quantify shortcut reliance.

The analysis yielded several critical findings. First, concept shortcuts exist even in widely used, curated benchmark datasets; for instance, high correlations between concepts like "style" or "music" and positive labels led to substantial baseline bias. Second, when datasets were filtered to make concept distributions fully biased, fine-tuned model accuracy on concept-bearing examples dropped substantially, falling from an average of 79.38% to 74.31%, while bias metrics increased dramatically. Third, prompt demonstrations in few-shot learning also induced concept shortcuts, though the magnitude was smaller than in fine-tuned models. Finally, standard mitigation methods that simply masked associated words failed to resolve the bias because models cluster conceptually related terms together in their internal representations. In contrast, an upsampling technique that injected concepts into counterfactual training examples reduced average bias from 4.90% to 2.74% while increasing target concept accuracy to 80.38%.

These findings indicate that language models deployed for enterprise text classification carry hidden operational and decision-making risks if training datasets contain conceptual imbalances. Conventional preprocessing, such as removing specific trigger words, is insufficient to prevent models from learning high-level conceptual shortcuts. Standard accuracy metrics can also obscure severe underlying bias, as inflated accuracy on majority classes masks steep performance degradation on underrepresented categories.

To address this vulnerability, development teams should audit training pipelines for concept-level imbalances and implement counterfactual data upsampling to balance label distributions before deployment. Organizations relying on few-shot prompting must similarly ensure that in-context exemplars maintain balanced concept-to-label ratios.

These conclusions are supported by consistent results across multiple model sizes and dataset domains. However, some limitations remain: the core evaluations focused primarily on classification tasks, used parameter-efficient tuning for the largest local model rather than full fine-tuning, and relied on automated model-based concept annotation. While confidence in the primary findings is high, further validation is warranted before applying these mitigation techniques to more open-ended generative applications or multimodal systems.

No sufficiently relevant recommendations were found.

Cover for Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Abstract

Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distributions in training data or ICL exemplars. Previous research has primarily concentrated on word, phrase, and syntax features, neglecting the concept level, often due to the absence of concept labels and difficulty in identifying conceptual content in input texts. This paper introduces two main contributions. First, we employ ChatGPT to assign concept labels to texts, assessing concept bias in models during fine-tuning or ICL on test data. We find that LMs, when encountering spurious correlations between a concept and a label in training or prompts, resort to shortcuts for predictions. Second, we introduce a data re-balancing technique that incorporates ChatGPT-generated counterfactual data, thereby balancing label distribution and mitigating spurious correlations. Our method’s efficacy, surpassing traditional token removal approaches, is validated through extensive testing.

Table of Contents

  • 1 Introduction
  • 2 Exploring Concept-level Spurious Correlations
  • 2.1 Obtaining the Concept Labels
  • 2.2 Measuring Concept Spurious Correlations
  • 2.3 Evaluation of Model Robustness to Concept Shortcut in Fine-tuning
  • 2.4 Evaluation of Model Robustness to Concept Shortcut in ICL
  • 3 Dataset Construction and Analysis
  • 3.1 Biased Dataset Construction
  • 3.2 Embedding Analysis of Associated Tokens
  • 4 Results of Spurious Correlation Measurement
  • 4.1 Spurious Correlations in Fine-tuning
  • 4.2 Spurious Correlations in ICL
  • 5 Mitigate Spurious Correlations
  • 5.1 Mitigation via Rebalancing
  • 5.2 Results of Mitigation Methods
  • 6 Related Work
  • 7 Conclusions
  • 8 Limitations
  • Acknowledgments
  • References
  • A Implementation Details
  • A.1 Fine-tuning Experiments
  • A.2 ICL Setup
  • B Prompt Details and Supplementary Results

Knowls

  1. Knowl 1 — Concept-level spurious correlation

    definition

    A concept-level spurious correlation occurs when texts containing a semantic concept are disproportionately associated with a label in training data or in-context demonstrations, and a language model uses that association as a shortcut when classifying new texts. The wording in a test text need not have appeared in training: for example, different expressions such as “salsa” and “Thai steak” may evoke the concept “food” and trigger a learned association between food-related text and a sentiment label. The paper hypothesizes that semantically related expressions can have similar model representations, making the concept-level shortcut possible.

  2. Knowl 2 — Bias@C measures concept-dependent prediction differences

    equation

    Let CC be a concept, Y\mathcal{Y} a finite set of classification labels, and MM a classifier that assigns probabilities to predicted labels. For each true label i∈Yi\in\mathcal{Y}, define ΔC,i\Delta_{C,i} as the difference between the mean probability assigned to the correct label for examples that contain CC and the corresponding mean for examples that do not contain CC, conditioning both groups on having true label ii:

    ΔC,i=E[pM(y^=i∣y=i,C present)]−E[pM(y^=i∣y=i,C absent)].\Delta_{C,i}=\mathbb{E}[p_M(\hat y=i\mid y=i, C\text{ present})]-\mathbb{E}[p_M(\hat y=i\mid y=i, C\text{ absent})].

    The paper’s concept-bias score averages differences in these class-conditional gaps over all unordered label pairs, with the larger label written first:

    Bias@C=1(∣Y∣2)∑i,j∈Y, i>j(ΔC,i−ΔC,j).\mathrm{Bias@C}=\frac{1}{\binom{|\mathcal{Y}|}{2}}\sum_{i,j\in\mathcal{Y},\ i>j}(\Delta_{C,i}-\Delta_{C,j}).

    For binary labels 00 and 11, this reduces to Bias@C=ΔC,1−ΔC,0\mathrm{Bias@C}=\Delta_{C,1}-\Delta_{C,0}. A score near zero indicates little measured reliance on the concept shortcut; a positive score indicates a tendency toward larger labels when the concept is present, and a negative score indicates the opposite. The score compares class-conditional prediction behavior, rather than relying only on the accuracy difference between texts with and without the concept.

  3. Knowl 3 — LLM-assisted concept-label construction

    model/method

    The study assigns concept labels to texts using GPT-3.5 because most of the evaluated datasets lack concept annotations. For datasets without existing concept labels, the researchers first use ChatGPT to identify concepts mentioned in texts and count their frequencies, then select frequent concepts judged not to determine the task label. They use a prompt containing instructions, five labeled examples, the text, and the fixed candidate-concept set; the output can contain no candidate concept or several of them. Each text is annotated twice with temperature 0.7, and an annotation is retained only when both runs agree. CeBaB’s human-annotated concepts are used directly. As a reliability check on CeBaB’s “service” concept, GPT-3.5 annotation achieved 90.4% accuracy against the human labels; the reported average agreement among five human annotators was 92.9%.

  4. Knowl 4 — Evaluation datasets and model settings

    experimental setup

    The experiments cover five classification datasets. Amazon Shoe has 70,117 training and 8,000 test examples, five labels, and the concepts size, color, and style. IMDB has 14,956 training and 4,000 test examples, two labels, and acting, comedy, and music. Yelp has 34,184 training and 4,000 test examples, two labels, and food, price, and service. CeBaB has 7,350 training and 2,000 test examples, five labels, and food, service, and ambiance. BoolQ has 2,393 training and 2,000 test examples, two labels, and country and history. Amazon Shoe and CeBaB use five sentiment labels; the binary sentiment datasets use negative label 0 and positive label 1, and BoolQ uses 0 for No and 1 for Yes. DistilBERT is fully fine-tuned, while LLaMA 2 7B uses LoRA for fine-tuning; GPT-3.5 is evaluated with in-context learning. For group-level utility evaluation, the test data are downsampled to balance labels, and accuracy is reported separately for examples with the concept (Acc@C) and without it (Acc@NoC).

  5. Knowl 5 — Fine-tuning comparison of original and concept-biased data

    model/method

    To test whether training-data imbalance induces concept shortcuts, the researchers fine-tune models separately on the original dataset and on a concept-biased version constructed for each concept. The biased version retains only examples containing that concept whose labels belong to the majority class for that concept in the original data. For the sentiment datasets, the retained class is negative for Amazon Shoe “size” and Yelp “service,” and positive for the other selected concepts; for BoolQ, it is negative for “country” and positive for “history.” They assess the resulting models on test data using Bias@C and the label-balanced Acc@C and Acc@NoC measures. This design tests both whether bias increases in the direction of the deliberately skewed concept-label association and whether that shift affects test accuracy.

  6. Knowl 6 — Fine-tuned models use concept shortcuts

    empirical result

    DistilBERT exhibits concept bias when trained on both original and deliberately concept-biased data. Even the original benchmark datasets produce notable scores for some concepts, including Bias@C = 11.56 for “style” in Amazon Shoe and 12.07 for “music” in IMDB. Training on the concept-biased versions increases the absolute Bias@C in almost every reported concept, with the direction generally following the imposed label association. The paper reports that, across its sentiment-classification comparisons, mean Acc@C falls from 79.38% with original training data to 74.31% with concept-biased data, while mean Acc@NoC falls from 78.08% to 76.56%. Thus, the observed accuracy loss is larger for examples containing the concept. The authors report similar patterns for LLaMA 2 7B, although the size and direction of individual effects vary by concept.

  7. Knowl 7 — In-context prompt construction and evaluation

    experimental setup

    The in-context-learning experiments use GPT-3.5 for sentiment classification with eight labeled demonstrations per prompt. A biased prompt places four examples containing a selected concept with its majority label and four examples without the concept with the contrasting label. A balanced prompt distributes labels evenly within the groups with and without the concept. The experiments compare Bias@C and label-balanced Acc@C and Acc@NoC under the two prompt conditions, repeating each experiment with three different selections of demonstrations and reporting averages. Binary sentiment labels are verbalized as “negative” and “positive”; the original numerical labels are used for tasks with multiple sentiment classes.

  8. Knowl 8 — Biased demonstrations affect in-context predictions

    empirical result

    Across the tested in-context settings, concept-biased demonstrations generally increase the magnitude of Bias@C relative to balanced demonstrations, although the change is not uniform across every concept. In some cases the score changes sign in the direction expected from the demonstrations; for example, Bias@C for Yelp “service” changes from 0.39 to 0.92 when moving from balanced to biased prompts. The paper reports that, on average over its sentiment-classification comparisons, Acc@C decreases from 76.80% to 76.42% and Acc@NoC from 76.37% to 75.58% under biased prompts. Balanced prompts do not eliminate concept bias. The reported changes in bias and accuracy are smaller than in the fine-tuning comparisons, which the authors suggest may be because a small number of demonstrations is less able to establish a strong shortcut.

  9. Knowl 9 — Counterfactual concept injection for data rebalancing

    algorithm

    The proposed upsampling method adds LLM-generated counterfactual examples to a training set so that labels are balanced among texts containing a chosen concept. Its input is an original dataset, a target concept CC, the desired numbers of concept-containing examples for each label, and examples labeled with or without CC. Its output is an augmented dataset with generated concept-containing examples intended to preserve the source examples’ task labels.

    The method identifies label counts needed to balance the examples containing CC. For each label, it samples the required number of examples without CC, then prompts ChatGPT to inject CC into each sampled text while preserving its sentiment or answer. The injection prompt includes five examples with the concept and five without it, along with the text to transform. The generated counterfactuals are added to the original training data. The reported generation setup uses greedy decoding. The paper does not state a computational-complexity bound or a separate automatic acceptance criterion for generated examples.

  10. Knowl 10 — Counterfactual upsampling gives the strongest reported mitigation

    empirical result

    The mitigation comparison evaluates downsampling to balance concept-label groups, counterfactual upsampling, and masking each concept’s top ten associated tokens. For the selected DistilBERT concept cases, downsampling reduces mean absolute Bias@C from 4.90% to 3.43% relative to original-data training, but loses accuracy in 4 of 8 Acc@C comparisons and 5 of 8 Acc@NoC comparisons. Counterfactual upsampling reduces mean absolute Bias@C to 2.74%; mean Acc@C rises from 79.24% to 80.38%, and Acc@NoC is reported as comparable to the original-data result. Masking associated tokens reduces bias in most cases but does not eliminate it, and its Acc@C is lower than counterfactual upsampling in 6 of 8 comparisons. The paper reports similar overall mitigation patterns for LLaMA 2 7B.

  11. Knowl 11 — Concept-associated expressions cluster in fine-tuned representations

    empirical result

    The paper examines whether different expressions of one concept have similar internal representations, a proposed explanation for concept-level shortcuts. For Amazon Shoe, it ranks tokens by pointwise mutual information with each concept, where p(t)p(t) is the fraction of texts containing token tt, p(C)p(C) the fraction containing concept CC, and p(t,C)p(t,C) the fraction containing both: PMI(t,C)=log⁡p(t,C)p(t)p(C)\mathrm{PMI}(t,C)=\log\frac{p(t,C)}{p(t)p(C)}. Examples of highly associated size tokens include “small,” “sizing,” and “9m.” The researchers extract final-layer embeddings for the associated tokens from fine-tuned DistilBERT, average the embeddings of subword pieces when needed, and cluster them using cosine distance and hierarchical clustering. The page 5 clustering visualization shows groups of tokens associated with the same concept, supporting the paper’s hypothesis that semantically related expressions may share representations through which an imbalanced label association can be learned.

  12. Knowl 12 — Scope and limitations of the evidence

    limitation

    The study evaluates five classification datasets and measures concept bias through class-conditional accuracy differences; its results therefore do not establish that the same measurement or effects extend to more complex tasks or modalities. The fine-tuning experiments use DistilBERT and LLaMA 2 7B, with LoRA rather than full fine-tuning for LLaMA 2, and concept annotation relies on GPT-3.5. The reported GPT-3.5 concept-label accuracy is slightly below the human-annotator agreement used for comparison. In the in-context setting, balanced demonstrations do not remove all measured bias, and the paper does not propose a mitigation method for that residual in-context bias.

Coverage note — No other substantial contributed material was omitted; the token-masking baseline and its comparison are included with the mitigation results, while supplementary per-dataset tables are summarized by the paper’s reported aggregate findings and representative values.

References

  1. 1.Eldar D Abraham, Karel D’Oosterlinck, Amir Feder, Yair Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu. 2022. Cebab: Estimating the causal effects of real-world concepts on nlp model behavior. Advances in Neural Information Processing Systems, 35:17582–17596.
  2. 2.Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998.
  3. 3.Ziv Bar-Joseph, David K Gifford, and Tommi S Jaakkola. 2001. Fast optimal leaf ordering for hierarchical clustering. Bioinformatics, 17(suppl_1):S22–S29.
  4. 4.David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022.
  5. 5.Angana Borah, Daria Pylypenko, Cristina Espana-Bonet, and Josef van Genabith. 2023. Measuring spurious correlation in classification:’clever hans’ in translationese. arXiv preprint arXiv:2308.13170.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Oscar Chew, Kuan-Hao Huang, Kai-Wei Chang, and Hsuan-Tien Lin. 2023. Understanding and mitigating spurious correlations in text classification. arXiv preprint arXiv:2305.13654.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  9. 9.Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2020. Learning to model and ignore dataset bias with mixed capacity ensembles. arXiv preprint arXiv:2011.03856.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  11. 11.Chandra Kiran Reddy Evuru, Sreyan Ghosh, Sonal Kumar, Ramaneswaran S, Utkarsh Tyagi, and Dinesh Manocha. 2024. Coda: Constrained generation based data augmentation for low-resource nlp.
  12. 12.Yanbo Fang and Yongfeng Zhang. 2022. Data-efficient concept extraction from pre-trained language models for commonsense explanation generation. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5883–5893.
  13. 13.Dan Friedman, Alexander Wettig, and Danqi Chen. 2022. Finding dataset shortcuts with grammar induction. arXiv preprint arXiv:2210.11560.
  14. 14.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.
  15. 15.Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673.
  16. 16.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056.
  17. 17.He He, Sheng Zha, and Haohan Wang. 2019. Unlearn dataset bias in natural language inference by fitting the residual. arXiv preprint arXiv:1908.10763.
  18. 18.Ruining He and Julian McAuley. 2016. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In proceedings of the 25th international conference on world wide web, pages 507–517.
  19. 19.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  20. 20.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. arXiv preprint arXiv:1707.07328.
  21. 21.Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 8018–8025.
  22. 22.Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. 2019. Learning the difference that makes a difference with counterfactually-augmented data. arXiv preprint arXiv:1909.12434.
  23. 23.Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan, and Ashesh Rambachan. 2018. Algorithmic fairness. In Aea papers and proceedings, volume 108, pages 22–27. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203.
  24. 24.Yuxuan Lai, Chen Zhang, Yansong Feng, Quzhe Huang, and Dongyan Zhao. 2021. Why machine reading comprehension models learn shortcuts? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 989–1002, Online. Association for Computational Linguistics.
  25. 25.Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. 2020. Adversarial filters of dataset biases. In International conference on machine learning, pages 1078–1088. PMLR.
  26. 26.Zongxia Li, Andrew Mao, Daniel Stephens, Pranav Goel, Emily Walpole, Alden Dima, Juan Fung, and Jordan Boyd-Graber. 2024a. Improving the tenor of labeling: Re-evaluating topic models for content analysis.
  27. 27.Zongxia Li, Ishani Mondal, Yijun Liang, Huy Nghiem, and Jordan Lee Boyd-Graber. 2024b. Panda (pedantic answer-correctness determination and adjudication):improving automatic evaluation for question answering and text generation.
  28. 28.Evan Z Liu, Behzad Haghgoo, Annie S Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn. 2021. Just train twice: Improving group robustness without training group information. In International Conference on Machine Learning, pages 6781–6792. PMLR.
  29. 29.Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. 2023a. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566.
  30. 30.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35.
  31. 31.Xiaoyu Liu, Hanlin Lu, Jianbo Yuan, and Xinyu Li. 2023c. Cat: Causal audio transformer for audio classification. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
  32. 32.Xiaoyu Liu, Paiheng Xu, Junda Wu, Jiaxin Yuan, Yifan Yang, Yuhang Zhou, Fuxiao Liu, Tianrui Guan, Haoliang Wang, Tong Yu, et al. 2024. Large language models and causal inference in collaboration: A comprehensive survey. arXiv preprint arXiv:2403.09606.
  33. 33.Xiaoyu Liu, Jiaxin Yuan, Bang An, Yuancheng Xu, Yifan Yang, and Furong Huang. 2023d. C-disentanglement: Discovering causally-independent generative factors under an inductive bias of confounder. arXiv preprint arXiv:2310.17325.
  34. 34.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  35. 35.Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150.
  36. 36.R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. arXiv preprint arXiv:1902.01007.
  37. 37.Matthias Minderer, Olivier Bachem, Neil Houlsby, and Michael Tschannen. 2020. Automatic shortcut removal for self-supervised representation learning. In International Conference on Machine Learning, pages 6927–6937. PMLR.
  38. 38.Xing Niu, Prashant Mathur, Georgiana Dinu, and Yaser Al-Onaizan. 2020. Evaluating robustness to input perturbations for neural machine translation. arXiv preprint arXiv:2005.00580.
  39. 39.OpenAI. 2023. Gpt-4 technical report.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  41. 41.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180–191.
  42. 42.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of nlp models with checklist. arXiv preprint arXiv:2005.04118.
  43. 43.Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang. 2020. An investigation of why overparameterization exacerbates spurious correlations. In International Conference on Machine Learning, pages 8346–8356. PMLR.
  44. 44.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  45. 45.Timo Schick and Hinrich Schütze. 2020. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676.
  46. 46.Rishi Sharma, James Allen, Omid Bakhshandeh, and Nasrin Mostafazadeh. 2018. Tackling the story ending biases in the story cloze test. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 752–757.
  47. 47.Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng, Danqi Chen, and He He. 2023. Measuring inductive biases of in-context learning with underspecified demonstrations. arXiv preprint arXiv:2305.13299.
  48. 48.Ruixiang Tang, Dehan Kong, Longtao Huang, and Hui Xue. 2023. Large language models can be lazy learners: Analyze shortcuts in in-context learning. arXiv preprint arXiv:2305.17256.
  49. 49.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  50. 50.Lifu Tu, Garima Lalwani, Spandana Gella, and He He. 2020. An empirical study on robustness to spurious correlations using pre-trained language models. Transactions of the Association for Computational Linguistics, 8:621–633.
  51. 51.Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych. 2020. Mind the trade-off: Debiasing nlu models without degrading the in-distribution performance. arXiv preprint arXiv:2005.00315.
  52. 52.Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2022. Identifying and mitigating spurious correlations for improving robustness in NLP models. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1719–1729, Seattle, United States. Association for Computational Linguistics.
  53. 53.Xiyao Wang, Jiuhai Chen, Zhaoyang Wang, Yuhang Zhou, Yiyang Zhou, Huaxiu Yao, Tianyi Zhou, Tom Goldstein, Parminder Bhatia, Furong Huang, and Cao Xiao. 2024a. Enhancing visual-language modality alignment in large vision language models via self-improvement.
  54. 54.Xiyao Wang, Yuhang Zhou, Xiaoyu Liu, Hongjin Lu, Yuancheng Xu, Feihong He, Jaehong Yoon, Taixi Lu, Gedas Bertasius, Mohit Bansal, et al. 2024b. Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences. arXiv preprint arXiv:2401.10529.
  55. 55.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  56. 56.Robert Wolfe and Aylin Caliskan. 2021. Low frequency names exhibit bias and overfitting in contextualizing language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 518–532, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  57. 57.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2023. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502.
  58. 58.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  59. 59.Michael Zhang, Nimit S Sohoni, Hongyang R Zhang, Chelsea Finn, and Christopher Ré. 2022. Correct-n-contrast: A contrastive approach for improving robustness to spurious correlations. arXiv preprint arXiv:2203.01517.
  60. 60.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28.
  61. 61.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876.
  62. 62.Jieyu Zhao, Xuezhi Wang, Yao Qin, Jilin Chen, and Kai-Wei Chang. 2022. Investigating ensemble methods for model robustness improvement of text classifiers. arXiv preprint arXiv:2210.16298.
  63. 63.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.
  64. 64.Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023. On large language models’ selection bias in multi-choice questions. arXiv preprint arXiv:2309.03882.
  65. 65.Xiang Zhou and Mohit Bansal. 2020. Towards robustifying NLI models against lexical dataset biases. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8759–8771, Online. Association for Computational Linguistics.
  66. 66.Yuhang Zhou, Suraj Maharjan, and Beiye Liu. 2023. Scalable prompt generation for semi-supervised learning with language models. In Findings of the Association for Computational Linguistics: EACL 2023, pages 758–769.
  67. 67.Jing Zhu, Yuhang Zhou, Vassilis N Ioannidis, Shengyi Qian, Wei Ai, Xiang Song, and Danai Koutra. 2023. Spottarget: Rethinking the effect of target edges for link prediction in graph neural networks. arXiv preprint arXiv:2306.00899.

Citation

MLA
(周宇航), Y. Z., et al. “Explore Spurious Correlations at the Concept Level in Language Models for Text Classification”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 478–92, https://doi.org/10.18653/v1/2024.acl-long.28.
APA
(周宇航), Y. Z., Xu, P., Liu, X., An, B., Ai, W., & Huang, F. (2024). Explore Spurious Correlations at the Concept Level in Language Models for Text Classification. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 478–492. https://doi.org/10.18653/v1/2024.acl-long.28
Chicago
(周宇航), Y. Z., P. Xu, X. Liu, B. An, W. Ai, and F. Huang. 2024. “Explore Spurious Correlations at the Concept Level in Language Models for Text Classification”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 478–92. https://doi.org/10.18653/v1/2024.acl-long.28.
Harvard
(周宇航), Y.Z. et al. (2024) “Explore Spurious Correlations at the Concept Level in Language Models for Text Classification”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 478–492. Available at: https://doi.org/10.18653/v1/2024.acl-long.28.
Vancouver
1. (周宇航) YZ, Xu P, Liu X, An B, Ai W, Huang F (2024) Explore Spurious Correlations at the Concept Level in Language Models for Text Classification. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 478–492

BibTeX

@inproceedings{zhou-etal-2024-explore,
    title = "Explore Spurious Correlations at the Concept Level in Language Models for Text Classification",
    author = "Zhou, Yuhang  and
      Xu, Paiheng  and
      Liu, Xiaoyu  and
      An, Bang  and
      Ai, Wei  and
      Huang, Furong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.28/",
    doi = "10.18653/v1/2024.acl-long.28",
    pages = "478--492"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/