Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations

Chenglei SiDan FriedmanNitish JoshiShi FengDanqi ChenHe He

article2023ACL65 citationsOutstanding Paper Award

Reveals how large language models resolve ambiguous in-context demonstrations by favoring semantic features over shallow lexical cues and evaluates the limits of prompt interventions in steering models away from their default feature biases.

Listen

Large language models frequently rely on in-context learning, a method where the model adapts to a new task purely from a handful of demonstration examples included in the prompt without updating underlying parameters. Because context windows limit demonstrations to small sample sizes, prompts often remain underspecified, meaning multiple distinct decision rules could explain the provided examples. This dynamic creates uncertainty regarding which patterns models prioritize and poses risks when an AI system relies on unintended shortcuts rather than human intent.

The article evaluates the inherent feature biases of language models during in-context learning and examines the effectiveness of prompting interventions designed to steer models toward intended task features.

To conduct this evaluation, the researchers constructed ambiguous classification prompts using 16 demonstration examples drawn from four natural language processing benchmarks: sentiment analysis, toxicity classification, natural language inference, and question answering. In these prompts, two features perfectly predicted the labels—a primary task feature and an alternative distractor feature, such as text length, word overlap, or punctuation. The researchers evaluated GPT-3 models (the standard base version and an instruction-tuned variant) on 1,200 disambiguating test examples where the two hypotheses conflicted, measuring how often each feature drove the model's predictions. They subsequently evaluated four intervention techniques to redirect model focus: providing natural language task instructions, adding structured step-by-step explanations, using semantically meaningful label words, and introducing partially disambiguating demonstration examples.

The experiments revealed several critical findings. First, language models exhibit distinct, innate feature preferences; for instance, both evaluated models favored sentiment over superficial markers like punctuation in over 90% of baseline evaluations. Second, instruction tuning significantly improves alignment with high-level task goals: the instruction-tuned model preferred natural language inference and question-answering logic over shallow lexical cues by 60% to 75%, whereas the base model favored superficial word overlap in over 58% of cases. Third, while prompt interventions substantially improved performance—raising target feature accuracy by 20 to 34 percentage points on the instruction-tuned model—they were largely successful only when reinforcing weak biases or operating on neutral tasks. Overriding a strong entrenched bias proved difficult; for example, the instruction-tuned model continued prioritizing sentiment even when explicit instructions and disambiguating data favored punctuation. Finally, data-independent interventions such as informative label words and instructions frequently outperformed providing additional unambiguous examples.

These findings indicate that in-context learning operates primarily as a task-recognition mechanism rather than a traditional data-driven learning process. For decision-makers and system architects, this demonstrates that relying solely on few-shot examples creates operational and safety risks, as models may lock onto unintended correlations. Pre-existing inductive biases can significantly reduce the engineering cost of prompt development when aligned with user goals, but they introduce non-trivial compliance and reliability risks when tasks diverge from common pretraining associations.

Organizations deploying few-shot language models should implement standardized prompt designs that combine explicit instructions with semantically descriptive label words instead of generic numerical or arbitrary labels. When utilizing base foundation models, teams should supply structured explanations or unambiguous demonstrations, as raw instructions alone have minimal impact on non-tuned models. For applications with strict accuracy and compliance requirements, practitioners should audit decision pathways across ambiguous edge cases rather than assuming prompt demonstrations suffice.

These conclusions are primarily bounded by the scope of the article, which evaluated two GPT-3 model checkpoints across four binary classification domains using hand-crafted features. While confidence in the observed behavioral patterns is high, further validation is necessary to evaluate newer model families, open-source architectures, multi-class settings, and generative workflows before extending these specific steering strategies into production environments.

Cover for Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations

Abstract

In-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood. We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of under-specified demonstrations in which two features are equally predictive of the labels. First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations. We find that LLMs exhibit clear feature biases—for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation. Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words. We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases. Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task.

Table of Contents

  • 1 Introduction
  • 2 Setup
  • 2.1 Measuring Feature Biases
  • 2.2 In-Context Learning
  • 3 Data Construction
  • 4 Measuring Feature Biases
  • 4.1 Experiment Details
  • 4.2 Results
  • 5 Comparing Interventions
  • 5.1 Experiment Details
  • 5.2 Results
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Acknowledgement
  • References
  • A Appendix
  • A.1 Intervention Results Across Features
  • A.2 List of Semantic Verbalizers
  • A.3 List of Task Instructions
  • A.4 List of Template Explanations
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — Underspecified Demonstration Framework for Measuring In-Context Feature Bias

    model/method

    To evaluate the inductive feature biases of large language models during in-context learning (ICL), a behavioral testing framework presents the model with demonstrations where two distinct feature hypotheses predict the labels equally well, and then evaluates the induced classifier on examples where the two hypotheses make opposing predictions.

    Formally, let x∈Xx \in \mathcal{X} denote a text input and let h1,h2:X→{0,1}h_1, h_2 : \mathcal{X} \to \{0, 1\} denote two binary feature functions. A learning algorithm ℓ\ell maps an in-context demonstration dataset Ddemo⊂X×{0,1}\mathcal{D}_{\text{demo}} \subset \mathcal{X} \times \{0, 1\} to an induced classifier f:X→{0,1}f : \mathcal{X} \to \{0, 1\}. The demonstration dataset Ddemo\mathcal{D}_{\text{demo}} is constructed such that for every demonstration example (x,y)∈Ddemo(x, y) \in \mathcal{D}_{\text{demo}}, the label satisfies y=h1(x)=h2(x)y = h_1(x) = h_2(x), and the demonstration labels are balanced between 00 and 11. Consequently, both hypotheses achieve identical 100%100\% training accuracy on Ddemo\mathcal{D}_{\text{demo}}.

    To identify which feature hypothesis the model prefers, the resulting classifier f=ℓ(Ddemo)f = \ell(\mathcal{D}_{\text{demo}}) is tested on a balanced disambiguating dataset Dtest\mathcal{D}_{\text{test}} where the two hypotheses strictly conflict:

    h1(x)≠h2(x)∀x∈Dtesth_1(x) \neq h_2(x) \quad \forall x \in \mathcal{D}_{\text{test}}

    The degree to which f(x)f(x) matches h1(x)h_1(x) versus h2(x)h_2(x) on Dtest\mathcal{D}_{\text{test}} reveals the model's inductive feature bias.

  2. Knowl 2 — In-Context Feature Accuracy Metric

    equation

    Given a classifier f:X→{0,1}f: \mathcal{X} \to \{0, 1\} induced via in-context learning on underspecified demonstrations, and a binary feature function h:X→{0,1}h: \mathcal{X} \to \{0, 1\}, the alignment between the predictions of ff and the feature hh on a disambiguating test dataset Dtest\mathcal{D}_{\text{test}} is quantified by the feature accuracy (hh-accuracy):

    h-accuracy=1∣Dtest∣∑x∈Dtest1[f(x)=h(x)]h\text{-accuracy} = \frac{1}{|\mathcal{D}_{\text{test}}|} \sum_{x \in \mathcal{D}_{\text{test}}} \mathbf{1}[f(x) = h(x)]

    where 1[⋅]\mathbf{1}[\cdot] is the indicator function.

    In a binary feature setting with two competing feature functions h1h_1 and h2h_2, the disambiguating dataset satisfies h1(x)≠h2(x)h_1(x) \neq h_2(x) for all x∈Dtestx \in \mathcal{D}_{\text{test}}, which guarantees that:

    h1-accuracy+h2-accuracy=1.0h_1\text{-accuracy} + h_2\text{-accuracy} = 1.0

    When Dtest\mathcal{D}_{\text{test}} is balanced between instances where (h1(x)=0,h2(x)=1)(h_1(x)=0, h_2(x)=1) and instances where (h1(x)=1,h2(x)=0)(h_1(x)=1, h_2(x)=0), an h1-accuracyh_1\text{-accuracy} of 0.50.5 (or 50%50\%) indicates that the model displays no inductive bias toward either feature. An h1-accuracy>0.5h_1\text{-accuracy} > 0.5 indicates a preference for feature h1h_1, while h1-accuracy<0.5h_1\text{-accuracy} < 0.5 indicates a preference for feature h2h_2.

  3. Knowl 3 — Experimental Benchmark for In-Context Feature Bias Evaluation

    experimental setup

    The benchmark evaluates feature biases across four NLP classification domains by pairing a default semantic task label (h1h_1) against alternative distractor features (h2h_2):

    1. Sentiment Analysis (IMDb and Yelp): Default h1h_1 is review sentiment (positive vs. negative). Distractors (h2h_2) include domain source (IMDb vs. Yelp), text length (<< or ≥40\ge 40 words), terminal punctuation (exclamation mark vs. period), keyword occurrence ("nice" or "food"), and word capitalization (all lowercase vs. contains all-uppercase words).
    2. Toxicity Classification (CivilComments): Default h1h_1 is comment toxicity (toxic vs. non-toxic). Distractors (h2h_2) include targeted identity groups based on human metadata (gender [female vs. male], sexuality [LGBTQ vs. non-LGBTQ], religion [Muslim vs. Christian/Jewish], race [Black vs. White, Asian vs. White]), text length, and word capitalization.
    3. Natural Language Inference (MultiNLI): Default h1h_1 is premise-hypothesis entailment (entailment vs. non-entailment, collapsing neutral and contradiction into non-entailment). Distractors (h2h_2) include text genre/domain (government vs. fiction/telephone), lexical overlap (whether the premise contains all hypothesis words), hypothesis length relative to the premise, and presence of negation words in the hypothesis.
    4. Question Answering (BoolQ): Default h1h_1 is the answer (yes vs. no). Distractors (h2h_2) include question initial word ("is/was" vs. "do/does/did"), question-passage lexical overlap, question structure (containing comparison phrase "same as"), and passage length (<< or ≥50\ge 50 words).

    Each evaluation prompt contains 1616 balanced demonstration examples (88 where h1=h2=0h_1=h_2=0 and 88 where h1=h2=1h_1=h_2=1), formatted with template t(x)="Input: $x Label: "t(x)=\text{"Input: \$x Label: "} and verbalizer v(0)="0",v(1)="1"v(0)=\text{"0"}, v(1)=\text{"1"}. Demonstrations are randomly ordered and label-calibrated. Performance is evaluated over 33 random demonstration seeds on 1,2001{,}200 balanced test examples using greedy decoding (temperature 0) on GPT-3 base (davinci) and instruction-tuned (text-davinci-002) models.

  4. Knowl 4 — Baseline Feature Biases in Pretrained Versus Instruction-Tuned GPT-3

    empirical result

    When presented with underspecified prompts without any prompt modifications or steering interventions, GPT-3 models demonstrate clear, distinct inductive biases:

    1. Strong Sentiment Bias: On sentiment analysis (IMDb + Yelp), both base davinci and instruction-tuned text-davinci-002 strongly favor sentiment (h1h_1) over distractor features (h2h_2). Specifically, against length, davinci and text-davinci-002 achieve 98.4%98.4\% and 98.6%98.6\% h1h_1-accuracy; against punctuation, 98.3%98.3\% and 99.1%99.1\%; against keywords, 95.5%95.5\% and 97.8%97.8\%; and against capitalization, 92.0%92.0\% and 98.4%98.4\%.
    2. Low Feature Bias in Toxicity: On CivilComments, both models exhibit near-neutral feature preferences, with aggregate h1h_1-accuracy of 45.8%45.8\% for davinci and 45.9%45.9\% for text-davinci-002, indicating no strong intrinsic preference between toxicity and demographic identity mentions, length, or capitalization.
    3. Divergence on Sentence-Pair Tasks: On MultiNLI and BoolQ, base davinci exhibits a strong preference for shallow heuristic distractors rather than the default semantic task, scoring an aggregate h1h_1-accuracy of only 31.9%31.9\% on MultiNLI (preferring genre at 88.5%88.5\% and lexical overlap at 69.5%69.5\%) and 38.2%38.2\% on BoolQ (preferring question structure at 66.0%66.0\% and question word at 60.2%60.2\%). In contrast, text-davinci-002 strongly prefers the semantic task labels, achieving 66.1%66.1\% aggregate h1h_1-accuracy on MultiNLI and 73.4%73.4\% on BoolQ.
  5. Knowl 5 — Prompt Intervention Strategies for Steering Inductive Biases

    model/method

    To steer a large language model toward an intended feature hypothesis (h∗∈{h1,h2}h^* \in \{h_1, h_2\}) when demonstrations are underspecified, four separate prompt interventions are applied:

    1. Semantic Verbalizers: The verbalizer function v(y)v(y) is changed from abstract digits ("0", "1") to semantically meaningful words corresponding to h∗h^*. For example, when steering towards sentiment, v(0)="negative"v(0)=\text{"negative"} and v(1)="positive"v(1)=\text{"positive"}; when steering towards length, v(0)="long"v(0)=\text{"long"} and v(1)="short"v(1)=\text{"short"}; or a short question prefix is added (e.g., "Negation? "v(0)="no",v(1)="yes"\text{"Negation? "} v(0)=\text{"no"}, v(1)=\text{"yes"}).
    2. Natural Language Instructions: A task instruction string ss is prepended to the prompt, explicitly instructing the model on the intended feature decision rule (e.g., "Given a review, you need to predict whether the review ends with an exclamation mark ('!') or period ('.')").
    3. Template-Based Explanations: A short explanatory sentence e(x)e(x) is inserted after each demonstration input t(x)t(x) and before the label v(y)v(y), articulating why the label was assigned according to the intended feature (e.g., "The review ends with an exclamation mark. Therefore, the answer is 1.").
    4. Disambiguating Demonstrations (Supervised Data Baseline): Half of the 1616 demonstration examples are replaced with disambiguating instances where h1(x)≠h2(x)h_1(x) \neq h_2(x) and y=h∗(x)y = h^*(x), providing unambiguous supervised evidence for the intended feature.
  6. Knowl 6 — Comparative Steerability of Base Versus Instruction-Tuned Models Across Interventions

    data/table

    The impact of different prompt intervention strategies on steering large language models toward either the default feature (h1h_1) or an alternative distractor feature (h2h_2) is measured by the average change in feature accuracy relative to baseline (Δh\Delta h-accuracy, in percentage points), aggregated across features and datasets:

    Steer towards h1h_1 Steer towards h2h_2
    Intervention Davinci TD002 Davinci TD002
    Baseline (accuracy %) 39.5 59.1 46.9 30.5
    Verbalizer (Δ\Delta) +11.9 +7.1 +15.6 +24.4
    Instruction (Δ\Delta) +1.6 +12.2 -2.4 +24.2
    Explanation (Δ\Delta) +14.4 +6.9 +14.3 +33.8
    Disambig (Δ\Delta) +12.9 +9.4 +18.6 +21.1

    Note: TD002 refers to text-davinci-002. The "Steer towards h1h_1" column excludes sentiment analysis datasets because both models already exhibit a near-ceiling baseline preference for h1h_1.

    Key takeaways:

    • On the non-instruction-tuned model (davinci), natural language instructions are ineffective for steering (+1.6% for h1h_1, -2.4% for h2h_2), whereas template explanations (+14.4% for h1h_1, +14.3% for h2h_2) and semantic verbalizers (+11.9% for h1h_1, +15.6% for h2h_2) successfully guide the model.
    • On the instruction-tuned model (text-davinci-002), all interventions are effective, with template explanations producing the largest overall improvement (+33.8% when steering towards h2h_2) and instructions providing substantial gains (+12.2% for h1h_1, +24.2% for h2h_2).
    • Prompt-based interventions (semantic verbalizers, explanations) frequently yield comparable or greater shifts in feature bias than providing unambiguous data demonstrations (Disambig).
  7. Knowl 7 — Asymmetry in Overriding Strong Prior Inductive Biases

    empirical result

    Prompt interventions exhibit marked asymmetry in their ability to steer model predictions: they succeed when aligning with an existing bias or when initial bias is low, but struggle to override strong conflicting prior feature biases.

    1. Reinforcing Existing or Neutral Biases: When steering text-davinci-002 toward the default semantic label h1h_1 on MultiNLI (baseline 66.1%66.1\%) and BoolQ (baseline 73.4%73.4\%), interventions readily reinforce the bias, pushing h1h_1-accuracy to 81.0%81.0\% (instructions) and 82.0%82.0\% (explanations). Similarly, on CivilComments, where the baseline bias is neutral (45.9%45.9\% h1h_1-accuracy / 54.1%54.1\% h2h_2-accuracy), explanations successfully steer the model to 60.6%60.6\% h1h_1-accuracy or 86.5%86.5\% h2h_2-accuracy.
    2. Failure to Override Strong Semantic Priors: When an intervention attempts to steer text-davinci-002 toward a surface feature (h2h_2) on sentiment reviews against a strong sentiment prior (baseline h1h_1-accuracy >97%>97\%), no intervention consistently establishes h2h_2 as the majority rule. For instance, when steering toward terminal punctuation, instructions achieve only 15.6%15.6\% h2h_2-accuracy and template explanations reach 49.5%49.5\% h2h_2-accuracy (up from a baseline of 0.9%0.9\%), failing to reliably overturn the sentiment hypothesis.
  8. Knowl 8 — Differential Steerability of Semantic Versus Non-Semantic Surface Features

    empirical result

    The effectiveness of prompt interventions depends strongly on the semantic relevance of the target feature. On MultiNLI, interventions are substantially more effective at steering models toward semantic attributes than toward non-semantic surface or lexical rules:

    • Semantic Distractors (Genre/Domain): For text-davinci-002, where baseline h2h_2-accuracy is 28.7%28.7\% for Government vs. Fiction and 39.5%39.5\% for Government vs. Telephone, applying semantic verbalizers increases h2h_2-accuracy by +66.8%+66.8\% (to 96.6%96.6\%) and +58.2%+58.2\% (to 99.5%99.5\%), respectively. Template explanations increase h2h_2-accuracy by +47.7%+47.7\% and +51.8%+51.8\%.
    • Non-Semantic Surface Distractors (Length and Lexical Overlap): For Lexical Overlap on text-davinci-002 (baseline 39.2%39.2\%), verbalizers yield Δ=−3.2%\Delta = -3.2\%, instructions yield Δ=−3.8%\Delta = -3.8\%, and explanations yield Δ=+3.1%\Delta = +3.1\%. For Hypothesis Length (baseline 27.3%27.3\%), verbalizers yield Δ=−4.5%\Delta = -4.5\% and instructions yield Δ=+6.8%\Delta = +6.8\%, although explanations achieve +26.5%+26.5\%.
    • Similarly, on base davinci, verbalizers and instructions fail to increase preference for lexical overlap (Δ=−6.6%\Delta = -6.6\% and −6.3%-6.3\%) or hypothesis length (Δ=−2.0%\Delta = -2.0\% and −1.1%-1.1\%).

    Features unrelated to the semantic meaning of sentences (such as token overlap and sentence length) are consistently harder to activate through prompting than semantic domain or topic features.

  9. Knowl 9 — Stated Limitations of the Feature Bias Study

    limitation

    The study notes several specific limitations regarding its methodology and scope:

    1. Model Scope: Evaluations are conducted exclusively on GPT-3 checkpoints (davinci and text-davinci-002); models from other families (such as OPT or BLOOM) or varying scales and pretraining mixtures may exhibit different feature bias profiles and intervention sensitivities.
    2. Task and Label Dimensionality: Analysis is restricted to four standard classification datasets converted into binary classification tasks (h1,h2:X→{0,1}h_1, h_2 : \mathcal{X} \to \{0, 1\}); multi-class classification and generation tasks are left unaddressed.
    3. Feature Selection: Competing feature hypotheses are manually designed based on known heuristic artifacts rather than discovered automatically from raw data distributions.
    4. Causal Attribution to Pretraining: The experiments observe behavioral biases but do not establish formal causal links back to specific corpora or pretraining distributions.
    5. Generalization to Fairness: Low measured feature bias on CivilComments demographic attributes should not be extrapolated as general evidence of demographic neutrality or fairness in downstream deployments.

Coverage note — The full verbatim text strings for all task instructions and template explanations in Appendix Tables 6-10 were omitted as representative examples were incorporated into the method and empirical knowls.

References

  1. 1.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2023. What learning algorithm is in-context learning? Investigations with linear models. In International Conference on Learning Representations (ICLR).
  2. 2.Nabiha Asghar. 2016. Yelp dataset challenge: Review rating prediction. arXiv preprint arXiv:1605.05362.
  3. 3.Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019. Nuanced metrics for measuring unintended bias with real data for text classification. In Companion Proceedings of the 2019 World Wide Web Conference.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  5. 5.Stephanie CY Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K Lampinen, and Felix Hill. 2022. Transformers generalize differently from information stored in context vs in weights. In Advances in Neural Information Processing Systems (NeurIPS).
  6. 6.Noam Chomsky. 1980. Rules and representations. Behavioral and brain sciences, 3(1):1–15.
  7. 7.Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  9. 9.Ishita Dasgupta, Erin Grant, and Tom Griffiths. 2022. Distinguishing rule and exemplar-based generalization in learning systems. In International Conference on Machine Learning (ICML), pages 4816–4830. PMLR.
  10. 10.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Association for Computational Linguistics (ACL).
  11. 11.Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673.
  12. 12.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In North American Chapter of the Association for Computational Linguistics (NAACL).
  13. 13.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019. Using pre-training can improve model robustness and uncertainty. In International Conference on Machine Learning (ICML).
  14. 14.Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Xiaodong Song. 2020. Pretrained transformers improve out-of-distribution robustness. In Association for Computational Linguistics (ACL).
  15. 15.Nitish Joshi, Xiang Pan, and He He. 2022. Are all spurious features in natural language alike? an analysis through a causal lens. In Empirical Methods in Natural Language Processing (EMNLP).
  16. 16.Andrew Kyle Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, and Felix Hill. 2022. Can language models learn from explanations in context? In Findings of Empirical Methods in Natural Language Processing (EMNLP).
  17. 17.Charles Lovering, Rohan Jha, Tal Linzen, and Ellie Pavlick. 2021. Predicting inductive biases of pre-trained models. In International Conference on Learning Representations (ICLR).
  18. 18.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Association for Computational Linguistics (ACL).
  19. 19.Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Association for Computational Linguistics: Human Language Technologies (ACL-HLT).
  20. 20.R Thomas McCoy, Robert Frank, and Tal Linzen. 2018. Revisiting the poverty of the stimulus: Hierarchical generalization without a hierarchical bias in recurrent neural networks. In CogSci.
  21. 21.R. Thomas McCoy, Robert Frank, and Tal Linzen. 2020. Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks. Transactions of the Association for Computational Linguistics, 8:125–140.
  22. 22.R. Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Association for Computational Linguistics (ACL).
  23. 23.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In Empirical Methods in Natural Language Processing (EMNLP).
  24. 24.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Natural instructions: Benchmarking generalization to new tasks from natural language instructions. ArXiv, abs/2104.08773.
  25. 25.Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022. Coloring the blank slate: Pre-training imparts a hierarchical inductive bias to sequence-to-sequence models. In Findings of Association for Computational Linguistics (ACL).
  26. 26.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155.
  27. 27.Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen. 2023. What in-context learning 'learns' in-context: Disentangling task recognition and task learning. In Findings of Association for Computational Linguistics (ACL).
  28. 28.Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron C. Wallace. 2022. Combining feature and instance attribution to detect artifacts. In Findings of Association for Computational Linguistics (ACL).
  29. 29.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  30. 30.Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang. 2020. An investigation of why over-parameterization exacerbates spurious correlations. In International Conference on Machine Learning, pages 8346–8356. PMLR.
  31. 31.Nikunj Saunshi, Sadhika Malladi, and Sanjeev Arora. 2021. A mathematical exploration of why language models help solve downstream tasks. In International Conference on Learning Representations.
  32. 32.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  33. 33.Weijia Shi, Julian Michael, Suchin Gururangan, and Luke Zettlemoyer. 2022. kNN-Prompt: Nearest neighbor zero-shot inference. In Empirical Methods in Natural Language Processing (EMNLP).
  34. 34.Kazutoshi Shinoda, Saku Sugawara, and Akiko Aizawa. 2022. Which shortcut solution do question answering models prefer to learn? ArXiv, abs/2211.16220.
  35. 35.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan L. Boyd-Graber, and Lijuan Wang. 2023. Prompting GPT-3 to be reliable. In International Conference on Learning Representations (ICLR).
  36. 36.Alex Tamkin, Kunal Handa, Avash Shrestha, and Noah Goodman. 2023. Task ambiguity in humans and language models. In International Conference on Learning Representations (ICLR).
  37. 37.Lifu Tu, Garima Lalwani, Spandana Gella, and He He. 2020. An empirical study on robustness to spurious correlations using pre-trained language models. Transactions of the Association for Computational Linguistics, 8:621–633.
  38. 38.Elena Voita and Ivan Titov. 2020. Information-theoretic probing with minimum description length. In Empirical Methods in Natural Language Processing (EMNLP).
  39. 39.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2022. Transformers learn in-context by gradient descent. arXiv preprint arXiv:2212.07677.
  40. 40.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, M. Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddharth Deepak Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hanna Hajishirzi, and Daniel Khashabi. 2022. Supernaturalinstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Empirical Methods in Natural Language Processing (EMNLP).
  41. 41.Alex Warstadt and Samuel R. Bowman. 2020. Can neural networks acquire a structural bias from raw linguistic data? In CogSci.
  42. 42.Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, and Samuel Bowman. 2020. Learning which features matter: RoBERTa acquires a preference for linguistic generalizations (eventually). In Empirical Methods in Natural Language Processing (EMNLP), pages 217–235.
  43. 43.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In North American Chapter of the Association for Computational Linguistics (NAACL).
  44. 44.Colin Wei, Sang Michael Xie, and Tengyu Ma. 2021. Why do pretrained language models help in downstream tasks? An analysis of head and prompt tuning. Advances in Neural Information Processing Systems (NeurIPS).
  45. 45.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS).
  46. 46.Jerry W. Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma. 2023. Larger language models do in-context learning differently. ArXiv, abs/2303.03846.
  47. 47.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  48. 48.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations (ICLR).
  49. 49.Xi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot prompting for textual reasoning. In Advances in Neural Information Processing Systems (NeurIPS).
  50. 50.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open pre-trained transformer language models. ArXiv, abs/2205.01068.
  51. 51.Tianyi Zhang and Tatsunori B Hashimoto. 2021. On the inductive bias of masked language modeling: From statistical to syntactic dependencies. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  52. 52.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning (ICML).

Citation

MLA
Si, C., et al. “Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11289–310, https://doi.org/10.18653/v1/2023.acl-long.632.
APA
Si, C., Friedman, D., Joshi, N., Feng, S., Chen, D., & He, H. (2023). Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11289–11310. https://doi.org/10.18653/v1/2023.acl-long.632
Chicago
Si, C., D. Friedman, N. Joshi, S. Feng, D. Chen, and H. He. 2023. “Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11289–310. https://doi.org/10.18653/v1/2023.acl-long.632.
Harvard
Si, C. et al. (2023) “Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11289–11310. Available at: https://doi.org/10.18653/v1/2023.acl-long.632.
Vancouver
1. Si C, Friedman D, Joshi N, Feng S, Chen D, He H (2023) Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11289–11310

BibTeX

@inproceedings{si-etal-2023-measuring,
    title = "Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations",
    author = "Si, Chenglei  and
      Friedman, Dan  and
      Joshi, Nitish  and
      Feng, Shi  and
      Chen, Danqi  and
      He, He",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.632/",
    doi = "10.18653/v1/2023.acl-long.632",
    pages = "11289--11310"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/