Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations
Chenglei SiDan FriedmanNitish JoshiShi FengDanqi ChenHe He
Reveals how large language models resolve ambiguous in-context demonstrations by favoring semantic features over shallow lexical cues and evaluates the limits of prompt interventions in steering models away from their default feature biases.
Large language models frequently rely on in-context learning, a method where the model adapts to a new task purely from a handful of demonstration examples included in the prompt without updating underlying parameters. Because context windows limit demonstrations to small sample sizes, prompts often remain underspecified, meaning multiple distinct decision rules could explain the provided examples. This dynamic creates uncertainty regarding which patterns models prioritize and poses risks when an AI system relies on unintended shortcuts rather than human intent.
The article evaluates the inherent feature biases of language models during in-context learning and examines the effectiveness of prompting interventions designed to steer models toward intended task features.
To conduct this evaluation, the researchers constructed ambiguous classification prompts using 16 demonstration examples drawn from four natural language processing benchmarks: sentiment analysis, toxicity classification, natural language inference, and question answering. In these prompts, two features perfectly predicted the labels—a primary task feature and an alternative distractor feature, such as text length, word overlap, or punctuation. The researchers evaluated GPT-3 models (the standard base version and an instruction-tuned variant) on 1,200 disambiguating test examples where the two hypotheses conflicted, measuring how often each feature drove the model's predictions. They subsequently evaluated four intervention techniques to redirect model focus: providing natural language task instructions, adding structured step-by-step explanations, using semantically meaningful label words, and introducing partially disambiguating demonstration examples.
The experiments revealed several critical findings. First, language models exhibit distinct, innate feature preferences; for instance, both evaluated models favored sentiment over superficial markers like punctuation in over 90% of baseline evaluations. Second, instruction tuning significantly improves alignment with high-level task goals: the instruction-tuned model preferred natural language inference and question-answering logic over shallow lexical cues by 60% to 75%, whereas the base model favored superficial word overlap in over 58% of cases. Third, while prompt interventions substantially improved performance—raising target feature accuracy by 20 to 34 percentage points on the instruction-tuned model—they were largely successful only when reinforcing weak biases or operating on neutral tasks. Overriding a strong entrenched bias proved difficult; for example, the instruction-tuned model continued prioritizing sentiment even when explicit instructions and disambiguating data favored punctuation. Finally, data-independent interventions such as informative label words and instructions frequently outperformed providing additional unambiguous examples.
These findings indicate that in-context learning operates primarily as a task-recognition mechanism rather than a traditional data-driven learning process. For decision-makers and system architects, this demonstrates that relying solely on few-shot examples creates operational and safety risks, as models may lock onto unintended correlations. Pre-existing inductive biases can significantly reduce the engineering cost of prompt development when aligned with user goals, but they introduce non-trivial compliance and reliability risks when tasks diverge from common pretraining associations.
Organizations deploying few-shot language models should implement standardized prompt designs that combine explicit instructions with semantically descriptive label words instead of generic numerical or arbitrary labels. When utilizing base foundation models, teams should supply structured explanations or unambiguous demonstrations, as raw instructions alone have minimal impact on non-tuned models. For applications with strict accuracy and compliance requirements, practitioners should audit decision pathways across ambiguous edge cases rather than assuming prompt demonstrations suffice.
These conclusions are primarily bounded by the scope of the article, which evaluated two GPT-3 model checkpoints across four binary classification domains using hand-crafted features. While confidence in the observed behavioral patterns is high, further validation is necessary to evaluate newer model families, open-source architectures, multi-class settings, and generative workflows before extending these specific steering strategies into production environments.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). This paper establishes that in-context learning relies heavily on task recognition and formatting rather than learning ground-truth mappings, providing the foundational conceptual premise for evaluating underspecified demonstration biases.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). It identifies innate prediction biases and prompt sensitivities in few-shot large language models, directly preceding the investigation of competing inductive feature preferences.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). It shows that prompting serves primarily to locate pre-existing learned tasks within a model rather than teach new rules at runtime, motivating the study's framing of in-context learning as task recognition.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). It introduces the framework of diagnosing superficial heuristics like lexical overlap versus semantic reasoning in NLP models, which the source adapts to test competing in-context hypotheses.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). It introduces the core paradigm of few-shot in-context learning in large autoregressive language models that the source paper critically evaluates.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It details the instruction-tuning and alignment methodologies that create the InstructGPT variants benchmarked against base models in the source study.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It investigates how subtle input biases systematically sway model reasoning while generated step-by-step explanations remain unfaithful, extending the study of hidden inductive biases to chain-of-thought outputs.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). It evaluates how superficial distractors and irrelevant context derail multi-step reasoning in language models, building directly upon the source's findings regarding model susceptibility to non-target features.
- Paper: Feedback Loops With Language Models Drive In-Context Reward Hacking, Alexander Pan et al. (2024). It explores how interactive feedback loops lead models to exploit underspecified proxy objectives and amplify unintended behaviors during test-time deployment.
- Paper: Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning, Shuai Zhao et al. (2024). It demonstrates how the vulnerability of in-context learning to shortcut features can be maliciously exploited through clean-label demonstration backdoor attacks.
- Paper: In-context Convergence of Transformers, Yu Huang et al. (2024). It provides a theoretical convergence analysis for attention mechanisms under imbalanced feature distributions, offering mathematical insight into why models prioritize dominant features during in-context learning.
