Prompting is not a substitute for probability measurements in large language models
Jennifer HuRoger Levy
Demonstrates that prompting large language models for metalinguistic judgments misrepresents their underlying linguistic capabilities compared to direct probability measurements, showing why evaluating models through closed APIs without logit access can lead to false conclusions about their knowledge.
Natural language prompting has become the standard method for evaluating the linguistic capabilities of large language models. However, using prompts to ask models about linguistic properties tests a distinct behavioral skill—metalinguistic judgment—rather than reading out what the model actually represents. The article evaluates whether zero-shot metalinguistic prompting is a valid substitute for direct probability measurements of a model's internal vocabulary distributions across a variety of linguistic tasks.
To test this, the researchers evaluated six language models—three open-source Flan-T5 models and three proprietary GPT-3/3.5 models—across four core tasks: next-word prediction, semantic plausibility comparison, isolated sentence acceptability judgments, and comparative sentence evaluations. They tested models using established linguistic benchmarks and recent news data, comparing direct probability measurements against three prompt variations that framed the evaluation as questions or instructions.
Across all experiments, the article found that metalinguistic prompt responses systematically diverge from the quantities directly derived from internal model representations. Direct probability measurements consistently matched or exceeded prompting methods in task accuracy. Furthermore, internal consistency degraded significantly as tasks moved further from direct next-word prediction: correlation coefficients dropped from approximately 0.78–0.79 in simple word prediction down to 0.20–0.24 for complex sentence-level syntactic judgments. The analysis also showed that presenting minimal pairs side by side substantially improved prompting performance compared to asking for isolated sentence evaluations.
These findings demonstrate that negative results from prompt-based evaluations do not provide conclusive evidence that a model lacks a specific linguistic generalization. The failure often stems from an inability to retrieve and verbalize internal representations in response to a prompt, reflecting a gap between underlying model competence and prompt-dependent performance. The trend among commercial providers toward closed application programming interfaces that block access to token probabilities poses a major risk to accurate model evaluation and scientific interpretability.
Decision-makers and evaluation teams should prioritize open models and direct probability measurements over black-box prompting when assessing core linguistic capabilities. When prompt-based methods are unavoidable, practitioners should present comparative pairs rather than isolated items to improve diagnostic reliability. While these findings provide high confidence regarding zero-shot evaluation limitations, further research is needed to determine whether multi-shot examples or reasoning techniques can improve the alignment between internal representations and prompted outputs.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This paper establishes the foundational methodology of querying language models' internal knowledge through fill-in-the-blank probability measurements and reveals how manual prompt phrasing underestimates model capabilities.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This study introduces the LAMA benchmark for extracting internal factual knowledge directly from next-word probability distributions, providing the essential baseline against which metalinguistic prompting is contrasted.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). This work demonstrates the inherent output probability biases of few-shot prompting and formalizes how direct probability calibration is required to assess true model capabilities.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This systematic survey categorizes prompting paradigms and output probability formulation techniques across natural language processing, offering the foundational framework evaluated in the source.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This foundational paper presents diagnostic evaluations of whether models exhibit genuine linguistic knowledge or rely on superficial heuristics, motivating the need for rigorous linguistic evaluation methods.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). This paper develops an unsupervised method to extract latent knowledge directly from internal activations and representations, directly addressing the limitations of prompt-based evaluations identified in the source.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This work critiques the reliability of models' explicit verbalizations by demonstrating that generated reasoning traces unfaithfully diverge from true internal decision mechanisms.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This study investigates the discrepancies between models' metalinguistic evaluation abilities and actual task performance by analyzing how automated LLM evaluators are swayed by superficial output cues.
- Paper: Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?, Nishant Balepur et al. (2024). This research explores how prompt formulations can obscure true model reasoning by testing language models on multiple-choice options in the absence of questions.
