Prompting is not a substitute for probability measurements in large language models

Jennifer HuRoger Levy

article2023EMNLP60 citations

Demonstrates that prompting large language models for metalinguistic judgments misrepresents their underlying linguistic capabilities compared to direct probability measurements, showing why evaluating models through closed APIs without logit access can lead to false conclusions about their knowledge.

Listen

Natural language prompting has become the standard method for evaluating the linguistic capabilities of large language models. However, using prompts to ask models about linguistic properties tests a distinct behavioral skill—metalinguistic judgment—rather than reading out what the model actually represents. The article evaluates whether zero-shot metalinguistic prompting is a valid substitute for direct probability measurements of a model's internal vocabulary distributions across a variety of linguistic tasks.

To test this, the researchers evaluated six language models—three open-source Flan-T5 models and three proprietary GPT-3/3.5 models—across four core tasks: next-word prediction, semantic plausibility comparison, isolated sentence acceptability judgments, and comparative sentence evaluations. They tested models using established linguistic benchmarks and recent news data, comparing direct probability measurements against three prompt variations that framed the evaluation as questions or instructions.

Across all experiments, the article found that metalinguistic prompt responses systematically diverge from the quantities directly derived from internal model representations. Direct probability measurements consistently matched or exceeded prompting methods in task accuracy. Furthermore, internal consistency degraded significantly as tasks moved further from direct next-word prediction: correlation coefficients dropped from approximately 0.78–0.79 in simple word prediction down to 0.20–0.24 for complex sentence-level syntactic judgments. The analysis also showed that presenting minimal pairs side by side substantially improved prompting performance compared to asking for isolated sentence evaluations.

These findings demonstrate that negative results from prompt-based evaluations do not provide conclusive evidence that a model lacks a specific linguistic generalization. The failure often stems from an inability to retrieve and verbalize internal representations in response to a prompt, reflecting a gap between underlying model competence and prompt-dependent performance. The trend among commercial providers toward closed application programming interfaces that block access to token probabilities poses a major risk to accurate model evaluation and scientific interpretability.

Decision-makers and evaluation teams should prioritize open models and direct probability measurements over black-box prompting when assessing core linguistic capabilities. When prompt-based methods are unavoidable, practitioners should present comparative pairs rather than isolated items to improve diagnostic reliability. While these findings provide high confidence regarding zero-shot evaluation limitations, further research is needed to determine whether multi-shot examples or reasoning techniques can improve the alignment between internal representations and prompted outputs.

Cover for Prompting is not a substitute for probability measurements in large language models

Abstract

Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models’ probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgement. In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models’ linguistic knowledge. Broadly, we find that LLMs’ metalinguistic judgments are inferior to quantities directly derived from representations. Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities. Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization. Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited.

Table of Contents

  • 1 Introduction
  • 3 General methods
  • 3.1 Overview of tasks
  • 3.2 Overview of prompts
  • 2 Related work
  • 3.3 Models
  • 4 Details of experiments
  • 4.1 Experiment 1: Word prediction
  • 4.2 Experiment 2: Word comparison
  • 4.3 Experiment 3a: Sentence judgment
  • 4.4 Experiment 3b: Sentence comparison
  • 5 Results
  • 5.1 Task performance
  • 5.2 Internal consistency
  • 6 Discussion
  • 6.1 Competence vs. performance in LLMs
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Details of syntactic phenomena
  • B Dataset-level task performance
  • C Relationship between metalinguistic and direct predictions
  • D Experiments in Mandarin Chinese

Knowls

  1. Knowl 1 — Direct probability measurement versus metalinguistic prompting

    definition

    The paper distinguishes two ways to evaluate linguistic knowledge in a large language model (LLM). Direct probability measurement reads probabilities from the model’s output distribution for the linguistic object itself—for example, comparing P(“is”∣“The keys to the cabinet”)P(\text{“is”}\mid\text{“The keys to the cabinet”}) with P(“are”∣“The keys to the cabinet”)P(\text{“are”}\mid\text{“The keys to the cabinet”}). Metalinguistic prompting places the same linguistic material inside a natural-language question or instruction and measures the probability of the model’s answer, such as the token corresponding to “Yes,” “No,” “1,” or “2.”

    Metalinguistic prompting therefore tests both whether the relevant linguistic generalization is represented and whether the model can access, interpret, and report it under the prompt. The two methods would be equivalent only if the prompt-conditioned answer distribution preserved the relevant ordering in the original linguistic probability distribution; the experiments show that this equivalence is not guaranteed.

  2. Knowl 2 — Prompting methods ordered by distance from direct prediction

    model/method

    The evaluation compares a direct method with three zero-shot metalinguistic prompt formats. MetaQuestionSimple asks a question whose linguistic object is placed close to the position where the model must make its prediction. MetaInstruct gives an imperative such as “Tell me what word is most likely to come next,” placing the answer within an instruction-following context. MetaQuestionComplex embeds the linguistic object in a more elaborate question and requests an explicit answer after the object, making it farthest from a direct next-word probability measurement.

    The study treats these prompt formats as an approximate ordering of increasing distance from direct measurement: Direct, MetaQuestionSimple, MetaInstruct, and MetaQuestionComplex. In word and sentence comparison tasks, metalinguistic answer options were presented in two counterbalanced orders, and results were averaged across the two orders.

  3. Knowl 3 — Models and extraction of direct probabilities

    model/method

    The study evaluates six LLMs: Flan-T5-small with 80 million parameters, Flan-T5-large with 780 million parameters, Flan-T5-XL with 3 billion parameters, text-curie-001, text-davinci-002, and text-davinci-003. The Flan-T5 models are encoder-decoder models pretrained with span corruption and subsequently instruction-finetuned. The OpenAI models differ in training regime: text-curie-001 is an autoregressive GPT-3 model, text-davinci-002 has supervised instruction fine-tuning, and text-davinci-003 additionally uses reinforcement learning from human feedback.

    For an autoregressive model, direct next-word probabilities are read from the vocabulary distribution following the sentence prefix. For Flan-T5, a prefix ⟨w1,…,wn−1⟩\langle w_1,\ldots,w_{n-1}\rangle is followed by the sentinel token <extra_id_0>, and the output sequence <extra_id_0> w_n is scored; probabilities for the tokens composing the target word wnw_n are summed. For a full sentence s=(w1,…,wm)s=(w_1,\ldots,w_m), Flan-T5 is scored with a pseudo-likelihood obtained by masking each word in turn and summing the log probabilities assigned to the true words:

    PLL⁡θ(s)=∑i=1mlog⁡Pθ ⁣(wi∣s with wi replaced by a sentinel mask),\operatorname{PLL}_\theta(s)=\sum_{i=1}^{m}\log P_\theta\!\left(w_i\mid s\text{ with }w_i\text{ replaced by a sentinel mask}\right),

    where θ\theta denotes the model parameters, wiw_i is the iith word, and PθP_\theta is the model’s token probability distribution.

  4. Knowl 4 — Four-task experimental suite

    experimental setup

    The study evaluates word-level and sentence-level knowledge in four English tasks. Word prediction uses 384 fact-like declarative sentences from Pereira et al. (2018) and 222 newly collected news items. Each news item consists of a United States news headline and the first sentence of an article published from March 20–26, 2023, with the final word withheld. Word comparison uses 395 minimal sentence pairs from Vassallo et al. (2018), where the two final words produce semantically plausible versus implausible events, such as “The archer released the arrow/interview.”

    Sentence judgment uses 345 minimal pairs derived from SyntaxGym and 390 items sampled from the 13 BLiMP categories. Each item contains a grammatical and an ungrammatical sentence differing in a targeted syntactic feature. In this task, each sentence is evaluated separately under metalinguistic prompting. Sentence comparison uses the same syntactic items but presents both sentences together and asks which is the better English sentence. Thus, the suite varies both the linguistic level—word versus sentence—and the comparison structure—isolated judgment versus minimal-pair comparison.

  5. Knowl 5 — Evaluation and internal-consistency measures

    model/method

    For word prediction, task performance is the average log probability assigned to the ground-truth final word. For word comparison, accuracy is the proportion of items for which the model assigns greater probability to the semantically plausible continuation than to the implausible continuation; chance performance is 50%. For sentence judgment, direct accuracy is the proportion of minimal pairs for which the grammatical sentence receives greater probability, while metalinguistic performance is balanced accuracy over “Yes” responses to grammatical sentences and “No” responses to ungrammatical sentences. For sentence comparison, accuracy is the proportion of pairs for which the grammatical sentence, or its corresponding answer option, receives greater probability; chance is 50%.

    Internal consistency is measured by Pearson correlation between direct and metalinguistic item-level responses. For word prediction, the responses are the log probabilities of the ground-truth continuation. For word comparison, they are the plausible-minus-implausible log-probability differentials. For sentence judgment, they are grammatical-minus-ungrammatical sentence differentials for the direct method and “Yes”-probability differentials between grammatical and ungrammatical prompts for the metalinguistic methods. For sentence comparison, they are grammatical-minus-ungrammatical differentials versus the corresponding answer-option differentials.

  6. Knowl 6 — Direct measurements generally outperform metalinguistic judgments

    empirical result

    Across the four English tasks and six evaluated models, direct probability measurements generally achieve the best or a comparable level of task performance relative to all three metalinguistic prompting methods. The methods nevertheless produce different performance scores, showing that metalinguistic judgments are not simply readouts of the underlying direct probabilities.

    The main exceptions are limited: Flan-T5-small performs best with MetaInstruct in word prediction, Flan-T5-XL performs relatively well with MetaQuestionComplex in word prediction, and the text-davinci models perform relatively well with metalinguistic prompts in word comparison. The overall result supports using direct probabilities as the primary measurement when the goal is to assess information encoded in a model’s linguistic representations. A failure under a metalinguistic prompt is not, by itself, conclusive evidence that the model lacks the corresponding linguistic generalization.

  7. Knowl 7 — Minimal-pair comparisons improve metalinguistic syntax evaluation

    empirical result

    Presenting grammatical and ungrammatical sentences as a minimal pair improves metalinguistic performance compared with asking for an isolated acceptability judgment. In isolated sentence judgment, each sentence is separately placed in a prompt asking whether it is a good English sentence. In sentence comparison, both members of the minimal pair are presented together and the model chooses which is better.

    Direct performance is identical in the two tasks by construction, because both direct conditions compare the same full-sentence probabilities. In contrast, metalinguistic accuracy increases from isolated judgments to minimal-pair comparisons for every model that performs above chance. This indicates that explicitly contrasting two controlled alternatives helps models reveal syntactic generalizations that are less reliably expressed in isolated metalinguistic judgments.

  8. Knowl 8 — Alignment declines as tasks and prompts depart from direct prediction

    data/table

    The correlation between direct and metalinguistic responses decreases as either the linguistic task or the prompt becomes more distant from direct next-word prediction. Averaged over models and datasets, the Pearson correlations for MetaQuestionSimple, MetaInstruct, and MetaQuestionComplex were:

    • Word prediction: 0.790.79, 0.780.78, and 0.630.63.
    • Word comparison: 0.500.50, 0.500.50, and 0.380.38.
    • Sentence judgment: 0.200.20, 0.200.20, and 0.200.20.
    • Sentence comparison: 0.230.23, 0.240.24, and 0.220.22.

    The strongest alignment occurs for word prediction, especially with prompts structurally close to the direct continuation prediction. Alignment is weaker for word comparison and is near zero-to-moderate for the two sentence-level syntax tasks. The result is not merely a difference in average accuracy: even item-level response magnitudes diverge, so the metalinguistic methods do not consistently preserve the relative strengths of the linguistic preferences encoded by direct probabilities.

  9. Knowl 9 — Preliminary evidence beyond English

    empirical result

    A preliminary evaluation with GPT-3.5 text-davinci-003 tested Mandarin Chinese using recent-news word prediction and controlled minimal pairs covering semantic and syntactic phenomena. As in English, direct probability measurements and metalinguistic prompts produced different results, direct measurement achieved the highest ground-truth continuation probability in word prediction, and minimal-pair presentation improved performance relative to isolated sentence judgments.

    The Chinese results also reveal an important qualification: in sentence comparison, direct accuracy was 0.60.6, whereas each metalinguistic condition achieved 0.80.8. The authors suggest that the model may be better suited to Chinese when additional Chinese text is supplied in the metalinguistic prompt than when it must assign probabilities to an isolated Chinese sentence. Thus, the broad mismatch between evaluation methods extends beyond English, but the relative advantage of direct measurement is not universal across languages and tasks.

  10. Knowl 10 — Scope limitations of the evidence

    limitation

    The evidence is limited to three zero-shot metalinguistic prompt formats and six models. The study does not test few-shot prompting or in-context examples, which could improve metalinguistic task performance or alignment with direct probabilities. It also does not evaluate chat-based models, largely because the tested API access did not expose the token probabilities required for direct measurement.

    The OpenAI models were accessed through a closed API, so their results may not be fully reproducible as the service changes. The authors therefore do not claim that prompting should never be used: prompting remains useful for open-ended responses and questions that are difficult to translate into direct probability measurements. Their narrower conclusion is that researchers should state the assumptions behind an evaluation method and should retain access to model-internal probabilities whenever scientific measurement of linguistic knowledge is the goal.

Coverage note — Detailed per-phenomenon syntactic item lists, model-by-model scatterplots, and ancillary ethics and reference material were omitted because they do not add distinct load-bearing findings beyond the summarized experimental design and results.

References

  1. 1.Marco Baroni. 2022. On the proper role of linguistically-oriented deep net analysis in linguistic theorizing. In Shalom Lappin and Jean-Philippe Bernardy, editors, Algebraic Structures in Natural Language. Taylor & Francis.
  2. 2.Gašper Beguš, Maksymilian D ˛abkowski, and Ryan Rhodes. 2023. Large Linguistic Models: Analyzing theoretical linguistic abilities of LLMs.
  3. 3.Anne Beyer, Sharid Loáiciga, and David Schlangen. 2021. Is Incoherence Surprising? Targeted Evaluation of Coherence Prediction from Language Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4164–4173, Online. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Kristy Choi, Chris Cundy, Sanjari Srivastava, and Stefano Ermon. 2022. LMPriors: Pre-Trained Language Models as Task-Specific Priors.
  6. 6.Noam Chomsky. 1965. Aspects of the Theory of Syntax. MIT Press.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling Instruction-Finetuned Language Models.
  8. 8.Pablo Contreras Kallens, Ross Deans Kristensen-McLachlan, and Morten H. Christiansen. 2023. Large Language Models Demonstrate the Potential of Statistical Learning in Language. Cognitive Science, 47(3):e13256. Publisher: John Wiley & Sons, Ltd.
  9. 9.Vittoria Dentella, Elliot Murphy, Gary Marcus, and Evelina Leivada. 2023. Testing AI performance on less frequent aspects of language reveals insensitivity to underlying meaning.
  10. 10.Gabe Dupre. 2021. (What) Can Deep Learning Contribute to Theoretical Linguistics? Minds and Machines, 31(4):617–635.
  11. 11.Chaz Firestone. 2020. Performance vs. competence in human–machine comparisons. Proceedings of the National Academy of Sciences, 117(43):26562–26571. Publisher: Proceedings of the National Academy of Sciences.
  12. 12.Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019. Neural language models as psycholinguistic subjects: Representations of syntactic state. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 32–42, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020. SyntaxGym: An Online Platform for Targeted Evaluation of Language Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 70–76, Online. Association for Computational Linguistics.
  14. 14.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70, pages 1321–1330.
  15. 15.Michael Hahn, Richard Futrell, Roger P. Levy, and Edward Gibson. 2022. A resource-rational model of human processing of recursive linguistic structure. 119(43).
  16. 16.Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020. A Systematic Assessment of Syntactic Generalization in Neural Language Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1725–1744, Online. Association for Computational Linguistics.
  17. 17.Janus. 2022. Mysteries of mode collapse.
  18. 18.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language Models (Mostly) Know What They Know.
  19. 19.Roni Katzir. 2023. Why large language models are poor theories of human linguistic cognition: A reply to Piantadosi (2023).
  20. 20.Carina Kauf, Anna A. Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan S. She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2022. Event knowledge in large language models: the gap between the impossible and the unlikely.
  21. 21.Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, and Yejin Choi. 2022. Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3631–3643, Seattle, United States. Association for Computational Linguistics.
  22. 22.Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large Language Models are Zero-Shot Reasoners. In Advances in Neural Information Processing Systems, volume 35, pages 22199–22213. Curran Associates, Inc.
  23. 23.Andrew Kyle Lampinen. 2023. Can language models handle recursively nested grammatical structures? A case study on comparing models and humans.
  24. 24.Nur Lan, Emmanuel Chemla, and Roni Katzir. 2022. Large Language Models and the Argument From the Poverty of the Stimulus.
  25. 25.Belinda Z. Li, William Chen, Pratyusha Sharma, and Jacob Andreas. 2023. LaMPP: Language Models as Probabilistic Priors for Perception and Action.
  26. 26.Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies. Transactions of the Association for Computational Linguistics, 4:521–535.
  27. 27.Benjamin Lipkin, Lionel Wong, Gabriel Grand, and Joshua B. Tenenbaum. 2023. Evaluating statistical language models as pragmatic reasoners. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 45.
  28. 28.R. Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L. Griffiths. 2023. Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve.
  29. 29.Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing Conversational Agents’ Overconfidence Through Linguistic Calibration. Transactions of the Association for Computational Linguistics, 10:857–872.
  30. 30.Daniel Milway. 2023. A Response to Piantadosi (2023).
  31. 31.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  32. 32.Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic. 2021. Revisiting the Calibration of Modern Neural Networks. In Advances in Neural Information Processing Systems, volume 34, pages 15682–15694. Curran Associates, Inc.
  33. 33.Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. 2023. The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain.
  34. 34.Elliot Murphy. 2023. Notes on Large Language Models and Linguistic Theory.
  35. 35.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021. Show Your Work: Scratchpads for Intermediate Computation with Language Models.
  36. 36.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  37. 37.Roma Patel and Ellie Pavlick. 2022. Mapping Language Models to Grounded Conceptual Spaces. In International Conference on Learning Representations.
  38. 38.Francisco Pereira, Bin Lou, Brianna Pritchett, Samuel Ritter, Samuel J. Gershman, Nancy Kanwisher, Matthew Botvinick, and Evelina Fedorenko. 2018. Toward a universal decoder of linguistic meaning from brain activation. Nature Communications, 9(1):963.
  39. 39.Steven T. Piantadosi. 2023. Modern language models refute Chomsky’s approach to language.
  40. 40.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2023. GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 3845–3864, Dubrovnik, Croatia. Association for Computational Linguistics.
  41. 41.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked Language Model Scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
  42. 42.Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.
  43. 43.Tomer Ullman. 2023. Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.
  44. 44.Paolo Vassallo, Emmanuele Chersoni, Enrico Santus, Alessandro Lenci, and Philippe Blache. 2018. Event Knowledge in Sentence Processing: A New Dataset for the Evaluation of Argument Typicality. In LREC 2018 Workshop on Linguistic and Neurocognitive Resources (LiNCR), Miyazaki, Japan.
  45. 45.Yiwen Wang, Jennifer Hu, Roger Levy, and Peng Qian. 2021. Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5604–5620, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Alex Warstadt and Samuel R. Bowman. 2022. What artificial neural networks can tell us about human language acquisition. In Shalom Lappin and Jean-Philippe Bernardy, editors, Algebraic Structures in Natural Language. Taylor & Francis.
  47. 47.Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics, 8. Publisher: MIT Press.
  48. 48.Taylor Webb, Keith J. Holyoak, and Hongjing Lu. 2023. Emergent Analogical Reasoning in Large Language Models.
  49. 49.Albert Webson, Alyssa Marie Loo, Qinan Yu, and Ellie Pavlick. 2023. Are Language Models Worse than Humans at Following Prompts? It’s Complicated.
  50. 50.Albert Webson and Ellie Pavlick. 2022. Do Prompt-Based Models Really Understand the Meaning of Their Prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  51. 51.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent Abilities of Large Language Models. Transactions on Machine Learning Research.
  52. 52.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022b. Chain of Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems.
  53. 53.Ethan Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger P. Levy. 2022. Learning syntactic structures from string input. In Shalom Lappin and Jean-Philippe Bernardy, editors, Algebraic Structures in Natural Language. Taylor & Francis.
  54. 54.Victor H. Yngve. 1960. A Model and an Hypothesis for Language Structure. Proceedings of the American Philosophical Society, 104(5):444–466. Publisher: American Philosophical Society.

Citation

MLA
Hu, J., and R. Levy. “Prompting Is Not a Substitute for Probability Measurements in Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 5040–60, https://doi.org/10.18653/v1/2023.emnlp-main.306.
APA
Hu, J., & Levy, R. (2023). Prompting is not a substitute for probability measurements in large language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5040–5060. https://doi.org/10.18653/v1/2023.emnlp-main.306
Chicago
Hu, J., and R. Levy. 2023. “Prompting Is Not a Substitute for Probability Measurements in Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5040–60. https://doi.org/10.18653/v1/2023.emnlp-main.306.
Harvard
Hu, J. and Levy, R. (2023) “Prompting is not a substitute for probability measurements in large language models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5040–5060. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.306.
Vancouver
1. Hu J, Levy R (2023) Prompting is not a substitute for probability measurements in large language models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5040–5060

BibTeX

@inproceedings{hu-levy-2023-prompting,
    title = "Prompting is not a substitute for probability measurements in large language models",
    author = "Hu, Jennifer  and
      Levy, Roger",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.306/",
    doi = "10.18653/v1/2023.emnlp-main.306",
    pages = "5040--5060"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/