Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Kaitlyn ZhouDan JurafskyTatsunori Hashimoto

article2023EMNLP136 citations

Reveals that prompt expressions of high confidence paradoxically degrade language model accuracy by up to seven percent compared to expressions of uncertainty, exposing how pretraining data patterns cause models to mimic surface linguistic habits rather than genuinely track factual truth.

Listen

As large language models are increasingly deployed in real-world environments requiring factual accuracy, understanding how they process human expressions of certainty, uncertainty, and evidential attribution is critical. Humans naturally use linguistic cues to frame the reliability of knowledge, yet standard model evaluations often overlook these nuanced expressions.

The main objective of the article is to evaluate how linguistic markers of uncertainty and certainty injected into input prompts affect the accuracy and probability distributions of language model outputs in question-answering tasks.

To conduct this evaluation, the researchers established a linguistic taxonomy categorizing 50 distinct epistemic markers into weakeners, strengtheners, factive verbs, and evidential citations. They applied zero-shot and few-shot prompting across four question-answering datasets—TriviaQA, Natural Questions, CountryQA, and Jeopardy—and tested multiple OpenAI GPT models ranging from early base models to GPT-4. They also performed quantitative and qualitative analyses on pretraining data from The Pile to investigate the linguistic origins of the observed model behaviors.

The article establishes several key findings. First, language models are extraordinarily sensitive to epistemic phrasing in prompts, causing accuracy to fluctuate by up to 80% on identical questions. Second, expressions of high certainty paradoxically degrade model performance: weakeners (hedges) outperformed strengtheners (boosters) across all datasets, yielding an average accuracy of 47% compared to 40% for strengtheners. This accuracy loss was largely driven by factive verbs that presuppose truth. Third, evidential markers that cite sources significantly improved accuracy, frequently outperforming standard unprompted baselines. Fourth, injecting numerical confidence revealed poor linguistic calibration; accuracy peaked between 70% and 90% but dropped sharply when prompts claimed 100% certainty. Pretraining data analysis revealed that human writers predominantly use certainty phrases in questions to frame problems or admit ignorance, while using uncertainty phrases in answers to remain polite or cautious.

These findings indicate that language models do not possess true epistemological awareness; instead, they mimic conversational patterns observed during pretraining. For decision-makers and system developers, this introduces notable operational risks. Users who write confident prompts expecting better factual retrieval may unintentionally cause model hallucinations or errors. Furthermore, while grounding prompts with evidential markers like "Wikipedia says" boosts accuracy, blindly generating simulated attributions creates safety and compliance hazards by presenting fabricated sources persuasively.

Organizations deploying language models should avoid assuming that confident inputs or outputs correlate with factual accuracy. Developers should design prompts that avoid extreme certainty assertions and factive verbs, favor evidential grounding only when sources are verified, and position questions to elicit direct answers before any framing expressions. Further work is required to teach models how to navigate idiomatic versus literal confidence expressions and to verify external attributions automatically before integration into critical workflows.

Confidence in these findings is supported by consistent replication across six distinct model architectures and multiple datasets. However, certain limitations remain. The empirical evaluation relied exclusively on English-language trivia and discrete question-answering benchmarks, excluding long-form dialogue, continuous numerical targets, and multilingual cultural variations in hedging.

  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its experiments on whether language models can recognize what they know establish the model-calibration context that this paper extends by testing how epistemic wording changes answers.
Cover for Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Abstract

The increased deployment of LMs for real-world tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs. Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures. We develop a typology of epistemic markers and inject 50 markers into prompts for question answering. We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%. Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance. Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions. These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.

Table of Contents

  • 1 Introduction
  • 2 Expressions of Certainty and Uncertainty: Linguistic Background
  • 2.1 Expressions of Uncertainty Typology
  • 6 The Impact of the Degree of Uncertainty on Performance
  • 100% Certainty is not 100% Accurate

Knowls

  1. Knowl 1 — Typology of Epistemic Markers for Language Model Prompting

    definition

    A linguistic typology categorizes naturalistic expressions of epistemic uncertainty and certainty injected into question-answering prompts across six orthogonal or overlapping dimensions:

    1. Weakeners vs. Strengtheners: Weakeners reduce speaker commitment or soften propositional claims (e.g., "I think it's...", "It could be..."), whereas strengtheners assert or presuppose high certainty and strong speaker commitment (e.g., "I'm certain it's...", "Undoubtedly it's...").
    2. Plausibility Shields: A subclass of weakeners that explicitly convey the speaker's lower subjective commitment (e.g., "I suppose it's...", "I feel like it should be...").
    3. Factive Verbs: Verbs whose semantic structure presupposes the truth of their complement clauses (e.g., "realize", "know", "understand", "Wikipedia confirms"), in contrast to non-factive verbs (e.g., "believe", "suggests").
    4. Evidential Markers: Linguistic devices that specify the source or mode of information (e.g., "According to...", "I heard it's...", "Online says it's...").
    5. Source Mentions: Expressions that explicitly cite an identifiable source (e.g., "Wikipedia says", "The textbook says") versus unspecified or indirect attributions (e.g., "They told me", "Rumor says").
    6. First-Person Pronouns (1P): Markers that incorporate first-person personal pronouns ("I", "we", "our") to index subjectivity or personal stance.
  2. Knowl 2 — Experimental Protocol for Epistemic Prompt Injection in Trivia Question Answering

    experimental setup

    Open-ended question answering is evaluated by injecting 50 epistemic marker templates into question-answer prompts across four discrete QA datasets:

    • Datasets: TriviaQA (n=200n=200 subset), Natural Questions closed-book (n=200n=200 subset), Jeopardy questions crawled from J! Archive (n=200n=200 subset), and CountryQA (n=53n=53 country-capital queries). All questions are filtered to ensure discrete, single-token or in-vocabulary answers to prevent out-of-vocabulary and multi-token paraphrase confounders.
    • Prompt Formatting: Prompts follow the structure Q: <question>\nA: <epistemic template> <answer> without trailing whitespace.
    • Evaluated Models: OpenAI GPT models including GPT-3 Davinci, Ada, Babbage, Curie, Instruct (text-davinci-003), and GPT-4 (32k context). Sampling temperature is set to 1.01.0 for log-probability access, and 0.00.0 for deterministic GPT-4 evaluation.
    • Metrics:
      • Top-1 Accuracy: Generation is executed for 10 tokens; an output is scored correct if any generated token sequence matches the gold target or its valid aliases.
      • Probability-on-Gold: The sum of probability mass assigned by the model to all accepted aliases of the correct answer.
      • Alternative Token Entropy: The Shannon entropy of the probability distribution over candidate predictions excluding the argmax prediction.
  3. Knowl 3 — Accuracy Degradation from Injected Expressions of High Certainty and Factive Verbs

    empirical result

    Language models exhibit substantial sensitivity to epistemic cues in zero-shot prompts: expressions of high certainty (strengtheners and factive verbs) consistently cause performance drops compared to expressions of uncertainty (weakeners and hedges).

    • Across TriviaQA, Natural Questions, Jeopardy, and CountryQA on GPT-3 Davinci, weakeners achieve an average accuracy of 47%, compared to 40% for strengtheners. In CountryQA, this gap reaches 17 percentage points (with templates using factive verbs like "We realize it's..." falling to 14% accuracy, compared to near 100% for various weakener templates).
    • Factive verbs (which presuppose the truth of their complement clause) systematically reduce accuracy across all evaluation datasets and model scales.
    • Template perplexity does not explain this degradation: the correlation between template perplexity and QA accuracy is negligible (Pearson's ρ=−0.03\rho = -0.03), demonstrating that accuracy variations are driven by the epistemic semantics rather than token frequency or surface fluency.
  4. Knowl 4 — Accuracy Improvements from Evidential Markers and External Source Grounding

    empirical result

    Prompting language models with evidential markers that attribute information to external sources (e.g., "Online says it's...", "Wikipedia says it's...", "Wikipedia claims it's...") significantly improves question answering accuracy compared to non-evidential prompts.

    • On GPT-3 Davinci, evidential markers provide statistically significant accuracy gains over non-evidentials in three out of four evaluation datasets (TriviaQA, CountryQA, and Natural Questions).
    • Certain evidential and uncertainty templates outperform the standard unadorned baseline prompt (Q: <question> A:). On TriviaQA, "Online says it's..." achieves 66.0% accuracy compared to 62.5% for the standard prompt. Across all GPT models on Natural Questions, multiple evidential templates exceed baseline performance (e.g., "The internet says it's..." reaches 41.6% average top-1 accuracy vs. 27.5% for standard prompting).
  5. Knowl 5 — TriviaQA Accuracy Comparison Across Linguistic Epistemic Categories for Six GPT Models

    data/table

    The table compares top-1 question-answering accuracy on TriviaQA (n=200n=200) across 50 epistemic marker templates grouped into paired linguistic categories, evaluated on six OpenAI language models:

    Epistemic Category ada babbage curie davinci instruct gpt-4
    Boosters 0.091 0.257 0.313 0.392 0.589 0.793
    Hedges 0.079 0.272 0.333*** 0.468*** 0.642*** 0.822***
    Factive Verbs 0.078 0.237 0.293 0.347 0.555 0.771
    Non-Factive Verbs 0.085* 0.276*** 0.336*** 0.468*** 0.641*** 0.821***
    Evidentials 0.087** 0.281*** 0.347*** 0.449* 0.640*** 0.820***
    Non-evidentials 0.080 0.250 0.301 0.433 0.601 0.799

    Instruct denotes text-davinci-003; gpt-4 uses a 32K context window. Asterisks denote statistical significance from two-sample tt-tests against the paired category: * p<0.05p < 0.05, ** p<0.01p < 0.01, *** p<0.001p < 0.001. Across model parameter scales from Ada to GPT-4, hedges outperform boosters, non-factive verbs outperform factive verbs, and evidential constructions outperform non-evidentials.

  6. Knowl 6 — Probability Mass Redistribution and Tail Entropy Flattening Under Weakeners

    empirical result

    The accuracy advantage of weakener prompts over strengthener prompts is not driven by increased probability on the correct answer. The average probability-on-gold for correct answers is slightly lower under weakeners than strengtheners across three datasets on GPT-3 Davinci: NaturalQA (42% vs 45%), Jeopardy (47% vs 51%), and TriviaQA (53% vs 55%).

    Instead, weakeners induce a significant increase in the entropy of the probability distribution over alternative candidate tokens (excluding the top predicted token):

    • TriviaQA: 2.980±0.012.980 \pm 0.01 (weakeners) vs. 2.917±0.012.917 \pm 0.01 (strengtheners)
    • CountryQA: 3.078±0.023.078 \pm 0.02 (weakeners) vs. 2.875±0.032.875 \pm 0.03 (strengtheners)
    • Jeopardy: 3.170±0.013.170 \pm 0.01 (weakeners) vs. 3.089±0.013.089 \pm 0.01 (strengtheners)
    • NaturalQA: 3.167±0.013.167 \pm 0.01 (weakeners) vs. 3.106±0.013.106 \pm 0.01 (strengtheners)

    This demonstrates that weakeners flatten the tail probability distribution across competing plausible candidates rather than over-concentrating probability on a single potentially incorrect choice.

  7. Knowl 7 — Pragmatic Asymmetry of Certainty and Uncertainty in Pretraining Corpora

    empirical result

    Analysis of the Stack Exchange subset of The Pile pretraining corpus reveals an asymmetry between questions and answers regarding certainty markers:

    • Certainty in Questions: High-certainty expressions ("I'm sure", "It must be", "I know") occur at more than double the rate in question posts (279.6 instances per million words; 1,907,691 total occurrences) compared to answer posts (103.7 instances per million words; 526,374 total occurrences). Qualitative coding indicates that question authors use certainty markers primarily to establish group membership, acknowledge ignorance ("I'm sure it's a simple error"), or isolate hypotheses.
    • Uncertainty in Answers: Uncertainty expressions ("I think", "It could be", "Maybe it's", "It should be") occur at roughly twice the rate in answer posts (436.2 instances per million words; 2,214,539 total occurrences) compared to question posts (222.3 instances per million words; 1,516,776 total occurrences), serving pragmatic functions such as politeness and hedging corrections.

    Because autoregressive language models capture co-occurrence distributions from training data, certainty markers in prompts prime the linguistic patterns of confused question-askers rather than confident, accurate answers.

  8. Knowl 8 — Non-Monotonic Accuracy and 100% Certainty Collapse Under Numerical Epistemic Prompts

    empirical result

    When verbal epistemic templates are modified with explicit numerical confidence percentages (0%,10%,30%,50%,70%,90%,100%0\%, 10\%, 30\%, 50\%, 70\%, 90\%, 100\%), language model generation exhibits severe miscalibration and non-monotonic scaling:

    • Linguistic Miscalibration: Expected Calibration Error (ECE) between prompt-stated probabilities and answer accuracy is poor, ranging from 0.30 to 0.50 (e.g., "I'm 90% sure it's..." on TriviaQA produces correct answers only 57% of the time).
    • Extreme Certainty Drop: Performance peaks between 70% and 90% confidence but drops sharply when 100% certainty is stated. Across four datasets and seven templates, 21 out of 28 templates with 100% numerical values yielded lower probability-on-gold than their 90% counterparts. Accuracy also drops severely at 0% certainty.
    • Hyperbolic Non-Literal Usage in Pretraining Data: In The Pile, 44% of occurrences of "100%" co-occur with negation (e.g., "never 100% accurate"), and "100%" is used to express uncertainty in questions at nearly twice the rate of answers (14 vs 8). Literal expressions of complete certainty (e.g., "I'm 100% sure that") account for only 3% of sampled occurrences.
  9. Knowl 9 — In-Context Learning Calibration of Emitted Epistemic Markers and Template Placement Effects

    empirical result

    Few-shot in-context learning (50 balanced examples) designed to teach GPT-3 to emit verbal epistemic markers when its confidence exceeds or falls below a threshold (p=0.5p = 0.5) reveals poor verbal calibration and strong positional sensitivity:

    • Calibration Deficit: GPT-3 achieves a macro-F1 of only 0.56 for emitting weakeners and 0.53 for emitting strengtheners (close to the 0.45 random baseline). Emitting a weakener correctly correlates with lower accuracy (74% when emitted vs. 83% when withheld), but emitting a strengthener fails to reflect higher truthfulness (79.9% accuracy when emitted vs. 78.9% when withheld).
    • Positional Sensitivity (Prefix vs. Suffix): Appending epistemic markers before the answer as a prefix (e.g., "I think it's Paris") degrades top-1 accuracy to 63% and probability-on-gold to 40%. Appending epistemic markers after the answer as a suffix (e.g., "Paris. I think") preserves 80% accuracy and 67% probability-on-gold.

Coverage note — Omitted minor crowdsourcing compensation details, full rankings of all 50 individual templates, and general ethics statements, keeping all primary methodologies, linguistic definitions, quantitative results, mechanistic analyses, and tabular data.

References

  1. 1.Alexandra Y Aikhenvald. 2004. Evidentiality. OUP Oxford.
  2. 2.Penelope Brown and Stephen C Levinson. 1987. Politeness: Some universals in language usage, volume 4. Cambridge university press.
  3. 3.Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The role of explanations on trust and reliance in clinical decision support systems. In 2015 International Conference on Healthcare Informatics, pages 160–169.
  4. 4.Annabelle Carrell, Neil Mallinar, James Lucas, and Preetum Nakkiran. 2022. The calibration generalization gap. International Conference on Machine Learning (Workshop on Distribution-Free Uncertainty Quantification).
  5. 5.Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012. Did it happen? the pragmatic complexity of veridicality assessment. Computational Linguistics, 38(2):301–333.
  6. 6.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124.
  7. 7.Judith Degen and Judith Tonhauser. 2022. Are there factive predicates? an empirical investigation. Language, 98(3):552–591.
  8. 8.Judith Degen, Andreas Trotzke, Gregory Scontras, Eva Wittenberg, and Noah D Goodman. 2019. Definitely, maybe: A new experimental paradigm for investigating the pragmatics of evidential devices across languages. Journal of Pragmatics, 140:33–48.
  9. 9.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  10. 10.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  11. 11.Adam Gleave and Geoffrey Irving. 2022. Uncertainty estimation for language reward models. arXiv preprint arXiv:2203.07472.
  12. 12.Hila Gonen, Srini Iyer, Terra Blevins, Noah A Smith, and Luke Zettlemoyer. 2022. Demystifying prompts in language models via perplexity estimation. arXiv preprint arXiv:2212.04037.
  13. 13.Ken Hyland. 1998. Boosting, hedging and the negotiation of academic knowledge. Text & Talk, 18(3):349–382.
  14. 14.Ken Hyland. 2005. Stance and engagement: A model of interaction in academic discourse. Discourse studies, 7(2):173–192.
  15. 15.Ken Hyland. 2014. Disciplinary discourses: Writer stance in research articles. In Writing: Texts, processes and practices, pages 99–121. Routledge.
  16. 16.Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. 2021. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry, 11(1):108.
  17. 17.Abhyuday Jagannatha and Hong Yu. 2020. Calibrating structured output predictors for natural language processing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2078–2092, Online. Association for Computational Linguistics.
  18. 18.Nanjiang Jiang and Marie-Catherine de Marneffe. 2021. He thinks he knows better than the doctors: BERT for event factuality fails on pragmatics. Transactions of the Association for Computational Linguistics, 9:1081–1097.
  19. 19.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  20. 20.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  21. 21.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  22. 22.Amita Kamath, Robin Jia, and Percy Liang. 2020. Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5684–5696, Online. Association for Computational Linguistics.
  23. 23.Lauri Karttunen. 1971. Some observations on factivity. Research on Language & Social Interaction, 4(1):55–69.
  24. 24.Paul Kiparsky and Carol Kiparsky. 1970. FACT, pages 143–173. De Gruyter Mouton, Berlin, Boston.
  25. 25.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in- and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1326–1340, Online. Association for Computational Linguistics.
  26. 26.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  27. 27.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019a. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  28. 28.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019b. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics.
  29. 29.George Lakoff. 1975. Hedges: A study in meaning criteria and the logic of fuzzy concepts. In Contemporary Research in Philosophical Logic and Linguistic Semantics: Proceedings of a Conference Held at the University of Western Ontario, London, Canada, pages 221–271. Springer.
  30. 30.Young-Jun Lee, Chae-Gyun Lim, Yunsu Choi, Ji-Hui Lm, and Ho-Jin Choi. 2022. PERSONACHATGEN: Generating personalized dialogues using GPT-3. In Proceedings of the 1st Workshop on Customized Chat Grounding Persona and Knowledge, pages 29–48, Gyeongju, Republic of Korea. Association for Computational Linguistics.
  31. 31.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R’e, Diana Acosta-Navas, Drew A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan S. Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas F. Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023. Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525:140 – 146.
  32. 32.Stephanie C. Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. Trans. Mach. Learn. Res., 2022.
  33. 33.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872.
  35. 35.Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic. 2021. Revisiting the calibration of modern neural networks. Advances in Neural Information Processing Systems, 34:15682–15694.
  36. 36.Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology, pages 1–18.
  37. 37.Roma Patel and Ellie Pavlick. 2021. “was it “stated” or was it “claimed”?: How linguistic bias affects generative language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10080–10095, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  38. 38.E. Prince, C. Bosk, and J. Frader. 1982. On hedging in physician-physician discourse. Di Pietro, R.J., Ed., Linguistics and the Professions, pages 83–97.
  39. 39.Anna Prokofieva and Julia Hirschberg. 2014. Hedging and speaker commitment. In 5th Intl. Workshop on Emotion, Social Signals, Sentiment & Linked Open Data, Reykjavik, Iceland.
  40. 40.Reid Pryzant, Richard Diehl Martinez, Nathan Dass, Sadao Kurohashi, Dan Jurafsky, and Diyi Yang. 2020. Automatically neutralizing subjective bias in text. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 480–489.
  41. 41.Valentina Pyatkin, Shoval Sadde, Aynat Rubinstein, Paul Portner, and Reut Tsarfaty. 2021. The possible, the plausible, and the desirable: Event-based modality detection for language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 953–965.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Yann Raphalen, Chloé Clavel, and Justine Cassell. 2022. “You might think about slightly revising the title”: Identifying hedges in peer-tutoring interactions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2160–2174, Dublin, Ireland. Association for Computational Linguistics.
  44. 44.Rachel Rudinger, Aaron Steven White, and Benjamin Van Durme. 2018. Neural models of factuality. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 731–744, New Orleans, Louisiana. Association for Computational Linguistics.
  45. 45.Roser Saurí and James Pustejovsky. 2009. Factbank: a corpus annotated with event factuality. Language resources and evaluation, 43:227–268.
  46. 46.Mandy Simons, Judith Tonhauser, David Beaver, and Craige Roberts. 2010. What projects and why. In Semantics and linguistic theory, volume 20, pages 309–327.
  47. 47.Gabriel Stanovsky, Judith Eckle-Kohler, Yevgeniy Puzikov, Ido Dagan, and Iryna Gurevych. 2017. Integrating deep linguistic features in factuality prediction over unified datasets. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 352–357, Vancouver, Canada. Association for Computational Linguistics.
  48. 48.Elias Stengel-Eskin and Benjamin Van Durme. 2023. Calibrated Interpretation: Confidence Estimation in Semantic Parsing. Transactions of the Association for Computational Linguistics, 11:1213–1231.
  49. 49.Meiqi Sun, Wilson Yan, Pieter Abbeel, and Igor Mordatch. 2022. Quantifying uncertainty in foundation models via ensembles. In NeurIPS 2022 Workshop on Robustness in Sequence Modeling.
  50. 50.Mirac Suzgun, Luke Melas-Kyriazi, and Dan Jurafsky. 2022. Prompt-and-rerank: A method for zero-shot and few-shot arbitrary textual style transfer with small language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2195–2222, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  51. 51.Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020. On the inference calibration of neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3070–3079, Online. Association for Computational Linguistics.
  52. 52.Orion Weller, Marc Marone, Nathaniel Weir, Dawn J. Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2023. "according to ..." prompting language models improves quoting from pre-training data. CoRR, abs/2305.13252.

Citation

MLA
Zhou, K., et al. “Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 5506–24, https://doi.org/10.18653/v1/2023.emnlp-main.335.
APA
Zhou, K., Jurafsky, D., & Hashimoto, T. B. (2023). Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5506–5524. https://doi.org/10.18653/v1/2023.emnlp-main.335
Chicago
Zhou, K., D. Jurafsky, and T. B. Hashimoto. 2023. “Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5506–24. https://doi.org/10.18653/v1/2023.emnlp-main.335.
Harvard
Zhou, K., Jurafsky, D. and Hashimoto, T.B. (2023) “Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5506–5524. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.335.
Vancouver
1. Zhou K, Jurafsky D, Hashimoto TB (2023) Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5506–5524

BibTeX

@inproceedings{zhou-etal-2023-navigating,
    title = "Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models",
    author = "Zhou, Kaitlyn  and
      Jurafsky, Dan  and
      Hashimoto, Tatsunori",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.335/",
    doi = "10.18653/v1/2023.emnlp-main.335",
    pages = "5506--5524"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/