Uncertainty Quantification for In-Context Learning of Large Language Models

Chen LingXujiang ZhaoXuchao ZhangWei ChengYanchi LiuYiyou SunMika OishiTakao OsakiKatsushi MatsudaJie Ji

article2024NAACL62 citations

Presents a Bayesian framework that decomposes predictive uncertainty in large language model in-context learning into prompt-induced aleatoric and model-induced epistemic components, enabling unsupervised diagnostic evaluation of output reliability across both white-box and black-box settings.

Listen

Large Language Models often produce unreliable outputs or hallucinations, posing significant operational and decision-making risks in practical applications. While prompting models with a few example demonstrations (in-context learning) improves performance, existing uncertainty metrics fail to distinguish whether model confusion stems from poor demonstration examples or limitations in the model itself.

The article aims to formulate and evaluate an unsupervised framework that quantifies and separates predictive uncertainty in in-context learning into two distinct sources: data-driven variation from the demonstrations and parameter-level variation from the model configurations.

The researchers evaluated this approach using natural language understanding benchmarks across sentiment analysis, topic classification, and linguistic acceptability. They tested open-source models, primarily LLaMA-2 (7B, 13B, and 70B parameter versions) and OPT-13B, across thousands of test cases. For accessible white-box models, the framework aggregates prediction entropy using beam search decoding across varying demonstration sets, while also defining a variance-based alternative for black-box systems.

The evaluation revealed several critical findings. First, decomposing uncertainty into data-level and model-level components identified misclassified samples more accurately than standard baseline metrics, such as semantic and raw token entropy. Second, ensuring demonstrations cover every label class consistently reduced error rates and uncertainty compared to purely random selection. Third, scaling up model capacity from 7B to 70B parameters significantly reduced overall uncertainty and improved error-detection accuracy. Finally, model-level uncertainty proved to be the most reliable indicator for detecting out-of-distribution or irrelevant prompt demonstrations, achieving area-under-the-curve scores exceeding 0.93.

These findings provide immediate practical value for managing generative AI risks. Organizations deploying language models can use these decomposed metrics as automated guardrails to reject low-confidence outputs, diagnose faulty prompt designs, and identify when queries fall outside the system's operational domain. Rather than treating uncertainty as a single uninterpretable number, teams can determine whether to fix prompt examples or upgrade the underlying model.

Stakeholders should implement balanced class-sampling strategies when designing few-shot prompts and integrate model-level uncertainty monitoring into production pipelines to catch out-of-domain inputs. Before deploying this approach on open-ended generation tasks like summarization or drafting, teams should run additional pilot tests, as the current framework relies on structured classification and closed-form outputs where specific answer tokens can be isolated.

Confidence in these findings is strong for deterministic natural language classification tasks across standard model architectures. However, decision-makers should remain cautious regarding open-ended text generation, where linguistic redundancy makes identifying core predictive tokens more challenging and requires further research.

arXiv: 2402.10189lingchen0331/UQ_ICL

No sufficiently relevant recommendations were found.

Cover for Uncertainty Quantification for In-Context Learning of Large Language Models

Abstract

In-context learning has emerged as a ground-breaking ability of Large Language Models (LLMs) and revolutionized various fields by providing a few task-relevant demonstrations in the prompt. However, trustworthy issues with LLM’s response, such as hallucination, have also been actively discussed. Existing works have been devoted to quantifying the uncertainty in LLM’s response, but they often overlook the complex nature of LLMs and the uniqueness of in-context learning. In this work, we delve into the predictive uncertainty of LLMs associated with in-context learning, highlighting that such uncertainties may stem from both the provided demonstrations (aleatoric uncertainty) and ambiguities tied to the model’s configurations (epistemic uncertainty). We propose a novel formulation and corresponding estimation method to quantify both types of uncertainties. The proposed method offers an unsupervised way to understand the prediction of in-context learning in a plug-and-play fashion. Extensive experiments are conducted to demonstrate the effectiveness of the decomposition. The code and data are available at: https://github.com/lingchen0331/UQ_ICL.

Table of Contents

  • 1 Introduction
  • 2 Uncertainty Decomposition of In-context Learning
  • 2.1 Background
  • 2.2 Predictive Uncertainty Formulation of In-context Learning
  • 2.3 Entropy-based Decomposition
  • 2.4 Entropy Approximation
  • 3 Related Works
  • 4 Experiments
  • 4.1 Experiment Setup
  • 4.2 Quantitative Analysis
  • 4.3 Generalization Capability
  • 4.4 Misclassification Rate with Out of Domain Demonstration
  • 4.5 Out-of-domain Demonstration Detection
  • 4.6 Semantic Out-of-distribution Detection
  • 5 Conclusion
  • Limitations
  • References
  • A Appendix
  • A.1 Variance-based Decomposition
  • A.2 Dataset Description
  • A.3 Experiment Setup
  • A.4 Prompt Template
  • A.5 Case Study

Knowls

  1. Knowl 1 — Latent-concept Bayesian formulation of in-context prediction

    model/method

    The paper models in-context learning as prediction under two sources of variation: a latent task concept inferred from demonstrations and model parameters or configurations. The prompt x1:Tx_{1:T} consists of demonstrations x1,…,xT−1x_1,\ldots,x_{T-1} and a test input xTx_T; each demonstration is assumed to be drawn independently from data associated with the same latent concept zz. The test output is yTy_T, and Θ\Theta denotes model parameters or decoding configurations. The proposed predictive distribution is

    p(yT∣x1:T)≈∫ ⁣∫p(yT∣Θ,x1:T,z) p(z∣x1:T) q(Θ) dz dΘ.p(y_T\mid x_{1:T})\approx \int\!\int p(y_T\mid \Theta,x_{1:T},z)\,p(z\mid x_{1:T})\,q(\Theta)\,dz\,d\Theta.

    Here q(Θ)q(\Theta) is an approximate distribution over model parameters or configurations, and p(z∣x1:T)p(z\mid x_{1:T}) represents uncertainty over the concept associated with the prompt. The paper approximates the conditional output likelihood using a Gaussian centered on a deterministic generation function, with covariance representing variation associated with model parameters. Treating in-context learning as Bayesian inference over latent concepts is presented as a probabilistic hypothesis about how demonstrations guide a model, not as an established account of the mechanism.

  2. Knowl 2 — Entropy-based epistemic and aleatoric decomposition

    equation

    For a fixed prompt x1:Tx_{1:T} and model setting Θ\Theta, the paper defines total predictive uncertainty using the entropy H(yT∣x1:T,Θ)H(y_T\mid x_{1:T},\Theta). Averaging the conditional output entropy over latent concepts (which are approximated through different demonstration sets) gives the paper's epistemic uncertainty (EU); the difference between total entropy and that average is its aleatoric uncertainty (AU):

    EU=Ez ⁣[H(yT∣x1:T,z,Θ)],AU=H(yT∣x1:T,Θ)−Ez ⁣[H(yT∣x1:T,z,Θ)].\mathrm{EU}=\mathbb{E}_{z}\!\left[H(y_T\mid x_{1:T},z,\Theta)\right],\qquad \mathrm{AU}=H(y_T\mid x_{1:T},\Theta)-\mathbb{E}_{z}\!\left[H(y_T\mid x_{1:T},z,\Theta)\right].

    Equivalently, the paper identifies AU with the mutual information between the output and latent concept conditional on the model setting, I(yT;z∣Θ)I(y_T;z\mid\Theta). In this decomposition, the first term is the paper's measure of EU and the entropy difference is its measure of AU; the labels follow the paper's formulation.

  3. Knowl 3 — Answer-focused entropy estimator for free-form model outputs

    model/method

    To estimate the entropy components for a task with KK possible answer labels, the method discards generated text that does not express the answer and uses the probabilities of answer-relevant tokens. For each of LL demonstration sets, it generates multiple candidate responses, using beam search as a practical approximation to sampling model outputs, and aggregates the answer-token probabilities into a KK-dimensional column. Let Ak,jA_{k,j} be the accumulated probability for label kk under demonstration set jj, so AA is a K×LK\times L matrix. Define σ(v)k=vk/∑rvr\sigma(v)_k=v_k/\sum_r v_r to normalize a nonnegative vector, and let H(p)=−∑k=1Kpklog⁡pkH(p)=-\sum_{k=1}^{K}p_k\log p_k be discrete entropy. The estimates are

    EU^=1L∑j=1LH ⁣(σ(A:,j)),AU^=H ⁣(σ ⁣(∑j=1LA:,j))−EU^.\widehat{\mathrm{EU}}=\frac{1}{L}\sum_{j=1}^{L}H\!\left(\sigma(A_{:,j})\right),\qquad \widehat{\mathrm{AU}}=H\!\left(\sigma\!\left(\sum_{j=1}^{L}A_{:,j}\right)\right)-\widehat{\mathrm{EU}}.

    Thus, EU is the average entropy of the answer distribution for individual demonstration sets, while AU is the entropy of the pooled answer distribution minus that average. The procedure is designed for structured-answer tasks such as classification: the model is instructed to return only the desired label, or answer tokens are extracted from longer responses. This entropy estimator requires access to generated-token probabilities, so it is intended for white-box models.

  4. Knowl 4 — Variance decomposition for black-box language models

    equation

    When token probabilities are unavailable, the paper proposes a variance-based alternative using the law of total variance. For output random variable yTy_T, prompt x1:Tx_{1:T}, and model-setting distribution q(Θ)q(\Theta), it writes

    Var⁡(yT∣x1:T)=Var⁡q(Θ) ⁣(E[yT∣x1:T,Θ])+Eq(Θ) ⁣[Var⁡(yT∣x1:T,Θ)].\operatorname{Var}(y_T\mid x_{1:T})= \operatorname{Var}_{q(\Theta)}\!\left(\mathbb{E}[y_T\mid x_{1:T},\Theta]\right) +\mathbb{E}_{q(\Theta)}\!\left[\operatorname{Var}(y_T\mid x_{1:T},\Theta)\right].

    The variance across conditional mean outputs as Θ\Theta varies is designated epistemic uncertainty; the expected conditional output variance is designated aleatoric uncertainty. In black-box use, outputs are collected by querying the model with different demonstrations and configurations, such as temperature or top-pp settings, and the corresponding variances are estimated from those outputs. Unlike the entropy estimator, this decomposition does not require access to token probabilities.

  5. Knowl 5 — Evaluation datasets and inference protocol

    experimental setup

    The main experiments use LLaMA-2 chat models with 7B, 13B, and 70B parameters; OPT-13B is used for a cross-backbone comparison. Test sets comprise EMOTION (2,000 examples, six classes), Financial Phrasebank (850, three classes), SST-2 (872, two classes), CoLA (1,040, two classes), and AG_News (1,160, four classes). Demonstrations are sampled either randomly or with at least one example from each class. The numbers of demonstrations for random/class sampling are: EMOTION, 6/1 per class; Financial Phrasebank, 6/2 per class; SST-2, 4/2 per class; CoLA, 2/1 per class; and AG_News, 4/1 per class. Demonstrations are sampled four times; decoding uses beam search with width 10 and a maximum of 16 new tokens.

    For misclassification assessment, each test prediction is labeled correct or incorrect, and uncertainty scores are evaluated using AUPR and AUROC, with larger scores intended to identify errors. Comparisons include likelihood-based uncertainty, token entropy, semantic uncertainty, and the proposed EU and AU scores.

  6. Knowl 6 — Uncertainty scores identify errors across tasks, with dataset-dependent results

    empirical result

    Across the five natural-language-understanding datasets and LLaMA-2 model sizes, the proposed EU and AU measures generally provide stronger AUPR and AUROC for distinguishing misclassified from correct predictions than likelihood, token-entropy, and semantic-uncertainty baselines, though the advantage is not uniform across every dataset, model, and metric. For example, on Financial Phrasebank with LLaMA-2-70B and class-balanced demonstrations, EU reaches AUPR 0.893 and AUROC 0.804, compared with semantic uncertainty at 0.774 and 0.649, respectively. Class-balanced demonstration sampling generally performs better than random sampling, and larger models tend to improve uncertainty-assessment performance, but the paper reports exceptions to both trends. A separate EMOTION comparison reports similar qualitative precision-recall and ROC behavior for OPT-13B and LLaMA-2-13B, supporting use across backbones.

  7. Knowl 7 — Demonstration shifts reduce AUROC more for AU than for EU

    empirical result

    On EMOTION with LLaMA-2-13B, the paper evaluates how well EU and AU rank misclassified examples when demonstrations come from the task's original training data, a related sentiment dataset (Financial Phrasebank), or an unrelated acceptability dataset (CoLA). The reported AUROCs for random/class demonstration sampling are:

    Random Class
    Demonstration source EU AU EU AU
    Original 0.681 0.585 0.686 0.599
    Relevant 0.688 (+1.0%) 0.541 (-7.5%) 0.671 (-2.2%) 0.524 (-12.5%)
    OOD 0.671 (-1.4%) 0.501 (-13.3%) 0.673 (-1.8%) 0.497 (-17.0%)

    Relative changes are reported against the corresponding original-demonstration score. EU changes comparatively little, whereas AUROC for AU declines substantially with shifted demonstrations, by as much as 17.0% for class-balanced sampling with CoLA demonstrations.

  8. Knowl 8 — EU detects out-of-domain demonstrations in the tested sentiment task

    empirical result

    For LLaMA-2-13B on EMOTION, the paper tests whether uncertainty scores distinguish in-domain demonstrations from relevant demonstrations drawn from Financial Phrasebank or out-of-domain (OOD) demonstrations drawn from CoLA. In-domain demonstrations are labeled 0 and shifted demonstrations 1. The reported AUPR/AUROC values are:

    Demonstration comparison Score AUPR AUROC
    Relevant vs. in-domain Semantic 0.702 0.644
    Relevant vs. in-domain EU 0.742 0.935
    Relevant vs. in-domain AU 0.657 0.682
    OOD vs. in-domain Semantic 0.698 0.712
    OOD vs. in-domain EU 0.784 0.941
    OOD vs. in-domain AU 0.773 0.607

    EU has the highest AUPR and AUROC in both comparisons. The result indicates that, in these experiments, EU is a stronger indicator of demonstration-domain shift than AU or semantic uncertainty.

  9. Knowl 9 — EU identifies semantic out-of-distribution test examples

    empirical result

    The semantic out-of-distribution (SOOD) experiment uses EMOTION, masks the sadness and anger classes from the task description, and asks the model to classify test inputs among the remaining four classes. Samples from the masked classes are labeled SOOD; other samples are in-distribution. For LLaMA-2-7B and LLaMA-2-13B, the reported AUPR/AUROC values are:

    Model Semantic EU AU
    AUPR AUROC AUPR AUROC AUPR AUROC
    7B 0.477 0.532 0.548 0.658 0.461 0.570
    13B 0.417 0.468 0.525 0.592 0.414 0.437

    EU gives the strongest AUPR and AUROC for both model sizes in this setup. The authors attribute increased uncertainty for SOOD cases to the mismatch between the examples and demonstrations and to the prompt excluding the correct class.

  10. Knowl 10 — Scope and access limitations of the uncertainty methods

    limitation

    The paper's evaluation and proposed answer-focused entropy estimation target natural-language-understanding tasks with identifiable answer tokens, including classification and multiple-choice questions. For open-ended generation, the method may not reliably determine which parts of a generated sequence carry semantic importance, limiting its applicability. Entropy-based estimation also needs token probabilities and therefore applies to white-box models; the variance-based alternative is offered for black-box models but does not remove the difficulty of identifying meaningful answers in unconstrained generation.

Coverage note — The detailed prompt template and the single-query case study are omitted because they illustrate the evaluated procedure rather than contribute a separate method or generalizable result.

References

  1. 1.Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U Rajendra Acharya, et al. 2021. A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information fusion, 76:243–297.
  2. 2.Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang. 2023. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models. arXiv preprint arXiv:2305.13712.
  3. 3.Guangji Bai, Zheng Chai, Chen Ling, Shiyu Wang, Jiaying Lu, Nan Zhang, Tingwei Shi, Ziyang Yu, Mengdan Zhu, Yifei Zhang, et al. 2024. Beyond efficiency: A systematic survey of resource-efficient large language models. arXiv preprint arXiv:2401.00625.
  4. 4.Pei Chen, Haibo Ding, Jun Araki, and Ruihong Huang. 2021. Explicitly capturing relations between entity mentions via graph neural networks for domain-specific named entity recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 735–742.
  5. 5.Pei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan, Sheng Zha, Ruihong Huang, and George Karypis. 2024. Hytrel: Hypergraph-enhanced tabular data representation learning. Advances in Neural Information Processing Systems, 36.
  6. 6.Pei Chen, Haotian Xu, Cheng Zhang, and Ruihong Huang. 2022. Crossroads, buildings and neighborhoods: A dataset for fine-grained location recognition. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3329–3339.
  7. 7.Kamaljit Chowdhary and Paul Dupuis. 2013. Distinguishing and integrating aleatoric and epistemic variation in uncertainty quantification. ESAIM: Mathematical Modelling and Numerical Analysis-Modélisation Mathématique et Analyse Numérique, 47(3):635–662.
  8. 8.Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft. 2017. Uncertainty decomposition in bayesian neural networks with latent variables. arXiv preprint arXiv:1706.08495.
  9. 9.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. arXiv preprint arXiv:2003.07892.
  10. 10.Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, et al. 2023. Lmpolygraph: Uncertainty estimation for language models. arXiv preprint arXiv:2311.07383.
  11. 11.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  12. 12.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  13. 13.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664.
  14. 14.Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models. arXiv preprint arXiv:2305.19187.
  15. 15.Zi Lin, Jeremiah Zhe Liu, and Jingbo Shang. 2022. Towards collaborative neural-symbolic graph semantic parsing via uncertainty. Findings of the Association for Computational Linguistics: ACL 2022.
  16. 16.Chen Ling, Junji Jiang, Junxiang Wang, and Zhao Liang. 2022. Source localization of graph diffusion via variational autoencoders for graph inverse problems. In Proceedings of the 28th ACM SIGKDD, pages 1010–1020.
  17. 17.Chen Ling, Xuchao Zhang, Xujiang Zhao, Yanchi Liu, Wei Cheng, Mika Oishi, Takao Osaki, Katsushi Matsuda, Haifeng Chen, and Liang Zhao. 2023a. Open-ended commonsense reasoning with unrestricted answer candidates. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8035–8047.
  18. 18.Chen Ling, Xujiang Zhao, Jiaying Lu, Chengyuan Deng, Can Zheng, Junxiang Wang, Tanmoy Chowdhury, Yun Li, Hejie Cui, Xuchao Zhang, Tianjiao Zhao, et al. 2023b. Domain specialization as the key to make large language models disruptive: A comprehensive survey. arXiv preprint arXiv:2305.18703.
  19. 19.Chen Ling, Xujiang Zhao, Xuchao Zhang, Yanchi Liu, Wei Cheng, Haoyu Wang, Zhengzhang Chen, Takao Osaki, Katsushi Matsuda, Haifeng Chen, et al. 2023c. Improving open information extraction with large language models: A study on demonstration uncertainty. arXiv preprint arXiv:2309.03433.
  20. 20.Andrey Malinin and Mark Gales. 2020. Uncertainty estimation in autoregressive structured prediction. arXiv preprint arXiv:2002.07650.
  21. 21.P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. 2014. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65.
  22. 22.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  23. 23.Vipula Rawte, Amit Sheth, and Amitava Das. 2023. A survey of hallucination in large foundation models. arXiv preprint arXiv:2309.05922.
  24. 24.Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018. CARER: Contextualized affect representations for emotion recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3687–3697.
  25. 25.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  26. 26.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  27. 27.Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  28. 28.Lisa Wimmer, Yusuf Sale, Paul Hofman, Bernd Bischl, and Eyke Hüllermeier. 2023. Quantifying aleatoric and epistemic uncertainty in machine learning: Are conditional entropy and mutual information appropriate measures? In Uncertainty in Artificial Intelligence, pages 2282–2292.
  29. 29.Yijun Xiao and William Yang Wang. 2019. Quantifying uncertainties in natural language processing tasks. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 7322–7329.
  30. 30.Yijun Xiao and William Yang Wang. 2021. On hallucination and predictive uncertainty in conditional language generation. arXiv preprint arXiv:2103.15025.
  31. 31.Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022. Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. arXiv preprint arXiv:2210.04714.
  32. 32.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080.
  33. 33.Jialin Yu, Alexandra I Cristea, Anoushka Harit, Zhongtian Sun, Olanrewaju Tahir Aduragba, Lei Shi, and Noura Al Moubayed. 2022. Efficient uncertainty quantification for multilabel text classification. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8.
  34. 34.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pretrained transformer language models.
  35. 35.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NIPS.
  36. 36.Yifei Zhang, Bo Pan, Chen Ling, Yuntong Hu, and Liang Zhao. 2024. Elad: Explanation-guided large language models active distillation. arXiv preprint arXiv:2402.13098.
  37. 37.Xujiang Zhao, Feng Chen, Shu Hu, and Jin-Hee Cho. 2020. Uncertainty aware semi-supervised learning on graph data. Advances in Neural Information Processing Systems, 33:12827–12836.
  38. 38.Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023. Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439.

Citation

MLA
Ling, C., et al. “Uncertainty Quantification for In-Context Learning of Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 3357–70, https://doi.org/10.18653/v1/2024.naacl-long.184.
APA
Ling, C., Zhao, X., Zhang, X., Cheng, W., Liu, Y., Sun, Y., Oishi, M., Osaki, T., Matsuda, K., Ji, J., Bai, G., (赵亮), L. Z., & Chen, H. (2024). Uncertainty Quantification for In-Context Learning of Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3357–3370. https://doi.org/10.18653/v1/2024.naacl-long.184
Chicago
Ling, C., X. Zhao, X. Zhang, et al. 2024. “Uncertainty Quantification for In-Context Learning of Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3357–70. https://doi.org/10.18653/v1/2024.naacl-long.184.
Harvard
Ling, C. et al. (2024) “Uncertainty Quantification for In-Context Learning of Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3357–3370. Available at: https://doi.org/10.18653/v1/2024.naacl-long.184.
Vancouver
1. Ling C, Zhao X, Zhang X, et al (2024) Uncertainty Quantification for In-Context Learning of Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3357–3370

BibTeX

@inproceedings{ling-etal-2024-uncertainty,
    title = "Uncertainty Quantification for In-Context Learning of Large Language Models",
    author = "Ling, Chen  and
      Zhao, Xujiang  and
      Zhang, Xuchao  and
      Cheng, Wei  and
      Liu, Yanchi  and
      Sun, Yiyou  and
      Oishi, Mika  and
      Osaki, Takao  and
      Matsuda, Katsushi  and
      Ji, Jie  and
      Bai, Guangji  and
      Zhao, Liang  and
      Chen, Haifeng",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.184/",
    doi = "10.18653/v1/2024.naacl-long.184",
    pages = "3357--3370"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/