Self-Prompting Large Language Models for Zero-Shot Open-Domain QA

Junlong LiJinyuan WangZhuosheng ZhangHai Zhao

article2024NAACL72 citations

Proposes a zero-shot open-domain question-answering framework that prompts large language models to generate synthetic passages, QA pairs, and explanations from scratch, using clustering-based retrieval to assemble in-context demonstrations that match supervised retrieval-augmented models.

Listen

Open-domain question answering requires answering broad factual questions without being provided specific reference documents. Traditional systems rely on extensive labeled datasets and external text corpora to train retrieval and answering pipelines, which creates substantial costs and operational overhead. While large language models can answer general questions directly from their internal knowledge, standard zero-shot prompting techniques underutilize their capabilities and typically lag behind customized, fine-tuned models.

The article evaluates a framework called Self-Prompting, which aims to improve zero-shot open-domain question answering without using any human-annotated training data or external document databases. The approach demonstrates how a large language model can autonomously generate its own high-quality reference examples and leverage them to guide its final answers.

The evaluated framework operates in two main phases. In the offline preparation phase, the language model generates around 5,000 synthetic question-answer pairs across 29 topics by creating short factual passages, extracting key entities, generating matching questions, verifying answer accuracy, and drafting brief explanations. In the inference phase, the system uses a clustering-based retrieval algorithm to dynamically select a diverse yet semantically relevant set of 10 synthetic demonstrations to prepend to each incoming test question. The article evaluated this methodology across three standard benchmark datasets: Web Questions, Natural Questions, and TriviaQA.

The key findings highlight substantial performance gains. First, Self-Prompting outperformed direct prompting baselines by an average of 15.5 Exact Match points and surpassed the previous state-of-the-art zero-shot method by 8.8 points across the three benchmarks. Second, the system achieved performance comparable to strong fine-tuned and retrieval-augmented models, as well as few-shot baselines that rely on real training data. Third, structuring demonstrations to provide the answer first followed by a concise explanation delivered superior accuracy compared to standard reasoning chains or multi-step generation approaches. Fourth, the framework proved consistently effective across multiple model architectures and parameter sizes, including smaller open-source models.

These results demonstrate that large language models contain sufficient internal world knowledge to serve as their own reference source for open-domain questions. For organizations, this approach eliminates the need to build, index, and maintain multi-gigabyte external corpora, while removing the recurring costs of human annotation. Additionally, generating brief explanatory statements alongside answers enhances system transparency and trustworthiness without introducing the high latency of multi-step reasoning chains.

Decision-makers should consider adopting offline synthetic demonstration pools combined with clustered retrieval when deploying language models for factual question answering where labeled data is unavailable. A single offline generation pipeline—costing approximately $120 and taking six hours in the evaluated setup—offers a more cost-effective and accurate operational strategy than real-time contextual generation or maintaining large external indexes.

Confidence in these findings is supported by consistent gains across multiple benchmarks and models. However, leaders should note several limitations: initial prompt engineering requires iterative adjustment, commercial application programming interface costs during the data generation phase can be significant, and smaller language models exhibit higher rates of factual inaccuracies in synthetic data generation compared to top-tier commercial models.

arXiv: 2212.08635lockon-n/self-prompting
Cover for Self-Prompting Large Language Models for Zero-Shot Open-Domain QA

Abstract

Large Language Models (LLMs) have demonstrated impressive proficiency in zero-shot open-domain question answering (ODQA), yet their performance remains limited by the knowledge-intensive nature of the task. In this paper, we propose Self-Prompting, a fully automated pipeline that leverages LLMs' inherent ability to generate data. Specifically, we prompt LLMs to imagine diverse user prompts (i.e., questions) via in-context sampling, and then use them to elicit knowledge from LLMs in order to generate synthetic passages. We demonstrate that our pipeline improves zero-shot ODQA performance by up to 9.6 F1 points across multiple base models and evaluation datasets, achieving comparable performance to supervised retrieval-augmented methods. Furthermore, we introduce self-variational prompting to mitigate hallucinations in knowledge-intensive tasks by generating multiple reasoning chains and selecting the most consistent one. Our results highlight the promise of harnessing LLMs' generative capabilities to address knowledge-intensive tasks in the absence of labeled data.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Approach
  • 3.1 Pseudo QA Dataset Generation
  • 3.2 Dynamic In-context Demonstrations Selection for Inference
  • 4 Experiments
  • 4.1 Datasets and Settings
  • 4.2 Main Results
  • 5 Analysis
  • 5.1 How to Use Generated Passages and Explanations
  • 5.2 Ways of Demonstration Selection
  • 5.3 Different Number of Demonstrations
  • 5.4 Impact of the size of generated pseudo data for Self-Prompting
  • 5.5 General Effectiveness with LLMs of Different Sizes
  • 5.6 Data Generation Quality Analysis
  • 5.7 Comparison between Self-Prompting and Using Training Data
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics
  • References
  • A Statistics of the datasets
  • B Names and Required Number of Examples for Topics
  • C Details in Generating Pseudo QA Dataset
  • D Statistics of the Generated Pseudo QA datasets
  • E Specific Templates in Different Formats
  • F Training Small Models with the Pseudo QA dataset
  • G Real-time Pseudo Data Generation Variant for Self-Prompting
  • H Cost and Running Efficiency
  • I More Data Quality Analysis

Knowls

  1. Knowl 1 — Self-Prompting turns an LLM into a generator and user of its own QA demonstrations

    model/method

    Self-Prompting is a zero-shot open-domain question answering method that uses no task-specific training data or external knowledge corpus. It has two stages: first, an LLM creates a pool of pseudo question–answer pairs, each with a background passage and a short explanation; second, examples are selected from that pool for each test question and supplied as in-context demonstrations to the LLM. The model then answers the test question using the demonstrations. The method is designed to elicit the LLM’s knowledge and language-understanding capabilities by having it construct the material used to prompt itself.

  2. Knowl 2 — Pseudo-QA construction combines topic coverage, answer checks, and explanations

    algorithm

    The preparation stage builds a pseudo open-domain QA dataset from scratch using an LLM. The authors define 29 topics under broad themes informed by the entity distribution of TriviaQA answers, then prompt the LLM to generate examples for those topics and write short Wikipedia-style passages about them. For each passage, the LLM extracts named entities as candidate answers and generates a question for each candidate. To filter question–answer pairs, the LLM is asked to answer each generated question from the passage; the pair is retained only if this answer matches the candidate entity after normalization. The LLM then writes a one-sentence explanation grounded in the passage and required to include the answer. The generation process allows at most 10 pairs per passage, bans selected pronouns to reduce context-dependent ambiguity, and filters out answers longer than five words. This process produced 977 passages and 4,883 question–answer–explanation triples.

  3. Knowl 3 — Clustering-based retrieval selects demonstrations that are both relevant and varied

    model/method

    For inference, each pseudo QA pair is encoded with Sentence-BERT (all-mpnet-base-v2), and the encoded pool is partitioned by kk-means into kk clusters, where kk is the desired number of demonstrations. The test question is encoded with the same model. From each cluster, the method selects the pseudo QA pair with the greatest cosine similarity to the test-question embedding. The selected examples therefore combine question-level semantic relevance with coverage across clusters. In the main experiments, the method used 10 demonstrations and presented each as question, answer, then explanation.

  4. Knowl 4 — Answer-before-explanation is the most effective tested demonstration format

    empirical result

    On randomly selected subsets of 1,000 test examples from each of WebQ, NQ, and TriviaQA, the authors compared formats using 10 demonstrations and report average Exact Match (EM) and F1 across the three datasets. Q, A, E, and P denote question, answer, explanation, and passage. The answer-then-explanation format (QAE) performed best: 48.2 EM and 58.3 F1. Formats that placed the explanation before the answer (QEA) or passage before the answer (QPA) performed worse, as did formats that added passages to answer-before-explanation demonstrations. Two-call formats that first generated a passage or explanation also underperformed the one-call formats.

    Demonstration format Prediction format EM F1
    QAE Q AE 48.2 58.3
    QAP Q AP 46.9 57.5
    QEA Q EA 43.3 53.6
    QPA Q PA 40.3 49.9
    QAEP Q AEP 46.9 57.3
    QAPE Q APE 46.4 57.1
    QP then PQA, two calls PQ A 41.5 51.8
    QE then EQA, two calls EQ A 43.7 54.7
  5. Knowl 5 — Self-Prompting substantially improves zero-shot benchmark performance

    empirical result

    On WebQ, NQ, and TriviaQA, the authors evaluate Self-Prompting with InstructGPT and Codex without training data or an external corpus. The reported EM scores show that InstructGPT Self-Prompting exceeds direct InstructGPT prompting by 15.5 points on average and GENREAD by 8.8 points on average. Its average EM of 46.2 is comparable to the reported retrieval-augmented fine-tuned systems. Codex Self-Prompting scores 84.3 EM on the reported TriviaQA subset, above the 83.5 score of RECITE (Codex) on that subset. A dagger marks results evaluated on a TriviaQA subset; the InstructGPT Self-Prompting entry reports both the full-set and subset scores.

    System WebQ EM NQ EM TriviaQA EM Average EM
    T5-SSM 40.8 34.8 51.0 42.2
    REALM 40.7 40.4 55.8 45.6
    DPR 41.1 41.5 56.8 46.5
    RAG 45.2 44.5 56.1 48.6
    Google+InstructGPT 19.9 27.8 58.7 35.5
    DPR+InstructGPT 20.1 29.9 55.3 35.1
    InstructGPT, direct 18.6 20.9 52.6 30.7
    GENREAD (InstructGPT) 24.8 28.2 59.3 37.4
    RECITE (Codex) – 35.8 83.5†^{\dagger} –
    Self-Prompting (InstructGPT) 35.6 36.2 66.8 / 79.4†^{\dagger} 46.2
    Self-Prompting (Codex) 38.9 40.7 84.3†^{\dagger} –

    The comparison includes fine-tuned systems, retrieval-augmented prompting, and direct prompting; the paper reports that Self-Prompting uses neither training data nor an external knowledge corpus. The 79.4, 83.5, and 84.3 TriviaQA scores are subset results.

  6. Knowl 6 — Retrieving within clusters beats global similarity and random selection

    empirical result

    With 10 demonstrations in QAE format, the authors compared four selection strategies on the three 1,000-example test subsets and report average EM and F1. Selecting the most similar example within each cluster achieved the best scores, supporting the combination of relevance and diversity. Random selection also had the largest reported run-to-run variability. In a separate NQ-subset comparison, generating a question-specific passage and QA set in real time scored below offline Self-Prompting, consistent with the authors’ account that narrowing examples to a highly relevant context can reduce useful diversity.

    Selection strategy EM F1
    Random 46.4±1.346.4\pm1.3 56.7±1.056.7\pm1.0
    Global cosine-similarity retrieval 46.7 57.1
    Closest example to each cluster center 47.1 57.2
    Most similar example within each cluster 48.2 58.3

    For the separate NQ-subset comparison, offline Self-Prompting scored 37.2 EM and 47.6 F1, while the real-time generation variant scored 35.3 EM and 45.5 F1.

  7. Knowl 7 — Performance largely saturates with 10 demonstrations and a smaller pseudo-data pool

    empirical result

    In analyses using Codex and the three 1,000-example test subsets, average EM generally improved as the number of demonstrations increased from 2 to 10; using more than 10 did not yield significant further gains. The authors therefore used 10 demonstrations in the main experiments, balancing performance and cost. A separate pool-size experiment found that near-maximum average EM could be obtained with substantially fewer than the approximately 5,000 generated QA pairs used in the full pool. These pool-size scores are average EM across the three analysis subsets.

    Number of generated QA pairs 100 500 1,000 2,000 Full pool
    Average EM 51.8 51.7 52.2 52.4 52.6
  8. Knowl 8 — Self-Prompting improves average scores across four LLMs of different sizes

    empirical result

    The authors compared direct answering with Self-Prompting on InstructGPT and Codex (175B parameters), GPT-NeoX (20B), and Alpaca (7B). The table reports average scores across WebQ, NQ, and TriviaQA: EM and F1 are answer metrics, while IE is the Instruct-Eval score judged by GPT-4. Self-Prompting raises all three reported average metrics for each model, including the smaller models, and narrows the gap between model sizes under direct prompting.

    Model Method Average EM Average F1 Average IE
    InstructGPT Direct 34.0 45.1 54.5
    InstructGPT Self-Prompting 48.2 58.3 62.6
    Codex Direct 40.8 52.1 59.3
    Codex Self-Prompting 52.6 62.3 65.9
    GPT-NeoX Direct 13.3 21.9 35.1
    GPT-NeoX Self-Prompting 28.4 37.0 38.3
    Alpaca Direct 19.0 31.9 42.4
    Alpaca Self-Prompting 30.1 41.0 46.8
  9. Knowl 9 — Generated demonstrations approach the performance of training-set demonstrations

    empirical result

    On non-overlapping subsets of WebQ (425 examples), NQ (357), and TriviaQA (254), the authors compared Self-Prompting with in-context learning from real training examples. To avoid training–test overlap, they used subsets identified as non-overlapping; both approaches used 10 demonstrations and the same clustering-based selection procedure. Because training examples lack explanations, this comparison used question–answer demonstrations rather than the QAE format. Averaged across the three datasets, training-data in-context learning scored 36.0 EM and 48.7 F1, while Self-Prompting scored 34.2 EM and 46.0 F1—a gap of 1.8 EM and 2.7 F1 despite the much larger available training-example pool.

  10. Knowl 10 — Generated data can contain factual errors, and API dependence constrains reproducibility

    limitation

    A manual quality assessment sampled 100 generated examples from each of four LLMs, matching topics across models, and used GPT-4 via New Bing to assess passage flaws, question ambiguity, answer correctness, and explanation quality. All models produced some outdated information; the smaller GPT-NeoX and Alpaca models made substantially more passage-level factual errors than InstructGPT and Codex. The authors report approximately one factual error per five InstructGPT-generated passages and one per ten Codex-generated passages, and found InstructGPT’s question–answer–explanation triples stronger on nearly all assessed measures. Thus, the generated pool is not fully reliable even though the authors’ filtering and checking steps improve its quality. In practical terms, dataset construction relied on paid OpenAI APIs, cost about $120, and took about six hours; identifying effective prompt templates also required trial and error, limiting reproducibility and leaving room to reduce manual effort.

Coverage note — The exploratory attempt to train a smaller RAG reader on the generated data is omitted because its fixed-retriever setup introduced confounds and was secondary to the paper’s in-context-learning contribution; detailed per-topic allocations and API prompt templates are implementation details rather than separate findings.

References

  1. 1.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  2. 2.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, Seattle, Washington, USA. Association for Computational Linguistics.
  3. 3.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879, Vancouver, Canada. Association for Computational Linguistics.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International Conference on Machine Learning, pages 3929–3938. PMLR.
  10. 10.Zhen Huang, Shiyi Xu, Minghao Hu, Xinyi Wang, Jinyan Qiu, Yongquan Fu, Yuncai Zhao, Yuxing Peng, and Changjian Wang. 2020. Recent trends in deep learning based open-domain textual question answering systems. IEEE Access, 8:94341–94356.
  11. 11.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  12. 12.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  13. 13.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  14. 14.Ehsan Kamalloo, Nouha Dziri, Charles Clarke, and Davood Rafiei. 2023. Evaluating open-domain question answering in the era of large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5591–5606, Toronto, Canada. Association for Computational Linguistics.
  15. 15.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  18. 18.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  19. 19.Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021a. Question and answer test-train overlap in open-domain question answering datasets. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1000–1008, Online. Association for Computational Linguistics.
  20. 20.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021b. PAQ: 65 million probably-asked questions and what you can do with them. Transactions of the Association for Computational Linguistics, 9:1098–1115.
  21. 21.Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2023. Rain: Your language models can align themselves without finetuning. arXiv preprint arXiv:2309.07124.
  22. 22.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021. Dexperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706.
  23. 23.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022a. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  24. 24.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2022b. Generated knowledge prompting for commonsense reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3154–3169, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022a. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS).
  26. 26.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022b. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  28. 28.OpenAI. 2023a. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  29. 29.OpenAI. 2023b. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  30. 30.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  31. 31.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  33. 33.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  34. 34.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  35. 35.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  36. 36.Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models. arXiv preprint arXiv:2210.01296.
  37. 37.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  38. 38.Ellen M Voorhees et al. 1999. The trec-8 question answering track report. In Trec, volume 99, pages 77–82.
  39. 39.Wenya Wang, Vivek Srikumar, Hanna Hajishirzi, and Noah A Smith. 2022. Elaboration-generating commonsense question answering at scale. arXiv preprint arXiv:2209.01232.
  40. 40.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  42. 42.Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022. Zerogen: Efficient zero-shot learning via dataset generation. arXiv preprint arXiv:2202.07922.
  43. 43.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2022. Generate rather than retrieve: Large language models are strong context generators. arXiv preprint arXiv:2209.10063.
  44. 44.Qin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, and Meng Fang. 2022a. A survey for efficient open domain question answering. arXiv preprint arXiv:2211.07886.
  45. 45.Qin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, and Meng Fang. 2023. A survey for efficient open domain question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14447–14465, Toronto, Canada. Association for Computational Linguistics.
  46. 46.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022b. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  47. 47.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022c. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  48. 48.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.
  49. 49.Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv preprint arXiv:2101.00774.

Citation

MLA
Li, J., et al. “Self-Prompting Large Language Models for Zero-Shot Open-Domain QA”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 296–310, https://doi.org/10.18653/v1/2024.naacl-long.17.
APA
Li, J., Wang, J., Zhang, Z., & Zhao, H. (2024). Self-Prompting Large Language Models for Zero-Shot Open-Domain QA. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 296–310. https://doi.org/10.18653/v1/2024.naacl-long.17
Chicago
Li, J., J. Wang, Z. Zhang, and H. Zhao. 2024. “Self-Prompting Large Language Models for Zero-Shot Open-Domain QA”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 296–310. https://doi.org/10.18653/v1/2024.naacl-long.17.
Harvard
Li, J. et al. (2024) “Self-Prompting Large Language Models for Zero-Shot Open-Domain QA”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 296–310. Available at: https://doi.org/10.18653/v1/2024.naacl-long.17.
Vancouver
1. Li J, Wang J, Zhang Z, Zhao H (2024) Self-Prompting Large Language Models for Zero-Shot Open-Domain QA. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 296–310

BibTeX

@inproceedings{li-etal-2024-self-prompting,
    title = "Self-Prompting Large Language Models for Zero-Shot Open-Domain {QA}",
    author = "Li, Junlong  and
      Wang, Jinyuan  and
      Zhang, Zhuosheng  and
      Zhao, Hai",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.17/",
    doi = "10.18653/v1/2024.naacl-long.17",
    pages = "296--310"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/