Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
Junlong LiJinyuan WangZhuosheng ZhangHai Zhao
Proposes a zero-shot open-domain question-answering framework that prompts large language models to generate synthetic passages, QA pairs, and explanations from scratch, using clustering-based retrieval to assemble in-context demonstrations that match supervised retrieval-augmented models.
Open-domain question answering requires answering broad factual questions without being provided specific reference documents. Traditional systems rely on extensive labeled datasets and external text corpora to train retrieval and answering pipelines, which creates substantial costs and operational overhead. While large language models can answer general questions directly from their internal knowledge, standard zero-shot prompting techniques underutilize their capabilities and typically lag behind customized, fine-tuned models.
The article evaluates a framework called Self-Prompting, which aims to improve zero-shot open-domain question answering without using any human-annotated training data or external document databases. The approach demonstrates how a large language model can autonomously generate its own high-quality reference examples and leverage them to guide its final answers.
The evaluated framework operates in two main phases. In the offline preparation phase, the language model generates around 5,000 synthetic question-answer pairs across 29 topics by creating short factual passages, extracting key entities, generating matching questions, verifying answer accuracy, and drafting brief explanations. In the inference phase, the system uses a clustering-based retrieval algorithm to dynamically select a diverse yet semantically relevant set of 10 synthetic demonstrations to prepend to each incoming test question. The article evaluated this methodology across three standard benchmark datasets: Web Questions, Natural Questions, and TriviaQA.
The key findings highlight substantial performance gains. First, Self-Prompting outperformed direct prompting baselines by an average of 15.5 Exact Match points and surpassed the previous state-of-the-art zero-shot method by 8.8 points across the three benchmarks. Second, the system achieved performance comparable to strong fine-tuned and retrieval-augmented models, as well as few-shot baselines that rely on real training data. Third, structuring demonstrations to provide the answer first followed by a concise explanation delivered superior accuracy compared to standard reasoning chains or multi-step generation approaches. Fourth, the framework proved consistently effective across multiple model architectures and parameter sizes, including smaller open-source models.
These results demonstrate that large language models contain sufficient internal world knowledge to serve as their own reference source for open-domain questions. For organizations, this approach eliminates the need to build, index, and maintain multi-gigabyte external corpora, while removing the recurring costs of human annotation. Additionally, generating brief explanatory statements alongside answers enhances system transparency and trustworthiness without introducing the high latency of multi-step reasoning chains.
Decision-makers should consider adopting offline synthetic demonstration pools combined with clustered retrieval when deploying language models for factual question answering where labeled data is unavailable. A single offline generation pipeline—costing approximately $120 and taking six hours in the evaluated setup—offers a more cost-effective and accurate operational strategy than real-time contextual generation or maintaining large external indexes.
Confidence in these findings is supported by consistent gains across multiple benchmarks and models. However, leaders should note several limitations: initial prompt engineering requires iterative adjustment, commercial application programming interface costs during the data generation phase can be significant, and smaller language models exhibit higher rates of factual inaccuracies in synthetic data generation compared to top-tier commercial models.
- Paper: Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations, Xinxi Lyu et al. (2023). Z-ICL establishes how automatically constructed pseudo-demonstrations can lift zero-shot performance, preparing you for Self-Prompting’s use of synthetic examples for factual QA.
- Paper: Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models, Zhihong Shao et al. (2023). Synthetic Prompting introduces model-generated demonstrations and clustered selection, the key ideas Self-Prompting adapts from reasoning tasks to open-domain QA.
- Paper: UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation, Daixuan Cheng et al. (2023). UPRISE explains how retrieving demonstrations can improve zero-shot inference, providing the retrieval framework that makes Self-Prompting’s example-selection phase easier to follow.
- Paper: Unified Demonstration Retriever for In-Context Learning, Xiaonan Li et al. (2023). The Unified Demonstration Retriever develops task-aware example ranking, clarifying the demonstration-selection problem that Self-Prompting addresses with clustering.
- Paper: Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL, Xuhan Tong et al. (2026). This later theoretical analysis generalizes how demonstrations and prompt design affect in-context learning, giving a formal lens on the synthetic-example strategy used by Self-Prompting.
- Paper: In-Context Learning with Long-Context Models: An In-Depth Exploration, Amanda Bertsch et al. (2025). This later study tests demonstration retrieval against thousands of in-context examples, extending Self-Prompting’s small, dynamically selected demonstration pool into the long-context regime.
