Ask Me Anything: A simple strategy for prompting language models

Simran AroraAvanika NarayanMayee F. ChenLaurel J. OrrNeel GuhaKush BhatiaInes ChamiChristopher Ré

article2023ICLR265 citations

Proposes a prompting strategy that converts inputs into open-ended question-answering prompts and combines their outputs with weak supervision, enabling small open-source models like GPT-J-6B to outperform few-shot GPT-3 across standard benchmarks.

Listen

Large language models can perform diverse tasks out of the box through natural language instructions, known as prompting. However, prompting is notoriously brittle, as slight phrasing changes often cause dramatic swings in accuracy. As a result, practitioners spend significant time and effort manually tuning prompts for specific models and tasks. While aggregating multiple prompt outputs has been proposed to reduce variance, standard majority voting fails to handle the wide variations in prompt accuracy and the complex error dependencies between different prompts.

To address this challenge, the article introduces and evaluates Ask Me Anything (AMA) prompting, a method designed to systematically generate effective prompts and reliably aggregate their predictions without requiring labeled data or model fine-tuning. The approach aims to show that combining multiple imperfect prompts can match or exceed the performance of much larger language models.

The evaluated approach uses a multi-step pipeline across 20 standard language processing benchmarks and 14 open-source language models spanning multiple model families and parameter sizes from 125 million to 175 billion parameters. First, the method uses functional prompt chains to recursively transform input statements into open-ended question-answering formats, which align better with pretraining objectives than restrictive true-or-false prompts. Next, it applies weak supervision—a statistical framework that estimates prompt accuracies and error dependencies using unlabeled data—to combine intermediate model responses into a final prediction.

The evaluation produced several key findings. Across all tested open-source models, AMA achieved an average absolute performance gain of 10.2% over standard few-shot baselines, with average relative gains exceeding 21%. Most notably, applying AMA to the open-source 6-billion parameter GPT-J model allowed it to match or outperform the much larger 175-billion parameter GPT-3 model on 15 of 20 benchmarks, outperforming it on average across all tasks despite having roughly 30 times fewer parameters. Furthermore, weak supervision aggregation outperformed standard majority voting by up to 8.7 accuracy points, successfully identifying error correlations across multiple prompts without labeled supervision.

These findings indicate that organizations can achieve state-of-the-art task performance using significantly smaller, open-source models that can be hosted locally and securely. This drastically lowers computational and financial costs, reduces reliance on proprietary cloud application programming interfaces, and simplifies compliance requirements for applications involving sensitive or proprietary data.

Organizations developing or deploying language models should consider adopting open-ended question-answering prompt structures and multi-prompt aggregation pipelines in place of extensive manual prompt engineering. Before full production deployment, technical teams should pilot prompt-chaining and weak supervision aggregation on their specific domain workflows to measure inference latency trade-offs, as running multiple prompt chains increases the number of inference passes.

While AMA demonstrates substantial gains on reading comprehension and context-heavy reasoning tasks, the results show limitations on closed-book tasks that require memorized factual recall, changing real-time data, or narrow domain knowledge. Decision-makers should have high confidence in the method for context-based understanding tasks, but exercise caution and consider integrating external retrieval tools when applications depend heavily on external world knowledge.

Cover for Ask Me Anything: A simple strategy for prompting language models

Abstract

Large language models (LLMs) transfer well to new tasks out-of-the-box simply given a natural language prompt that demonstrates how to perform the task and no additional training. Prompting is a brittle process wherein small modifications to the prompt can cause large variations in the model predictions, and therefore significant effort is dedicated towards designing a painstakingly "perfect prompt" for a task. To mitigate the high degree of effort involved in prompt-design, we instead ask whether producing multiple effective, yet imperfect, prompts and aggregating them can lead to a high quality prompting strategy. Our observations motivate our proposed prompting method, ASK ME ANYTHING (AMA). We first develop an understanding of the effective prompt formats, finding that question-answering (QA) prompts, which encourage open-ended generation ("Who went to the park?") tend to outperform those that restrict the model outputs ("John went to the park. Output True or False."). Our approach recursively uses the LLM itself to transform task inputs to the effective QA format. We apply the collected prompts to obtain several noisy votes for the input's true label. We find that the prompts can have very different accuracies and complex dependencies and thus propose to use weak supervision, a procedure for combining the noisy predictions, to produce the final predictions for the inputs. We evaluate AMA across open-source model families (e.g., EleutherAI, BLOOM, OPT, and T0) and model sizes (125M-175B parameters), demonstrating an average performance lift of 10.2% over the few-shot baseline. This simple strategy enables the open-source GPT-J-6B model to match and exceed the performance of few-shot GPT3-175B on 15 of 20 popular benchmarks. Averaged across these tasks, the GPT-J-6B model outperforms few-shot GPT3-175B. We release our code here: this https URL

Citation

MLA
Arora, S., et al. “Ask Me Anything: A Simple Strategy for Prompting Language Models”. arXiv, 2022, http://arxiv.org/abs/2210.02441v3.
APA
Arora, S., Narayan, A., Chen, M. F., Orr, L., Guha, N., Bhatia, K., Chami, I., Sala, F., & Ré, C. (2022). Ask Me Anything: A simple strategy for prompting language models. arXiv. http://arxiv.org/abs/2210.02441v3
Chicago
Arora, S., A. Narayan, M. F. Chen, et al. 2022. “Ask Me Anything: A Simple Strategy for Prompting Language Models”. arXiv. http://arxiv.org/abs/2210.02441v3.
Harvard
Arora, S. et al. (2022) “Ask Me Anything: A simple strategy for prompting language models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.02441v3.
Vancouver
1. Arora S, Narayan A, Chen MF, Orr L, Guha N, Bhatia K, Chami I, Sala F, Ré C (2022) Ask Me Anything: A simple strategy for prompting language models. arXiv

BibTeX

@article{arora2022ask,
  title = {Ask Me Anything: A simple strategy for prompting language models},
  author = {Arora, Simran and Narayan, Avanika and Chen, Mayee F. and Orr, Laurel and Guha, Neel and Bhatia, Kush and Chami, Ines and Sala, Frederic and Ré, Christopher},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.02441v3},
  eprint = {2210.02441}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors