Ask Me Anything: A simple strategy for prompting language models
Simran AroraAvanika NarayanMayee F. ChenLaurel J. OrrNeel GuhaKush BhatiaInes ChamiChristopher Ré
Proposes a prompting strategy that converts inputs into open-ended question-answering prompts and combines their outputs with weak supervision, enabling small open-source models like GPT-J-6B to outperform few-shot GPT-3 across standard benchmarks.
Large language models can perform diverse tasks out of the box through natural language instructions, known as prompting. However, prompting is notoriously brittle, as slight phrasing changes often cause dramatic swings in accuracy. As a result, practitioners spend significant time and effort manually tuning prompts for specific models and tasks. While aggregating multiple prompt outputs has been proposed to reduce variance, standard majority voting fails to handle the wide variations in prompt accuracy and the complex error dependencies between different prompts.
To address this challenge, the article introduces and evaluates Ask Me Anything (AMA) prompting, a method designed to systematically generate effective prompts and reliably aggregate their predictions without requiring labeled data or model fine-tuning. The approach aims to show that combining multiple imperfect prompts can match or exceed the performance of much larger language models.
The evaluated approach uses a multi-step pipeline across 20 standard language processing benchmarks and 14 open-source language models spanning multiple model families and parameter sizes from 125 million to 175 billion parameters. First, the method uses functional prompt chains to recursively transform input statements into open-ended question-answering formats, which align better with pretraining objectives than restrictive true-or-false prompts. Next, it applies weak supervision—a statistical framework that estimates prompt accuracies and error dependencies using unlabeled data—to combine intermediate model responses into a final prediction.
The evaluation produced several key findings. Across all tested open-source models, AMA achieved an average absolute performance gain of 10.2% over standard few-shot baselines, with average relative gains exceeding 21%. Most notably, applying AMA to the open-source 6-billion parameter GPT-J model allowed it to match or outperform the much larger 175-billion parameter GPT-3 model on 15 of 20 benchmarks, outperforming it on average across all tasks despite having roughly 30 times fewer parameters. Furthermore, weak supervision aggregation outperformed standard majority voting by up to 8.7 accuracy points, successfully identifying error correlations across multiple prompts without labeled supervision.
These findings indicate that organizations can achieve state-of-the-art task performance using significantly smaller, open-source models that can be hosted locally and securely. This drastically lowers computational and financial costs, reduces reliance on proprietary cloud application programming interfaces, and simplifies compliance requirements for applications involving sensitive or proprietary data.
Organizations developing or deploying language models should consider adopting open-ended question-answering prompt structures and multi-prompt aggregation pipelines in place of extensive manual prompt engineering. Before full production deployment, technical teams should pilot prompt-chaining and weak supervision aggregation on their specific domain workflows to measure inference latency trade-offs, as running multiple prompt chains increases the number of inference passes.
While AMA demonstrates substantial gains on reading comprehension and context-heavy reasoning tasks, the results show limitations on closed-book tasks that require memorized factual recall, changing real-time data, or narrow domain knowledge. Decision-makers should have high confidence in the method for context-based understanding tasks, but exercise caution and consider integrating external retrieval tools when applications depend heavily on external world knowledge.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This foundational work establishes the core concept of generating diverse prompt paraphrases and ensembling their outputs to extract knowledge from language models more reliably than manual prompting.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). It analyzes the extreme sensitivity and systematic biases of few-shot prompting, providing key motivations for developing automated multi-prompt aggregation strategies.
- Paper: Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, Yao Lu et al. (2021). It demonstrates how prompt variations and sample ordering cause severe prediction variance, underpinning the need for methods that mitigate prompt brittleness.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It investigates what makes in-context prompting work, establishing why prompt formatting often matters more than ground-truth demonstration labels.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). It introduces the concept of using the language model itself to generate and optimize effective prompt candidates, a central mechanism in the AMA framework.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This comprehensive survey categorizes the mechanics of prompt engineering, multi-prompt ensembling, and output aggregation across language model architectures.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). It introduces the standard few-shot in-context learning paradigm that AMA benchmarks against and aims to improve upon with open-source models.
- Paper: Mixture-of-Agents Enhances Large Language Model Capabilities, Junlin Wang et al. (2024). It generalizes multi-prompt aggregation by orchestrating collaborative layered networks of distinct language model agents to synthesize high-quality collective responses.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It advances prompt aggregation concepts into structured graph-based networks that combine, evaluate, and refine intermediate thoughts across multiple LLM reasoning paths.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). It explores alternative training-free enhancement by using iterative self-generated feedback and refinement rather than multi-prompt weak supervision.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). It extends multi-prompt questioning strategies into an interactive cross-examination protocol between models to detect factual inconsistencies and errors.
- Paper: Prompt Repetition Improves Non-Reasoning LLMs, Yaniv Leviathan et al. (2025). It builds upon prompt structure optimization by demonstrating how simple textual query repetition enhances output accuracy in direct non-reasoning LLM queries.
