An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels
Taylor SorensenJoshua RobinsonChristopher Michael RyttingAlexander Glenn ShawKyle Jeffrey RogersAlexia Pauline DeloreyMahmoud KhalilNancy FuldaDavid Wingate
Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.
Large language models store vast linguistic and factual knowledge, but their performance depends heavily on the exact wording of the natural language prompt templates used to query them. Minor prompt rephrasings can cause severe swings in accuracy. Currently, identifying effective prompt templates requires either large quantities of expensive labeled data for validation or direct access to internal model weights for gradient-based optimization, which is often impossible when using proprietary or commercial application programming interfaces.
The article demonstrates a method for selecting high-performing prompt templates without requiring any labeled ground-truth data or access to model parameters. Specifically, the researchers evaluate whether an information-theoretic metric called mutual information can serve as a dependable surrogate for task accuracy when choosing among candidate prompt templates.
The researchers developed a five-step framework centered on maximizing the mutual information between the model's inputs and outputs across candidate templates. In high-level terms, this metric favors prompt designs where the model exhibits high output confidence on individual inputs while maintaining an unbiased, balanced distribution of predictions overall. The team evaluated 20 candidate templates per dataset across eight benchmark datasets covering seven standard language processing tasks, including question-answering, sentiment analysis, and reading comprehension. The evaluation spanned eight causal language models ranging in scale from 124 million to 175 billion parameters.
The analysis yielded several key findings. First, mutual information strongly correlates with actual test accuracy, particularly in larger language models. Second, on the largest 175-billion-parameter model, prompt selection via mutual information closed approximately 90% of the gap between the average candidate prompt accuracy and the optimal prompt accuracy, achieving near-perfect selection performance without labels. Third, mutual information proved effective even in extremely low-data settings, selecting better prompts using as few as two unlabeled examples than standard methods could select using two labeled examples. Fourth, creating a combined ensemble of the top five prompt templates ranked by mutual information matched or outperformed an ensemble of all twenty candidate prompts while reducing computational costs by about 75%. Finally, prompt templates chosen via mutual information demonstrated strong cross-model transferability, showing that prompts selected on large models generalize well to others.
These findings indicate that organizations can effectively optimize language model performance without incurring the timeline delays and high labor costs associated with manual data labeling. The approach also enables high-quality prompt selection when interacting with closed-source, black-box systems where model parameters are hidden. Furthermore, prompt ensembling using top mutual information scores offers a cost-effective hedge against occasionally deceptive or erratic prompt templates.
Practitioners should adopt mutual information-based selection when deploying models on tasks lacking labeled validation data, starting with a diverse set of candidate templates. For mission-critical applications requiring risk mitigation, teams should ensemble predictions across the top-ranked templates. However, decision-makers should note that this method relies on the underlying model already possessing the capability to perform the given task; mutual information cannot compensate if all candidate prompts are poor or if the task exceeds the model's baseline reasoning capabilities. Confidence remains highest when applying the technique to large, well-calibrated models.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This survey establishes the prompt-template and prompt-selection landscape that the source’s label-free template-scoring method addresses.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). This later method carries information-theoretic prompt selection into in-context learning by using an information criterion to select and order examples without labeled validation data.
