keyword
prompt selection
Prompt selection is the process of identifying and choosing the most effective prompt, such as a specific instruction template, phrasing, or set of in-context demonstration examples, from a pool of candidates to optimize the performance of a pretrained language or multimodal model on a given task. Because model outputs are often highly sensitive to subtle variations in wording, format, reasoning complexity, and example ordering, prompt selection aims to find configurations that maximize task accuracy, consistency, and reasoning quality. This selection can be performed manually or through automated techniques, utilizing evaluation methods such as validation performance, heuristic scoring, information-theoretic metrics, search algorithms, or retrieval mechanisms to identify high-performing prompts without updating the underlying model parameters.
5 items

Complexity-Based Prompting for Multi-Step Reasoning
Yao Fu, Hao-Chun Peng, Ashish Sabharwal, Peter Clark, Tushar Khot
Why you should read this
Introduces complexity-based prompting, a simple strategy showing that selecting in-context examples and decoding paths with more reasoning steps substantially boosts large language model performance on multi-step mathematical and logical reasoning tasks.
We study the task of prompting large-scale language models to perform multi-step reasoning. Existing work shows that when prompted with a chain of thoughts (CoT), sequences of short sentences describing intermediate reasoning steps towards a final answer, large language models can generate new reasoning chains and predict answers for new inputs. A central question is which reasoning examples make the most effective prompts. In this work, we propose complexity-based prompting, a simple and effective example selection scheme for multi-step reasoning. We show that prompts with higher reasoning complexity, i.e., chains with more reasoning steps, achieve substantially better performance on multi-step reasoning tasks over strong baselines. We further extend our complexity-based criteria from prompting (selecting inputs) to decoding (selecting outputs), where we sample multiple reasoning chains from the model, then choose the majority of generated answers from complex reasoning chains (over simple chains). When used to prompt GPT-3 and Codex, our approach substantially improves multi-step reasoning accuracy and achieves new state-of-the-art (SOTA) performance on three math benchmarks (GSM8K, MultiArith, and MathQA) and two BigBenchHard tasks (Date Understanding and Penguins), with an average +5.3 and up to +18 accuracy improvements. Compared with existing example selection schemes like manual tuning or retrieval-based selection, selection based on reasoning complexity is intuitive, easy to implement, and annotation-efficient. Further results demonstrate the robustness of performance gains from complex prompts under format perturbation and distribution shift.
Added
2026-10-05

GPS: Genetic Prompt Search for Efficient Few-Shot Learning
Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, Zhilin Yang
Why you should read this
Proposes a gradient-free genetic algorithm that automatically discovers high-performing, fluent natural language prompts using only a tiny validation set, outperforming both manual prompt engineering and parameter-efficient tuning methods.
Prompt-based techniques have demonstrated great potential for improving the few-shot generalization of pretrained language models. However, their performance heavily relies on the manual design of prompts and thus requires a lot of human efforts. In this paper, we introduce Genetic Prompt Search (GPS) to improve few-shot learning with prompts, which utilizes a genetic algorithm to automatically search for high-performing prompts. GPS is gradient-free and requires no update of model parameters but only a small validation set. Experiments on diverse datasets proved the effectiveness of GPS, which outperforms manual prompts by a large margin of 2.6 points. Our method is also better than other parameter-efficient tuning methods such as prompt tuning.
Added
2026-10-03

A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models
James Urquhart Allingham, Jie Ren, Michael W. Dusenberry, Xiuye Gu, Yin Cui, Dustin Tran, Jeremiah Zhe Liu, Balaji Lakshminarayanan
Why you should read this
Proposes a bias-corrected zero-shot prompt weighting algorithm that automatically scores and ensembles prompts for text-image models without needing labeled validation data or manual prompt engineering.
Added
2026-10-03

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels
Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, David Wingate
Why you should read this
Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.
Pre-trained language models derive substantial linguistic and factual knowledge from the massive corpora on which they are trained, and prompt engineering seeks to align these models to specific tasks. Unfortunately, existing prompt engineering methods require significant amounts of labeled data, access to model parameters, or both. We introduce a new method for selecting prompt templates without labeled examples and without direct access to the model. Specifically, over a set of candidate templates, we choose the template that maximizes the mutual information between the input and the corresponding model output. Across 8 datasets representing 7 distinct NLP tasks, we show that when a template has high mutual information, it also has high accuracy on the task. On the largest model, selecting prompts with our method gets 90% of the way from the average prompt accuracy to the best prompt accuracy and requires no ground truth labels.
Added
2026-10-01

Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp
Why you should read this
Demonstrates that few-shot language models are highly sensitive to prompt example ordering—causing performance to swing between state-of-the-art and random chance—and introduces an entropy-based method to find optimal permutations without extra labeled data.
When primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, fine-tuned, large, pretrained language models. We demonstrate that the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance: essentially some permutations are "fantastic" and some not. We analyse this phenomenon in detail, establishing that: it is present across model sizes (even for the largest current models), it is not related to a specific subset of samples, and that a given good permutation for one model is not transferable to another. While one could use a development set to determine which permutations are performant, this would deviate from the true few-shot setting as it requires additional annotated data. Instead, we use the generative nature of language models to construct an artificial development set and based on entropy statistics of the candidate permutations on this set, we identify performant prompts. Our method yields a 13% relative improvement for GPT-family models across eleven different established text classification tasks.
Added
2026-09-24
