Built independently by an author, for readers. Read the story and support ChapterPal

keyword

prompt selection

Prompt selection is the process of identifying and choosing the most effective prompt, such as a specific instruction template, phrasing, or set of in-context demonstration examples, from a pool of candidates to optimize the performance of a pretrained language or multimodal model on a given task. Because model outputs are often highly sensitive to subtle variations in wording, format, reasoning complexity, and example ordering, prompt selection aims to find configurations that maximize task accuracy, consistency, and reasoning quality. This selection can be performed manually or through automated techniques, utilizing evaluation methods such as validation performance, heuristic scoring, information-theoretic metrics, search algorithms, or retrieval mechanisms to identify high-performing prompts without updating the underlying model parameters.

5 items

Complexity-Based Prompting for Multi-Step Reasoning

Complexity-Based Prompting for Multi-Step Reasoning

Yao Fu, Hao-Chun Peng, Ashish Sabharwal, Peter Clark, Tushar Khot

OrganizationsAllen Institute for AIUniversity of Edinburgh

Why you should read this

Introduces complexity-based prompting, a simple strategy showing that selecting in-context examples and decoding paths with more reasoning steps substantially boosts large language model performance on multi-step mathematical and logical reasoning tasks.

We study the task of prompting large-scale language models to perform multi-step reasoning. Existing work shows that when prompted with a chain of thoughts (CoT), sequences of short sentences describing intermediate reasoning steps towards a final answer, large language models can generate new reasoning chains and predict answers for new inputs. A central question is which reasoning examples make the most effective prompts. In this work, we propose complexity-based prompting, a simple and effective example selection scheme for multi-step reasoning. We show that prompts with higher reasoning complexity, i.e., chains with more reasoning steps, achieve substantially better performance on multi-step reasoning tasks over strong baselines. We further extend our complexity-based criteria from prompting (selecting inputs) to decoding (selecting outputs), where we sample multiple reasoning chains from the model, then choose the majority of generated answers from complex reasoning chains (over simple chains). When used to prompt GPT-3 and Codex, our approach substantially improves multi-step reasoning accuracy and achieves new state-of-the-art (SOTA) performance on three math benchmarks (GSM8K, MultiArith, and MathQA) and two BigBenchHard tasks (Date Understanding and Penguins), with an average +5.3 and up to +18 accuracy improvements. Compared with existing example selection schemes like manual tuning or retrieval-based selection, selection based on reasoning complexity is intuitive, easy to implement, and annotation-efficient. Further results demonstrate the robustness of performance gains from complex prompts under format perturbation and distribution shift.

Added

2026-10-05

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, David Wingate

OrganizationsBrigham Young University

Why you should read this

Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.

Pre-trained language models derive substantial linguistic and factual knowledge from the massive corpora on which they are trained, and prompt engineering seeks to align these models to specific tasks. Unfortunately, existing prompt engineering methods require significant amounts of labeled data, access to model parameters, or both. We introduce a new method for selecting prompt templates without labeled examples and without direct access to the model. Specifically, over a set of candidate templates, we choose the template that maximizes the mutual information between the input and the corresponding model output. Across 8 datasets representing 7 distinct NLP tasks, we show that when a template has high mutual information, it also has high accuracy on the task. On the largest model, selecting prompts with our method gets 90% of the way from the average prompt accuracy to the best prompt accuracy and requires no ground truth labels.

Added

2026-10-01

Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp

OrganizationsMishcon de Reya LLPUniversity College London

Why you should read this

Demonstrates that few-shot language models are highly sensitive to prompt example ordering—causing performance to swing between state-of-the-art and random chance—and introduces an entropy-based method to find optimal permutations without extra labeled data.

When primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, fine-tuned, large, pretrained language models. We demonstrate that the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance: essentially some permutations are "fantastic" and some not. We analyse this phenomenon in detail, establishing that: it is present across model sizes (even for the largest current models), it is not related to a specific subset of samples, and that a given good permutation for one model is not transferable to another. While one could use a development set to determine which permutations are performant, this would deviate from the true few-shot setting as it requires additional annotated data. Instead, we use the generative nature of language models to construct an artificial development set and based on entropy statistics of the candidate permutations on this set, we identify performant prompts. Our method yields a 13% relative improvement for GPT-family models across eleven different established text classification tasks.

Added

2026-09-24