Instruction Induction: From Few Examples to Natural Language Task Descriptions
Or HonovichUri ShahamSamuel R. BowmanOmer Levy
Demonstrates that instruction-tuned large language models can explicitly infer underlying tasks from a few input-output demonstrations by generating precise natural language instructions, establishing instruction induction as an interpretable learning paradigm evaluated on a new 24-task benchmark.
Organizations increasingly deploy large language models to automate complex workflows using few-shot demonstrations, where models implicitly infer how to complete a task from a handful of examples. However, this process remains opaque and difficult to verify because models typically execute tasks without explaining the underlying rule. Addressing this challenge, the article investigates whether language models can explicitly infer and describe an underlying task in plain, natural language when presented with input-output examples, framing this capability as "instruction induction."
The main objective of the article is to demonstrate that large, instruction-aligned language models can infer natural language instructions from a few examples in a zero-shot setting and to evaluate the quality and functionality of these generated instructions across diverse tasks.
To evaluate this capability, the authors established a benchmark across 24 distinct language tasks covering spelling, syntax, semantic similarity, and style transfer. They created a dataset with 100 induction instances per task, each presenting five input-output pairs framed within a puzzle-style prompt. The study evaluated model outputs through two primary approaches: reference-based text similarity (using BERTScore against human-written instructions) and a novel functional metric termed "execution accuracy." Execution accuracy tests whether a generated instruction can successfully guide an independent language model to execute the task on 100 held-out test examples without demonstrations. The evaluation compared eight versions of GPT-3 and InstructGPT against human baselines collected from college-graduate annotators.
The findings reveal several key outcomes. First, the ability to induce accurate instructions emerges strongly in the largest instruction-aligned model (InstructGPT, 175 billion parameters), which achieved an average execution accuracy of 43.6%—representing approximately 65.7% of the human control performance of 66.4%. Second, instruction induction failed in non-aligned and smaller models; the original 175-billion-parameter GPT-3 achieved an execution accuracy of only 6.5% (roughly 9.8% of human performance), while smaller models scored below 9% regardless of instruction tuning. Third, on 12 of the 24 tasks, InstructGPT achieved at least 75% of human execution performance, matching or exceeding human descriptions on straightforward tasks such as formality transfer and first-letter extraction. Finally, qualitative error analyses showed that models occasionally generate imprecise or oversimplified instructions—such as prompting "Reverse the input" for an antonym task—which can lead to execution failures.
These results demonstrate that instruction induction represents a viable new learning paradigm, replacing opaque continuous parameter tuning with human-interpretable natural language hypotheses. For technology leaders, grounding model behavior in natural language rules significantly improves explainability, compliance, and system verifiability, while reducing risks related to spurious correlations and deployment failures.
Organizations evaluating or deploying language models should consider using instruction induction to automatically extract and audit rules from user demonstrations rather than relying solely on black-box prompting. Before deploying instruction induction in production, decision-makers should pilot the approach with human-in-the-loop validation, and researchers should extend the benchmark to complex, multi-step tasks with more than five demonstrations.
Confidence in these findings is strong for simple-to-moderate language tasks, but readers should exercise caution regarding broader generalization. The capabilities were observed solely in the proprietary InstructGPT model, whose closed architecture limits granular understanding of why the capability emerged. Additionally, the execution accuracy metric relies on the proficiency of the executing model, meaning performance on complex, highly nuanced workflows may exhibit greater uncertainty.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Introduces Automatic Prompt Engineer (APE) to treat prompt generation as program synthesis using execution feedback on the same 24 instruction induction tasks.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Presents InstructGPT and instruction alignment via RLHF, which is directly evaluated as the key emergent driver of instruction induction capabilities.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Establishes few-shot in-context learning in GPT-3, providing the foundational paradigm from which instruction induction inverts demonstrations into explicit text descriptions.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Introduces instruction tuning to enable zero-shot task execution from explicit natural language descriptions.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Analyzes the essential mechanisms of few-shot demonstrations in context, clarifying how language models interpret input-output exemplars.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). Explores meta-training language models to infer task semantics directly from concatenated input-output pairs.
- Paper: Meta-learning via Language Model In-context Tuning, Yanda Chen et al. (2022). Studies how tuning language models on few-shot prompts conditions them to learn and generalize task structures from demonstrations.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Analyzes prompt formulation and the relationship between runtime few-shot exemplars and explicit zero-shot task descriptions.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). Leverages language models' ability to generate instructions and instances to bootstrap synthetic instruction-tuning datasets without human supervision.
- Paper: Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor, Or Honovich et al. (2023). Scales automated instruction generation and synthetic prompt bootstrapping to build large-scale instruction-following datasets with minimal seed demonstrations.
- Paper: Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations, Chenglei Si et al. (2023). Investigates the inductive biases and hypothesis-selection behaviors of language models when interpreting underspecified few-shot demonstrations.
- Paper: WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, Can Xu et al. (2023). Extends automated instruction generation by iteratively evolving and complexifying generated prompts for advanced model alignment.
- Paper: Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks, Po-Nien Kung et al. (2023). Applies prompt uncertainty and sensitivity analysis to optimize task selection and generalization in instruction tuning.
- Paper: From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning, Xuansheng Wu et al. (2024). Analyzes internal representations to uncover how instruction-tuning shifts a model's operational mechanisms toward following induced instructions.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Evaluates the capability and reliability of large language models when acting as automated judges to score instruction adherence.
- Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). Provides a comprehensive survey synthesizing the mechanics, evaluation frameworks, and emergent paradigms of in-context learning and instruction conditioning.
