Instruction Induction: From Few Examples to Natural Language Task Descriptions

Or HonovichUri ShahamSamuel R. BowmanOmer Levy

article2023ACL225 citations

Demonstrates that instruction-tuned large language models can explicitly infer underlying tasks from a few input-output demonstrations by generating precise natural language instructions, establishing instruction induction as an interpretable learning paradigm evaluated on a new 24-task benchmark.

Listen

Organizations increasingly deploy large language models to automate complex workflows using few-shot demonstrations, where models implicitly infer how to complete a task from a handful of examples. However, this process remains opaque and difficult to verify because models typically execute tasks without explaining the underlying rule. Addressing this challenge, the article investigates whether language models can explicitly infer and describe an underlying task in plain, natural language when presented with input-output examples, framing this capability as "instruction induction."

The main objective of the article is to demonstrate that large, instruction-aligned language models can infer natural language instructions from a few examples in a zero-shot setting and to evaluate the quality and functionality of these generated instructions across diverse tasks.

To evaluate this capability, the authors established a benchmark across 24 distinct language tasks covering spelling, syntax, semantic similarity, and style transfer. They created a dataset with 100 induction instances per task, each presenting five input-output pairs framed within a puzzle-style prompt. The study evaluated model outputs through two primary approaches: reference-based text similarity (using BERTScore against human-written instructions) and a novel functional metric termed "execution accuracy." Execution accuracy tests whether a generated instruction can successfully guide an independent language model to execute the task on 100 held-out test examples without demonstrations. The evaluation compared eight versions of GPT-3 and InstructGPT against human baselines collected from college-graduate annotators.

The findings reveal several key outcomes. First, the ability to induce accurate instructions emerges strongly in the largest instruction-aligned model (InstructGPT, 175 billion parameters), which achieved an average execution accuracy of 43.6%—representing approximately 65.7% of the human control performance of 66.4%. Second, instruction induction failed in non-aligned and smaller models; the original 175-billion-parameter GPT-3 achieved an execution accuracy of only 6.5% (roughly 9.8% of human performance), while smaller models scored below 9% regardless of instruction tuning. Third, on 12 of the 24 tasks, InstructGPT achieved at least 75% of human execution performance, matching or exceeding human descriptions on straightforward tasks such as formality transfer and first-letter extraction. Finally, qualitative error analyses showed that models occasionally generate imprecise or oversimplified instructions—such as prompting "Reverse the input" for an antonym task—which can lead to execution failures.

These results demonstrate that instruction induction represents a viable new learning paradigm, replacing opaque continuous parameter tuning with human-interpretable natural language hypotheses. For technology leaders, grounding model behavior in natural language rules significantly improves explainability, compliance, and system verifiability, while reducing risks related to spurious correlations and deployment failures.

Organizations evaluating or deploying language models should consider using instruction induction to automatically extract and audit rules from user demonstrations rather than relying solely on black-box prompting. Before deploying instruction induction in production, decision-makers should pilot the approach with human-in-the-loop validation, and researchers should extend the benchmark to complex, multi-step tasks with more than five demonstrations.

Confidence in these findings is strong for simple-to-moderate language tasks, but readers should exercise caution regarding broader generalization. The capabilities were observed solely in the proprietary InstructGPT model, whose closed architecture limits granular understanding of why the capability emerged. Additionally, the execution accuracy metric relies on the proficiency of the executing model, meaning performance on complex, highly nuanced workflows may exhibit greater uncertainty.

Cover for Instruction Induction: From Few Examples to Natural Language Task Descriptions

Abstract

Large language models are able to perform a task by conditioning on a few input-output demonstrations – a paradigm known as in-context learning. We show that language models can explicitly infer an underlying task from a few demonstrations by prompting them to generate a natural language instruction that fits the examples. To explore this ability, we introduce the instruction induction challenge, compile a dataset consisting of 24 tasks, and define a novel evaluation metric based on executing the generated instruction. We discover that, to a large extent, the ability to generate instructions does indeed emerge when using a model that is both large enough and aligned to follow instructions; InstructGPT achieves 65.7% of human performance in our execution-based metric, while the original GPT-3 model reaches only 9.8% of human performance. This surprising result suggests that instruction induction might be a viable learning paradigm in and of itself, where instead of fitting a set of latent continuous parameters to the data, one searches for the best description in the natural language hypothesis space.

Citation

MLA
Honovich, O., et al. “Instruction Induction: From Few Examples to Natural Language Task Descriptions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 1935–52, https://doi.org/10.18653/v1/2023.acl-long.108.
APA
Honovich, O., Shaham, U., Bowman, S. R., & Levy, O. (2023). Instruction Induction: From Few Examples to Natural Language Task Descriptions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1935–1952. https://doi.org/10.18653/v1/2023.acl-long.108
Chicago
Honovich, O., U. Shaham, S. R. Bowman, and O. Levy. 2023. “Instruction Induction: From Few Examples to Natural Language Task Descriptions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1935–52. https://doi.org/10.18653/v1/2023.acl-long.108.
Harvard
Honovich, O. et al. (2023) “Instruction Induction: From Few Examples to Natural Language Task Descriptions”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1935–1952. Available at: https://doi.org/10.18653/v1/2023.acl-long.108.
Vancouver
1. Honovich O, Shaham U, Bowman SR, Levy O (2023) Instruction Induction: From Few Examples to Natural Language Task Descriptions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1935–1952

BibTeX

@inproceedings{honovich-etal-2023-instruction,
    title = "Instruction Induction: From Few Examples to Natural Language Task Descriptions",
    author = "Honovich, Or  and
      Shaham, Uri  and
      Bowman, Samuel R.  and
      Levy, Omer",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.108/",
    doi = "10.18653/v1/2023.acl-long.108",
    pages = "1935--1952"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/