Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations

Xinxi LyuSewon MinIz BeltagyLuke ZettlemoyerHannaneh Hajishirzi

article2023ACL90 citations

Introduces a zero-shot prompting method that constructs pseudo-demonstrations from raw, unlabelled text using nearest neighbors and label synonyms, matching the performance of standard few-shot in-context learning without requiring any task annotations.

Listen

Large language models deliver high performance when guided by few-shot demonstrations, but their accuracy drops significantly in zero-shot settings where no task-specific examples are provided. Because recent research indicates that demonstrations primarily communicate formatting and domain context rather than explicit training rules, standard zero-shot benchmarks substantially underestimate what these models can achieve on their own. The article evaluates a new zero-shot framework called Z-ICL, which automatically constructs synthetic demonstration examples from unannotated text to bridge the performance gap between zero-shot and few-shot classification.

To establish this framework, the article examines the copying effect: a vulnerability where language models blindly reproduce the labels of demonstration examples that closely resemble the test input. To prevent this distortion while still supplying useful context, the Z-ICL method implements a three-step procedure. It first searches a broad text repository to find close matches to the test prompt, selects adjacent sentences rather than the exact matches to ensure contextual relevance without excessive similarity, and pairs these sentences with random synonyms of the target labels. The model then evaluates the target input alongside these artificial examples without requiring manual prompt engineering.

Across nine text classification datasets evaluated on models ranging from 6 billion to 175 billion parameters, Z-ICL consistently outperformed traditional zero-shot methods by an absolute gain of 5 to 30 percentage points. On tasks where the text repository covered the relevant subject matter, the method achieved accuracy on par with standard few-shot learning that relies on human-annotated training data. Ablation experiments confirmed that using adjacent sentences and label synonyms are both essential to suppress label copying, while simply providing text without structured label pairings degraded accuracy.

These findings indicate that organizations can achieve few-shot performance levels without the expense, operational delays, and governance risks associated with sourcing human-labeled datasets. By demonstrating that unannotated text corpora can effectively activate model capabilities, the article highlights an opportunity to lower artificial intelligence deployment costs and simplify prompt engineering pipelines across routine text categorization workflows.

Organizations evaluating large-scale text classification should test retrieval-based synthetic prompts as a direct alternative to manual labeling or intensive prompt tuning. For optimal performance, practitioners should ensure that reference corpora include domains relevant to their target applications, as expanding topic coverage by just 2% produced consistent accuracy gains on previously unsupported domains. Future work should validate this approach on multi-sentence reasoning and open-ended text generation, while automating synonym selection to remove remaining manual steps.

Confidence in these findings is high for standard single-sentence text classification across multiple model architectures. However, decision-makers should recognize existing boundaries: the evaluations focused exclusively on classification tasks, relied on manually chosen label synonyms, and showed reduced performance when domain coverage in the underlying corpus was absent.

No sufficiently relevant recommendations were found.

Cover for Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations

Abstract

Large language models (LLMs) can learn new tasks through in-context learning (ICL), where demonstrations condition their predictions. However, collecting high-quality labeled examples across diverse downstream applications may be expensive.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Copying Effect Hypothesis
  • 4 Our Method: Z-ICL
  • 4.1 Step 1: Retrieve Relevant Sentences
  • 4.2 Step 2: Construct pseudo-demonstrations
  • 4.3 Step 3: Inference
  • 5 Experimental Setup
  • 5.1 Data
  • 5.2 Baselines
  • 5.3 Experimental Details
  • 6 Experimental Results
  • 6.1 Main results
  • 6.2 Ablations
  • 7 Conclusion
  • Limitation
  • Better construction of pseudo-demonstrations.
  • Acknowledgements
  • References
  • A Data Statistics
  • B Implementation Details

Knowls

  1. Knowl 1 — Z-ICL constructs zero-shot pseudo-demonstrations from raw text

    model/method

    Z-ICL predicts a class for a test sentence without using labeled or unlabeled task-specific training examples. Given a test sentence xx, a raw text corpus CC, a candidate class set YY, and a demonstration count kk, it proceeds as follows:

    1. Retrieve the kk corpus sentences most similar to xx.
    2. For each retrieved sentence, use a sentence physically adjacent to it in the corpus as a pseudo-demonstration input. This retains related input-domain content while avoiding use of the sentences closest to xx itself.
    3. Independently sample a class uniformly from YY for each pseudo-demonstration, and replace that class with a manually selected synonym. Pair each corpus sentence with its corresponding synonym.
    4. Prompt the language model with the resulting kk input–synonym pairs followed by xx. Choose the prediction from the original class set YY, not from the synonym words.

    The physical-neighbor selection and synonym labels are intended to reduce copying of labels from demonstrations while preserving information about the input distribution and label semantics. In the experiments, sentence similarity was cosine similarity between SimCSE embeddings.

  2. Knowl 2 — Z-ICL improves zero-shot classification and approaches few-shot performance

    empirical result

    On nine single-sentence classification datasets, Z-ICL was evaluated with GPT-J and GPT-NeoX using direct and channel inference. The following are the reported accuracy averages (%) across the six datasets whose domains were represented in the text corpus and the three whose domains were not. In each comparison, the figures are No-demonstrations → Z-ICL; the two oracle figures are ICL with gold labels and ICL with random labels, respectively.

    • GPT-J, channel inference: corpus-covered 61.0 → 65.8; oracle 67.9 and 67.4. Not covered 51.3 → 67.7; oracle 72.6 and 71.5.
    • GPT-J, direct inference: corpus-covered 54.8 → 61.6; oracle 65.4 and 63.2. Not covered 43.8 → 64.8; oracle 72.7 and 68.6.
    • GPT-NeoX, channel inference: corpus-covered 43.7 → 62.5; oracle 65.9 and 63.7. Not covered 48.1 → 60.2; oracle 72.1 and 71.4.
    • GPT-NeoX, direct inference: corpus-covered 42.8 → 65.4; oracle 64.4 and 63.3. Not covered 38.8 → 68.4; oracle 73.5 and 65.2.

    The results show that Z-ICL improves substantially over the no-demonstration baseline and is broadly comparable to training-example-based ICL on corpus-covered domains. It generally remains behind the oracle methods on domains absent from the corpus, highlighting the importance of corpus-domain coverage.

  3. Knowl 3 — Similar demonstrations induce a copying effect

    empirical result

    The paper’s copying-effect hypothesis is that a language model’s prediction is strongly biased toward the label paired with a demonstration input that is very similar to the test input. In experiments with GPT-J on CR, Amazon reviews, and Yelp, gold labels and random labels gave relatively comparable performance when demonstrations were sampled randomly. With nearest-neighbor demonstrations, random labels performed substantially worse than gold labels, indicating that label correctness matters more when demonstration inputs are close to the test input.

    A separate CR experiment placed a demonstration identical to the test input in the prompt, with either its correct or incorrect label. The measured outcome was the percentage of predictions matching the label paired with that identical input:

    • GPT-J: 82.3% overall, 90.8% when the paired label was correct, and 73.9% when it was incorrect.
    • GPT-NeoX: 88.0% overall, 94.2% when the paired label was correct, and 81.7% when it was incorrect.

    Thus, in these conditions, both models often followed even an incorrect label attached to an identical demonstration input.

  4. Knowl 4 — GPT-3 results show gains over no-demonstration prompting, with domain-dependent parity

    empirical result

    For GPT-3 (175B), the paper evaluated CR, Amazon reviews, Yelp, Tweet-Eval, and SST2. Reported average accuracies (%) for corpus-covered datasets and the single dataset not covered by the corpus, respectively, were:

    • Channel inference: No-demonstrations 69.5 / 80.8; Z-ICL 73.4 / 82.4; oracle ICL with gold labels 73.9 / 88.1; oracle ICL with random labels 72.3 / 84.8.
    • Direct inference: No-demonstrations 72.7 / 73.2; Z-ICL 72.7 / 78.1; oracle ICL with gold labels 79.3 / 94.2; oracle ICL with random labels 77.4 / 93.9.

    These averages show gains over no-demonstration prompting for channel inference on covered datasets and for the uncovered dataset in both inference settings; direct inference on covered datasets has the same average with and without Z-ICL. The reported results therefore support the paper’s broader finding that Z-ICL can approach oracle performance on covered domains, while performance on an uncovered domain remains more difficult.

  5. Knowl 5 — Evaluation uses raw-domain corpora, nine classification datasets, and three language models

    experimental setup

    The retrieval corpus was the English Demix corpus, which contains 16 domains, including news, web text, Reddit, reviews, academic papers, tweets, and Yelp reviews. The authors subsampled up to 10 million paragraphs per domain, split paragraphs into sentences, removed duplicate sentences while preserving their order, and retrieved at sentence level.

    Evaluation used CR, Amazon reviews (Amz), Amz5, Yelp, Yelp5, Tweet-Eval, MR, SST2, and SST5. Six datasets were considered covered by corpus domains (CR, Amz, Amz5, Yelp, Yelp5, and Tweet-Eval); MR, SST2, and SST5 were considered uncovered. The models were GPT-J (6B), GPT-NeoX (20B), and GPT-3 (175B), evaluated with direct and channel inference. Unless varied in an ablation, Z-ICL used k=16k=16 demonstrations. GPT-J and GPT-NeoX results averaged five random seeds; GPT-3 results averaged two seeds and covered five datasets. Test sets larger than 2,000 examples were subsampled uniformly without replacement to 2,000 examples.

  6. Knowl 6 — Input–label pairs, not retrieved text alone, are central to Z-ICL’s gains

    empirical result

    An ablation compared Z-ICL with a variant that conditions on retrieved inputs but omits the randomly assigned labels (“Inputs-only”). It also varied the number of demonstrations from k=1k=1 to k=64k=64. Across the tested values of kk, Z-ICL remained substantially better than zero-shot baselines and on par with oracle ICL, whereas the Inputs-only variant performed significantly worse. This comparison indicates that the paired input–label format contributes to the gains; the effect is not explained merely by supplying relevant retrieved text.

  7. Knowl 7 — Physical neighbors reduce the cost of retrieving inputs closest to the test sentence

    empirical result

    The retrieval ablation compared three ways to select pseudo-demonstration inputs: directly using the kk nearest sentences; sampling kk sentences from the K=4,096K=4{,}096 nearest candidates; and using sentences physically adjacent in the corpus to the nearest sentences. On the tested datasets, the direct-nearest method performed worst across the evaluated language models, while diverse-nearest and physical-neighbor retrieval performed well, with physical neighbors reported as the best retrieval method. The authors interpret this pattern as evidence that retaining input relevance while increasing distance or diversity helps mitigate copying.

  8. Knowl 8 — Synonym labels balance label semantics against direct label copying

    empirical result

    The label ablation compared pseudo-demonstrations using the original test labels, random English words, or synonyms of the test labels. Random words consistently outperformed the original labels, showing that avoiding the exact test-label words matters. Synonyms consistently outperformed random words, indicating that preserving the labels’ semantic space also matters. These trends held across retrieval methods; the gap between original labels and synonyms was smaller with physical-neighbor retrieval than with direct-nearest retrieval, consistent with physical neighbors already reducing copying to some extent.

  9. Knowl 9 — Corpus size and domain coverage affect Z-ICL performance

    empirical result

    When the retrieval corpus was reduced from its full size (160 million paragraphs in the reported scaling experiment) to 0.03% of that size (48,000 paragraphs), Z-ICL performance declined. The decline was clearer on datasets whose domains were represented in the corpus, consistent with smaller corpora providing fewer relevant retrieval candidates.

    The authors also added unlabeled IMDB reviews to the corpus. This increased corpus size by only 2% but added coverage for MR, SST2, and SST5, which had previously been classified as uncovered. Z-ICL performance on those datasets improved across all evaluated language models. Together, these experiments show that domain coverage can matter independently of total corpus size.

  10. Knowl 10 — Evaluation is limited to single-sentence classification with manually selected synonyms

    limitation

    The experiments cover only single-sentence classification, because the retrieval procedure selects single sentences; extending the method to multi-sentence tasks would require constructing multi-sentence pseudo-demonstrations. They also cover only classification with a fixed set of candidate labels. Applying the approach to multiple-choice or generation tasks would require handling outputs that are not a shared fixed label set. In addition, the synonym labels used in the experiments were selected manually; the paper identifies better pseudo-demonstration construction as an open direction and only hypothesizes that it could further improve performance.

Coverage note — Omitted only non-contributory implementation details such as prompt templates and compute costs; the substantive method, findings, ablations, experimental scope, and stated limitations are represented.

References

  1. 1.Anonymous. 2023. Overthinking the truth: Understanding how language models process false demonstrations. In Submitted to The Eleventh International Conference on Learning Representations.
  2. 2.Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa-Anke, and Leonardo Neves. 2020. TweetEval:Unified Benchmark and Comparative Evaluation for Tweet Classification. In EMNLP.
  3. 3.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In NeurIPS.
  5. 5.Xiaowen Ding, Bing Liu, and Philip S Yu. 2008. A holistic lexicon-based approach to opinion mining. In Proceedings of the 2008 international conference on web search and data mining.
  6. 6.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In EMNLP.
  7. 7.Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021. Demix layers: Disentangling domains for modular language modeling. In NAACL.
  8. 8.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data.
  9. 9.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures.
  10. 10.Yanchen Liu, Timo Schick, and Hinrich Schütze. 2022. Semantic-oriented unlabeled priming for large-scale language models. arXiv preprint arXiv:2202.06133.
  11. 11.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies.
  12. 12.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Noisy channel language model prompting for few-shot text classification. In ACL.
  13. 13.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In EMNLP.
  14. 14.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022. In-context learning and induction heads. arXiv preprint arXiv:2209.11895.
  15. 15.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In ACL.
  16. 16.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS.
  17. 17.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  18. 18.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot reasoning. In EMNLP.
  19. 19.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems.
  20. 20.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2021. Learning to retrieve prompts for in-context learning. In NAACL.
  21. 21.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP.
  22. 22.Ben Wang and Aran Komatsuzaki. 2021. Gpt-j-6b: A 6 billion parameter autoregressive language model.
  23. 23.Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. 2020. Unsupervised data augmentation for consistency training. In NeurIPS.
  24. 24.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In ICLR.
  25. 25.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. In ICLR.
  26. 26.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NeurIPS.
  27. 27.Tony Z Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In ICML.

Citation

MLA
Lyu, X., et al. “Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 2304–17, https://doi.org/10.18653/v1/2023.acl-long.129.
APA
Lyu, X., Min, S., Beltagy, I., Zettlemoyer, L., & Hajishirzi, H. (2023). Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2304–2317. https://doi.org/10.18653/v1/2023.acl-long.129
Chicago
Lyu, X., S. Min, I. Beltagy, L. Zettlemoyer, and H. Hajishirzi. 2023. “Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2304–17. https://doi.org/10.18653/v1/2023.acl-long.129.
Harvard
Lyu, X. et al. (2023) “Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2304–2317. Available at: https://doi.org/10.18653/v1/2023.acl-long.129.
Vancouver
1. Lyu X, Min S, Beltagy I, Zettlemoyer L, Hajishirzi H (2023) Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2304–2317

BibTeX

@inproceedings{lyu-etal-2023-z,
    title = "{Z}-{ICL}: Zero-Shot In-Context Learning with Pseudo-Demonstrations",
    author = "Lyu, Xinxi  and
      Min, Sewon  and
      Beltagy, Iz  and
      Zettlemoyer, Luke  and
      Hajishirzi, Hannaneh",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.129/",
    doi = "10.18653/v1/2023.acl-long.129",
    pages = "2304--2317"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/