On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model

Seongjin ShinSang-Woo LeeHwijeen AhnSungdong KimHyoungSeok KimBoseop KimKyunghyun ChoGichang LeeWoo-Myoung ParkJung-Woo Ha

article2022NAACL113 citations

Reveals that large language model in-context learning capabilities depend heavily on pretraining data sources and combinations rather than corpus size alone, showing that low validation perplexity does not reliably predict few-shot performance.

Listen

Large language models can perform downstream tasks through in-context learning—solving tasks using task prompts and a few reference examples without updating internal model weights. However, training these models requires massive computational and financial resources, and engineering teams currently lack clear insights into which pretraining text sources drive this emergent capability.

The article systematically analyzes how the source domain, volume, and combinations of pretraining text affect the emergence of zero- and few-shot in-context learning in large language models. It also investigates whether standard validation loss measurements reliably predict how well a model will perform on target tasks.

To conduct this evaluation, the researchers trained multiple 1.3-billion parameter variants of the Korean-centric HyperCLOVA model across seven distinct text domains (blogs, community forums, news, comments, question-and-answer boards, a curated government corpus, and encyclopedias) as well as combinations of these datasets, alongside baseline tests on a 6.9-billion parameter model. The models were trained under controlled token budgets (up to 150 billion tokens) and evaluated across four standardized tasks: sentiment classification, reading comprehension, machine translation, and topic classification.

The investigation produced four central findings. First, the pretraining source domain heavily governs in-context learning capability regardless of size; for example, a model trained on 150 billion blog tokens achieved competitive few-shot performance comparable to the full multi-domain model, whereas models trained on large volumes of news or forum text failed to develop few-shot capabilities. Second, combining ineffective single-domain datasets can trigger the emergence of strong in-context learning; mixing question-and-answer and encyclopedia texts produced strong reading comprehension and translation abilities that neither corpus could achieve alone. Third, domain relevance between pretraining text and target tasks improves zero-shot performance but fails to guarantee few-shot capability. Fourth, validation perplexity—a standard measure of how well a model predicts words—is not a reliable cross-model predictor of few-shot task accuracy, as models with low perplexity frequently failed downstream few-shot evaluations.

These findings challenge the common industry assumption that simply gathering more data or minimizing general validation loss will ensure high-performing few-shot models. For organizations building and deploying language models, strategic curation and diverse combination of text domains are far more critical than raw volume, offering substantial opportunities to lower compute costs, energy usage, and training timelines.

Practitioners should prioritize domain diversity and strategic data mixing rather than relying solely on single-domain scaling or validation loss metrics during pretraining. When preparing models for specialized applications, teams should pilot combinations of complementary text types and evaluate downstream task performance directly.

These conclusions are primarily bounded by evaluations conducted on a Korean-language architecture and models up to 6.9 billion parameters. While confidence in these empirical results is high, stakeholders should exercise caution before generalizing specific domain behaviors to other languages or massive models exceeding tens of billions of parameters without preliminary validation.

arXiv: 2204.13509
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Brown et al. establish GPT-3’s zero-, one-, and few-shot in-context learning paradigm, the foundation this study probes by varying pretraining corpora.
  • Paper: Pre-Training to Learn in Context, Yuxian Gu et al. (2023). PICL turns the source’s finding that pretraining data shapes in-context learning into a training method, using retrieved passages from unstructured text to cultivate the capability.
Cover for On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model

Abstract

Many recent studies on large-scale language models have reported successful in-context zero- and few-shot learning ability. However, the in-depth analysis of when in-context learning occurs is still lacking. For example, it is unknown how in-context learning performance changes as the training corpus varies. Here, we investigate the effects of the source and size of the pretraining corpus on in-context learning in HyperCLOVA, a Korean-centric GPT-3 model. From our in-depth investigation, we introduce the following observations: (1) in-context learning performance heavily depends on the corpus domain source, and the size of the pretraining corpus does not necessarily determine the emergence of in-context learning, (2) in-context learning ability can emerge when a language model is trained on a combination of multiple corpora, even when each corpus does not result in in-context learning on its own, (3) pretraining with a corpus related to a downstream task does not always guarantee the competitive in-context learning performance of the downstream task, especially in the few-shot setting, and (4) the relationship between language modeling (measured in perplexity) and in-context learning does not always correlate: e.g., low perplexity does not always imply high in-context few-shot learning performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 In-context Learning
  • 2.2 Domain Relevance on Pretraining Corpus
  • 2.3 Quantity and Quality of Pretraining Corpus
  • 2.4 Multi-task Learning
  • 3 Task Definition
  • 3.1 Model
  • 3.2 Pretraining with Different Corpus
  • 3.3 Downstream Tasks
  • 3.4 Experimental Details
  • 3.5 Measuring Validation Perplexity
  • 4 Experimental Results
  • 4.1 Main Results
  • 4.2 Effect of Corpus Source
  • 4.3 Effect of Corpus Size
  • 4.4 Effect of Combining Corpora
  • 4.5 Effect of Domain Relevance
  • 4.6 Perplexity and Downstream Task
  • 5 Discussion
  • 6 Conclusion
  • Broader Impact Statement
  • Acknowledgment
  • References
  • A Details on Experimental Results
  • B Details on Pretraining Corpus
  • B.1 Deduplication Preprocess
  • C Experiments on LoRA
  • D Examples of Few-shot Prompt
  • E Generalization to Other Languages

Knowls

  1. Knowl 1 — Pretraining source strongly changes few-shot performance

    empirical result

    The authors compared 1.3B-parameter HyperCLOVA models pretrained on individual Korean text domains and on the full corpus, then evaluated them by in-context few-shot learning. Performance varied substantially by source: Blog approached the full-corpus model on several tasks, while Cafe and News generally performed poorly despite having much more training data than Modu. KiN was a notable exception on Korean-to-English translation. Scores below are reported as NSMC accuracy, KorQuAD exact match (EM) and F1, AI Hub translation BLEU for Korean-to-English and English-to-Korean, and YNAT F1; corpus sizes are the training-token amounts used for these models.

    Pretraining corpusTrain tokensValidation PPLNSMC Acc.KorQuAD EMKorQuAD F1AI Hub Ko→En BLEUAI Hub En→Ko BLEUYNAT F1
    ALL150B119.9984.5956.1773.476.1523.3659.57
    Blog150B152.4083.5050.7469.343.8220.1160.68
    Cafe82.5B170.8557.773.1214.262.8316.5311.04
    News73.1B234.7850.720.149.961.1015.8814.36
    Comments40.7B225.3979.7814.6933.330.795.0636.17
    KiN27.0B187.8054.734.8518.996.8118.169.23
    Modu5.9B226.0169.9130.2049.291.216.1343.27
    Ency1.7B549.4053.810.7111.880.580.6927.99

    Blog is close to ALL on NSMC, KorQuAD, and YNAT, while Modu—trained on less than one tenth as many tokens as Cafe or News—outperforms both on most listed tasks. The results show that corpus size alone does not determine whether competitive in-context few-shot performance emerges.

  2. Knowl 2 — Combining individually weak corpora can produce few-shot ability

    empirical result

    The authors tested mixed-domain HyperCLOVA models using the same 1.3B-parameter setup and few-shot tasks. KiN+Ency and Cafe+KiN achieved useful performance on tasks where each component corpus, when used alone, performed poorly; combining corpora therefore could produce few-shot ability that was not evident in either individual model. Mixing corpora was not sufficient by itself, however: Cafe+News remained weak, and adding News to KiN+Ency reduced YNAT F1.

    Pretraining corpusTrain tokensValidation PPLNSMC Acc.KorQuAD EMKorQuAD F1AI Hub Ko→En BLEUAI Hub En→Ko BLEUYNAT F1
    ALL150B119.9984.5956.1773.476.1523.3659.57
    KiN+Ency28.7B164.6959.1742.0961.008.9923.1242.84
    Cafe+KiN109.5B141.9276.4238.4559.008.4123.4156.96
    Cafe+News150B154.2054.158.9522.724.4517.778.19
    Blog+Comments+Modu150B144.6782.8254.9472.274.0921.1765.01
    News+KiN+Ency101.8B142.1375.9635.4255.608.7023.3827.54

    Scores use the same metric ordering as follows: NSMC accuracy, KorQuAD EM and F1, AI Hub Ko→En and En→Ko BLEU, and YNAT F1. For example, KiN+Ency reaches 42.09 KorQuAD EM and 61.00 F1, compared with 4.85 EM and 18.99 F1 for KiN alone and 0.71 EM and 11.88 F1 for Ency alone. Conversely, News+KiN+Ency has lower YNAT F1 than KiN+Ency, despite the addition of a large News corpus.

  3. Knowl 3 — Reducing corpus size has a nonlinear effect on few-shot performance

    empirical result

    For the 1.3B-parameter ALL model, the authors compared random samples of the original pretraining corpus while keeping training at 72K steps. Smaller corpora were therefore revisited for more epochs: the 56B-token model trained for about 3 epochs and the 6B-token model for about 25. A tenfold reduction from 150B to 56B tokens caused modest changes on most metrics, whereas reducing to 6B caused much larger degradation.

    Pretraining tokensNSMC accuracyKorQuAD EMAI Hub Ko→En BLEUAI Hub En→Ko BLEUYNAT F1
    150B84.5956.176.1523.3659.57
    56B84.3555.135.4722.9851.89
    6B74.7036.723.9717.8130.24

    The 6B-token ALL model still substantially outperformed the 150B-token Cafe+News model on these few-shot tasks. Blog models showed a similar pattern: Blog and Blog 54B performed similarly, while Blog 27B degraded more noticeably. Thus, the effect of reducing data depends on both the amount and source of the corpus, rather than following a simple proportional relationship.

  4. Knowl 4 — Task-domain relevance behaves differently in few-shot and zero-shot evaluation

    empirical result

    A close topical or vocabulary relationship between pretraining material and an evaluation task did not guarantee strong few-shot performance. News is related to YNAT headline topic classification, yet the News-pretrained model scored 14.36 YNAT F1 in few-shot evaluation, compared with 59.57 for ALL. Likewise, Ency contains Korean Wikipedia text and KiN contains question-answer material, but their individual models performed poorly on KorQuAD: Ency scored 0.71 EM and 11.88 F1, and KiN scored 4.85 EM and 18.99 F1, versus ALL at 56.17 EM and 73.47 F1. High vocabulary overlap also did not reliably predict task performance; for example, Modu had substantial overlap with AI Hub vocabulary but weak translation scores.

    The pattern differed in zero-shot evaluation. News-related pretraining was associated with stronger YNAT scores: News scored 48.03 F1, Cafe+News 47.34, and News+KiN+Ency 51.89, compared with 42.79 for ALL. KiN-related mixtures also exceeded ALL on zero-shot AI Hub translation. The table reports zero-shot scores for the relevant task dimensions.

    Pretraining corpusYNAT F1AI Hub Ko→En BLEUAI Hub En→Ko BLEU
    ALL42.797.4324.81
    News48.031.2810.21
    Cafe+News47.343.4915.77
    News+KiN+Ency51.8910.1824.13
    KiN+Ency37.7111.5124.93
    Cafe+KiN45.4410.1224.95

    These observations indicate that domain relevance may help in some zero-shot settings, but it is not a general guarantee of competitive few-shot performance; the direction and size of the effect depend on the task and corpus combination.

  5. Knowl 5 — Validation perplexity is not a reliable cross-corpus predictor of few-shot ability

    empirical result

    The authors calculated validation perplexity on the same 70,000-example validation set, covering seven domains, with a shared tokenizer across models. Across models trained on different sources, lower perplexity did not consistently imply stronger in-context few-shot performance. Blog had the best validation perplexity among the single-domain models and strong few-shot scores, while Ency had the worst perplexity and generally poor scores; these matching extremes did not hold for intermediate models. Cafe and KiN had the second- and third-lowest single-domain perplexities, respectively, yet showed weak few-shot performance on several tasks.

    ModelValidation PPLNSMC accuracyKorQuAD EMKorQuAD F1
    Blog152.4083.5050.7469.34
    Cafe170.8557.773.1214.26
    KiN187.8054.734.8518.99
    Ency549.4053.810.7111.88
    ALL119.9984.5956.1773.47

    Perplexity therefore distinguishes some extremes but is not a strong predictor for comparing few-shot performance across different pretraining corpora. The authors also observed that corpus size could affect few-shot ability more strongly than perplexity: for example, Blog 27B performed substantially worse than Blog while its perplexity changed less markedly.

  6. Knowl 6 — Perplexity and few-shot performance tend to improve together during training one model

    empirical result

    When the authors tracked the 1.3B-parameter ALL model across pretraining steps, validation perplexity and in-context few-shot performance showed a clear co-movement across the five evaluated task settings. This within-model training trajectory contrasts with the weaker relationship observed when comparing models trained on different corpus sources. The paper therefore treats the relationship as conditional on the comparison: progress in training a fixed model may track few-shot improvement, but perplexity alone does not reliably rank models with different pretraining data.

  7. Knowl 7 — Experimental design controls tokenizer and training conditions across corpora

    experimental setup

    The study used Korean-centric HyperCLOVA autoregressive language models, primarily with 1.3B parameters and with an additional 6.9B comparison. All models shared a 2,048-token maximum sequence length and the same morpheme-aware byte-level BPE vocabulary. The seven analyzed corpus sources were Blog, Cafe, News, Comments, KiN, Modu, and Ency; the heterogeneous Others portion was excluded from source-specific analysis. For corpora below 150B tokens, 99% was assigned to training; larger sources were randomly sampled to 150B. Models were generally trained for 72K steps with global batch size 1,024, using AdamW at learning rate 2.0×10−42.0\times10^{-4} and cosine scheduling. Ency used its 12K-step minimum-validation-loss checkpoint because of overfitting; other reported models generally used the 72K-step checkpoint.

    Evaluation covered NSMC movie-review sentiment classification, KorQuAD machine reading comprehension, AI Hub Korean-English translation in both directions, and YNAT seven-class news-topic classification. Few-shot prompts used 70 examples for NSMC, 4 for KorQuAD, 4 for AI Hub, and 70 for YNAT; scores were averaged over 12, 1, 3, and 6 random trials, respectively. Classification used rank comparison of candidate labels, while KorQuAD and translation used greedy free-form generation. Perplexity was evaluated on a common 70,000-example validation set assembled from 10,000 examples for each of the seven domains.

  8. Knowl 8 — Fine-tuning reduced the large corpus-source differences seen in in-context learning

    empirical result

    As a comparison with in-context learning, the authors fine-tuned selected HyperCLOVA models using LoRA. On the two classification tasks reported, LoRA scores were much more tightly grouped across pretraining sources than the few-shot scores: NSMC accuracy ranged from 86.93 to 92.02, and YNAT F1 ranged from 82.37 to 87.52. In contrast, few-shot source-specific scores varied widely—for example, Cafe and News had very low YNAT F1, while Blog and ALL were near 60. This comparison suggests that the strong corpus-source sensitivity documented in the study is especially pronounced for in-context learning in this experimental setup.

    Pretraining modelLoRA NSMC accuracyLoRA YNAT F1
    ALL91.8386.47
    Blog91.9386.21
    Cafe91.5785.45
    News90.6286.57
    Comments92.0284.07
    KiN90.8984.46
    Modu90.6086.31
    Ency86.9382.37
    KiN+Ency90.9284.00
    Cafe+KiN90.9987.52
    Cafe+News91.3786.37
    Blog+Comments+Modu88.8387.07
    News+KiN+Ency91.1386.62
  9. Knowl 9 — The analyzed corpus sources span distinct Korean text domains and sizes

    definition

    The study partitions the HyperCLOVA pretraining corpus into seven domain sources, with token counts describing their corpus inventory: Blog contains blog posts (273.6B tokens); Cafe is online community content (83.3B); News is online news articles (73.8B); Comments contains crawled comment threads (41.1B); KiN is a Korean question-and-answer service (27.3B); Modu combines five public datasets covering news, written and spoken language, web text, and messenger text (6.0B); and Ency is encyclopedia text, including Korean Wikipedia (1.7B). The full inventory also contains 55.0B tokens of heterogeneous Others data, for a total of 561.8B tokens. Others was excluded from source-specific experiments; ALL denotes the original corpus including Others. The training amounts in individual model comparisons can be smaller than these inventory totals because of the study’s sampling and token limits.

  10. Knowl 10 — Generalization beyond Korean and larger models remains untested

    limitation

    The experiments analyze a Korean-centric HyperCLOVA corpus and do not establish that the observed source, mixture, and perplexity effects generalize to other languages. The study also does not validate its corpus findings at scales of tens of billions of parameters; its main experiments use 1.3B-parameter models, with a limited 6.9B comparison. The authors identify testing other languages and larger models as future work.

Coverage note — No substantial contributed material was omitted; supplementary prompt examples and per-trial standard deviations were left out because they document evaluation details rather than constitute separate findings.

References

  1. 1.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  2. 2.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In NeurIPS.
  3. 3.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, et al. 2021a. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  4. 4.Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, and He He. 2021b. Meta-learning via language model in-context tuning.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  7. 7.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In ACL.
  8. 8.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  9. 9.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  11. 11.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  12. 12.Boseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee, Donghyun Kwak, Dong Hyeon Jeon, Sunghyun Park, Sungju Kim, Seonhoon Kim, Dongpil Seo, et al. 2021. What changes can large-scale language models bring? intensive study on hyperclova: Billions-scale korean generative pretrained transformers. In EMNLP.
  13. 13.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  14. 14.Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019. Korquad1.0: Korean qa dataset for machine reading comprehension. arXiv preprint arXiv:1909.07005.
  15. 15.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668.
  16. 16.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR.
  17. 17.Vincent Micheli, Martin d’Hoffschmidt, and François Fleuret. 2020. On the importance of pre-training data volume for compact language models. In EMNLP.
  18. 18.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021. Metaicl: Learning to learn in context.
  19. 19.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  20. 20.Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Jiyoon Han, Jangwon Park, Chisung Song, Junseong Kim, Yongsook Song, Taehwan Oh, et al. 2021. Klue: Korean language understanding evaluation. arXiv preprint arXiv:2105.09680.
  21. 21.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  23. 23.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In EMNLP.
  24. 24.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization.
  25. 25.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners.
  26. 26.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In ICLR.
  27. 27.Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al. 2021. Pangu-α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369.
  28. 28.Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. ArXiv, abs/2102.09690.

Citation

MLA
Shin, S., et al. “On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5168–86, https://doi.org/10.18653/v1/2022.naacl-main.380.
APA
Shin, S., Lee, S.-W., Ahn, H., Kim, S., Kim, H., Kim, B., Cho, K., Lee, G., Park, W., Ha, J.-W., & Sung, N. (2022). On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5168–5186. https://doi.org/10.18653/v1/2022.naacl-main.380
Chicago
Shin, S., S.-W. Lee, H. Ahn, et al. 2022. “On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5168–86. https://doi.org/10.18653/v1/2022.naacl-main.380.
Harvard
Shin, S. et al. (2022) “On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5168–5186. Available at: https://doi.org/10.18653/v1/2022.naacl-main.380.
Vancouver
1. Shin S, Lee S-W, Ahn H, et al (2022) On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5168–5186

BibTeX

@inproceedings{shin-etal-2022-effect,
    title = "On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model",
    author = "Shin, Seongjin  and
      Lee, Sang-Woo  and
      Ahn, Hwijeen  and
      Kim, Sungdong  and
      Kim, HyoungSeok  and
      Kim, Boseop  and
      Cho, Kyunghyun  and
      Lee, Gichang  and
      Park, Woomyoung  and
      Ha, Jung-Woo  and
      Sung, Nako",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.380/",
    doi = "10.18653/v1/2022.naacl-main.380",
    pages = "5168--5186"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/