Instruction-tuned Language Models are Better Knowledge Learners

Zhengbao JiangZhiqing SunWeijia ShiPedro RodríguezChunting ZhouGraham NeubigXi Victoria LinWen-tau YihSrini Iyer

article2024ACL87 citations

Proposes pre-instruction-tuning to train language models on question-answer pairs before continued pre-training on new documents, significantly improving their ability to retain and accurately answer questions about newly acquired factual knowledge.

Listen

Updating the factual knowledge of large language models is essential for keeping AI assistants accurate as new information emerges or when deploying them in specialized domains. The conventional strategy updates models by continuing pre-training on new documents and then performing instruction-tuning using question-answer pairs. However, models trained under this standard pipeline struggle to recall and answer questions about the new documents, even after being trained to the point of near-perfect document memorization, a challenge known as the perplexity curse.

The article evaluates how effectively large language models absorb new factual knowledge through continued training and demonstrates a new training strategy called pre-instruction-tuning. The main goal is to test whether changing the sequence of training stages enables models to better encode and retrieve facts from complex, newly introduced documents.

To conduct this evaluation, researchers built Wiki2023, a dedicated dataset containing Wikipedia articles from 2023 across multiple domains along with associated question-answer pairs, ensuring minimal overlap with the models' original training data. The team performed systematic experiments using Llama-2 models (7-billion and 70-billion parameters) to evaluate factual recall across various training sequences, including standard training, data mixing, and pre-instruction-tuning variants, using exact match accuracy on factual questions as the primary benchmark.

The article reports several key findings. First, under the standard approach of training on documents followed by instruction-tuning, models achieved only 30.3% exact match accuracy on the 7-billion model and 46.4% on the 70-billion model. Second, pre-instruction-tuning—training on question-answer pairs before or alongside continued document training—substantially improved performance, reaching 48.1% on the 7-billion model (a 17.8 percentage point increase) and 62.7% on the 70-billion model (a 16.3 percentage point increase). Third, the best-performing variant, pre-instruction-tuning++, established that learning how knowledge is accessed through questions before learning to encode complex documents is the key driver of success. Finally, models trained with pre-instruction-tuning successfully generalized across different domains, non-Wikipedia texts, and questions posed by real search engine users.

These findings indicate that large language models absorb complex facts much more effectively when they are first taught the structure of how knowledge will be queried. Relying on the standard post-training instruction-tuning recipe creates severe performance bottlenecks and increases the risk of deploying under-informed models. In contrast, restructuring the training sequence provides a direct, cost-effective way to enhance parametric knowledge storage without increasing retrieval latency or runtime infrastructure costs.

Organizations seeking to continuously update language models should adopt pre-instruction-tuning workflows by prioritizing question-answer data ahead of raw document pre-training. When preparing updates, practitioners should first train models on query patterns and subsequently train on combined question and document corpora before final document ingestion. Further work should focus on developing automated question generation pipelines to scale this workflow across larger corporate and technical document repositories.

A primary limitation of the study is its primary reliance on Wikipedia articles and concise, fact-based questions, leaving uncertainties about how the approach generalizes to complex reasoning tasks, unstructured web scrapes, or dense scientific literature. Nevertheless, the consistent improvements across model sizes and distinct evaluation benchmarks provide high confidence in the fundamental efficacy of pre-instruction-tuning for factual knowledge absorption.

Jiang et al (2024).pdf
  • Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). This foundational study establishes instruction tuning as a way to teach models to answer new task formats, clarifying the training stage whose order the source reexamines.
Cover for Instruction-tuned Language Models are Better Knowledge Learners

Abstract

In order for large language model (LLM)-based assistants to effectively adapt to evolving information needs, it must be possible to update their factual knowledge through continued training on new data. The standard recipe for doing so involves continued pre-training on new documents followed by instruction-tuning on question-answer (QA) pairs. However, we find that LLMs trained with this recipe struggle to answer questions, even though the perplexity of documents is minimized. We found that QA pairs are generally straightforward, while documents are more complex, weaving many factual statements together in an intricate manner. Therefore, we hypothesize that it is beneficial to expose LLMs to QA pairs before continued pre-training on documents so that the process of encoding knowledge from complex documents takes into account how this knowledge is accessed through questions. Based on this, we propose pre-instruction-tuning (PIT), a method that instruction-tunes on questions prior to training on documents. This contrasts with standard instruction-tuning, which learns how to extract knowledge after training on documents. Extensive experiments and ablation studies demonstrate that PIT significantly enhances the ability of LLMs to absorb knowledge from new documents, outperforming standard instruction-tuning by 17.8%.

Table of Contents

  • Instruction-tuned Language Models are Better Knowledge Learners
  • Abstract

Knowls

  1. Knowl 1 — Pre-instruction-tuning++ orders question learning before document encoding

    model/method

    Pre-instruction-tuning++ (PIT++) is a continued-training strategy for helping a language model learn facts from documents so it can later answer questions about them without document context. On training examples, it first trains on question-answer (QA) pairs alone, then trains on QA pairs mixed with their associated documents. It subsequently trains on the held-out document set, with evaluation performed on questions about those documents. The first phase teaches how facts are accessed through questions; the mixed phase maintains that access pattern while the model learns to encode information from documents. In the Llama-2 7B Wiki2023-film experiment, PIT++ achieved 48.1% exact match (EM), compared with 45.4% for the paper's PIT variant and 30.3% for standard instruction-tuning. Training on QA only after the mixed QA-and-document phase did not provide the same gain.

  2. Knowl 2 — PIT jointly trains on training questions and documents before the target documents

    model/method

    Pre-instruction-tuning (PIT) exposes a model to question-based access to knowledge before it is trained on documents whose facts must later be retrieved from its parameters. In the paper's main PIT setting, the model first trains on training QA pairs and their associated training documents together, then continues training on separate target documents; evaluation uses questions about the target documents in a closed-book setting. The QA examples are therefore available before the target-document training phase, rather than being used only afterward as in standard instruction-tuning. This training order is motivated by the authors' hypothesis that learning how facts are queried helps the model encode facts embedded in complex, information-dense documents.

  3. Knowl 3 — PIT improves closed-book question answering over standard continued training

    empirical result

    On the Wiki2023-film test questions, the paper compares Llama-2 7B and 70B using EM, answer recall, and ROUGE-L, all reported as percentages. Each triplet below gives these metrics in that order.

    • Original closed-book models: 7B, 9.5 / 10.0 / 21.2; 70B, 17.2 / 18.1 / 31.4.
    • Open-book evaluation with the relevant document supplied: 7B, 72.2 / 75.4 / 91.5; 70B, 78.2 / 80.6 / 94.9.
    • Continued pre-training on target documents: 7B, 27.6 / 31.6 / 43.8; 70B, 41.7 / 45.8 / 60.2.
    • Continued pre-training followed by standard instruction-tuning: 7B, 30.3 / 34.7 / 47.4; 70B, 46.4 / 50.9 / 64.1.
    • Training on all QA pairs and documents mixed together: 7B, 39.4 / 44.6 / 56.7; 70B, 57.1 / 63.4 / 72.4.
    • PIT using QA pairs only before target-document training: 7B, 28.6 / 32.7 / 45.2; 70B, 49.7 / 53.7 / 67.9.
    • PIT using QA pairs and associated documents sequentially before target-document training: 7B, 32.5 / 37.2 / 49.0; 70B, 54.6 / 60.0 / 73.8.
    • Main PIT setting: 7B, 45.4 / 51.2 / 63.2; 70B, 62.7 / 68.6 / 78.8.

    The main PIT setting outperforms standard instruction-tuning on all three metrics for both model sizes. The large difference between closed-book and open-book scores also shows that answering from parameter-stored knowledge remains substantially harder than answering with the relevant document provided.

  4. Knowl 4 — The order of QA and document examples affects PIT performance

    empirical result

    Ablations with Llama-2 7B on Wiki2023-film show that performance depends on presenting QA examples before their associated documents. With three training epochs, grouping repeated QA examples before repeated documents achieved 38.2% EM, while grouping documents before QA examples achieved 27.2%. When examples were interleaved over three passes, putting each QA example before its associated document achieved 45.9% EM, compared with 43.2% when the document came first. The main PIT arrangement scored 45.4% EM. These comparisons support the paper's account that learning how knowledge is accessed through questions should precede or accompany learning to encode it from documents.

    Training duration also mattered: the one-epoch variant scored 33.3% EM, compared with 45.4% at three epochs, 45.8% at five epochs, and 46.5% at ten epochs. The gains from five to ten epochs were small, and ROUGE-L declined from 63.6% at five epochs to 61.9% at ten epochs.

  5. Knowl 5 — Wiki2023 provides document and QA splits for continual knowledge acquisition

    experimental setup

    The paper constructs Wiki2023 from Wikipedia articles classified under the 2023 category, using only each article's first section as the training or evaluation document. Publicly available language models generate questions and concise answers from each article; the generated questions are intended to cover the article's facts and include the article's subject. On average, 4.93 QA pairs are generated per article.

    The in-domain evaluation uses the film category: 1,720 training articles with 11,603 QA pairs, and 256 test articles with 1,743 QA pairs. Models train on the test documents and are evaluated on their corresponding test questions, which are not used as training QA. For cross-domain experiments, the other-domain training split contains 8,675 documents and 39,114 QA pairs; models train on these data and are evaluated on the film test split. The authors use low initial model accuracy on Wiki2023 questions as evidence that much of the evaluated information was not readily available in the original model, while acknowledging that complete factual separation from pretraining data is difficult to guarantee.

  6. Knowl 6 — Document perplexity can fall to one without strong closed-book recall

    empirical result

    In continued pre-training experiments with Llama-2 7B, exact-match accuracy on questions about the training documents generally rose as document-token perplexity fell, including as perplexity approached its minimum value of 1. This indicates that, in these experiments, learning the document facts for later question answering required extensive next-token loss minimization rather than stopping early. However, even after document perplexity was minimized, closed-book question-answering accuracy remained limited relative to the open-book setting; the authors call this gap the “perplexity curse.”

    The training-dynamics experiments also found a cost to this improvement: accuracy on Natural Questions, used as a measure of retention of previously acquired knowledge, declined during continued training. Among settings that reached minimum document perplexity, more training or a larger learning rate usually improved target-question performance when the learning rate remained reasonable; the tested 5e-5 rate was described as too large and prone to overfitting. The authors present reduced overfitting to misleading document patterns as a hypothesis for why more aggressive training can help after perplexity is minimized.

  7. Knowl 7 — Training uses all document tokens but only answer tokens in QA examples

    model/method

    Document training minimizes average next-token negative log likelihood over the document, while QA training minimizes it only over answer tokens conditioned on the question and preceding answer tokens:

    Ldoc=−1∣d∣∑t=1∣d∣log⁡P(dt∣d<t),LQA=−1∣a∣∑t=1∣a∣log⁡P(at∣q,a<t).L_{\mathrm{doc}}=-\frac{1}{|d|}\sum_{t=1}^{|d|}\log P(d_t\mid d_{<t}),\qquad L_{\mathrm{QA}}=-\frac{1}{|a|}\sum_{t=1}^{|a|}\log P(a_t\mid q,a_{<t}).

    Here, d=(d1,…,d∣d∣)d=(d_1,\ldots,d_{|d|}) is a document token sequence, qq is a question token sequence, a=(a1,…,a∣a∣)a=(a_1,\ldots,a_{|a|}) is its answer token sequence, and PP is the language model's conditional next-token distribution. A beginning-of-sequence token is prepended to documents. The reported setup uses batches of 256 documents with an initial learning rate of 3×10−53\times10^{-5} for document training, and batches of 256 QA pairs with an initial learning rate of 5×10−65\times10^{-6} for QA training. It uses AdamW with β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, weight decay 0.10.1, and a cosine learning-rate schedule decaying to 10% of the initial rate without warm-up. PIT uses three epochs unless an ablation specifies otherwise.

  8. Knowl 8 — PIT transfers to other domains and two additional question settings

    empirical result

    In cross-domain Wiki2023 experiments, models train on non-film domains and answer questions about film documents. Standard instruction-tuning versus PIT yields the following EM / answer recall / ROUGE-L percentages: for Llama-2 7B, in-domain 30.3 / 34.7 / 47.4 versus 45.4 / 51.2 / 63.2, and cross-domain 23.6 / 28.2 / 38.4 versus 36.9 / 43.2 / 54.9; for Llama-2 70B, in-domain 46.4 / 50.9 / 64.1 versus 62.7 / 68.6 / 78.8, and cross-domain 42.8 / 49.7 / 58.5 versus 55.2 / 66.7 / 74.0. PIT therefore retains an advantage over standard instruction-tuning in the cross-domain test, although both methods score lower cross-domain than in-domain.

    Two further Llama-2 7B evaluations also favored PIT. After training on non-Wikipedia Wiki2023 data and then on synthetic biographies, continued pre-training scored 29.6 / 29.8 / 38.7, while PIT scored 58.1 / 58.4 / 61.9; supplying the biography document directly scored 95.2 / 95.4 / 95.6, and the unadapted closed-book model scored 2.9 / 2.9 / 11.0. On 93 questions collected from Google's “People Also Ask” results, standard instruction-tuning scored 21.5 / 30.1 / 36.8 and PIT scored 29.0 / 35.5 / 48.2. Metrics in these triplets are EM, answer recall, and ROUGE-L, respectively, in percentages.

  9. Knowl 9 — Control experiments distinguish PIT from forgetting prevention and token reweighting

    empirical result

    The paper tests whether standard instruction-tuning performs worse merely because it forgets facts in target documents, and whether PIT's gain is equivalent to emphasizing salient document tokens. For Llama-2 7B on Wiki2023-film, standard instruction-tuning scored 30.3% EM; adding target documents to its instruction-tuning phase to reduce forgetting scored 30.2% EM. Instruction-tuning on training QA without first training on the training documents scored 27.1% EM. Thus, the tested forgetting-prevention modification did not improve on the standard sequence.

    In a weighted-document baseline, tokens appearing in answers received weight 1.0 and other document tokens received weight 0.5 during document training. This scored 27.7% EM, close to 27.6% for ordinary continued pre-training and well below PIT. These controls show that the reported PIT advantage was not reproduced by either of these tested alternatives.

  10. Knowl 10 — The evaluation corpus and task scope limit the demonstrated generality

    limitation

    Wiki2023 is built from Wikipedia, so the experiments do not establish that the same results hold for other document sources such as general web pages or scientific papers. The authors also note that factual overlap with Llama-2's original pretraining corpus cannot be completely excluded; for example, some information about a film released in 2023 may have been available earlier. Finally, the study focuses on eliciting factual knowledge through QA instruction-tuning and does not establish whether pre-instruction-tuning helps other capabilities, such as reasoning or comprehension, or works with other kinds of training data.

Coverage note — No substantial contributed material was deliberately omitted; background and related-work discussion are excluded because they are not contributions of this paper.

References

  1. 1.Uri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig. 2022. Neuro-symbolic language modeling with automaton-augmented retrieval. In International Conference on Machine Learning.
  2. 2.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. CoRR, abs/2310.11511.
  3. 3.Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2023. The reversal curse: Llms trained on "a is b" fail to learn "b is a". CoRR, abs/2309.12288.
  4. 4.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2022. Improving language models by retrieving from trillions of tokens. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 2206–2240. PMLR.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. CoRR, abs/2202.07646.
  7. 7.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1870–1879. Association for Computational Linguistics.
  8. 8.Daixuan Cheng, Shaohan Huang, and Furu Wei. 2023. Adapting large language models via reading comprehension. CoRR, abs/2309.09530.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  11. 11.Lucio M. Dery, Paul Michel, Ameet Talwalkar, and Graham Neubig. 2022. Should we be pre-training? an argument for end-task aware training as an alternative. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  12. 12.Gemini Team. 2023. Gemini: A family of highly capable multimodal models.
  13. 13.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  14. 14.Tianyu Han, Lisa C. Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K. Bressem. 2023. Medalpaca - an open-source collection of medical conversational AI models and training data. CoRR, abs/2304.08247.
  15. 15.Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick. 2021. Efficient nearest neighbor language models. In Conference on Empirical Methods in Natural Language Processing.
  16. 16.Nathan Hu, Eric Mitchell, Christopher D. Manning, and Chelsea Finn. 2023. Meta-learning online adaptation of language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 4418–4432. Association for Computational Linguistics.
  17. 17.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. Camels in a changing climate: Enhancing LM adaptation with tulu 2. CoRR, abs/2311.10702.
  18. 18.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, Xian Li, Brian O’Horo, Gabriel Pereyra, Jeff Wang, Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, and Ves Stoyanov. 2022. OPT-IML: scaling language model instruction meta learning through the lens of generalization. CoRR, abs/2212.12017.
  19. 19.Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. J. Mach. Learn. Res., 24:251:1–251:43.
  20. 20.Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2022. Towards continual knowledge learning of language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  21. 21.Robin Jia, Mike Lewis, and Luke Zettlemoyer. 2022. Question answering infused pre-training of general-purpose contextualized representations. In ACL (Findings), pages 711–728. Association for Computational Linguistics.
  22. 22.Zhengbao Jiang, Luyu Gao, Zhiruo Wang, Jun Araki, Haibo Ding, Jamie Callan, and Graham Neubig. 2022. Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 2336–2349. Association for Computational Linguistics.
  23. 23.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know. Trans. Assoc. Comput. Linguistics, 8:423–438.
  24. 24.Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 7969–7992. Association for Computational Linguistics.
  25. 25.Andreas Kopf, Yannic Kilcher, Dimitri von Rutte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Rich’ard Nagyfi, ES Shahul, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. ArXiv, abs/2304.07327.
  26. 26.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: a benchmark for question answering research. Trans. Assoc. Comput. Linguistics, 7:452–466.
  27. 27.Haejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape, Christopher D. Manning, and Kyoung-Gu Woo. 2022. You only need one model for open-domain question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 3047–3060. Association for Computational Linguistics.
  28. 28.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  29. 29.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020b. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  30. 30.Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Rich James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, and Scott Yih. 2023. RA-DIT: retrieval-augmented dual instruction tuning. CoRR, abs/2310.01352.
  31. 31.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  32. 32.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3470–3487. Association for Computational Linguistics.
  33. 33.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. Webgpt: Browser-assisted question-answering with human feedback. CoRR, abs/2112.09332.
  34. 34.Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciuca, Charlie O’Neill, Ze-Chang Sun, Maja Jablonska, Sandor Kruk, Ernest Perkowski, Jack W. Miller, Jason Li, Josh Peek, Kartheik Iyer, Tomasz Rózanski, Pranav Khetarpal, Sharaf Zaman, David Brodrick, Sergio J. Rodríguez Méndez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill P. Naiman, Jesse Cranney, Kevin Schawinski, and UniverseTBD. 2023. Astrollama: Towards specialized foundation models in astronomy. CoRR, abs/2309.06126.
  35. 35.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  36. 36.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. CoRR, abs/2203.02155.
  37. 37.Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. 2023. Fine-tuning or retrieval? comparing knowledge injection in llms. CoRR, abs/2312.05934.
  38. 38.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2463–2473. Association for Computational Linguistics.
  39. 39.Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2023. Webcpm: Interactive web search for chinese long-form question answering. CoRR, abs/2305.06849.
  40. 40.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8).
  41. 41.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  43. 43.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 5418–5426. Association for Computational Linguistics.
  44. 44.Devendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer, and Dani Yogatama. 2021. End-to-end training of multi-document reader and retriever for open-domain question answering. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 25968–25981.
  45. 45.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  46. 46.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. REPLUG: retrieval-augmented black-box language models. CoRR, abs/2301.12652.
  47. 47.Zhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou, Zhenfang Chen, David D. Cox, Yiming Yang, and Chuang Gan. 2023a. SALMON: self-alignment with principle-following reward models. CoRR, abs/2310.05910.
  48. 48.Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David D. Cox, Yiming Yang, and Chuang Gan. 2023b. Principle-driven self-alignment of language models from scratch with minimal human supervision. CoRR, abs/2305.03047.
  49. 49.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  50. 50.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn. 2023. Fine-tuning language models for factuality. CoRR, abs/2311.08401.
  51. 51.Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. Memorization without overfitting: Analyzing the training dynamics of large language models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  52. 52.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  53. 53.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  54. 54.Boxin Wang, Wei Ping, Peng Xu, Lawrence McAfee, Zihan Liu, Mohammad Shoeybi, Yi Dong, Oleksii Kuchaiev, Bo Li, Chaowei Xiao, Anima Anandkumar, and Bryan Catanzaro. 2023a. Shall we pretrain autoregressive language models with retrieval? A comprehensive study. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 7763–7786. Association for Computational Linguistics.
  55. 55.Cunxiang Wang, Pai Liu, and Yue Zhang. 2021. Can generative pre-trained language models serve as knowledge bases for closed-book qa? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3241–3251. Association for Computational Linguistics.
  56. 56.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023b. How far can camels go? exploring the state of instruction tuning on open resources. CoRR, abs/2306.04751.
  57. 57.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  58. 58.Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Pmc-llama: Towards building open-source language models for medicine.
  59. 59.Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Veselin Stoyanov. 2023. Training trajectories of language models across scales. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 13711–13738. Association for Computational Linguistics.
  60. 60.Ruohong Zhang, Luyu Gao, Chen Zheng, Zhen Fan, Guokun Lai, Zheng Zhang, Fangzhou Ai, Yiming Yang, and Hongxia Yang. 2023. A self-enhancement approach for domain-specific chatbot training via knowledge mining and digest. CoRR, abs/2311.10614.
  61. 61.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pre-trained transformer language models. ArXiv, abs/2205.01068.
  62. 62.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.
  63. 63.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: less is more for alignment. CoRR, abs/2305.11206.
  64. 64.Zeyuan Allen Zhu and Yuanzhi Li. 2023a. Physics of language models: Part 3.1, knowledge storage and extraction. CoRR, abs/2309.14316.
  65. 65.Zeyuan Allen Zhu and Yuanzhi Li. 2023b. Physics of language models: Part 3.2, knowledge manipulation. CoRR, abs/2309.14402.

Citation

MLA
Jiang, Z., et al. “Instruction-tuned Language Models Are Better Knowledge Learners”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5421–34, https://doi.org/10.18653/v1/2024.acl-long.296.
APA
Jiang, Z., Sun, Z., Shi, W., Rodriguez, P., Zhou, C., Neubig, G., Lin, X. V., Yih, W.-. tau ., & Iyer, S. (2024). Instruction-tuned Language Models are Better Knowledge Learners. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5421–5434. https://doi.org/10.18653/v1/2024.acl-long.296
Chicago
Jiang, Z., Z. Sun, W. Shi, et al. 2024. “Instruction-tuned Language Models Are Better Knowledge Learners”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5421–34. https://doi.org/10.18653/v1/2024.acl-long.296.
Harvard
Jiang, Z. et al. (2024) “Instruction-tuned Language Models are Better Knowledge Learners”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5421–5434. Available at: https://doi.org/10.18653/v1/2024.acl-long.296.
Vancouver
1. Jiang Z, Sun Z, Shi W, Rodriguez P, Zhou C, Neubig G, Lin XV, Yih W-tau, Iyer S (2024) Instruction-tuned Language Models are Better Knowledge Learners. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5421–5434

BibTeX

@inproceedings{jiang-etal-2024-instruction,
    title = "Instruction-tuned Language Models are Better Knowledge Learners",
    author = "Jiang, Zhengbao  and
      Sun, Zhiqing  and
      Shi, Weijia  and
      Rodriguez, Pedro  and
      Zhou, Chunting  and
      Neubig, Graham  and
      Lin, Xi Victoria  and
      Yih, Wen-tau  and
      Iyer, Srinivasan",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.296/",
    doi = "10.18653/v1/2024.acl-long.296",
    pages = "5421--5434"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/