Instruction-tuned Language Models are Better Knowledge Learners
Zhengbao JiangZhiqing SunWeijia ShiPedro RodríguezChunting ZhouGraham NeubigXi Victoria LinWen-tau YihSrini Iyer
Proposes pre-instruction-tuning to train language models on question-answer pairs before continued pre-training on new documents, significantly improving their ability to retain and accurately answer questions about newly acquired factual knowledge.
Updating the factual knowledge of large language models is essential for keeping AI assistants accurate as new information emerges or when deploying them in specialized domains. The conventional strategy updates models by continuing pre-training on new documents and then performing instruction-tuning using question-answer pairs. However, models trained under this standard pipeline struggle to recall and answer questions about the new documents, even after being trained to the point of near-perfect document memorization, a challenge known as the perplexity curse.
The article evaluates how effectively large language models absorb new factual knowledge through continued training and demonstrates a new training strategy called pre-instruction-tuning. The main goal is to test whether changing the sequence of training stages enables models to better encode and retrieve facts from complex, newly introduced documents.
To conduct this evaluation, researchers built Wiki2023, a dedicated dataset containing Wikipedia articles from 2023 across multiple domains along with associated question-answer pairs, ensuring minimal overlap with the models' original training data. The team performed systematic experiments using Llama-2 models (7-billion and 70-billion parameters) to evaluate factual recall across various training sequences, including standard training, data mixing, and pre-instruction-tuning variants, using exact match accuracy on factual questions as the primary benchmark.
The article reports several key findings. First, under the standard approach of training on documents followed by instruction-tuning, models achieved only 30.3% exact match accuracy on the 7-billion model and 46.4% on the 70-billion model. Second, pre-instruction-tuning—training on question-answer pairs before or alongside continued document training—substantially improved performance, reaching 48.1% on the 7-billion model (a 17.8 percentage point increase) and 62.7% on the 70-billion model (a 16.3 percentage point increase). Third, the best-performing variant, pre-instruction-tuning++, established that learning how knowledge is accessed through questions before learning to encode complex documents is the key driver of success. Finally, models trained with pre-instruction-tuning successfully generalized across different domains, non-Wikipedia texts, and questions posed by real search engine users.
These findings indicate that large language models absorb complex facts much more effectively when they are first taught the structure of how knowledge will be queried. Relying on the standard post-training instruction-tuning recipe creates severe performance bottlenecks and increases the risk of deploying under-informed models. In contrast, restructuring the training sequence provides a direct, cost-effective way to enhance parametric knowledge storage without increasing retrieval latency or runtime infrastructure costs.
Organizations seeking to continuously update language models should adopt pre-instruction-tuning workflows by prioritizing question-answer data ahead of raw document pre-training. When preparing updates, practitioners should first train models on query patterns and subsequently train on combined question and document corpora before final document ingestion. Further work should focus on developing automated question generation pipelines to scale this workflow across larger corporate and technical document repositories.
A primary limitation of the study is its primary reliance on Wikipedia articles and concise, fact-based questions, leaving uncertainties about how the approach generalizes to complex reasoning tasks, unstructured web scrapes, or dense scientific literature. Nevertheless, the consistent improvements across model sizes and distinct evaluation benchmarks provide high confidence in the fundamental efficacy of pre-instruction-tuning for factual knowledge absorption.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). This foundational study establishes instruction tuning as a way to teach models to answer new task formats, clarifying the training stage whose order the source reexamines.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). Building on the source’s finding that training order shapes new-fact learning, this work tests how multiple continual-learning mechanisms can preserve knowledge across long sequences of updates.
