LLaMA Pro: Progressive LLaMA with Block Expansion
Chengyue WuYukang GanYixiao GeZeyu LuJiahao WangYe FengYing ShanPing Luo
Proposes a block-expansion post-pretraining method that adds and tunes zero-initialized Transformer blocks on domain-specific data while freezing the base model, enabling large language models to master specialized coding and math skills without suffering from catastrophic forgetting of their general capabilities.
Adapting general-purpose large language models to specialized domains like mathematics and programming typically requires continued pretraining on domain-specific data. However, standard adaptation techniques suffer from catastrophic forgetting, where the model loses its initial general language and reasoning capabilities as it learns new material. Training specialized models entirely from scratch prevents this loss but requires massive computational budgets and vast datasets, limiting practical deployment and rapid iteration.
To solve this challenge, the article evaluates a post-pretraining technique called block expansion. The method inserts new structural layers into an existing foundation model and freezes the original layers, updating only the added components with domain-specific text to prevent the erosion of preexisting knowledge.
Evaluating this framework, the authors expanded the base LLaMA-2 7B model by adding eight interleaved blocks—growing it to an 8.3-billion parameter model called LLAMA PRO—and trained the new layers on 80 billion tokens of code and mathematics data using 2,830 GPU hours across 16 NVIDIA H800 units. The authors then applied standard instruction tuning to create LLAMA PRO - INSTRUCT and conducted comparative evaluations across standard natural language, mathematics, programming, conversational, and multi-turn tool-use benchmarks, alongside an ablation study on legal text.
The evaluation yielded several key findings:
- LLAMA PRO preserved general natural language performance while substantially improving specialized capabilities, raising the average benchmark score from 39.62 in base LLaMA-2 to 44.23, and more than doubling baseline coding scores (e.g., HumanEval rose from 13.05% to 28.66%).
- The instruction-tuned version, LLAMA PRO - INSTRUCT, achieved an overall average score of 53.85 across core benchmarks, outperforming comparable tuned models such as LLaMA-2-7B-Chat (40.04%) and CodeLLaMA-7B-Instruct (39.35%).
- In interactive agent settings (MINT-Bench), the model demonstrated superior tool augmentation and code execution abilities, achieving a 14.68% overall success rate across multi-turn tasks.
- Expanding the model by eight interleaved blocks delivered the optimal balance between computational efficiency and specialized domain accuracy, outperforming Low-Rank Adaptation (LoRA), Mixture-of-Experts (MoE), and top- or bottom-stacking layer configurations.
These results demonstrate that block expansion provides a cost-effective pathway to build versatile foundation models. Organizations can specialize existing open-source models for complex technical workflows without sacrificing conversational fluency or incurring the high expense of training domain models from scratch. Unlike standard parameter-efficient methods that struggle to absorb extensive new domain distributions or full fine-tuning that degrades general capabilities, block expansion balances both requirements.
Decision-makers should consider block expansion when adapting existing base models to technical domains that require strong reasoning and coding alongside general language proficiency. When implementing this approach, teams should use interleaved block insertion rather than purely top- or bottom-layer additions to prevent performance degradation, and follow post-pretraining with full instruction tuning.
The findings are constrained to English text and programming languages, and the expanded architecture introduces a modest increase in inference compute requirements due to the larger parameter count. Nevertheless, because the framework was validated across multiple foundation backbones (including Mistral-7B) and diverse domains (including legal corpora), there is high confidence in the reliability and generalizability of the block expansion method.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). Because LLaMA Pro starts from LLaMA-family weights, this report establishes the base model and training context its expansion method depends on.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). Code Llama shows how code-specialized models were adapted from LLaMA 2, providing a direct point of comparison for LLaMA Pro’s code-focused expansion.
- Paper: Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora, Xisen Jin et al. (2022). Its account of continual pretraining, parameter expansion, and catastrophic forgetting clarifies the adaptation challenge that LLaMA Pro addresses.
- Paper: Continual Sequence Generation with Adaptive Compositional Modules, Yanzhe Zhang et al. (2022). Its compositional-module approach introduces the continual-learning trade-off between adding capacity for new tasks and preserving prior skills that motivates block expansion.
No sufficiently relevant recommendations were found.
