LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue WuYukang GanYixiao GeZeyu LuJiahao WangYe FengYing ShanPing Luo

article2024ACL129 citations

Proposes a block-expansion post-pretraining method that adds and tunes zero-initialized Transformer blocks on domain-specific data while freezing the base model, enabling large language models to master specialized coding and math skills without suffering from catastrophic forgetting of their general capabilities.

Listen

Adapting general-purpose large language models to specialized domains like mathematics and programming typically requires continued pretraining on domain-specific data. However, standard adaptation techniques suffer from catastrophic forgetting, where the model loses its initial general language and reasoning capabilities as it learns new material. Training specialized models entirely from scratch prevents this loss but requires massive computational budgets and vast datasets, limiting practical deployment and rapid iteration.

To solve this challenge, the article evaluates a post-pretraining technique called block expansion. The method inserts new structural layers into an existing foundation model and freezes the original layers, updating only the added components with domain-specific text to prevent the erosion of preexisting knowledge.

Evaluating this framework, the authors expanded the base LLaMA-2 7B model by adding eight interleaved blocks—growing it to an 8.3-billion parameter model called LLAMA PRO—and trained the new layers on 80 billion tokens of code and mathematics data using 2,830 GPU hours across 16 NVIDIA H800 units. The authors then applied standard instruction tuning to create LLAMA PRO - INSTRUCT and conducted comparative evaluations across standard natural language, mathematics, programming, conversational, and multi-turn tool-use benchmarks, alongside an ablation study on legal text.

The evaluation yielded several key findings:

  1. LLAMA PRO preserved general natural language performance while substantially improving specialized capabilities, raising the average benchmark score from 39.62 in base LLaMA-2 to 44.23, and more than doubling baseline coding scores (e.g., HumanEval rose from 13.05% to 28.66%).
  2. The instruction-tuned version, LLAMA PRO - INSTRUCT, achieved an overall average score of 53.85 across core benchmarks, outperforming comparable tuned models such as LLaMA-2-7B-Chat (40.04%) and CodeLLaMA-7B-Instruct (39.35%).
  3. In interactive agent settings (MINT-Bench), the model demonstrated superior tool augmentation and code execution abilities, achieving a 14.68% overall success rate across multi-turn tasks.
  4. Expanding the model by eight interleaved blocks delivered the optimal balance between computational efficiency and specialized domain accuracy, outperforming Low-Rank Adaptation (LoRA), Mixture-of-Experts (MoE), and top- or bottom-stacking layer configurations.

These results demonstrate that block expansion provides a cost-effective pathway to build versatile foundation models. Organizations can specialize existing open-source models for complex technical workflows without sacrificing conversational fluency or incurring the high expense of training domain models from scratch. Unlike standard parameter-efficient methods that struggle to absorb extensive new domain distributions or full fine-tuning that degrades general capabilities, block expansion balances both requirements.

Decision-makers should consider block expansion when adapting existing base models to technical domains that require strong reasoning and coding alongside general language proficiency. When implementing this approach, teams should use interleaved block insertion rather than purely top- or bottom-layer additions to prevent performance degradation, and follow post-pretraining with full instruction tuning.

The findings are constrained to English text and programming languages, and the expanded architecture introduces a modest increase in inference compute requirements due to the larger parameter count. Nevertheless, because the framework was validated across multiple foundation backbones (including Mistral-7B) and diverse domains (including legal corpora), there is high confidence in the reliability and generalizability of the block expansion method.

No sufficiently relevant recommendations were found.

Cover for LLaMA Pro: Progressive LLaMA with Block Expansion

Abstract

Humans generally acquire new skills without compromising the old; however, the opposite holds for Large Language Models (LLMs), e.g., from LLaMA to CodeLLaMA. To this end, we propose a new post-pretraining method for LLMs with an expansion of Transformer blocks. We tune the expanded blocks using only new corpus, efficiently and effectively improving the model’s knowledge while mitigating forgetting. In this paper, we experiment on the corpus of code and math, yielding LLaMA Pro-8.3B, a versatile foundation model initialized from LLaMA2-7B, excelling in general tasks, programming, and mathematics. LLaMA Pro and its instruction-following counterpart (LLaMA Pro - Instruct) achieve advanced performance among various benchmarks, demonstrating superiority over existing open models in the LLaMA family and the immense potential of reasoning and addressing diverse tasks as an intelligent agent. Our findings provide valuable insights into integrating natural and programming languages, laying a solid foundation for developing advanced language agents that operate effectively in various environments.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminaries: The LLaMA Block
  • 3.2 Block Expansion
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Pretrain Results
  • 4.3 SFT Results
  • 4.4 Mistral-Pro Results
  • 4.5 Ablation Study
  • 5 Scope and Limitations
  • 6 Conclusion
  • 7 Ethical Statement
  • References
  • A Gradient Derivation
  • B Dataset Details
  • C Mistral-Pro Details
  • D Evaluation Benchmark
  • E MINT-Bench
  • F MT-Bench
  • G Token Distribution
  • H Domain of Law

Knowls

  1. Knowl 1 — Block Expansion Post-Pretraining Framework

    model/method

    Block expansion is a post-pretraining method designed to inject domain-specific knowledge into a pre-trained Transformer language model without suffering from catastrophic forgetting of its original general capabilities.

    Given an initial model composed of LL sequential Transformer blocks (ϕ0,ϕ1,…,ϕL−1)(\phi_0, \phi_1, \dots, \phi_{L-1}), the architecture is expanded to L′L' blocks using an interleaved grouping scheme:

    1. The original LL blocks are partitioned into NN equal groups, where each group contains M=L/NM = L / N consecutive blocks.
    2. For each group, identity copies of the top PP blocks within that group are created and stacked directly on top of the group, increasing the depth of each group from MM to M+PM + P blocks and the total depth to L′=L+N×PL' = L + N \times P.
    3. The newly added N×PN \times P blocks are initialized to compute an identity mapping ϕid(x)=x\phi_{\text{id}}(x) = x, ensuring the expanded model initially produces outputs identical to the base pre-trained model.
    4. During domain-adaptive post-pretraining on a domain corpus, all original LL blocks are kept frozen, and only the newly added N×PN \times P blocks are tuned.
    5. In subsequent Supervised Fine-Tuning (SFT), all L′L' blocks in the expanded model are unfrozen and fine-tuned jointly using standard instruction tuning pipelines.
  2. Knowl 2 — Identity Initialization and Gradient Vanishing Prevention in SwiGLU Blocks

    theoretical result

    In a LLaMA Transformer block, given input x∈Rn×dx \in \mathbb{R}^{n \times d} (sequence length nn, hidden size dd), the forward pass consists of Multi-Head Self-Attention (MHSA) and a Swish-Gated Linear Unit (SwiGLU) Feed-Forward Network (FFN) with RMSNorm and residual connections without bias terms:

    x′=x+MHSA(RMSNorm(x))x' = x + \text{MHSA}(\text{RMSNorm}(x)) y=x′+FFN(RMSNorm(x′))y = x' + \text{FFN}(\text{RMSNorm}(x'))

    where MHSA(Q,K,V)=Concat(head1,…,headh)WO\text{MHSA}(Q, K, V) = \text{Concat}(\text{head}_1, \dots, \text{head}_h)W^O, FFN(t)=SwiGLU(t,W1,W2)W3\text{FFN}(t) = \text{SwiGLU}(t, W_1, W_2)W_3, and SwiGLU(t,W1,W2)=SiLU(tW1)⊗(tW2)\text{SwiGLU}(t, W_1, W_2) = \text{SiLU}(tW_1) \otimes (tW_2) with SiLU(u)=u⊗σ(u)\text{SiLU}(u) = u \otimes \sigma(u).

    If one attempts to create an identity block ϕid(x)=x\phi_{\text{id}}(x) = x by zero-initializing the scale/weight vector ww of the RMSNorm module (such that RMSNorm(x′)=t=0\text{RMSNorm}(x') = t = 0), the gradient of the loss L\mathcal{L} with respect to ww is:

    ∂L∂w=∂L∂y∂FFN(t)∂t∂RMSNorm(x′)∂w\frac{\partial \mathcal{L}}{\partial w} = \frac{\partial \mathcal{L}}{\partial y} \frac{\partial \text{FFN}(t)}{\partial t} \frac{\partial \text{RMSNorm}(x')}{\partial w}

    Because the derivative of the SwiGLU operation satisfies:

    ∂SwiGLU(t,W1,W2)∂t=(∂SiLU(tW1)∂t)⊗(tW2)+SiLU(tW1)⊗(∂(tW2)∂t)\frac{\partial \text{SwiGLU}(t, W_1, W_2)}{\partial t} = \left(\frac{\partial \text{SiLU}(tW_1)}{\partial t}\right) \otimes (tW_2) + \text{SiLU}(tW_1) \otimes \left(\frac{\partial (tW_2)}{\partial t}\right)

    evaluating this derivative at t=0t = 0 yields ∂FFN(t)∂t=0\frac{\partial \text{FFN}(t)}{\partial t} = 0, causing ∂L∂w=0\frac{\partial \mathcal{L}}{\partial w} = 0. Consequently, zero-initializing RMSNorm permanently blocks gradient propagation to the normalization layer.

    To establish an exact identity mapping while preserving gradient flow, the RMSNorm weights are left untouched and the linear projection matrices WOW^O (in MHSA) and W3W_3 (in SwiGLU FFN) in the newly added blocks are initialized to zero. Because there are no bias terms, this enforces MHSA(RMSNorm(x))=0\text{MHSA}(\text{RMSNorm}(x)) = 0 and FFN(RMSNorm(x′))=0\text{FFN}(\text{RMSNorm}(x')) = 0, reducing the block to y=xy = x via residual connections while allowing full gradient flow to all internal parameters during post-pretraining.

  3. Knowl 3 — Performance of LLaMA PRO Across General, Math, and Code Benchmarks

    data/table

    LLaMA PRO (8.3B parameters, expanded from LLaMA2-7B with 8 blocks and post-pretrained on 80B code and math tokens) and its supervised fine-tuned variant LLaMA PRO - INSTRUCT are evaluated against open-source foundation and chat models across general NLP tasks (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande), mathematical reasoning (GSM8K, GSM8K-PoT), and code generation (HumanEval pass@1, MBPP pass@1 with greedy decoding).

    Model ARC HellaSwag MMLU TruthfulQA Winogrande GSM8K GSM8K-PoT HumanEval MBPP Avg.
    Pretrained
    LLAMA PRO (8B) 54.10 77.94 47.88 39.04 73.95 17.89 25.42 28.66 33.20 44.23
    CrystalCoder (7B) 47.01 71.97 48.78 35.91 67.17 10.77 24.96 28.38 36.38 41.26
    LLaMA2-7B 53.07 78.59 46.87 38.76 74.03 14.48 17.68 13.05 20.09 39.62
    CodeLLaMA-7B 39.93 60.80 31.12 37.82 64.01 5.16 25.20 33.50 41.40 37.66
    StarCoder-15B 30.38 47.93 29.96 41.28 56.12 9.48 25.09 33.63 43.28 35.24
    LLaMA-7B 50.94 77.81 35.69 34.33 71.43 8.04 10.46 10.61 17.04 35.15
    OpenLLaMA-v2-7B 43.69 72.20 41.29 35.54 69.38 3.49 5.46 15.32 12.69 33.23
    Falcon-7B 47.87 78.13 27.79 34.26 72.38 4.62 4.32 9.42 13.39 32.46
    SFT
    LLAMA PRO - INSTRUCT 52.30 76.88 52.57 48.80 72.53 43.59 55.61 44.51 37.88 53.85
    LLaMA2-7B-Chat 52.90 78.55 48.32 45.57 71.74 7.35 19.73 14.63 21.60 40.04
    CodeLLaMA-7B-Instruct 36.52 55.44 34.54 41.25 64.56 7.96 34.67 34.80 44.40 39.35
    WizardCoder-Python-7B 41.81 65.06 32.29 36.32 61.72 4.70 17.60 42.07 47.20 38.75
    WizardMath-7B 54.10 79.55 45.97 43.65 72.69 2.73 25.57 12.20 18.00 39.38

    The results demonstrate that traditional domain post-pretraining (e.g., CodeLLaMA-7B) severely degrades general language reasoning (e.g., ARC drops to 39.93, MMLU drops to 31.12). In contrast, LLaMA PRO maintains or slightly improves general capabilities while doubling code performance compared to base LLaMA2-7B.

  4. Knowl 4 — Ablation on Expansion Capacity, Block Placement, and Adaptation Strategies

    data/table

    An ablation study evaluated on general language tasks (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande) and a domain-specific legal task (UNFAIR-ToS 4-shot classification from LexGLUE, using the FreeLaw subset of the Pile for post-pretraining) compares the number of added blocks, placement strategies, Mixture-of-Experts (MoE), full fine-tuning, and LoRA.

    Method ARC HellaSwag MMLU TruthfulQA Winogrande Lang. Avg. Unfair-ToS Overall Avg.
    Add 1 Block 52.30 77.92 38.62 37.80 73.16 55.96 67.45 61.71
    Add 2 Block 53.16 77.91 39.62 38.92 73.01 56.52 69.57 63.05
    Add 4 Block 52.39 76.92 37.30 40.53 72.22 55.87 71.31 63.59
    Add 8 Block 52.90 76.63 41.74 39.83 72.38 56.70 75.11 65.91
    Add 16 Block 51.88 76.59 41.35 40.13 71.82 56.35 75.17 65.76
    Add 32 Block 50.77 76.72 40.68 41.66 72.77 56.52 73.93 65.23
    MoE Expansion 51.45 76.51 42.47 40.13 72.23 56.56 67.27 61.92
    Fine-tuning 48.81 74.49 41.13 41.49 69.14 55.01 70.63 62.82
    LoRA (r=1024r=1024) 53.50 78.12 44.30 40.96 73.88 58.15 65.34 61.75
    Prefix Stacking (8 Blocks) 27.82 26.12 23.12 22.52 47.20 29.36 0.81 15.08
    Suffix Stacking (8 Blocks) 52.56 77.89 39.10 39.03 72.38 56.19 60.98 58.59

    Key empirical findings from the ablation:

    • Block Depth Scaling: Adding 8 blocks provides the optimal trade-off between computational cost and domain adaptation (75.11% on Unfair-ToS), matching 16 blocks (75.17%) while scaling past 16 blocks shows diminishing returns.
    • Block Placement: Stacking blocks at the bottom (Prefix Stacking) destroys model performance entirely (29.36% Lang. Avg, 0.81% Unfair-ToS) due to disrupting early representation layers. Stacking at the top (Suffix Stacking) preserves general language (56.19%) but underperforms interleaved expansion on domain adaptation (60.98% vs. 75.11%).
    • Adaptation Strategy: LoRA preserves general language ability well (58.15%) but falls short on domain learning (65.34%). Full fine-tuning degrades general capabilities (55.01%). Interleaved block expansion achieves strong domain transfer (75.11%) with high general retention (56.70%).
  5. Knowl 5 — LLaMA PRO-8.3B Architecture and Post-Pretraining Setup

    experimental setup

    LLaMA PRO-8.3B is constructed and trained using the following specifications:

    • Architecture: Initialized from LLaMA2-7B (L=32L = 32 Transformer blocks). It uses N=8N = 8 groups, M=4M = 4 base blocks per group, and P=1P = 1 expanded block per group, yielding L′=40L' = 40 blocks (8.3B parameters total, with ~1B newly added trainable parameters while the remaining ~7B are frozen).
    • Domain Pretraining Corpus: An ~80 billion token mix comprising:
      1. Proof-Pile-2 (55B tokens, weight 1.00): AlgebraicStack (11B), OpenWebMath (15B), ArXiv (29B).
      2. The-Stack-Dedup (Python split, 22B tokens, weight 1.50).
    • Post-Pretraining Hyperparameters: 15,900 training steps (~2,830 NVIDIA H800 GPU hours across 16 GPUs), global batch size of 1024, sequence length of 4096 tokens, learning rate of 2×10−42 \times 10^{-4} with a Cosine schedule and 6% warmup ratio, AdamW optimizer with weight decay 0.1, gradient clipping at 1.0, bfloat16 mixed precision, and FlashAttention.
    • Instruction Fine-Tuning (LLaMA PRO - INSTRUCT): Supervised fine-tuning of all 40 blocks on ~1M instances (~80M tokens) merged from ShareGPT (63,817 instances), WizardLM_evol_instruct_V2 (143,000 instances), SlimOrca (517,982 instances), MetaMath (395,000 instances), and Evol-CodeAlpaca (111,272 instances). SFT hyperparameters: batch size 128, sequence length 4096, warmup ratio 0.03, learning rate 2×10−52 \times 10^{-5} with Cosine decay, and bfloat16 mixed precision.
  6. Knowl 6 — Generalization of Block Expansion to Mistral-7B (Mistral-Pro)

    empirical result

    Block expansion was applied to Mistral-7B by adding 8 interleaved Transformer blocks to construct Mistral-Pro. Pretraining was conducted on code and math datasets (the automath subset of Cosmopedia, Proof-Pile-2, and the Python split of The Stack) using a learning rate of 2×10−52 \times 10^{-5} (1/101/10 of LLaMA PRO's learning rate).

    Model ARC HellaSwag MMLU TruthfulQA Winogrande GSM8K HumanEval
    Gemma-7B 61.9 82.2 64.6 44.8 79.0 50.9 32.3
    Mistral-7B 60.8 83.3 62.7 42.6 78.0 39.2 28.7
    Mistral-Pro 63.2 82.6 60.6 48.3 78.9 50.6 32.9

    When subsequently fine-tuned on the MetaMath dataset, MetaMath-Mistral-Pro achieved 78.4% Pass@1 on GSM8k and 30.3% Pass@1 on MATH, outperforming MetaMath-Mistral-7B (77.7% and 28.2%), MetaMath-Llemma-7B (69.2% and 30.0%), and MetaMath-13B (72.3% and 22.4%).

  7. Knowl 7 — Interactive Agent and Multi-Turn Tool Use Performance

    empirical result

    LLaMA PRO - INSTRUCT was evaluated on multi-turn tool interaction using MINT-Bench (586 instances across Code Generation, Decision Making, and Reasoning, allowing LLMs to generate Python code for tool execution up to k=5k=5 interaction turns) and multi-turn open-ended conversational ability using MT-Bench (evaluated via GPT-4 judging).

    Model Code Gen. Decision Making Reasoning MINT Micro Avg. (k=5k=5) MT-Bench Score
    AgentLM-7B 1.47 9.70 8.86 7.34 -
    CodeLLaMA-7B-Instruct 2.21 17.16 7.91 8.70 5.71
    LLaMA2-7B-Chat 0.00 0.00 13.61 7.34 6.27
    Vicuna-7B - - - - 6.17
    Alpaca-13B - - - - 4.53
    Mistral-Instruct-v0.1 6.62 34.33 8.54 13.99 -
    LLAMA PRO - INSTRUCT 11.76 29.10 9.81 14.68 6.32

    LLaMA PRO - INSTRUCT achieved a micro-average success rate of 14.68% on MINT-Bench at turn k=5k=5, outperforming similar-sized open models, and attained an MT-Bench score of 6.32, exceeding LLaMA2-7B-Chat (6.27), Vicuna-7B (6.17), and CodeLLaMA-7B-Instruct (5.71).

  8. Knowl 8 — Token Distribution Shift and Domain Perplexity Analysis

    empirical result

    To evaluate behavioral drift after block expansion post-pretraining, next-token prediction distributions are analyzed across user queries q=(q1,q2,… )q = (q_1, q_2, \dots) from the Alpaca dataset. Given context xt=q+(o1,…,ot−1)x_t = q + (o_1, \dots, o_{t-1}) and greedy output token oto_t predicted by LLaMA PRO (PalignP_{\text{align}}), the rank η\eta of oto_t under the base LLaMA2-7B model (PbaseP_{\text{base}}) is categorized as:

    1. Unshifted (η=1\eta = 1): oto_t is the top-1 ranked token in both PalignP_{\text{align}} and PbaseP_{\text{base}}.
    2. Marginally shifted (1<η≤31 < \eta \le 3): oto_t is the 2nd or 3rd highest probability token in PbaseP_{\text{base}}.
    3. Shifted (η>3\eta > 3): oto_t has a rank greater than 3 in PbaseP_{\text{base}}.

    Across general Alpaca queries, LLaMA PRO exhibits 92.6% unshifted tokens, 6.8% marginally shifted tokens, and only 0.7% significantly shifted tokens, confirming that base generation behavior is preserved.

    Furthermore, perplexity evaluations on general (LAMBADA) and code (The Stack smol-xs Python split) benchmarks demonstrate retention of general capabilities alongside domain acquisition:

    Model LAMBADA OpenAI LAMBADA Standard The Stack (Python)
    LLaMA2-7B 3.39 4.13 9.46
    LLaMA PRO 3.46 4.30 5.25
  9. Knowl 9 — Limitations and Scope of Block Expansion

    limitation

    The block expansion methodology presents two specific limitations:

    1. Modality and Language Scope: The evaluation and empirical verification are confined to text-based language modalities, specifically English natural language, mathematics, legal text, and Python source code. Its behavior on multilingual datasets or multimodal foundation models (e.g., vision-language architectures) remains unverified.
    2. Inference Latency Overhead: Although tuning only the newly added blocks significantly reduces training cost compared to full domain pretraining from scratch or standard continued pretraining, the resulting model has increased depth (L′>LL' > L, e.g., 40 blocks vs. 32 blocks), which incurs higher inference compute and memory requirements than the original base model.

Coverage note — Qualitative dialogue and code-completion case studies (Figures 8, 10, 11, and 12) were omitted as their quantitative conclusions are comprehensively captured in the benchmark tables.

References

  1. 1.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  2. 2.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2023. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631.
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966.
  4. 4.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021. Lexglue: A benchmark dataset for legal language understanding in english. arXiv preprint arXiv:2110.00976.
  7. 7.Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, and Qun Liu. 2021a. bert2bert: Towards reusable pretrained language models. arXiv preprint arXiv:2110.07143.
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021b. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  9. 9.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023a. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research.
  10. 10.Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. 2023b. Theoremqa: A theorem-driven question answering dataset. arXiv preprint arXiv:2305.12524.
  11. 11.Daixuan Cheng, Shaohan Huang, and Furu Wei. 2023. Adapting large language models via reading comprehension. arXiv preprint arXiv:2309.09530.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  13. 13.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  14. 14.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  15. 15.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2021. A continual learning survey: Defying forgetting in classification tasks. IEEE transactions on pattern analysis and machine intelligence, 44(7):3366–3385.
  16. 16.Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2023. How abilities in large language models are affected by supervised fine-tuning data composition. arXiv preprint arXiv:2310.05492.
  17. 17.William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research, 23(1):5232–5270.
  18. 18.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  19. 19.Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li, Xintao Wang, and Ying Shan. 2023. Making llama see and draw with seed tokenizer. arXiv preprint arXiv:2310.01218.
  20. 20.Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tieyan Liu. 2019. Efficient training of bert by progressively stacking. In International conference on machine learning, pages 2337–2346. PMLR.
  21. 21.Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han. 2020. On the transformer growth for progressive bert training. arXiv preprint arXiv:2010.12562.
  22. 22.Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort. 2023. Continual pre-training of large language models: How to (re) warm your model? arXiv preprint arXiv:2308.04014.
  23. 23.Suchin Gururangan, Ana Marasovic, Swabha ´ Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:2004.10964.
  24. 24.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  25. 25.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. NeurIPS.
  26. 26.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  28. 28.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  29. 29.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  30. 30.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  31. 31.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  32. 32.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023a. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161.
  33. 33.Xiang Li, Yiqun Yao, Xin Jiang, Xuezhi Fang, Xuying Meng, Siqi Fan, Peng Han, Jing Li, Li Du, Bowen Qin, et al. 2023b. Flm-101b: An open llm and how to train it with $100 k budget. arXiv preprint arXiv:2309.03852.
  34. 34.Wing Lian, Guan Wang, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". 2023. Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification.
  35. 35.Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023. The unlocking spell on base llms: Rethinking alignment via in-context learning. arXiv preprint arXiv:2312.01552.
  36. 36.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
  37. 37.Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Micklitz, Giovanni Sartor, and Paolo Torroni. 2019. Claudette: an automated detector of potentially unfair clauses in online terms of service. Artificial Intelligence and Law, 27:117–139.
  38. 38.Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550.
  39. 39.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct.
  40. 40.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  41. 41.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950.
  44. 44.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.
  45. 45.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  46. 46.Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan. 2022. Fine-tuned language models are continual learners. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6107–6122.
  47. 47.Sheng Shen, Pete Walsh, Kurt Keutzer, Jesse Dodge, Matthew Peters, and Iz Beltagy. 2022. Staged training for transformer language models. In International Conference on Machine Learning, pages 19893–19908. PMLR.
  48. 48.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768.
  49. 49.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  50. 50.Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295.
  51. 51.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. Bert rediscovers the classical nlp pipeline. arXiv preprint arXiv:1905.05950.
  52. 52.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  53. 53.Betty Van Aken, Benjamin Winter, Alexander Löser, and Felix A Gers. 2019. How does bert answer questions? a layer-wise analysis of transformer representations. In Proceedings of the 28th ACM international conference on information and knowledge management, pages 1823–1832.
  54. 54.Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard, Leonid Karlinsky, Rogerio Feris, David Daniel Cox, Zhangyang Wang, and Yoon Kim. 2023a. Learning to grow pretrained models for efficient transformer training. arXiv preprint arXiv:2303.00980.
  55. 55.Xiao Wang, Yuansen Zhang, Tianze Chen, Songyang Gao, Senjie Jin, Xianjun Yang, Zhiheng Xi, Rui Zheng, Yicheng Zou, Tao Gui, et al. 2023b. Trace: A comprehensive benchmark for continual learning in large language models. arXiv preprint arXiv:2310.06762.
  56. 56.Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023c. Mint: Evaluating llms in multi-turn interaction with tools and language feedback.
  57. 57.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023d. How far can camels go? exploring the state of instruction tuning on open resources.
  58. 58.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  59. 59.Chengyue Wu, Teng Wang, Yixiao Ge, Zeyu Lu, Ruisong Zhou, Ying Shan, and Ping Luo. 2023a. π-tuning: Transferring multimodal foundation models with optimal multi-task interpolation. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 37713–37727. PMLR.
  60. 60.Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023b. BloombergGPT: A large language model for finance. arXiv preprint arXiv:2303.17564.
  61. 61.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  62. 62.Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu, Weijia Shi, Binyuan Hui, Fan Zhou, Yitao Liu, Tianbao Xie, et al. 2023b. Lemur: Harmonizing natural language and code for language agents. arXiv preprint arXiv:2310.06830.
  63. 63.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600.
  64. 64.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  65. 65.Yiqun Yao, Zheng Zhang, Jing Li, and Yequan Wang. 2023. 2x faster language model pre-training via masked structural growth. arXiv preprint arXiv:2305.02869.
  66. 66.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284.
  67. 67.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  68. 68.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847.
  69. 69.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  70. 70.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  71. 71.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.

Citation

MLA
Wu, C., et al. “LLaMA Pro: Progressive LLaMA with Block Expansion”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6518–37, https://doi.org/10.18653/v1/2024.acl-long.352.
APA
Wu, C., Gan, Y., Ge, Y., Lu, Z., Wang, J., Feng, Y., Shan, Y., & Luo, P. (2024). LLaMA Pro: Progressive LLaMA with Block Expansion. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6518–6537. https://doi.org/10.18653/v1/2024.acl-long.352
Chicago
Wu, C., Y. Gan, Y. Ge, et al. 2024. “LLaMA Pro: Progressive LLaMA with Block Expansion”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6518–37. https://doi.org/10.18653/v1/2024.acl-long.352.
Harvard
Wu, C. et al. (2024) “LLaMA Pro: Progressive LLaMA with Block Expansion”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6518–6537. Available at: https://doi.org/10.18653/v1/2024.acl-long.352.
Vancouver
1. Wu C, Gan Y, Ge Y, Lu Z, Wang J, Feng Y, Shan Y, Luo P (2024) LLaMA Pro: Progressive LLaMA with Block Expansion. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6518–6537

BibTeX

@inproceedings{wu-etal-2024-llama,
    title = "{LL}a{MA} Pro: Progressive {LL}a{MA} with Block Expansion",
    author = "Wu, Chengyue  and
      Gan, Yukang  and
      Ge, Yixiao  and
      Lu, Zeyu  and
      Wang, Jiahao  and
      Feng, Ye  and
      Shan, Ying  and
      Luo, Ping",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.352/",
    doi = "10.18653/v1/2024.acl-long.352",
    pages = "6518--6537"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/