LLM Augmented LLMs: Expanding Capabilities through Composition

Rachit BansalBidisha SamantaSiddharth DalmiaNitish GuptaSriram GanapathyAbhishek BapnaPrateek JainPartha Talukdar

article2024ICLR51 citations

Introduces CALM, an efficient framework that combines frozen large language models with specialized smaller models via cross-attention to acquire new capabilities in coding and low-resource translation without retraining the base model.

Listen

Large language models excel at broad reasoning, world knowledge, and fluent communication, but expanding their capabilities to new specialized domains remains difficult and expensive. Retraining or fine-tuning massive models requires vast computing budgets, risks catastrophic forgetting of existing skills, and often faces data privacy constraints across organizational boundaries. Simpler alternatives like basic model merging or output routing fall short when neither model can solve a composite task independently.

The article demonstrates an efficient framework called Composition to Augment Language Models (CALM). The objective is to combine a large, general-purpose anchor model with a smaller, domain-specialized augmenting model to enable new composite capabilities while keeping both original models completely frozen.

The approach introduces a lightweight set of learnable projection and cross-attention parameters between selected intermediate layers of the two models. This structure allows the anchor model to dynamically attend to internal representations from the specialized model during text generation. To evaluate this method, the authors conducted experiments across three diverse domains: synthetic key-value arithmetic, low-resource language translation and mathematical reasoning, and software code generation and explanation. Composition training required only a small fraction—approximately 5% to 7%—of domain examples, avoiding the need for extensive retraining.

The findings show that model composition significantly enhances capabilities without degrading baseline performance. In synthetic tests requiring both key lookup and arithmetic, the combined system achieved an 84.3% success rate, whereas both standalone base models failed completely. For multilingual expansion across low-resource languages, composition improved translation to English across 175 of 192 languages and raised mathematical reasoning accuracy by up to 13% in absolute terms. In coding evaluations, the composed system delivered a relative improvement of about 40% over the base model on code generation benchmarks like HumanEval and MBPP, performing on par with fully fine-tuned alternatives.

These results demonstrate that organizations can systematically reuse existing specialized models rather than incurring the massive expense of training monolithic models from scratch. The method adds only about 1.5% in parameter overhead and operates at roughly 2% to 12% of the compute cost of full model retraining. Furthermore, it protects established capabilities, such as high-resource language understanding and natural language generation, which often degrade during standard fine-tuning.

Organizations seeking to extend language models to proprietary or highly specialized domains should pilot representation-level composition frameworks before investing in costly full-scale retraining. Teams should focus next on testing composition setups across multiple augmenting models simultaneously and measuring runtime inference latency in production environments.

Confidence in these findings is high for the tested domains of translation, arithmetic, and coding using standard transformer architectures. However, decision-makers should note that the approach assumes full access to the internal layer representations of both models, making it unsuitable for proprietary black-box systems accessed strictly through external text interfaces.

arXiv: 2401.02412
Cover for LLM Augmented LLMs: Expanding Capabilities through Composition

Abstract

Foundational models with billions of parameters which have been trained on large corpora of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to augment them or impart new skills. On the other hand, due to their adaptation abilities, several new instances of these models are being trained towards new domains and tasks. In this work, we study the problem of efficient and practical composition of existing foundation models with more specific models to enable newer capabilities. To this end, we propose CALM -- Composition to Augment Language Models -- which introduces cross-attention between models to compose their representations and enable new capabilities. Salient features of CALM are: (i) Scales up LLMs on new tasks by 're-using' existing LLMs along with a few additional parameters and data, (ii) Existing model weights are kept intact, and hence preserves existing capabilities, and (iii) Applies to diverse domains and settings. We illustrate that augmenting PaLM2-S with a smaller model trained on low-resource languages results in an absolute improvement of up to 13% on tasks like translation into English and arithmetic reasoning for low-resource languages. Similarly, when PaLM2-S is augmented with a code-specific model, we see a relative improvement of 40% over the base model for code generation and explanation tasks -- on-par with fully fine-tuned counterparts.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Composition to Augment Language Models (CALM)
  • 3.1 Learning to Compose (Θ𝐂\Theta_{\mathbf{C}})
  • 3.2 Composition Training Data (𝐃𝐂\mathbf{D}_{\mathbf{C}}^{\text{}})
  • 4 Experiments
  • 4.1 Key-value Arithmetic
  • 4.2 Low-resource Language Inclusivity
  • 4.3 Code Understanding and Generation
  • 4.4 Ablations
  • 5 Conclusion
  • References
  • A Supplementary Material for NTL
  • A.1 FLORES-200
  • A.2 GSM-8K
  • B Qualitative Analysis
  • C Overhead with CALM
  • C.1 Parametric Overhead
  • C.2 Training Overhead

Knowls

  1. Knowl 1 — Composition to Augment Language Models Architecture

    model/method

    Composition to Augment Language Models (CALM) composes a frozen anchor language model mBm_B with a frozen domain-specific augmenting model mAm_A via cross-attention over intermediate layer representations to introduce new capabilities without altering base model parameters.

    Let mAm_A have NAN_A layers and token representation dimension DAD_A, and let mBm_B have NBN_B layers and representation dimension DBD_B. A subset of nn compositional layers LA={l1A,…,lnA}\mathbb{L}_A = \{l_1^A, \dots, l_n^A\} and LB={l1B,…,lnB}\mathbb{L}_B = \{l_1^B, \dots, l_n^B\} are selected with uniform spacing NA/n=NB/nN_A/n = N_B/n.

    For each selected layer index i∈{1,…,n}i \in \{1, \dots, n\}, CALM introduces learnable composition parameters ΘC\Theta_C consisting of:

    1. A linear projection function fproj:RDA→RDBf_{\text{proj}}: \mathbb{R}^{D_A} \to \mathbb{R}^{D_B} projecting representation HAi∈RL×DAH_{Ai} \in \mathbb{R}^{L \times D_A} of mAm_A into the representation space of mBm_B: fproj(HAi)∈RL×DBf_{\text{proj}}(H_{Ai}) \in \mathbb{R}^{L \times D_B}

    2. A multi-head cross-attention layer fcrossf_{\text{cross}} with NHN_H attention heads. The anchor representation HBj∈RL×DBH_{Bj} \in \mathbb{R}^{L \times D_B} acts as the query, and the projected augmenting representation fproj(HAi)f_{\text{proj}}(H_{Ai}) acts as keys and values: QB,k=HBjWkQ,KA,k=fproj(HAi)WkK,VA,k=fproj(HAi)WkVQ_{B,k} = H_{Bj} W_k^Q, \quad K_{A,k} = f_{\text{proj}}(H_{Ai}) W_k^K, \quad V_{A,k} = f_{\text{proj}}(H_{Ai}) W_k^V headk=Attention(QB,k,KA,k,VA,k)\text{head}_k = \text{Attention}(Q_{B,k}, K_{A,k}, V_{A,k}) fcross(fproj(HAi),HBj)=Concat(head1,…,headNH)WOf_{\text{cross}}(f_{\text{proj}}(H_{Ai}), H_{Bj}) = \text{Concat}(\text{head}_1, \dots, \text{head}_{N_H}) W^O where WkQ,WkK,WkV∈RDB×(DB/NH)W_k^Q, W_k^K, W_k^V \in \mathbb{R}^{D_B \times (D_B / N_H)} and WO∈RDB×DBW^O \in \mathbb{R}^{D_B \times D_B} are learnable weights.

    The output of cross-attention is added as a residual connection to the anchor model representation before passing to the next layer (j+1)(j+1) in mBm_B: HA⊕B,j=HBj+fcross(fproj(HAi),HBj)H_{A \oplus B, j} = H_{Bj} + f_{\text{cross}}(f_{\text{proj}}(H_{Ai}), H_{Bj})

    During autoregressive generation at step tt, the predicted token yty_t is appended to the input sequence xt+1=xt⊕ytx_{t+1} = x_t \oplus y_t. This updated sequence is passed through both mAm_A and mBm_B, refreshing all internal intermediate representations in both models at each decoding step.

  2. Knowl 2 — Problem Formulation and Assumptions for Frozen Model Composition

    definition

    The CALM model composition framework operates under the following conditions and constraints:

    1. Model Access: Forward passes, backward passes, model weights, and intermediate layer representations are accessible for both an anchor model mBm_B and an augmenting model mAm_A.
    2. Weight Preservation: The weights of both mAm_A and mBm_B are kept frozen during composition training to preserve existing capabilities and eliminate catastrophic forgetting.
    3. Pre-training Data Agnosticism: The composition process does not require access to the original pre-training datasets, hyperparameters, or training states of either base model.
    4. Composition Data: The composition function mA⊕B=f(mA,mB,ΘC,DC)m_{A \oplus B} = f(m_A, m_B, \Theta_C, D_C) is trained on a small dataset DCD_C representing the joint task skills or a mix of source task domains (tA∪tBt_A \cup t_B). This allows learning joint cross-attention parameters ΘC\Theta_C using a fraction of the data and computational resources required for end-to-end retraining or full fine-tuning.
  3. Knowl 3 — Generalization in Synthetic Key-Value Arithmetic via Representation Composition

    empirical result

    To evaluate whether CALM can compose disjoint parametric skills from two separate models, a synthetic benchmark was created combining key-value (KV) memorization with arithmetic expression evaluation.

    • Knowledge Artifact (DKVD_{KV}): 25,000 key-value pairs mapping sampled 2-6 character English strings to unique integers in [1,25000][1, 25000].
    • KV-Substitution (DKV-SUBSD_{\text{KV-SUBS}}): Maps arithmetic expressions over keys (e.g., <K1> + <K2> - <K3>) to numerical expressions (e.g., 10 + 22 - 24).
    • Numeric-Arithmetic (DNUM-MATHD_{\text{NUM-MATH}}): Maps numerical expressions to their evaluated integer values (e.g., 10 + 22 - 24 to 8).
    • KV-Arithmetic (DKV-MATHD_{\text{KV-MATH}}): Directly maps key expressions to evaluated integers (e.g., <K1> + <K2> - <K3> to 8).

    An augmenting model mAm_A (PaLM2-XXS) was trained on DKV-SUBSD_{\text{KV-SUBS}} to memorize key mappings without learning arithmetic operations. An anchor model mBm_B (PaLM2-XS) possessed numeric arithmetic capabilities but had no exposure to DKVD_{KV}. CALM parameters ΘC\Theta_C were trained on composition examples from DKV-SUBSD_{\text{KV-SUBS}} spanning only 20% of the keys in DKVD_{KV}.

    Evaluation on held-out expressions containing keys unseen during composition training yielded:

    Dataset mAm_A mBm_B CALM (mA⊕Bm_{A \oplus B})
    DKV-SUBSD_{\text{KV-SUBS}} 98.1% 0.0% 92.9%
    DNUM-MATHD_{\text{NUM-MATH}} 4.2% 73.7% 72.0%
    DKV-MATHD_{\text{KV-MATH}} 0.7% 0.0% 84.3%

    CALM achieved 84.3% accuracy on DKV-MATHD_{\text{KV-MATH}}, demonstrating that representation composition generalizes arithmetic execution to 100% of the key dictionary despite being trained on arithmetic examples covering only 20% of keys.

  4. Knowl 4 — Low-Resource Language Translation Performance via CALM Composition

    data/table

    CALM enables an anchor LLM to translate low-resource languages by composing it with an augmenting model trained on low-resource corpora.

    • Augmenting model mAm_A: PaLM2-XXS trained on the Next Thousand Languages (NTL) dataset DNTLD_{\text{NTL}} (~1000 languages).
    • Anchor model mBm_B: Pre-trained PaLM2-S with no training on DNTLD_{\text{NTL}}.
    • Composition dataset DCD_C: ~5% of DNTLD_{\text{NTL}}.
    • Continued pre-training skyline mBNTLm_B^{\text{NTL}}: PaLM2-S fully continued pre-trained on DNTLD_{\text{NTL}}.

    Evaluations were conducted on the FLORES-200 benchmark in a 5-shot in-context learning setting for non-English to English translation, measured using chrF1 score.

    Model lij mr taq nn su ban pl th min acm avg.
    PaLM2-XXS 24.0 16.5 21.6 33.3 20.6 2.1 5.3 63.2 44.0 59.8 29.0
    mAm_A (+ NTL) 32.0 21.6 46.9 50.0 40.6 4.1 4.0 63.8 47.8 61.1 37.2
    mBm_B (PaLM2-S) 32.6 24.2 44.6 50.8 50.9 5.4 9.5 69.0 61.0 68.6 41.7
    CALM (mA⊕Bm_{A \oplus B}) 44.1 30.4 55.1 54.6 54.4 11.8 11.3 69.4 61.1 68.9 46.1
    mBNTLm_B^{\text{NTL}} 48.1 39.1 59.2 57.5 57.3 11.4 9.9 69.4 61.4 69.0 48.2

    The 10 low-resource languages displayed are: Ligurian (lij), Marathi (mr), Tamasheq (taq), Norwegian Nynorsk (nn), Sundanese (su), Balinese (ban), Polish (pl), Thai (th), Minangkabau (min), and Mesopotamian Arabic (acm). Across the complete set of 192 FLORES-200 languages evaluated, CALM (mA⊕Bm_{A \oplus B}) outperformed both underlying models mAm_A and mBm_B on 175 languages, approaching the performance of the fully pre-trained model mBNTLm_B^{\text{NTL}} at a fraction of the training cost.

  5. Knowl 5 — Multilingual Math Reasoning and Catastrophic Forgetting Prevention

    data/table

    When evaluated on grade-school math word problems (GSM8K) in low-resource languages (LRL) and high-resource languages (HRL), CALM transfers reasoning ability from the anchor LLM (mBm_B, PaLM2-S) to low-resource languages provided by the augmenting model (mAm_A, PaLM2-XXS trained on DNTLD_{\text{NTL}}), while preventing catastrophic forgetting of high-resource capabilities.

    GSM8K Low-resource Languages (Accuracy %)
    Model meo mfa pcm efi min ilo ady mai nso mzn avg.
    PaLM2-XXS 5.2 6.8 6.8 4.0 5.6 7.2 6.0 3.6 7.2 6.8 5.9
    mAm_A (+ NTL) 7.6 4.0 4.4 3.2 6.0 4.8 6.4 3.2 6.0 4.8 5.0
    mBm_B (PaLM2-S) 28.8 14.0 34.4 14.8 25.2 14.8 30.0 22.8 8.4 31.6 22.5
    CALM (mA⊕Bm_{A \oplus B}) 34.0 17.6 33.6 18.0 23.6 16.8 36.4 24.8 8.4 36.4 25.0
    mBNTLm_B^{\text{NTL}} 33.2 20.4 31.6 14.0 24.8 14.0 29.2 21.2 9.6 27.6 22.6
    GSM8K High-resource Languages (Accuracy %)
    Model en te bn sw ja zh th fr es de avg.
    PaLM2-XXS 5.6 4.0 2.0 7.6 2.0 4.4 6.0 6.8 5.6 9.2 5.3
    mAm_A (+ NTL) 4.8 3.6 3.2 4.8 3.2 7.6 6.4 9.2 5.6 7.2 5.6
    mBm_B (PaLM2-S) 36.8 19.2 23.2 16.0 2.0 39.2 29.6 38.0 32.4 43.2 28.0
    CALM (mA⊕Bm_{A \oplus B}) 37.2 28.0 27.2 18.0 2.4 43.6 33.2 42.8 36.0 49.2 31.8
    mBNTLm_B^{\text{NTL}} 36.0 17.6 18.4 14.4 0.8 33.6 27.2 34.8 31.2 42.0 25.6

    Key findings:

    1. Across 25 low-resource languages, CALM (mA⊕Bm_{A \oplus B}) outperformed the anchor model mBm_B on 20 of 25 languages (and on 18 of 25 compared to continued pre-training mBNTLm_B^{\text{NTL}}), lifting average accuracy across all 25 languages from 18.6% to 21.4%.
    2. On high-resource languages, continued pre-training on DNTLD_{\text{NTL}} degraded performance in mBNTLm_B^{\text{NTL}} from 28.0% to 25.6% due to catastrophic forgetting. CALM increased average accuracy to 31.8%, outperforming mBm_B on 9 of 10 high-resource languages.
  6. Knowl 6 — Code Generation and Understanding via LLM Composition

    data/table

    CALM was evaluated on code generation and code understanding tasks by composing an anchor LLM (mBm_B, PaLM2-S) with a code-specialized model (mAm_A, PaLM2-XXS trained on GitHub corpus DCodeD_{\text{Code}}) using only 7% of DCodeD_{\text{Code}} for composition training.

    Evaluations covered three tasks:

    1. Code Completion (CC): Zero-shot Pass@1 on HumanEval.
    2. Text-to-Code (T2C): 3-shot Pass@1 on MBPP.
    3. Code-to-Text (C2T): 3-shot chrF1 on CodeXGlue across 6 languages (Python, PHP, Go, Java, JavaScript, Ruby).
    Model CC (P@1) T2C (P@1) C2T (chrF1)
    HumanEval MBPP Python PHP Go Java JS Ruby
    mAm_A (PaLM2-XXS + Code) 19.5 28.0 28.0 34.7 32.6 29.6 26.5 26.0
    mBm_B (PaLM2-S) 16.4 28.6 30.4 35.5 40.4 31.0 28.8 27.9
    CALM (mA⊕Bm_{A \oplus B}) 22.5 32.2 30.5 35.8 40.6 31.4 29.3 29.0
    mBCodem_B^{\text{Code}} (Fine-tuned mBm_B) 24.3 43.0 18.9 35.0 41.1 31.1 20.2 27.6

    Key results:

    • CALM improved code generation over mBm_B by +6.1% absolute (+37.2% relative) on HumanEval and +3.6% absolute (+12.6% relative) on MBPP.
    • For Code-to-Text natural language explanation, directly fine-tuning mBm_B on code (mBCodem_B^{\text{Code}}) causes severe catastrophic forgetting of natural language generation capabilities (Python chrF1 drops from 30.4 to 18.9; JS drops from 28.8 to 20.2). CALM maintains text generation ability and slightly outperforms mBm_B across all six programming languages.
  7. Knowl 7 — Ablations on Model Specialization, Autoregressive Refresh, and LoRA Comparison

    empirical result

    Ablation experiments evaluated the relative contribution of augmenting model specialization, autoregressive representation refreshing, and parameter-matched Low-Rank Adaptation (LoRA) under identical composition parameter budgets and data DCD_C.

    Configurations tested:

    • Vanilla mAm_A: The specialized mAm_A is replaced with an unspecialized base PaLM2-XXS checkpoint.
    • Random mAm_A: mAm_A is replaced with a randomly initialized, untrained model.
    • mAm_A as Encoder: Autoregressive token appending is disabled for mAm_A, passing only initial prompt prefix representations without updating mAm_A on decoded tokens.
    • Parameter-Matched LoRA: LoRA layers trained on mBm_B over dataset DCD_C where the rank is adjusted such that trainable parameters equal the parameter budget of CALM.
    Task Metric mBNTL/Codem_B^{\text{NTL/Code}} CALM (mA⊕Bm_{A \oplus B}) Vanilla mAm_A Random mAm_A mAm_A as encoder LoRA
    FLORES-200 (XX-En) chrF1 62.1 60.5 59.2 58.8 59.3 59.2
    #(>mB> m_B) 171/192 175/192 115/192 43/192 102/192 82/192
    GSM-8K (LRL) Accuracy 19.8 21.4 19.0 17.8 19.1 20.9
    #(>mB> m_B) 15/25 20/25 15/25 9/25 12/25 15/25
    GSM-8K (HRL) Accuracy 27.1 33.1 29.7 28.5 29.1 31.2
    #(>mB> m_B) 1/11 11/11 8/11 4/11 6/11 9/11
    HumanEval Pass@1 24.3 22.5 20.0 20.1 16.0 18.3
    MBPP Pass@1 43.0 32.2 28.0 27.0 27.0 28.7
    CodeXGLUE chrF1 29.0 32.6 32.2 32.1 32.0 32.6

    Results demonstrate:

    1. Gains in CALM stem from the domain representations of mAm_A rather than the capacity of composition parameters ΘC\Theta_C, as shown by performance drops when mAm_A is replaced with vanilla or random models.
    2. Iterative token decoding through mAm_A is critical; disabling it (mAm_A as encoder) reduces HumanEval Pass@1 from 22.5% to 16.0%.
    3. CALM consistently outperforms parameter-matched LoRA on the same adaptation data across translation (175 vs 82 improved languages), reasoning, and code generation (22.5% vs 18.3% on HumanEval; 32.2% vs 28.7% on MBPP).
  8. Knowl 8 — Parametric and Computational Training Overhead of CALM

    theoretical result

    The parametric and computational cost of composing an anchor model mBm_B and augmenting model mAm_A via CALM scales as a small fraction of the anchor model's parameter count and training budget.

    Parametric Overhead: Let mAm_A and mBm_B have NAN_A and NBN_B transformer layers with hidden dimensions DAD_A and DBD_B, respectively. For n=nA=nBn = n_A = n_B selected composition layers:

    • Parameters per projection layer fprojf_{\text{proj}}: DA⋅DBD_A \cdot D_B
    • Parameters per cross-attention layer fcrossf_{\text{cross}}: 3DB23 D_B^2
    • Total added composition parameters ΘC\Theta_C: Params(ΘC)=n⋅(DA⋅DB+3DB2)\text{Params}(\Theta_C) = n \cdot \left(D_A \cdot D_B + 3 D_B^2\right)

    In comparison, an anchor model mBm_B with vocabulary size VBV_B and feed-forward hidden multiplier KBK_B has parameters: Params(mB)=NB⋅(VBDB+3DB2+2DB2KB)\text{Params}(m_B) = N_B \cdot \left(V_B D_B + 3 D_B^2 + 2 D_B^2 K_B\right)

    For standard transformer configurations (e.g., mAm_A with NA=4,DA=512N_A = 4, D_A = 512 and mBm_B with NB=24,DB=1024,VB=30K,KB=4N_B = 24, D_B = 1024, V_B = 30\text{K}, K_B = 4 choosing n=4n = 4): Params(ΘC)=4⋅(512⋅1024+3⋅10242)≈1.5×107≈15M\text{Params}(\Theta_C) = 4 \cdot (512 \cdot 1024 + 3 \cdot 1024^2) \approx 1.5 \times 10^7 \approx 15\text{M} Params(mB)≈1B\text{Params}(m_B) \approx 1\text{B} Params(ΘC)Params(mB)≈1.5%\frac{\text{Params}(\Theta_C)}{\text{Params}(m_B)} \approx 1.5\%

    Training Cost Overhead: Composition training requires only 5-7% of the data volume needed for full anchor fine-tuning. Assuming the parameter count of mAm_A is ≈10%\approx 10\% of mBm_B, and composition uses 2%2\% of the data needed for anchor training:

    • Compute cost to train ΘC\Theta_C: ≈0.02×X=2%\approx 0.02 \times X = 2\% of anchor model fine-tuning cost XX.
    • Total compute cost to train both mAm_A and ΘC\Theta_C: ≈(0.10⋅X+0.02⋅X)=12%\approx (0.10 \cdot X + 0.02 \cdot X) = 12\% of the compute cost of fine-tuning the anchor model on the full domain data.

Coverage note — No substantial contributed material was omitted; back-translation quality checks for synthetic dataset creation (Appendix A.2) and cherry-picked code generation qualitative examples (Appendix B) were condensed into the main empirical knowls.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a Visual Language Model for Few-Shot Learning, 2022. URL https://arxiv.org/abs/2204.14198.
  2. 2.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. ArXiv preprint, abs/2108.07732, 2021. URL https://arxiv.org/abs/2108.07732.
  3. 3.Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, Theresa Breiner, Vera Axelrod, Jason Riesa, Yuan Cao, Mia Xu Chen, Klaus Macherey, Maxim Krikun, Pidong Wang, Alexander Gutkin, Apurva Shah, Yanping Huang, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. Building machine translation systems for the next thousand languages, 2022.
  4. 4.Sebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with GPT-4. ArXiv preprint, abs/2303.12712, 2023. URL https://arxiv.org/abs/2303.12712.
  5. 5.Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Proceedings of the 28th International Conference on Computational Linguistics, pp. 6588–6608, Barcelona, Spain (Online), 2020. International Committee on Computational Linguistics. doi: 10.18653/v1/2020.coling-main.579. URL https://aclanthology.org/2020.coling-main.579.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. ArXiv preprint, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. ArXiv preprint, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  8. 8.Marta R. Costa-jussa, James Cross, Onur C¸ elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Lo¨ıc Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzman, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. No language left behind: Scaling human-centered machine translation. ArXiv preprint, abs/2207.04672, 2022. URL https://arxiv.org/abs/2207.04672.
  9. 9.Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe, Florian Metze, Luke Zettlemoyer, and Abdelrahman Mohamed. LegoNN: Building Modular Encoder-Decoder Models, 2022. URL https://arxiv.org/abs/2206.03318.
  10. 10.Xavier Garcia, Aditya Siddhant, Orhan Firat, and Ankur Parikh. Harnessing multilinguality in unsupervised machine translation for rare languages. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1126–1137, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.89. URL https://aclanthology.org/2021.naacl-main.89.
  11. 11.Google, Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Katlorahy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clement Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Dıaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. Palm 2 technical report, 2023.
  12. 12.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pp. 2790–2799. PMLR, 2019. URL http://proceedings.mlr.press/v97/houlsby19a.html.
  13. 13.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  14. 14.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. ArXiv preprint, abs/2212.04089, 2022. URL https://arxiv.org/abs/2212.04089.
  15. 15.Samuel Kessler, Bethan Thomas, and Salah Karout. An Adapter Based Pre-Training for Efficient and Scalable Self-Supervised Speech Representation Learning, 2021. URL https://arxiv.org/abs/2107.13530.
  16. 16.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models. In NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/18abbeef8cfe9203fdf9053c9c4fe191-Abstract-Conference.html.
  17. 17.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. Codexglue: A machine learning benchmark dataset for code understanding and generation. ArXiv preprint, abs/2102.04664, 2021. URL https://arxiv.org/abs/2102.04664.
  18. 18.Jiaqi Ma, Zhe Zhao, Jilin Chen, Ang Li, Lichan Hong, and Ed H. Chi. SNR: sub-network routing for flexible parameter sharing in multi-task learning. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pp. 216–223. AAAI Press, 2019. doi: 10.1609/aaai.v33i01.3301216. URL https://doi.org/10.1609/aaai.v33i01.3301216.
  19. 19.Michael S Matena and Colin A Raffel. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems, 35:17703–17716, 2022.
  20. 20.Mohammed Muqeeth, Haokun Liu, and Colin Raffel. Soft merging of experts with adaptive routing. ArXiv preprint, abs/2306.03745, 2023. URL https://arxiv.org/abs/2306.03745.
  21. 21.Jonas Pfeiffer, Aishwarya Kamath, Andreas Ruckl e, Kyunghyun Cho, and Iryna Gurevych. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 487–503, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.39. URL https://aclanthology.org/2021.eacl-main.39.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  23. 23.Timo Schick, Jane Dwivedi-Yu, Roberto Dess`ı, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. ArXiv preprint, abs/2302.04761, 2023. URL https://arxiv.org/abs/2302.04761.
  24. 24.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace, 2023. URL https://arxiv.org/abs/2303.17580.
  25. 25.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=fR3wGCk-IXp.
  26. 26.Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. Towards the next 1000 languages in multilingual machine translation: Exploring the synergy between supervised and self-supervised learning. ArXiv preprint, abs/2201.03110, 2022. URL https://arxiv.org/abs/2201.03110.
  27. 27.Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y Arcas, Nenad Tomasev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara  Mahdavi, Joelle K. Barral, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan. Towards expert-level medical question answering with large language models. ArXiv preprint, abs/2305.09617, 2023. URL https://arxiv.org/abs/2305.09617.
  28. 28.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
  29. 29.Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou. K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 1405–1418, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.121. URL https://aclanthology.org/2021.findings-acl.121.
  30. 30.Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.),  International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 23965–23998. PMLR, 2022. URL https://proceedings.mlr.press/v162/wortsman22a.html.
  31. 31.Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. Resolving interference when merging models. ArXiv preprint, abs/2306.01708, 2023. URL https://arxiv.org/abs/2306.01708.
  32. 32.Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language, 2022. URL https://arxiv.org/abs/2204.00598.

Citation

MLA
Bansal, R., et al. “LLM Augmented LLMs: Expanding Capabilities Through Composition”. arXiv, 2024, http://arxiv.org/abs/2401.02412v1.
APA
Bansal, R., Samanta, B., Dalmia, S., Gupta, N., Vashishth, S., Ganapathy, S., Bapna, A., Jain, P., & Talukdar, P. (2024). LLM Augmented LLMs: Expanding Capabilities through Composition. arXiv. http://arxiv.org/abs/2401.02412v1
Chicago
Bansal, R., B. Samanta, S. Dalmia, et al. 2024. “LLM Augmented LLMs: Expanding Capabilities Through Composition”. arXiv. http://arxiv.org/abs/2401.02412v1.
Harvard
Bansal, R. et al. (2024) “LLM Augmented LLMs: Expanding Capabilities through Composition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.02412v1.
Vancouver
1. Bansal R, Samanta B, Dalmia S, Gupta N, Vashishth S, Ganapathy S, Bapna A, Jain P, Talukdar P (2024) LLM Augmented LLMs: Expanding Capabilities through Composition. arXiv

BibTeX

@article{bansal2024llm,
  title = {LLM Augmented LLMs: Expanding Capabilities through Composition},
  author = {Bansal, Rachit and Samanta, Bidisha and Dalmia, Siddharth and Gupta, Nitish and Vashishth, Shikhar and Ganapathy, Sriram and Bapna, Abhishek and Jain, Prateek and Talukdar, Partha},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.02412v1},
  eprint = {2401.02412}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission