Towards Modular LLMs by Building and Reusing a Library of LoRAs

Oleksiy OstapenkoZhan SuEdoardo M. PontiLaurent CharlinNicolas Le RouxLucas CacciaAlessandro Sordoni

article2024ICML90 citations

Proposes a modular framework that clusters LoRA adapters by parameter similarity and dynamically routes hidden states to the most relevant adapters at inference time, achieving superior zero-shot and supervised generalization without requiring joint retraining.

Listen

Adapting large language models to new tasks typically requires extensive computational resources and centralized datasets. While parameter-efficient adapters, such as Low-Rank Adaptation (LoRA), allow models to be fine-tuned cheaply on specific tasks, organizations face significant challenges when trying to combine independently developed adapters to solve new, unseen problems without retraining entire models.

The article investigates how to efficiently build and reuse libraries of modular adapters. Specifically, it demonstrates how to group training tasks based on model parameters to maximize knowledge transfer and introduces an automated routing system to direct new inputs to the most relevant adapters without requiring access to original training data.

To construct the adapter library, the authors developed Model-Based Clustering (MBC), which measures the mathematical similarity between independently trained adapter weights and clusters related tasks together before training a single consolidated adapter per cluster. To reuse these libraries on new tasks without further training, the authors designed Arrow, a routing mechanism that identifies the primary direction of variation within each adapter's weights and dynamically matches token representations to the most suitable adapters at inference time. The methods were evaluated on multi-task benchmarks using 256 training tasks across models including Phi-2 (2.8 billion parameters) and Mistral (7 billion parameters), testing on multiple held-out reasoning, coding, and question-answering tasks.

The findings show that weight similarity between adapters serves as an effective proxy for task compatibility and positive transfer. Building libraries with Model-Based Clustering outperformed standard multi-task baselines, achieving a 67.4% zero-shot accuracy on Phi-2 held-out benchmarks compared to 65.6% for full model fine-tuning and 63.8% for the base model. Furthermore, when routing large collections of independently trained private adapters, Arrow outperformed uniform adapter averaging by 1.8 percentage points on Phi-2 and 2.5 percentage points on Mistral, matching or exceeding the performance of computationally intensive joint training. In supervised adaptation settings with limited data, initializing from clustered adapter libraries consistently accelerated learning and improved task scores over non-modular baselines.

These results indicate that organizations can achieve state-of-the-art model adaptation in a decentralized manner, avoiding the privacy risks, infrastructure requirements, and energy costs of continuous centralized multi-task training. Practitioners looking to deploy modular language models should adopt weight-based clustering when shared data is accessible, or use Arrow-style prototype routing when integrating decentralized libraries of adapters. However, because the study focused exclusively on linear LoRA adapters and tested models up to 7 billion parameters, teams should conduct pilot evaluations before applying these routing strategies to larger models or non-linear adapter architectures.

arXiv: 2405.11157
Cover for Towards Modular LLMs by Building and Reusing a Library of LoRAs

Abstract

Given the increasing number of parameter-efficient adapters of large language models (LLMs), how can we reuse them to improve LLM performance on new tasks? We study how to best build a library of adapters given multi-task data and devise techniques for both zero-shot and supervised task generalization through routing in such library. We benchmark existing approaches to build this library and introduce model-based clustering, MBC, a method that groups tasks based on the similarity of their adapter parameters, indirectly optimizing for transfer across tasks. In order to reuse the library, we present a novel zero-shot routing mechanism, Arrow, which enables dynamic selection of the most relevant adapters for new inputs without the need for retraining. We experiment with several LLMs, such as Phi-2 and Mistral, on a wide array of held-out tasks, verifying that MBC-based adapters and Arrow routing lead to superior generalization to new tasks. Thus, we make steps towards creating modular, adaptable LLMs that can match or outperform traditional joint training.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Building the LoRA Library
  • 4. Reusing the LoRA Library
  • 4.1. Routing
  • 4.1.1. ZERO-SHOT ROUTING
  • 4.1.2. SUPERVISED TASK ROUTING
  • 4.2. LoRA Composition
  • 5. Experiments
  • 5.1. Zero-Shot Results
  • 5.2. Upstream Performance
  • 5.3. Supervised Adaptation
  • 5.4. Summary of Results
  • 6. Related Work
  • 7. Conclusions and Future Work
  • Impact Statement
  • Acknowledgement
  • References
  • 8. Appendix
  • 8.1. Analyzing ∥ AB T v ∥ 2 for in-distribution and out-of-distribution samples
  • 8.2. Few-shot adaptation
  • 9. Implementation details and hyperparameters

Knowls

  1. Knowl 1 — Arrow Zero-Shot Adapter Routing

    algorithm

    Arrow (↗\nearrow) is a zero-shot, data-free routing mechanism designed to dynamically route token hidden representations to pre-trained Low-Rank Adaptation (LoRA) adapters at inference time without requiring joint training, routing re-training, or access to the training data of the adapters.

    For a transformer layer ℓ\ell and an adapter library L\mathcal{L} of low-rank adapter pairs {(Ai,Bi)}i=1∣L∣\{(A_i, B_i)\}_{i=1}^{|\mathcal{L}|} where Ai,Bi∈Rd×rA_i, B_i \in \mathbb{R}^{d \times r} modify base model activations via hi=AiBi⊤hh_i = A_i B_i^\top h, Arrow constructs a routing prototype vector Wℓ[i]∈RdW_\ell[i] \in \mathbb{R}^d for each expert ii. The prototype corresponds to the first right singular vector of the outer product matrix AiBi⊤A_i B_i^\top, identifying the direction in activation space that produces the maximum variance under adapter ii. When r=1r=1, this simplifies to the normalized vector Bi/∥Bi∥2B_i / \|B_i\|_2.

    Weight Initialization:
    Input: LoRA adapter library L\mathcal{L}, layer index ℓ\ell
    Output: Routing parameter matrix for layer ℓ\ell: Wℓ∈R∣L∣×dW_\ell \in \mathbb{R}^{|\mathcal{L}| \times d}
    for i=1i = 1 to ∣L∣|\mathcal{L}| do
      Ai,Bi=L[i,ℓ]A_i, B_i = \mathcal{L}[i, \ell]
      U,D,V=SVD(AiBi⊤)U, D, V = \text{SVD}(A_i B_i^\top)
      Wℓ[i]=V[:,0]W_\ell[i] = V[:, 0]
    end for
    return WℓW_\ell
    Routing Computation:
    Input: Routing parameters Wℓ∈R∣L∣×dW_\ell \in \mathbb{R}^{|\mathcal{L}| \times d}, token hidden state in layer ℓ\ell: hℓ∈Rdh_\ell \in \mathbb{R}^d, top-k parameter kk
    Output: Routing probability vector for layer ℓ\ell: pℓ∈R∣L∣p_\ell \in \mathbb{R}^{|\mathcal{L}|}
    logits=∣Wℓhℓ∣\text{logits} = |W_\ell h_\ell|
    for i=1i = 1 to ∣L∣|\mathcal{L}| do
      if i∈arg⁡top-k(logits)i \in \arg\text{top-}k(\text{logits}) then
        pℓ[i]=logits[i]p_\ell[i] = \text{logits}[i]
      else
        pℓ[i]=−∞p_\ell[i] = -\infty
      end if
    end for
    return softmax(pℓ)\text{softmax}(p_\ell)

    Routing decisions are computed per token and per layer. Taking the absolute value ∣Wℓhℓ∣|W_\ell h_\ell| accounts for sign ambiguity in the singular vector decomposition.

  2. Knowl 2 — Model-Based Clustering for Adapter Library Construction

    algorithm

    Model-Based Clustering (MBC) is a two-stage method for constructing a compact and transferable library of KK adapters from a multi-task dataset D={D1,…,DT}\mathcal{D} = \{D_1, \dots, D_T\} across TT tasks (K<TK < T), while maintaining the same training compute budget as individual task training.

    In the first stage, independent private LoRA adapters are partially trained on each task dataset DtD_t for nn steps out of a total computational budget of NN steps per task. The parameter vectors of these adapters are flattened, concatenated, and dimensionally reduced via Singular Value Decomposition (SVD). Pairwise cosine similarities between task representations are computed, and tasks are partitioned into KK disjoint clusters c1,…,cKc_1, \dots, c_K using kk-means clustering on the similarity matrix. In the second stage, for each cluster kk, the constituent task datasets are pooled (Dk=⋃t∈ckDtD_k = \bigcup_{t \in c_k} D_t) and a single cluster adapter is trained for the remaining N−nN - n steps.

    Input: Multi-task data D1,…,DTD_1, \dots, D_T, base language model LLMθ\text{LLM}_\theta, number of target library adapters KK
    Output: Adapter library L\mathcal{L}
    L=∅\mathcal{L} = \emptyset
    A=∅\mathcal{A} = \emptyset
    for t=1t = 1 to TT do
      At,Bt=train(Dt,LLMθ)A_t, B_t = \text{train}(D_t, \text{LLM}_\theta) (train for nn steps)
      A=A∪{concat(flatten(At),flatten(Bt))}\mathcal{A} = \mathcal{A} \cup \{\text{concat}(\text{flatten}(A_t), \text{flatten}(B_t))\}
    end for
    U=SVD(A)U = \text{SVD}(\mathcal{A})
    S=cosine_similarity(U,U)S = \text{cosine\_similarity}(U, U)
    c1,…,cK=k-means(S,K)c_1, \dots, c_K = k\text{-means}(S, K)
    for k=1k = 1 to KK do
      Dk=⋃t∈ckDtD_k = \bigcup_{t \in c_k} D_t
      Ak,Bk=train(Dk,LLMθ)A_k, B_k = \text{train}(D_k, \text{LLM}_\theta) (train for N−nN - n steps)
      L=L∪{(Ak,Bk)}\mathcal{L} = \mathcal{L} \cup \{(A_k, B_k)\}
    end for
    return L\mathcal{L}
  3. Knowl 3 — Correlation Between LoRA Parameter Similarity and Multi-Task Transfer

    empirical result

    The cosine similarity between the flattened weight vectors of task-specific LoRA adapters trained independently correlates positively with the performance delta gained from training a single joint LoRA on the combined data of both tasks versus training them separately.

    For two tasks ti,tjt_i, t_j, let ΔL\Delta \mathcal{L} be the difference in average log-likelihood on the test sets when evaluated with a LoRA trained jointly on Di∪DjD_i \cup D_j versus two private LoRAs trained on DiD_i and DjD_j individually. Across task pairs sampled from Flan-v2:

    • GPT-Neo shows a Pearson correlation coefficient of r=0.75r = 0.75 between LoRA cosine similarity and joint training performance delta ΔL\Delta \mathcal{L}.
    • Phi-2 (2.8B) shows a Pearson correlation coefficient of r=0.51r = 0.51 between LoRA cosine similarity and joint training performance delta ΔL\Delta \mathcal{L}.

    This positive correlation provides the foundation for using parameter-space similarity as a surrogate distance metric for task clustering prior to multi-task adapter training.

  4. Knowl 4 — Linear Composition of LoRA Adapters Across Layers

    equation

    Given an adapter library L={(Ai,Bi)}i=1∣L∣\mathcal{L} = \{(A_i, B_i)\}_{i=1}^{|\mathcal{L}|} with low-rank matrices Ai,Bi∈Rd×rA_i, B_i \in \mathbb{R}^{d \times r}, and a layer-specific routing distribution w=p(⋅∣h∗,x)=(w1,…,w∣L∣)w = p(\cdot \mid h^*, x) = (w_1, \dots, w_{|\mathcal{L}|}) over the library modules, the composed LoRA adapter matrices (A∗,B∗)(A^*, B^*) applied at that layer are given by linear parameter aggregation:

    A∗=∑i=1∣L∣wiAi,B∗=∑i=1∣L∣wiBiA^* = \sum_{i=1}^{|\mathcal{L}|} w_i A_i, \quad B^* = \sum_{i=1}^{|\mathcal{L}|} w_i B_i

    The modified transformer layer transformation on input hidden state xx with frozen base model weights W∈Rd×dW \in \mathbb{R}^{d \times d} and scaling factor s≥1s \ge 1 is:

    h=Wx+s⋅A∗(B∗)⊤xh = W x + s \cdot A^* (B^*)^\top x

    In top-kk zero-shot routing, ww is non-zero only for the kk experts that yield the largest routing logits.

  5. Knowl 5 — Zero-Shot Evaluation of Adapter Libraries and Routing Baselines

    data/table

    Downstream zero-shot accuracy across 10 held-out evaluation tasks (PIQA, BoolQ, WinoGrande [WG], HellaSwag [HSWAG], ARC-Easy [ARCE], ARC-Challenge [ARCC], HumanEval [HE], OpenBookQA [OQA], Big-Bench Hard [BBH], and MBPP) for Phi-2 (2.8B) and Mistral (7B) backbones comparing unadapted base models, fully fine-tuned models, shared multi-task LoRA, and modular libraries under uniform (μ\mu) and Arrow (↗\nearrow) routing:

    Library Route ∣L∣|\mathcal{L}| PIQA BOOLQ WG HSWAG ARCE ARCC HE OQA BBH MBPP Acc.
    Phi-2 (2.8B)
    Base - - 79.2 82.7 75.7 72.5 77.5 52.9 45.1 49.8 48.0 56.0 63.8
    FullFT - - 80.3 80.8 77.0 73.2 83.5 57.9 50.0 48.0 47.7 57.2 65.6
    Shared - 1 80.4 82.4 76.6 73.4 83.2 55.8 46.3 50.4 48.4 58.4 65.5
    Poly μ\mu 8 80.6 82.3 76.7 71.7 82.7 55.3 48.2 50.4 49.8 59.1 65.7
    MHR μ\mu 8 80.1 83.0 77.1 70.4 83.2 55.5 46.3 53.4 52.0 58.0 65.9
    Private μ\mu 256 79.5 83.2 76.0 73.1 81.4 54.5 43.9 47.8 48.5 59.9 64.8
    Private ↗\nearrow 256 80.2 84.3 77.6 72.6 84.2 56.4 50.6 52.2 47.7 59.9 66.6
    MBC μ\mu 10 80.3 85.1 77.3 73.1 84.3 57.7 48.8 50.2 51.6 62.3 67.1
    MBC ↗\nearrow 10 79.9 84.7 77.7 72.9 84.8 57.9 51.8 50.2 52.2 62.3 67.4
    Mistral (7B)
    Base - - 81.1 82.2 66.5 78.8 68.9 49.6 28.0 44.6 47.9 47.5 59.5
    Shared - 1 50.4 84.6 68.6 79.5 84.8 60.0 24.4 50.4 49.2 47.5 63.1
    Private μ\mu 256 82.1 82.7 67.2 79.6 78.7 54.8 29.9 45.2 49.0 49.4 61.9
    Private ↗\nearrow 256 82.8 86.6 66.6 81.1 85.7 60.8 30.5 50.6 49.5 49.4 64.4
    MBC μ\mu 10 83.0 87.6 68.5 80.8 86.2 60.9 28.7 48.6 51.5 50.2 64.6
    MBC ↗\nearrow 10 82.8 87.3 70.6 80.9 84.5 59.6 28.0 52.8 45.5 47.1 63.9

    These results demonstrate that:

    1. Reusing independently trained private LoRAs with Arrow routing (Private-↗\nearrow) achieves 66.6%66.6\% on Phi-2 and 64.4%64.4\% on Mistral, outperforming full multi-task fine-tuning (FullFT: 65.6%65.6\%) and shared LoRA training without requiring simultaneous data access.
    2. Model-Based Clustering (MBC) with 10 clusters yields the overall highest accuracy (67.4%67.4\% on Phi-2).
  6. Knowl 6 — Validation of Arrow Norm-Maximization Hypothesis

    empirical result

    Arrow routing relies on the hypothesis that for an input activation hlh_l originating from task ii at layer ll, the in-distribution LoRA adapter (AB⊤)li(AB^\top)_l^i produces an output transformation of greater norm than an out-of-distribution LoRA adapter (AB⊤)lj(AB^\top)_l^j from task j≠ij \neq i:

    ∥(AB⊤)lihl∥2>∥(AB⊤)ljhl∥2\|(AB^\top)_l^i h_l\|_2 > \|(AB^\top)_l^j h_l\|_2

    Evaluating across 5,000 samples from the multi-task dataset and averaging across all LL layers using the average norm ratio:

    r=1L∑l=1L∥(AB⊤)lihli∥2∥(AB⊤)ljhli∥2r = \frac{1}{L} \sum_{l=1}^L \frac{\|(AB^\top)_l^i h_l^i\|_2}{\|(AB^\top)_l^j h_l^i\|_2}

    yields r>1.0r > 1.0 for all sampled points (empirical density centered between 1.0021.002 and 1.0101.010). This confirms that in-distribution LoRA adapters consistently produce larger transformation norms on matching inputs than randomly selected out-of-distribution adapters, validating the right-singular-vector prototype initialization used in Arrow routing.

  7. Knowl 7 — Ablation of Multi-Task Clustering Strategies for Library Construction

    data/table

    Comparison of different grouping strategies to partition 256 Flan-v2 tasks into K=10K=10 clusters for LoRA library building on Phi-2, evaluated via mean zero-shot accuracy across 10 held-out tasks and average pairwise cosine similarity between final cluster adapter weights:

    Clustering Method Mean Acc. (%) Pairwise Similarity
    RandExamples∗^*-μ\mu 64.8 0.82
    RandTask∗^*-μ\mu 66.4 0.58
    RandTask-μ\mu 66.4 0.58
    Embeddings∗^*-μ\mu 66.1 0.37
    MBC∗^*-μ\mu 66.7 0.37
    MBC-μ\mu 67.1 0.27

    Note: $*$ denotes 1 epoch of training to save computation; methods without $*$ use standard training schedules.

    Key findings:

    1. Partitioning tasks (RandTask: 66.4%66.4\%) significantly outperforms partitioning random examples across the joint dataset (RandExamples: 64.8%64.8\%), showing task grouping prevents negative multi-task interference.
    2. Clustering based on initial LoRA weight similarity (MBC: 67.1%67.1\%) outperforms clustering based on model penultimate-layer activations (Embeddings: 66.1%66.1\%).
    3. Higher downstream performance correlates with lower average pairwise cosine similarity between cluster adapters (MBC achieves the highest diversity with similarity 0.270.27).
  8. Knowl 8 — Supervised Adaptation Performance Across Data Regimes

    empirical result

    Supervised adaptation evaluated on 12 held-out SuperNatural Instructions (SNI) categories for Phi-2 (2.8B) under full data (100%100\%, up to 10,000 samples/task), limited data (10%10\%, up to 1,000 samples/task), and few-shot regimes (5%5\% and 0.5%0.5\%, approx. 40 samples/task) using Rouge-L:

    1. Full Data (100%) vs. Limited Data (10%):

      • Base model (unadapted): 22.2 Rouge-L (100% and 10%).
      • No Library (random init LoRA per task): 75.5 Rouge-L (100%) / 53.9 Rouge-L (10%).
      • Shared multi-task LoRA: 75.8 Rouge-L (100%) / 56.4 Rouge-L (10%).
      • Poly: 73.4 Rouge-L (100%) / 61.7 Rouge-L (10%).
      • MHR: 74.8 Rouge-L (100%) / 64.5 Rouge-L (10%).
      • π\pi-tuning: 76.7 Rouge-L (100%) / 64.6 Rouge-L (10%).
      • Private-μ\mu: 76.9 Rouge-L (100%) / 62.5 Rouge-L (10%).
      • RandTask-Poly: 76.7 Rouge-L (100%) / 67.6 Rouge-L (10%).
      • MBC-μ\mu: 78.8 Rouge-L (100%) / 67.0 Rouge-L (10%).
      • MBC-Poly: 78.8 Rouge-L (100%) / 68.2 Rouge-L (10%).
    2. Few-Shot Adaptation (Gradient-based vs. Gradient-free):

      • In data-scarce settings (0.5% data regime, ~40 examples), gradient-based adaptation of both routing coefficients and adapter weights (MBC-Poly) achieves 50.2 Rouge-L, whereas gradient-free routing coefficient optimization (LoraHub) achieves only 45.3 Rouge-L (full data) and is severely constrained because it cannot update underlying expert weights.
      • Updating only the routing parameters without updating LoRA weights (MBC-PolyZ) reaches 45.3 Rouge-L on full data and 44.8 Rouge-L on 0.5% data.
  9. Knowl 9 — Layer Selection Sensitivity for LoRA Patching

    empirical result

    When applying LoRA adapters across transformer layers, patching attention projection layers (query, key, value, and output projection: ATT+O) yields substantially better parameter efficiency and downstream accuracy than patching feed-forward / multi-layer perceptron fully-connected layers (FC).

    Patching FC layers alone severely underperforms ATT+O layers. Patching both attention and MLP layers (BOTH) yields only marginal accuracy gains over ATT+O while incurring a large compute and memory penalty due to the wide intermediate hidden projection dimension (4×d4 \times d) of the transformer MLP block.

  10. Knowl 10 — Effect of Adapter Library Size on Zero-Shot Routing Expressivity

    empirical result

    The necessity of non-uniform routing depends strongly on the adapter library size ∣L∣|\mathcal{L}|:

    1. Large libraries (∣L∣=256|\mathcal{L}| = 256 private experts): Dynamic token-level routing is critical. Arrow (↗\nearrow) improves mean downstream zero-shot accuracy over uniform averaging (μ\mu) by +1.8%+1.8\% on Phi-2 (66.6%66.6\% vs. 64.8%64.8\%) and by +2.5%+2.5\% on Mistral 7B (64.4%64.4\% vs. 61.9%61.9\%). On the 12 SNI zero-shot evaluation tasks, Private-↗\nearrow outperforms Private-μ\mu by +10.5+10.5 Rouge-L points on Mistral 7B (46.946.9 vs. 36.436.4). Furthermore, Arrow closely tracks Oracle routing on upstream training validation sets without knowing task identities.

    2. Compact libraries (∣L∣=10|\mathcal{L}| = 10 MBC experts): Dynamic routing yields diminishing returns over uniform averaging (MBC-↗\nearrow yields 67.4%67.4\% vs. MBC-μ\mu 67.1%67.1\% on Phi-2, and 63.9%63.9\% vs. 64.6%64.6\% on Mistral). This occurs because MBC groups correlated tasks into robust generalist clusters whose linear average already retains strong transfer properties without high-variance token routing.

Coverage note — None was omitted; all primary methodological contributions (MBC, Arrow, linear LoRA composition) and experimental evaluations (zero-shot benchmarks, supervised SNI benchmarks, clustering ablations, layer selection ablations, and routing hypotheses) are covered.

References

  1. 1.Ainsworth, S. K., Hayase, J., and Srinivasa, S. Git re-basin: Merging models modulo permutation symmetries. arXiv preprint arXiv:2209.04836, 2022.
  2. 2.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  3. 3.Beck, T., Bohlender, B., Viehmann, C., Hane, V., Adamson, Y., Khuri, J., Brossmann, J., Pfeiffer, J., and Gurevych, I. Adapterhub playground: Simple and flexible few-shot learning with adapters. arXiv preprint arXiv:2108.08103, 2021.
  4. 4.Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  5. 5.Belofsky, J. Token-level adaptation of lora adapters for downstream task generalization, 2023.
  6. 6.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, pp. 7432–7439, 2020.
  7. 7.Bragman, F. J., Tanno, R., Ourselin, S., Alexander, D. C., and Cardoso, J. Stochastic filter groups for multi-task cnns: Learning specialist and generalist convolution kernels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1385–1394, 2019.
  8. 8.Breiman, L. Bagging predictors. Machine learning, 24: 123–140, 1996.
  9. 9.Caccia, L., Ponti, E., Su, Z., Pereira, M., Roux, N. L., and Sordoni, A. Multi-head adapter routing for cross-task generalization, 2023.
  10. 10.Caron, M., Bojanowski, P., Joulin, A., and Douze, M. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV), pp. 132–149, 2018.
  11. 11.Caruana, R. Multitask learning. Machine learning, 28: 41–75, 1997.
  12. 12.Chai, Z., Wang, G., Su, J., Zhang, T., Huang, X., Wang, X., Xu, J., Yuan, J., Yang, H., Wu, F., et al. An expert is worth one token: Synergizing multiple expert llms as generalist via expert token routing. arXiv preprint arXiv:2403.16854, 2024.
  13. 13.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  14. 14.Chen, Z., Shen, Y., Ding, M., Chen, Z., Zhao, H., Learned-Miller, E., and Gan, C. Mod-squad: Designing mixture of experts as modular multi-task learners, 2022.
  15. 15.Chronopoulou, A., Peters, M. E., and Dodge, J. Efficient hierarchical domain adaptation for pretrained language models. arXiv preprint arXiv:2112.08786, 2021.
  16. 16.Chronopoulou, A., Peters, M. E., Fraser, A., and Dodge, J. Adaptersoup: Weight averaging to improve generalization of pretrained language models. arXiv preprint arXiv:2302.07027, 2023a.
  17. 17.Chronopoulou, A., Pfeiffer, J., Maynez, J., Wang, X., Ruder, S., and Agrawal, P. Language and task arithmetic with parameter-efficient layers for zero-shot summarization. arXiv preprint arXiv:2311.09344, 2023b.
  18. 18.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. In NAACL, 2019.
  19. 19.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  20. 20.Daheim, N., Möllenhoff, T., Ponti, E., Gurevych, I., and Khan, M. E. Model merging by uncertainty-based gradient matching. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=D7KJmfEDQP.
  21. 21.Dietterich, T. G. Ensemble methods in machine learning. In International workshop on multiple classifier systems, pp. 1–15. Springer, 2000.
  22. 22.Ding, C., Lu, Z., Wang, S., Cheng, R., and Boddeti, V. N. Mitigating task interference in multi-task learning via explicit task routing with non-learnable primitives. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7756–7765, 2023.
  23. 23.Duan, Z., Zhang, H., Wang, C., Wang, Z., Chen, B., and Zhou, M. Enslm: Ensemble language model for data diversity by semantic clustering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 2954–2967, 2021.
  24. 24.EleutherAI. Multiple-choice normalization. https://blog.eleuther.ai/multiple-choice-normalization/, 2021. Accessed: 2024-05-12.
  25. 25.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1, 2021.
  26. 26.Ermis, B., Zappella, G., Wistuba, M., Rawal, A., and Archambeau, C. Memory efficient continual learning with transformers. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=U07d1Y-x2E.
  27. 27.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. URL http://jmlr.org/papers/v23/21-0998.html.
  28. 28.Fifty, C., Amid, E., Zhao, Z., Yu, T., Anil, R., and Finn, C. Efficiently identifying task groupings for multi-task learning. Advances in Neural Information Processing Systems, 34:27503–27516, 2021.
  29. 29.Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning, pp. 3259–3269. PMLR, 2020.
  30. 30.Gou, Y., Liu, Z., Chen, K., Hong, L., Xu, H., Li, A., Yeung, D.-Y., Kwok, J. T., and Zhang, Y. Mixture of cluster-conditional lora experts for vision-language instruction tuning. arXiv preprint arXiv:2312.12379, 2023.
  31. 31.Gross, S., Ranzato, M., and Szlam, A. Hard mixtures of experts for large scale weakly supervised vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6865–6873, 2017.
  32. 32.Gupta, S., Mukherjee, S., Subudhi, K., Gonzalez, E., Jose, D., Awadallah, A. H., and Gao, J. Sparsely activated mixture-of-experts are robust multi-task learners. arXiv preprint arXiv:2204.07689, 2022.
  33. 33.Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L. Scaling expert language models with unsupervised domain discovery. arXiv preprint arXiv:2303.14177, 2023.
  34. 34.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, pp. 2790–2799, 2019. URL http://proceedings.mlr.press/v97/houlsby19a/houlsby19a.pdf.
  35. 35.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  36. 36.Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M. Lorahub: Efficient cross-task generalization via dynamic lora composition. arXiv preprint arXiv:2307.13269, 2023.
  37. 37.Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M. Lorahub: Efficient cross-task generalization via dynamic lora composition, 2024.
  38. 38.Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089, 2022.
  39. 39.Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M. Exploring the benefits of training expert language models over instruction tuning, 2023.
  40. 40.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  41. 41.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts, 2024.
  42. 42.Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P. Dataless knowledge fusion by merging weights of language models. arXiv preprint arXiv:2212.09849, 2022.
  43. 43.Jolicoeur-Martineau, A., Gervais, E., Fatras, K., Zhang, Y., and Lacoste-Julien, S. Population parameter averaging (papa), 2023.
  44. 44.Karimi Mahabadi, R., Ruder, S., Dehghani, M., and Henderson, J. Parameter-efficient multi-task fine-tuning for Transformers via shared hypernetworks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pp. 565–576, August 2021. URL https://aclanthology.org/2021.acl-long.47.
  45. 45.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning, 2021. URL https://arxiv.org/pdf/2104.08691.pdf.
  46. 46.Leviathan, Y., Kalman, M., and Matias, Y. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pp. 19274–19286. PMLR, 2023.
  47. 47.Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L. Branch-train-merge: Embarrassingly parallel training of expert language models. arXiv preprint arXiv:2208.03306, 2022.
  48. 48.Li, X. L. and Eisner, J. Specializing word embeddings (for parsing) by information bottleneck. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2744–2754, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1276. URL https://www.aclweb.org/anthology/D19-1276.
  49. 49.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4582–4597, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.353. URL https://aclanthology.org/2021.acl-long.353.
  50. 50.Lin, C.-Y. and Hovy, E. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 human language technology conference of the North American chapter of the association for computational linguistics, pp. 150–157, 2003.
  51. 51.Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning, 2022. URL https://arxiv.org/abs/2205.05638.
  52. 52.Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023.
  53. 53.Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J. Routing to the expert: Efficient reward-guided ensemble of large language models. arXiv preprint arXiv:2311.08692, 2023.
  54. 54.Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  55. 55.Matena, M. S. and Raffel, C. A. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems, 35:17703–17716, 2022.
  56. 56.Microsoft Research. Phi-2: The Surprising Power of Small Language Models, 2023.
  57. 57.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  58. 58.Mireshghallah, F., Taram, M., Vepakomma, P., Singh, A., Raskar, R., and Esmaeilzadeh, H. Privacy in deep learning: A survey. arXiv preprint arXiv:2004.12254, 2020.
  59. 59.Muqeeth, M., Liu, H., and Raffel, C. Soft merging of experts with adaptive routing. arXiv preprint arXiv:2306.03745, 2023.
  60. 60.Muqeeth, M., Liu, H., Liu, Y., and Raffel, C. Learning to route among specialized experts for zero-shot generalization. arXiv preprint arXiv: 2402.05859, 2024.
  61. 61.Nakatsukasa, Y. The low-rank eigenvalue problem. arXiv preprint arXiv:1905.11490, 2019.
  62. 62.Ostapenko, O., Rodriguez, P., Caccia, M., and Charlin, L. Continual learning via local module composition. Advances in Neural Information Processing Systems, 34, 2021. URL https://proceedings.neurips.cc/paper/2021/file/fe5e7cb609bdbe6d62449d61849c38b0-Paper.pdf.
  63. 63.Ostapenko, O., Caccia, L., Su, Z., Le Roux, N., Charlin, L., and Sordoni, A. A case study of instruction tuning with mixture of parameter-efficient experts. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023.
  64. 64.Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pp. 487–503, April 2021. URL https://aclanthology.org/2021.eacl-main.39.
  65. 65.Pfeiffer, J., Ruder, S., Vulić, I., and Ponti, E. M. Modular deep learning. arXiv preprint arXiv:2302.11529, 2023. URL https://arxiv.org/pdf/2302.11529.pdf.
  66. 66.Ponti, E. M., Sordoni, A., Bengio, Y., and Reddy, S. Combining parameter-efficient modules for task-level generalisation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 687–702, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.eacl-main.49.
  67. 67.Poth, C., Sterz, H., Paul, I., Purkayastha, S., Engländer, L., Imhof, T., Vulić, I., Ruder, S., Gurevych, I., and Pfeiffer, J. Adapters: A unified library for parameter-efficient and modular transfer learning. arXiv preprint arXiv:2311.11077, 2023.
  68. 68.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21: 1–67, 2020. URL https://www.jmlr.org/papers/volume21/20-074/20-074.pdf.
  69. 69.Ramé, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D. Model ratatouille: Recycling diverse models for out-of-distribution generalization. In International Conference on Machine Learning, pp. 28656–28679. PMLR, 2023.
  70. 70.Rapin, J. and Teytaud, O. Nevergrad - A gradient-free optimization platform. https://GitHub.com/FacebookResearch/Nevergrad, 2018.
  71. 71.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  72. 72.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  73. 73.Shah, V., Ruiz, N., Cole, F., Lu, E., Lazebnik, S., Li, Y., and Jampani, V. Ziplora: Any subject in any style by effectively merging loras. arXiv preprint arXiv:2311.13600, 2023.
  74. 74.Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  75. 75.Shen, S. Z., Lang, H., Wang, B., Kim, Y., and Sontag, D. Learning to decode collaboratively with multiple language models. arXiv preprint arXiv:2403.03870, 2024.
  76. 76.Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M. Large language model routing with benchmark datasets. arXiv preprint arXiv:2309.15789, 2023.
  77. 77.Strezoski, G., Noord, N. v., and Worring, M. Many task learning with task routing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1375–1384, 2019.
  78. 78.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  79. 79.Tam, D., Bansal, M., and Raffel, C. Merging by matching models in task subspaces. arXiv preprint arXiv:2312.04339, 2023.
  80. 80.Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al. Towards generalist biomedical ai. arxiv. Preprint posted online, 26, 2023.
  81. 81.Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M. Exploring and predicting transferability across NLP tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7882–7926, Online, November 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.635. URL https://aclanthology.org/2020.emnlp-main.635.
  82. 82.Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M. Exploring and predicting transferability across nlp tasks. arXiv preprint arXiv:2005.00770, 2020b.
  83. 83.Vu, T., Lester, B., Constant, N., Al-Rfou, R., and Cer, D. Spot: Better frozen model adaptation through soft prompt transfer. arXiv preprint arXiv:2110.07904, 2021.
  84. 84.Wallingford, M., Li, H., Achille, A., Ravichandran, A., Fowlkes, C., Bhotika, R., and Soatto, S. Task adaptive parameter sharing for multi-task learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7561–7570, 2022.
  85. 85.Wang, Y., Agarwal, S., Mukherjee, S., Liu, X., Gao, J., Awadallah, A. H., and Gao, J. Adamix: Mixture-of-adaptations for parameter-efficient model tuning. arXiv preprint arXiv:2205.12410, 2022a.
  86. 86.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. arXiv preprint arXiv:2204.07705, 2022b.
  87. 87.Wang, Z., Tsvetkov, Y., Firat, O., and Cao, Y. Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=F1vEjWK-lH_.
  88. 88.Wang, Z., Zhang, Z., Lee, C.-Y., Zhang, H., Sun, R., Ren, X., Su, G., Perot, V., Dy, J., and Pfister, T. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 139–149, 2022c.
  89. 89.Wen, Y., Tran, D., and Ba, J. Batchensemble: An alternative approach to efficient ensemble and lifelong learning, 2020.
  90. 90.Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. arXiv preprint arXiv:2302.03668, 2023.
  91. 91.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning, pp. 23965–23998. PMLR, 2022.
  92. 92.Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., and Luo, P. π-tuning: Transferring multimodal foundation models with optimal multi-task interpolation. In International Conference on Machine Learning, pp. 37713–37727. PMLR, 2023.
  93. 93.Xun Wu, Shaohan Huang, F. W. Mole: Mixture of lora experts. In International Conference on Learning Representations, ICLR 2024, 2024. URL https://openreview.net/forum?id=uWvKBCYh4S.
  94. 94.Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M. Ties-merging: Resolving interference when merging models. Advances in Neural Information Processing Systems, 36, 2024.
  95. 95.Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D. Adamerging: Adaptive model merging for multi-task learning. arXiv preprint arXiv:2310.02575, 2023.
  96. 96.Ye, Q., Zha, J., and Ren, X. Eliciting and understanding cross-task skills with task-level mixture-of-experts. arXiv preprint arXiv:2205.12701, 2022.
  97. 97.Zadouri, T., Üstün, A., Ahmadian, A., Ermiş, B., Locatelli, A., and Hooker, S. Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning. arXiv preprint arXiv:2309.05444, 2023.
  98. 98.Zaremoodi, P., Buntine, W., and Haffari, G. Adaptive knowledge sharing in multi-task learning: Improving low-resource neural machine translation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 656–661, 2018.
  99. 99.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  100. 100.Zhang, J., Chen, S., Liu, J., and He, J. Composing parameter-efficient modules with arithmetic operations. arXiv preprint arXiv:2306.14870, 2023a.
  101. 101.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023b.
  102. 102.Zhao, X., Li, H., Shen, X., Liang, X., and Wu, Y. A modulation module for multi-task learning with applications in image retrieval. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 401–416, 2018.
  103. 103.Zhou, W., Xu, C., and McAuley, J. Efficiently tuned parameters are task embeddings. arXiv preprint arXiv:2210.11705, 2022.

Citation

MLA
Ostapenko, O., et al. “Towards Modular LLMs by Building and Reusing a Library of LoRAs”. arXiv, 2024, http://arxiv.org/abs/2405.11157v1.
APA
Ostapenko, O., Su, Z., Ponti, E. M., Charlin, L., Roux, N. L., Pereira, M., Caccia, L., & Sordoni, A. (2024). Towards Modular LLMs by Building and Reusing a Library of LoRAs. arXiv. http://arxiv.org/abs/2405.11157v1
Chicago
Ostapenko, O., Z. Su, E. M. Ponti, et al. 2024. “Towards Modular LLMs by Building and Reusing a Library of LoRAs”. arXiv. http://arxiv.org/abs/2405.11157v1.
Harvard
Ostapenko, O. et al. (2024) “Towards Modular LLMs by Building and Reusing a Library of LoRAs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.11157v1.
Vancouver
1. Ostapenko O, Su Z, Ponti EM, Charlin L, Roux NL, Pereira M, Caccia L, Sordoni A (2024) Towards Modular LLMs by Building and Reusing a Library of LoRAs. arXiv

BibTeX

@article{ostapenko2024towards,
  title = {Towards Modular LLMs by Building and Reusing a Library of LoRAs},
  author = {Ostapenko, Oleksiy and Su, Zhan and Ponti, Edoardo Maria and Charlin, Laurent and Roux, Nicolas Le and Pereira, Matheus and Caccia, Lucas and Sordoni, Alessandro},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.11157v1},
  eprint = {2405.11157}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/