Towards Modular LLMs by Building and Reusing a Library of LoRAs
Oleksiy OstapenkoZhan SuEdoardo M. PontiLaurent CharlinNicolas Le RouxLucas CacciaAlessandro Sordoni
Proposes a modular framework that clusters LoRA adapters by parameter similarity and dynamically routes hidden states to the most relevant adapters at inference time, achieving superior zero-shot and supervised generalization without requiring joint retraining.
Adapting large language models to new tasks typically requires extensive computational resources and centralized datasets. While parameter-efficient adapters, such as Low-Rank Adaptation (LoRA), allow models to be fine-tuned cheaply on specific tasks, organizations face significant challenges when trying to combine independently developed adapters to solve new, unseen problems without retraining entire models.
The article investigates how to efficiently build and reuse libraries of modular adapters. Specifically, it demonstrates how to group training tasks based on model parameters to maximize knowledge transfer and introduces an automated routing system to direct new inputs to the most relevant adapters without requiring access to original training data.
To construct the adapter library, the authors developed Model-Based Clustering (MBC), which measures the mathematical similarity between independently trained adapter weights and clusters related tasks together before training a single consolidated adapter per cluster. To reuse these libraries on new tasks without further training, the authors designed Arrow, a routing mechanism that identifies the primary direction of variation within each adapter's weights and dynamically matches token representations to the most suitable adapters at inference time. The methods were evaluated on multi-task benchmarks using 256 training tasks across models including Phi-2 (2.8 billion parameters) and Mistral (7 billion parameters), testing on multiple held-out reasoning, coding, and question-answering tasks.
The findings show that weight similarity between adapters serves as an effective proxy for task compatibility and positive transfer. Building libraries with Model-Based Clustering outperformed standard multi-task baselines, achieving a 67.4% zero-shot accuracy on Phi-2 held-out benchmarks compared to 65.6% for full model fine-tuning and 63.8% for the base model. Furthermore, when routing large collections of independently trained private adapters, Arrow outperformed uniform adapter averaging by 1.8 percentage points on Phi-2 and 2.5 percentage points on Mistral, matching or exceeding the performance of computationally intensive joint training. In supervised adaptation settings with limited data, initializing from clustered adapter libraries consistently accelerated learning and improved task scores over non-modular baselines.
These results indicate that organizations can achieve state-of-the-art model adaptation in a decentralized manner, avoiding the privacy risks, infrastructure requirements, and energy costs of continuous centralized multi-task training. Practitioners looking to deploy modular language models should adopt weight-based clustering when shared data is accessible, or use Arrow-style prototype routing when integrating decentralized libraries of adapters. However, because the study focused exclusively on linear LoRA adapters and tested models up to 7 billion parameters, teams should conduct pilot evaluations before applying these routing strategies to larger models or non-linear adapter architectures.
- Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). Introduces Low-Rank Adaptation (LoRA), the core parameter-efficient fine-tuning module and mathematical foundation upon which the source paper builds its adapter clustering and routing framework.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). Pioneers the concept of modular adapter tuning for transfer learning in frozen language models, establishing the paradigm of lightweight parameter reuse.
- Paper: Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy Constraints, Felix Sattler et al. (2019). Establishes model-weight similarity and geometric clustering as effective mechanisms for grouping compatible distributed tasks under privacy constraints.
- Paper: AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning, Yaqing Wang et al. (2022). Presents a mixture-of-adaptations approach that routes representations across multiple parameter-efficient modules, directly preceding dynamic adapter routing mechanisms.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Provides a comprehensive taxonomy and analysis of parameter-efficient transfer methods, framing the functional behavior of low-rank updates across transformer layers.
- Paper: S-LoRA: Serving Thousands of Concurrent LoRA Adapters, Ying Sheng et al. (2024). Addresses the downstream infrastructure and serving challenges of deploying thousands of concurrent LoRA adapters on shared hardware.
- Paper: On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters, Mind Lab et al. (2026). Scales the vision of modular, personalized low-rank adapters up to millions of individual instances layered over frontier foundation models.
- Paper: Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying, Adithya Renduchintala et al. (2024). Extends adapter parameter efficiency by combining weight tying with selective LoRA parameter training across multi-task deployments.
- Paper: VeRA: Vector-based Random Matrix Adaptation, Dawid J. Kopiczko et al. (2024). Offers an alternative vector-based parameterization to drastically compress adapter memory footprints when maintaining massive task libraries.
