Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying
Adithya RenduchintalaTugrul KonukOleksii Kuchaiev
Proposes Tied-LoRA, a parameter-efficient fine-tuning framework that combines layer-shared projection matrices with selective training to match standard LoRA performance across diverse tasks while using up to 87.5% fewer trainable parameters.
Customizing large language models for specific tasks and diverse user preferences is essential for practical enterprise deployment. However, maintaining distinct customized model weights for numerous user-task combinations introduces substantial computational expenses during training, as well as significant operational and storage costs during post-training deployment. While parameter-efficient fine-tuning techniques like Low-Rank Adaptation (LoRA) reduce training burdens, their parameter overhead remains non-trivial as base models scale in depth and size.
The article demonstrates an approach called Tied-LoRA, which integrates weight tying and selective parameter training to dramatically reduce the number of trainable parameters in LoRA while maintaining high task performance. The authors evaluate various parameter configurations across five distinct natural language processing tasks—extractive question answering, dialogue summarization, commonsense reasoning, machine translation, and mathematical reasoning—using two base language models of different scales (a 2-billion parameter model and a 7-billion parameter model).
The evaluation reveals three major findings. First, a specific configuration termed TL6 achieves performance within 1% to 2% of standard LoRA on average, while utilizing only a small fraction of the parameters. In a 7-billion parameter translation task, TL6 matched and slightly exceeded LoRA's performance while using only 12.5% of the trainable parameters. Second, the performance gap between TL6 and standard LoRA narrowed when moving from the smaller 2-billion parameter model to the larger 7-billion parameter model, suggesting that larger foundation models benefit more from parameter sharing. Third, applying shared low-rank updates across all model layers substantially outperformed applying unshared LoRA updates to any individual layer, even when using the exact same parameter budget.
These findings indicate that organizations can drastically reduce storage, memory, and operational serving costs for specialized language models without suffering meaningful drops in accuracy, particularly on tasks that align with the base model's core language capabilities. However, complex mathematical reasoning tasks still favor standard LoRA due to the higher parameter capacity required for arithmetic operations.
Organizations serving multiple customized language model variants should consider adopting TL6 as a drop-in replacement for standard LoRA to lower deployment costs and storage footprint. Prior to full deployment, engineering teams should validate performance on complex analytical tasks and conduct pilot tests, as further empirical work on ultra-large base models (such as 70-billion parameter models) is still needed.
- Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). This foundational work introduces Low-Rank Adaptation (LoRA), establishing the core low-rank parameterization and adapter architecture that Tied-LoRA directly extends through weight tying.
- Paper: AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning, Qingru Zhang et al. (2023). This paper establishes adaptive budget allocation across transformer layers for LoRA, providing critical background on the uneven layer-wise parameter sensitivity addressed in Tied-LoRA.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This work formulates a unified mathematical taxonomy of parameter-efficient fine-tuning methods, clarifying the design space of modular hidden-state modifications underlying low-rank parameter sharing.
- Paper: QLoRA: Efficient Finetuning of Quantized LLMs, Tim Dettmers et al. (2023). This study demonstrates how low-rank adapters interact with memory-compressed base models, establishing the standard baseline framework for parameter-efficient adaptation at scale.
- Paper: Predicting Parameters in Deep Learning, Misha Denil et al. (2013). This paper demonstrates foundational parameter redundancy in deep networks through low-rank decompositions and static parameter sharing, motivating Tied-LoRA's weight-tying strategy.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). This work introduces modular parameter-efficient bottleneck adapter tuning for transformers, defining the architectural paradigm that subsequent low-rank fine-tuning methods build upon.
- Paper: VeRA: Vector-based Random Matrix Adaptation, Dawid J. Kopiczko et al. (2024). VeRA pushes parameter sharing across transformer layers even further by freezing a single pair of random matrices across all layers and learning only layer-wise scaling vectors.
- Paper: S-LoRA: Serving Thousands of Concurrent LoRA Adapters, Ying Sheng et al. (2024). S-LoRA provides the specialized serving infrastructure and GPU memory-pooling mechanisms necessary to deploy and serve thousands of distinct parameter-efficient adapters concurrently.
- Paper: DoRA: Weight-Decomposed Low-Rank Adaptation, Shih-Yang Liu et al. (2024). DoRA enhances parameter-efficient low-rank adaptation by decomposing updates into magnitude and direction components, offering a complementary architectural refinement to parameter allocation.
- Paper: LoRA+: Efficient Low Rank Adaptation of Large Models, Soufiane Hayou et al. (2024). LoRA+ analyzes and rectifies optimization inefficiencies in low-rank adapter matrices via asymmetric learning rates, directly complementing structural parameter-reduction techniques like Tied-LoRA.
- Paper: PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models, Fanxu Meng et al. (2024). PiSSA improves low-rank adaptation performance by initializing adapter matrices via principal singular components of the pretrained weights rather than random initialization.
- Paper: LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models, Yaowei Zheng et al. (2024). LLaMA-Factory implements a unified open-source framework to train, benchmark, and deploy diverse parameter-efficient adaptation methods across hundreds of LLM architectures.
- Paper: On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters, Mind Lab et al. (2026). This work explores multi-axis scaling properties of modular parameter-efficient adapters to support millions of concurrent personal models on massive foundation models.
