Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying

Adithya RenduchintalaTugrul KonukOleksii Kuchaiev

article2024NAACL88 citations

Proposes Tied-LoRA, a parameter-efficient fine-tuning framework that combines layer-shared projection matrices with selective training to match standard LoRA performance across diverse tasks while using up to 87.5% fewer trainable parameters.

Listen

Customizing large language models for specific tasks and diverse user preferences is essential for practical enterprise deployment. However, maintaining distinct customized model weights for numerous user-task combinations introduces substantial computational expenses during training, as well as significant operational and storage costs during post-training deployment. While parameter-efficient fine-tuning techniques like Low-Rank Adaptation (LoRA) reduce training burdens, their parameter overhead remains non-trivial as base models scale in depth and size.

The article demonstrates an approach called Tied-LoRA, which integrates weight tying and selective parameter training to dramatically reduce the number of trainable parameters in LoRA while maintaining high task performance. The authors evaluate various parameter configurations across five distinct natural language processing tasks—extractive question answering, dialogue summarization, commonsense reasoning, machine translation, and mathematical reasoning—using two base language models of different scales (a 2-billion parameter model and a 7-billion parameter model).

The evaluation reveals three major findings. First, a specific configuration termed TL6 achieves performance within 1% to 2% of standard LoRA on average, while utilizing only a small fraction of the parameters. In a 7-billion parameter translation task, TL6 matched and slightly exceeded LoRA's performance while using only 12.5% of the trainable parameters. Second, the performance gap between TL6 and standard LoRA narrowed when moving from the smaller 2-billion parameter model to the larger 7-billion parameter model, suggesting that larger foundation models benefit more from parameter sharing. Third, applying shared low-rank updates across all model layers substantially outperformed applying unshared LoRA updates to any individual layer, even when using the exact same parameter budget.

These findings indicate that organizations can drastically reduce storage, memory, and operational serving costs for specialized language models without suffering meaningful drops in accuracy, particularly on tasks that align with the base model's core language capabilities. However, complex mathematical reasoning tasks still favor standard LoRA due to the higher parameter capacity required for arithmetic operations.

Organizations serving multiple customized language model variants should consider adopting TL6 as a drop-in replacement for standard LoRA to lower deployment costs and storage footprint. Prior to full deployment, engineering teams should validate performance on complex analytical tasks and conduct pilot tests, as further empirical work on ultra-large base models (such as 70-billion parameter models) is still needed.

arXiv: 2311.09578
  • Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). This foundational work introduces Low-Rank Adaptation (LoRA), establishing the core low-rank parameterization and adapter architecture that Tied-LoRA directly extends through weight tying.
  • Paper: AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning, Qingru Zhang et al. (2023). This paper establishes adaptive budget allocation across transformer layers for LoRA, providing critical background on the uneven layer-wise parameter sensitivity addressed in Tied-LoRA.
  • Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This work formulates a unified mathematical taxonomy of parameter-efficient fine-tuning methods, clarifying the design space of modular hidden-state modifications underlying low-rank parameter sharing.
  • Paper: QLoRA: Efficient Finetuning of Quantized LLMs, Tim Dettmers et al. (2023). This study demonstrates how low-rank adapters interact with memory-compressed base models, establishing the standard baseline framework for parameter-efficient adaptation at scale.
  • Paper: Predicting Parameters in Deep Learning, Misha Denil et al. (2013). This paper demonstrates foundational parameter redundancy in deep networks through low-rank decompositions and static parameter sharing, motivating Tied-LoRA's weight-tying strategy.
  • Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). This work introduces modular parameter-efficient bottleneck adapter tuning for transformers, defining the architectural paradigm that subsequent low-rank fine-tuning methods build upon.
Cover for Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying

Abstract

We introduce Tied-LoRA, a novel paradigm leveraging weight tying and selective training to enhance the parameter efficiency of Low-rank Adaptation (LoRA). Our exploration encompasses different plausible combinations of parameter training and freezing, coupled with weight tying, aimed at identifying the optimal trade-off between performance and the count of trainable parameters. Across 5 diverse tasks and two foundational language models with different parameter counts, our experiments provide comprehensive insights into the inherent trade-offs between efficiency and performance.

Our findings reveal a specific Tied-LoRA configuration that distinguishes itself by showcasing comparable performance to LoRA across multiple tasks while utilizing only a fraction of the parameters employed by the standard LoRA method, particularly at elevated ranks. This underscores the efficacy of Tied-LoRA in achieving impressive results with significantly reduced model complexity.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Formulation
  • 2.2 Weight Tying
  • 2.3 Selective Training
  • 3 Experiments
  • 3.1 Tasks & Datasets
  • 3.2 Base Language Models
  • 3.3 Implementation Details
  • 4 Results
  • 4.1 Task-Dependent Optimal Rank
  • 4.2 Layer Selection Vs. Tied-LoRA
  • 4.3 Stability Across Ranks
  • 5 Related Work
  • 6 Conclusion & Future Work
  • Limitations
  • References

Knowls

  1. Knowl 1 — Tied-LoRA low-rank update formulation

    model/method

    Tied-LoRA adapts a frozen transformer by approximating the update to each layer’s joint query, key, and value projection. For layer ℓ∈{1,…,L}\ell \in \{1,\ldots,L\}, let Wℓ∈Rd×3dW_\ell \in \mathbb{R}^{d\times 3d} be the frozen projection, x∈Rdx\in\mathbb{R}^{d} the layer input, and zℓ∈R3dz_\ell\in\mathbb{R}^{3d} the output. The low-rank factors are A∈Rd×rA\in\mathbb{R}^{d\times r} and B∈Rr×3dB\in\mathbb{R}^{r\times 3d}, while uℓ∈Rru_\ell\in\mathbb{R}^{r} and vℓ∈R3dv_\ell\in\mathbb{R}^{3d} are scaling vectors with diagonal matrices Λuℓ\Lambda_{u_\ell} and Λvℓ\Lambda_{v_\ell}. The update is

    zℓ=Wℓx+αrΛvℓBΛuℓAx,z_\ell = W_\ell x + \frac{\alpha}{r}\Lambda_{v_\ell}B\Lambda_{u_\ell}Ax,

    where rr is the bottleneck rank and the method sets α=r\alpha=r, so the explicit scaling factor is one. Tied-LoRA shares the matrices AA and BB across all LL layers, while the scaling vectors may be trained, frozen, or layer-specific depending on the configuration. Because the adaptation path is linear, its update can be merged into the frozen projection after training and therefore does not add inference latency.

  2. Knowl 2 — Spectrum of Tied-LoRA training configurations

    model/method

    The paper defines eight configurations by choosing whether AA, BB, uu, and vv are trainable and whether AA and BB are shared across the LL transformer layers. Here dd is the hidden size, rr is the low-rank dimension, and the parameter count is the number of trainable scalar parameters.

    • LoRA trains untied layer-specific AA, BB, uu, and vv: 4Ldr4Ldr parameters.
    • VeRA freezes randomly initialized tied AA and BB and trains layer-specific uu and vv: L(r+3d)L(r+3d) parameters.
    • TL1 trains only tied AA while tied BB and both scaling vectors are frozen: drdr parameters.
    • TL2 trains tied AA plus layer-specific uu and vv, while tied BB is frozen: dr+L(r+3d)dr+L(r+3d) parameters.
    • TL3 trains only tied BB while tied AA and both scaling vectors are frozen: 3dr3dr parameters.
    • TL4 trains tied BB and a layer-specific rank-dimensional scaling vector, with tied AA and the other scaling vector frozen: (L+3d)r(L+3d)r parameters.
    • TL5 trains tied AA and tied BB, with both scaling vectors frozen: 4dr4dr parameters.
    • TL6 trains tied AA and tied BB together with layer-specific uu and vv: 4dr+L(r+3d)4dr+L(r+3d) parameters.

    The initialization follows the LoRA convention of random-normal AA, zero BB, and unit scaling vectors for configurations that train the corresponding factors. VeRA and the configurations that freeze a projection use random-normal frozen projections; TL6 initializes its tied projections randomly, with uu initialized to one and vv initialized to zero.

  3. Knowl 3 — Weight tying produces large parameter reductions

    empirical result

    For the 7-billion-parameter LLaMA-2 model with L=32L=32 layers and rank r=8r=8, standard LoRA requires approximately 4.24.2 million trainable parameters. Sharing the low-rank matrices across layers reduces this to approximately 131131 thousand parameters for the tied-matrix configuration, a 96.875%96.875\% reduction. VeRA, which additionally freezes the shared projections and trains layer-specific scaling vectors, gives a 90.6%90.6\% reduction relative to LoRA. The reduction grows with model depth because the tied matrices do not acquire a separate copy in every layer.

  4. Knowl 4 — TL6 is the strongest efficiency-performance trade-off

    data/table

    Across five task-customization benchmarks and both evaluated base models, TL6 was the strongest Tied-LoRA configuration and generally approached standard LoRA while training far fewer parameters. The following are the best scores found after varying the rank; parameter percentages are relative to the best LoRA parameter count for the same base model and task.

    • For LLaMA-2 7B, LoRA scored 40.76 Rouge-L on DialogSum, 32.75 EM on GSM8K, 91.97 accuracy on HellaSwag, 41.30 BLEU on IWSLT 2017, and 88.52 EM on SQuAD. TL6 scored 39.71 at rank 16 using 15.6% of LoRA’s parameters, 31.77 at rank 64 using 4.3%, 91.90 at rank 64 using 17.2%, 41.37 at rank 32 using 21.9%, and 88.49 at rank 4 using 43.8%, respectively.

    • For GPT-2B-001, LoRA scored 38.59 Rouge-L, 12.28 EM, 85.64 accuracy, 40.19 BLEU, and 83.58 EM on the same five tasks. TL6 scored 37.81 at rank 32 using 52.2% of LoRA’s parameters, 10.31 at rank 16 using 2.2%, 85.13 at rank 32 using 3.3%, 39.74 at rank 128 using 4.8%, and 83.56 at rank 64 using 10.8%, respectively.

    At the rank where LoRA is optimal for LLaMA-2 7B translation, TL6 obtains 41.33 BLEU at rank r=8r=8 using 12.5% of LoRA’s parameters, compared with LoRA’s 41.30 BLEU. The paper reports an average TL6 performance decline of 1.36% relative to LoRA for LLaMA-2 7B and 1.95% for GPT-2B-001.

  5. Knowl 5 — Evaluation protocol spans five tasks and two language models

    experimental setup

    The experiments evaluate task-specific customization rather than general instruction following. The five tasks and datasets are extractive question answering on SQuADv1, dialogue summarization on DialogSum, commonsense natural-language inference on HellaSwag, German-to-English translation on IWSLT 2017, and mathematical reasoning on GSM8K. The metrics are SQuAD exact match, DialogSum Rouge-L, HellaSwag accuracy, IWSLT BLEU, and GSM8K exact match. Because the official SQuAD test split lacks answers, the authors sample 4,800 training examples as a validation-based test set.

    The base models are NVIDIA GPT-2B-001 and Meta LLaMA-2 7B, both autoregressive models with a 4,096-token context. GPT-2B-001 was pretrained on 1.1 trillion multilingual tokens, whereas LLaMA-2 7B was pretrained on 2 trillion predominantly English tokens. Each configuration is trained for at most 2,000 steps with early stopping after 10 validation checks without improvement, using AdamW with weight decay 0.010.01, a cosine learning-rate schedule, 50 warm-up steps, and learning rates 10−410^{-4} and 10−510^{-5}. The tested ranks are r∈{2,4,8,16,32,64,128}r\in\{2,4,8,16,32,64,128\}. The global batch size is 256 and validation occurs every 30 steps, except for translation, which uses batch size 1,024 and validation every 60 steps. Predictions use greedy decoding with a 500-token limit, and no extensive hyperparameter search is performed.

  6. Knowl 6 — Optimal rank is task-dependent, with high-rank tasks favoring TL6

    empirical result

    The rank that maximizes LoRA performance varies substantially by task and is not monotonic with performance. For LLaMA-2 7B, the best LoRA ranks are 8 for DialogSum, 64 for GSM8K, 16 for HellaSwag, 8 for IWSLT 2017, and 2 for SQuAD. For GPT-2B-001, the corresponding ranks are 4, 64, 64, 128, and 32. Thus, increasing rank does not universally improve adaptation.

    Tied-LoRA is particularly competitive when standard LoRA requires a high rank. On LLaMA-2 7B GSM8K, TL6 reaches 31.77 EM at rank 64 while using 4.3% of LoRA’s parameters and scoring close to LoRA’s 32.75 EM. This behavior is important because tying makes the parameter savings more pronounced as rank increases: the tied projection cost scales with rr but does not carry the multiplicative factor of the number of layers.

  7. Knowl 7 — Tied-LoRA outperforms most single-layer LoRA controls at equal budget

    empirical result

    The authors compare TL5 with LoRA applied to only one LLaMA-2 7B transformer layer. Both methods use rank r=16r=16 and the same number of trainable parameters. Single-layer LoRA is tested at layers 1, 4, 8, 12, 16, 20, 24, 28, and 32. The scores for single-layer LoRA, followed by the TL5 score, are:

    • IWSLT BLEU: 37.94,38.99,39.47,38.68,38.10,35.33,33.24,28.90,22.4037.94, 38.99, 39.47, 38.68, 38.10, 35.33, 33.24, 28.90, 22.40; TL5: 41.3741.37.
    • SQuAD EM: 85.30,86.50,87.55,85.51,80.01,71.71,65.90,60.70,56.2285.30, 86.50, 87.55, 85.51, 80.01, 71.71, 65.90, 60.70, 56.22; TL5: 87.7387.73.
    • GSM8K EM: 13.04,19.33,19.56,14.93,11.67,6.14,4.92,3.56,2.8013.04, 19.33, 19.56, 14.93, 11.67, 6.14, 4.92, 3.56, 2.80; TL5: 27.0727.07.
    • HellaSwag accuracy: 77.70,88.36,87.92,85.75,85.46,77.44,27.96,49.52,47.7977.70, 88.36, 87.92, 85.75, 85.46, 77.44, 27.96, 49.52, 47.79; TL5: 91.7691.76.
    • DialogSum Rouge-L: 37.75,37.80,39.17,38.73,37.26,34.55,33.54,33.02,31.6437.75, 37.80, 39.17, 38.73, 37.26, 34.55, 33.54, 33.02, 31.64; TL5: 38.7338.73.

    Sharing one low-rank update across all layers is therefore substantially more effective than concentrating the same parameter budget in most individual layers. The control results also show that lower transformer layers, especially layers 4 and 8, are generally stronger single-layer locations than higher layers.

  8. Knowl 8 — TL6 is the most stable configuration across ranks

    empirical result

    When performance is averaged across the five tasks, nearly every Tied-LoRA configuration degrades as the rank increases. TL3 and TL4 show the sharpest high-rank declines. TL5 remains close to TL6 at typical ranks r=4r=4 through 1616, but also loses performance at larger ranks; VeRA displays a similar high-rank deterioration. TL6 is the exception: it maintains high performance across a broad rank range and remains the Tied-LoRA curve closest to standard LoRA for both the 7B and 2B base models.

  9. Knowl 9 — Efficiency depends on the task’s reliance on pretrained capability

    empirical result

    The experiments indicate that Tied-LoRA is most competitive on tasks that can exploit capabilities already present in the frozen base model, including commonsense NLI, extractive question answering, and summarization. Mathematical reasoning and arithmetic benefit more from LoRA’s larger trainable capacity, making GSM8K the clearest case where the parameter-heavy baseline retains an advantage. Translation is an important outlier: on LLaMA-2 7B, TL6 slightly exceeds LoRA despite using only 12.5% of LoRA’s parameters. The results therefore support strong efficiency claims for the tested tasks but do not establish that one Tied-LoRA configuration will generalize uniformly to new tasks.

  10. Knowl 10 — Scope limitations of the empirical evidence

    limitation

    The evaluation is limited to two base language models and five task datasets because of computational cost. Parameter-efficient fine-tuning is sensitive to both the base model and the customization task, and the paper reports that performance on an observed task did not reliably predict behavior on another task; in particular, the translation result was unexpectedly favorable to Tied-LoRA. The study therefore cautions against broad claims of task generalization. The authors identify evaluation on larger base models and weight tying for other parameter-efficient methods such as adapters and prefix tuning as directions for further work.

Coverage note — The full rank-by-rank scores for every configuration and every task were not reproduced; they were omitted to avoid repeating the same empirical pattern, while the defining configurations, key numerical comparisons, control experiment, rank analysis, stability findings, and stated limitations are retained.

References

  1. 1.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–9, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian Stüker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. 2017. Overview of the IWSLT 2017 evaluation campaign. In Proceedings of the 14th International Conference on Spoken Language Translation, pages 2–14, Tokyo, Japan. International Workshop on Spoken Language Translation.
  3. 3.Arnav Chavan, Zhuang Liu, Deepak Gupta, Eric Xing, and Zhiqiang Shen. 2023. One-for-all: Generalized lora for parameter-efficient fine-tuning. arXiv preprint arXiv:2306.07967.
  4. 4.Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang. 2021. DialogSum: A real-life scenario dialogue summarization dataset. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 5062–5074, Online. Association for Computational Linguistics.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  6. 6.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  7. 7.Wei Dong, Dawei Yan, Zhijun Lin, and Peng Wang. 2023. Efficient adaptation of large vision transformer via adapter re-composing. arXiv preprint arXiv:2310.06234.
  8. 8.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  9. 9.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  10. 10.Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017. Tying word vectors and word classifiers: A loss framework for language modeling. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  11. 11.Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki Markus Asano. 2023. Vera: Vector-based random matrix adaptation. arXiv preprint arXiv:2310.11454.
  12. 12.Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish. 2023. Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b. arXiv preprint arXiv:2310.20624.
  13. 13.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  14. 14.Chin-Yew Lin and Franz Josef Och. 2004. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 605–612, Barcelona, Spain.
  15. 15.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems, 35:1950–1965.
  16. 16.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023. Gpt understands, too. AI Open.
  17. 17.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  18. 18.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  19. 19.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503, Online. Association for Computational Linguistics.
  20. 20.Ofir Press and Lior Wolf. 2017. Using the output embedding to improve language models. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 157–163, Valencia, Spain. Association for Computational Linguistics.
  21. 21.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  22. 22.Xianghui Sun, Yunjie Ji, Baochang Ma, and Xiangang Li. 2023. A comparative study between full-parameter and lora-based fine-tuning on chinese instruction data for instruction following large language model. arXiv preprint arXiv:2304.08109.
  23. 23.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  24. 24.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy. Association for Computational Linguistics.
  25. 25.Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023. Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512.

Citation

MLA
Renduchintala, A., et al. “Tied-LoRA: Enhancing Parameter Efficiency of LoRA with Weight Tying”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 8694–705, https://doi.org/10.18653/v1/2024.naacl-long.481.
APA
Renduchintala, A., Konuk, T., & Kuchaiev, O. (2024). Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8694–8705. https://doi.org/10.18653/v1/2024.naacl-long.481
Chicago
Renduchintala, A., T. Konuk, and O. Kuchaiev. 2024. “Tied-LoRA: Enhancing Parameter Efficiency of LoRA with Weight Tying”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8694–8705. https://doi.org/10.18653/v1/2024.naacl-long.481.
Harvard
Renduchintala, A., Konuk, T. and Kuchaiev, O. (2024) “Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8694–8705. Available at: https://doi.org/10.18653/v1/2024.naacl-long.481.
Vancouver
1. Renduchintala A, Konuk T, Kuchaiev O (2024) Tied-LoRA: Enhancing parameter efficiency of LoRA with Weight Tying. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 8694–8705

BibTeX

@inproceedings{renduchintala-etal-2024-tied,
    title = "Tied-{L}o{RA}: Enhancing parameter efficiency of {L}o{RA} with Weight Tying",
    author = "Renduchintala, Adithya  and
      Konuk, Tugrul  and
      Kuchaiev, Oleksii",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.481/",
    doi = "10.18653/v1/2024.naacl-long.481",
    pages = "8694--8705"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/