Full Parameter Fine-tuning for Large Language Models with Limited Resources

Kai LvYuqing YangTengxiao LiuQipeng GuoXipeng Qiu

article2024ACL249 citations

Proposes a fused optimizer called LOMO that drastically cuts training memory to just over ten percent of standard DeepSpeed solutions, enabling full-parameter fine-tuning of a 65-billion-parameter language model on consumer GPUs.

Listen

Adapting large language models to specialized tasks through full-parameter fine-tuning typically demands immense computing resources, such as high-end graphics processing unit clusters. Standard optimizers like Adam store massive amounts of intermediate calculation states and gradients, creating hardware barriers that prevent smaller research labs and organizations from customizing large-scale artificial intelligence models.

The article demonstrates and evaluates a memory-efficient optimization approach called LOw-Memory Optimization, or LOMO. The objective is to enable full-parameter fine-tuning of multi-billion-parameter language models on budget-friendly, consumer-grade hardware without compromising the adaptation process.

To achieve this, the authors replace memory-heavy optimizers with basic stochastic gradient descent and fuse gradient computation directly with parameter updates during the backpropagation process. This eliminates the need to store intermediate optimizer states and minimizes gradient memory overhead to that of just a single parameter tensor. The researchers integrated this method with activation checkpointing and mixed-precision stabilization techniques, testing it on LLaMA models ranging from 7 billion to 65 billion parameters across standard language understanding benchmarks using consumer RTX 3090 graphics cards.

The findings show that LOMO reduces total GPU memory usage to approximately 10.8% of standard industry solutions. When fine-tuning a 7-billion parameter model, memory usage plummeted from 102.2 gigabytes under standard configurations to 14.58 gigabytes, boosting throughput on a single card by roughly elevenfold compared to traditional setups due to minimized communication overhead. Furthermore, the approach successfully enabled the full fine-tuning of a 65-billion parameter model on a single machine equipped with eight 24-gigabyte consumer graphics cards. On downstream task evaluations, LOMO consistently outperformed zero-shot baselines and generally surpassed parameter-efficient techniques like Low-Rank Adaptation, achieving an 89.9% average accuracy across test tasks on the 65-billion parameter scale.

These results demonstrate that organizations can conduct full model adaptation without purchasing high-end enterprise computing clusters, significantly lowering capital expenditures, operational barriers, and infrastructure timelines. The study also confirms that basic first-order optimization is viable for fine-tuning because large pre-trained models already possess relatively smooth parameter landscapes, challenging the long-standing assumption that complex adaptive optimizers are essential for fine-tuning.

Organizations seeking cost-effective customization should evaluate LOMO as an alternative or complementary method to parameter-efficient tuning. Because parameters now constitute the vast majority of remaining memory overhead, subsequent technical initiatives should explore parameter quantization to lower hardware requirements even further. Additional development is also recommended to streamline gradient normalization into a single backward pass to improve overall training speed.

Decision-makers should note that evaluations were conducted using a limited sample size of one thousand training examples per task across a subset of benchmark datasets, and tests were not run on high-end enterprise accelerators like A100 systems. While confidence in the memory-saving mechanism is high, training speed may experience slight latency in configurations that require two backward passes for gradient clipping and dynamic loss scaling.

Cover for Full Parameter Fine-tuning for Large Language Models with Limited Resources

Abstract

Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but demand massive GPU resources for training. Lowering the threshold for LLMs training would encourage greater participation from researchers, benefiting both academia and society. While existing approaches have focused on parameter-efficient fine-tuning, which tunes or adds a small number of parameters, few have addressed the challenge of tuning the full parameters of LLMs with limited resources. In this work, we propose a new optimizer, LOw-Memory Optimization (LOMO), which fuses the gradient computation and the parameter update in one step to reduce memory usage. By integrating LOMO with existing memory saving techniques, we reduce memory usage to 10.8% compared to the standard approach (DeepSpeed solution). Consequently, our approach enables the full parameter fine-tuning of a 65B model on a single machine with 8×RTX 3090, each with 24GB memory.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Rethink the Functionality of Optimizer
  • 3.1.1 Using SGD
  • 3.1.2 Implicit Batch Size
  • 3.2 LOMO: LOw-Memory Optimization
  • 3.3 Stabilize Training with LOMO
  • 3.3.1 Alternatives to Gradient Normalization and Clipping
  • 3.3.2 Mitigating Precision Degradation
  • 4 Experiment
  • 4.1 Memory Profile
  • 4.2 Throughput
  • 4.3 Downstream Performance
  • 4.3.1 Main results
  • 4.3.2 LoRA with LOMO
  • 5 Conclusion
  • Limitations
  • Ethics statement
  • Acknowledgments
  • References
  • A Hyperparameters
  • B Training Dynamics

Knowls

  1. Knowl 1 — LOMO Fused Backward Algorithm

    algorithm

    LOw-Memory Optimization (LOMO) updates parameters in place during the backward pass via backward hooks, avoiding storing gradients for all parameters simultaneously.

    Input: Model f(⋅)f(\cdot) with LL layers and parameters θ∈Rp\theta \in \mathbb{R}^p, learning rate α\alpha, maximum steps TT, training dataset D\mathcal{D}, loss function L\mathcal{L}
    Output: Fine-tuned model parameters θ\theta
    for t=1,…,Tt = 1, \dots, T do
        Sample batch B=(x,y)⊂DB = (x, y) \subset \mathcal{D}
        y^←f(x,θ)\hat{y} \leftarrow f(x, \theta)
        ℓ←L(y,y^)\ell \leftarrow \mathcal{L}(y, \hat{y})
        for l=L,…,1l = L, \dots, 1 do
            θl←[θi for θi∈layer l]\theta_l \leftarrow [\theta_i \text{ for } \theta_i \in \text{layer } l]
            gl←∂ℓ∂θlg_l \leftarrow \frac{\partial \ell}{\partial \theta_l}
            θl←θl−α⋅gl\theta_l \leftarrow \theta_l - \alpha \cdot g_l
            gl←Noneg_l \leftarrow \text{None}
        end for
    end for
    return θ\theta

    In standard deep learning frameworks, backpropagation accumulates gradient tensors for all model parameters before an optimizer step updates the weights, demanding O(n)O(n) gradient memory where nn is total parameter count. LOMO replaces the separate optimization step by executing parameter updates immediately as gradients for each tensor are computed along the reverse topological order of the computation graph. PyTorch hook functions are registered to trigger immediately upon gradient calculation for each parameter tensor θl\theta_l. Once θl\theta_l is updated via the SGD update rule θl←θl−α⋅gl\theta_l \leftarrow \theta_l - \alpha \cdot g_l, the gradient buffer glg_l is freed (gl←Noneg_l \leftarrow \text{None}). This reduces the peak gradient memory requirement to O(1)O(1), which corresponds only to the size of the single largest parameter tensor in the model.

  2. Knowl 2 — Implicit Batch Size Effect in LLM Fine-Tuning with SGD

    theoretical result

    When fine-tuning large language models with a smooth loss surface using stochastic gradient descent (SGD), sequential updates on individual samples approximate a mini-batch update with an implicitly enlarged batch size.

    Let f(⋅,θ)f(\cdot, \theta) denote a neural network parameterized by θ∈Rp\theta \in \mathbb{R}^p, L\mathcal{L} denote the loss function, and di,djd_i, d_j denote two training samples. A single simultaneous SGD update on the two-sample batch with learning rate α\alpha is: θ′=θ−α[∇L(di,f(di,θ))+∇L(dj,f(dj,θ))]\theta' = \theta - \alpha \left[ \nabla \mathcal{L}(d_i, f(d_i, \theta)) + \nabla \mathcal{L}(d_j, f(d_j, \theta)) \right]

    In contrast, performing two sequential SGD updates on sample did_i followed by sample djd_j yields: θ1=θ−α∇L(di,f(di,θ))\theta_1 = \theta - \alpha \nabla \mathcal{L}(d_i, f(d_i, \theta)) θ2=θ1−α∇L(dj,f(dj,θ1))\theta_2 = \theta_1 - \alpha \nabla \mathcal{L}(d_j, f(d_j, \theta_1))

    Applying the differential mean value theorem to L(dj,f(dj,θ1))\mathcal{L}(d_j, f(d_j, \theta_1)) around θ\theta gives: L(dj,f(dj,θ1))=L(dj,f(dj,θ))+∇L(dj,ξ)(f(dj,θ1)−f(dj,θ))\mathcal{L}(d_j, f(d_j, \theta_1)) = \mathcal{L}(d_j, f(d_j, \theta)) + \nabla \mathcal{L}(d_j, \xi) \left( f(d_j, \theta_1) - f(d_j, \theta) \right) for some intermediate point ξ\xi between f(dj,θ)f(d_j, \theta) and f(dj,θ1)f(d_j, \theta_1). Expanding θ2\theta_2 yields: θ2=θ−α[∇L(di,f(di,θ))+∇L(dj,f(dj,θ))]−α∇[∇L(dj,ξ)(f(dj,θ1)−f(dj,θ))]\theta_2 = \theta - \alpha \left[ \nabla \mathcal{L}(d_i, f(d_i, \theta)) + \nabla \mathcal{L}(d_j, f(d_j, \theta)) \right] - \alpha \nabla \left[ \nabla \mathcal{L}(d_j, \xi) \left( f(d_j, \theta_1) - f(d_j, \theta) \right) \right]

    The discrepancy between the sequential update θ2\theta_2 and batch update θ′\theta' is the residual term: −α∇[∇L(dj,ξ)(f(dj,θ1)−f(dj,θ))]-\alpha \nabla \left[ \nabla \mathcal{L}(d_j, \xi) \left( f(d_j, \theta_1) - f(d_j, \theta) \right) \right]

    Under the assumption that the loss surface of pre-trained LLMs is sufficiently smooth (small curvature) on natural language tasks, this second-order difference term is negligible. Thus, sequential SGD steps closely approximate larger batch updates, contributing to the training stability of SGD in fine-tuning large models even when SGD struggles on smaller models.

  3. Knowl 3 — Theoretical Justifications for SGD in LLM Fine-Tuning

    assumption

    While stochastic gradient descent (SGD) traditionally faces challenges with large curvature, suboptimal local minima, and saddle points when training deep networks from scratch, these challenges are mitigated during full-parameter fine-tuning of large language models (LLMs) due to three domain-specific properties:

    1. Smoother Loss Surface: The parameter loss landscape of pre-trained LLMs on natural language tasks has lower curvature compared to training from scratch or synthetic tasks. Small perturbations to pre-trained weights cause minimal loss fluctuations, mitigating curvature issues.
    2. Sufficiency of Local Optima: The objective of fine-tuning is task adaptation within the neighborhood of pre-trained weights rather than global optimization over the entire parameter space. Because downstream fine-tuning data is limited compared to pre-training corpora, local optima near the pre-trained weights provide strong task performance without needing to reach distant global minima.
    3. Distant Saddle Points: Pre-trained LLM checkpoints reside in loss valleys rather than on ridge structures where saddle points typically occur. Because fine-tuning updates do not deviate far from initial valley representations (especially for models pre-trained on diverse instructions), the optimization trajectory avoids high-loss saddle points.
  4. Knowl 4 — Stabilizing Mixed-Precision LOMO with Two-Pass Backpropagation

    model/method

    Because LOMO immediately updates weights during backpropagation, global gradient properties (such as the global gradient ℓ2\ell_2-norm and floating-point overflow status) are not known at the time an individual layer's gradient is computed. To preserve training stability and mixed-precision accuracy, two alternative mechanisms are used:

    1. Gradient Clipping Alternatives:
      • Value-based Clipping: Gradients are clipped element-wise to a fixed threshold rather than by global norm. This operates in a single backward pass without memory overhead and is effective for small to medium learning rates (α≤1×10−3\alpha \le 1\times 10^{-3}).
      • Two-Pass Gradient Normalization: In scenarios requiring exact global ℓ2\ell_2-norm normalization, training executes two backward passes per step. The first backward pass computes and accumulates the squared norm across all parameter gradients. The second backward pass computes the gradients again and immediately performs the scaled update.
    2. Mitigating Precision Degradation in FP16:
      • Dynamic Loss Scaling: To prevent FP16 underflow, the loss is scaled before the backward pass. A two-pass procedure checks for gradient overflow in the first pass; if an overflow is detected, the optimizer skips the update and halves the scale factor; otherwise, the second pass updates parameters and periodically doubles the scale factor.
      • Full-Precision Transitions: During gradient normalization, scaling, and parameter update arithmetic, parameters and gradients are temporarily cast to 32-bit floating point (FP32) to eliminate rounding and precision degradation before storing updated weights in FP16.
  5. Knowl 5 — Memory Usage Breakdown of LOMO on LLaMA-7B

    data/table

    When fine-tuning LLaMA-7B (sequence length 512, batch size 8), LOMO drastically reduces total GPU memory footprint compared to AdamW and standard SGD, both with and without Activation Checkpointing (AC).

    Optimizer AC Params (GB) Gradients (GB) Optim States (GB) Activations (GB) Total Memory (GB)
    AdamW ×\times 12.55 12.55 75.31 45.61 147.02
    AdamW ✓\checkmark 12.55 12.55 75.31 1.79 102.20
    SGD ×\times 12.55 12.55 25.10 45.61 96.81
    SGD ✓\checkmark 12.55 12.55 25.10 1.79 51.99
    LOMO ×\times 12.55 0.24 0.00 45.61 59.40
    LOMO ✓\checkmark 12.55 0.24 0.00 1.79 14.58

    Under FP16 mixed-precision training:

    • AdamW with AC requires 102.20 GB memory, where optimizer states (FP32 master weights, momentum, variance) account for 73.7% (75.31 GB) of total memory.
    • SGD with AC eliminates momentum and variance buffers, reducing optimizer states to 25.10 GB (master weights) and total memory to 51.99 GB.
    • LOMO with AC fuses gradient calculation and parameter updating, eliminating all persistent optimizer states (0.00 GB) and reducing gradient memory from 12.55 GB to 0.24 GB (the single largest parameter tensor). This achieves a total memory footprint of 14.58 GB (10.8% of the un-checkpointed AdamW baseline and 14.3% of AdamW with AC), allowing full parameter fine-tuning of a 7B model within a single 24GB GPU.
  6. Knowl 6 — Throughput and Scalability of LOMO across LLaMA Model Scales

    data/table

    Full parameter fine-tuning throughput and peak memory per GPU evaluated on a single server equipped with 8×RTX 30908 \times \text{RTX 3090} GPUs (24GB VRAM each, PCIe interconnection) using sequence length 1024 and per-GPU batch size 1 with ZeRO-3 parameter partitioning:

    Model Params Optimizer Hardware Peak Memory / GPU (GB) Throughput (TGS)
    7B AdamW 8×RTX 30908 \times \text{RTX 3090} 15.76 67.37
    7B SGD 8×RTX 30908 \times \text{RTX 3090} 9.49 69.66
    7B LOMO 1×RTX 30901 \times \text{RTX 3090} 13.61 769.92
    13B SGD 8×RTX 30908 \times \text{RTX 3090} 15.74 32.51
    13B LOMO 2×RTX 30902 \times \text{RTX 3090} 15.92 66.19
    30B LOMO 4×RTX 30904 \times \text{RTX 3090} 19.78 11.61
    65B LOMO 8×RTX 30908 \times \text{RTX 3090} 19.18 4.93

    Throughput is measured in tokens processed per GPU per second (TGS).

    • For LLaMA-7B, LOMO fits on a single GPU without inter-GPU communication overhead, achieving 769.92 TGS—an ≈11×\approx 11\times throughput increase compared to 8-GPU distributed AdamW (67.37 TGS) and SGD (69.66 TGS).
    • For LLaMA-13B, AdamW cannot run on 8×24GB8 \times 24\text{GB} GPUs due to out-of-memory (OOM), whereas LOMO trains on only 2 GPUs with 66.19 TGS (double the throughput of 8-GPU SGD).
    • For LLaMA-30B, SGD encounters OOM on 8 GPUs, while LOMO trains successfully on 4 GPUs.
    • For LLaMA-65B, LOMO enables full parameter fine-tuning on 8×RTX 30908 \times \text{RTX 3090} GPUs (19.18 GB peak allocated memory per GPU) at 4.93 TGS, completing 1,000 samples of 512 tokens in approximately 3.6 hours.
  7. Knowl 7 — SuperGLUE Downstream Fine-Tuning Performance Across Model Scales

    data/table

    Downstream performance comparison of Zero-shot, LoRA, and LOMO full-parameter fine-tuning across LLaMA model scales (7B, 13B, 30B, 65B) on six SuperGLUE classification tasks using 1,000 sampled training examples and 1,000 validation evaluation examples:

    Method Params RTE BoolQ WSC WIC MultiRC COPA Avg.
    Zero-shot 7B 57.0 66.5 36.5 49.7 42.3 85.0 56.2
    LoRA 7B 85.9 85.2 64.4 65.5 84.8 87.0 78.8
    LOMO 7B 86.6 87.5 66.4 71.2 84.0 89.0 80.8
    Zero-shot 13B 60.6 65.0 36.5 49.5 43.4 88.0 57.2
    LoRA 13B 89.9 87.1 63.5 69.9 86.1 92.0 81.4
    LOMO 13B 89.9 87.3 75.0 74.3 85.7 93.0 84.2
    Zero-shot 30B 53.4 74.6 36.5 50.0 46.9 89.0 58.4
    LoRA 30B 91.0 89.7 83.7 74.0 87.0 93.0 86.4
    LOMO 30B 92.8 89.3 85.6 74.1 87.9 93.0 87.1
    Zero-shot 65B 59.6 73.6 44.2 51.3 48.3 91.0 61.3
    LoRA 65B 93.1 90.9 88.5 74.5 90.0 97.0 89.0
    LOMO 65B 93.9 90.7 92.3 75.4 89.9 97.0 89.9

    Evaluation metric is Accuracy (%). Key findings include:

    • LOMO consistently improves over the Zero-shot baseline across all model sizes (e.g., +27.0+27.0 points average accuracy gain on LLaMA-13B).
    • LOMO outperforms parameter-efficient fine-tuning via LoRA on average accuracy across all evaluated model sizes: +2.0%+2.0\% on 7B (80.880.8 vs 78.878.8), +2.8%+2.8\% on 13B (84.284.2 vs 81.481.4), +0.7%+0.7\% on 30B (87.187.1 vs 86.486.4), and +0.9%+0.9\% on 65B (89.989.9 vs 89.089.0).
    • Performance gains are especially pronounced on tasks requiring deeper language reasoning such as WSC (+11.5+11.5 points over LoRA on 13B) and WIC (+4.4+4.4 points over LoRA on 13B).
  8. Knowl 8 — Complementary Integration of LoRA and LOMO

    empirical result

    LoRA and LOMO operate on orthogonal mechanisms: LOMO fine-tunes pre-trained backbone model weights with low memory overhead, whereas LoRA freezes backbone weights and trains added low-rank adapter modules. Combining both techniques—applying LOMO to the pre-trained model weights while simultaneously training LoRA modules with AdamW—yields additive accuracy improvements on downstream tasks.

    On LLaMA-13B with 1,000 training examples:

    • BoolQ Accuracy:
      • Baseline LoRA without backbone tuning achieves 87.3% (r=1r=1), 87.1% (r=2r=2), 87.5% (r=4r=4), and 87.5% (r=8r=8).
      • Combining LoRA + LOMO increases accuracy to 88.2% (r=1r=1), 87.5% (r=2r=2), 88.0% (r=4r=4), and 88.1% (r=8r=8), outperforming LoRA alone at every attention rank r∈{1,2,4,8}r \in \{1, 2, 4, 8\} as well as LOMO alone (87.7%).
    • MultiRC Accuracy:
      • Baseline LoRA achieves 85.5% (r=1r=1), 86.1% (r=2r=2), 86.1% (r=4r=4), and 86.0% (r=8r=8).
      • Combining LoRA + LOMO increases accuracy to 86.3% (r=1r=1), 86.3% (r=2r=2), 86.7% (r=4r=4), and 86.3% (r=8r=8).
  9. Knowl 9 — Computational Overhead of Exact Gradient Normalization in LOMO

    limitation

    A primary limitation of LOMO is the computational overhead associated with exact gradient normalization and dynamic loss scaling. Because LOMO clears gradient tensors immediately after updating each parameter to maintain O(1)O(1) gradient memory, computing the global ℓ2\ell_2 gradient norm ∑i∥gi∥2\sqrt{\sum_i \|g_i\|^2} or checking global FP16 overflow across all layers requires executing an initial full backward pass over the network prior to the second parameter-updating backward pass. Although this two-pass strategy maintains constant gradient memory, it introduces extra computational cost, effectively doubling backward pass latency when exact global gradient normalization or dynamic overflow checking is active.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174.
  2. 2.Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. 2018. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 793–802. PMLR.
  3. 3.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 2924–2936. Association for Computational Linguistics.
  4. 4.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The PASCAL recognising textual entailment challenge. In Machine Learning Challenges, Evaluating Predictive Uncertainty, Visual Object Classification and Recognizing Textual Entailment, First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers, volume 3944 of Lecture Notes in Computer Science, pages 177–190. Springer.
  5. 5.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022. 8-bit optimizers via block-wise quantization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  6. 6.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, Jing Yi, Weilin Zhao, Xiaozhi Wang, Zhiyuan Liu, Hai-Tao Zheng, Jianfei Chen, Yang Liu, Jie Tang, Juanzi Li, and Maosong Sun. 2022. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. CoRR, abs/2203.06904.
  7. 7.Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019. Visualizing and understanding the effectiveness of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 4141–4150. Association for Computational Linguistics.
  8. 8.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  9. 9.Kenji Kawaguchi, Jiaoyang Huang, and Leslie Pack Kaelbling. 2019. Every local minimum value is the global minimum value of induced model in dropconvex machine learning. Neural Comput., 31(12):2293–2323.
  10. 10.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 252–262. Association for Computational Linguistics.
  11. 11.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  12. 12.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Principles of Knowledge Representation and Reasoning: Proceedings of the Thirteenth International Conference, KR 2012, Rome, Italy, June 10-14, 2012. AAAI Press.
  13. 13.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 4582–4597. Association for Computational Linguistics.
  14. 14.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  15. 15.Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. 2023. Fine-tuning language models with just forward passes. CoRR, abs/2305.17333.
  16. 16.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed precision training. In International Conference on Learning Representations.
  17. 17.Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. 2021. Efficient large-scale language model training on gpu clusters using megatron-lm. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15.
  18. 18.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in pytorch.
  19. 19.Mohammad Taher Pilehvar and José Camacho-Collados. 2019. Wic: the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 1267–1273. Association for Computational Linguistics.
  20. 20.Bharadwaj Pudipeddi, Maral Mesmakhosroshahi, Jinwen Xi, and Sujeeth Bharadwaj. 2020. Training large neural networks with constant memory using a new execution algorithm. arXiv preprint arXiv:2002.05645.
  21. 21.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory optimizations toward training trillion parameter models. SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16.
  22. 22.Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–14.
  23. 23.Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, and Dong Li. 2021a. Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 598–611.
  24. 24.Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021b. Zero-offload: Democratizing billion-scale model training. USENIX Annual Technical Conference, pages 551–564.
  25. 25.Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler. 2016. vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 1–13.
  26. 26.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In Logical Formalizations of Commonsense Reasoning, Papers from the 2011 AAAI Spring Symposium, Technical Report SS-11-06, Stanford, California, USA, March 21-23, 2011. AAAI.
  27. 27.Sebastian Ruder. 2016. An overview of gradient descent optimization algorithms. CoRR, abs/1609.04747.
  28. 28.Shiliang Sun, Zehui Cao, Han Zhu, and Jing Zhao. 2020a. A survey of optimization methods from a machine learning perspective. IEEE Trans. Cybern., 50(8):3668–3681.
  29. 29.Xianghui Sun, Yunjie Ji, Baochang Ma, and Xiangang Li. 2023. A comparative study between full-parameter and lora-based fine-tuning on chinese instruction data for instruction following large language model. CoRR, abs/2304.08109.
  30. 30.Xiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni, Ankur Agrawal, Xiaodong Cui, Swagath Venkataramani, Kaoutar El Maghraoui, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. 2020b. Ultra-low precision 4-bit training of deep neural networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  31. 31.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  32. 32.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 3261–3275.
  33. 33.Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Song, Zenglin Xu, and Tim Kraska. 2018. Superneurons: dynamic gpu memory management for training deep neural networks. ACM SIGPLAN Notices, 53:41–53.
  34. 34.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent abilities of large language models. Trans. Mach. Learn. Res., 2022.
  35. 35.Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian. 2024. Galore: Memory-efficient LLM training by gradient low-rank projection. CoRR, abs/2403.03507.

Citation

MLA
Lv, K., et al. “Full Parameter Fine-tuning for Large Language Models with Limited Resources”. arXiv, 2023, http://arxiv.org/abs/2306.09782v2.
APA
Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., & Qiu, X. (2023). Full Parameter Fine-tuning for Large Language Models with Limited Resources. arXiv. http://arxiv.org/abs/2306.09782v2
Chicago
Lv, K., Y. Yang, T. Liu, Q. Gao, Q. Guo, and X. Qiu. 2023. “Full Parameter Fine-tuning for Large Language Models with Limited Resources”. arXiv. http://arxiv.org/abs/2306.09782v2.
Harvard
Lv, K. et al. (2023) “Full Parameter Fine-tuning for Large Language Models with Limited Resources”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.09782v2.
Vancouver
1. Lv K, Yang Y, Liu T, Gao Q, Guo Q, Qiu X (2023) Full Parameter Fine-tuning for Large Language Models with Limited Resources. arXiv

BibTeX

@article{lv2023full,
  title = {Full Parameter Fine-tuning for Large Language Models with Limited Resources},
  author = {Lv, Kai and Yang, Yuqing and Liu, Tengxiao and Gao, Qinghui and Guo, Qipeng and Qiu, Xipeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.09782v2},
  eprint = {2306.09782}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/