Full Parameter Fine-tuning for Large Language Models with Limited Resources
Kai LvYuqing YangTengxiao LiuQipeng GuoXipeng Qiu
Proposes a fused optimizer called LOMO that drastically cuts training memory to just over ten percent of standard DeepSpeed solutions, enabling full-parameter fine-tuning of a 65-billion-parameter language model on consumer GPUs.
Adapting large language models to specialized tasks through full-parameter fine-tuning typically demands immense computing resources, such as high-end graphics processing unit clusters. Standard optimizers like Adam store massive amounts of intermediate calculation states and gradients, creating hardware barriers that prevent smaller research labs and organizations from customizing large-scale artificial intelligence models.
The article demonstrates and evaluates a memory-efficient optimization approach called LOw-Memory Optimization, or LOMO. The objective is to enable full-parameter fine-tuning of multi-billion-parameter language models on budget-friendly, consumer-grade hardware without compromising the adaptation process.
To achieve this, the authors replace memory-heavy optimizers with basic stochastic gradient descent and fuse gradient computation directly with parameter updates during the backpropagation process. This eliminates the need to store intermediate optimizer states and minimizes gradient memory overhead to that of just a single parameter tensor. The researchers integrated this method with activation checkpointing and mixed-precision stabilization techniques, testing it on LLaMA models ranging from 7 billion to 65 billion parameters across standard language understanding benchmarks using consumer RTX 3090 graphics cards.
The findings show that LOMO reduces total GPU memory usage to approximately 10.8% of standard industry solutions. When fine-tuning a 7-billion parameter model, memory usage plummeted from 102.2 gigabytes under standard configurations to 14.58 gigabytes, boosting throughput on a single card by roughly elevenfold compared to traditional setups due to minimized communication overhead. Furthermore, the approach successfully enabled the full fine-tuning of a 65-billion parameter model on a single machine equipped with eight 24-gigabyte consumer graphics cards. On downstream task evaluations, LOMO consistently outperformed zero-shot baselines and generally surpassed parameter-efficient techniques like Low-Rank Adaptation, achieving an 89.9% average accuracy across test tasks on the 65-billion parameter scale.
These results demonstrate that organizations can conduct full model adaptation without purchasing high-end enterprise computing clusters, significantly lowering capital expenditures, operational barriers, and infrastructure timelines. The study also confirms that basic first-order optimization is viable for fine-tuning because large pre-trained models already possess relatively smooth parameter landscapes, challenging the long-standing assumption that complex adaptive optimizers are essential for fine-tuning.
Organizations seeking cost-effective customization should evaluate LOMO as an alternative or complementary method to parameter-efficient tuning. Because parameters now constitute the vast majority of remaining memory overhead, subsequent technical initiatives should explore parameter quantization to lower hardware requirements even further. Additional development is also recommended to streamline gradient normalization into a single backward pass to improve overall training speed.
Decision-makers should note that evaluations were conducted using a limited sample size of one thousand training examples per task across a subset of benchmark datasets, and tests were not run on high-end enterprise accelerators like A100 systems. While confidence in the memory-saving mechanism is high, training speed may experience slight latency in configurations that require two backward passes for gradient clipping and dynamic loss scaling.
- Paper: ZeRO: Memory optimizations Toward Training Trillion Parameter Models, Samyam Rajbhandari et al. (2020). ZeRO establishes the model-state partitioning and memory-accounting framework that clarifies the resource bottleneck LOMO redesigns for single-machine full-parameter fine-tuning.
- Paper: Adafactor: Adaptive Learning Rates with Sublinear Memory Cost, Noam Shazeer et al. (2018). Adafactor provides the foundational sublinear-memory optimizer perspective against which LOMO’s fused gradient computation and update can be understood.
- Paper: GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection, Jiawei Zhao et al. (2024). GaLore introduces low-memory full-parameter optimization through gradient projection, making LOMO’s alternative approach to reducing optimizer-state memory easier to assess.
- Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). Training Deep Nets with Sublinear Memory Cost supplies the activation-recomputation and memory–compute trade-offs relevant to LOMO’s broader limited-resource training objective.
- Paper: ZeRO-Offload: Democratizing Billion-Scale Model Training, Jie Ren et al. (2021). ZeRO-Offload shows how moving optimizer work and state across devices can democratize large-model training, providing an important systems baseline for LOMO’s on-device memory reduction.
- Paper: QLoRA: Efficient Finetuning of Quantized LLMs, Tim Dettmers et al. (2023). QLoRA defines the parameter-efficient and quantized fine-tuning baseline that motivates the source’s effort to make full-parameter tuning feasible instead.
- Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). LoRA frames the contrast between parameter-efficient adaptation and updating all model parameters, a central distinction in the source’s contribution.
- Paper: Flora: Low-Rank Adapters Are Secretly Gradient Compressors, Yongchang Hao et al. (2024). FLORA continues the low-memory full-update direction by replacing LOMO’s fused-update mechanism with resampled random gradient compression for sublinear optimizer memory.
- Paper: LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning, Rui Pan et al. (2024). LISA extends resource-constrained full-parameter fine-tuning through layerwise importance sampling, applying selective updates to reach much larger models under tight memory budgets.
- Paper: Adam-mini: Use Fewer Learning Rates To Gain More, Yushun Zhang et al. (2025). Adam-mini advances optimizer-state reduction beyond LOMO by exploiting parameter-block structure to retain adaptive training with substantially less memory.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). LQ-LoRA carries the memory-efficient adaptation agenda into sub-four-bit quantized regimes, extending the practical resource savings enabled by the source’s full-tuning focus.
