RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation

Mahdi NikdanSoroush TabeshElvir CrncevicDan Alistarh

article2024ICML51 citations

Proposes RoSA, a parameter-efficient fine-tuning method combining parallel low-rank and sparse adapters with custom GPU kernels to match full fine-tuning performance on challenging language tasks under tight parameter and memory constraints.

Listen

Adapting large language models to specific enterprise tasks typically requires fine-tuning, but adjusting all model parameters demands immense memory and compute resources that are often cost-prohibitive. Existing parameter-efficient fine-tuning methods like Low-Rank Adaptation (LoRA) cut resource costs by training compact adapter layers, but they often suffer a noticeable drop in accuracy on complex, specialized tasks such as mathematical reasoning and code generation.

The article introduces and evaluates Robust Adaptation (RoSA), a parameter-efficient fine-tuning method designed to match the accuracy of full-parameter fine-tuning while retaining the low memory footprint of lightweight adapters. Inspired by robust principal component analysis, RoSA operates on the premise that weight updates during fine-tuning are best approximated as a combination of a low-rank matrix and a sparse matrix to capture important outlier updates.

The researchers evaluated RoSA by fine-tuning the 7-billion-parameter LLaMA-2 model across three challenging benchmarks: math word problems (GSM8k), dialogue generation (ViGGO), and text-to-SQL query generation. They tested RoSA against full fine-tuning, standard LoRA, pure sparse adaptation, and alternative hybrid baselines across multiple parameter budgets ranging from 40 million to 160 million trainable parameters. To make sparse computation practical on graphics processing units (GPUs), the team engineered custom sparse GPU kernels tailored to the specific sparsity patterns observed during training.

The evaluation produced four central findings. First, RoSA consistently outperformed standard LoRA and pure sparse adaptation across challenging tasks at identical parameter budgets. Second, RoSA matched or surpassed full fine-tuning accuracy—for example, reaching up to 97.3% on ViGGO compared to 95.0% for full fine-tuning, and achieving comparable or superior performance on GSM8k—while training 40 to 100 times fewer parameters. Third, when combined with 4-bit base model quantization (QRoSA), the method maintained high accuracy while reducing GPU memory consumption below 12 GB, compared to over 60 GB required for full fine-tuning. Fourth, custom GPU backward-pass kernels achieved an average 1.36-fold speedup (and up to 3-fold peak speedup) over existing state-of-the-art sparse GPU libraries by exploiting mask structures where approximately 50% of parameter rows and columns are completely empty.

These findings indicate that organizations can achieve full fine-tuning quality for demanding downstream applications using a single commodity GPU rather than expensive, multi-GPU infrastructure. This capability substantially lowers capital expenditures and operational risks associated with deploying specialized models. However, testing on general instruction-tuning benchmarks (such as Alpaca and OpenPlatypus) revealed that RoSA does not provide an advantage over LoRA on simpler tasks that closely resemble the pre-training data, meaning its benefits are concentrated in complex, specialized domains.

Organizations aiming to deploy specialized reasoning or coding models should adopt RoSA—or QRoSA for severe hardware constraints—as a drop-in replacement for LoRA or full fine-tuning. As a default implementation strategy, practitioners can split parameter budgets equally between low-rank and sparse components (such as a rank of 16 and a sparsity density of 0.6%), which reliably delivers top-tier accuracy without extensive hyperparameter search. In terms of limitations, RoSA is currently 1.7 to 2 times slower per training step than standard LoRA due to sparse matrix computation overheads, and the evidence base is concentrated on 7-billion-parameter architectures. Further testing on larger model scales (such as 70-billion-parameter models) and further kernel optimization will help generalize these high-confidence findings.

Cover for RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation

Abstract

We investigate parameter-efficient fine-tuning (PEFT) methods that can provide good accuracy under limited computational and memory budgets in the context of large language models (LLMs). We present a new PEFT method called Robust Adaptation (RoSA) inspired by robust principal component analysis that jointly trains low-rank and highly-sparse components on top of a set of fixed pretrained weights to efficiently approximate the performance of a full-fine-tuning (FFT) solution. Across a series of challenging generative tasks such as grade-school math and SQL query generation, which require fine-tuning for good performance, we show that RoSA outperforms LoRA, pure sparse fine-tuning, and alternative hybrid methods at the same parameter budget, and can even recover the performance of FFT on some tasks. We provide system support for RoSA to complement the training algorithm, specifically in the form of sparse GPU kernels which enable memory- and computationally-efficient training, and show that it is also compatible with low-precision base weights, resulting in the first joint representation combining quantization, low-rank and sparse approximations. Our code is available at https://github.com/IST-DASLab/RoSA.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Adaptation of Large Language Models
  • 3.1. Notation
  • 3.2. RoSA: Robust Adaptation
  • 4. System Implementation
  • 5. Experiments
  • 5.1. Settings
  • 5.2. Results
  • 6. Discussion
  • Acknowledgments
  • Impact Statement
  • References
  • A. System Details
  • A.1. SDDMM Kernel
  • A.2. CSR-ADD Kernel
  • A.3. Other Details
  • B. Runtime
  • C. Comparison with IA3
  • D. Singular Value Analysis on Full Fine-Tuning
  • E. Instruction-tuning Results
  • F. Qualitative Results

Knowls

  1. Knowl 1 — Robust Adaptation (RoSA) Optimization Formulation

    model/method

    Robust Adaptation (RoSA) is a parameter-efficient fine-tuning (PEFT) framework that approximates the parameter update of full fine-tuning (FFT) by decomposing the weight perturbation of each adapted linear layer into the sum of a low-rank adapter matrix and a highly sparse adapter matrix.

    Let a pretrained language model contain kk adapted fully connected weight matrices W={W1,W2,…,Wk}W = \{W_1, W_2, \dots, W_k\} with Wi∈Rmi×niW_i \in \mathbb{R}^{m_i \times n_i}, and let wˉ∈Rdˉ\bar{w} \in \mathbb{R}^{\bar{d}} represent non-adapted parameters (such as layer normalizations and biases). Given a training dataset D\mathcal{D} and loss function L\mathcal{L}, RoSA freezes WW and wˉ\bar{w}, and solves the joint optimization problem:

    min⁡ΔL,ΔSL(D;W+ΔL+ΔS,wˉ)\min_{\Delta^L, \Delta^S} \mathcal{L}(\mathcal{D}; W + \Delta^L + \Delta^S, \bar{w})

    subject to the layer-wise rank and sparsity constraints:

    ∀1≤i≤k:{rank(ΔiL)≤r∥ΔiS∥0≤d⋅mini\forall 1 \le i \le k : \begin{cases} \text{rank}(\Delta_i^L) \le r \\ \|\Delta_i^S\|_0 \le d \cdot m_i n_i \end{cases}

    where r∈Nr \in \mathbb{N} denotes the maximum adapter rank, d∈(0,1)d \in (0, 1) denotes the fixed target density parameter for the sparse adapter, and ∥⋅∥0\|\cdot\|_0 is the ℓ0\ell_0 pseudo-norm counting non-zero entries. The low-rank update is parameterized as ΔiL=BiLAiL\Delta_i^L = B_i^L A_i^L with BiL∈Rmi×rB_i^L \in \mathbb{R}^{m_i \times r} and AiL∈Rr×niA_i^L \in \mathbb{R}^{r \times n_i}, while ΔiS∈Rmi×ni\Delta_i^S \in \mathbb{R}^{m_i \times n_i} maintains a fixed non-zero coordinate support mask identified prior to joint optimization.

  2. Knowl 2 — RoSA Mask Generation and Warm-Up Training Procedure

    algorithm

    The RoSA fine-tuning pipeline constructs task-adaptive sparsity masks using a two-stage process: first warming up a low-rank adapter instance to capture principal low-rank directions, and then generating fixed coordinate masks by accumulating gradient magnitudes over a small calibration sample set before restarting joint training.

    Input: Pretrained weights W={W1,…,Wk}W = \{W_1, \dots, W_k\}, non-adapted parameters wˉ\bar{w}, training dataset D\mathcal{D}, loss function L\mathcal{L}, rank rr, density dd, calibration set size mm, accumulation exponent α∈{1,2}\alpha \in \{1, 2\}, warmup steps Twarmup=64T_{\text{warmup}} = 64
    Output: Trained low-rank adapters ΔL∗\Delta^{L*}, sparse adapters ΔS∗\Delta^{S*}
    Initialize low-rank parameters ΔL={BiLAiL}i=1k\Delta^L = \{B_i^L A_i^L\}_{i=1}^k with rank rr
    Warm up ΔL\Delta^L by training on D\mathcal{D} for TwarmupT_{\text{warmup}} batches to obtain Δ~L\tilde{\Delta}^L
    Sample a calibration subset DM⊂D\mathcal{D}_M \subset \mathcal{D} containing mm samples (m=32m=32)
    Initialize gradient accumulator Gi←0G_i \leftarrow 0 for each i∈{1,…,k}i \in \{1, \dots, k\}
    for each sample s∈DMs \in \mathcal{D}_M do
        Compute gradient Gis←∇WiL(s;W+Δ~L,wˉ)G_i^s \leftarrow \nabla_{W_i} \mathcal{L}(s; W + \tilde{\Delta}^L, \bar{w}) for each layer ii
        Transfer GisG_i^s to host CPU memory
        Accumulate Gi←Gi+∣Gis∣αG_i \leftarrow G_i + |G_i^s|^\alpha
    end for
    for each layer i∈{1,…,k}i \in \{1, \dots, k\} do
        ki←⌊d⋅mini⌋k_i \leftarrow \lfloor d \cdot m_i n_i \rfloor
        Compute binary mask Mi←TopK-Mask(Gi,ki)M_i \leftarrow \text{TopK-Mask}(G_i, k_i)
        Initialize sparse adapter ΔiS\Delta_i^S with zeros on support MiM_i
        Re-initialize low-rank adapter ΔiL=BiLAiL\Delta_i^L = B_i^L A_i^L
    end for
    Jointly optimize ΔL\Delta^L and non-zero entries of ΔS\Delta^S on D\mathcal{D} using AdamW
    return ΔL∗,ΔS∗\Delta^{L*}, \Delta^{S*}
  3. Knowl 3 — Forward and Backward Propagation Formulations for RoSA Layers

    equation

    For an adapted linear layer with input activation tensor X∈Rb×mX \in \mathbb{R}^{b \times m} (where bb is batch size times sequence length), pretrained frozen weights W∈Rm×nW \in \mathbb{R}^{m \times n}, low-rank adapter factors BL∈Rm×rB^L \in \mathbb{R}^{m \times r} and AL∈Rr×nA^L \in \mathbb{R}^{r \times n}, and sparse adapter ΔS∈Rm×n\Delta^S \in \mathbb{R}^{m \times n} in Compressed Sparse Row (CSR) format, the forward pass evaluation is:

    O=X(W+ΔS)+(XBL)ALO = X(W + \Delta^S) + (X B^L) A^L

    Given the output loss gradient ∂L∂O∈Rb×n\frac{\partial \mathcal{L}}{\partial O} \in \mathbb{R}^{b \times n}, the backward pass computes parameter and input gradients as:

    ∂L∂X=∂L∂O(W+ΔS)T+(∂L∂O(AL)T)(BL)T\frac{\partial \mathcal{L}}{\partial X} = \frac{\partial \mathcal{L}}{\partial O} (W + \Delta^S)^T + \left(\frac{\partial \mathcal{L}}{\partial O} (A^L)^T\right) (B^L)^T

    ∂L∂BL=XT(∂L∂O(AL)T)\frac{\partial \mathcal{L}}{\partial B^L} = X^T \left(\frac{\partial \mathcal{L}}{\partial O} (A^L)^T\right)

    ∂L∂AL=((BL)TXT)∂L∂O\frac{\partial \mathcal{L}}{\partial A^L} = \left((B^L)^T X^T\right) \frac{\partial \mathcal{L}}{\partial O}

    ∂L∂ΔS=XT∂L∂O\frac{\partial \mathcal{L}}{\partial \Delta^S} = X^T \frac{\partial \mathcal{L}}{\partial O}

    where ∂L∂ΔS\frac{\partial \mathcal{L}}{\partial \Delta^S} is evaluated strictly at the non-zero coordinates specified by the sparsity mask via Sampled Dense-Dense Matrix Multiplication (SDDMM).

  4. Knowl 4 — Sparsity-Adaptive SDDMM and CSR Addition GPU Kernels

    model/method

    RoSA employs custom CUDA kernels optimized for the structural sparsity present in language model adaptation masks, where across trained models an average of 46.74% of rows or columns in attention and MLP linear projections contain zero active coordinates.

    1. Sparsity-Adaptive SDDMM Kernel: Evaluates XT∂L∂OX^T \frac{\partial \mathcal{L}}{\partial O} restricted to the CSR representation of ΔS\Delta^S. Unlike standard Sputnik kernels that launch a complete thread grid across the output matrix and terminate idle threads, the RoSA SDDMM kernel uses sorted row-indices to limit thread block launches exclusively to rows and columns containing non-zero elements, and introduces native 16-bit integer indexing. On LLaMA2-7B dimension benchmarks (K=512K=512), this kernel achieves a geometric mean speedup of 1.36x and a peak speedup of 3.0x relative to Sputnik.

    2. CSR-ADD Kernel: Computes the dense-sparse in-place addition W+ΔSW + \Delta^S during the forward pass and its transposed equivalent in the backward pass. The kernel distributes thread blocks across the rows of the sparse matrix via warps, writing sparse CSR non-zero entries directly into the dense representation, with native support for float32, float16, and bfloat16 data types.

  5. Knowl 5 — Fine-Tuning Performance Comparison across Downstream Tasks

    data/table

    Evaluation of single-epoch and extended fine-tuning across three specialized tasks (GSM8k for grade-school math reasoning, ViGGO for video game data-to-text generation, and Spider/Seq2SQL for text-to-SQL generation) using LLaMA2-7B. LoRA and RoSA use bfloat16 parameters; FFT uses float32 (single-epoch float32 FFT achieves 31.8 on GSM8k, 94.0 on ViGGO, and 89.4 on SQL). All adapter experiments adapt all fully connected layers and execute within 20.3–21.8 GB GPU VRAM on a single 24 GB GPU, whereas FFT requires >60 GB.

    Method #Params Memory GSM8k ViGGO SQL
    1 Epoch Extended 1 Epoch Extended 1 Epoch
    FFT 6.7 B > 60 GB 32.3 38.8 82.1 95.0 89.0
    LoRA r=16r = 16 41.1 M 20.6 GB 28.4 37.8 90.5 95.8 88.7
    RoSA r=12,d=0.15%r = 12, d = 0.15\% 41.0 M 20.3 GB 31.2 36.0 95.0 96.5 88.3
    RoSA r=8,d=0.3%r = 8, d = 0.3\% 40.8 M 20.3 GB 29.2 37.5 94.5 97.1 77.6
    RoSA r=4,d=0.45%r = 4, d = 0.45\% 40.6 M 20.3 GB 30.6 35.5 93.4 96.6 89.7
    SpA d=0.6%d = 0.6\% 40.4 M 20.3 GB 26.2 29.5 72.6 89.8 83.2
    LoRA r=32r = 32 82.3 M 20.9 GB 29.6 36.2 87.0 96.8 89.1
    RoSA r=24,d=0.3%r = 24, d = 0.3\% 81.9 M 20.6 GB 30.5 37.8 94.4 95.8 88.9
    RoSA r=16,d=0.6%r = 16, d = 0.6\% 81.6 M 20.7 GB 32.2 38.6 95.2 97.1 88.3
    RoSA r=8,d=0.9%r = 8, d = 0.9\% 81.2 M 20.7 GB 30.3 37.2 94.5 96.9 88.9
    SpA d=1.2%d = 1.2\% 80.9 M 20.7 GB 21.9 29.9 45.8 95.7 74.2
    LoRA r=64r = 64 164.5 M 21.7 GB 27.4 35.5 76.9 95.0 88.7
    RoSA r=48,d=0.6%r = 48, d = 0.6\% 163.8 M 21.3 GB 30.5 38.2 93.0 96.6 88.1
    RoSA r=32,d=1.2%r = 32, d = 1.2\% 163.1 M 21.4 GB 32.2 36.2 93.4 97.3 89.2
    RoSA r=16,d=1.8%r = 16, d = 1.8\% 162.4 M 21.5 GB 32.8 38.4 95.1 96.5 84.6
    SpA d=2.4%d = 2.4\% 161.7 M 21.8 GB 29.6 37.2 92.3 95.7 87.8

    Across parameter budgets (40M, 80M, 160M), RoSA outperforms both pure LoRA and pure sparse adaptation (SpA). In single-epoch regimes, RoSA matches or surpasses FFT on all three benchmarks (e.g., GSM8k single-epoch 32.8 vs. 32.3 FFT; ViGGO 95.2 vs. 82.1 FFT). In extended multi-epoch training, RoSA reaches 38.6 on GSM8k (matching FFT at 38.8) and 97.3 on ViGGO (exceeding FFT at 95.0).

  6. Knowl 6 — Quantized Robust Adaptation (QRoSA) Performance

    data/table

    QRoSA combines 4-bit double-quantized base model weights with joint low-rank and sparse FP16/bfloat16 adapters. Because automatic differentiation does not operate natively on 4-bit quantized base tensors in PyTorch, QRoSA computes gradient accumulation during mask generation via manual output-gradient by input outer products on CPU.

    Evaluation on LLaMA2-7B across GSM8k, ViGGO, and SQL datasets under 4-bit NormalFloat base quantization:

    Method Memory GSM8k ViGGO SQL
    FFT (unquantized) > 60 GB 32.3 82.1 89.0
    QLoRA r=16r = 16 12.6 GB 29.8 88.0 88.2
    QRoSA r=12,d=0.15%r = 12, d = 0.15\% 10.7 GB 31.8 93.8 88.5
    QRoSA r=8,d=0.3%r = 8, d = 0.3\% 10.7 GB 30.9 95.0 88.6
    QRoSA r=4,d=0.45%r = 4, d = 0.45\% 10.7 GB 30.3 92.4 86.7
    QSpA d=0.6%d = 0.6\% 10.8 GB 22.8 89.5 79.2
    QLoRA r=32r = 32 13.0 GB 25.6 74.7 89.0
    QRoSA r=24,d=0.3%r = 24, d = 0.3\% 11.0 GB 30.4 93.3 88.3
    QRoSA r=16,d=0.6%r = 16, d = 0.6\% 11.1 GB 33.1 93.8 86.6
    QRoSA r=8,d=0.9%r = 8, d = 0.9\% 11.1 GB 32.8 95.4 83.7
    QSpA d=1.2%d = 1.2\% 11.3 GB 28.0 93.0 85.0
    QLoRA r=64r = 64 13.8 GB 30.6 88.1 89.4
    QRoSA r=48,d=0.6%r = 48, d = 0.6\% 11.9 GB 30.5 93.6 81.6
    QRoSA r=32,d=1.2%r = 32, d = 1.2\% 11.9 GB 32.3 94.3 88.2
    QRoSA r=16,d=1.8%r = 16, d = 1.8\% 12.0 GB 30.8 95.0 88.5
    QSpA d=2.4%d = 2.4\% 12.2 GB 28.9 90.8 42.9

    QRoSA operates under 12 GB GPU VRAM, allowing execution on consumer-grade GPUs. On GSM8k, QRoSA achieves an accuracy of 33.1, outperforming unquantized FFT (32.3) and QLoRA (25.6–30.6).

  7. Knowl 7 — Mask Selection Strategy Ablation for Sparse Adapters

    data/table

    Comparison of sparse mask generation techniques on LLaMA2-7B trained on GSM8k for 1 epoch using an 80M trainable parameter budget with target density dd. Let τd(⋅)\tau_d(\cdot) denote the TopK magnitude mask operator selecting the largest entries to meet density dd.

    Masking Method Mask Formula MM GSM8k Accuracy (%)
    Lottery Ticket Mask (LTM, Oracle) τd(ΔS∗)\tau_d(\Delta_S^*) from RPCA on Δ∗=WFFT−WBASE\Delta^* = W_{\text{FFT}} - W_{\text{BASE}} 33.66
    GradMag-LW (RoSA, Ours) τd(∇W+Δ~LL)\tau_d(\nabla_{W+\tilde{\Delta}^L} \mathcal{L}) after LoRA warmup 32.16
    GradMag / GradFish (FISH Mask) τd(∇WL)\tau_d(\nabla_W \mathcal{L}) accumulated at init 30.10
    WeightRPCA (DSEE) τd(WS)\tau_d(W_S) from RPCA on pretrained weights WW 30.71
    GradRPCA τd(∇WS)\tau_d(\nabla W_S) from RPCA on weight gradients 29.87
    RND Uniform random coordinates 30.25

    GradMag-LW (RoSA's method) comes within 1.5% accuracy of the hindsight oracle mask (LTM). Mask generation at initialization without a LoRA warmup period (FISH Mask, GradRPCA, WeightRPCA) performs on par with random coordinate masks (~30%), demonstrating that the initial task-adaptive warmup period is essential to isolate complementary non-low-rank gradient directions.

  8. Knowl 8 — Optimal Adapter Budget Allocation Between Low-Rank and Sparse Components

    empirical result

    Singular value decomposition of full fine-tuning weight perturbations Δ∗=WFFT−WBASE\Delta^* = W_{\text{FFT}} - W_{\text{BASE}} on LLaMA2-7B indicates that while the update matrix is rank-deficient, it is not strictly low-rank: it possesses a small number of large singular values followed by a substantial tail of non-zero singular values. Consequently, increasing LoRA rank beyond a threshold (rank r≈12–16r \approx 12\text{--}16) yields diminishing returns.

    Under a fixed parameter budget PP:

    1. Approximating Δ∗\Delta^* via Robust PCA into low-rank matrix LL and sparse matrix SS achieves lower Frobenius norm reconstruction error than dedicating the entire parameter budget exclusively to low-rank (S=0S=0) or sparse (L=0L=0) updates.
    2. In fine-tuning experiments, splitting the parameter budget roughly equally between low-rank and sparse adapters (Prank≈PsparseP_{\text{rank}} \approx P_{\text{sparse}}) provides a consistent heuristic that outperforms pure LoRA and pure SpA across tasks.
  9. Knowl 9 — Task Complexity Scope and Instruction-Tuning Limitations

    limitation

    RoSA provides substantial accuracy gains over LoRA specifically on complex downstream reasoning and structured generation tasks (such as mathematical problem solving, SQL parsing, and data-to-text generation) where a large accuracy gap exists between LoRA and full fine-tuning. Conversely, on simpler general instruction-tuning tasks where downstream data is closely aligned with pretraining distributions, RoSA does not outperform LoRA.

    Evaluation on LLaMA2-7B using 5-shot MMLU accuracy:

    • Base pretrained LLaMA2-7B: 45.75%
    • Fine-tuned on OpenPlatypus: LoRA r=16r=16 achieves 49.92%, whereas RoSA (r=16,d=0.6%r=16, d=0.6\%) achieves 46.54%.
    • Fine-tuned on Alpaca: LoRA r=16r=16 achieves 45.80%, whereas RoSA (r=16,d=0.6%r=16, d=0.6\%) achieves 46.52%.

    On these instruction datasets, LoRA already captures the task update adequately, making the additional sparse component in RoSA unnecessary.

  10. Knowl 10 — Training Throughput and Runtime Characteristics of RoSA

    empirical result

    Runtime benchmarking of fine-tuning LLaMA2-7B on a single NVIDIA RTX A6000 GPU under an 80M parameter budget shows that RoSA exhibits a training throughput of approximately 0.0575–0.0602 batches/second, compared to 0.1149 batches/second for LoRA (r=32r=32), representing a 1.7x to 2.0x slowdown.

    Under 4-bit base quantization, QRoSA achieves 0.0515–0.0531 batches/second versus 0.0911 batches/second for QLoRA (r=32r=32). The runtime overhead stems from unstructured sparsity memory access patterns and the requirement of FP32 accumulation precision in custom sparse kernels compared to highly optimized dense FP16/Tensor Core matrix multiplications in LoRA.

Coverage note — Qualitative mathematical text generation examples from Appendix F and few-shot comparisons against IA3 on the RAFT dataset from Appendix C were omitted as they represent illustrative samples and peripheral baselines rather than core contributions.

References

  1. 1.Alex, N., Lifland, E., Tunstall, L., Thakur, A., Maham, P., Riedel, C. J., Hine, E., Ashurst, C., Sedille, P., Carlier, A., et al. Raft: A real-world few-shot text classification benchmark. arXiv preprint arXiv:2109.14076, 2021.
  2. 2.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  3. 3.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  4. 4.Candès, E. J., Li, X., Ma, Y., and Wright, J. Robust principal component analysis? Journal of the ACM (JACM), 58(3): 1–37, 2011.
  5. 5.Castro, R. L., Ivanov, A., Andrade, D., Ben-Nun, T., Fraguela, B. B., and Hoefler, T. Venom: A vectorized n: M format for unleashing the power of sparse tensor cores. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–14, 2023.
  6. 6.Chen, X., Chen, T., Chen, W., Awadallah, A. H., Wang, Z., and Cheng, Y. Dsee: Dually sparsity-embedded efficient tuning of pre-trained language models. arXiv preprint arXiv:2111.00160, 2021.
  7. 7.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.De La Torre, F. and Black, M. J. A framework for robust subspace learning. International Journal of Computer Vision, 54:117–142, 2003.
  9. 9.Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al. Large scale distributed deep networks. Advances in neural information processing systems, 25, 2012.
  10. 10.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. LLM.int8(): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, 2022.
  11. 11.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023a.
  12. 12.Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078, 2023b.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Edalati, A., Tahaei, M., Kobyzev, I., Nia, V. P., Clark, J. J., and Rezagholizadeh, M. Krona: Parameter efficient tuning with kronecker adapter. arXiv preprint arXiv:2212.10650, 2022.
  15. 15.Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning, pp. 2943–2952. PMLR, 2020.
  16. 16.Fischler, M. A. and Bolles, R. C. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
  17. 17.Frantar, E. and Alistarh, D. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  18. 18.Gale, T., Elsen, E., and Hooker, S. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574, 2019.
  19. 19.Gale, T., Zaharia, M., Young, C., and Elsen, E. Sparse GPU kernels for deep learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, 2020.
  20. 20.Gnanadesikan, R. and Kettenring, J. R. Robust estimates, residuals, and outlier detection with multiresponse data. Biometrics, pp. 81–124, 1972.
  21. 21.Gray, S., Radford, A., and Kingma, D. P. Gpu kernels for block-sparse weights. arXiv preprint arXiv:1711.09224, 3(2):2, 2017.
  22. 22.He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G. Towards a unified view of parameter-efficient transfer learning. In Proceedings of the 10th International Conference on Learning Representations (ICLR-2022), 2022.
  23. 23.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  24. 24.Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. The Journal of Machine Learning Research, 22(1):10882–11005, 2021.
  25. 25.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  26. 26.Hubara, I., Chmiel, B., Island, M., Banner, R., Naor, J., and Soudry, D. Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks. Advances in neural information processing systems, 34: 21099–21111, 2021.
  27. 27.Huber, P. J. Robust statistics, volume 523. John Wiley & Sons, 2004.
  28. 28.Hyeon-Woo, N., Ye-Bin, M., and Oh, T.-H. Fedpara: Low-rank hadamard product for communication-efficient federated learning. arXiv preprint arXiv:2108.06098, 2021.
  29. 29.Ivanov, A., Dryden, N., and Hoefler, T. Sten: An interface for efficient sparsity in pytorch. 2022.
  30. 30.Jiang, P., Hu, L., and Song, S. Exposing and exploiting fine-grained block structures for fast and accurate sparse training. Advances in Neural Information Processing Systems, 35:38345–38357, 2022.
  31. 31.Juraska, J., Bowden, K., and Walker, M. ViGGO: A video game corpus for data-to-text generation in open-domain conversation. In Proceedings of the 12th International Conference on Natural Language Generation, pp. 164–172, Tokyo, Japan, October–November 2019. Association for Computational Linguistics. doi: 10.18653/v1/W19-8623. URL https://aclanthology.org/W19-8623.
  32. 32.Ke, Q. and Kanade, T. Robust l/sub 1/norm factorization in the presence of outliers and missing data by alternative convex programming. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), volume 1, pp. 739–746. IEEE, 2005.
  33. 33.Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259, 2022.
  34. 34.Kurtic, E., Kuznedelev, D., Frantar, E., Goin, M., and Alistarh, D. Sparse finetuning for inference acceleration of large language models. arXiv preprint arXiv:2310.06927, 2023.
  35. 35.Lee, A. N., Hunter, C. J., and Ruiz, N. Platypus: Quick, cheap, and powerful refinement of llms. arXiv preprint arXiv:2308.07317, 2023.
  36. 36.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  37. 37.Li, S., Osawa, K., and Hoefler, T. Efficient quantized sparse matrix operations on tensor cores. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–15. IEEE, 2022.
  38. 38.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  39. 39.Li, Y., Yu, Y., Liang, C., He, P., Karampatziakis, N., Chen, W., and Zhao, T. Loftq: Lora-fine-tuning-aware quantization for large language models. arXiv preprint arXiv:2310.08659, 2023.
  40. 40.Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems, 35: 1950–1965, 2022.
  41. 41.Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., and Tang, J. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021.
  42. 42.Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. Gpt understands, too. AI Open, 2023.
  43. 43.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  44. 44.Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  45. 45.Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H. Metaicl: Learning to learn in context. arXiv preprint arXiv:2110.15943, 2021.
  46. 46.MosaicML. LLM Foundry, 2023a. URL https://github.com/mosaicml/llm-foundry.
  47. 47.MosaicML. Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023b. URL www.mosaicml.com/blog/mpt-7b. Accessed: 2023-12-22.
  48. 48.Niederfahrenhorst, A., Hakhamaneshi, K., and Ahmad, R. Fine-Tuning LLMs: LoRA or Full-Parameter?, 2023. URL https://www.anyscale.com/blog/fine-tuning-llms-lora-or-full-parameter-an-in-depth-analysis-with-llama-2.
  49. 49.Nikdan, M., Pegolotti, T., Iofinova, E., Kurtic, E., and Alistarh, D. Sparseprop: Efficient sparse backpropagation for faster training of neural networks at the edge. In International Conference on Machine Learning, pp. 26215–26227. PMLR, 2023.
  50. 50.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  51. 51.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, 2019.
  52. 52.Peste, A., Iofinova, E., Vladu, A., and Alistarh, D. Ac/dc: Alternating compressed/decompressed training of deep neural networks. Advances in neural information processing systems, 34:8557–8570, 2021.
  53. 53.Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., and Schölkopf, B. Controlling text-to-image diffusion by orthogonal finetuning. arXiv preprint arXiv:2306.07280, 2023.
  54. 54.Sanh, V., Wolf, T., and Rush, A. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389, 2020.
  55. 55.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  56. 56.Singh, S. P. and Alistarh, D. Woodfisher: Efficient second-order approximation for neural network compression. Advances in Neural Information Processing Systems, 33: 18098–18109, 2020.
  57. 57.Sung, Y.-L., Nair, V., and Raffel, C. A. Training neural networks with fixed sparse masks. Advances in Neural Information Processing Systems, 34:24193–24205, 2021.
  58. 58.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  59. 59.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  60. 60.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  61. 61.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022a.
  62. 62.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., et al. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 5085–5109, 2022b.
  63. 63.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  64. 64.Wright, J., Ganesh, A., Rao, S., Peng, Y., and Ma, Y. Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization. Advances in neural information processing systems, 22, 2009.
  65. 65.Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887, 2018.
  66. 66.Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T. Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512, 2023.
  67. 67.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  68. 68.Zhong, V., Xiong, C., and Socher, R. Seq2sql: Generating structured queries from natural language using reinforcement learning. CoRR, abs/1709.00103, 2017.
  69. 69.Zhou, T. and Tao, D. Greedy bilateral sketch, completion & smoothing. In Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics, volume 31 of Proceedings of Machine Learning Research, pp. 650–658. PMLR, 2013.

Citation

MLA
Nikdan, M., et al. “RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation”. arXiv, 2024, http://arxiv.org/abs/2401.04679v7.
APA
Nikdan, M., Tabesh, S., Crnčević, E., & Alistarh, D. (2024). RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation. arXiv. http://arxiv.org/abs/2401.04679v7
Chicago
Nikdan, M., S. Tabesh, E. Crnčević, and D. Alistarh. 2024. “RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation”. arXiv. http://arxiv.org/abs/2401.04679v7.
Harvard
Nikdan, M. et al. (2024) “RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.04679v7.
Vancouver
1. Nikdan M, Tabesh S, Crnčević E, Alistarh D (2024) RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation. arXiv

BibTeX

@article{nikdan2024rosa,
  title = {RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation},
  author = {Nikdan, Mahdi and Tabesh, Soroush and Crnčević, Elvir and Alistarh, Dan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.04679v7},
  eprint = {2401.04679}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/