LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning

Han GuoPhilip GreengardEric P. XingYoon Kim

article2024ICLR91 citations

Presents LQ-LoRA, a matrix decomposition technique that coordinates low-rank adaptation with dynamically budgeted quantization to enable sub-3-bit language model finetuning and compression with minimal performance loss.

Listen

Adapting large language models to new tasks typically requires massive computational and hardware memory resources, creating a major barrier for organizations seeking cost-effective artificial intelligence deployment. Parameter-efficient fine-tuning methods like Low-Rank Adaptation (LoRA) reduce memory demands by freezing original model weights and training compact low-rank matrices. When combined with quantization—a technique that compresses weights into lower bit-widths—methods such as QLoRA have become popular. However, conventional quantization introduces severe mathematical errors when pushed below 4 bits, and standard zero-initialization fails to compensate for these discrepancies.

The main objective of the article is to demonstrate that an alternating matrix decomposition approach, termed LQ-LoRA, can effectively adapt and compress language models into sub-4-bit and sub-3-bit regimes while maintaining strong performance and adhering to flexible target memory budgets.

The authors evaluate this method through empirical experiments on RoBERTa-Large and LLaMA-2 models (7 billion and 70 billion parameters) across continual language modeling, instruction following, and standard natural language understanding benchmarks. The technique iteratively decomposes pretrained weight matrices into fixed, memory-efficient quantized components and trainable low-rank components that capture high-variance parameters. It incorporates integer linear programming to dynamically allocate varying quantization bit-widths across individual model layers according to a global memory budget, alongside an optional data-aware approach that weights matrix reconstruction using an empirical Fisher information matrix.

The primary findings show that LQ-LoRA consistently outperforms standard QLoRA and GPTQ-LoRA baselines across similar bit allocations. Notably, 3.5-bit LQ-LoRA matches the performance of standard 4-bit QLoRA, while 2.75-bit LQ-LoRA performs competitively with 3-bit QLoRA. When applied as a standalone compression technique, a 2.75-bit LLaMA-2-70B model (effective 2.85 bits) achieves perplexity comparable to the uncompressed 16-bit baseline while fitting within 27 gigabytes of storage, allowing full execution on a single commercial GPU. In addition, incorporating Fisher weighting significantly reduces performance loss on smaller 7B models, and increasing the low-rank capacity directly improves model reconstruction quality under LQ-LoRA.

These findings indicate that organizations can substantially reduce infrastructure and operational costs by fine-tuning and running 70-billion-parameter models on single GPUs rather than multi-GPU clusters. By dynamically assigning precision across layers rather than applying uniform quantization, engineering teams can maximize task performance within strict hardware limits without requiring proprietary CUDA extensions.

Decision-makers should consider adopting LQ-LoRA when memory constraints prevent the deployment of standard 4-bit models, particularly for 70B-scale models where 2.75- to 3.5-bit allocations provide substantial memory savings with minimal quality loss. For sub-3-bit configurations, practitioners should integrate Fisher weighting using generic calibration text to prevent performance drops. However, teams evaluating aggressive quantization below 3 bits should run task-specific pilot benchmarks, as complex reasoning and math tasks (such as GSM8K) exhibit noticeable degradation even when perplexity metrics appear strong.

Readers should note that the decomposition algorithm is a heuristic method lacking theoretical convergence guarantees. Furthermore, performance degrades steeply when pushing quantization to 2.5 bits or lower. Despite these boundary limits, the empirical results provide high confidence that LQ-LoRA is a robust, practical solution for sub-4-bit model adaptation and compression.

Cover for LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning

Abstract

We propose a simple approach for memory-efficient adaptation of pretrained language models. Our approach uses an iterative algorithm to decompose each pretrained matrix into a high-precision low-rank component and a memory-efficient quantized component. During finetuning, the quantized component remains fixed and only the low-rank component is updated. We present an integer linear programming formulation of the quantization component which enables dynamic configuration of quantization parameters (e.g., bit-width, block size) for each matrix given an overall target memory budget. We further explore a data-aware version of the algorithm which uses an approximation of the Fisher information matrix to weight the reconstruction objective during matrix decomposition. Experiments on finetuning RoBERTa and LLaMA-2 (7B and 70B) demonstrate that our low-rank plus quantized matrix decomposition approach (LQ-LoRA) outperforms strong QLoRA and GPTQ-LoRA baselines and enables aggressive quantization to sub-3 bits with only minor performance degradations. When finetuned on a language modeling calibration dataset, LQ-LoRA can also be used for model compression; in this setting our 2.75-bit LLaMA-2-70B model (which has 2.85 bits on average when including the low-rank components and requires 27GB of GPU memory) performs respectably compared to the 16-bit baseline.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Low-rank Adaptation of Large Language Models
  • 2.2 Weight Quantization of Large Language Models
  • 3 Method: LQ-LoRA
  • 3.1 Low-rank Plus Quantized Matrix Decomposition
  • 3.2 Mixed-Configuration Quantization via an Integer Linear Program
  • 3.3 Data-Aware Matrix Decomposition via Fisher-weighted SVD
  • 4 Empirical Study
  • 4.1 Results
  • 4.2 LQ-LoRA for Model Compression
  • 4.3 Analysis
  • 5 Discussion and Limitations
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Implementation Details
  • B Full Results

Knowls

  1. Knowl 1 — Low-Rank Plus Quantized Matrix Decomposition Formulation

    model/method

    In low-rank adaptation (LoRA), a pretrained linear layer weight matrix W∈Rd×kW \in \mathbb{R}^{d \times k} is parameterized as W+L1L2W + L_1 L_2, where L1∈Rd×rL_1 \in \mathbb{R}^{d \times r} and L2∈Rr×kL_2 \in \mathbb{R}^{r \times k} with rank r≪min⁡(d,k)r \ll \min(d, k). Conventional LoRA initializes L1∼N(0,σ2)L_1 \sim \mathcal{N}(0, \sigma^2) and L2=0L_2 = 0 so that the model's initial output is unchanged (X(W+L1L2)=XWX(W + L_1 L_2) = XW). When adapting quantized models where WW is replaced by a quantized matrix q(W)q(W), initializing L2=0L_2 = 0 yields q(W)+L1L2=q(W)≠Wq(W) + L_1 L_2 = q(W) \neq W, preserving substantial quantization error, particularly in sub-4-bit regimes.

    LQ-LoRA resolves this by formulating the initialization as a matrix decomposition that jointly optimizes for an easily quantizable component QQ and a low-rank component L1L2L_1 L_2 capturing the highest-variance directions of WW:

    arg⁡min⁡Q,L1,L2∥W−(Q+L1L2)∥Fsubject toQ∈Qbd×k,  L1∈Rd×r,  L2∈Rr×k\arg\min_{Q, L_1, L_2} \|W - (Q + L_1 L_2)\|_F \quad \text{subject to} \quad Q \in \mathcal{Q}_b^{d \times k}, \; L_1 \in \mathbb{R}^{d \times r}, \; L_2 \in \mathbb{R}^{r \times k}

    where Qbd×k⊂Rd×k\mathcal{Q}_b^{d \times k} \subset \mathbb{R}^{d \times k} denotes the set of matrices representable under bb-bit NormalFloat (NF) quantization. During downstream fine-tuning, the quantized component QQ remains completely fixed in low-precision storage, and only the low-rank parameters L1L_1 and L2L_2 are updated.

  2. Knowl 2 — Alternating Iterative Matrix Decomposition for LQ-LoRA

    algorithm

    The alternating iterative decomposition heuristic finds the low-rank factors L1,L2L_1, L_2 and quantized matrix QQ for a pretrained weight matrix WW by alternating between randomized Singular Value Decomposition (SVD) and quantization.

    Input: Pretrained weight matrix W∈Rd×kW \in \mathbb{R}^{d \times k}, rank rr, quantization configuration cc, optional Fisher weighting matrix F∈Rd×kF \in \mathbb{R}^{d \times k}, maximum iterations TT
    Output: Quantized matrix Q∈Qbd×kQ \in \mathcal{Q}_b^{d \times k}, low-rank factors L1∈Rd×r,L2∈Rr×kL_1 \in \mathbb{R}^{d \times r}, L_2 \in \mathbb{R}^{r \times k}, final error ϵ\epsilon
    Initialize Q←0Q \leftarrow 0 and ϵ0←∞\epsilon_0 \leftarrow \infty
    for t←1t \leftarrow 1 to TT do
        L1,L2←Factorize(W−Q,F,r)L_1, L_2 \leftarrow \text{Factorize}(W - Q, F, r)
        Q←Quantize(W−L1L2,c)Q \leftarrow \text{Quantize}(W - L_1 L_2, c)
        if FF is None then
            ϵt←∥W−(Q+L1L2)∥F\epsilon_t \leftarrow \|W - (Q + L_1 L_2)\|_F
        else
            ϵt←∥F⊙(W−(Q+L1L2))∥F\epsilon_t \leftarrow \|\sqrt{F} \odot (W - (Q + L_1 L_2))\|_F
        end if
        if ϵt>ϵt−1\epsilon_t > \epsilon_{t-1} then
            break
        end if
    end for
    return Q,L1,L2,ϵtQ, L_1, L_2, \epsilon_t

    Here, Factorize(A,F,r)\text{Factorize}(A, F, r) computes a rank-rr truncated SVD on AA (or Fisher-weighted SVD if FF is present), and Quantize(R,c)\text{Quantize}(R, c) applies NormalFloat quantization under configuration cc. The procedure terminates early if the reconstruction error strictly increases.

  3. Knowl 3 — Data-Aware Fisher-Weighted Low-Rank Decomposition

    model/method

    To weight reconstruction by parameter sensitivity on actual data distributions, data-aware LQ-LoRA uses the diagonal of the empirical Fisher information matrix F∈Rd×kF \in \mathbb{R}^{d \times k} computed across DD calibration sequences:

    Fij=1D∑d=1D(∂∂Wijlog⁡pLM(x(d)))2F_{ij} = \frac{1}{D} \sum_{d=1}^D \left( \frac{\partial}{\partial W_{ij}} \log p_{\text{LM}}(x^{(d)}) \right)^2

    The weighted low-rank decomposition objective for residual matrix E=W−QE = W - Q is:

    arg⁡min⁡L1,L2∥F⊙(E−L1L2)∥F2\arg\min_{L_1, L_2} \|\sqrt{F} \odot (E - L_1 L_2)\|_F^2

    Because general element-weighted low-rank approximation is NP-hard, LQ-LoRA approximates FF by separable row- and column-averaged scaling factors Drow∈Rd×dD_{\text{row}} \in \mathbb{R}^{d \times d} and Dcol∈Rk×kD_{\text{col}} \in \mathbb{R}^{k \times k}:

    Drow=diag(avg(F1,⋅),…,avg(Fd,⋅)),Dcol=diag(avg(F⋅,1),…,avg(F⋅,k))D_{\text{row}} = \text{diag}\left(\text{avg}\left(\sqrt{F_{1,\cdot}}\right), \dots, \text{avg}\left(\sqrt{F_{d,\cdot}}\right)\right), \quad D_{\text{col}} = \text{diag}\left(\text{avg}\left(\sqrt{F_{\cdot,1}}\right), \dots, \text{avg}\left(\sqrt{F_{\cdot,k}}\right)\right)

    Under this approximation, the objective simplifies to arg⁡min⁡L1,L2∥Drow(E−L1L2)Dcol∥F2\arg\min_{L_1, L_2} \|D_{\text{row}} (E - L_1 L_2) D_{\text{col}}\|_F^2, which is solved via standard SVD:

    [U,Σ,V⊤]←SVD(DrowEDcol,r)[U, \Sigma, V^\top] \leftarrow \text{SVD}(D_{\text{row}} E D_{\text{col}}, r) L1=Drow−1UΣ,L2=ΣV⊤Dcol−1L_1 = D_{\text{row}}^{-1} U \sqrt{\Sigma}, \quad L_2 = \sqrt{\Sigma} V^\top D_{\text{col}}^{-1}

    The Fisher matrix is estimated once from generic calibration data (e.g., 10,000 C4 sequences) and reused across downstream fine-tuning tasks.

  4. Knowl 4 — Dynamic Mixed-Configuration Quantization via Integer Linear Programming

    model/method

    LQ-LoRA employs double NormalFloat (NF) quantization parameterized by a configuration tuple c=(b0,b1,b2,B0,B1)c = (b_0, b_1, b_2, B_0, B_1), where b0b_0 is the base NF bit-width for blocks of size B0B_0, b1b_1 is the uniform integer bit-width used to quantize the first-level absolute maximum scale factors ss in blocks of size B1B_1, and b2b_2 is the floating-point precision of the second-level scale factors vv. The storage requirement in bits for matrix AA under configuration cc is:

    storage(A,c)=sizeof(A)⋅(b0+b1B0+b2B0B1)\text{storage}(A, c) = \text{sizeof}(A) \cdot \left( b_0 + \frac{b_1}{B_0} + \frac{b_2}{B_0 B_1} \right)

    Given a set of NN linear layer weight matrices {W(i)}i=1N\{W^{(i)}\}_{i=1}^N in a language model and a discrete configuration search space C\mathcal{C} (e.g., b0,b1∈{2,3,4}b_0, b_1 \in \{2, 3, 4\}, b2∈{bf16,fp16,fp32}b_2 \in \{\text{bf16}, \text{fp16}, \text{fp32}\}, B0∈{16,32,64}B_0 \in \{16, 32, 64\}, B1∈{16,64,256}B_1 \in \{16, 64, 256\}), layer-specific configurations are assigned by solving an integer linear program (ILP):

    min⁡X∈{0,1}N×∣C∣∑i=1N∑c∈Cerror(W(i),c)⋅X[i,c]\min_{X \in \{0, 1\}^{N \times |\mathcal{C}|}} \sum_{i=1}^N \sum_{c \in \mathcal{C}} \text{error}(W^{(i)}, c) \cdot X[i, c] subject to ∑i=1N∑c∈Cstorage(W(i),c)⋅X[i,c]≤budget,∑c∈CX[i,c]=1,  ∀i∈{1,…,N}\text{subject to } \sum_{i=1}^N \sum_{c \in \mathcal{C}} \text{storage}(W^{(i)}, c) \cdot X[i, c] \le \text{budget}, \quad \sum_{c \in \mathcal{C}} X[i, c] = 1, \; \forall i \in \{1, \dots, N\}

    where error(W(i),c)=∥W(i)−(Q+L1L2)∥F2\text{error}(W^{(i)}, c) = \|W^{(i)} - (Q + L_1 L_2)\|_F^2 is the reconstruction error resulting from running iterative LQ decomposition on W(i)W^{(i)} under configuration cc.

  5. Knowl 5 — End-to-End LQ-LoRA Optimization and Adaptation Algorithm

    algorithm

    The overall LQ-LoRA pipeline combines the ILP configuration search with alternating matrix decomposition across all model parameters.

    Input: Model weight matrices {W(i)}i=1N\{W^{(i)}\}_{i=1}^N, optional Fisher matrices {F(i)}i=1N\{F^{(i)}\}_{i=1}^N, candidate configurations C\mathcal{C}, rank rr, target quantization budget BQB_Q
    Output: Decomposed model parameters {Q(i),L1(i),L2(i)}i=1N\{Q^{(i)}, L_1^{(i)}, L_2^{(i)}\}_{i=1}^N
    Initialize error matrix E←zeros(N,∣C∣)E \leftarrow \text{zeros}(N, |\mathcal{C}|) and storage matrix S←zeros(N,∣C∣)S \leftarrow \text{zeros}(N, |\mathcal{C}|)
    for i←1i \leftarrow 1 to NN do
        for c∈Cc \in \mathcal{C} do
            Q(i),L1(i),L2(i),ϵ←LQ(W(i),F(i),c,r)Q^{(i)}, L_1^{(i)}, L_2^{(i)}, \epsilon \leftarrow \text{LQ}(W^{(i)}, F^{(i)}, c, r)
            E[i,c]←ϵ2E[i, c] \leftarrow \epsilon^2
            S[i,c]←storage(W(i),c)S[i, c] \leftarrow \text{storage}(W^{(i)}, c)
        end for
    end for
    {c(i)}i=1N←ILPSolve(S,E,BQ)\{c^{(i)}\}_{i=1}^N \leftarrow \text{ILPSolve}(S, E, B_Q)
    for i←1i \leftarrow 1 to NN do
        Q(i),L1(i),L2(i),_←LQ(W(i),F(i),c(i),r)Q^{(i)}, L_1^{(i)}, L_2^{(i)}, \_ \leftarrow \text{LQ}(W^{(i)}, F^{(i)}, c^{(i)}, r)
    end for
    return {Q(i),L1(i),L2(i)}i=1N\{Q^{(i)}, L_1^{(i)}, L_2^{(i)}\}_{i=1}^N

    The pre-computation of reconstruction errors across the configuration grid is performed once in parallel across GPUs. After obtaining the optimal layer configurations {c(i)}\{c^{(i)}\}, final decomposition matrices are computed and used directly for fine-tuning.

  6. Knowl 6 — LLaMA-2 Language Modeling and Instruction Tuning Benchmarks

    data/table

    Performance of LQ-LoRA (rank r=64r=64) evaluated on LLaMA-2 7B and 70B across C4 validation perplexity, WikiText-2 perplexity, 5-shot MMLU accuracy, and Vicuna-style instruction tuning evaluation (pairwise GPT-4 win rate against GPT-3.5) after adaptation on C4 or OpenAssistant.

    Method Bits/param C4 (PPL) WikiText (PPL) MMLU (acc.) Vicuna Eval
    70B 7B 70B 7B 70B 7B 70B 7B
    Dense (no training) 16.0 6.50 8.22 3.68 6.10 0.70 0.46 - -
    Dense (full fine-tune) 16.0 - - - - - - OOM 0.41
    QLoRA 3-bit 3.127 6.23 8.21 4.12 6.76 0.68 0.43 0.46 0.33
    QLoRA 4-bit 4.127 6.01 7.61 3.78 6.25 0.70 0.46 0.47 0.41
    GPTQ-LoRA 3-bit 3.148 6.34 8.48 4.33 7.09 0.67 0.39 - -
    GPTQ-LoRA 4-bit 4.156 6.03 7.68 3.82 6.29 0.69 0.45 - -
    QLoRA + ILP 2.50 2223.2 2996.3 3319.4 4084.3 0.23 0.23 0.00 0.00
    2.75 2193.9 2736.5 3292.6 3932.2 0.23 0.27 0.00 0.00
    3.00 1781.5 1969.3 2587.0 3091.0 0.23 0.23 0.44 0.33
    3.25 6.15 8.04 3.99 6.66 0.69 0.44 0.50 0.41
    3.50 6.10 7.91 3.93 6.51 0.69 0.45 0.47 0.36
    3.75 6.06 7.76 3.85 6.39 0.69 0.44 0.55 0.35
    4.00 6.02 7.65 3.80 6.29 0.70 0.45 0.49 0.49
    LQ-LoRA 2.50 6.83 10.00 4.95 8.44 0.62 0.31 0.57 0.23
    2.75 6.42 8.95 4.44 7.55 0.66 0.31 0.56 0.38
    3.00 6.18 8.09 4.08 6.73 0.68 0.41 0.59 0.47
    3.25 6.10 7.83 3.95 6.44 0.69 0.44 0.56 0.56
    3.50 6.06 7.75 3.88 6.39 0.69 0.46 0.55 0.45
    3.75 6.02 7.64 3.80 6.27 0.69 0.45 0.65 0.40
    4.00 5.99 7.57 3.77 6.23 0.69 0.46 0.66 0.44
    LQ-LoRA (Fisher) 2.50 6.72 9.03 4.80 7.42 0.67 0.39 0.59 0.45
    2.75 6.35 8.25 4.32 6.78 0.67 0.43 0.56 0.44
    3.00 6.14 7.88 4.02 6.48 0.68 0.44 0.65 0.51
    3.25 6.08 7.76 3.92 6.40 0.69 0.46 0.54 0.49
    3.50 6.04 7.66 3.86 6.31 0.69 0.45 0.62 0.49
    3.75 6.01 7.57 3.80 6.24 0.69 0.47 0.59 0.47
    4.00 5.98 7.53 3.76 6.20 0.70 0.46 0.66 0.51

    LQ-LoRA consistently matches or outperforms baseline QLoRA and GPTQ-LoRA at lower bit-widths (e.g., 3.50-bit LQ-LoRA is comparable to 4.127-bit NF-4 QLoRA, and 2.75-bit LQ-LoRA is competitive with 3.127-bit NF-3 QLoRA). At ≤3.0\le 3.0 bits, standard QLoRA+ILP collapses completely (PPL >1700> 1700), whereas LQ-LoRA maintains low perplexity and strong downstream instruction following.

  7. Knowl 7 — Sub-4-Bit Post-Training Quantization and LLM Compression Benchmarks

    data/table

    Comparison of LQ-LoRA against sub-4-bit post-training quantization (PTQ) methods on LLaMA-1 and LLaMA-2 models on C4 and WikiText-2 perplexity. Effective bits include storage overhead from quantization metadata (scaling factors) and, for LQ-LoRA, the low-rank component quantized to 8 bits (NF-8).

    Method Effective Bits C4 (PPL) WikiText (PPL)
    (7B, 65B/70B) 7B 65B/70B 7B 65B/70B
    LLaMA-1 Uncompressed 16.0 7.08 5.62 5.68 3.53
    SpQR 3.94, 3.90 7.28 5.70 5.87 3.68
    RTN (3-bits, g128) 3.15 8.62 6.10 7.01 4.24
    GPTQ (3-bits, g128) 3.15 7.85 6.00 6.55 4.17
    AWQ (3-bits, g128) 3.15 7.92 5.94 6.46 3.99
    PEQA (3-bits, g128) 3.15 - - 5.91 -
    OWQ (3-bits) 3.10 8.15 6.16 6.39 4.08
    SqueezeLLM (3-bits, 0.45%) 3.24 7.56 - 6.13 -
    SqueezeLLM (3-bits) 3.02 7.75 - 6.32 -
    OmniQuant (3-bits, g128) 3.15 7.75 5.93 6.15 3.94
    OmniQuant (2-bits, g64) 2.28 11.78 7.60 8.90 5.65
    LREC (3-bits, g128) 3.35 8.24 - 5.52 -
    LREC (2-bits, g128) 2.24 12.52 - 8.74 -
    LLaMA-2 Uncompressed 16.0 6.97 5.52 5.47 3.31
    RTN (3-bits, g128) 3.15 8.40 6.02 6.66 3.97
    GPTQ (3-bits, g128) 3.15 7.89 5.85 6.29 3.85
    AWQ (3-bits, g128) 3.15 7.84 - 6.24 -
    OmniQuant (3-bits, g128) 3.15 7.75 5.85 6.03 3.78
    OmniQuant (2-bits, g64) 2.28 12.72 7.88 9.62 6.11
    LQ-LoRA (2.75-bits, 64-rank, Fisher) 2.95, 2.85 7.60 5.88 5.67 3.65

    When fine-tuned on a calibration set for post-training compression, LQ-LoRA at 2.85 effective bits on LLaMA-2-70B achieves 5.88 C4 perplexity and 3.65 WikiText perplexity, outperforming 3-bit PTQ baselines (such as GPTQ, AWQ, and OmniQuant at 3.15 effective bits).

  8. Knowl 8 — RoBERTa-Large Finetuning Performance on GLUE Benchmark

    data/table

    Average GLUE benchmark score across all tasks for RoBERTa-Large adapted with full fine-tuning, QLoRA, and LQ-LoRA under various bit-width budgets.

    Method Target Bits GLUE Average
    Full Fine-Tuning 16.0 88.5
    QLoRA 3-bit 3.127 86.1
    QLoRA (ILP) 2.50 75.4
    2.75 80.7
    3.00 85.5
    3.25 86.1
    LQ-LoRA 2.50 85.7
    2.75 87.1
    3.00 87.3
    3.25 88.1
    LQ-LoRA (Fisher) 2.50 87.3
    2.75 86.4
    3.00 87.3
    3.25 88.3

    LQ-LoRA outperforms QLoRA at every evaluated bit budget. At the aggressive 2.5-bit budget, standard QLoRA drops to 75.4, whereas Fisher-weighted LQ-LoRA preserves a score of 87.3, within 1.2 points of the unquantized 16-bit full fine-tuning score (88.5).

  9. Knowl 9 — Zero- and Few-Shot Open LLM Leaderboard Evaluation of Compressed Models

    data/table

    Zero- and few-shot evaluation across 6 core benchmarks from the Open LLM Leaderboard (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, GSM8K) using EleutherAI LM Evaluation Harness on LLaMA-2 models compressed with LQ-LoRA (Fisher, 2.75 target bits, rank 64, low-rank adapters quantized to NF-8).

    Method Size ARC HellaSwag MMLU TruthfulQA Winogrande GSM8K Average
    Uncompressed (16 bits) 7B 53.2 78.6 39.0 46.6 73.6 14.9 51.0
    LQ-LoRA (2.95 bits) 7B 49.8 75.9 39.3 43.0 72.4 7.4 48.0
    Uncompressed (16 bits) 70B 67.2 87.3 44.8 69.6 83.7 53.7 67.7
    LQ-LoRA (2.85 bits) 70B 65.8 86.2 44.5 66.9 83.2 45.6 65.3

    On the 70B model, compressing to 2.85 effective bits incurs only a 2.4-point drop in aggregate benchmark performance (67.7 to 65.3). However, mathematical reasoning capabilities (GSM8K) show noticeable degradation (53.7 down to 45.6 on 70B, and 14.9 down to 7.4 on 7B), demonstrating that downstream reasoning tasks can be more sensitive to aggressive quantization than standard perplexity metrics suggest.

  10. Knowl 10 — Sensitivity of LoRA Rank on Factorization Error and Language Modeling Perplexity

    data/table

    Evaluation of the impact of LoRA rank r∈{32,64,128}r \in \{32, 64, 128\} on LLaMA-2-7B with a fixed NF-3 quantization configuration (3.127 bits/parameter across all layers). Total error corresponds to ∑i∥W(i)−Quantize(W(i))∥F2\sum_i \|W^{(i)} - \text{Quantize}(W^{(i)})\|_F^2 for QLoRA and ∑i∥W(i)−(Q(i)+L1(i)L2(i))∥F2\sum_i \|W^{(i)} - (Q^{(i)} + L_1^{(i)} L_2^{(i)})\|_F^2 for LQ-LoRA across all matrices.

    Method LoRA rank (rr) C4 (PPL) WikiText (PPL) Reconstruction Error
    QLoRA 3-bit 32 8.21 6.75 9.83×1049.83 \times 10^4
    (3.127 bits/param) 64 8.21 6.76 9.83×1049.83 \times 10^4
    128 8.21 6.76 9.83×1049.83 \times 10^4
    LQ-LoRA 3-bit 32 8.02 6.61 7.99×1047.99 \times 10^4
    (3.127 bits/param) 64 7.93 6.51 7.12×1047.12 \times 10^4
    128 7.84 6.46 5.98×1045.98 \times 10^4

    Because standard QLoRA initializes L2=0L_2 = 0, the initialization error equals the unmitigated quantization error regardless of rank, resulting in invariant perplexity across rr. In contrast, LQ-LoRA leverages additional rank capacity during alternating decomposition to reduce Frobenius reconstruction error at initialization (7.99×104→5.98×1047.99 \times 10^4 \to 5.98 \times 10^4), yielding consistent perplexity improvements.

  11. Knowl 11 — Layer-Wise and Projection-Wise Quantization Sensitivity in LLMs

    empirical result

    Integer Linear Programming configuration search on LLaMA-2 reveals non-uniform quantization sensitivity across layer types and depths:

    1. Output projections (o_projo\_\text{proj}) and value projections (v_projv\_\text{proj}) become increasingly sensitive to quantization at deeper layers, with the ILP allocating higher bit rates (up to ~3.2 bits/parameter) to deeper layers for these matrices.
    2. Key projections (k_projk\_\text{proj}) and query projections (q_projq\_\text{proj}) become easier to quantize at deeper layers, receiving smaller bit allocations.
    3. MLP matrices (\text{gate_proj}, \text{down_proj}, \text{up_proj}) exhibit distinct sensitivity curves across layers.
    4. Incorporating Fisher information weighting alters the optimal ILP bit allocation relative to unweighted reconstruction, allocating higher precision to parameters with larger gradient variances on calibration data.
  12. Knowl 12 — PyTorch-Based Tensor Dispatch and CPU Optimizer Offloading for Sub-3-Bit LLM Finetuning

    model/method

    LQ-LoRA avoids hard-coded CUDA quantization kernels by implementing mixed-quantization entirely in PyTorch using torch.__torch_dispatch__ to duck-type torch.Tensor. This allows dynamic overloading of linear algebra operations (such as matrix multiplications) to trigger just-in-time dequantization, while the PyTorch full-graph compiler compiles bits-unpacking, dequantization, casting, and transposed matrix operations into efficient GPU kernels.

    To further reduce memory during adapter fine-tuning, LQ-LoRA implements CPU optimizer offloading: low-rank trainable parameters are mirrored on the CPU, parameter gradients are transferred from GPU to CPU after the backward pass, optimizer steps are executed on CPU, and updated parameters are asynchronously copied back to GPU memory overlapping with forward computations. On LLaMA-2-70B, this provides a 14% memory saving with less than 2% runtime overhead, enabling full forward/backward training of sub-3-bit 70B models on a single 80GB GPU (sequence length 2048, batch size 2).

  13. Knowl 13 — Limitations and Ineffective Architectural Extensions of LQ-LoRA

    limitation

    The LQ-LoRA methodology exhibits specific constraints and tested variations that failed to yield improvements:

    1. Lack of Theoretical Convergence: The alternating minimization between NormalFloat quantization and randomized SVD is a heuristic without formal convergence guarantees, requiring empirical early-stopping based on monitored reconstruction error.
    2. Periodic Re-Factorization Ineffectiveness: Re-decomposing weight matrices into low-rank and quantized components periodically during fine-tuning (e.g., every KK gradient steps) does not improve performance.
    3. Hybrid LoRA Initialization Ineffectiveness: Allocating half of the low-rank capacity to LQ-LoRA initialization and the remaining half to standard LoRA zero/Gaussian initialization fails to outperform pure LQ-LoRA initialization.
    4. Task-Agnostic ILP Objective: The ILP configuration search optimizes Frobenius reconstruction error rather than downstream task loss.
    5. Coupling to Low-Rank Reparameterization: The method fundamentally relies on low-rank updates to absorb initialization errors, precluding its direct application to other parameter-efficient fine-tuning paradigms like prompt tuning or prefix tuning.

Coverage note — None was omitted; all core methods, formulations, algorithms, experimental tables, analyses, and limitations from the paper are fully covered.

References

  1. 1.A. Aravkin, S. Becker, V. Cevher, and P. Olsen. A variational approach to stable principal component pursuit. In Conference on Uncertainty in Artificial Intelligence (UAI), July 2014.
  2. 2.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv:2303.12712, 2023.
  3. 3.HanQin Cai, Jialin Liu, and Wotao Yin. Learned Robust PCA: A Scalable Deep Unfolding Approach for High-Dimensional Outlier Detection. arXiv:2110.05649, 2021.
  4. 4.Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright. Robust principal component analysis? Journal of the ACM (JACM), 58(3):1–37, 2011.
  5. 5.Yuji Chai, John Gkountouras, Glenn G Ko, David Brooks, and Gu-Yeon Wei. Int2. 1: Towards fine-tunable quantized large language models with error correction through low-rank adaptation. arXiv preprint arXiv:2306.08162, 2023.
  6. 6.Jiahao Chen and Rene Ranftl. Deep robust pca using convolutional autoencoders. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2836–2840. IEEE, 2018.
  7. 7.Patrick Chen, Hsiang-Fu Yu, Inderjit Dhillon, and Cho-Jui Hsieh. DRONE: Data-aware Low-rank Compression for Large NLP Models. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 29321–29334. Curran Associates, Inc., 2021.
  8. 8.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  9. 9.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  10. 10.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. arXiv:2208.07339, 2022.
  11. 11.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023a.
  12. 12.Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078, 2023b.
  13. 13.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, Jing Yi, Weilin Zhao, Xiaozhi Wang, Zhiyuan Liu, Hai-Tao Zheng, Jianfei Chen, Yang Liu, Jie Tang, Juanzi Li, and Maosong Sun. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence, 2023.
  14. 14.Elias Frantar and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  15. 15.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training compression for generative pretrained transformers. arXiv preprint arXiv:2210.17323, 2022.
  16. 16.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  17. 17.Demi Guo, Alexander M. Rush, and Yoon Kim. Parameter-efficient transfer learning with diff pruning. In Proceedings of ACL, 2021.
  18. 18.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  19. 19.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  20. 20.Michael Hintermuller and Tao Wu. Robust Principal Component Pursuit via Inexact Alternating Minimization on Matrix Manifolds. Journal of Mathematical Imaging, 51:361–377, 2014.
  21. 21.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pp. 2790–2799. PMLR, 2019.
  22. 22.Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin. Language model compression with weighted low-rank factorization. In International Conference on Learning Representations, 2022.
  23. 23.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of ICLR, 2022.
  24. 24.Jeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, and Dongsoo Lee. Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization. arXiv preprint arXiv:2305.14152, 2023a.
  25. 25.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023b.
  26. 26.Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In Proceedings of ICLR, 2015.
  27. 27.Souvik Kundu, Shikai Wang, Qirui Sun, Peter A Beerel, and Massoud Pedram. Bmpq: bit-gradient sensitivity-driven mixed-precision quantization of dnns from scratch. In 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE), pp. 588–591. IEEE, 2022.
  28. 28.Se Jung Kwon, Jeonghoon Kim, Jeongin Bae, Kang Min Yoo, Jin-Hwa Kim, Baeseong Park, Byeongwook Kim, Jung-Woo Ha, Nako Sung, and Dongsoo Lee. AlphaTuning: Quantization-aware parameter-efficient adaptation of large-scale pre-trained language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 3288–3305, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.240.
  29. 29.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richard Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. OpenAssistant Conversations – Democratizing Large Language Model Alignment. arXiv:2304.07327, 2023.
  30. 30.Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. Owq: Lessons learned from activation outliers for weight quantization in large language models. arXiv preprint arXiv:2306.02272, 2023.
  31. 31.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045–3059, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.emnlp-main.243.
  32. 32.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of ACL, August 2021.
  33. 33.Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. Loftq: Lora-fine-tuning-aware quantization for large language models. arXiv preprint arXiv:2310.08659, 2023.
  34. 34.Yuanzhi Li, Yingyu Liang, and Andrej Risteski. Recovery guarantee of weighted low-rank approximation via alternating minimization. In International Conference on Machine Learning, pp. 2358–2367. PMLR, 2016.
  35. 35.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  36. 36.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, 2022.
  37. 37.Zhouchen Lin, Minming Chen, and Yi Ma. The Augmented Lagrange Multiplier Method for Exact Recovery of Corrupted Low-Rank Matrices. arXiv:1009.5055, 2010.
  38. 38.Guangcan Liu, Zhouchen Lin, Shuicheng Yan, Ju Sun, Yong Yu, and Yi Ma. Robust recovery of subspace structures by low-rank representation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1):171–184, 2013.
  39. 39.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  40. 40.Shiqian Ma and Necdet Serhat Aybat. Efficient Optimization Algorithms for Robust Principal Component Analysis and Its Variants. arXiv:1806.03430, 2018.
  41. 41.Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 565–576, 2021.
  42. 42.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016.
  43. 43.OpenAI. GPT-4 Technical Report. arXiv:2303.08774, 2023.
  44. 44.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, et al. Training language models to follow instructions with human feedback. arXiv:2203.02155, 2022.
  45. 45.Ilya Razenshteyn, Zhao Song, and David P Woodruff. Weighted low rank approximations with provable guarantees. In Proceedings of the forty-eighth annual ACM symposium on Theory of Computing, pp. 250–263, 2016.
  46. 46.Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. {ZeRO-Offload}: Democratizing {Billion-Scale} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pp. 551–564, 2021.
  47. 47.Rajarshi Saha, Varun Srivastava, and Mert Pilanci. Matrix compression via randomized low rank and low precision factorization. arXiv preprint arXiv:2310.11028, 2023.
  48. 48.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  49. 49.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, et al. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. arXiv:2211.05100, 2022.
  50. 50.Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. Omniquant: Omnidirectionally calibrated quantization for large language models. arXiv preprint arXiv:2308.13137, 2023.
  51. 51.Nathan Srebro and Tommi Jaakkola. Weighted low-rank approximations. In Proceedings of the 20th international conference on machine learning (ICML-03), pp. 720–727, 2003.
  52. 52.Yi-Lin Sung, Varun Nair, and Colin A Raffel. Training neural networks with fixed sparse masks. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 24193–24205. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/cb2653f548f8709598e8b5156738cc51-Paper.pdf.
  53. 53.Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia, Ali Ghodsi, and Mehdi Rezagholizadeh. KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation. arXiv:2109.06243, 2021.
  54. 54.Chen Tang, Kai Ouyang, Zhi Wang, Yifei Zhu, Wen Ji, Yaowei Wang, and Wenwu Zhu. Mixed-precision neural network quantization via learned layer-wise importance. In European Conference on Computer Vision, pp. 259–275. Springer, 2022.
  55. 55.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  56. 56.The Vicuna Team. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. 2023. https://lmsys.org/blog/2023-03-30-vicuna/.
  57. 57.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971, 2023a.
  58. 58.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  59. 59.Murad Tukan, Alaa Maalouf, Matan Weksler, and Dan Feldman. No fine-tuning, no cry: Robust svd for compressing deep networks. Sensors, 21(16):5599, August 2021. ISSN 1424-3210. doi: 10.3390/s21165599.
  60. 60.Elena Tuzhilina and Trevor Hastie. Weighted low rank matrix approximation and acceleration. arXiv preprint arXiv:2109.11057, 2021.
  61. 61.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2018.
  62. 62.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13484–13508, Toronto, Canada, July 2023.
  63. 63.John Wright, Arvind Ganesh, Shankar Rao, Yigang Peng, and Yi Ma. Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization. In Advances in Neural Information Processing Systems, volume 22, pp. 2080–2088, 2009.
  64. 64.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. arXiv:2211.10438, 2022.
  65. 65.Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael Mahoney, et al. Hawq-v3: Dyadic neural network quantization. In International Conference on Machine Learning, pp. 11875–11886. PMLR, 2021.
  66. 66.Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers. arXiv:2206.01861, 2022.
  67. 67.Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, and Yuxiong He. Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation. arXiv preprint arXiv:2303.08302, 2023.
  68. 68.Xinyang Yi, Dohyung Park, Yudong Chen, and Constantine Caramanis. Fast Algorithms for Robust PCA via Gradient Descent. arXiv:1605.07784, 2016. v1, last revised v2.
  69. 69.Davis Yoshida. Nf4 isn’t information theoretically optimal (and that’s good). arXiv preprint arXiv:2306.06965, 2023.
  70. 70.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 1–9, 2022.
  71. 71.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, 2019.
  72. 72.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, et al. Opt: Open pre-trained transformer language models. arXiv:2205.01068, 2022.
  73. 73.Teng Zhang and Yi Yang. Robust PCA by Manifold Optimization. arXiv:1708.00257, 2017.
  74. 74.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206, 2023.
  75. 75.Tianyi Zhou and Dacheng Tao. Godec: Randomized low-rank & sparse matrix decomposition in noisy case. Proceedings of the 28th International Conference on Machine Learning (ICML-11), 2011.
  76. 76.Tianyi Zhou and Dacheng Tao. Greedy bilateral sketch, completion & smoothing. In Carlos M. Carvalho and Pradeep Ravikumar (eds.), Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics, volume 31, 2013.

Citation

MLA
Guo, H., et al. “LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning”. arXiv, 2023, http://arxiv.org/abs/2311.12023v4.
APA
Guo, H., Greengard, P., Xing, E. P., & Kim, Y. (2023). LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning. arXiv. http://arxiv.org/abs/2311.12023v4
Chicago
Guo, H., P. Greengard, E. P. Xing, and Y. Kim. 2023. “LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning”. arXiv. http://arxiv.org/abs/2311.12023v4.
Harvard
Guo, H. et al. (2023) “LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.12023v4.
Vancouver
1. Guo H, Greengard P, Xing EP, Kim Y (2023) LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning. arXiv

BibTeX

@article{guo2023lora,
  title = {LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning},
  author = {Guo, Han and Greengard, Philip and Xing, Eric P. and Kim, Yoon},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.12023v4},
  eprint = {2311.12023}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors