DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Haokun LinHaobo XuYichen WuJingzhi CuiYingtao ZhangLinzhan MouLinqi SongZhenan SunYing Wei

article2024NeurIPS143 citations

Develops DuQuant, a quantization method combining block-wise rotation and zigzag permutation to redistribute extreme activation outliers across channels, enabling accurate and efficient 4-bit weight-activation large language model compression without heavy optimization.

Listen

Deploying massive large language models into commercial operations and edge hardware is severely constrained by substantial memory footprints and high inference latency. Quantization—the process of compressing floating-point numerical values into low-bit integers—offers an effective path to reduce operational overhead. However, when compressing both model weights and activation states to aggressive 4-bit formats, models routinely suffer severe performance degradation. This loss of accuracy is primarily driven by activation outliers, particularly newly recognized "massive outliers" that exhibit extreme magnitudes in feed-forward down-projection layers and resist conventional smoothing techniques.

The article introduces and evaluates DuQuant (Dual transformations Quantization), an innovative post-training quantization method designed to eliminate both persistent normal outliers and sparse massive outliers without requiring complex model retraining. DuQuant operates by redistributing extreme values across feature dimensions using structured mathematical transformations: it applies diagonal block-wise rotation matrices to locally smooth activations, alternates these with a zigzag permutation that balances activation magnitudes globally across blocks, and applies a secondary rotation to achieve a uniform activation landscape while simultaneously smoothing model weights.

The authors conducted extensive empirical evaluations across multiple open-source model families (LLaMA, LLaMA2, LLaMA3, Mistral, Phi-2, and Vicuna) spanning sizes from 2.8 billion to 70 billion parameters across diverse benchmarks, including language generation, commonsense reasoning, multitasking understanding, and long-context processing. Across standard 4-bit weight-activation benchmarks, DuQuant consistently outperformed existing state-of-the-art compression techniques. On commonsense reasoning tasks, it improved zero-shot accuracy by approximately 5% over Atom and 9% over QLLM across all LLaMA model sizes. In multitasking benchmarks on Vicuna-13B, it achieved up to a 10% gain over top baselines. In operational tests on LLaMA2-7B, DuQuant accelerated the initial prompt processing phase by up to 2.08× and reduced peak memory consumption during token generation by 3.50×, suffering only a 2.71% drop in accuracy relative to uncompressed full-precision models. Furthermore, its execution runtime is remarkably low, quantizing a 13-billion parameter model in approximately 71 seconds compared to hours required by optimization-based baselines.

These findings indicate that dual transformation post-training quantization resolves a critical bottleneck in deploying highly compressed language models at scale. By avoiding expensive gradient-based parameter training and maintaining near-lossless performance at 4-bit precision, organizations can drastically reduce enterprise hosting costs, lower latency, and facilitate deployment on resource-constrained edge devices.

Technical leaders and practitioners looking to optimize model deployment should consider adopting DuQuant pipelines for low-bit serving infrastructure. Organizations can implement the standard round-to-nearest formulation for maximum efficiency, or integrate learnable weight clipping if minor accuracy recovery is required for mission-critical tasks.

Decision-makers should note that the evaluation relied on standardized calibration sets using fixed sample sequences, although preliminary tests indicate the approach remains robust under randomly generated calibration data. The results provide high confidence in standard text-generation workloads, though application to specialized domains or non-transformer architectures should be validated via initial pilot testing.

Cover for DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Abstract

Quantization of large language models (LLMs) faces significant challenges, particularly due to the presence of outlier activations that impede efficient low-bit representation. Traditional approaches predominantly address Normal Outliers, which are activations across all tokens with relatively large magnitudes. However, these methods struggle with smoothing Massive Outliers that display significantly larger values, which leads to significant performance degradation in low-bit quantization. In this paper, we introduce DuQuant, a novel approach that utilizes rotation and permutation transformations to more effectively mitigate both massive and normal outliers. First, DuQuant starts by constructing the rotation matrix, using specific outlier dimensions as prior knowledge, to redistribute outliers to adjacent channels by block-wise rotation. Second, We further employ a zigzag permutation to balance the distribution of outliers across blocks, thereby reducing block-wise variance. A subsequent rotation further smooths the activation landscape, enhancing model performance. DuQuant simplifies the quantization process and excels in managing outliers, outperforming the state-of-the-art baselines across various sizes and types of LLMs on multiple tasks, even with 4-bit weight-activation quantization. Our code is available at https://github.com/Hsu1023/DuQuant.

Table of Contents

  • 1 Introduction
  • 2 Motivation
  • 3 Method
  • 3.1 Preliminaries
  • 3.2 The proposed DuQuant Method
  • 3.3 Theoretical Analysis
  • 4 Experiment
  • 4.1 Main Results
  • 4.2 Ablation Study
  • 5 Conclusion
  • References
  • A Related Work
  • A.1 Network Quantization
  • A.2 Post Training Quantization of LLM
  • B Proofs
  • C Additional Implementation Details
  • D More Empirical Results
  • E More Ablation Studies
  • E.1 Time Speedup and Memory Saving
  • E.2 Effects of Rotation Matrix
  • E.3 Effects of Permutation Algorithm.
  • E.4 Effects of Calibration Datasets
  • F Detailed Comparison with QuaRot
  • G Algorithm for Rotation Matrix
  • H Limitations and Broader Impacts
  • I More Visualizations

Knowls

  1. Knowl 1 — The DuQuant Dual Transformation Framework for LLM Quantization

    model/method

    DuQuant (Dual transformations Quantization) is a post-training weight-activation quantization framework designed to mitigate both normal and massive activation outliers in large language models (LLMs). For any linear layer Y=XWY = XW, where X∈RT×CinX \in \mathbb{R}^{T \times C_{\text{in}}} is the activation input across TT tokens and CinC_{\text{in}} channels, and W∈RCin×CoutW \in \mathbb{R}^{C_{\text{in}} \times C_{\text{out}}} is the weight matrix, DuQuant applies an invertible transformation matrix G∈RCin×CinG \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} to smooth activations and weights simultaneously:

    Y=XW=(XG)(G−1W)=X^W^Y = XW = (X G)(G^{-1} W) = \hat{X} \hat{W}

    The overall transformation matrix GG is composed of a diagonal smoothing matrix Λ\Lambda, a first block-diagonal orthogonal rotation matrix R^(1)\hat{R}_{(1)}, an orthogonal zigzag permutation matrix PP, and a second block-diagonal orthogonal rotation matrix R^(2)\hat{R}_{(2)}:

    G=Λ−1R^(1)PR^(2)G = \Lambda^{-1} \hat{R}_{(1)} P \hat{R}_{(2)} G−1=R^(2)⊤P⊤R^(1)⊤ΛG^{-1} = \hat{R}_{(2)}^\top P^\top \hat{R}_{(1)}^\top \Lambda

    Here:

    1. Λ=diag(Λ1,…,ΛCin)\Lambda = \text{diag}(\Lambda_1, \dots, \Lambda_{C_{\text{in}}}) is a per-channel smoothing diagonal matrix with Λj=max⁡(∣Xj∣)α/max⁡(∣Wj∣)1−α\Lambda_j = \max(|X_j|)^\alpha / \max(|W_j|)^{1-\alpha}, where XjX_j and WjW_j represent the jj-th column of XX and jj-th row of WW, respectively, and α∈[0,1]\alpha \in [0, 1] is a migration strength hyperparameter.
    2. R^(1),R^(2)∈RCin×Cin\hat{R}_{(1)}, \hat{R}_{(2)} \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} are orthogonal block-diagonal rotation matrices (hatRR^⊤=I\\hat{R}\hat{R}^\top = I, ∣det⁡(R^)∣=1|\det(\hat{R})| = 1) constructed block-wise to redistribute outlier values locally to adjacent channels within each block.
    3. P∈RCin×CinP \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} is an orthogonal permutation matrix (PP⊤=IPP^\top = I) that reorders activation channels using a zigzag pattern to balance outlier magnitudes uniformly across all blocks.

    Because G−1G^{-1} is multiplied into WW offline, the weights are simultaneously rotated and smoothed, mitigating weight outliers induced by Λ\Lambda without requiring runtime decompression overhead or slow optimization techniques such as GPTQ.

  2. Knowl 2 — Greedy Construction of Outlier-Mitigating Rotation Matrices

    algorithm

    In DuQuant, the block-diagonal rotation matrix R^∈RCin×Cin\hat{R} \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} is composed of diagonal blocks R^=BlockDiag(R^b1,…,R^bK)\hat{R} = \text{BlockDiag}(\hat{R}_{b_1}, \dots, \hat{R}_{b_K}), where each block R^bi∈R2n×2n\hat{R}_{b_i} \in \mathbb{R}^{2^n \times 2^n} operates on a sub-vector of dimension 2n2^n (K=Cin/2nK = C_{\text{in}} / 2^n). To minimize computation and memory overhead, DuQuant constructs the rotation matrix on the block containing the largest outlier and shares it across all KK blocks (setting R^bi=R^bk\hat{R}_{b_i} = \hat{R}_{b_k} for all 1≤i≤K1 \le i \le K).

    The greedy search algorithm identifies the feature dimension containing the highest activation outlier, swaps it to the first column via a permutation matrix EdE_{d}, applies an orthogonal initialized rotation matrix R~\tilde{R} whose first row is uniformly distributed (such as a normalized Hadamard matrix), rotates the remaining 2n−12^n - 1 columns with a random orthogonal matrix Q′Q', and iteratively accepts the rotation if it reduces the maximum activation magnitude max⁡i,j∣Xij∣\max_{i,j} |X_{ij}|.

    Input: Pre-initialized rotation matrix R~∈RCin×Cin\tilde{R} \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}}, greedy search steps NN, activation matrix X∈RT×CinX \in \mathbb{R}^{T \times C_{\text{in}}}
    Output: Approximated rotation matrix R^\hat{R}
    function get_rotation_matrix (XX, R~\tilde{R}, NN)
        T,Cin=X.shapeT, C_{\text{in}} = X\text{.shape}
        R=eye(Cin)R = \text{eye}(C_{\text{in}})
        a=max⁡i,j∣Xij∣a = \max_{i,j} |X_{ij}|
        R^=R\hat{R} = R
        for k=1,…,Nk = 1, \dots, N do
            channel_max=X.abs().max(dim=0).values\text{channel\_max} = X\text{.abs().max(dim}=0\text{).values}
            outlier_channel=arg⁡max⁡(channel_max)\text{outlier\_channel} = \arg\max(\text{channel\_max})
            Obtain randomly initialized orthogonal matrix Q′∈R(Cin−1)×(Cin−1)Q' \in \mathbb{R}^{(C_{\text{in}}-1) \times (C_{\text{in}}-1)}
            Q′=concat([zeros(Cin−1,1),Q′],dim=1)Q' = \text{concat}([\text{zeros}(C_{\text{in}}-1, 1), Q'], \text{dim}=1)
            Q=concat([zeros(1,Cin),Q′],dim=0)Q = \text{concat}([\text{zeros}(1, C_{\text{in}}), Q'], \text{dim}=0)
            Q[0,0]=1Q[0, 0] = 1
            R′=matmul(R~,Q)R' = \text{matmul}(\tilde{R}, Q)
            R′[:,outlier_channel],R′[:,0]=R′[:,0],R′[:,outlier_channel]R'[:, \text{outlier\_channel}], R'[:, 0] = R'[:, 0], R'[:, \text{outlier\_channel}]
            R′[outlier_channel,:],R′[0,:]=R′[0,:],R′[outlier_channel,:]R'[\text{outlier\_channel}, :], R'[0, :] = R'[0, :], R'[\text{outlier\_channel}, :]
            R=matmul(R,R′)R = \text{matmul}(R, R')
            X=matmul(X,R′)X = \text{matmul}(X, R')
            if max⁡i,j∣Xij∣<a\max_{i,j} |X_{ij}| < a then
                R^=R\hat{R} = R
                a=max⁡i,j∣Xij∣a = \max_{i,j} |X_{ij}|
            end if
        end for
        return R^\hat{R}
  3. Knowl 3 — Zigzag Permutation for Balancing Outliers Across Blocks

    model/method

    While block-diagonal rotation matrices efficiently redistribute outliers within isolated local blocks of size 2n2^n, they cannot exchange activation magnitudes across different blocks. This creates high variance in outlier severity across the K=Cin/2nK = C_{\text{in}} / 2^n blocks: some blocks contain extreme outliers while others have mild values.

    To minimize the inter-block variance Var([Mb1,Mb2,…,MbK])\text{Var}([M_{b_1}, M_{b_2}, \dots, M_{b_K}]), where MbiM_{b_i} is the mean of the maximum outlier magnitudes OjO_j in the ii-th block, DuQuant applies an orthogonal zigzag permutation matrix P∈RCin×CinP \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} (PP⊤=IPP^\top = I).

    The zigzag permutation proceeds as follows:

    1. Sort all CinC_{\text{in}} channels in descending order of their absolute maximum activation values: O(1)≥O(2)≥⋯≥O(Cin)O^{(1)} \ge O^{(2)} \ge \dots \ge O^{(C_{\text{in}})}.
    2. Distribute the sorted channels across the KK blocks in an alternating, back-and-forth pattern:
      • Assign the highest remaining activation channels sequentially to block 1,2,…,K1, 2, \dots, K.
      • Reverse direction and assign the next channels in ascending order to block K,K−1,…,1K, K-1, \dots, 1.
      • Repeat this serpentine assignment for 2n−12^{n-1} rounds until all Cin=K⋅2nC_{\text{in}} = K \cdot 2^n channels are assigned.

    Specifically, the ii-th block receives the set of channel indices:

    bi={Oc(2mK+i),Oc(2(m+1)K−i+1)∣m=0,1,…,2n−1−1}b_i = \{O_c^{(2mK + i)}, O_c^{(2(m+1)K - i + 1)} \mid m = 0, 1, \dots, 2^{n-1}-1\}

    where Oc(r)O_c^{(r)} denotes the channel index having the rr-th largest outlier.

    This distribution ensures that no single block receives a disproportionate concentration of high or low activation channels, equalizing the mean outlier magnitude across blocks and enabling an effective subsequent block-rotation R^(2)\hat{R}_{(2)}.

  4. Knowl 4 — Normal Outliers versus Massive Outliers in LLMs

    definition

    In large language models (LLMs), activation outliers are divided into two distinct categories based on their distribution, magnitude, and architectural location:

    1. Normal Outliers:

      • Magnitude: Moderately large values relative to typical activations.
      • Distribution: Persist consistently across all token positions in a small subset of feature channels.
      • Location: Present across multiple linear layers and attention modules (e.g., attention key projections, up projections).
      • Mitigation: Can be partially mitigated by per-channel scaling/smoothing methods (such as SmoothQuant).
    2. Massive Outliers:

      • Magnitude: Exceedingly high values (often exceeding 100 to 1400 in magnitude, roughly 1000 times larger than the median activation).
      • Distribution: Confined to very few tokens (isolated to one or a small subset of token positions across a sequence) rather than persisting across all tokens.
      • Location: Occur systematically at the input activations of the down-projection layer in the Feed-Forward Network (FFN) module across various Transformer architectures (e.g., LLaMA, Mistral, Vicuna).
      • Quantization Impact: Because massive outliers are confined to specific tokens, per-channel smoothing divides the activation by an excessively large scaling factor Λj\Lambda_j, which shifts the massive outlier directly into the weight matrix WW, causing severe quantization degradation or gradient instability in optimization-based quantization.
  5. Knowl 5 — Maximum Outlier Reduction via Block-Wise Rotation

    theoretical result

    Let X∈RT×CinX \in \mathbb{R}^{T \times C_{\text{in}}} denote an activation tensor, and let R^∈RCin×Cin\hat{R} \in \mathbb{R}^{C_{\text{in}} \times C_{\text{in}}} be a block-diagonal rotation matrix R^=BlockDiag(R^b1,…,R^bK)\hat{R} = \text{BlockDiag}(\hat{R}_{b_1}, \dots, \hat{R}_{b_K}) where each R^bi∈R2n×2n\hat{R}_{b_i} \in \mathbb{R}^{2^n \times 2^n} is an orthogonal matrix (R^biR^bi⊤=I\hat{R}_{b_i} \hat{R}_{b_i}^\top = I).

    For a given block bib_i spanning 2n2^n channels, let Xbi∈RT×2nX_{b_i} \in \mathbb{R}^{T \times 2^n} denote the activation sub-matrix, and let Oj(Xbi)=max⁡t=1T∣(Xbi)tj∣O_j(X_{b_i}) = \max_{t=1}^T |(X_{b_i})_{tj}| denote the maximum absolute activation outlier in the jj-th dimension of block bib_i. Under the greedy rotation construction where the largest outlier dimension d(1)d^{(1)} is swapped to the first column and transformed with an orthogonal matrix having a uniformly distributed first row (1/2n1/\sqrt{2^n}), the rotated maximum outlier satisfies:

    max⁡1≤j≤2nOj(XbiR^bi)≤max⁡1≤j≤2nOj(Xbi)\max_{1 \le j \le 2^n} O_j(X_{b_i} \hat{R}_{b_i}) \le \max_{1 \le j \le 2^n} O_j(X_{b_i})

    This confirms that the greedy rotation transformation strictly bounds the post-rotation maximum outlier within each block by the initial maximum outlier of that block.

  6. Knowl 6 — Balanced Mean Outlier Bound via Zigzag Permutation

    theoretical result

    Let X∈RT×CinX \in \mathbb{R}^{T \times C_{\text{in}}} be divided into KK blocks of dimension 2n2^n, where K=Cin/2nK = C_{\text{in}} / 2^n. Let OjO_j be the maximum absolute outlier of channel djd_j in XX, and let O(1)≥O(2)≥⋯≥O(Cin)O^{(1)} \ge O^{(2)} \ge \dots \ge O^{(C_{\text{in}})} represent the sorted sequence of maximum channel outliers.

    Define the maximum consecutive outlier gap as:

    δ:=max⁡i∈{1,…,Cin−1}∣O(i+1)−O(i)∣\delta := \max_{i \in \{1, \dots, C_{\text{in}}-1\}} |O^{(i+1)} - O^{(i)}|

    Let Mbi=12n∑j∈biOjM_{b_i} = \frac{1}{2^n} \sum_{j \in b_i} O_j denote the mean maximum outlier magnitude within the ii-th block after applying the zigzag permutation. For every block i∈{1,2,…,K}i \in \{1, 2, \dots, K\}, the block mean MbiM_{b_i} satisfies the uniform upper bound:

    Mbi≤O(1)+(2nK−1)(2n−1−1)2nδM_{b_i} \le O^{(1)} + \frac{(2^n K - 1)(2^{n-1} - 1)}{2^n} \delta

    Because all KK blocks share this identical theoretical upper bound, the zigzag permutation guarantees a balanced outlier distribution across blocks and minimizes the inter-block variance Var([Mb1,…,MbK])\text{Var}([M_{b_1}, \dots, M_{b_K}]).

  7. Knowl 7 — 4-Bit Weight-Activation Quantization Perplexity on LLaMA Models

    data/table

    The table below compares the language modeling perplexity (PPL, lower is better) under 4-bit weight and 4-bit activation (W4A4) post-training quantization on WikiText2 and C4 datasets across the LLaMA1 and LLaMA2 model families. DuQuant and DuQuant+LWC (which incorporates Learnable Weight Clipping) outperform prior post-training quantization baselines, maintaining low perplexity close to the full-precision FP16 models.

    Dataset Method L1-7B L1-13B L1-30B L1-65B L2-7B L2-13B L2-70B
    WikiText2 FP16 5.68 5.09 4.10 3.53 5.47 4.88 3.31
    (W4A4) SmoothQuant 25.25 40.05 192.40 275.53 83.12 35.88 26.01
    OmniQuant 11.26 10.87 10.33 9.17 14.26 12.30 NaN
    AffineQuant 10.28 10.32 9.35 - 12.69 11.45 -
    QLLM 9.65 8.41 8.37 6.87 11.75 9.09 7.00
    Atom 8.15 7.43 6.52 5.14 8.40 6.96 NaN
    DuQuant 6.40 5.65 4.72 4.13 6.28 5.42 3.79
    DuQuant+LWC 6.18 5.47 4.55 3.93 6.08 5.33 3.76
    C4 FP16 7.08 6.61 5.98 5.62 6.97 6.46 5.52
    (W4A4) SmoothQuant 32.32 47.18 122.38 244.35 77.27 43.19 34.61
    OmniQuant 14.51 13.78 12.49 11.28 18.02 14.55 NaN
    AffineQuant 13.64 13.44 11.58 - 15.76 13.97 -
    QLLM 12.29 10.58 11.51 8.98 13.26 11.13 8.89
    Atom 10.34 9.57 8.56 8.17 10.96 9.12 NaN
    DuQuant 7.84 7.16 6.45 6.03 7.90 7.05 5.87
    DuQuant+LWC 7.73 7.07 6.37 5.93 7.79 7.02 5.85

    On LLaMA2-7B, DuQuant achieves a WikiText2 perplexity of 6.28 (and 6.08 with LWC), compared to 83.12 for SmoothQuant, 14.26 for OmniQuant, and 8.40 for Atom. On LLaMA2-70B, where Atom and OmniQuant encounter numerical instability/NaN due to massive outliers and group-query attention, DuQuant achieves 3.79 PPL compared to 3.31 for FP16.

  8. Knowl 8 — Zero-Shot Commonsense QA Accuracy under 4-Bit WA Quantization

    data/table

    The table below presents the zero-shot accuracy (%) across six commonsense reasoning benchmarks (PIQA, ARC-E, ARC-C, BoolQ, HellaSwag, and WinoGrande) for LLaMA1 models quantized to 4-bit weights and 4-bit activations (W4A4).

    Model Method PIQA ARC-E ARC-C BoolQ HellaSwag WinoGrande Avg.
    LLaMA1-7B FP16 77.47 52.48 41.46 73.08 73.00 67.07 64.09
    SmoothQuant 49.80 30.40 25.80 49.10 27.40 48.00 38.41
    OS+ 62.73 39.98 30.29 60.21 44.39 52.96 48.43
    OmniQuant 66.15 45.20 31.14 63.51 56.44 53.43 52.65
    AffineQuant 69.37 42.55 31.91 63.73 57.65 55.33 53.42
    QLLM 68.77 45.20 31.14 - 57.43 56.67 51.84
    Atom 71.44 47.74 35.49 67.71 63.89 55.01 56.88
    DuQuant 76.44 50.04 38.99 70.98 69.39 64.72 61.76
    DuQuant+LWC 76.22 50.04 38.31 70.09 69.82 62.59 61.18
    LLaMA1-13B FP16 79.10 59.89 44.45 68.01 76.21 70.31 66.33
    Atom 71.38 49.07 36.69 64.53 68.00 58.56 58.04
    DuQuant 77.26 58.04 41.55 67.55 73.62 66.69 64.12
    DuQuant+LWC 77.64 57.32 41.21 66.79 74.12 65.98 63.84
    LLaMA1-30B FP16 80.08 58.92 45.47 68.44 79.21 72.53 67.44
    Atom 71.98 49.07 40.02 66.85 70.45 58.64 59.50
    DuQuant 78.56 56.99 42.32 66.73 76.70 69.61 65.15
    DuQuant+LWC 78.73 56.52 43.17 68.84 77.53 70.96 65.96
    LLaMA1-65B FP16 80.79 58.71 46.24 82.29 80.72 77.50 71.04
    Atom 74.48 51.60 40.61 73.76 73.78 62.12 62.73
    DuQuant 79.71 57.95 45.05 79.82 78.66 72.29 68.91
    DuQuant+LWC 79.98 58.29 44.80 77.89 79.22 72.21 68.73

    DuQuant improves average commonsense QA performance by approximately 5% over Atom across all model sizes (e.g., reaching 61.76% vs 56.88% on LLaMA1-7B, and 68.91% vs 62.73% on LLaMA1-65B).

  9. Knowl 9 — Robustness on LLaMA3-8B and LLaMA3-70B under W4A4 Quantization

    empirical result

    Low-bit quantization of LLaMA3 models typically causes severe accuracy collapse due to sensitive activation patterns. When evaluated on LLaMA3-8B under 4-bit weight-activation (W4A4) quantization, baseline methods degrade dramatically:

    • OmniQuant achieves a WikiText2 perplexity of 3.64×1033.64 \times 10^3 and an average commonsense QA accuracy of 36.08%.
    • AffineQuant achieves a WikiText2 perplexity of 21.21×10321.21 \times 10^3 and an average QA accuracy of 36.29%.
    • SmoothQuant yields a perplexity of 210.19 and an average QA accuracy of 40.97%.
    • Atom yields a perplexity of 22.14 and an average QA accuracy of 52.10%.

    In contrast, DuQuant maintains strong performance:

    • DuQuant achieves a WikiText2 perplexity of 8.56, C4 perplexity of 11.98, and an average QA accuracy of 66.21% (compared to 74.22% for FP16).
    • DuQuant+LWC achieves a WikiText2 perplexity of 8.06, C4 perplexity of 11.29, and an average QA accuracy of 67.72%.

    On LLaMA3-70B under W4A4:

    • SmoothQuant exhibits a WikiText2 perplexity of 9.6 and an average QA accuracy of 63.7%.
    • DuQuant achieves a WikiText2 perplexity of 4.9 (compared to 2.9 for FP16) and an average QA accuracy of 76.6% (compared to 80.1% for FP16), demonstrating robustness against low-bit degradation on modern architectures without needing learnable clipping.
  10. Knowl 10 — Downstream Evaluation on Instruction-Tuned Vicuna and LongBench

    empirical result

    When evaluated on instruction-tuned models (Vicuna-v1.5-7B and Vicuna-v1.5-13B) under W4A4 quantization:

    1. MMLU Benchmark:

      • On Vicuna-v1.5-13B, DuQuant achieves 50.94% (0-shot) and 51.96% (5-shot) average accuracy, while DuQuant+LWC achieves 51.08% (0-shot) and 51.61% (5-shot), closely tracking FP16 (54.54% 0-shot, 55.78% 5-shot).
      • In comparison, SmoothQuant achieves 22.82% (0-shot), OmniQuant achieves 28.12% (0-shot), and Atom achieves 41.07% (0-shot).
    2. Long-Context Generation (LongBench):

      • On LongBench (sequence length up to 3500 tokens across 13 datasets spanning single-doc QA, multi-doc QA, summarization, few-shot tasks, and code completion), Vicuna-v1.5-7B quantized with DuQuant reaches an average score of 40.37 (FP16 baseline: 41.83), outperforming Atom (34.34), SmoothQuant (7.56), and OmniQuant (3.43).
      • For Vicuna-v1.5-13B, DuQuant achieves an average score of 40.32 (FP16 baseline: 42.45), compared to Atom (37.36), OmniQuant (4.77), and SmoothQuant (3.83).
    3. MT-Bench (GPT-4 Evaluation):

      • In head-to-head GPT-4 evaluations on Vicuna-v1.5-7B responses, DuQuant beats Atom with 68 wins vs 16 losses (78 ties), and beats OmniQuant with 155 wins vs 1 loss (44 ties).
      • Against FP16 Vicuna-v1.5-7B, DuQuant achieves 36 wins, 56 ties, and 68 losses.
  11. Knowl 11 — Quantization Runtime, Memory Footprint, and Inference Speedup

    empirical result

    DuQuant offers substantial efficiency gains during both calibration/quantization and online inference:

    1. Quantization Runtime: On a single NVIDIA A100 GPU, DuQuant quantizes models without iterative backpropagation or gradient updates:

      • LLaMA2-7B is quantized in 50 seconds (compared to 20 minutes for Atom, 1.1 hours for QLLM, 2.0 hours for OmniQuant, and 9.1 hours for AffineQuant).
      • LLaMA2-13B is quantized in 71 seconds (compared to 36 minutes for Atom, 1.7 hours for QLLM, and 16.0 hours for AffineQuant).
      • LLaMA2-70B is quantized in 270 seconds (compared to 3.5 hours for Atom, 9.3 hours for QLLM, and 18.6 hours for AffineQuant).
    2. Peak Memory Consumption: On LLaMA2-7B running on an NVIDIA RTX 3090 GPU (batch size 1, sequence length 2048, 128 decoding steps):

      • Pre-filling peak memory is reduced from 15.282 GB (FP16) to 4.786 GB (3.193×\times memory reduction).
      • Decoding peak memory is reduced from 13.638 GB (FP16) to 3.893 GB (3.503×\times memory reduction).
    3. Inference Latency Speedup:

      • Layer-wise pre-filling speedup over FP16 reaches 1.95×\times (batch size 1), 2.03×\times (batch size 4), and 2.08×\times (batch size 16) on LLaMA2-7B. On LLaMA2-13B, pre-filling speedup reaches 2.34×\times (batch size 16).
      • End-to-end pre-filling latency on LLaMA2-7B drops from 1449 ms to 720 ms (2.01×\times speedup at batch size 3).
  12. Knowl 12 — Experimental Setup and Hyperparameters for DuQuant

    experimental setup

    The standard experimental configuration for DuQuant is as follows:

    1. Quantization Format:

      • Linear layer activations: per-token asymmetric integer uniform quantization.
      • Linear layer weights: per-channel asymmetric integer uniform quantization.
      • Self-attention matrix multiplications (Query ×\times Key, attention output ×\times Value): transformed via fixed Hadamard rotation matrices (±1/2n\pm 1 / \sqrt{2^n}) with per-token / per-head asymmetric quantization. SoftMax outputs are kept unquantized.
      • Default evaluation precision is 4-bit (W4A4) and 6-bit (W6A6).
    2. Calibration Configuration:

      • Calibration set: 128 randomly sampled sequences from WikiText2, each with a sequence length of 2048 tokens.
      • Activation clipping: clipping ratio of 0.9 on maximum activations across all projection blocks.
      • Weight clipping (DuQuant): clipping ratio of 0.8 on maximum weight magnitudes.
      • Smooth strength α\alpha: set to α=0.6\alpha = 0.6 for DuQuant and α=0.5\alpha = 0.5 for DuQuant+LWC.
    3. Transformation Parameters:

      • Block size: 2n=1282^n = 128.
      • Maximum greedy search steps: N=256N = 256.
      • Permutation frequency: 1 permutation step flanked by 2 rotation steps (R(1)→P→R(2)R_{(1)} \rightarrow P \rightarrow R_{(2)}).
    4. Learnable Weight Clipping (DuQuant+LWC):

      • When enabled, learnable parameters γ,β∈[0,1]\gamma, \beta \in [0, 1] parameterize step size Δ=γmax⁡(X)−βmin⁡(X)2b−1\Delta = \frac{\gamma \max(X) - \beta \min(X)}{2^b - 1}. Training runs for 20 epochs with batch size 1, learning rate 5×10−35 \times 10^{-3}, and zero weight decay.
  13. Knowl 13 — Component and Hyperparameter Ablations of DuQuant

    empirical result

    Ablation studies on LLaMA2-7B and LLaMA2-13B demonstrate the specific role of each element in DuQuant:

    1. Component Ablation (W4A4):

      • Smoothing alone (Λ\Lambda only): WikiText2 PPL is NaN on LLaMA2-7B and 160.30 on LLaMA2-13B.
      • Smoothing + Initial Rotation (R(1)R_{(1)}): PPL drops to 8.48 (7B) and 14.32 (13B).
      • Smoothing + Rotation 1 + Permutation (R(1)⋅PR_{(1)} \cdot P): PPL drops to 7.92 (7B) and 5.96 (13B).
      • Full DuQuant (R(1)⋅P⋅R(2)R_{(1)} \cdot P \cdot R_{(2)}): PPL achieves 6.28 (7B) and 5.42 (13B).
      • Without smoothing (Rotation + Permutation + Rotation only): PPL degrades to 6.79 (7B) and 6.06 (13B).
    2. Normal vs. Massive Outlier Isolation: Applying only the smooth technique to the down-projection layer (where massive outliers reside) while using full DuQuant elsewhere results in severe degradation: WikiText2 PPL spikes to 18.16 (7B) and 10.51 (13B). Conversely, applying smoothing only to normal outliers while using DuQuant on massive outliers yields 10.88 PPL (7B) and 7.87 PPL (13B).

    3. Rotation Block Size and Greedy Steps:

      • Varying block size 2n2^n from 4 to 128 monotonically decreases WikiText2 PPL from 18.69 to 6.28 (7B) and reduces runtime from 64.4 s to 48.6 s.
      • Varying greedy search steps NN improves PPL from 6.60 (N=1N=1) to 6.28 (N=256N=256), while N=1024N=1024 overfits (6.31 PPL).
    4. Permutation Method Comparison: Zigzag permutation achieves 6.28 PPL (7B) and a block variance of 3.0×10−43.0 \times 10^{-4} in 48.6 s, closely matching Simulated Annealing (6.26 PPL, variance 1.7×10−41.7 \times 10^{-4}) while being over 15×15\times faster (769.6 s for Simulated Annealing).

  14. Knowl 14 — Calibration Data Heuristic and Calibration-Free Robustness

    limitation

    The primary limitation of DuQuant is its reliance on a heuristic calibration selection: by default, it uses 128 randomly sampled sequences from WikiText2 to compute mean activation embeddings that guide the greedy rotation search and zigzag channel ordering.

    However, ablation studies indicate that DuQuant exhibits low sensitivity to the calibration set:

    • Varying calibration datasets between WikiText2 (6.28 WikiText2 PPL / 7.90 C4 PPL on LLaMA2-7B) and C4 (6.25 WikiText2 PPL / 7.87 C4 PPL) produces virtually identical quantization performance.
    • When evaluated in a completely calibration-free regime using randomly generated tokens sampled uniformly from the model's vocabulary, DuQuant achieves a WikiText2 PPL of 6.25 and C4 PPL of 7.86 on LLaMA2-7B, and 5.45 WikiText2 PPL / 7.05 C4 PPL on LLaMA2-13B (compared to 5.44 / 7.05 using real WikiText2 data).
    • Varying the number of calibration sequences from 16 to 256 yields stable WikiText2 PPL between 6.29 and 6.23.

    This insensitivity occurs because massive and normal outliers are structural properties of the model's weights (specifically the GLU down-projection layers) rather than artifacts of specific calibration prompts.

Coverage note — Detailed subtask breakdown tables for LongBench and 6-bit common-sense QA tables were omitted as the aggregated benchmark tables and 4-bit main results fully capture the core empirical conclusions.

References

  1. 1.Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, and Dan Alistarh. Towards end-to-end 4-bit inference on generative large language models. arXiv preprint arXiv:2310.09259, 2023.
  2. 2.Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. Quarot: Outlier-free 4-bit inference in rotated llms. arXiv preprint arXiv:2404.00456, 2024.
  3. 3.Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. Binarybert: Pushing the limit of bert quantization. arXiv preprint arXiv:2012.15701, 2020.
  4. 4.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508, 2023.
  5. 5.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  6. 6.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  8. 8.Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher M De Sa. Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems, 36, 2024.
  9. 9.Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Yu Qiao, and Ping Luo. Efficientqat: Efficient quantization-aware training for large language models. arXiv preprint arXiv:2407.11062, 2024.
  10. 10.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  11. 11.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2924–2936, 2019.
  12. 12.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  13. 13.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm.int8(): 8-bit matrix multiplication for transformers at scale. In Conference on Neural Information Processing Systems, 2022.
  14. 14.Tim Dettmers, Ruslan A. Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. SpQR: A sparse-quantized representation for near-lossless LLM weight compression. In The Twelfth International Conference on Learning Representations, 2024.
  15. 15.Peijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu, Xinglin Pan, Qiang Wang, and Xiaowen Chu. Pruner-zero: Evolving symbolic pruning metric from scratch for large language models. In Proceedings of the 41st International Conference on Machine Learning. PMLR, 2024. URL https://arxiv.org/abs/2406.02924. [arXiv: 2406.02924].
  16. 16.Dayou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo, Ting Cao, Xiaowen Chu, and Ningyi Xu. Bitdistiller: Unleashing the potential of sub-4-bit llms via self-distillation. arXiv preprint arXiv:2402.10631, 2024.
  17. 17.Haojie Duanmu, Zhihang Yuan, Xiuhong Li, Jiangfei Duan, Xingcheng Zhang, and Dahua Lin. Skvq: Sliding-window key and value cache quantization for large language models. arXiv preprint arXiv:2405.06219, 2024.
  18. 18.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  19. 19.Yuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng, Xiaoyi Bao, Shijie Ma, Wei Zou, and Yun Zheng. Crossmae: Cross-modality masked autoencoders for region-aware audio-visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26721–26731, 2024.
  20. 20.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015.
  21. 21.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  22. 22.Jung Hwan Heo, Jeonghoon Kim, Beomseok Kwon, Byeongwook Kim, Se Jung Kwon, and Dongsoo Lee. Rethinking channel dimensions to isolate outliers for low-bit weight quantization of large language models. arXiv preprint arXiv:2309.15531, 2023.
  23. 23.Lu Hou, Quanming Yao, and James T Kwok. Loss-aware binarization of deep networks. arXiv preprint arXiv:1611.01600, 2016.
  24. 24.Wei Huang, Xudong Ma, Haotong Qin, Xingyu Zheng, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, and Michele Magno. How good are low-bit quantized llama3 models? an empirical study. arXiv preprint arXiv:2404.14047, 2024.
  25. 25.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704–2713, 2018.
  26. 26.Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al. Phi-2: The surprising power of small language models. Microsoft Research Blog, 2023.
  27. 27.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  28. 28.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023.
  29. 29.Liang Li, Qingyuan Li, Bo Zhang, and Xiangxiang Chu. Norm tweaking: High-performance low-bit quantization of large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18536–18544, 2024.
  30. 30.Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang. Evaluating quantized large language models. arXiv preprint arXiv:2402.18158, 2024.
  31. 31.Haokun Lin, Haoli Bai, Zhili Liu, Lu Hou, Muyi Sun, Linqi Song, Ying Wei, and Zhenan Sun. Mope-clip: Structured pruning for efficient vision-language models with module-wise pruning error metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27370–27380, 2024.
  32. 32.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  33. 33.Akide Liu, Jing Liu, Zizheng Pan, Yefei He, Gholamreza Haffari, and Bohan Zhuang. Minicache: Kv cache compression in depth dimension for large language models. arXiv preprint arXiv:2405.14366, 2024.
  34. 34.Jing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong, Jianfei Cai, and Bohan Zhuang. QLLM: Accurate and efficient low-bitwidth quantization for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=FIplmUWdm3.
  35. 35.Ruikang Liu, Haoli Bai, Haokun Lin, Yuening Li, Han Gao, Zhengzhuo Xu, Lu Hou, Jun Yao, and Chun Yuan. IntactKV: Improving large language model quantization by keeping pivot tokens intact. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics ACL 2024, pages 7716–7741, Bangkok, Thailand and virtual meeting, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.460. URL https://aclanthology.org/2024.findings-acl.460.
  36. 36.Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888, 2023.
  37. 37.Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750, 2024.
  38. 38.Shijie Ma, Fei Zhu, Zhun Zhong, Wenzhuo Liu, Xu-Yao Zhang, and Cheng-Lin Liu. Happy: A debiased learning framework for continual generalized category discovery. arXiv preprint arXiv:2410.06535, 2024.
  39. 39.Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, Xuefeng Xiao, Rui Wang, Shilei Wen, Fei Chao, and Rongrong Ji. Affinequant: Affine transformation quantization for large language models. arXiv preprint arXiv:2403.12544, 2024.
  40. 40.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations, 2016.
  41. 41.Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pages 7197–7206. PMLR, 2020.
  42. 42.Davide Paglieri, Saurabh Dash, Tim Rocktäschel, and Jack Parker-Holder. Outliers and calibration sets have diminishing effect on quantization of modern llms. arXiv preprint arXiv:2405.20835, 2024.
  43. 43.Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative llm inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132. IEEE, 2024.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  45. 45.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  46. 46.Yuzhang Shang, Zhihang Yuan, and Zhen Dong. PB-LLM: Partially binarized large language models. In The Twelfth International Conference on Learning Representations, 2024.
  47. 47.Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. Omniquant: Omnidirectionally calibrated quantization for large language models. In The Twelfth International Conference on Learning Representations, 2023.
  48. 48.Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu. Massive activations in large language models. arXiv preprint arXiv:2402.17762, 2024.
  49. 49.Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong. Compression of generative pre-trained language models via quantization. arXiv preprint arXiv:2203.10705, 2022.
  50. 50.Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong. Scaling laws with vocabulary: Larger models deserve larger vocabularies. arXiv preprint arXiv:2407.13623, 2024.
  51. 51.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  52. 52.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  53. 53.Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks. arXiv preprint arXiv:2402.04396, 2024.
  54. 54.Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, et al. Efficient large language models: A survey. arXiv preprint arXiv:2312.03863, 1, 2023.
  55. 55.Zhongwei Wan, Xinjian Wu, Yu Zhang, Yi Xin, Chaofan Tao, Zhihong Zhu, Xin Wang, Siqi Luo, Jing Xiong, and Mi Zhang. D2o: Dynamic discriminative operations for efficient generative inference of large language models. arXiv preprint arXiv:2406.13035, 2024.
  56. 56.Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, and Li Yuan. Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference. arXiv preprint arXiv:2406.18139, 2024.
  57. 57.Haoxuan Wang, Yuzhang Shang, Zhihang Yuan, Junyi Wu, and Yan Yan. Quest: Low-bit diffusion model quantization via efficient selective finetuning. arXiv preprint arXiv:2402.03666, 2024.
  58. 58.Xin Wang, Yu Zheng, Zhongwei Wan, and Mi Zhang. Svd-llm: Truncation-aware singular value decomposition for large language model compression. arXiv preprint arXiv:2403.07378, 2024.
  59. 59.Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, and Xianglong Liu. Outlier suppression+: Accurate quantization of large language models by equivalent and effective shifting and scaling. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1648–1665, 2023.
  60. 60.Junyi Wu, Haoxuan Wang, Yuzhang Shang, Mubarak Shah, and Yan Yan. Ptq4dit: Post-training quantization for diffusion transformers. arXiv preprint arXiv:2405.16005, 2024.
  61. 61.Yichen Wu, Long-Kai Huang, and Ying Wei. Adversarial task up-sampling for meta-learning. Advances in Neural Information Processing Systems, 35:31102–31115, 2022.
  62. 62.Yichen Wu, Long-Kai Huang, Renzhen Wang, Deyu Meng, and Ying Wei. Meta continual learning revisited: Implicitly enhancing online hessian approximation via variance reduction. In The Twelfth International Conference on Learning Representations, 2024.
  63. 63.Haocheng Xi, Changhao Li, Jianfei Chen, and Jun Zhu. Training transformers with 4-bit integers. Advances in Neural Information Processing Systems, 36:49146–49168, 2023.
  64. 64.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pages 38087–38099. PMLR, 2023.
  65. 65.Haobo Xu, Yuchen Yan, Dingsu Wang, Zhe Xu, Zhichen Zeng, Tarek F Abdelzaher, Jiawei Han, and Hanghang Tong. Slog: An inductive spectral graph neural network beyond polynomial filter. In Forty-first International Conference on Machine Learning, 2024.
  66. 66.Jaewoo Yang, Hayun Kim, and Younghoon Kim. Mitigating quantization errors due to activation spikes in glu-based llms. arXiv preprint arXiv:2405.14428, 2024.
  67. 67.Lianwei Yang, Zhikai Li, Junrui Xiao, Haisong Gong, and Qingyi Gu. Mgrq: Post-training quantization for vision transformer with mixed granularity reconstruction. arXiv preprint arXiv:2406.09229, 2024.
  68. 68.Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. In Conference on Neural Information Processing Systems, 2022.
  69. 69.Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. Rptq: Reorder-based post-training quantization for large language models. arXiv preprint arXiv:2304.01089, 2023.
  70. 70.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.
  71. 71.Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. Ternarybert: Distillation-aware ultra-low bit bert. arXiv preprint arXiv:2009.12812, 2020.
  72. 72.Yingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao, Lu Hou, and Carlo Vittorio Cannistraci. Plug-and-play: An efficient post-training pruning method for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Tr0lPx9woF.
  73. 73.Tianchen Zhao, Xuefei Ning, Tongcheng Fang, Enshu Liu, Guyue Huang, Zinan Lin, Shengen Yan, Guohao Dai, and Yu Wang. Mixdq: Memory-efficient few-step text-to-image diffusion models with metric-decoupled mixed precision quantization. arXiv preprint arXiv:2405.17873, 2024.
  74. 74.Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. Atom: Low-bit quantization for efficient and accurate llm serving. arXiv preprint arXiv:2310.19102, 2023.
  75. 75.Zhenghao Zhao, Yuzhang Shang, Junyi Wu, and Yan Yan. Dataset quantization with active learning based adaptive sampling. arXiv preprint arXiv:2407.07268, 2024.
  76. 76.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  77. 77.Yinan Zhou, Yuxin Chen, Haokun Lin, Shuyu Yang, Li Zhu, Zhongang Qi, Chen Ma, and Ying Shan. Doge: Towards versatile visual document grounding and referring. arXiv preprint arXiv:2411.17125, 2024.

Citation

MLA
Lin, H., et al. “DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 87766–800, https://proceedings.neurips.cc/paper_files/paper/2024/file/9febda1c8344cc5f2d51713964864e93-Paper-Conference.pdf.
APA
Lin, H., Xu, H., Wu, Y., Cui, J., Zhang, Y., Mou, L., Song, L., Sun, Z., & Wei, Y. (2024). DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs. Advances in Neural Information Processing Systems, 37, 87766–87800. https://proceedings.neurips.cc/paper_files/paper/2024/file/9febda1c8344cc5f2d51713964864e93-Paper-Conference.pdf
Chicago
Lin, H., H. Xu, Y. Wu, et al. 2024. “DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs”. Advances in Neural Information Processing Systems 37: 87766–800. https://proceedings.neurips.cc/paper_files/paper/2024/file/9febda1c8344cc5f2d51713964864e93-Paper-Conference.pdf.
Harvard
Lin, H. et al. (2024) “DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 87766–87800. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/9febda1c8344cc5f2d51713964864e93-Paper-Conference.pdf.
Vancouver
1. Lin H, Xu H, Wu Y, Cui J, Zhang Y, Mou L, Song L, Sun Z, Wei Y (2024) DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 87766–87800

BibTeX

@inproceedings{lin2024duquant,
  title = {DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs},
  author = {Lin, Haokun and Xu, Haobo and Wu, Yichen and Cui, Jingzhi and Zhang, Yingtao and Mou, Linzhan and Song, Linqi and Sun, Zhenan and Wei, Ying},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {87766-87800},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/9febda1c8344cc5f2d51713964864e93-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors