DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory

Jerry CheeArturs BackursRainie HeckLi ZhangJanardhan (Jana) KulkarniThomas RothvossSivakanth Gopi

article2025Annual Conference Computational Learning Theory5 citations

Introduces DiscQuant, a weight-rounding algorithm grounded in discrepancy theory that guarantees bounded quantization error and significantly outperforms standard post-training methods like GPTQ on low-bit large language model compression.

Listen

Modern large language models require substantial memory and computational resources, creating high operational costs and serving bottlenecks during deployment. Post-training quantization reduces these costs by compressing model weights into lower-bit formats, but standard rounding methods often degrade model accuracy. Most prior research focused on designing low-bit grids rather than optimizing how continuous weights are rounded to discrete values. The article introduces DiscQuant, a data-dependent rounding method grounded in mathematical discrepancy theory, designed to compress entire models simultaneously with minimal loss in task performance.

To develop this method, the authors prove theoretically that model loss changes can be effectively approximated by first-order gradient terms and that sample gradients exhibit a rapidly decaying, low-rank structure. Drawing inspiration from discrepancy theory algorithms, the authors formulate a practical optimization objective that minimizes the discrepancy between original and compressed model outputs via knowledge distillation combined with a linear regularization term. The approach was evaluated by quantizing two popular open-source models, Phi-3-mini-4k-instruct and Meta-Llama-3.1-8B-Instruct, across standard benchmarks spanning mathematical reasoning, perplexity, and commonsense question answering across 3-bit to 4.5-bit compression levels.

DiscQuant substantially outperforms current standard rounding baselines, particularly at aggressive low-bit settings. On a challenging math benchmark, rounding Phi-3-mini to 3.25 bits per parameter with DiscQuant achieved 64.2% accuracy, compared to 54.3% for the leading baseline GPTQ and 31.0% for simple round-to-nearest. Across multiple commonsense reasoning benchmarks, DiscQuant recovered full baseline accuracy using at least 0.25 fewer bits per parameter than competing approaches. Furthermore, the experiments demonstrate that DiscQuant integrates seamlessly with orthogonal compression techniques, such as incoherence processing, to provide additional performance improvements at 3-bit precision.

These findings indicate that intelligent, global weight rounding can drastically lower deployment memory overhead without the steep accuracy penalties historically seen in ultra-low-bit quantization. Organizations can serve larger, more capable models on smaller hardware footprints, directly lowering serving infrastructure expenses while preserving generation quality. Because DiscQuant works across arbitrary pre-existing quantization grids, engineering teams can adopt the rounding method without rewriting existing hardware-optimized inference kernels.

Engineering teams preparing to deploy quantized models should adopt DiscQuant for post-training weight rounding, especially when targeting sub-4-bit compression. When implementing the algorithm, teams must carefully curate calibration datasets, as empirical tests show task performance is sensitive to the calibration data distribution. Future work should focus on developing principled guidelines for calibration data selection and extending DiscQuant to vector quantization formats. Confidence in the reported results is high across evaluated models and benchmarks, though practitioners should account for memory overhead during the optimization phase, which mirrors standard two-model knowledge distillation.

  • Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). KronQ extends post-training quantization for large language models by incorporating Kronecker-factored gradient covariance information to advance beyond standard second-order rounding solvers.
Cover for DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory

Abstract

Quantizing the weights of a neural network has two steps: (1) Finding a good low bit-complexity representation for weights (which we call the quantization grid) and (2) Rounding the original weights to values in the quantization grid. In this paper, we study the problem of rounding optimally given any quantization grid. The simplest and most commonly used way to round is Round-to-Nearest (RTN). By rounding in a data-dependent way instead, one can improve the quality of the quantized model significantly.

We study the rounding problem from the lens of \emph{discrepancy theory}, which studies how well we can round a continuous solution to a discrete solution without affecting solution quality too much. We prove that given m=poly(1/ϵ)m=\mathrm{poly}(1/\epsilon) samples from the data distribution, we can round all but O(m)O(m) model weights such that the expected approximation error of the quantized model on the true data distribution is ≤ϵ\le \epsilon as long as the space of gradients of the original model is approximately low rank (which we empirically validate).

Our proof, which is algorithmic, inspired a simple and practical rounding algorithm called \emph{DiscQuant}. In our experiments, we demonstrate that DiscQuant significantly improves over the prior state-of-the-art rounding method called GPTQ and the baseline RTN over a range of benchmarks on Phi3mini-3.8B and Llama3.1-8B. For example, rounding Phi3mini-3.8B to a fixed quantization grid with 3.25 bits per parameter using DiscQuant gets 64% accuracy on the GSM8k dataset, whereas GPTQ achieves 54% and RTN achieves 31% (the original model achieves 84%). We make our code available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Quantization Grids
  • 2.2 Rounding
  • 2.3 Discrepancy Theory
  • 3 Connections to Discrepancy Theory
  • 3.1 Bounding empirical error (Question )
  • 3.2 Bounding Generalization Error (Question )
  • 4 DiscQuant: Algorithm
  • 5 Experiments
  • 5.1 Block Scaling
  • 5.2 Incoherence Processing
  • 5.3 Effect of Data
  • References
  • A Additional Experiments
  • A.1 Experimental Setup Details
  • A.2 Incoherence Processing
  • A.3 Effect of Data
  • A.4 Ablations
  • B Rounding weights via Discrepancy Theory
  • B.1 The Lovett Meka algorithm
  • B.2 The main theoretical result
  • B.3 Analyzing the covariance estimator
  • C Non-uniform Quantization Grid
  • D Taylor Series for KL Divergence
  • E LoRA experiments

Knowls

  1. Knowl 1 — DiscQuant Post-Training Weight Rounding Objective and Linear Bias Regularization

    model/method

    Let w∈Rnw \in \mathbb{R}^n denote the unquantized weights of a pretrained neural network, and let Q=Q1×⋯×QnQ = Q_1 \times \dots \times Q_n be a scalar quantization grid where Qj⊂RQ_j \subset \mathbb{R} represents the finite set of allowable discrete values for the jj-th weight. For each coordinate jj, let wjdown=max⁡{q∈Qj:q≤wj}w^{\text{down}}_j = \max \{q \in Q_j : q \le w_j\} and wjup=min⁡{q∈Qj:q≥wj}w^{\text{up}}_j = \min \{q \in Q_j : q \ge w_j\} denote the immediate surrounding grid points. The rounding parameter x∈[0,1]nx \in [0, 1]^n parametrizes the interpolated weights via:

    wx=wdown⊙(1−x)+wup⊙xw^x = w^{\text{down}} \odot (1 - x) + w^{\text{up}} \odot x

    where ⊙\odot represents component-wise multiplication. The original continuous weight vector ww corresponds to the continuous interpolation coordinate y∈[0,1]ny \in [0, 1]^n defined by wj=wjdown(1−yj)+wjupyjw_j = w^{\text{down}}_j (1 - y_j) + w^{\text{up}}_j y_j.

    To find a quantized configuration that minimizes perturbation from the unquantized network while encouraging coordinates to snap to extreme values ({0,1}n\{0, 1\}^n), DiscQuant minimizes the objective:

    min⁡x∈[0,1]nλ⟨c∗,x⟩+Ez∼DdataEi[DKL(pw(⋅∣z<i) ∥ pwx(⋅∣z<i))]\min_{x \in [0, 1]^n} \lambda \langle c^*, x \rangle + \mathbb{E}_{z \sim \mathcal{D}_{\text{data}}} \mathbb{E}_i \left[ D_{\text{KL}}\left(p_w(\cdot \mid z_{<i}) \,\parallel\, p_{w^x}(\cdot \mid z_{<i})\right) \right]

    where pw(⋅∣z<i)p_w(\cdot \mid z_{<i}) is the next-token probability distribution predicted by the full-precision model given prefix tokens z<iz_{<i}, pwx(⋅∣z<i)p_{w^x}(\cdot \mid z_{<i}) is the distribution produced by the model parameterized by wxw^x, Ddata\mathcal{D}_{\text{data}} is the calibration data distribution, and λ>0\lambda > 0 is a regularization coefficient.

    The linear regularization direction is chosen specifically as c∗=1−2yc^* = 1 - 2y. Because xj2=xjx_j^2 = x_j for binary endpoints xj∈{0,1}x_j \in \{0, 1\}, the squared Euclidean distance between xx and yy satisfies:

    ∥x−y∥22=∑j=1n(xj2−2xjyj+yj2)≈∑j=1n(xj−2xjyj+yj2)=⟨1−2y,x⟩+∥y∥22=⟨c∗,x⟩+∥y∥22\|x - y\|_2^2 = \sum_{j=1}^n (x_j^2 - 2x_j y_j + y_j^2) \approx \sum_{j=1}^n (x_j - 2x_j y_j + y_j^2) = \langle 1 - 2y, x \rangle + \|y\|_2^2 = \langle c^*, x \rangle + \|y\|_2^2

    Minimizing ⟨c∗,x⟩\langle c^*, x \rangle over the hypercube [0,1]n[0, 1]^n biases the optimization toward the vertex of the feasible polytope closest to the original unquantized weights ww.

  2. Knowl 2 — Generalization Error Bound for Discrepancy-Based Rounding with Low-Rank Gradient Covariance

    theoretical result

    Let D\mathcal{D} be a β\beta-reasonable distribution over Rn\mathbb{R}^n with unknown covariance matrix Σ=Eg∼D[ggT]\Sigma = \mathbb{E}_{g \sim \mathcal{D}}[g g^T], meaning that for all directions θ∈Rn\theta \in \mathbb{R}^n, E[⟨g,θ⟩4]≤β(E[⟨g,θ⟩2])2\mathbb{E}[\langle g, \theta \rangle^4] \le \beta (\mathbb{E}[\langle g, \theta \rangle^2])^2 for some constant β≥1\beta \ge 1. Suppose the eigenvalues λ1≥λ2≥⋯≥λn≥0\lambda_1 \ge \lambda_2 \ge \dots \ge \lambda_n \ge 0 of Σ\Sigma satisfy a polynomial decay condition λk≤λ1/kα\lambda_k \le \lambda_1 / k^\alpha for a constant decay rate α>1\alpha > 1.

    Let y∈[0,1]ny \in [0, 1]^n represent the continuous fractional weights of a model, and let g1,…,gm∼Dg_1, \dots, g_m \sim \mathcal{D} be mm independent gradient samples drawn from D\mathcal{D}, with sample size 1≤m≤n/161 \le m \le n/16.

    Then there exists a randomized polynomial-time rounding algorithm that produces a vector x∈[0,1]nx \in [0, 1]^n such that with probability at least 0.990.99:

    1. At most 16m16m coordinates of xx are fractional, i.e., ∣{j∈[n]:xj∈(0,1)}∣≤16m|\{j \in [n] : x_j \in (0, 1)\}| \le 16m.
    2. The expected squared approximation error on unseen samples drawn from D\mathcal{D} satisfies:

    Eg∼D[⟨g,x−y⟩2]=(x−y)TΣ(x−y)≲α,βlog⁡(nm)⋅Fα(m,n)\mathbb{E}_{g \sim \mathcal{D}}\left[\langle g, x - y \rangle^2\right] = (x - y)^T \Sigma (x - y) \lesssim_{\alpha, \beta} \log\left(\frac{n}{m}\right) \cdot F_\alpha(m, n)

    where

    Fα(m,n)={m1−αif 1<α<32log⁡nmif α=321mif α>32F_\alpha(m, n) = \begin{cases} m^{1-\alpha} & \text{if } 1 < \alpha < \frac{3}{2} \\ \frac{\log n}{\sqrt{m}} & \text{if } \alpha = \frac{3}{2} \\ \frac{1}{\sqrt{m}} & \text{if } \alpha > \frac{3}{2} \end{cases}

    Consequently, all but O(m)O(m) parameters are fully rounded into {0,1}\{0, 1\}, while the expected generalization error on the true data distribution is bounded by O~(m−min⁡{1/2,α−1})\widetilde{O}(m^{-\min\{1/2, \alpha - 1\}}), demonstrating that sample complexity m=poly(log⁡(n)/ε)m = \text{poly}(\log(n) / \varepsilon) suffices to achieve expected generalization error at most ε\varepsilon.

  3. Knowl 3 — DiscQuant Algorithm for Neural Network Post-Training Quantization

    algorithm

    DiscQuant rounds all weights in a pretrained neural network simultaneously using projected stochastic gradient descent on a knowledge distillation loss regularized by a linear projection objective, followed by round-to-nearest on remaining fractional values.

    Input: Pretrained weight vector w∈Rnw \in \mathbb{R}^n, scalar quantization grid QQ, calibration dataset Ddata\mathcal{D}_{\text{data}}, regularization weight λ>0\lambda > 0, learning rate η\eta, step count TT, gradient clamp threshold γ\gamma
    Output: Fully quantized weight vector w^∈Q\hat{w} \in Q
    1. For each coordinate j∈{1,…,n}j \in \{1, \dots, n\}, determine bounding grid points wjdown=max⁡{q∈Qj:q≤wj}w^{\text{down}}_j = \max \{q \in Q_j : q \le w_j\} and wjup=min⁡{q∈Qj:q≥wj}w^{\text{up}}_j = \min \{q \in Q_j : q \ge w_j\}.
    2. Compute continuous interpolation target y∈[0,1]ny \in [0, 1]^n such that wj=wjdown(1−yj)+wjupyjw_j = w^{\text{down}}_j(1 - y_j) + w^{\text{up}}_j y_j.
    3. Set linear regularization direction c∗=1−2yc^* = 1 - 2y.
    4. Initialize interpolation vector x∈[0,1]nx \in [0, 1]^n uniformly at random: xj∼Uniform(0,1)x_j \sim \text{Uniform}(0, 1).
    5. for t=1t = 1 to TT do:
        a. Sample a minibatch of sequences z∼Ddataz \sim \mathcal{D}_{\text{data}}.
        b. Instantiate continuous student weights wx=wdown⊙(1−x)+wup⊙xw^x = w^{\text{down}} \odot (1 - x) + w^{\text{up}} \odot x.
        c. Compute teacher token distributions pw(⋅∣z<i)p_w(\cdot \mid z_{<i}) and student token distributions pwx(⋅∣z<i)p_{w^x}(\cdot \mid z_{<i}).
        d. Evaluate KL distillation loss Ldistill(x)=Ei[DKL(pw(⋅∣z<i)∥pwx(⋅∣z<i))]\mathcal{L}_{\text{distill}}(x) = \mathbb{E}_i [D_{\text{KL}}(p_w(\cdot \mid z_{<i}) \parallel p_{w^x}(\cdot \mid z_{<i}))].
        e. Compute stochastic gradient gdistill=∇xLdistill(x)g_{\text{distill}} = \nabla_x \mathcal{L}_{\text{distill}}(x) and apply entry-wise clipping: gdistill,j←clip(gdistill,j,−γ,γ)g_{\text{distill}, j} \leftarrow \text{clip}(g_{\text{distill}, j}, -\gamma, \gamma).
        f. Compute total step gradient G=λc∗+gdistillG = \lambda c^* + g_{\text{distill}}.
        g. Update xx using AdamW with cosine learning rate decay: x←AdamW(x,G,η)x \leftarrow \text{AdamW}(x, G, \eta).
        h. Project back to unit hypercube: xj←min⁡(1,max⁡(0,xj))x_j \leftarrow \min(1, \max(0, x_j)) for all j∈{1,…,n}j \in \{1, \dots, n\}.
    6. Apply Round-to-Nearest (RTN) to final coordinates: for each coordinate jj, set x^j=1\hat{x}_j = 1 if xj≥0.5x_j \ge 0.5 else 00.
    7. Compute final quantized weights w^=wdown⊙(1−x^)+wup⊙x^\hat{w} = w^{\text{down}} \odot (1 - \hat{x}) + w^{\text{up}} \odot \hat{x}.
    return w^\hat{w}

    Typical hyperparameters for block scaling quantization: total iterations T=1024T = 1024, warmup =128= 128 steps, regularization λ=200\lambda = 200, batch size ∈{4,8}\in \{4, 8\}, learning rate η∈{0.05,0.1}\eta \in \{0.05, 0.1\}, and gradient clamp threshold γ∈{0.5,1.0}\gamma \in \{0.5, 1.0\}.

  4. Knowl 4 — Second-Order Hessian Structure of Token Distillation Loss

    theoretical result

    Let pw(⋅∣z<i)p_w(\cdot \mid z_{<i}) denote the next-token probability distribution produced by a model with weights w∈Rnw \in \mathbb{R}^n on prefix z<iz_{<i}, where sequence z∼Ddataz \sim \mathcal{D}_{\text{data}}. Let the distillation error function be:

    error(w^)=Ez∼DdataEi[DKL(pw(⋅∣z<i) ∥ pw^(⋅∣z<i))]\text{error}(\hat{w}) = \mathbb{E}_{z \sim \mathcal{D}_{\text{data}}} \mathbb{E}_i \left[ D_{\text{KL}}\left(p_w(\cdot \mid z_{<i}) \,\parallel\, p_{\hat{w}}(\cdot \mid z_{<i})\right) \right]

    Consider the Taylor series expansion of error(w^)\text{error}(\hat{w}) around w^=w\hat{w} = w:

    error(w^)=⟨gw,w^−w⟩+(w^−w)THw(w^−w)+O(∥w^−w∥23)\text{error}(\hat{w}) = \langle g_w, \hat{w} - w \rangle + (\hat{w} - w)^T H_w (\hat{w} - w) + \mathcal{O}(\|\hat{w} - w\|_2^3)

    where gw=∇w^error(w^)∣w^=wg_w = \nabla_{\hat{w}} \text{error}(\hat{w})\big|_{\hat{w}=w} is the gradient and Hw=∇w^2error(w^)∣w^=wH_w = \nabla^2_{\hat{w}} \text{error}(\hat{w})\big|_{\hat{w}=w} is the Hessian matrix.

    1. The first-order gradient identically vanishes:

    gw=0g_w = 0

    1. The Hessian is positive semidefinite and equals the expected outer product of the log-likelihood score functions:

    Hw=Ez∼DdataEiEt∼pw(⋅∣z<i)[(∇wlog⁡pw(t∣z<i))(∇wlog⁡pw(t∣z<i))T]H_w = \mathbb{E}_{z \sim \mathcal{D}_{\text{data}}} \mathbb{E}_i \mathbb{E}_{t \sim p_w(\cdot \mid z_{<i})} \left[ \left(\nabla_w \log p_w(t \mid z_{<i})\right) \left(\nabla_w \log p_w(t \mid z_{<i})\right)^T \right]

    Therefore, the distillation loss around the unquantized weights is dominated by the second-order term:

    error(w^)≈Ez∼DdataEiEt∼pw(⋅∣z<i)[⟨∇wlog⁡pw(t∣z<i),w^−w⟩2]\text{error}(\hat{w}) \approx \mathbb{E}_{z \sim \mathcal{D}_{\text{data}}} \mathbb{E}_i \mathbb{E}_{t \sim p_w(\cdot \mid z_{<i})} \left[ \langle \nabla_w \log p_w(t \mid z_{<i}), \hat{w} - w \rangle^2 \right]

  5. Knowl 5 — Polytope Geometry of Weight Rounding under Linear Sample Constraints

    definition

    Let w∈Rnw \in \mathbb{R}^n be the unquantized weight vector of a model and Q=Q1×⋯×QnQ = Q_1 \times \dots \times Q_n be a quantization grid. Under the constraint that each weight wjw_j can only be rounded up to wjup∈Qjw^{\text{up}}_j \in Q_j or down to wjdown∈Qjw^{\text{down}}_j \in Q_j, the space of valid roundings lies at the extreme vertices of the hypercube [0,1]n[0, 1]^n via the parametrization wx=wdown⊙(1−x)+wup⊙xw^x = w^{\text{down}} \odot (1 - x) + w^{\text{up}} \odot x.

    Let y∈[0,1]ny \in [0, 1]^n be the continuous target such that wy=ww^y = w. For a set of mm calibration samples s1,…,sms_1, \dots, s_m, preserving the first-order loss change Δf(w;si)≈⟨∇wf(w;si),wx−w⟩=0\Delta f(w; s_i) \approx \langle \nabla_w f(w; s_i), w^x - w \rangle = 0 corresponds to mm linear constraints:

    M(x−y)=0M(x - y) = 0

    where M∈Rm×nM \in \mathbb{R}^{m \times n} has rows Mi,:=∇wf(w;si)⊙(wup−wdown)M_{i, :} = \nabla_w f(w; s_i) \odot (w^{\text{up}} - w^{\text{down}}). The subspace V={x∈Rn:Mx=My}V = \{x \in \mathbb{R}^n : Mx = My\} is an affine subspace of dimension at least n−mn - m.

    The feasible set K=[0,1]n∩VK = [0, 1]^n \cap V is a non-empty convex polytope. Because any basic feasible solution (vertex) of KK requires nn linearly independent tight constraints, and VV accounts for at most mm constraints, every vertex of KK has at least n−mn - m tight box constraints xj∈{0,1}x_j \in \{0, 1\}. When n≫mn \gg m, any vertex of KK represents an almost completely integral rounded weight vector.

  6. Knowl 6 — Schatten-1 Norm Covariance Estimation Error under Heavy-Tailed Spectrum

    theoretical result

    Let D\mathcal{D} be a β\beta-reasonable distribution over Rn\mathbb{R}^n whose covariance matrix Σ∈Rn×n\Sigma \in \mathbb{R}^{n \times n} has eigenvalues satisfying λk≤1/kα\lambda_k \le 1/k^\alpha for all k∈{1,…,n}k \in \{1, \dots, n\}, with constant α>1\alpha > 1. Let g1,…,gm∼Dg_1, \dots, g_m \sim \mathcal{D} be mm independent samples, and let X=1m∑ℓ=1mgℓgℓTX = \frac{1}{m} \sum_{\ell=1}^m g_\ell g_\ell^T be the empirical sample covariance estimator.

    The expected error of the empirical covariance in the Schatten-1 norm (trace norm / nuclear norm) ∥M∥S(1)=∑iσi(M)\|M\|_{S(1)} = \sum_i \sigma_i(M) is bounded by:

    E[∥X−Σ∥S(1)]≲α,βFα(m,n)\mathbb{E}\left[\|X - \Sigma\|_{S(1)}\right] \lesssim_{\alpha, \beta} F_\alpha(m, n)

    where

    Fα(m,n)={m1−αif 1<α<32log⁡nmif α=321mif α>32F_\alpha(m, n) = \begin{cases} m^{1-\alpha} & \text{if } 1 < \alpha < \frac{3}{2} \\ \frac{\log n}{\sqrt{m}} & \text{if } \alpha = \frac{3}{2} \\ \frac{1}{\sqrt{m}} & \text{if } \alpha > \frac{3}{2} \end{cases}

  7. Knowl 7 — Performance of DiscQuant on Phi-3-mini-4k-instruct under Block Scaling

    data/table

    The table below compares the post-training quantization performance of Round-to-Nearest (RTN), GPTQ, and DiscQuant across bitrates from 3.0 to 4.5 bits per parameter on Phi-3-mini-4k-instruct using symmetric linear block scaling with a 16-bit scale per block (e.g., 3.25 bits uses 3-bit quantization with groupsize 64). Evaluations cover Wikitext perplexity (Wiki), GSM8k chain-of-thought 8-shot accuracy, MMLU 5-shot accuracy, and zero-shot accuracies on ARC-Challenge (ArcC), PIQA, HellaSwag (Hella), and WinoGrande (Wino).

    Method Wbits Wiki ↓\downarrow GSM8k ↑\uparrow MMLU ↑\uparrow ArcC ↑\uparrow PIQA ↑\uparrow Hella ↑\uparrow Wino ↑\uparrow
    Baseline 16.0 9.5 84.4 ±\pm 1.0 70.4 ±\pm 0.4 56.7 ±\pm 1.4 80.8 ±\pm 0.9 77.4 ±\pm 0.4 73.5 ±\pm 1.2
    RTN 3.0 6.3E5 1.0 ±\pm 0.3 23.3 ±\pm 0.4 26.9 ±\pm 1.3 53.4 ±\pm 1.2 28.2 ±\pm 0.4 48.6 ±\pm 1.4
    GPTQ 3.0 28.2 2.3 ±\pm 0.4 37.7 ±\pm 0.4 34.8 ±\pm 1.4 64.3 ±\pm 1.1 56.5 ±\pm 0.5 52.6 ±\pm 1.4
    DiscQ 3.0 17.7 26.8 ±\pm 1.2 45.6 ±\pm 0.4 44.1 ±\pm 1.5 73.9 ±\pm 1.0 63.3 ±\pm 0.5 66.6 ±\pm 1.3
    RTN 3.25 22.5 31.0 ±\pm 1.3 53.2 ±\pm 0.4 48.4 ±\pm 1.5 72.5 ±\pm 1.0 68.3 ±\pm 0.5 62.6 ±\pm 1.4
    GPTQ 3.25 13.8 54.3 ±\pm 1.4 59.0 ±\pm 0.4 49.6 ±\pm 1.5 77.3 ±\pm 1.0 71.1 ±\pm 0.5 66.5 ±\pm 1.3
    DiscQ 3.25 12.6 64.2 ±\pm 1.3 60.7 ±\pm 0.4 53.5 ±\pm 1.5 78.7 ±\pm 1.0 72.3 ±\pm 0.4 72.5 ±\pm 1.3
    RTN 3.5 18.8 46.3 ±\pm 1.4 57.0 ±\pm 0.4 46.2 ±\pm 1.5 73.8 ±\pm 1.0 70.0 ±\pm 0.5 63.9 ±\pm 1.4
    GPTQ 3.5 12.8 54.6 ±\pm 1.4 61.7 ±\pm 0.4 51.6 ±\pm 1.5 78.9 ±\pm 1.0 72.3 ±\pm 0.4 68.3 ±\pm 1.3
    DiscQ 3.5 12.0 69.5 ±\pm 1.3 63.0 ±\pm 0.4 51.1 ±\pm 1.5 78.9 ±\pm 1.0 73.0 ±\pm 0.4 73.9 ±\pm 1.2
    RTN 4.0 14.6 62.2 ±\pm 1.3 61.2 ±\pm 0.4 53.6 ±\pm 1.5 76.3 ±\pm 1.0 72.9 ±\pm 0.4 65.3 ±\pm 1.3
    GPTQ 4.0 11.5 71.5 ±\pm 1.2 65.1 ±\pm 0.4 54.6 ±\pm 1.5 78.8 ±\pm 1.0 74.7 ±\pm 0.4 70.9 ±\pm 1.3
    DiscQ 4.0 11.2 77.3 ±\pm 1.2 65.7 ±\pm 0.4 56.8 ±\pm 1.4 79.5 ±\pm 0.9 74.5 ±\pm 0.4 72.0 ±\pm 1.3
    RTN 4.25 11.2 64.4 ±\pm 1.3 67.5 ±\pm 0.4 55.5 ±\pm 1.5 79.3 ±\pm 0.9 76.1 ±\pm 0.4 69.1 ±\pm 1.3
    GPTQ 4.25 10.3 81.0 ±\pm 1.1 68.5 ±\pm 0.4 56.9 ±\pm 1.4 79.7 ±\pm 0.9 76.1 ±\pm 0.4 72.1 ±\pm 1.3
    DiscQ 4.25 10.2 80.7 ±\pm 1.1 68.4 ±\pm 0.4 57.3 ±\pm 1.4 80.7 ±\pm 0.9 76.3 ±\pm 0.4 74.2 ±\pm 1.2
    RTN 4.5 10.8 71.6 ±\pm 1.2 67.7 ±\pm 0.4 57.5 ±\pm 1.4 79.3 ±\pm 0.9 76.6 ±\pm 0.4 72.2 ±\pm 1.3
    GPTQ 4.5 10.1 82.0 ±\pm 1.1 68.8 ±\pm 0.4 55.8 ±\pm 1.5 80.8 ±\pm 0.9 76.5 ±\pm 0.4 71.8 ±\pm 1.3
    DiscQ 4.5 10.0 82.1 ±\pm 1.1 68.5 ±\pm 0.4 56.6 ±\pm 1.4 80.2 ±\pm 0.9 76.7 ±\pm 0.4 74.2 ±\pm 1.2

    DiscQuant outperforms RTN and GPTQ across tasks, especially in lower bit-depth regimes (≤3.5\le 3.5 bits) and generative tasks. On GSM8k at 3.25 bits, DiscQuant scores 64.2% versus 54.3% for GPTQ and 31.0% for RTN. On ARC-Challenge, PIQA, and WinoGrande, DiscQuant achieves full recovery of unquantized performance using at least 0.25 fewer bits per parameter than GPTQ.

  8. Knowl 8 — Performance of DiscQuant on Meta-Llama-3.1-8B-Instruct under Block Scaling

    data/table

    The table below details quantization evaluations on Meta-Llama-3.1-8B-Instruct comparing RTN, GPTQ, and DiscQuant using block scaling across bit rates from 3.0 to 4.5 bits. Benchmarks include Wikitext perplexity (Wiki), GSM8k cot 8-shot accuracy, MMLU 5-shot accuracy, ARC-Challenge 0-shot (ArcC), PIQA 0-shot, HellaSwag 0-shot (Hella), and WinoGrande 0-shot (Wino).

    Method Wbits Wiki ↓\downarrow GSM8k ↑\uparrow MMLU ↑\uparrow ArcC ↑\uparrow PIQA ↑\uparrow Hella ↑\uparrow Wino ↑\uparrow
    Baseline 16.0 8.7 77.0 ±\pm 1.2 68.0 ±\pm 0.4 55.2 ±\pm 1.5 81.3 ±\pm 0.9 79.3 ±\pm 0.4 73.7 ±\pm 1.2
    RTN 3.0 4.4E3 0.5 ±\pm 0.2 23.2 ±\pm 0.4 22.3 ±\pm 1.2 52.4 ±\pm 1.2 29.1 ±\pm 0.5 50.0 ±\pm 1.4
    GPTQ 3.0 23.2 3.6 ±\pm 0.5 24.6 ±\pm 0.4 31.8 ±\pm 1.4 66.6 ±\pm 1.1 45.8 ±\pm 0.5 54.1 ±\pm 1.4
    DiscQ 3.0 15.2 14.3 ±\pm 1.0 44.6 ±\pm 0.4 39.4 ±\pm 1.4 73.2 ±\pm 1.0 64.4 ±\pm 0.5 62.8 ±\pm 1.4
    RTN 3.25 15.2 10.8 ±\pm 0.9 50.5 ±\pm 0.4 44.3 ±\pm 1.5 75.2 ±\pm 1.0 71.4 ±\pm 0.5 67.2 ±\pm 1.3
    GPTQ 3.25 10.7 56.3 ±\pm 1.4 60.5 ±\pm 0.4 46.3 ±\pm 1.5 76.7 ±\pm 1.0 74.4 ±\pm 0.4 68.7 ±\pm 1.3
    DiscQ 3.25 10.5 58.3 ±\pm 1.4 60.2 ±\pm 0.4 49.1 ±\pm 1.5 79.1 ±\pm 0.9 75.1 ±\pm 0.4 72.1 ±\pm 1.3
    RTN 3.5 12.7 35.9 ±\pm 1.3 51.4 ±\pm 0.4 48.4 ±\pm 1.5 76.7 ±\pm 1.0 73.0 ±\pm 0.4 69.1 ±\pm 1.3
    GPTQ 3.5 10.4 57.0 ±\pm 1.4 62.1 ±\pm 0.4 49.9 ±\pm 1.5 77.3 ±\pm 1.0 75.1 ±\pm 0.4 71.1 ±\pm 1.3
    DiscQ 3.5 10.3 60.7 ±\pm 1.3 60.9 ±\pm 0.4 51.7 ±\pm 1.5 79.2 ±\pm 0.9 76.3 ±\pm 0.4 72.5 ±\pm 1.3
    RTN 4.0 12.5 50.8 ±\pm 1.4 59.3 ±\pm 0.4 50.5 ±\pm 1.5 77.6 ±\pm 1.0 74.7 ±\pm 0.4 69.9 ±\pm 1.3
    GPTQ 4.0 9.9 63.2 ±\pm 1.3 64.4 ±\pm 0.4 52.4 ±\pm 1.5 78.4 ±\pm 1.0 75.9 ±\pm 0.4 71.7 ±\pm 1.3
    DiscQ 4.0 9.8 66.5 ±\pm 1.3 63.4 ±\pm 0.4 51.6 ±\pm 1.5 79.2 ±\pm 0.9 76.9 ±\pm 0.4 72.8 ±\pm 1.3
    RTN 4.25 9.4 70.6 ±\pm 1.3 65.7 ±\pm 0.4 54.2 ±\pm 1.5 80.1 ±\pm 0.9 78.0 ±\pm 0.4 73.9 ±\pm 1.2
    GPTQ 4.25 9.1 74.6 ±\pm 1.2 66.8 ±\pm 0.4 53.4 ±\pm 1.5 79.6 ±\pm 0.9 77.9 ±\pm 0.4 73.5 ±\pm 1.2
    DiscQ 4.25 9.1 74.9 ±\pm 1.2 66.9 ±\pm 0.4 53.6 ±\pm 1.5 79.9 ±\pm 0.9 78.4 ±\pm 0.4 72.6 ±\pm 1.3
    RTN 4.5 9.3 71.9 ±\pm 1.2 65.8 ±\pm 0.4 54.8 ±\pm 1.5 80.3 ±\pm 0.9 78.4 ±\pm 0.4 72.4 ±\pm 1.3
    GPTQ 4.5 9.0 73.8 ±\pm 1.2 66.9 ±\pm 0.4 53.6 ±\pm 1.5 79.6 ±\pm 0.9 78.1 ±\pm 0.4 73.7 ±\pm 1.2
    DiscQ 4.5 9.1 74.8 ±\pm 1.2 66.8 ±\pm 0.4 54.1 ±\pm 1.5 80.6 ±\pm 0.9 78.7 ±\pm 0.4 72.9 ±\pm 1.2

    DiscQuant yields superior compression across most bit levels. For instance, at 4.0 bits, DiscQuant obtains 66.5% on GSM8k compared to 63.2% for GPTQ and 50.8% for RTN. At 3.0 bits, DiscQuant preserves functional capability (GSM8k 14.3%, PIQA 73.2%) where RTN completely degrades (GSM8k 0.5%, PIQA 52.4%).

  9. Knowl 9 — Composing DiscQuant with Incoherence Processing

    empirical result

    DiscQuant is agnostic to the quantization grid format and can be directly combined with incoherence transformations, such as the Randomized Hadamard Transform (RHT). Incoherence processing multiplies weights by random orthogonal matrices prior to quantization to suppress outlier coordinates and flatten weight distributions.

    When evaluated on Phi-3-mini-4k-instruct and Meta-Llama-3.1-8B-Instruct at 3-bit and 4-bit uniform grids without group scaling:

    1. On Phi-3-mini-4k-instruct at 3.0 bits with incoherence processing, DiscQuant achieves 29.9% GSM8k, 48.0% MMLU, and 46.2% ARC-Challenge, outperforming GPTQ with incoherence (10.0% GSM8k, 43.8% MMLU, 39.2% ARC-Challenge) and RTN with incoherence (0.0% GSM8k, 23.4% MMLU, 28.5% ARC-Challenge).
    2. On Meta-Llama-3.1-8B-Instruct at 3.0 bits with incoherence processing, DiscQuant achieves 25.4% GSM8k and 51.5% MMLU, compared to GPTQ with incoherence (24.4% GSM8k, 49.7% MMLU).
    3. At 3.0 bits on Phi-3-mini, DiscQuant using standard block scaling alone (without incoherence processing) achieves 26.8% on GSM8k, remaining competitive with GPTQ even after GPTQ is augmented with incoherence processing (10.0%).
  10. Knowl 10 — Comparison of Knowledge Distillation Objectives for Weight Rounding

    data/table

    The table below compares different distillation objective formulations when rounding Phi-3-mini-4k-instruct to 3.25 bits using 1024 calibration samples from RedPajama. Tested objectives include standard next-token KL divergence, normalized L2L_2 intermediate feature distance per decoder layer (Layer), and normalized L2L_2 distance per linear projection (Linear), as well as affine linear combinations of KL and intermediate losses.

    KL Coeff Intermed Coeff Intermed Type Wiki ↓\downarrow GSM8k ↑\uparrow
    1.0 0.0 None 12.8 64.9 ±\pm 1.3
    0.0 1.0 Layer 14.7 54.1 ±\pm 1.4
    0.0 1.0 Linear 14.3 60.1 ±\pm 1.4
    0.1 0.9 Linear 13.1 61.4 ±\pm 1.3
    0.5 0.5 Linear 12.9 63.9 ±\pm 1.3
    0.9 0.1 Linear 12.8 63.8 ±\pm 1.3

    Direct optimization of the full-sequence next-token KL divergence alone yields the lowest perplexity (12.8 on Wikitext) and the highest generative math accuracy (64.9% on GSM8k). Incorporating intermediate activation losses degrades performance relative to pure KL divergence. Furthermore, optimizing ground-truth cross-entropy loss directly instead of distillation yields substantially worse results (52.7% on GSM8k and 13.6 on Wikitext).

  11. Knowl 11 — Non-Zero Per-Sample Gradients and Covariance Decay in Pretrained Language Models

    empirical result

    Prior weight quantization works often assume first-order loss perturbations Δf≈⟨∇wf(w;s),w^−w⟩\Delta f \approx \langle \nabla_w f(w; s), \hat{w} - w \rangle are negligible based on the assumption that pretrained weights sit near a local minimum where gradients vanish. Empirical evaluation reveals that while the expected gradient norm squared ∥E[g]∥22\|\mathbb{E}[g]\|_2^2 is near zero, per-sample gradients g=∇wf(w;s)g = \nabla_w f(w; s) have substantial variance and magnitude:

    1. On Phi-3-mini-128k (evaluated over 8192 samples of sequence length 2048 from RedPajama-1T), ∥E[g]∥22=0.1021\|\mathbb{E}[g]\|_2^2 = 0.1021, whereas the expected squared gradient norm E[∥g∥22]=4.7812\mathbb{E}[\|g\|_2^2] = 4.7812.
    2. On Llama-3.1-8B, ∥E[g]∥22=1.6328\|\mathbb{E}[g]\|_2^2 = 1.6328, whereas E[∥g∥22]=107.0\mathbb{E}[\|g\|_2^2] = 107.0.

    Moreover, the eigenvalues λk\lambda_k of the empirical gradient covariance matrix Σ=E[ggT]\Sigma = \mathbb{E}[g g^T] across transformer layers decay polynomially fast ({λk≤λ1/kα}\{\lambda_k \le \lambda_1 / k^\alpha\} with α≥1\alpha \ge 1). Because total gradient variance ∑i=1nλi=Tr(Σ)=E[∥g∥22]=O(1)\sum_{i=1}^n \lambda_i = \text{Tr}(\Sigma) = \mathbb{E}[\|g\|_2^2] = \mathcal{O}(1), the gradient space is approximately low rank, enabling generalization from a modest number of calibration samples.

Coverage note — Post-quantization LoRA adapter fine-tuning (Appendix E, Table 7) and data mixture variation experiments (Figure 6 / Appendix A.3) were omitted as secondary extension analyses.

References

  1. 1.Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, and Harkirat Behl. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv.org/abs/2404.14219.
  2. 2.Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. Quarot: Outlier-free 4-bit inference in rotated llms. In Thirty-either Conference on Neural Information Processing Systems, 2024.
  3. 3.Nikhil Bansal. Discrepancy theory and related algorithms. In Proc. Int. Cong. Math, volume 7, pages 5178–5210, 2022.
  4. 4.Kayhan Behdin, Ayan Acharya, Aman Gupta, Sathiya Keerthi, Rahul Mazumder, Zhu Siyu, and Song Qingquan. Quantease: Optimization-based quantization for language models–an efficient and intuitive algorithm. arXiv preprint arXiv:2309.01885, 2023.
  5. 5.Bernard Chazelle, William WL Chen, and Anand Srivastav. Discrepancy theory and its applications. Oberwolfach Reports, 1(1):673–722, 2004.
  6. 6.Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa. QuIP: 2-bit quantization of large language models with guarantees. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=xrk9g5vcXR.
  7. 7.Together Computer. Redpajama: An open source recipe to reproduce llama training dataset, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  8. 8.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm.int(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems, 2022.
  9. 9.Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless llm weight compression, 2024.
  10. 10.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  11. 11.Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, and Dan Alistarh. Extreme compression of large language models via additive quantization. In Forty-First International Conference on Machine Learning, 2024.
  12. 12.Elias Frantar and Dan Alistarh. Sparsegpt: Massive language models can be accurately pruned in one-shot. In Proceedings of the International Conference on Machine Learning, 2023.
  13. 13.Elias Frantar, Sidak Pal Singh, and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. In Advances in Neural Information Processing Systems, 2022.
  14. 14.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023.
  15. 15.Elias Frantar, Roberto L Castro, Jiale Chen, Torsten Hoefler, and Dan Alistarh. Marlin: Mixed-precision auto-regressive parallel inference on large language models. arXiv preprint arXiv:2408.11743, 2024.
  16. 16.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  17. 17.Babak Hassibi, Daivd G Stork, and Gregory J Wolff. optimal brain surgeon and general network pruning. In IEEE International Conference on Neural Networks, 1993.
  18. 18.Itay Hubara, Yury Nahshan, Yair Hanami, Ron Banner, and Daniel SOudry. Accurate post training quantization with small calibration sets. In Thirty-Eighth International Conference on Machine Learning, 2021.
  19. 19.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. In Forty-First International Conference on Machine Learning, 2024.
  20. 20.Eldar Kurtic, Denis Kuznedelev, Elias Frantar, Michael Goin, and Dan Alistarh. Sparse fine-tuning for inference acceleration of large language models, 2023. URL https://arxiv.org/abs/2310.06927.
  21. 21.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611–626, 2023.
  22. 22.Yann LeCun, John Denker, and Sara Solla. Optimal brain damage. Advances in neural information processing systems, 2, 1989.
  23. 23.Yuang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu. Brecq: Pushing the limit of post-training quantization by block reconstruction. In The Nineth International Conference on Learning Representations, 2021.
  24. 24.Jin Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. Awq: Acttivation-aware weight quantization for on-device llm compression and acceleration. In Seventh Conference on Machine Learning and Systems, 2024.
  25. 25.Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant–llm quantization with learned rotations. arXiv preprint arXiv:2405.16406, 2024a.
  26. 26.Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. In Forty-First International Conference on Machine Learning, 2024b.
  27. 27.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. The International Conference on Learning Representations, 2019.
  28. 28.L'aszl'o Lov'asz, Joel Spencer, and Katalin Vesztergombi. Discrepancy of set-systems and matrices. European Journal of Combinatorics, 7(2):151–160, 1986.
  29. 29.Shachar Lovett and Raghu Meka. Constructive discrepancy minimization by walking on the edges. In FOCS, pages 61–67. IEEE Computer Society, 2012.
  30. 30.Eric Lybrand and Rayan Saab. A greedy algorithm for quantizing neural networks. Journal of Machine Learning Research, 22(156):1–38, 2021.
  31. 31.Jiri Matousek. Geometric discrepancy: An illustrated guide, volume 18. Springer Science & Business Media, 2009.
  32. 32.Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? Adaptive rounding for post-training quantization. In Hal Daum'e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 7197–7206. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v119/nagel20a.html.
  33. 33.Pranav Ajit Nair and Arun Sai Suggala. Cdquant: Accurate post-training weight quantization of large pre-trained models using greedy coordinate descent, 2024. URL https://arxiv.org/abs/2406.17542.
  34. 34.Ryan O’Donnell. Analysis of Boolean Functions. Cambridge University Press, 2014.
  35. 35.Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. Omniquant: Omnidirectionally calibrated quantization for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=8Wuvhh0LYW.
  36. 36.Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. A simple and effective pruning approach for large language models. In Workshop on Efficient Systems for Foundation Models @ ICML2023, 2023. URL https://openreview.net/forum?id=tz9JV2PRSv.
  37. 37.Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. QuIP#: Even better llm quantization with hadamard incoherence and lattice codebooks. In Forty-First International Conference on Machine Learning, 2024a.
  38. 38.Albert Tseng, Qingyao Sun, David Hou, and Christopher De Sa. QTIP: Quantization with trellises and incoherence processing. In Advances in Neural Information Processing Systems, 2024b.
  39. 39.Mart van Baalen, Andrey Kuzmin, Markus Nagel, Peter Couperus, Cedric Bastoul, Eric Mahurin, Tijmen Blankevoort, and Paul Whatmough. Gptvq: The blessing of dimensionality in llm quantization. arXiv preprint arXiv:2402.15319, 2024.
  40. 40.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In Fortieth International Conference on Machine Learning, 2023.

Citation

MLA
Chee, J., et al. “DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory”. arXiv, 2025, http://arxiv.org/abs/2501.06417v1.
APA
Chee, J., Backurs, A., Heck, R., Zhang, L., Kulkarni, J., Rothvoss, T., & Gopi, S. (2025). DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory. arXiv. http://arxiv.org/abs/2501.06417v1
Chicago
Chee, J., A. Backurs, R. Heck, et al. 2025. “DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory”. arXiv. http://arxiv.org/abs/2501.06417v1.
Harvard
Chee, J. et al. (2025) “DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.06417v1.
Vancouver
1. Chee J, Backurs A, Heck R, Zhang L, Kulkarni J, Rothvoss T, Gopi S (2025) DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory. arXiv

BibTeX

@article{chee2025discquant,
  title = {DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory},
  author = {Chee, Jerry and Backurs, Arturs and Heck, Rainie and Zhang, Li and Kulkarni, Janardhan and Rothvoss, Thomas and Gopi, Sivakanth},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.06417v1},
  eprint = {2501.06417}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/