Extreme Compression of Large Language Models via Additive Quantization

Vage EgiazarianAndrei PanferovDenis KuznedelevElias FrantarArtem BabenkoDan Alistarh

article2024ICML282 citations

Introduces AQLM, a multi-codebook quantization method that compresses large language model weights down to 2 bits per parameter while achieving state-of-the-art accuracy and matching 16-bit floating-point inference speeds on standard hardware.

Listen

Deploying modern large language models locally or on commodity hardware is challenging due to massive memory and compute requirements. While post-training quantization reduces model sizes by lowering parameter precision, traditional methods face severe accuracy degradation at extreme compression levels below 3 bits per parameter. Consequently, existing 2-bit models have historically underperformed smaller baseline models quantized to 3 or 4 bits, limiting their real-world utility.

The article demonstrates that multi-codebook vector quantization, a technique adapted from information retrieval, can achieve extreme post-training compression for large language models while preserving generation quality. The primary objective is to evaluate a novel compression algorithm, called Additive Quantization for Language Models (AQLM), across 2 to 4 bits per parameter.

The approach generalizes classical additive quantization into a three-stage optimization framework calibrated on model token activations. First, the method optimizes discrete codes for weight groups using beam search. Second, it continuously updates learned codebooks via standard optimization algorithms. Third, it fine-tunes parameters across multi-layer transformer blocks to maintain end-to-end output fidelity. The evaluation assessed open model families, including LLAMA 2 (7B, 13B, and 70B parameters) and Mixtral, against leading post-training quantization baselines using language modeling perplexity, multi-domain benchmarks, and execution speed on consumer hardware.

The findings establish that AQLM outperforms existing compression methods across 2 to 4 bits, achieving the largest accuracy gains in the extreme 2-bit regime. For the first time, Pareto optimality is demonstrated below 3 bits per parameter: starting at roughly 2.5 bits per parameter, a compressed 13B model outperforms a smaller 7B model of equivalent total byte size. When paired with end-to-end distillation fine-tuning, accuracy improves further, matching or exceeding competing approaches across zero-shot evaluations. Moreover, the homogeneous weight format enables high-performance inference, delivering up to 30% speedups on GPUs and up to fourfold speedups on CPUs compared to original precision implementations while shrinking memory footprint by up to eightfold.

These results show that organizations can deploy higher-capacity language models within strictly constrained hardware environments, substantially reducing operational hosting costs and memory transfer bottlenecks without sacrificing core capabilities. The algorithm shifts the practical threshold of low-bit model compression, making previously unfeasible sub-3-bit deployments viable for production systems.

Decision-makers should consider AQLM when hardware memory limits prevent standard 16-bit or 4-bit model deployments. If selecting codebook configurations, engineering teams should weigh the trade-off between higher-precision codebooks for maximum predictive accuracy and smaller codebooks for improved inference latency. Because the primary limitation of this method is high calibration compute time (such as requiring several days on multi-GPU setups for 70B parameter models), organizations should plan quantization workflows as offline preparation tasks before production rollout.

arXiv: 2401.06118Vahe1994/AQLM
Cover for Extreme Compression of Large Language Models via Additive Quantization

Abstract

The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of “extreme” LLM compression—defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter—from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.

Table of Contents

  • 1. Introduction
  • 2. Background & Related Work
  • 2.1. LLM Quantization
  • 2.2. Quantization for Nearest Neighbor Search
  • 3. AQLM: Additive Quantization for LLMs
  • 3.1. Overview
  • 3.2. Phase 1: Beam search for codes
  • 3.3. Phase 2: Codebook update
  • 3.4. Phase 3: Fine-tuning for intra-layer cohesion
  • 4. Experiments
  • 4.1. Compression quality for modern LLMs
  • 4.2. End-to-end fine-tuning experiments
  • 4.3. Ablation analysis
  • 4.4. Inference Speed
  • 5. Conclusion and Future Work
  • Acknowledgements
  • Impact Statement
  • References
  • A. End-to-end fine-tuning
  • B. Code reproducibility
  • C. Experimental Configurations
  • D. Quantization time
  • E. Ablation analysis
  • F. Additional experiments
  • F.1. Mixtral
  • F.2. LLAMA 2
  • F.3. Mistral
  • G. Pareto optimality
  • H. Estimating model size
  • I. End-to-End Inference Speed
  • J. Codebook and codes distribution
  • K. Evaluation on MMLU and GSM8k
  • L. Block-wise tuning for scalar quantization

Knowls

  1. Knowl 1 — AQLM minimizes calibration-input layer-output error

    model/method

    For a linear layer with weights W∈Rdout×dinW\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}} and calibration activations X∈Rdin×nX\in\mathbb{R}^{d_{\mathrm{in}}\times n}, AQLM chooses a quantized weight matrix WqW_q to preserve the layer outputs on those activations, rather than minimizing weight reconstruction error alone. Here, dind_{\mathrm{in}} and doutd_{\mathrm{out}} are the input and output dimensions, and nn is the number of calibration activation vectors. Its layer objective is

    min⁡Wq  ∥WX−WqX∥F2.\min_{W_q}\;\|WX-W_qX\|_F^2.

    The objective is instance-aware: the calibration inputs determine which weight errors matter. Equivalently, the loss is ∥(W−Wq)X∥F2\|(W-W_q)X\|_F^2, so activation covariance influences the relative cost of errors in different input directions.

  2. Knowl 2 — Additive codebooks encode groups of weights

    model/method

    AQLM divides each output row of a weight matrix into groups of gg consecutive input weights. Each group is approximated by summing one vector from each of MM learned codebooks. A codebook contains 2B2^B vectors of dimension gg, and its selected vector is specified by a BB-bit index. If Cm[:,am]C_m[:,a_m] denotes the vector selected from codebook mm by index ama_m, the group reconstruction for output row ii is si∑m=1MCm[:,ai,m]s_i\sum_{m=1}^{M}C_m[:,a_{i,m}], where sis_i is a learned scale for that output row. The codebooks are additive and are not constrained to be orthogonal or to occupy separate subvectors.

    With codebooks and scales stored in 16-bit precision, a layer with dimensions dind_{\mathrm{in}} and doutd_{\mathrm{out}} uses 16gM2B16gM2^B bits for codebooks, dout(din/g)MBd_{\mathrm{out}}(d_{\mathrm{in}}/g)MB bits for indices, and 16dout16d_{\mathrm{out}} bits for scales. Thus, its average storage rate in bits per weight is

    bˉ=16gM2B+dout(din/g)MB+16doutdoutdin.\bar b=\frac{16gM2^B+d_{\mathrm{out}}(d_{\mathrm{in}}/g)MB+16d_{\mathrm{out}}}{d_{\mathrm{out}}d_{\mathrm{in}}}.

    The compressed representation stores the indices and shared codebooks; decoding selects and adds codebook vectors for each group, then applies the row scale.

  3. Knowl 3 — AQLM learns codes and codebooks by alternating optimization

    model/method

    For each linear layer, AQLM initializes its codebooks and indices with residual kk-means: it clusters the weight groups, subtracts the first-stage cluster assignments, and repeatedly clusters the residuals to initialize subsequent codebooks. It then alternates between discrete code assignment and continuous codebook-and-scale updates, evaluating the calibration-output squared-error objective.

    For code assignment, the method expresses the objective using the precomputed input Gram matrix XXTXX^T and approximately solves the resulting fully connected discrete Markov random field with beam search. Starting from existing assignments, the search changes one code at a time, considers the 2B2^B possible indices for that code, and retains the best kk configurations. The loss can be updated incrementally because a code change affects only a subset of its terms; the search processes output rows in parallel. The paper does not prescribe one universal beam width.

    With indices fixed, codebooks and row scales are updated using full-batch Adam. The reported implementation uses 100 Adam steps per update phase, learning rate 10−410^{-4}, and β1=0.90\beta_1=0.90, β2=0.95\beta_2=0.95. Scales are initialized from the corresponding weight-row norms. The discrete and continuous updates are repeated until the loss improvement falls below a stopping tolerance, reported between 10−210^{-2} and 10−310^{-3}.

  4. Knowl 4 — Block-level tuning coordinates quantization across transformer layers

    model/method

    After independently quantizing the linear layers in a transformer block, AQLM fine-tunes parameters jointly across that block using the calibration data. A block typically contains 4–8 linear layers. For each calibration input XblockX_{\mathrm{block}}, the original block output, recorded before quantization, is the target; the loss is the squared error between that target and the output of the quantized block.

    During this tuning, the code indices remain fixed, while the additive codebooks, row scales, and non-quantized block parameters such as RMSNorm scales and biases are optimized with Adam through backpropagation. This allows quantization errors in different layers to be adjusted together without full-model quantization-aware training. The paper reports that block tuning uses the same calibration data as layer quantization and takes a minority of calibration time—typically 10–30% or less.

  5. Knowl 5 — End-to-end distillation further improves very-low-bit models

    model/method

    AQLM models can be further tuned end to end by treating the original floating-point model as a teacher and the quantized model as a student. The optimized parameters are the codebooks, scales, and non-quantized parameters; the discrete code indices stay fixed. The objective is the mean Kullback–Leibler divergence between student and teacher output distributions over calibration sequences.

    The reported procedure uses RedPajama data, Adam with constant learning rate 10−510^{-5} and no weight decay, batch sizes of 8–16 sequences, and one epoch. LLAMA 2 tuning uses 1,000–4,000 sequences of length 4,096; Mixtral tuning uses 512 sequences of length 8,192. For LLAMA 2 13B at 2.19 bits per parameter, end-to-end tuning changes WikiText-2 perplexity from 5.37 to 5.22, C4 perplexity from 7.16 to 6.98, and mean accuracy on five zero-shot tasks from 61.80% to 62.67%. The paper reports larger benefits at 2 bits than at 3 bits and above.

  6. Knowl 6 — AQLM improves roughly 2-bit results on LLAMA 2 and Mixtral

    empirical result

    The paper evaluates post-training quantization using WikiText-2 and C4 perplexity (lower is better) and the mean accuracy on five zero-shot tasks (higher is better). LLAMA 2 methods are calibrated on RedPajama sequences of length 4,096; the Mixtral comparison uses sequences of length 8,192. The results show that AQLM generally improves on QuIP# at similar approximately 2-bit rates, including on Mixtral, although the exact rates differ slightly.

    ModelMethodAverage bits/parameterWikiText-2 PPLC4 PPLMean zero-shot accuracy
    LLAMA 2 7BFP16165.126.6362.35%
    LLAMA 2 7BAQLM2.026.598.5457.28%
    LLAMA 2 7BQuIP#2.028.2211.0152.23%
    LLAMA 2 13BFP16164.576.0565.38%
    LLAMA 2 13BAQLM1.975.607.4961.32%
    LLAMA 2 13BQuIP#2.016.068.0757.55%
    LLAMA 2 70BFP16163.124.9770.17%
    LLAMA 2 70BAQLM2.073.945.7268.75%
    LLAMA 2 70BQuIP#2.014.166.0167.67%
    Mixtral 8×7BFP16163.465.0272.33%
    Mixtral 8×7BAQLM1.984.615.7567.68%
    Mixtral 8×7BQuIP#2.014.755.8966.34%

    On LLAMA 2 13B and 70B, the table also reports that non-# QuIP performs substantially worse than these methods: its WikiText-2 perplexities are 13.48 and 5.90, respectively.

  7. Knowl 7 — AQLM outperforms competing methods around 3 bits

    empirical result

    This comparison evaluates LLAMA 2 models at approximately 3 bits per parameter after post-training quantization. Metrics are WikiText-2 and C4 perplexity (lower is better) and mean accuracy across five zero-shot tasks (higher is better). Calibration uses RedPajama sequences of length 4,096. At these rates, AQLM has the lowest perplexity among the listed methods for each model size; it also has the highest mean zero-shot accuracy for 7B and 13B, while the FP16 baseline is slightly higher for 70B.

    ModelMethodAverage bits/parameterWikiText-2 PPLC4 PPLMean zero-shot accuracy
    LLAMA 2 7BFP16165.126.6362.35%
    LLAMA 2 7BAQLM3.045.467.0860.88%
    LLAMA 2 7BGPTQ3.008.0610.6153.08%
    LLAMA 2 7BSpQR2.986.208.2059.07%
    LLAMA 2 13BFP16164.576.0565.38%
    LLAMA 2 13BAQLM3.034.826.3764.49%
    LLAMA 2 13BGPTQ3.005.857.8659.61%
    LLAMA 2 13BSpQR2.985.287.0661.99%
    LLAMA 2 13BQuIP3.005.126.7963.15%
    LLAMA 2 70BFP16163.124.9770.17%
    LLAMA 2 70BAQLM3.013.365.1769.86%
    LLAMA 2 70BGPTQ3.004.406.2665.41%
    LLAMA 2 70BSpQR2.983.855.6368.22%
    LLAMA 2 70BQuIP3.013.875.6766.96%
  8. Knowl 8 — AQLM reaches a sub-3-bit accuracy–size Pareto frontier

    empirical result

    The paper defines Pareto optimality as achieving the highest accuracy at the same or a smaller total model size. Its LLAMA 2 experiments show that the best AQLM rate for this trade-off is around 2.5 bits per parameter, rather than exactly 2 bits. For example, LLAMA 2 13B at 2.76 bits has WikiText-2 perplexity 4.94 and mean five-task accuracy 64.15%, compared with 5.12 and 62.35% for uncompressed LLAMA 2 7B. The quantized 13B model is also smaller in bytes than the uncompressed 7B model. The authors therefore report AQLM as the first method to attain Pareto-optimal compression below 3 bits per parameter in their evaluation.

  9. Knowl 9 — AQLM kernels accelerate generation while using less memory

    empirical result

    The authors benchmarked generation of 128 tokens from scratch at batch size 1 using compiled computational graphs. GPU measurements use one Nvidia RTX 3090; CPU measurements use an Intel i9 with 8 cores. The table reports generated tokens per second for LLAMA 2 models. The 2×8-bit codebook format is the fastest reported option on both devices, though the paper notes that using multiple smaller codebooks can reduce accuracy relative to the larger-codebook configuration.

    Device and formatLLAMA 2 7BLLAMA 2 13BLLAMA 2 70B
    RTX 3090, FP1654.229.55.8
    RTX 3090, AQLM 1×16-bit65.334.16.7
    RTX 3090, AQLM 2×8-bit114.168.114.3
    Intel i9, FP323.1061.5960.297
    Intel i9, AQLM 2×8-bit6.9614.1800.966
    Intel i9, AQLM 4×8-bit6.8374.0040.948
    Intel i9, AQLM 8×8-bit5.3193.1930.775

    These measurements support the paper's practical claim that AQLM's smaller weight footprint need not entail slower inference: the tested AQLM formats meet or exceed the corresponding floating-point generation rates.

  10. Knowl 10 — Quantization cost and comparative scope are limitations

    limitation

    AQLM's learned additive representation and iterative optimization make model quantization more computationally expensive than direct post-training methods such as RTN or GPTQ. With the reported default configuration, quantizing LLAMA 2 7B takes about one day on one A100 GPU; quantizing 70B takes 10–14 days on one GPU. Parallel runs take about 14 hours for 7B on two GPUs and 3–4 days for 70B on eight GPUs. These costs apply to compression, not inference.

    The accuracy advantage is not universal across model families and bit rates. For Mistral 7B at approximately 2 bits, QuIP# slightly outperforms AQLM on most reported benchmarks; the paper reports closer results at 4 bits. Thus, the broad LLAMA 2 and Mixtral results do not establish that AQLM wins for every model or quantization setting.

Coverage note — Detailed calibration-size and codebook-configuration ablations, plus the secondary MMLU/GSM8k and scalar-quantization experiments, are omitted because they do not change the central method or its main accuracy–size and inference findings.

References

  1. 1.Babenko, A. and Lempitsky, V. Additive quantization for extreme vector compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 931–938, 2014.
  2. 2.Besag, J. On the statistical analysis of dirty pictures. Journal of the Royal Statistical Society Series B: Statistical Methodology, 48(3):259–279, 1986.
  3. 3.Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. arXiv preprint arXiv:2304.01373, 2023.
  4. 4.Blalock, D. and Guttag, J. Multiplying matrices without multiplying. In International Conference on Machine Learning, pp. 992–1004. PMLR, 2021.
  5. 5.Burton, D., Shore, J., and Buck, J. A generalization of isolated word recognition using vector quantization. In ICASSP ’83. IEEE International Conference on Acoustics, Speech, and Signal Processing, volume 8, pp. 1021–1024, 1983. doi: 10.1109/ICASSP.1983.1171915.
  6. 6.Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D. Quip: 2-bit quantization of large language models with guarantees, 2023.
  7. 7.Chen, S., Wang, W., and Pan, S. J. Deep neural network quantization via layer-wise optimization using limited training data. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):3329–3336, Jul. 2019. doi: 10.1609/aaai.v33i01.33013329. URL https://ojs.aaai.org/index.php/AAAI/article/view/4206.
  8. 8.Chen, Y., Guan, T., and Wang, C. Approximate nearest neighbor search by residual vector quantization. Sensors, 10(12):11259–11273, 2010.
  9. 9.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  10. 10.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  11. 11.Computer, T. Redpajama: an open dataset for training large language models, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  12. 12.Dettmers, T. and Zettlemoyer, L. The case for 4-bit precision: k-bit inference scaling laws. arXiv preprint arXiv:2212.09720, 2022.
  13. 13.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. LLM.int8(): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, 2022.
  14. 14.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. QLoRA: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023a.
  15. 15.Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078, 2023b.
  16. 16.Fernández-Marqués, J., AbouElhamayed, A. F., Lane, N. D., and Abdelfattah, M. S. Are we there yet? product quantization and its hardware acceleration. ArXiv, abs/2305.18334, 2023. URL https://api.semanticscholar.org/CorpusID:258967539.
  17. 17.Frantar, E. and Alistarh, D. Qmoe: Practical sub-1-bit compression of trillion-parameter models. arXiv preprint arXiv:2310.16795, 2023.
  18. 18.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pretrained transformers. arXiv preprint arXiv:2210.17323, 2022a.
  19. 19.Frantar, E., Singh, S. P., and Alistarh, D. Optimal Brain Compression: A framework for accurate post-training quantization and pruning. arXiv preprint arXiv:2208.11580, 2022b. Accepted to NeurIPS 2022, to appear.
  20. 20.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  21. 21.Ge, T., He, K., Ke, Q., and Sun, J. Optimized product quantization. IEEE transactions on pattern analysis and machine intelligence, 36(4):744–755, 2013.
  22. 22.Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  23. 23.Gray, R. Vector quantization. IEEE ASSP Magazine, 1(2):4–29, 1984. doi: 10.1109/MASSP.1984.1162229.
  24. 24.Guo, R., Kumar, S., Choromanski, K., and Simcha, D. Quantization based fast inner product search. In Artificial intelligence and statistics, pp. 482–490. PMLR, 2016.
  25. 25.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. CoRR, abs/2009.03300, 2020. URL https://arxiv.org/abs/2009.03300.
  26. 26.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network, 2015.
  27. 27.Jegou, H., Douze, M., and Schmid, C. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128, 2010.
  28. 28.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  29. 29.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  30. 30.Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023.
  31. 31.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. International Conference on Learning Representations (ICLR), 2015.
  32. 32.Li, Z., Ni, B., Zhang, W., Yang, X., and Gao, W. Performance guaranteed network acceleration via high-order residual quantization, 2017.
  33. 33.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  34. 34.Malinovskii, V., Mazur, D., Ilin, I., Kuznedelev, D., Burlachenko, K., Yi, K., Alistarh, D., and Richtarik, P. Pv-tuning: Beyond straight-through estimation for extreme llm compression. arXiv preprint arXiv:2405.14852, 2024.
  35. 35.Martinez, J., Clement, J., Hoos, H. H., and Little, J. J. Revisiting additive quantization. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pp. 137–153. Springer, 2016.
  36. 36.Martinez, J., Zakhmi, S., Hoos, H. H., and Little, J. J. Lsq++: Lower running time and higher recall in multi-codebook quantization. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 491–506, 2018.
  37. 37.McCarter, C. and Dronen, N. Look-ups are not (yet) all you need for deep learning inference. ArXiv, abs/2207.05808, 2022. URL https://api.semanticscholar.org/CorpusID:250491319.
  38. 38.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  39. 39.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? Adaptive rounding for post-training quantization. In International Conference on Machine Learning (ICML), 2020.
  40. 40.Norouzi, M. and Fleet, D. J. Cartesian k-means. In Proceedings of the IEEE Conference on computer Vision and Pattern Recognition, pp. 3017–3024, 2013.
  41. 41.Ozan, E. C., Kiranyaz, S., and Gabbouj, M. Competitive quantization for approximate nearest neighbor search. IEEE Transactions on Knowledge and Data Engineering, 28(11):2884–2894, 2016. doi: 10.1109/TKDE.2016.2597834.
  42. 42.Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. nuQmm: Quantized matmul for efficient inference of large-scale generative language models. arXiv preprint arXiv:2206.09557, 2022.
  43. 43.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An imperative style, high-performance deep learning library. In Conference on Neural Information Processing Systems (NeurIPS). 2019.
  44. 44.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  45. 45.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: an adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106, 2021. doi: 10.1145/3474381. URL https://doi.org/10.1145/3474381.
  46. 46.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  47. 47.Shazeer, N. Glu variants improve transformer, 2020.
  48. 48.Tata, S. and Patel, J. M. PiQA: An algebra for querying protein data sets. In International Conference on Scientific and Statistical Database Management, 2003.
  49. 49.TII UAE. The Falcon family of large language models. https://huggingface.co/tiiuae/falcon-40b, May 2023.
  50. 50.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  51. 51.Tseng, A., Chee, J., Sun, Q., Kuleshov, V., and Sa, C. D. Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024.
  52. 52.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  53. 53.Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438, 2022.
  54. 54.Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. arXiv preprint arXiv:2206.01861, 2022.
  55. 55.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D. R., and Màrquez, L. (eds.), Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pp. 4791–4800. Association for Computational Linguistics, 2019. doi: 10.18653/v1/p19-1472. URL https://doi.org/10.18653/v1/p19-1472.
  56. 56.Zhang, B. and Sennrich, R. Root mean square layer normalization. CoRR, abs/1910.07467, 2019. URL http://arxiv.org/abs/1910.07467.
  57. 57.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  58. 58.Zhang, T., Du, C., and Wang, J. Composite quantization for approximate nearest neighbor search. In International Conference on Machine Learning, pp. 838–846. PMLR, 2014.
  59. 59.Zhou, S.-C., Wang, Y.-Z., Wen, H., He, Q.-Y., and Zou, Y.-H. Balanced quantization: An effective and efficient approach to quantized neural networks. Journal of Computer Science and Technology, 32(4):667–682, Jul 2017. ISSN 1860-4749. doi: 10.1007/s11390-017-1750-y. URL https://doi.org/10.1007/s11390-017-1750-y.

Citation

MLA
Egiazarian, V., et al. “Extreme Compression of Large Language Models via Additive Quantization”. arXiv, 2024, http://arxiv.org/abs/2401.06118v4.
APA
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., & Alistarh, D. (2024). Extreme Compression of Large Language Models via Additive Quantization. arXiv. http://arxiv.org/abs/2401.06118v4
Chicago
Egiazarian, V., A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh. 2024. “Extreme Compression of Large Language Models via Additive Quantization”. arXiv. http://arxiv.org/abs/2401.06118v4.
Harvard
Egiazarian, V. et al. (2024) “Extreme Compression of Large Language Models via Additive Quantization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.06118v4.
Vancouver
1. Egiazarian V, Panferov A, Kuznedelev D, Frantar E, Babenko A, Alistarh D (2024) Extreme Compression of Large Language Models via Additive Quantization. arXiv

BibTeX

@article{egiazarian2024extreme,
  title = {Extreme Compression of Large Language Models via Additive Quantization},
  author = {Egiazarian, Vage and Panferov, Andrei and Kuznedelev, Denis and Frantar, Elias and Babenko, Artem and Alistarh, Dan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.06118v4},
  eprint = {2401.06118}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/