Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Multi-codebook quantization

Multi-codebook quantization is a method of representing a vector or parameter group compactly by selecting one codeword from each of several codebooks and combining the selected codewords—often by adding them—to approximate the original values. The compact representation stores the codeword selections rather than the full-precision values, trading memory use for some approximation error.

1 item

Extreme Compression of Large Language Models via Additive Quantization

Extreme Compression of Large Language Models via Additive Quantization

Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh

OrganizationsHigher School of EconomicsInstitute of Science and Technology AustriaNeural MagicSkolkovo Institute of Science and TechnologyYandex

Why you should read this

Introduces AQLM, a multi-codebook quantization method that compresses large language model weights down to 2 bits per parameter while achieving state-of-the-art accuracy and matching 16-bit floating-point inference speeds on standard hardware.

The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of “extreme” LLM compression—defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter—from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.

Added

2026-09-30