Built independently by an author, for readers. Read the story and support ChapterPal

keyword

additive quantization

Additive quantization is a vector compression and representation technique that approximates a target data vector as the sum of multiple component vectors chosen from distinct codebooks. Rather than splitting an input vector into separate, lower-dimensional orthogonal segments as done in product quantization, additive quantization uses codebook entries that can span the full vector space and combine additively to reconstruct the original data with minimal distortion. By encoding high-dimensional points as compact sets of discrete indices referencing these shared codebooks, this method enables efficient storage, fast approximate nearest neighbor search, and high-fidelity compression in areas such as signal processing and deep neural network parameter reduction.

3 items

Autoregressive Image Generation using Residual Quantization

Autoregressive Image Generation using Residual Quantization

Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, Wook-Shin Han

OrganizationsKakao BrainPohang University of Science and Technology

Why you should read this

Proposes a two-stage residual quantization framework, comprising RQ-VAE and RQ-Transformer, that drastically shortens discrete code sequences to achieve faster sampling and lower computational costs in high-resolution autoregressive image generation without sacrificing fidelity.

For autoregressive (AR) modeling of high-resolution images, vector quantization (VQ) represents an image as a sequence of discrete codes. A short sequence length is important for an AR model to reduce its computational costs to consider long-range interactions of codes. However, we postulate that previous VQ cannot shorten the code sequence and generate high-fidelity images together in terms of the rate-distortion trade-off. In this study, we propose the two-stage framework, which consists of Residual-Quantized VAE (RQ-VAE) and RQ-Transformer, to effectively generate high-resolution images. Given a fixed codebook size, RQ-VAE can precisely approximate a feature map of an image and represent the image as a stacked map of discrete codes. Then, RQ-Transformer learns to predict the quantized feature vector at the next position by predicting the next stack of codes. Thanks to the precise approximation of RQ-VAE, we can represent a 256×\times256 image as 8×\times8 resolution of the feature map, and RQ-Transformer can efficiently reduce the computational costs. Consequently, our framework outperforms the existing AR models on various benchmarks of unconditional and conditional image generation. Our approach also has a significantly faster sampling speed than previous AR models to generate high-quality images.

Added

2026-10-05

Extreme Compression of Large Language Models via Additive Quantization

Extreme Compression of Large Language Models via Additive Quantization

Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh

OrganizationsHigher School of EconomicsInstitute of Science and Technology AustriaNeural MagicSkolkovo Institute of Science and TechnologyYandex

Why you should read this

Introduces AQLM, a multi-codebook quantization method that compresses large language model weights down to 2 bits per parameter while achieving state-of-the-art accuracy and matching 16-bit floating-point inference speeds on standard hardware.

The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices. In this paper, we revisit the problem of “extreme” LLM compression—defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter—from the point of view of classic methods in Multi-Codebook Quantization (MCQ). Our algorithm, called AQLM, generalizes the classic Additive Quantization (AQ) approach for information retrieval to advance the state-of-the-art in LLM compression, via two innovations: 1) learned additive quantization of weight matrices in input-adaptive fashion, and 2) joint optimization of codebook parameters across each transformer blocks. Broadly, AQLM is the first scheme that is Pareto optimal in terms of accuracy-vs-model-size when compressing to less than 3 bits per parameter, and significantly improves upon all known schemes in the extreme compression (2bit) regime. In addition, AQLM is practical: we provide fast GPU and CPU implementations of AQLM for token generation, which enable us to match or outperform optimized FP16 implementations for speed, while executing in a much smaller memory footprint.

Added

2026-09-30

Accelerating Large-Scale Inference with Anisotropic Vector Quantization

Accelerating Large-Scale Inference with Anisotropic Vector Quantization

Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, Sanjiv Kumar

OrganizationsGoogle

Why you should read this

Develops an anisotropic vector quantization approach that significantly improves large-scale maximum inner product search by prioritizing reconstruction accuracy for higher-scoring database points, achieving state-of-the-art results on public benchmarks.

Quantization based techniques are the current state-of-the-art for scaling maximum inner product search to massive databases. Traditional approaches to quantization aim to minimize the reconstruction error of the database points. Based on the observation that for a given query, the database points that have the largest inner products are more relevant, we develop a family of anisotropic quantization loss functions. Under natural statistical assumptions, we show that quantization with these loss functions leads to a new variant of vector quantization that more greatly penalizes the parallel component of a datapoint's residual relative to its orthogonal component. The proposed approach achieves state-of-the-art results on the public benchmarks available at \url{this http URL}.

Added

2026-05-04

Creative Commons License