Built independently by an author, for readers. Read the story and support ChapterPal

keyword

stochastic quantization

Stochastic quantization is a technique in machine learning and digital signal processing in which continuous or high-precision numerical values, such as neural network weights, activations, gradients, or latent representations, are mapped to a discrete set of low-precision values or codebook vectors through probabilistic sampling rather than deterministic rounding. In this approach, the probability of selecting a particular discrete state is typically proportional to its proximity to the original continuous value, which often allows the expected quantized output to serve as an unbiased estimator of the input. By introducing controlled randomness into the discretization process, stochastic quantization helps mitigate systematic rounding errors, improves codebook utilization, aids gradient estimation during model training, and enables efficient data compression in distributed optimization without collapsing representational diversity.

5 items

Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks

Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks

Minyoung Huh, Brian Cheung, Pulkit Agrawal, Phillip Isola

OrganizationsMassachusetts Institute of Technology

Why you should read this

Identifies internal codebook covariate shift as the fundamental cause of index collapse in vector-quantized networks and introduces an affine re-parameterization alongside alternating optimization to stabilize training across vision and generative architectures.

This work examines the challenges of training neural networks using vector quantization using straight-through estimation. We find that a primary cause of training instability is the discrepancy between the model embedding and the code-vector distribution. We identify the factors that contribute to this issue, including the codebook gradient sparsity and the asymmetric nature of the commitment loss, which leads to misaligned code-vector assignments. We propose to address this issue via affine re-parameterization of the code vectors. Additionally, we introduce an alternating optimization to reduce the gradient error introduced by the straight-through estimation. Moreover, we propose an improvement to the commitment loss to ensure better alignment between the codebook representation and the model embedding. These optimization methods improve the mathematical approximation of the straight-through estimation and, ultimately, the model performance. We demonstrate the effectiveness of our methods on several common model architectures, such as AlexNet, ResNet, and ViT, across various tasks, including image classification and generative modeling. Project page: minyoungg.github.io/vqtorch

Added

2026-09-26

Regularized Vector Quantization for Tokenized Image Synthesis

Regularized Vector Quantization for Tokenized Image Synthesis

Jiahui Zhang, Fangneng Zhan, Christian Theobalt, Shijian Lu

OrganizationsMax Planck Institute for InformaticsNanyang Technological University

Why you should read this

Proposes a dual-regularized vector quantization framework with a probabilistic contrastive loss that prevents codebook collapse and aligns training with stochastic sampling for superior image synthesis in autoregressive and diffusion models.

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distribution. However, deterministic quantization suffers from severe codebook collapse and misalignment with inference stage while stochastic quantization suffers from low codebook utilization and perturbed reconstruction objective. This paper presents a regularized vector quantization framework that allows to mitigate above issues effectively by applying regularization from two perspectives. The first is a prior distribution regularization which measures the discrepancy between a prior token distribution and the predicted token distribution to avoid codebook collapse and low codebook utilization. The second is a stochastic mask regularization that introduces stochasticity during quantization to strike a good balance between inference stage misalignment and unperturbed reconstruction objective. In addition, we design a probabilistic contrastive loss which serves as a calibrated metric to further mitigate the perturbed reconstruction objective. Extensive experiments show that the proposed quantization framework outperforms prevailing vector quantization methods consistently across different generative models including auto-regressive models and diffusion models.

Added

2026-09-26

Binarized Neural Networks

Binarized Neural Networks

Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio

OrganizationsCIFARColumbia UniversityTechnion – Israel Institute of TechnologyUniversité de Montréal

Why you should read this

Proposes a method to train deep neural networks with binary weights and activations, replacing standard arithmetic operations with fast bitwise operations to drastically reduce memory use and energy consumption while retaining competitive accuracy.

We introduce a method to train Binarized Neural Networks (BNNs) - neural networks with binary weights and activations at run-time. At training-time the binary weights and activations are used for computing the parameters gradients. During the forward pass, BNNs drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations, which is expected to substantially improve power-efficiency. To validate the effectiveness of BNNs we conduct two sets of experiments on the Torch7 and Theano frameworks. On both, BNNs achieved nearly state-of-the-art results over the MNIST, CIFAR-10 and SVHN datasets. Last but not least, we wrote a binary matrix multiplication GPU kernel with which it is possible to run our MNIST BNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy. The code for training and running our BNNs is available on-line.

Added

2026-09-26

License

Published with permission

QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

Dan Alistarh, Demjan Grubic, Jungshian Li, Ryota Tomioka, Milan Vojnovic

OrganizationsETH ZurichGoogleInstitute of Science and Technology AustriaLondon School of EconomicsMassachusetts Institute of TechnologyMicrosoft

Why you should read this

Proposes Quantized SGD (QSGD), a communication-efficient gradient compression scheme with provable convergence guarantees that significantly accelerates distributed deep learning without sacrificing model accuracy.

Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to excellent scalability properties of this algorithm, and to its efficiency in the context of training deep neural networks. A fundamental barrier for parallelizing large-scale SGD is the fact that the cost of communicating the gradient updates between nodes can be very large. Consequently, lossy compression heuristics have been proposed, by which nodes only communicate quantized gradients. Although effective in practice, these heuristics do not always provably converge, and it is not clear whether they are optimal. In this paper, we propose Quantized SGD (QSGD), a family of compression schemes which allow the compression of gradient updates at each node, while guaranteeing convergence under standard assumptions. QSGD allows the user to trade off compression and convergence time: it can communicate a sublinear number of bits per iteration in the model dimension, and can achieve asymptotically optimal communication cost. We complement our theoretical results with empirical data, showing that QSGD can significantly reduce communication cost, while being competitive with standard uncompressed techniques on a variety of real tasks. In particular, experiments show that gradient quantization applied to training of deep neural networks for image classification and automated speech recognition can lead to significant reductions in communication cost, and end-to-end training time. For instance, on 16 GPUs, we are able to train a ResNet-152 network on ImageNet 1.8x faster to full accuracy. Of note, we show that there exist generic parameter settings under which all known network architectures preserve or slightly improve their full accuracy when using quantization.

Added

2026-09-19

Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations

Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations

Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio

OrganizationsColumbia UniversityTechnion – Israel Institute of TechnologyUniversité de Montréal

Why you should read this

Demonstrates how to train neural networks using low-precision weights and activations down to 1 bit, replacing standard arithmetic with fast bitwise operations to drastically reduce memory footprint and power consumption without sacrificing competitive accuracy.

We introduce a method to train Quantized Neural Networks (QNNs) --- neural networks with extremely low precision (e.g., 1-bit) weights and activations, at run-time. At train-time the quantized weights and activations are used for computing the parameter gradients. During the forward pass, QNNs drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations. As a result, power consumption is expected to be drastically reduced. We trained QNNs over the MNIST, CIFAR-10, SVHN and ImageNet datasets. The resulting QNNs achieve prediction accuracy comparable to their 32-bit counterparts. For example, our quantized version of AlexNet with 1-bit weights and 2-bit activations achieves 51%51\% top-1 accuracy. Moreover, we quantize the parameter gradients to 6-bits as well which enables gradients computation using only bit-wise operation. Quantized recurrent neural networks were tested over the Penn Treebank dataset, and achieved comparable accuracy as their 32-bit counterparts using only 4-bits. Last but not least, we programmed a binary matrix multiplication GPU kernel with which it is possible to run our MNIST QNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy. The QNN code is available online.

Added

2026-09-18