Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Gradient estimators

Gradient estimators are computational techniques used in machine learning and optimization to approximate the gradient of an objective function with respect to its parameters when exact analytical derivatives are intractable, unavailable, or computationally prohibitive to calculate. They are widely applied when optimizing systems that involve stochastic expectations, discrete latent variables, black-box simulators, or non-differentiable operations such as data quantization. Common approaches include score function estimators that evaluate function values across random samples, pathwise derivative estimators that propagate gradients through reparameterized continuous distributions, and surrogate techniques that substitute non-differentiable operations with smooth approximations during backpropagation. A primary objective in designing gradient estimators is managing the fundamental trade-off between bias and variance to enable stable and efficient convergence in gradient-based optimization algorithms.

3 items

Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training

Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training

Charbel Sakr, Steve Dai, Rangharajan Venkatesan, Brian Zimmer, William J. Dally, Brucek Khailany

OrganizationsNVIDIA

Why you should read this

Proposes a fast Newton-Raphson-based algorithm to dynamically compute MSE-optimal clipping scalars alongside magnitude-aware differentiation, achieving state-of-the-art accuracy in low-precision quantization-aware training without altering standard baseline hyperparameters.

Data clipping is crucial in reducing noise in quantization operations and improving the achievable accuracy of quantization-aware training (QAT). Current practices rely on heuristics to set clipping threshold scalars and cannot be shown to be optimal. We propose Optimally Clipped Tensors And Vectors (OCTAV), a recursive algorithm to determine MSE-optimal clipping scalars. Derived from the fast Newton-Raphson method, OCTAV finds optimal clipping scalars on the fly, for every tensor, at every iteration of the QAT routine. Thus, the QAT algorithm is formulated with provably minimum quantization noise at each step. In addition, we reveal limitations in common gradient estimation techniques in QAT and propose magnitude-aware differentiation as a remedy to further improve accuracy. Experimentally, OCTAV-enabled QAT achieves state-of-the-art accuracy on multiple tasks. These include training-from-scratch and retraining ResNets and MobileNets on ImageNet, and Squad fine-tuning using BERT models, where OCTAV-enabled QAT consistently preserves accuracy at low precision (4-to-6-bits). Our results require no modifications to the baseline training recipe, except for the insertion of quantization operations where appropriate.

Added

2026-10-03

Do Differentiable Simulators Give Better Policy Gradients?

Do Differentiable Simulators Give Better Policy Gradients?

Hyung Ju Terry Suh, Max Simchowitz, Kaiqing Zhang, Russ Tedrake

OrganizationsMassachusetts Institute of Technology

Why you should read this

Explains why physical discontinuities and stiffness introduce severe bias and variance into differentiable simulator gradients, while developing a hybrid alpha-order estimator that effectively blends zeroth- and first-order techniques for reliable policy optimization.

Differentiable simulators promise faster computation time for reinforcement learning by replacing zeroth-order gradient estimates of a stochastic objective with an estimate based on first-order gradients. However, it is yet unclear what factors decide the performance of the two estimators on complex landscapes that involve long-horizon planning and control on physical systems, despite the crucial relevance of this question for the utility of differentiable simulators. We show that characteristics of certain physical systems, such as stiffness or discontinuities, may compromise the efficacy of the first-order estimator, and analyze this phenomenon through the lens of bias and variance. We additionally propose an α-order gradient estimator, with α ∈ [0, 1], which correctly utilizes exact gradients to combine the efficiency of first-order estimates with the robustness of zeroth-order methods. We demonstrate the pitfalls of traditional estimators and the advantages of the α-order estimator on some numerical examples.

Added

2026-10-01

Neural Variational Inference and Learning in Belief Networks

Neural Variational Inference and Learning in Belief Networks

Andriy Mnih, Karol Gregor

OrganizationsGoogle

Why you should read this

Demonstrates how to train complex probabilistic models with discrete and continuous variables efficiently by using a neural network to approximate posterior inference and applying practical variance reduction techniques to make gradient estimation feasible.

Highly expressive directed latent variable models, such as sigmoid belief networks, are difficult to train on large datasets because exact inference in them is intractable and none of the approximate inference methods that have been applied to them scale well. We propose a fast non-iterative approximate inference method that uses a feedforward network to implement efficient exact sampling from the variational posterior. The model and this inference network are trained jointly by maximizing a variational lower bound on the log-likelihood. Although the naive estimator of the inference model gradient is too high-variance to be useful, we make it practical by applying several straightforward model-independent variance reduction techniques. Applying our approach to training sigmoid belief networks and deep autoregressive networks, we show that it outperforms the wake-sleep algorithm on MNIST and achieves state-of-the-art results on the Reuters RCV1 document dataset.

Added

2026-02-21