keyword
gradient checkpointing
Gradient checkpointing, also known as activation checkpointing or rematerialization, is a deep learning training technique that reduces accelerator memory usage by storing only a subset of intermediate activations during the forward pass and recomputing the remaining ones on demand during backpropagation. In standard neural network training, all layer activations calculated during the forward pass are retained in memory to compute parameter gradients during the backward pass, which can cause out-of-memory errors in very large models, long context sequences, or large batch sizes. By dividing the network into segments and discarding intermediate activations within those segments until they are recalculated during the backward pass, gradient checkpointing trades a modest increase in computational overhead for a substantial decrease in peak memory consumption, enabling the training of larger architectures on memory-constrained hardware.
4 items

Diffeomorphic Optimization
Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel
Why you should read this
Introduces diffeomorphic optimization to perform Riemannian gradient descent through the base space of generative models, extending the framework to Lie groups to achieve superior accuracy and speed in computational protein design.
Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space. Optimizing differentiable objectives on this manifold is challenging: the ambient loss landscape is high-dimensional, rugged, and non-convex. Direct gradient descent, blind to the manifold's geometry, quickly drifts off it. Diffeomorphic optimization starts from the observation that diffusion and flow models provide a map from the data manifold to a much simpler base space in which we perform gradient descent. Using differential geometry, we show this is equivalent to Riemannian gradient descent on the data manifold up to corrections, keeping trajectories on-manifold by construction and yielding a smoother optimization surface. For protein design, we extend diffeomorphic optimization to the matrix Lie groups and , deriving an autograd-compatible gradient and a generalized adjoint-state method for backpropagation through Lie-group ODE solvers. Diffeomorphic optimization improves over tuned guidance on secondary-structure targeting with FrameFlow ( vs. of residues in the Ramachandran target), outperforms OC-Flow on peptide binding affinity at the speed, and reduces Rosetta energies by thousands of units across the PDB test set for structures with hundreds of residues.
Added
2026-09-29

Learning to (Learn at Test Time): RNNs with Expressive Hidden States
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin
Why you should read this
Introduces Test-Time Training layers that treat RNN hidden states as internal machine learning models updated via self-supervised gradient steps, achieving linear-time sequence modeling that continues to improve across long contexts where existing architectures plateau.
Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with linear complexity and expressive hidden states. The key idea is to make the hidden state a machine learning model itself, and the update rule a step of self-supervised learning. Since the hidden state is updated by training even on test sequences, our layers are called Test-Time Training (TTT) layers. We consider two instantiations: TTT-Linear and TTT-MLP, whose hidden state is a linear model and a two-layer MLP respectively. We evaluate our instantiations at the scale of 125M to 1.3B parameters, comparing with a strong Transformer and Mamba, a modern RNN. Similar to Transformer, TTT-Linear and TTT-MLP can keep reducing perplexity by conditioning on more tokens, while Mamba cannot after 16k context. TTT-MLP still faces challenges in memory I/O, but shows larger potential in long context, pointing to a promising direction for future research.
Added
2026-09-26

Glow: Generative Flow with Invertible 1x1 Convolutions
Diederik P. Kingma, Prafulla Dhariwal
Why you should read this
Introduces Glow, a flow-based generative model using invertible 1x1 convolutions that enables efficient high-resolution image synthesis and realistic semantic manipulation while maintaining exact log-likelihood computation.
Flow-based generative models (Dinh et al., 2014) are conceptually attractive due to tractability of the exact log-likelihood, tractability of exact latent-variable inference, and parallelizability of both training and synthesis. In this paper we propose Glow, a simple type of generative flow using an invertible 1x1 convolution. Using our method we demonstrate a significant improvement in log-likelihood on standard benchmarks. Perhaps most strikingly, we demonstrate that a generative model optimized towards the plain log-likelihood objective is capable of efficient realistic-looking synthesis and manipulation of large images. The code for our model is available at this https URL
Added
2026-09-11

Fine-Tuning Language Models with Just Forward Passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, Sanjeev Arora
Why you should read this
Explodes the traditional requirement for backpropagation by applying a zeroth-order optimizer that accurately estimates gradients using only forward evaluations to radically reduce memory consumption.
Fine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory. Zeroth-order (ZO) methods can in principle estimate gradients using only two forward passes but are theorized to be catastrophically slow for optimizing large models. In this work, we propose a memory-efficient zerothorder optimizer (MeZO), adapting the classical ZO-SGD method to operate in-place, thereby fine-tuning LMs with the same memory footprint as inference. For example, with a single A100 80GB GPU, MeZO can train a 30-billion parameter model, whereas fine-tuning with backpropagation can train only a 2.7B LM with the same budget. We conduct comprehensive experiments across model types (masked and autoregressive LMs), model scales (up to 66B), and downstream tasks (classification, multiple-choice, and generation). Our results demonstrate that (1) MeZO significantly outperforms in-context learning and linear probing; (2) MeZO achieves comparable performance to fine-tuning with backpropagation across multiple tasks, with up to 12x memory reduction and up to 2x GPU-hour reduction in our implementation; (3) MeZO is compatible with both full-parameter and parameter-efficient tuning techniques such as LoRA and prefix tuning; (4) MeZO can effectively optimize non-differentiable objectives (e.g., maximizing accuracy or F1). We support our empirical findings with theoretical insights, highlighting how adequate pre-training and task prompts enable MeZO to fine-tune huge models, despite classical ZO analyses suggesting otherwise.
Added
2026-06-22
