Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep equilibrium models

Deep equilibrium models are a class of implicit neural network architectures that compute representations by finding the stable fixed point, or equilibrium state, of a parameterized transformation rather than passing data through a predefined sequence of discrete layers. By treating the forward pass as finding a steady-state solution to an infinite-depth weight-tied network, these models determine hidden states using numerical root-finding or fixed-point iteration algorithms. During training, gradients are calculated analytically at the equilibrium state using implicit differentiation based on the implicit function theorem, which eliminates the need to store intermediate activations and allows the network to train with constant memory overhead regardless of effective depth.

6 items

SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation

SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation

Malyaban Bal, Abhronil Sengupta

OrganizationsPennsylvania State UniversitySchool of Electrical Engineering and Computer Science

Why you should read this

Proposes a scalable spiking language model trained via equilibrium-state implicit differentiation and BERT knowledge distillation, bypassing non-differentiability challenges without surrogate gradients to deliver energy-efficient NLP performance on the GLUE benchmark.

Large language Models (LLMs), though growing exceedingly powerful, comprises of orders of magnitude less neurons and synapses than the human brain. However, it requires significantly more power/energy to operate. In this work, we propose a novel bio-inspired spiking language model (LM) which aims to reduce the computational cost of conventional LMs by drawing motivation from the synaptic information flow in the brain. In this paper, we demonstrate a framework that leverages the average spiking rate of neurons at equilibrium to train a neuromorphic spiking LM using implicit differentiation technique, thereby overcoming the non-differentiability problem of spiking neural network (SNN) based algorithms without using any type of surrogate gradient. The steady-state convergence of the spiking neurons also allows us to design a spiking attention mechanism, which is critical in developing a scalable spiking LM. Moreover, the convergence of average spiking rate of neurons at equilibrium is utilized to develop a novel ANN-SNN knowledge distillation based technique wherein we use a pre-trained BERT model as “teacher” to train our “student” spiking architecture. While the primary architecture proposed in this paper is motivated by BERT, the technique can be potentially extended to different kinds of LLMs. Our work is the first one to demonstrate the performance of an operational spiking LM architecture on multiple different tasks in the GLUE benchmark. Our implementation source code is available at https://github.com/NeuroCompLab-psu/SpikingBERT.

Added

2026-09-26

Deep Equilibrium Approaches to Diffusion Models

Deep Equilibrium Approaches to Diffusion Models

Ashwini Pokle, Zhengyang Geng, J. Zico Kolter

OrganizationsBosch Center for AICarnegie Mellon University

Why you should read this

Formulates the diffusion sampling chain as a joint fixed-point system using deep equilibrium models, enabling parallel multi-GPU image generation and memory-efficient backpropagation for faster model inversion.

Diffusion-based generative models are extremely effective in generating high-quality images, with generated samples often surpassing the quality of those produced by other models under several metrics. One distinguishing feature of these models, however, is that they typically require long sampling chains to produce high-fidelity images. This presents a challenge not only from the lenses of sampling time, but also from the inherent difficulty in backpropagating through these chains in order to accomplish tasks such as model inversion, i.e., approximately finding latent states that generate known images. In this paper, we look at diffusion models through a different perspective, that of a (deep) equilibrium (DEQ) fixed point model. Specifically, we extend the recent denoising diffusion implicit model (DDIM) [68], and model the entire sampling chain as a joint, multi-variate fixed point system. This setup provides an elegant unification of diffusion and equilibrium models, and shows benefits in 1) single image sampling, as it replaces the fully-serial typical sampling process with a parallel one; and 2) model inversion, where we can leverage fast gradients in the DEQ setting to much more quickly find the noise that generates a given image. The approach is also orthogonal and thus complementary to other methods used to reduce the sampling time, or improve model inversion. We demonstrate our method’s strong performance across several datasets, including CIFAR10, CelebA, and LSUN Bedroom and Churches.

Added

2026-09-26

Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning

Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning

Benhao Huang, Zhengyang Geng, Zico Kolter

OrganizationsCarnegie Mellon University

Why you should read this

Introduces Equilibrium Reasoners, demonstrating that scalable test-time reasoning emerges from learning task-conditioned dynamical attractors, which allows neural networks to adaptively scale computation depth and breadth without external verifiers to boost Sudoku-Extreme accuracy from 2.6% to over 99%.

Scaling test-time compute by iteratively updating a latent state has emerged as a powerful paradigm for reasoning. Yet the internal mechanisms that enable these iterative models to generalize beyond memorized patterns remain unclear. We hypothesize that generalizable reasoning arises from learning task-conditioned attractors: latent dynamical systems whose stable fixed points correspond to valid solutions. We formalize this process through Equilibrium Reasoners (EqR), which enable test-time scaling without external verifiers or task-specific priors. EqR scales internal dynamics along two axes: depth, by running more iterations, and breadth, by aggregating stochastic trajectories from multiple initializations. Empirically, gains from test-time scaling are tightly coupled with stronger convergence toward solution-aligned attractors. This attractor perspective allows neural networks to adaptively allocate test-time compute based on task difficulty. While simple cases converge within 1 to 5 iteration steps, harder cases benefit from massive test-time scaling. By unrolling up to the equivalent of 40,000 layers, scalable latent reasoning boosts accuracy from 2.6% for feedforward models to over 99% on Sudoku-Extreme. These results suggest that learned attractor landscapes provide a useful mechanistic lens for understanding scalable reasoning in iterative latent models.

Added

2026-09-07

Creative Commons License
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Amr Hegazy, Amr Alanwar, Mostafa Elhoushi

OrganizationsCerebras SystemsGerman University in CairoTechnical University of Munich

Why you should read this

Introduces a gated recurrent transformer architecture that uses state-conditioned modulation across iterated shared layers to match standard language model performance while reducing parameters and peak decoding memory by roughly sixty percent.

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.

Added

2026-08-29

Creative Commons License
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto

OrganizationsELLIS Institute TübingenETH ZurichLiquid AIMax Planck Institute for Intelligent SystemsSwiss Institute of BioinformaticsTübingen AI CenterUniversité Paris Cité

Why you should read this

Presents FPRM, a Transformer-based Fixed-Point Reasoning Model that leverages fixed-point convergence to dynamically adapt its computational effort to task difficulty in looped architectures, addressing signal propagation issues through architectural modifications.

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these models find. Like deep architectures, looped architectures are prone to a signal propagation problem induced by depth as the halting decision is postponed. In this paper, we address this signal propagation issue using pre-norm layers and residual scaling. Building on these architectural modifications, we propose FPRM, a Transformer-based Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to task difficulty. FPRM is effective on common reasoning benchmarks, namely Sudoku, Maze, state-tracking, and ARC-AGI.

Added

2026-06-21

Creative Commons License
Less is More: Recursive Reasoning with Tiny Networks

Less is More: Recursive Reasoning with Tiny Networks

Alexia Jolicoeur-Martineau

Why you should read this

Presents a groundbreaking approach, Tiny Recursive Model (TRM), that achieves significantly higher generalization on complex tasks like ARC-AGI than LLMs with orders of magnitude fewer parameters, demonstrating that "less is more" in AI reasoning.

Hierarchical Reasoning Model (HRM) is a novel approach using two small neural networks recursing at different frequencies. This biologically inspired method beats Large Language models (LLMs) on hard puzzle tasks such as Sudoku, Maze, and ARC-AGI while trained with small models (27M parameters) on small data (around 1000 examples). HRM holds great promise for solving hard problems with small networks, but it is not yet well understood and may be suboptimal. We propose Tiny Recursive Model (TRM), a much simpler recursive reasoning approach that achieves significantly higher generalization than HRM, while using a single tiny network with only 2 layers. With only 7M parameters, TRM obtains 45% test-accuracy on ARC-AGI-1 and 8% on ARC-AGI-2, higher than most LLMs (e.g., Deepseek R1, o3-mini, Gemini 2.5 Pro) with less than 0.01% of the parameters.

Added

2025-11-23

License

Published with permission