Built independently by an author, for readers. Read the story and support ChapterPal

keyword

neuron pruning

Neuron pruning is a technique in machine learning that compresses artificial neural networks by identifying and removing redundant or less important neurons along with all of their incoming and outgoing connections. Unlike unstructured weight pruning, which sets individual connection weights to zero across a model, neuron pruning eliminates entire computational nodes, directly reducing the dimensional size of network layers. This form of structured pruning decreases memory consumption, reduces computational complexity, and accelerates training and inference on standard hardware, while also acting as a regularizer that can prevent overfitting and maintain or improve generalization performance.

1 item

Learning Sparse Neural Networks through L0 Regularization

Learning Sparse Neural Networks through L0 Regularization

Christos Louizos, Max Welling, Diederik P. Kingma

OrganizationsCIFAROpenAITNOUniversity of Amsterdam

Why you should read this

Develops a continuous relaxation method using hard concrete stochastic gates to make L0L_0 regularization differentiable, enabling neural networks to automatically prune weights during training via standard gradient descent.

We propose a practical method for L0L_0 norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known model selection criteria, are special cases of L0L_0 regularization. However, since the L0L_0 norm of weights is non-differentiable, we cannot incorporate it directly as a regularization term in the objective function. We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected L0L_0 norm of the resulting gated weights is differentiable with respect to the distribution parameters. We further propose the \emph{hard concrete} distribution for the gates, which is obtained by "stretching" a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.

Added

2026-09-25