Learning Sparse Neural Networks through L0 Regularization

Christos LouizosMax WellingDiederik P. Kingma

article2017ICLR1,354 citations

Develops a continuous relaxation method using hard concrete stochastic gates to make L0L_0 regularization differentiable, enabling neural networks to automatically prune weights during training via standard gradient descent.

Listen

Modern deep neural networks deliver strong performance across many domains, but they are often heavily overparameterized. This excessive size leads to unnecessary computational and memory costs during training and deployment, while also increasing the risk of overfitting by memorizing random patterns in data. Regularizing models by penalizing the exact count of non-zero parameters—known as L0 regularization—is theoretically ideal for model compression and generalization because it directly prunes away unneeded connections without artificially shrinking useful weights. However, optimizing an L0 penalty has historically been computationally intractable because counting non-zero parameters is discrete and non-differentiable.

The article aims to introduce and validate an efficient, practical mathematical framework that enables direct L0 regularization during standard neural network training. It demonstrates that continuous stochastic gating mechanisms can prune redundant weights and neurons to exact zeros while maintaining compatibility with standard gradient-based optimization algorithms.

To achieve this, the authors introduce continuous random variables that act as stochastic gates on network weights. By transforming these variables with a stretched distribution and a hard-sigmoid rectification—termed the hard concrete distribution—the gates can take exact values of zero or one while remaining differentiable in expectation. The approach was evaluated across standard image recognition benchmarks, including multilayer perceptrons and convolutional networks on the MNIST dataset, as well as Wide Residual Networks on the CIFAR-10 and CIFAR-100 datasets.

The experimental findings show that the proposed method prunes networks dynamically during training without sacrificing model accuracy. On MNIST, the method successfully reduced the parameter footprint across fully connected and convolutional architectures, achieving competitive compression rates compared to complex multi-stage pruning techniques. On CIFAR-10 and CIFAR-100, the method improved classification accuracy over standard dropout baselines while steadily reducing the expected floating-point operations required per training step. Furthermore, the approach allows group sparsity, enabling the simultaneous pruning of entire neurons or convolutional feature maps to deliver practical hardware speedups.

These results demonstrate that model pruning does not need to be an expensive, post-hoc procedure requiring full-model pre-training followed by threshold-based fine-tuning. By pruning directly during training, organizations can reduce computational and infrastructure costs, lower energy consumption, and deploy smaller models to edge devices without degrading accuracy. The method also provides a principled, continuous relaxation of discrete model selection criteria, avoiding the biased gradients or heuristic workarounds found in earlier approaches.

Engineering and research teams should consider adopting this gating formulation as a drop-in replacement for standard dropout or heuristic pruning pipelines when model efficiency and inference latency are critical requirements. When implementing the approach, practitioners should tailor regularization strengths per layer to prioritize pruning computationally intensive convolutional feature maps over less demanding layers. Future efforts should focus on integrating this method with specialized hardware implementations to fully capitalize on dynamic runtime speedups and extending the gating mechanism to broader discrete latent variable models.

While the theoretical foundation and empirical results are strong across benchmark vision tasks, readers should note certain practical limitations. The realized execution speedups depend on whether runtime software and hardware efficiently exploit structured group sparsity rather than unstructured, element-wise weight sparsity. Additionally, selecting the proper regularization strength requires standard hyperparameter tuning to balance model compression against task accuracy.

arXiv: 1712.01312
Cover for Learning Sparse Neural Networks through L0 Regularization

Abstract

We propose a practical method for L0L_0 norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known model selection criteria, are special cases of L0L_0 regularization. However, since the L0L_0 norm of weights is non-differentiable, we cannot incorporate it directly as a regularization term in the objective function. We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected L0L_0 norm of the resulting gated weights is differentiable with respect to the distribution parameters. We further propose the \emph{hard concrete} distribution for the gates, which is obtained by "stretching" a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.

Table of Contents

  • 1 Introduction
  • 2 Minimizing the L0L_{0} norm of parametric models
  • 2.1 A general recipe for efficiently minimizing L0L_{0} norms
  • 2.2 The hard concrete distribution
  • 2.3 Combining the L0L_{0} norm with other norms
  • 2.4 Group sparsity under an L0L_{0} norm
  • 3 Related work
  • 4 Experiments
  • 4.1 MNIST classification and sparsification
  • 4.2 CIFAR classification
  • 5 Discussion
  • References
  • A Relation to variational inference
  • B The hard concrete distribution
  • C Negative KL-divergence for hard concrete distributions

Knowls

  1. Knowl 1 — Hard Concrete Distribution

    definition

    The Hard Concrete distribution is a continuous distribution on the closed interval [0,1][0, 1] that assigns non-zero discrete probability mass to exact zeros and exact ones while maintaining a continuous density on (0,1)(0, 1). It is defined by stretching a binary Concrete (Gumbel-Softmax) random variable s∈(0,1)s \in (0, 1) to an extended interval (γ,ζ)(\gamma, \zeta) with γ<0\gamma < 0 and ζ>1\zeta > 1, and subsequently applying a hard-sigmoid rectification.

    Let ϕ=(log⁡α,β)\phi = (\log \alpha, \beta) denote the distribution parameters, where log⁡α∈R\log \alpha \in \mathbb{R} is the location parameter and β>0\beta > 0 is the temperature. Sampling a hard concrete random variable zz is performed via: u∼U(0,1)u \sim \mathcal{U}(0, 1) s=Sigmoid(log⁡u−log⁡(1−u)+log⁡αβ)s = \text{Sigmoid}\left(\frac{\log u - \log(1 - u) + \log \alpha}{\beta}\right) sˉ=s(ζ−γ)+γ\bar{s} = s(\zeta - \gamma) + \gamma z=min⁡(1,max⁡(0,sˉ))z = \min(1, \max(0, \bar{s}))

    The cumulative distribution function (CDF) of the stretched pre-rectified variable sˉ\bar{s} is: Qsˉ(sˉ∣ϕ)=Sigmoid(β(log⁡(sˉ−γ)−log⁡(ζ−sˉ))−log⁡α)Q_{\bar{s}}(\bar{s}|\phi) = \text{Sigmoid}\left(\beta \left(\log(\bar{s} - \gamma) - \log(\zeta - \bar{s})\right) - \log \alpha\right)

    Applying the hard-sigmoid clamps the probability mass of negative values Qsˉ(0∣ϕ)Q_{\bar{s}}(0|\phi) to a Dirac delta peak at z=0z = 0, and the mass of values exceeding one, 1−Qsˉ(1∣ϕ)1 - Q_{\bar{s}}(1|\phi), to a Dirac delta peak at z=1z = 1. The probability that the gate is non-zero (active) is given in closed form by: q(z≠0∣ϕ)=1−Qsˉ(0∣ϕ)=Sigmoid(log⁡α−βlog⁡−γζ)q(z \neq 0 | \phi) = 1 - Q_{\bar{s}}(0|\phi) = \text{Sigmoid}\left(\log \alpha - \beta \log \frac{-\gamma}{\zeta}\right)

  2. Knowl 2 — Continuous Stochastic Gate Relaxation for Expected L0 Regularization

    model/method

    To minimize the non-differentiable L0L_0 regularized empirical risk: R(θ)=1N∑i=1NL(h(xi;θ),yi)+λ∑j=1∣θ∣I[θj≠0]\mathcal{R}(\theta) = \frac{1}{N} \sum_{i=1}^N \mathcal{L}(h(x_i; \theta), y_i) + \lambda \sum_{j=1}^{|\theta|} \mathbb{I}[\theta_j \neq 0] where θ\theta are the parameters of a model h(⋅;θ)h(\cdot; \theta), L\mathcal{L} is the task loss over NN input-output pairs (xi,yi)(x_i, y_i), and λ>0\lambda > 0 is the regularization strength, the parameters are reparameterized as θj=θ~jzj\theta_j = \tilde{\theta}_j z_j with continuous un-gated weights θ~j∈R\tilde{\theta}_j \in \mathbb{R} and stochastic gates zj∈[0,1]z_j \in [0, 1].

    Each gate is computed as zj=g(sj)=min⁡(1,max⁡(0,sj))z_j = g(s_j) = \min(1, \max(0, s_j)) from a continuous random variable sj∼q(sj∣ϕj)s_j \sim q(s_j | \phi_j) with parameters ϕj\phi_j and CDF Q(sj∣ϕj)Q(s_j | \phi_j). The expected L0L_0 penalty is differentiable with respect to ϕj\phi_j because Eq(zj∣ϕj)[I[zj≠0]]=1−Q(sj≤0∣ϕj)\mathbb{E}_{q(z_j|\phi_j)}[\mathbb{I}[z_j \neq 0]] = 1 - Q(s_j \le 0 | \phi_j).

    Using the reparameterization trick s=f(ϕ,ϵ)s = f(\phi, \epsilon) with parameter-free noise ϵ∼p(ϵ)\epsilon \sim p(\epsilon), the expected regularized objective becomes differentiable: R(θ~,ϕ)=Ep(ϵ)[1N∑i=1NL(h(xi;θ~⊙g(f(ϕ,ϵ))),yi)]+λ∑j=1∣θ∣(1−Q(sj≤0∣ϕj))\mathcal{R}(\tilde{\theta}, \phi) = \mathbb{E}_{p(\epsilon)}\left[ \frac{1}{N} \sum_{i=1}^N \mathcal{L}\left(h(x_i; \tilde{\theta} \odot g(f(\phi, \epsilon))), y_i\right) \right] + \lambda \sum_{j=1}^{|\theta|} (1 - Q(s_j \le 0 | \phi_j)) where ⊙\odot denotes the elementwise product. Using LL Monte Carlo samples ϵ(l)∼p(ϵ)\epsilon^{(l)} \sim p(\epsilon) and z(l)=g(f(ϕ,ϵ(l)))z^{(l)} = g(f(\phi, \epsilon^{(l)})), the objective is approximated as R^(θ~,ϕ)=LE(θ~,ϕ)+λLC(ϕ)\hat{\mathcal{R}}(\tilde{\theta}, \phi) = \mathcal{L}_E(\tilde{\theta}, \phi) + \lambda \mathcal{L}_C(\phi) and optimized end-to-end via stochastic gradient descent.

  3. Knowl 3 — Test-Time Parameter Estimation Under Hard Concrete Gates

    equation

    At test time, the stochasticity of the hard concrete gates is removed to produce deterministic and sparse model weights θ∗\theta^*. The final parameters are obtained by multiplying the learned continuous weights θ~∗\tilde{\theta}^* by the deterministic gate estimate z^\hat{z}: z^j=min⁡(1,max⁡(0,Sigmoid(log⁡αj)(ζ−γ)+γ))\hat{z}_j = \min\left(1, \max\left(0, \text{Sigmoid}(\log \alpha_j)(\zeta - \gamma) + \gamma\right)\right) θj∗=θ~j∗⋅z^j\theta_j^* = \tilde{\theta}_j^* \cdot \hat{z}_j where log⁡αj\log \alpha_j is the optimized location parameter for the jj-th gate, and γ<0\gamma < 0 and ζ>1\zeta > 1 are the fixed stretch hyperparameters (typically set to γ=−0.1\gamma = -0.1 and ζ=1.1\zeta = 1.1). If Sigmoid(log⁡αj)≤−γζ−γ\text{Sigmoid}(\log \alpha_j) \le \frac{-\gamma}{\zeta - \gamma}, the gate z^j\hat{z}_j evaluates to exactly 0, pruning the corresponding parameter θj∗\theta_j^* completely.

  4. Knowl 4 — Combined Expected L0 and L2 Regularization Under Stochastic Gating

    equation

    To combine L0L_0 sparsity regularization with L2L_2 weight decay without inducing extra shrinkage on pruned weights, the standard L2L_2 penalty is formulated as the negative log density of a zero-mean Gaussian prior whose standard deviation σj\sigma_j is governed by the gate zjz_j: σj=1\sigma_j = 1 when zj=0z_j = 0, and σj=zj\sigma_j = z_j when zj>0z_j > 0.

    Under this formulation, the expected L2L_2 penalty for normalized parameters θ^j=θjσj=θ~jzjσj\hat{\theta}_j = \frac{\theta_j}{\sigma_j} = \frac{\tilde{\theta}_j z_j}{\sigma_j} simplifies to: Eq(z∣ϕ)[∥θ^∥22]=∑j=1∣θ∣(Qsˉj(0∣ϕj)⋅0+(1−Qsˉj(0∣ϕj))Eq(zj∣ϕj,sˉj>0)[θ~j2zj2zj2])=∑j=1∣θ∣(1−Qsˉj(0∣ϕj))θ~j2\mathbb{E}_{q(z|\phi)}\left[ \|\hat{\theta}\|_2^2 \right] = \sum_{j=1}^{|\theta|} \left( Q_{\bar{s}_j}(0|\phi_j) \cdot 0 + (1 - Q_{\bar{s}_j}(0|\phi_j)) \mathbb{E}_{q(z_j|\phi_j, \bar{s}_j > 0)}\left[ \frac{\tilde{\theta}_j^2 z_j^2}{z_j^2} \right] \right) = \sum_{j=1}^{|\theta|} (1 - Q_{\bar{s}_j}(0|\phi_j)) \tilde{\theta}_j^2 where θ~j\tilde{\theta}_j is the un-gated parameter value and 1−Qsˉj(0∣ϕj)1 - Q_{\bar{s}_j}(0|\phi_j) is the probability that gate jj is non-zero. Active weights are penalized proportionally to their squared un-gated values, while pruned weights contribute zero to the L2L_2 penalty.

  5. Knowl 5 — Group Sparsity and Structured Pruning Under L0 Regularization

    model/method

    Structured neuron or feature map pruning is achieved under L0L_0 regularization by assigning a single shared hard concrete gate zgz_g to all parameters in group g∈Gg \in G. For fully connected layers, a single gate is assigned per input neuron; for convolutional layers, a single gate is assigned per output feature map and shared across all spatial positions.

    The expected group L0L_0 complexity loss and corresponding expected L2L_2 penalty are given by: Eq(z∣ϕ)[∥θ∥0]=∑g=1∣G∣∣g∣(1−Q(sg≤0∣ϕg))\mathbb{E}_{q(z|\phi)}[\|\theta\|_0] = \sum_{g=1}^{|G|} |g| \left(1 - Q(s_g \le 0 | \phi_g)\right) Eq(z∣ϕ)[∥θ^∥22]=∑g=1∣G∣(1−Q(sg≤0∣ϕg))∑j=1∣g∣θ~g,j2\mathbb{E}_{q(z|\phi)}\left[ \|\hat{\theta}\|_2^2 \right] = \sum_{g=1}^{|G|} \left(1 - Q(s_g \le 0 | \phi_g)\right) \sum_{j=1}^{|g|} \tilde{\theta}_{g,j}^2 where ∣G∣|G| is the total number of groups, ∣g∣|g| is the number of weights in group gg, and ϕg\phi_g are the distribution parameters of the gate for group gg. Because gates are shared spatially across feature maps, inactive channels are skipped entirely during forward and backward passes, enabling computational FLOP savings directly during training.

  6. Knowl 6 — Variational Free Energy Interpretation of Expected L0 Regularization

    theoretical result

    The expected L0L_0 regularized objective corresponds to an upper bound on the variational free energy of a Bayesian neural network under a spike-and-slab prior: p(zj)=Bernoulli(πj),p(θj∣zj=0)=δ(θj),p(θj∣zj=1)=N(θj∣0,1)p(z_j) = \text{Bernoulli}(\pi_j), \quad p(\theta_j | z_j = 0) = \delta(\theta_j), \quad p(\theta_j | z_j = 1) = \mathcal{N}(\theta_j | 0, 1)

    Using a factorized spike-and-slab approximate posterior q(θ,z)=∏jq(zj)q(θj∣zj)q(\theta, z) = \prod_j q(z_j) q(\theta_j | z_j) with q(θj∣zj=0)=δ(θj)q(\theta_j | z_j = 0) = \delta(\theta_j), the exact variational free energy is: F=−Eq(z)q(θ∣z)[log⁡p(D∣θ)]+∑j=1∣θ∣DKL(q(zj)∥p(zj))+∑j=1∣θ∣q(zj=1)DKL(q(θj∣zj=1)∥p(θj∣zj=1))\mathcal{F} = -\mathbb{E}_{q(z)q(\theta|z)}[\log p(\mathcal{D}|\theta)] + \sum_{j=1}^{|\theta|} D_{KL}(q(z_j) \parallel p(z_j)) + \sum_{j=1}^{|\theta|} q(z_j = 1) D_{KL}(q(\theta_j | z_j = 1) \parallel p(\theta_j | z_j = 1))

    Assuming the parameters θ\theta are point-optimized (θ=θ~⊙z\theta = \tilde{\theta} \odot z) and that the coding cost for each active parameter is fixed to a constant DKL(q(θj∣zj=1)∥p(θj∣zj=1))=λD_{KL}(q(\theta_j | z_j = 1) \parallel p(\theta_j | z_j = 1)) = \lambda, the free energy simplifies to: F=−Eq(z)[log⁡p(D∣θ~⊙z)]+∑j=1∣θ∣DKL(q(zj)∥p(zj))+λ∑j=1∣θ∣q(zj=1)\mathcal{F} = -\mathbb{E}_{q(z)}[\log p(\mathcal{D} | \tilde{\theta} \odot z)] + \sum_{j=1}^{|\theta|} D_{KL}(q(z_j) \parallel p(z_j)) + \lambda \sum_{j=1}^{|\theta|} q(z_j = 1)

    By non-negativity of the gate Kullback-Leibler divergence ∑jDKL(q(zj)∥p(zj))≥0\sum_j D_{KL}(q(z_j) \parallel p(z_j)) \ge 0, minimizing the expected L0L_0 regularized empirical loss minimizes a tight upper bound on this variational free energy.

  7. Knowl 7 — Kullback-Leibler Divergence Between Hard Concrete Distributions

    equation

    When optimizing variational objectives where prior uncertainty over the hard concrete gates zz is explicitly retained, the Kullback-Leibler divergence from a hard concrete prior p(z)p(z) (with underlying pre-rectified CDF PsˉP_{\bar{s}}) to a hard concrete variational posterior q(z)q(z) (with underlying pre-rectified CDF QsˉQ_{\bar{s}}) is computed via the chain rule of relative entropy: DKL(q(z)∥p(z))=Qsˉ(0)log⁡Qsˉ(0)Psˉ(0)+(1−Qsˉ(1))log⁡1−Qsˉ(1)1−Psˉ(1)+(Qsˉ(1)−Qsˉ(0))Eqsˉ(z∣sˉ∈(0,1))[log⁡qsˉ(z)−log⁡psˉ(z)]D_{KL}(q(z) \parallel p(z)) = Q_{\bar{s}}(0) \log \frac{Q_{\bar{s}}(0)}{P_{\bar{s}}(0)} + (1 - Q_{\bar{s}}(1)) \log \frac{1 - Q_{\bar{s}}(1)}{1 - P_{\bar{s}}(1)} + \left(Q_{\bar{s}}(1) - Q_{\bar{s}}(0)\right) \mathbb{E}_{q_{\bar{s}}(z | \bar{s} \in (0, 1))}\left[ \log q_{\bar{s}}(z) - \log p_{\bar{s}}(z) \right]

    If the expectation under the truncated distribution over (0,1)(0, 1) is not available analytically, it is approximated with Monte Carlo samples obtained using the inverse CDF transform: u∼U(0,1)u \sim \mathcal{U}(0, 1) z=Qsˉ−1(Qsˉ(γ)+u(Qsˉ(ζ)−Qsˉ(γ)))z = Q_{\bar{s}}^{-1}\left( Q_{\bar{s}}(\gamma) + u \left( Q_{\bar{s}}(\zeta) - Q_{\bar{s}}(\gamma) \right) \right) where Qsˉ−1(⋅)Q_{\bar{s}}^{-1}(\cdot) is the quantile function of the pre-rectified random variable sˉ\bar{s}.

  8. Knowl 8 — Structured Sparsification and Classification Performance on MNIST

    data/table

    Structured neuron and feature map pruning using hard concrete L0L_0 regularization (L0hcL_{0hc}) was evaluated on MNIST using a multilayer perceptron (MLP 784-300-100) and LeNet-5-Caffe (20-50-800-500). Models were optimized with Adam for 200 epochs using concrete temperature β=2/3\beta = 2/3, stretch boundaries γ=−0.1\gamma = -0.1, ζ=1.1\zeta = 1.1, and one gate sample per minibatch. In the table below, pruned architectures indicate the number of retained neurons/feature maps per layer, and test error is measured after 200 epochs (NN is the number of training datapoints).

    Network Size Method Pruned Architecture Error (%)
    MLP Sparse VD 512-114-72 1.8
    784-300-100 BC-GNJ 278-98-13 1.8
    BC-GHS 311-86-14 1.8
    L0hcL_{0hc}, λ=0.1/N\lambda = 0.1/N 219-214-100 1.4
    L0hcL_{0hc}, λ\lambda sep. 266-88-33 1.8
    LeNet-5-Caffe Sparse VD 14-19-242-131 1.0
    20-50-800-500 GL 3-12-192-500 1.0
    GD 7-13-208-16 1.1
    SBP 3-18-284-283 0.9
    BC-GNJ 8-13-88-13 1.0
    BC-GHS 5-10-76-16 1.0
    L0hcL_{0hc}, λ=0.1/N\lambda = 0.1/N 20-25-45-462 0.9
    L0hcL_{0hc}, λ\lambda sep. 9-18-65-25 1.0

    A uniform penalty λ=0.1/N\lambda = 0.1/N preferentially prunes layers with the largest parameter footprints (reducing MLP inputs to 219 and LeNet-5 dense layer to 45) while improving accuracy. Specifying separate layer penalties (λ\lambda sep., with λ=10/N\lambda = 10/N and 0.5/N0.5/N for convolutional layers) forces heavy sparsification in computation-heavy convolutional layers (e.g. 9-18-65-25 on LeNet-5), achieving drastic reductions in training FLOPs.

  9. Knowl 9 — Wide Residual Network Sparsification and Performance on CIFAR

    data/table

    Hard concrete L0L_0 regularization (L0hcL_{0hc}) was evaluated on Wide Residual Networks (WRN-28-10) on CIFAR-10 and CIFAR-100. Hard concrete gates were applied to the hidden convolutional layers of residual blocks. Training used a minibatch size of 128 split across two GPUs (one gate sample per GPU), β=2/3\beta = 2/3, γ=−0.1\gamma = -0.1, ζ=1.1\zeta = 1.1, and weight decay scaled by 1/0.71 / 0.7. Median test error over 5 runs after 200 epochs is reported (NN is the training set size).

    Network CIFAR-10 Error (%) CIFAR-100 Error (%)
    original-ResNet-110 6.43 25.16
    pre-act-ResNet-110 6.37 –
    WRN-28-10 4.00 21.18
    WRN-28-10-dropout 3.89 18.85
    WRN-28-10-L0hcL_{0hc}, λ=0.001/N\lambda = 0.001/N 3.83 18.75
    WRN-28-10-L0hcL_{0hc}, λ=0.002/N\lambda = 0.002/N 3.93 19.04

    With λ=0.001/N\lambda = 0.001/N, L0hcL_{0hc} achieves lower test error than standard dropout on both CIFAR-10 (3.83% vs 3.89%) and CIFAR-100 (18.75% vs 18.85%). Concurrently, training FLOPs decrease steadily during training as inactive gates reach exact zero, without degrading convergence speed relative to dropout.

Coverage note — All primary methodological, theoretical, and empirical contributions of the paper—including the continuous gate relaxation, the hard concrete distribution, combined L0/L2 regularization, group sparsity, the variational connection, and MNIST/CIFAR benchmark experiments—have been fully extracted. General background discussions on previous discrete relaxations (REINFORCE, Straight-Through) were deliberately omitted.

References

  1. 1.Hirotogu Akaike. Information theory and an extension of the maximum likelihood principle. In Selected Papers of Hirotugu Akaike, pp. 199–213. Springer, 1998.
  2. 2.Matthew James Beal. Variational algorithms for approximate Bayesian inference. 2003.
  3. 3.Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup. Conditional computation in neural networks for faster models. arXiv preprint arXiv:1511.06297, 2015.
  4. 4.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  5. 5.Thomas M Cover and Joy A Thomas. Elements of information theory. John Wiley & Sons, 2012.
  6. 6.Yarin Gal, Jiri Hron, and Alex Kendall. Concrete dropout. arXiv preprint arXiv:1705.07832, 2017.
  7. 7.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015.
  8. 8.Markus Harva and Ata Kabán. Variational learning for rectified factor analysis. Signal Processing, 87(3):509–527, 2007.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016a.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In European Conference on Computer Vision, pp. 630–645. Springer, 2016b.
  11. 11.John R Hershey and Peder A Olsen. Approximating the kullback leibler divergence between gaussian mixture models. In Acoustics, Speech and Signal Processing, 2007. ICASSP 2007. IEEE International Conference on, volume 4, pp. IV–317. IEEE, 2007.
  12. 12.Geoffrey E Hinton and Zoubin Ghahramani. Generative models for discovering sparse distributed representations. Philosophical Transactions of the Royal Society of London B: Biological Sciences, 352(1358):1177–1190, 1997.
  13. 13.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016.
  14. 14.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  15. 15.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. International Conference on Learning Representations (ICLR), 2014.
  16. 16.Diederik P Kingma, Tim Salimans, and Max Welling. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems, pp. 2575–2583, 2015.
  17. 17.Yann LeCun, John S Denker, and Sara A Solla. Optimal brain damage. In Advances in neural information processing systems 2, NIPS 1989, volume 2, pp. 598–605. Morgan-Kaufmann Publishers, 1990.
  18. 18.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  19. 19.Christos Louizos, Karen Ullrich, and Max Welling. Bayesian compression for deep learning. arXiv preprint arXiv:1705.08665, 2017.
  20. 20.Chris J Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. arXiv preprint arXiv:1611.00712, 2016.
  21. 21.Toby J Mitchell and John J Beauchamp. Bayesian variable selection in linear regression. Journal of the American Statistical Association, 83(404):1023–1032, 1988.
  22. 22.Andriy Mnih and Karol Gregor. Neural variational inference and learning in belief networks. arXiv preprint arXiv:1402.0030, 2014.
  23. 23.Andriy Mnih and Danilo Rezende. Variational inference for monte carlo objectives. In International Conference on Machine Learning, pp. 2188–2196, 2016.
  24. 24.Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. Variational dropout sparsifies deep neural networks. arXiv preprint arXiv:1701.05369, 2017.
  25. 25.Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. Structured bayesian pruning via log-normal multiplicative noise. arXiv preprint arXiv:1705.07283, 2017.
  26. 26.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the 31th International Conference on Machine Learning, ICML 2014, Beijing, China, 21-26 June 2014, pp. 1278–1286, 2014.
  27. 27.Jason Tyler Rolfe. Discrete variational autoencoders. arXiv preprint arXiv:1609.02200, 2016.
  28. 28.Tim Salimans. A structured variational auto-encoder for learning deep hierarchies of sparse features. arXiv preprint arXiv:1602.08734, 2016.
  29. 29.Gideon Schwarz et al. Estimating the dimension of a model. The annals of statistics, 6(2):461–464, 1978.
  30. 30.Suraj Srinivas and R Venkatesh Babu. Generalized dropout. arXiv preprint arXiv:1611.06791, 2016.
  31. 31.Suraj Srinivas, Akshayvarun Subramanya, and R Venkatesh Babu. Training sparse neural networks. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2017 IEEE Conference on, pp. 455–462. IEEE, 2017.
  32. 32.Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of machine learning research, 15(1):1929–1958, 2014.
  33. 33.Robert Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological), pp. 267–288, 1996.
  34. 34.Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, and Christoph Bregler. Efficient object localization using convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 648–656, 2015.
  35. 35.George Tucker, Andriy Mnih, Chris J Maddison, and Jascha Sohl-Dickstein. Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models. arXiv preprint arXiv:1703.07370, 2017.
  36. 36.Karen Ullrich, Edward Meeds, and Max Welling. Soft weight-sharing for neural network compression. ICLR, 2017.
  37. 37.Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Learning structured sparsity in deep neural networks. In Advances in Neural Information Processing Systems, pp. 2074–2082, 2016.
  38. 38.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3-4):229–256, 1992.
  39. 39.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  40. 40.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. arXiv preprint arXiv:1611.03530, 2016.

Citation

MLA
Louizos, C., et al. “Learning Sparse Neural Networks Through $L_0$ Regularization”. arXiv, 2017, http://arxiv.org/abs/1712.01312v2.
APA
Louizos, C., Welling, M., & Kingma, D. P. (2017). Learning Sparse Neural Networks through $L_0$ Regularization. arXiv. http://arxiv.org/abs/1712.01312v2
Chicago
Louizos, C., M. Welling, and D. P. Kingma. 2017. “Learning Sparse Neural Networks Through $L_0$ Regularization”. arXiv. http://arxiv.org/abs/1712.01312v2.
Harvard
Louizos, C., Welling, M. and Kingma, D.P. (2017) “Learning Sparse Neural Networks through $L_0$ Regularization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1712.01312v2.
Vancouver
1. Louizos C, Welling M, Kingma DP (2017) Learning Sparse Neural Networks through $L_0$ Regularization. arXiv

BibTeX

@article{louizos2017learning,
  title = {Learning Sparse Neural Networks through $L_0$ Regularization},
  author = {Louizos, Christos and Welling, Max and Kingma, Diederik P.},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1712.01312v2},
  eprint = {1712.01312}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission