SNIP: Single-shot Network Pruning based on Connection Sensitivity

Namhoon LeeThalaiyasingam AjanthanPhilip H. S. Torr

article2018ICLR1,530 citations

Proposes a connection sensitivity metric to prune neural networks in a single step at initialization, eliminating costly pretraining and iterative pruning schedules while preserving accuracy across diverse architectures.

Listen

Modern deep learning models achieve impressive accuracy but are heavily overparameterized, demanding significant computational power and memory. Compressing these models through network pruning makes them viable for resource-constrained environments, yet standard techniques require expensive cycles of pretraining, pruning, and fine-tuning governed by complex schedules and tuning parameters.

The article introduces and evaluates Single-Shot Network Pruning (SNIP), a method that prunes redundant connections in a neural network in a single step at random initialization prior to training. The objective is to demonstrate that structurally vital connections can be identified and isolated before any standard learning occurs, removing the computational overhead of iterative retraining.

The approach calculates a connection sensitivity score for each parameter using a single forward-backward pass over a small batch of training data at variance scaling initialization. This score quantifies how much a connection influences the loss function. Connections with the lowest sensitivity are removed immediately, and the remaining sparse network is trained using standard optimization techniques. The researchers evaluated this method across several benchmark image classification datasets—MNIST, CIFAR-10, and Tiny-ImageNet—using fully connected, convolutional, residual, and recurrent network architectures.

The findings show that SNIP removes 90% to 99% of parameters across architectures while maintaining classification error rates within roughly 1% of dense baseline models. On LeNet architectures, removing up to 99% of parameters caused only negligible performance loss (0.3% to 0.7%), matching or exceeding the accuracy of methods requiring complex iterative retraining. Across modern convolutional, residual, and recurrent networks, pruning 90% to 95% of connections yielded comparable accuracy. Layer visualizations confirmed that retained connections directly align with discriminative input features, and models pruned with SNIP were unable to fit randomized labels, proving that the algorithm preserves task-relevant structural capacity rather than arbitrary connections.

These results demonstrate that extensive pretraining is unnecessary for effective network compression. Implementing single-shot pruning upfront significantly reduces the computational expense, training duration, and engineering complexity associated with model compression pipelines. Furthermore, because it functions as an architecture-agnostic preprocessor, it minimizes the risk and implementation friction of deploying compact deep learning models to embedded or edge devices.

Engineering and research teams should consider adopting single-shot pruning as a lightweight preprocessing step before training deep neural networks. Practitioners should ensure proper variance scaling initialization, as experiments indicate it is essential for model stability, particularly in recurrent networks. Before large-scale deployment, teams should run pilot evaluations on their specific datasets and target architectures to confirm appropriate sparsity thresholds.

The conclusions are supported with high confidence across standard computer vision benchmarks and common network architectures. However, decision-makers should note that evaluations were limited to classification tasks on benchmark datasets rather than production-scale workloads such as full ImageNet, large language models, or real-time control systems. Performance tradeoffs on highly customized architectures or non-classification tasks should be validated independently.

Cover for SNIP: Single-shot Network Pruning based on Connection Sensitivity

Abstract

Pruning large neural networks while maintaining their performance is often desirable due to the reduced space and time complexity. In existing methods, pruning is done within an iterative optimization procedure with either heuristically designed pruning schedules or additional hyperparameters, undermining their utility. In this work, we present a new approach that prunes a given network once at initialization prior to training. To achieve this, we introduce a saliency criterion based on connection sensitivity that identifies structurally important connections in the network for the given task. This eliminates the need for both pretraining and the complex pruning schedule while making it robust to architecture variations. After pruning, the sparse network is trained in the standard way. Our method obtains extremely sparse networks with virtually the same accuracy as the reference network on the MNIST, CIFAR-10, and Tiny-ImageNet classification tasks and is broadly applicable to various architectures including convolutional, residual and recurrent networks. Unlike existing methods, our approach enables us to demonstrate that the retained connections are indeed relevant to the given task.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Neural network pruning
  • 4 Single-shot network pruning based on connection sensitivity
  • 4.1 Connection sensitivity: architectural perspective
  • 4.2 Single-shot pruning at initialization
  • 5 Experiments
  • 5.1 Pruning LeNets with varying levels of sparsity
  • 5.2 Comparisons to existing approaches
  • 5.3 Various modern architectures
  • 5.4 Understanding which connections are being pruned
  • 5.5 Effects of data and weight initialization
  • 5.6 Fitting random labels
  • 6 Discussion and future work
  • References
  • A Visualizing pruned parameters on (inverted) (fashion-)mnist
  • B Fitting random labels: varying sparsity levels
  • C Tiny-imagenet
  • D Architecture details

Knowls

  1. Knowl 1 — Masked optimization formulation for single-shot pruning

    model/method

    SNIP separates whether a connection exists from the numerical value of its weight. For a training set D={(xi,yi)}i=1nD=\{(x_i,y_i)\}_{i=1}^{n}, a neural network with mm parameters, weights w∈Rmw\in\mathbb{R}^{m}, and target number of retained parameters κ\kappa, let c∈{0,1}mc\in\{0,1\}^{m} be a binary connectivity mask and let ⊙\odot denote elementwise multiplication. Pruning is formulated as

    min⁡c,wL(c⊙w;D)=1n∑i=1nℓ(c⊙w;(xi,yi))subject to∥c∥0≤κ,\min_{c,w} L(c\odot w;D)=\frac{1}{n}\sum_{i=1}^{n}\ell(c\odot w;(x_i,y_i)) \quad\text{subject to}\quad \|c\|_0\leq\kappa,

    where ℓ\ell is the per-example training loss, LL is the empirical loss, and ∥c∥0\|c\|_0 counts active connections. The mask cc is selected before training; the surviving weights are then learned by ordinary optimization with the mask fixed.

  2. Knowl 2 — Connection sensitivity as a data-dependent saliency criterion

    equation

    For a network with weights w∈Rmw\in\mathbb{R}^{m}, binary mask c∈{0,1}mc\in\{0,1\}^{m}, loss L(c⊙w;D)L(c\odot w;D), and dataset DD, the exact loss change from removing connection jj is

    ΔLj(w;D)=L(1⊙w;D)−L((1−ej)⊙w;D),\Delta L_j(w;D)=L(\mathbf{1}\odot w;D)-L((\mathbf{1}-e_j)\odot w;D),

    where 1∈Rm\mathbf{1}\in\mathbb{R}^{m} is the all-ones vector and ej∈Rme_j\in\mathbb{R}^{m} is one at index jj and zero elsewhere. SNIP approximates this discrete removal influence by the derivative with respect to a multiplicative connectivity variable:

    gj(w;D)=∂L(c⊙w;D)∂cj∣c=1,ΔLj(w;D)≈gj(w;D).g_j(w;D)=\left.\frac{\partial L(c\odot w;D)}{\partial c_j}\right|_{c=\mathbf{1}}, \qquad \Delta L_j(w;D)\approx g_j(w;D).

    The connection-sensitivity score is the normalized absolute influence

    sj=∣gj(w;D)∣∑k=1m∣gk(w;D)∣,j∈{1,…,m}.s_j=\frac{|g_j(w;D)|}{\sum_{k=1}^{m}|g_k(w;D)|}, \qquad j\in\{1,\ldots,m\}.

    The mask retains the κ\kappa connections with the largest sjs_j values, breaking ties arbitrarily. Taking the absolute value preserves connections with large influence regardless of whether an infinitesimal reduction in connectivity would increase or decrease the loss. Unlike ∂L/∂wj\partial L/\partial w_j, gjg_j measures sensitivity to a multiplicative removal of a connection rather than to an additive change in its weight.

  3. Knowl 3 — SNIP single-shot pruning procedure

    algorithm

    SNIP takes a differentiable loss function LL, training dataset DD, desired retained-connection count κ\kappa, and a network with mm parameters. It returns a fixed binary mask and a trained sparse network. The saliency computation uses one mini-batch and one automatic-differentiation pass.

    Input: Loss function LL, training dataset DD, retained-connection count κ\kappa
    Output: Binary mask cc and trained sparse weights w∗w^*
    Initialize weights ww with variance scaling.
    Sample one pruning mini-batch DbD_b from DD.
    Set all connectivity variables to one: c←1c \leftarrow \mathbf{1}.
    Compute gj←∂L(c⊙w;Db)/∂cjg_j \leftarrow \partial L(c \odot w;D_b)/\partial c_j for every connection jj in one forward-backward pass.
    Compute sj←∣gj∣/∑k∣gk∣s_j \leftarrow |g_j|/\sum_k |g_k| for every connection jj.
    Find the κ\kappa-th largest score s~κ\tilde{s}_{\kappa}.
    Set cj←1c_j \leftarrow 1 if sj≥s~κs_j \geq \tilde{s}_{\kappa} and zero otherwise; break ties to retain exactly κ\kappa connections.
    Train ww by ordinary network training on the full training data while applying the fixed mask cc.
    Return w∗←c⊙ww^* \leftarrow c \odot w.

    The reported MNIST and CIFAR-10 training used SGD with momentum 0.90.9, batch sizes 100100 and 128128, weight decay 0.00050.0005, initial learning rate 0.10.1, and learning-rate reductions by 0.10.1 every 25,00025{,}000 or 30,00030{,}000 iterations, respectively. These are ordinary training settings rather than pruning schedules; SNIP itself adds no pretraining, iterative prune–retrain cycles, or pruning-specific hyperparameters.

  4. Knowl 4 — Variance-scaled initialization makes the criterion architecture-robust

    assumption

    SNIP computes sensitivity before learning, so the initialization must produce informative, non-saturated activations and gradients. The paper therefore recommends variance-scaling initialization, which keeps signal variance approximately stable across layers. A fixed-variance initialization can make the scores depend strongly on layer width and depth, while overly large weights can saturate nonlinearities and yield uninformative derivatives.

    The scores are task-dependent because they use the loss function and examples from the task. A practitioner may compute them from the full training set, one mini-batch, several accumulated batches, or an exponential moving average. The paper's experiments show that a single reasonable training mini-batch is often sufficient, enabling pruning before any pretraining. This design is intended to make the same pruning procedure applicable without architectural modifications to fully connected, convolutional, residual, and recurrent networks.

  5. Knowl 5 — Extreme sparsity on MNIST LeNets

    empirical result

    SNIP was evaluated on MNIST using LeNet-300-100, containing approximately 267267k parameters, and LeNet-5-Caffe, containing approximately 431431k parameters. Each sparsity setting was repeated 2020 times with different dataset and initialization seeds. Across sparsity levels, LeNet-300-100 had test performance close to its dense reference through 90%90\% sparsity, while the degradation for LeNet-5-Caffe was nearly invisible.

    At more extreme sparsity, the dense reference errors were 1.7%1.7\% for LeNet-300-100 and 0.9%0.9\% for LeNet-5-Caffe. SNIP retained only 2%2\% of LeNet-300-100's parameters (98%98\% sparsity) with error 2.4%2.4\%, and only 1%1\% of LeNet-5-Caffe's parameters (99%99\% sparsity) with error 1.1%1.1\%. At less extreme settings, 95%95\% sparsity on LeNet-300-100 produced 1.6%1.6\% error and 98%98\% sparsity on LeNet-5-Caffe produced 0.8%0.8\% error, both numerically below the corresponding dense-reference errors. These results show that connections selected before training can provide nearly all of the task-relevant capacity of the dense networks.

  6. Knowl 6 — Comparison with iterative and Bayesian pruning methods

    data/table

    On MNIST LeNets, SNIP was compared with magnitude-based LWC, DNS, and LC; Bayesian SWS and SVD; and Hessian-based OBD and L-OBS. The dense reference errors were 1.7%1.7\% for LeNet-300-100 and 0.9%0.9\% for LeNet-5-Caffe. The comparison reported the following representative operating points:

    • LWC: 91.7%91.7\% sparsity and 1.6%1.6\% error on LeNet-300-100; 91.7%91.7\% and 0.8%0.8\% on LeNet-5-Caffe.
    • DNS: 98.2%98.2\% and 2.0%2.0\%; 99.1%99.1\% and 0.9%0.9\%.
    • LC: 99.0%99.0\% and 3.2%3.2\%; 99.0%99.0\% and 1.1%1.1\%.
    • SWS: 95.6%95.6\% and 1.9%1.9\%; 99.5%99.5\% and 1.0%1.0\%.
    • SVD: 98.5%98.5\% and 1.9%1.9\%; 99.6%99.6\% and 0.8%0.8\%.
    • OBD: 92.0%92.0\% and 2.0%2.0\%; 92.0%92.0\% and 2.7%2.7\%.
    • L-OBS: 98.5%98.5\% and 2.0%2.0\%; 99.0%99.0\% and 2.1%2.1\%.
    • SNIP: 95.0%95.0\% sparsity and 1.6%1.6\% error, or 98.0%98.0\% and 2.4%2.4\%, on LeNet-300-100; and 98.0%98.0\% and 0.8%0.8\%, or 99.0%99.0\% and 1.1%1.1\%, on LeNet-5-Caffe.

    Unlike the comparison methods, SNIP used no pretraining, one pruning operation, no additional pruning hyperparameters, no augmented training objective, and no architecture-dependent pruning constraints. The comparison therefore demonstrates not only competitive accuracy at extreme sparsity but also the substantially simpler training procedure claimed by the paper.

  7. Knowl 7 — Applicability to convolutional, residual, and recurrent networks

    data/table

    The paper applied the unchanged SNIP procedure to CIFAR-10 convolutional and residual models and to sequential-MNIST recurrent models. Each entry below gives sparsity, parameter count before and after pruning, and classification error before and after pruning; Δ\Delta is the after-minus-before error in percentage points.

    • AlexNet-s: 90.0%90.0\% sparsity; 5.15.1m →\rightarrow 507507k parameters; error 14.12%→14.99%14.12\%\rightarrow14.99\%; Δ=+0.87\Delta=+0.87.
    • AlexNet-b: 90.0%90.0\%; 8.58.5m →\rightarrow 849849k; 13.92%→14.50%13.92\%\rightarrow14.50\%; Δ=+0.58\Delta=+0.58.
    • VGG-C: 95.0%95.0\%; 10.510.5m →\rightarrow 526526k; 6.82%→7.27%6.82\%\rightarrow7.27\%; Δ=+0.45\Delta=+0.45.
    • VGG-D: 95.0%95.0\%; 15.215.2m →\rightarrow 762762k; 6.76%→7.09%6.76\%\rightarrow7.09\%; Δ=+0.33\Delta=+0.33.
    • VGG-like: 97.0%97.0\%; 15.015.0m →\rightarrow 449449k; 8.26%→8.00%8.26\%\rightarrow8.00\%; Δ=−0.26\Delta=-0.26.
    • WRN-16-8: 95.0%95.0\%; 10.010.0m →\rightarrow 548548k; 6.21%→6.63%6.21\%\rightarrow6.63\%; Δ=+0.42\Delta=+0.42.
    • WRN-16-10: 95.0%95.0\%; 17.117.1m →\rightarrow 856856k; 5.91%→6.43%5.91\%\rightarrow6.43\%; Δ=+0.52\Delta=+0.52.
    • WRN-22-8: 95.0%95.0\%; 17.217.2m →\rightarrow 858858k; 6.14%→5.85%6.14\%\rightarrow5.85\%; Δ=−0.29\Delta=-0.29.
    • LSTM-s: 95.0%95.0\%; 137137k →\rightarrow 6.86.8k; 1.88%→1.57%1.88\%\rightarrow1.57\%; Δ=−0.31\Delta=-0.31.
    • LSTM-b: 95.0%95.0\%; 535535k →\rightarrow 26.826.8k; 1.15%→1.35%1.15\%\rightarrow1.35\%; Δ=+0.20\Delta=+0.20.
    • GRU-s: 95.0%95.0\%; 104104k →\rightarrow 5.25.2k; 1.87%→2.41%1.87\%\rightarrow2.41\%; Δ=+0.54\Delta=+0.54.
    • GRU-b: 95.0%95.0\%; 404404k →\rightarrow 20.220.2k; 1.71%→1.52%1.71\%\rightarrow1.52\%; Δ=−0.19\Delta=-0.19.

    The convolutional models were evaluated on CIFAR-10 and the recurrent models on sequential MNIST, where each recurrent cell processed one image row at a time. Every model retained at least 90%90\% sparsity, and every reported error change was below 11 percentage point, supporting the claim that the score is not tied to a particular architecture type.

  8. Knowl 8 — Performance on the more difficult Tiny-ImageNet task

    empirical result

    SNIP was also tested on Tiny-ImageNet, which has 200200 classes, 500500 training images and 5050 validation images per class, with 64×6464\times64 input resolution. The reported parameter counts and errors were:

    • AlexNet-s at 90.0%90.0\% sparsity: 5.15.1m →\rightarrow 507507k parameters and error 62.52%→65.27%62.52\%\rightarrow65.27\% (Δ=+2.75\Delta=+2.75 percentage points).
    • AlexNet-b at 90.0%90.0\%: 8.58.5m →\rightarrow 849849k and 62.76%→65.54%62.76\%\rightarrow65.54\% (Δ=+2.78\Delta=+2.78).
    • VGG-C at 95.0%95.0\%: 10.510.5m →\rightarrow 526526k and 56.49%→57.48%56.49\%\rightarrow57.48\% (Δ=+0.99\Delta=+0.99).
    • VGG-D at 95.0%95.0\%: 15.215.2m →\rightarrow 762762k and 56.85%→57.00%56.85\%\rightarrow57.00\% (Δ=+0.15\Delta=+0.15).
    • VGG-like at 95.0%95.0\%: 15.015.0m →\rightarrow 749749k and 54.86%→55.73%54.86\%\rightarrow55.73\% (Δ=+0.87\Delta=+0.87).

    Thus, even on this more complex classification problem, the method removed 9090–95%95\% of parameters with small error increases for the VGG models. The paper notes that the larger AlexNet degradation may be related to its first convolution using stride 44 on the higher-resolution images, which can lose information when connections are pruned.

  9. Knowl 9 — Retained connections expose task-relevant image structure

    empirical result

    To inspect what SNIP preserves, the paper used the first fully connected layer of LeNet-300-100, whose weights connect 784784 input pixels to 300300 hidden units. For a class-specific mini-batch, the binary mask was averaged across the 300300 input-to-hidden columns and reshaped from 784784 entries into a 28×2828\times28 image. The visualizations on page 9 show that, as sparsity increases from 10%10\% to 90%90\%, retained connections concentrate on discriminative regions: digit foregrounds for MNIST and clothing silhouettes for Fashion-MNIST, while much of the background is removed.

    This behavior is class-dependent: a batch of digit-0 examples preserves connections covering the foreground of the zero, and the Fashion-MNIST masks reveal the identity and shape of the clothing class. The appendix visualization on page 13 shows that inverting image intensities produces essentially the same SNIP masks, whereas directly pruning by the weight gradient ∂L/∂w\partial L/\partial w gives different and inconsistent patterns across sparsity levels and between original and inverted data. These observations provide the paper's interpretability evidence that SNIP selects connections relevant to the task rather than merely pruning blindly.

  10. Knowl 10 — Effects of pruning-batch size and weight initialization

    empirical result

    The paper measured how the pruning data and initialization affect SNIP on LeNet-300-100 at 90%90\% sparsity. For pruning batches of sizes ∣Db∣=1,10,100,1000,10000|D_b|=1,10,100,1000,10000, the resulting classification errors were respectively 1.94%1.94\%, 1.72%1.72\%, 1.64%1.64\%, 1.56%1.56\%, and 1.40%1.40\%. The ∣Db∣=1|D_b|=1 mask closely reflected the individual sampled example, while larger batches produced masks increasingly similar to the average pattern over the training set and lower errors. The training batch size remained 100100 in this analysis.

    Initialization was evaluated over 2020 runs at 90%90\% sparsity using random normal (RN), truncated random normal (TN), Glorot variance scaling (VS-X), and He variance scaling (VS-H). The reported mean classification errors with standard deviations were:

    • LeNet-300-100: RN 1.90±0.09%1.90\pm0.09\%, TN 1.96±0.11%1.96\pm0.11\%, VS-X 1.91±0.10%1.91\pm0.10\%, VS-H 1.88±0.10%1.88\pm0.10\%.
    • LeNet-5-Caffe: RN 0.89±0.04%0.89\pm0.04\%, TN 0.87±0.05%0.87\pm0.05\%, VS-X 0.88±0.07%0.88\pm0.07\%, VS-H 0.85±0.05%0.85\pm0.05\%.
    • LSTM-s: RN 2.93±0.20%2.93\pm0.20\%, TN 3.03±0.17%3.03\pm0.17\%, VS-X 1.48±0.09%1.48\pm0.09\%, VS-H 1.47±0.08%1.47\pm0.08\%.
    • GRU-s: RN 47.61±20.49%47.61\pm20.49\%, TN 46.48±22.25%46.48\pm22.25\%, VS-X 1.80±0.10%1.80\pm0.10\%, VS-H 1.80±0.14%1.80\pm0.14\%.

    Variance scaling made little difference for the LeNets but was essential for the recurrent models, especially the GRU, where non-variance-scaled initializations often produced unusable sparse networks.

  11. Knowl 11 — SNIP-selected networks resist memorizing random labels

    empirical result

    The paper tested whether SNIP's sparse subnetworks merely retain enough capacity to memorize arbitrary labels. LeNet-5-Caffe was trained on MNIST with true labels and with randomly shuffled labels; the sparse model used 99%99\% sparsity, and its SNIP mask was always computed using the true labels. With true labels, both the dense reference and the SNIP-pruned model rapidly reached nearly zero training loss. With random labels, the dense reference also reached very low training loss, even when an explicit L2L_2 regularizer was added, whereas the SNIP-pruned model failed to fit the random labels and retained high training error.

    The appendix's varying-sparsity experiment showed that less sparse SNIP models, which retain more parameters, achieve lower training loss on random labels, while all tested sparse models still learn the genuine MNIST classification task without substantial accuracy loss. The paper interprets this contrast as evidence that SNIP preserves task-relevant capacity while removing capacity useful for memorizing arbitrary labels, although it also notes that insufficient capacity is a possible explanation.

Coverage note — Low-level architecture specifications, the full related-work discussion, and supplementary visualization details were omitted because they support the experiments without adding separate load-bearing method or result claims.

References

  1. 1.Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. Stronger generalization bounds for deep nets via a compression approach. ICML, 2018.
  2. 2.Leo Breiman. Better subset regression using the nonnegative garrote. Technometrics, 1995.
  3. 3.Miguel A. Carreira-Perpiñan and Yerlan Idelbayev. “Learning-compression” algorithms for neural net pruning. CVPR, 2018.
  4. 4.Yves Chauvin. A back-propagation algorithm with optimal use of hidden units. NIPS, 1989.
  5. 5.Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. EMNLP, 2014.
  6. 6.Xin Dong, Shangyu Chen, and Sinno Pan. Learning to prune deep neural networks via layer-wise optimal brain surgeon. NIPS, 2017.
  7. 7.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. AISTATS, 2010.
  8. 8.Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. Compressing deep convolutional networks using vector quantization. arXiv preprint arXiv:1412.6115, 2014.
  9. 9.Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. Deep learning. MIT press Cambridge, 2016.
  10. 10.Yiwen Guo, Anbang Yao, and Yurong Chen. Dynamic network surgery for efficient dnns. NIPS, 2016.
  11. 11.Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. Deep learning with limited numerical precision. ICML, 2015.
  12. 12.Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network. NIPS, 2015.
  13. 13.Babak Hassibi, David G Stork, and Gregory J Wolff. Optimal brain surgeon and general network pruning. Neural Networks, 1993.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. ICCV, 2015.
  15. 15.Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks. NIPS, 2016.
  16. 16.Masumi Ishikawa. Structural learning with forgetting. Neural Networks, 1996.
  17. 17.Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman. Speeding up convolutional neural networks with low rank expansions. BMVC, 2014.
  18. 18.Ehud D Karnin. A simple procedure for pruning back-propagation trained neural networks. Neural Networks, 1990.
  19. 19.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. ICLR, 2015.
  20. 20.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. ICML, 2017.
  21. 21.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. 2012.
  22. 22.Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton. A simple way to initialize recurrent networks of rectified linear units. CoRR, 2015.
  23. 23.Yann LeCun, John S Denker, and Sara A Solla. Optimal brain damage. NIPS, 1990.
  24. 24.Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller. Efficient backprop. Neural Networks: Tricks of the Trade, 1998.
  25. 25.Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. Pruning filters for efficient convnets. ICLR, 2017.
  26. 26.Christos Louizos, Max Welling, and Diederik P Kingma. Learning sparse neural networks through l0 regularization. ICLR, 2018.
  27. 27.Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature Communications, 2018.
  28. 28.Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. Variational dropout sparsifies deep neural networks. ICML, 2017a.
  29. 29.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. ICLR, 2017b.
  30. 30.Michael C Mozer and Paul Smolensky. Skeletonization: A technique for trimming the fat from a network via relevance assessment. NIPS, 1989.
  31. 31.Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta. Exploring sparsity in recurrent neural networks. ICLR, 2017.
  32. 32.Alexander Novikov, Dmitrii Podoprikhin, Anton Osokin, and Dmitry P Vetrov. Tensorizing neural networks. NIPS, 2015.
  33. 33.Steven J Nowlan and Geoffrey E Hinton. Simplifying neural networks by soft weight-sharing. Neural Computation, 1992.
  34. 34.Ameya Prabhu, Girish Varma, and Anoop Namboodiri. Deep expander networks: Efficient deep networks from graph theory. ECCV, 2018.
  35. 35.Russell Reed. Pruning algorithms-a survey. Neural Networks, 1993.
  36. 36.Abigail See, Minh-Thang Luong, and Christopher D Manning. Compression of neural machine translation models via pruning. CoNLL, 2016.
  37. 37.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. ICLR, 2015.
  38. 38.Karen Ullrich, Edward Meeds, and Max Welling. Soft weight-sharing for neural network compression. ICLR, 2017.
  39. 39.Andreas S Weigend, David E Rumelhart, and Bernardo A Huberman. Generalization by weight-elimination with application to forecasting. NIPS, 1991.
  40. 40.Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Learning structured sparsity in deep neural networks. NIPS, 2016.
  41. 41.Sergey Zagoruyko. 92.45% on cifar-10 in torch. Torch Blog, 2015.
  42. 42.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. BMVC, 2016.
  43. 43.Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. Recurrent neural network regularization. arXiv preprint arXiv:1409.2329, 2014.
  44. 44.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. ICLR, 2017.

Citation

MLA
Lee, N., et al. “SNIP: Single-shot Network Pruning Based on Connection Sensitivity”. arXiv, 2018, http://arxiv.org/abs/1810.02340v2.
APA
Lee, N., Ajanthan, T., & Torr, P. H. S. (2018). SNIP: Single-shot Network Pruning based on Connection Sensitivity. arXiv. http://arxiv.org/abs/1810.02340v2
Chicago
Lee, N., T. Ajanthan, and P. H. S. Torr. 2018. “SNIP: Single-shot Network Pruning Based on Connection Sensitivity”. arXiv. http://arxiv.org/abs/1810.02340v2.
Harvard
Lee, N., Ajanthan, T. and Torr, P.H.S. (2018) “SNIP: Single-shot Network Pruning based on Connection Sensitivity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1810.02340v2.
Vancouver
1. Lee N, Ajanthan T, Torr PHS (2018) SNIP: Single-shot Network Pruning based on Connection Sensitivity. arXiv

BibTeX

@article{lee2018snip,
  title = {SNIP: Single-shot Network Pruning based on Connection Sensitivity},
  author = {Lee, Namhoon and Ajanthan, Thalaiyasingam and Torr, Philip H. S.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1810.02340v2},
  eprint = {1810.02340}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors