Predicting Parameters in Deep Learning

Misha DenilBabak ShakibiLaurent DinhMarc'Aurelio RanzatoNando de Freitas

article2013NeurIPS1,414 citations

Demonstrates that deep neural networks contain substantial parameter redundancy by showing that more than 95% of their weights can be accurately predicted from a small subset of learned values without degrading predictive accuracy.

Listen

Modern artificial intelligence relies on scaling up deep neural networks to tens of millions or even billions of parameters. However, training these massive models across distributed clusters of computers causes severe communication and synchronization bottlenecks, leading to diminishing returns as more machines are added. The article addresses this inefficiency by investigating parameter redundancy within deep neural networks to determine whether large models can be trained with significantly fewer learned weights, thereby reducing computational overhead and hardware requirements.

The main objective of the article is to demonstrate that deep neural networks contain substantial redundancy and to introduce a general framework that trains models by learning only a small fraction of their parameters while accurately predicting the remainder. The researchers evaluated this approach by decomposing the weight matrices of neural networks into low-rank factorizations, splitting the weights into dynamic parameters (which are actively updated during training) and static parameters (which are fixed beforehand using known data structures or data-driven methods). They conducted high-level experimental evaluations across multiple standard image-processing architectures and benchmarks, including convolutional neural networks and multilayer perceptrons trained on datasets such as MNIST, CIFAR-10, and STL-10.

The findings show that deep learning parameter sets exhibit strong internal structure, meaning individual parameter values can be reliably predicted from a sparse subset. Most notably, the experimental results indicate that in the best cases, more than 95% of a network's weights can be predicted rather than learned dynamically, without causing any drop in predictive accuracy. The study also found that the method integrates seamlessly with standard deep learning techniques and optimizations without requiring structural redesigns of existing models.

These results demonstrate that organizations can substantially lower the communication overhead, memory footprint, and synchronization costs of distributed training. Because static parameters never change during learning, they can be distributed across machines without coordination penalties, making it feasible to train larger models on fewer machines or single workstations. The primary limitation highlighted is that successful implementation depends on effectively capturing data smoothness, either through domain knowledge or preliminary data-driven calculations. Practitioners are advised to evaluate low-rank parameter prediction strategies to optimize compute infrastructure, while further work should focus on automating static factor design across non-visual domains.

  • Paper: Optimal Brain Damage, Yann LeCun et al. (1989). Provides the foundational theoretical framework for identifying and pruning redundant neural network parameters by analyzing second-order derivative saliency.
  • Paper: Keeping Neural Networks Simple by Minimizing the Description Length of the Weights, Geoffrey E. Hinton et al. (1993). Introduces the principle of penalizing weight information content to keep models compact, establishing the information-theoretic basis for predicting parameter redundancy.
  • Paper: Model Compression, Cristian Bucila et al. (2006). Establishes early foundational methodology for model compression by demonstrating that large parameter spaces can be effectively distilled and reconstructed using compact representations.
  • Paper: Practical Variational Inference for Neural Networks, Alex Graves (2011). Formulates variational techniques for learning weight distributions and pruning, directly motivating probabilistic representations of network parameters.
  • Paper: A Simple Weight Decay Can Improve Generalization, Anders Krogh et al. (1991). Demonstrates how weight decay suppresses unnecessary parameter complexity to improve generalization, underpinning the assumption of parameter over-specification.
Cover for Predicting Parameters in Deep Learning

Abstract

We demonstrate that there is significant redundancy in the parameterization of several deep learning models. Given only a few weight values for each feature it is possible to accurately predict the remaining values. Moreover, we show that not only can the parameter values be predicted, but many of them need not be learned at all. We train several different architectures by learning only a small number of weights and predicting the rest. In the best case we are able to predict more than 95% of the weights of a network without any drop in accuracy.

Table of Contents

  • 1 Introduction
  • 2 Low rank weight matrices
  • 3 Feature prediction
  • 3.1 Choice of dictionary
  • 3.2 A concrete example
  • 3.3 Interpretation as pooling
  • 3.4 Columnar architecture
  • 3.5 Constructing dictionaries
  • 4 Experiments
  • 4.1 Multilayer perceptron
  • 4.2 Convolutional network
  • 4.3 Reconstruction ICA
  • 5 Related work and future directions
  • 6 Conclusion

Knowls

  1. Knowl 1 — Parameter Prediction via Fixed Dictionary Factorization

    model/method

    In deep neural networks with layer transformations h=g(vW)h = g(v W), where v∈Rnvv \in \mathbb{R}^{n_v} is the input, h∈Rnhh \in \mathbb{R}^{n_h} is the hidden representation, g(⋅)g(\cdot) is an activation function, and W∈Rnv×nhW \in \mathbb{R}^{n_v \times n_h} is the parameter matrix, parameter redundancy can be reduced by decomposing WW into a low-rank factorization:

    W=UVW = U V

    where U∈Rnv×nαU \in \mathbb{R}^{n_v \times n_\alpha} represents a dictionary matrix and V∈Rnα×nhV \in \mathbb{R}^{n_\alpha \times n_h} represents feature coordinates, with nα≪min⁡(nv,nh)n_\alpha \ll \min(n_v, n_h).

    Standard joint optimization of both UU and VV exhibits rotational ambiguity W=(UQ)(Q−1V)W = (U Q)(Q^{-1} V) for any invertible Q∈Rnα×nαQ \in \mathbb{R}^{n_\alpha \times n_\alpha} and empirically degrades classification performance. To eliminate this redundancy and retain high expressive capacity, the dictionary factor UU is fixed as a static parameter matrix constructed to encode prior structure (such as spatial smoothness or data correlations), while only the reduced factor VV is learned dynamically via gradient-based optimization.

  2. Knowl 2 — Feature Prediction via Kernel Ridge Regression

    model/method

    Individual neural network feature filters (columns w∈Rnvw \in \mathbb{R}^{n_v} of weight matrix W∈Rnv×nhW \in \mathbb{R}^{n_v \times n_h}) can be modeled as continuous functions w:W→Rw: \mathcal{W} \to \mathbb{R} evaluated over an input index space W\mathcal{W}. Given a randomly selected subset of input coordinates α⊂W\alpha \subset \mathcal{W} of size ∣α∣=nα≪nv|\alpha| = n_\alpha \ll n_v with explicitly stored weights wα∈Rnαw_\alpha \in \mathbb{R}^{n_\alpha}, the full feature vector is predicted via kernel ridge regression:

    w=kαT(Kα+λI)−1wαw = k_\alpha^T (K_\alpha + \lambda I)^{-1} w_\alpha

    where:

    • Kα∈Rnα×nαK_\alpha \in \mathbb{R}^{n_\alpha \times n_\alpha} is the kernel Gram matrix with entries (Kα)ij=k(i,j)(K_\alpha)_{ij} = k(i, j) for i,j∈αi, j \in \alpha,
    • kα∈Rnα×nvk_\alpha \in \mathbb{R}^{n_\alpha \times n_v} has entries (kα)ij=k(i,j)(k_\alpha)_{ij} = k(i, j) for i∈αi \in \alpha and j∈Wj \in \mathcal{W},
    • λ>0\lambda > 0 is a ridge regularization parameter, and
    • I∈Rnα×nαI \in \mathbb{R}^{n_\alpha \times n_\alpha} is the identity matrix.

    The entire weight matrix WW is predicted in parallel across all nhn_h hidden units as W=UαWαW = U_\alpha W_\alpha, where the static dictionary is Uα=kαT(Kα+λI)−1∈Rnv×nαU_\alpha = k_\alpha^T (K_\alpha + \lambda I)^{-1} \in \mathbb{R}^{n_v \times n_\alpha} and the dynamic learnable parameter matrix is Wα=[(w1)α,…,(wnh)α]∈Rnα×nhW_\alpha = [(w_1)_\alpha, \dots, (w_{n_h})_\alpha] \in \mathbb{R}^{n_\alpha \times n_h}.

    For p×pp \times p 2D image patches with pixel coordinates i=(ix,iy)i = (i_x, i_y), spatial smoothness is modeled using the squared exponential kernel:

    k(i,j)=exp⁡(−(ix−jx)2+(iy−jy)22σ2)k(i, j) = \exp\left( - \frac{(i_x - j_x)^2 + (i_y - j_y)^2}{2\sigma^2} \right)

    where σ\sigma is a length-scale parameter governing the degree of spatial smoothness.

  3. Knowl 3 — Equivalence of Parameter Prediction to Fixed Linear Pooling

    theoretical result

    Predicting a neural network layer's weight matrix W∈Rnv×nhW \in \mathbb{R}^{n_v \times n_h} from a reduced representation Wα∈Rnα×nhW_\alpha \in \mathbb{R}^{n_\alpha \times n_h} via a fixed dictionary Uα∈Rnv×nαU_\alpha \in \mathbb{R}^{n_v \times n_\alpha} (W=UαWαW = U_\alpha W_\alpha) is algebraically equivalent to a linear pooling operation on the layer inputs.

    For an input vector v∈Rnvv \in \mathbb{R}^{n_v} and activation function g(⋅)g(\cdot), the pre-activation g−1(h)=vWg^{-1}(h) = v W can be regrouped via matrix associativity:

    g−1(h)=v(UαWα)=(vUα)Wα=vαWαg^{-1}(h) = v (U_\alpha W_\alpha) = (v U_\alpha) W_\alpha = v_\alpha W_\alpha

    where vα=vUα∈Rnαv_\alpha = v U_\alpha \in \mathbb{R}^{n_\alpha} is a lower-dimensional projection of the input.

    Under this formulation, a parameter-predicted layer is functionally equivalent to a sequence of two operations:

    1. A fixed linear pooling layer that maps the nvn_v-dimensional input vv to an nαn_\alpha-dimensional pooled vector vαv_\alpha using the fixed matrix UαU_\alpha.
    2. A standard linear fully connected layer mapping vαv_\alpha to nhn_h hidden units with learnable parameter matrix WαW_\alpha.

    For convolutional layers where g−1(h)=v∗W∗g^{-1}(h) = v * W^* with vectorized filter bank W=UαwαW = U_\alpha w_\alpha, re-ordering yields g−1(h)=(v∗Uα)wαg^{-1}(h) = (v * U_\alpha) w_\alpha, where the input is first convolved with fixed dictionary filters UαU_\alpha and then linearly combined via the dynamic parameters wαw_\alpha.

  4. Knowl 4 — Columnar Parameter Prediction Architecture

    model/method

    To enable different subsets of hidden units to specialize in different receptive regions or frequency bands without sharing a single global basis dictionary, the weight matrix can be split into JJ independent columns.

    Given JJ distinct coordinate subsets α1,α2,…,αJ⊂W\alpha_1, \alpha_2, \dots, \alpha_J \subset \mathcal{W} and corresponding static sub-dictionaries Uαj∈Rnv×∣αj∣U_{\alpha_j} \in \mathbb{R}^{n_v \times |\alpha_j|}, each column jj predicts its own feature block:

    Wj=UαjWαjW_j = U_{\alpha_j} W_{\alpha_j}

    where Wαj∈R∣αj∣×nh,jW_{\alpha_j} \in \mathbb{R}^{|\alpha_j| \times n_{h,j}} is the dynamic parameter matrix for the nh,jn_{h,j} hidden units allocated to column jj, with ∑j=1Jnh,j=nh\sum_{j=1}^J n_{h,j} = n_h.

    The complete layer parameter matrix is the block concatenation:

    W=[W1,W2,…,WJ]=[Uα1Wα1,Uα2Wα2,…,UαJWαJ]W = [W_1, W_2, \dots, W_J] = [U_{\alpha_1} W_{\alpha_1}, U_{\alpha_2} W_{\alpha_2}, \dots, U_{\alpha_J} W_{\alpha_J}]

    Each column processes the shared input independently. While adding columns increases the number of static parameters stored in memory (because each column has a distinct dictionary UαjU_{\alpha_j}), the total number of dynamic (learnable) parameters remains unchanged for a fixed total number of hidden units nhn_h.

  5. Knowl 5 — Data-Driven Dictionary and Kernel Construction for Deep Layers

    model/method

    When weight spaces lack an intrinsic spatial topology (such as in intermediate or deep hidden layers of a network, or non-spatial input domains), dictionaries cannot be defined by analytical geometric kernels. Data-driven methods construct static dictionaries or kernel matrices:

    1. Autoencoder Dictionary (AE): A shallow unsupervised autoencoder is trained on the layer's input data, and its learned feature weights serve directly as the dictionary columns of UU.
    2. Empirical Covariance Kernel (Emp): The empirical covariance matrix Σ\Sigma of the layer's inputs is computed across the training dataset:

    Σ=E[(x−E[x])(x−E[x])T]\Sigma = \mathbb{E}\left[(x - \mathbb{E}[x])(x - \mathbb{E}[x])^T\right]

    The kernel function is defined as k(i,j)=Σijk(i, j) = \Sigma_{ij}, setting covariance between unit ii and unit jj as their kernel similarity in kernel ridge regression. 3. Squared Empirical Covariance Kernel (Emp2): The kernel function is defined as k(i,j)=(Σij)2k(i, j) = (\Sigma_{ij})^2.

    For deep architectures, layers are sequentially pre-trained in a greedy, layer-wise fashion using autoencoders to compute activation statistics. After pre-training, the empirical covariance kernels and dictionaries UU are computed and frozen (static), and the dynamic parameters of the entire network are fine-tuned via backpropagation.

  6. Knowl 6 — Distinction Between Dynamic and Static Parameters

    definition

    In parameter-reduced deep learning architectures, network parameters are categorized into two distinct classes:

    • Dynamic Parameters: Parameters whose numerical values are actively updated throughout optimization (e.g., via backpropagation or L-BFGS). In distributed training systems, dynamic parameters must be communicated and synchronized across worker nodes at each iteration or mini-batch, representing the primary communication overhead and scalability bottleneck.
    • Static Parameters: Parameters whose numerical values are computed once (or predefined via analytical kernels) and remain fixed during training. Because static parameters never change during optimization, duplicate copies can be distributed to compute workers without requiring synchronization mechanisms or network bandwidth during training.
  7. Knowl 7 — Parameter Prediction Performance in Multilayer Perceptrons

    empirical result

    Parameter prediction was evaluated on feedforward Multilayer Perceptrons (MLPs) across image classification (MNIST) and speech recognition (TIMIT):

    • MNIST Architecture and Results: A 784-500-500-10 MLP with sigmoid activations and a final uncompressed softmax layer was evaluated across dictionary construction methods (using 10 columns per layer). Dictionaries combining squared exponential kernels in the first layer with empirical covariance kernels in the second layer (SE-Emp and SE-Emp2) maintained classification error rates near the baseline model without prediction (~1.8% to 2.5% error) even when the proportion of learned parameters was reduced to 20%-30%.
    • Comparison to Baselines: Naive joint optimization of both factor matrices (LowRank) and random projection dictionaries (RandFixU and RandCon) exhibited severe performance degradation at low parameter budgets. Autoencoder pre-training for the LowRank baseline was found to be highly unstable.
    • TIMIT Speech Results: Evaluated on 12th-order MFCC plus energy features with first and second temporal derivatives using a two-hidden-layer MLP (1024 units per layer). Using empirical covariance dictionaries across both layers (Emp-Emp), the Phone Error Rate (PER) evaluated via Viterbi decoding remained stable at approximately 31% as the proportion of dynamic parameters learned was decreased from 100% down to 20%.
  8. Knowl 8 — Parameter Prediction in Convolutional Networks and Reconstruction ICA

    empirical result

    The feature prediction method was evaluated on convolutional neural networks and Reconstruction Independent Component Analysis (RICA):

    • Convolutional Networks on CIFAR-10: Evaluated on a 3-layer convolutional network (48 filters of size 8×8×38 \times 8 \times 3, 64 filters of size 8×8×488 \times 8 \times 48, and 64 filters of size 5×5×645 \times 5 \times 64) followed by a 500-unit fully connected layer (5 columns) and a softmax classifier. Using squared exponential kernel dictionaries to predict weights, learning only 25% of the parameters achieved classification error (~20%) indistinguishable from training 100% of the parameters.
    • RICA on CIFAR-10 and STL-10: Single-layer RICA models using a squared exponential kernel dictionary (length scale σ=1.0\sigma = 1.0) were able to predict over 50% of the dynamic parameters without noticeable accuracy degradation.
    • Parameter-Constrained RICA Scaling: When compared at fixed total budgets of dynamic parameters, an RICA model with 50% parameter prediction (which doubles its total number of features relative to standard RICA) achieved substantially lower classification error across all parameter budget sizes on both CIFAR-10 and STL-10 (e.g., error dropped from ~38% to ~28% on CIFAR-10 at low parameter counts).

Coverage note — No substantial contributed material was omitted. All models, mathematical formulations, dictionary construction strategies, pooling and columnar extensions, and experimental evaluations across MLPs, ConvNets, and RICA are fully covered.

References

  1. 1.Y. Bengio. Deep learning of representations: Looking forward. Technical Report arXiv:1305.0445, Universite de Montreal, 2013.
  2. 2.D. Cireŕan, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In IEEE Computer Vision and Pattern Recognition, pages 3642–3649, 2012.
  3. 3.D. Cireŕan, U. Meier, and J. Masci. High-performance neural networks for visual object classification. arXiv:1102.0183, 2011.
  4. 4.A. Coates, A. Karpathy, and A. Ng. Emergence of object-selective features in unsupervised feature learning. In Advances in Neural Information Processing Systems, pages 2690–2698, 2012.
  5. 5.A. Coates and A. Y. Ng. Selecting receptive fields in deep networks. In Advances in Neural Information Processing Systems, pages 2528–2536, 2011.
  6. 6.A. Coates, A. Y. Ng, and H. Lee. An analysis of single-layer networks in unsupervised feature learning. In Artificial Intelligence and Statistics, 2011.
  7. 7.J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng. Large scale distributed deep networks. In Advances in Neural Information Processing Systems, pages 1232–1240, 2012.
  8. 8.L. Deng, D. Yu, and J. Platt. Scalable stacking and learning for building deep architectures. In International Conference on Acoustics, Speech, and Signal Processing, pages 2133–2136, 2012.
  9. 9.I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio. Maxout networks. In International Conference on Machine Learning, 2013.
  10. 10.K. Gregor and Y. LeCun. Emergence of complex-like cells in a temporal product network with local receptive fields. arXiv preprint arXiv:1006.0448, 2010.
  11. 11.C. Gulcŕehre and Y. Bengio. Knowledge matters: Importance of prior information for optimization. In International Conference on Learning Representations, 2013.
  12. 12.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. CoRR, abs/1207.0580, 2012.
  13. 13.A. Krizhevsky, I. Sutskever, and G. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, pages 1106–1114, 2012.
  14. 14.K. Lang and G. Hinton. Dimensionality reduction and prior knowledge in e-set recognition. In Advances in Neural Information Processing Systems, 1990.
  15. 15.Q. V. Le, A. Karpenko, J. Ngiam, and A. Y. Ng. ICA with reconstruction cost for efficient overcomplete feature learning. Advances in Neural Information Processing Systems, 24:1017–1025, 2011.
  16. 16.Q. V. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. Corrado, J. Dean, and A. Ng. Building high-level features using large scale unsupervised learning. In International Conference on Machine Learning, 2012.
  17. 17.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  18. 18.Y. LeCun, J. S. Denker, S. Solla, R. E. Howard, and L. D. Jackel. Optimal brain damage. In Advances in Neural Information Processing Systems, pages 598–605, 1990.
  19. 19.K.-F. Lee and H.-W. Hon. Speaker-independent phone recognition using hidden markov models. Acoustics, Speech and Signal Processing, IEEE Transactions on, 37(11):1641–1648, 1989.
  20. 20.V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In Proc. 27th International Conference on Machine Learning, pages 807–814. Omnipress Madison, WI, 2010.
  21. 21.M. Ranzato, A. Krizhevsky, and G. E. Hinton. Factored 3-way restricted Boltzmann machines for modeling natural images. In Artificial Intelligence and Statistics, 2010.
  22. 22.R. Rigamonti, A. Sironi, V. Lepetit, and P. Fua. Learning separable filters. In IEEE Computer Vision and Pattern Recognition, 2013.
  23. 23.R. Rubinstein, M. Zibulevsky, and M. Elad. Double sparsity: learning sparse dictionaries for sparse signal approximation. IEEE Transactions on Signal Processing, 58:1553–1564, 2010.
  24. 24.D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Learning representations by back-propagating errors. Nature, 323(6088):533–536, 1986.
  25. 25.J. Shawe-Taylor and N. Cristianini. Kernel Methods for Pattern Analysis. Cambridge University Press, New York, NY, USA, 2004.
  26. 26.K. Swersky, M. Ranzato, D. Buchman, B. Marlin, and N. Freitas. On autoencoders and score matching for energy based models. In International Conference on Machine Learning, pages 1201–1208, 2011.
  27. 27.P. Vincent and Y. Bengio. A neural support vector network architecture with adaptive kernels. In International Joint Conference on Neural Networks, pages 187–192, 2000.

Citation

MLA
Denil, M., et al. “Predicting Parameters in Deep Learning”. arXiv, 2013, http://arxiv.org/abs/1306.0543v2.
APA
Denil, M., Shakibi, B., Dinh, L., Ranzato, M., & Freitas, N. de . (2013). Predicting Parameters in Deep Learning. arXiv. http://arxiv.org/abs/1306.0543v2
Chicago
Denil, M., B. Shakibi, L. Dinh, M. Ranzato, and N. de . Freitas. 2013. “Predicting Parameters in Deep Learning”. arXiv. http://arxiv.org/abs/1306.0543v2.
Harvard
Denil, M. et al. (2013) “Predicting Parameters in Deep Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1306.0543v2.
Vancouver
1. Denil M, Shakibi B, Dinh L, Ranzato M, Freitas N de (2013) Predicting Parameters in Deep Learning. arXiv

BibTeX

@article{denil2013predicting,
  title = {Predicting Parameters in Deep Learning},
  author = {Denil, Misha and Shakibi, Babak and Dinh, Laurent and Ranzato, Marc'Aurelio and Freitas, Nando de},
  year = {2013},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1306.0543v2},
  eprint = {1306.0543}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission