Practical Variational Inference for Neural Networks

Alex Graves

article2011NeurIPS1,894 citations

Proposes a stochastic variational inference method with a diagonal Gaussian posterior that scales practical Bayesian learning, regularisation, and weight pruning to general differentiable neural network architectures.

Listen

Complex neural networks often struggle with overfitting and high computational demands, while traditional Bayesian methods that quantify uncertainty have remained mathematically intractable for all but the simplest architectures. This article develops and evaluates a practical, stochastic variational inference method that can be applied to virtually any standard neural network trained with gradient descent, framing the learning problem through a Minimum Description Length lens.

The approach replaces intractable analytical derivations with numerical sampling over a diagonal Gaussian weight distribution, allowing simultaneous learning of weight means and variances. The article evaluated this framework alongside an information-theoretic pruning rule on a 15-layer recurrent neural network using the benchmark TIMIT speech dataset (3,696 training utterances).

Key findings demonstrate that adaptive weight noise achieved a 23.8% phoneme error rate on the test set, noticeably outperforming standard maximum likelihood training (27.1%) and fixed regularizers like weight decay (27.4%). Because the network actively compressed the data, it proved resistant to overfitting, eliminating the need to sacrifice training data for early stopping validation. Furthermore, applying the proposed pruning rule allowed removing between 55% and 78% of the network weights while slightly reducing the error rate further to 23.3% after retraining.

These results provide a straightforward method for practitioners to train compact, robust neural models that generalize better without requiring custom Bayesian derivations or complex tuning. The method directly reduces hardware memory overhead and inferencing costs by eliminating unneeded weights.

Organizations deploying large neural architectures should consider adopting adaptive weight distributions to improve generalization and using the signal-to-noise pruning heuristic to streamline model size. Next steps should explore validating the method across additional domains, though teams should note that training with stochastic derivatives increases training times and introduces noisy optimization curves.

Graves (2011).pdf
  • Paper: Keeping Neural Networks Simple by Minimizing the Description Length of the Weights, Geoffrey E. Hinton et al. (1993). This seminal paper introduced Minimum Description Length (MDL) and variational weight penalties to keep neural network weights simple, which the source directly builds upon and generalises to arbitrary architectures.
  • Paper: Optimal Brain Damage, Yann LeCun et al. (1989). This foundational work establishes second-order saliency pruning heuristics for neural network compression, motivating the variational pruning criteria developed in the source.
  • Paper: A Simple Weight Decay Can Improve Generalization, A. Krogh et al. (1991). This paper provides the classical understanding of weight decay regularization, which the source re-interprets and analyses through a variational lens.
  • Paper: An Introduction to Variational Methods for Graphical Models, MICHAEL I. JORDAN et al. (1999). This comprehensive overview establishes the foundational mathematical principles of variational inference and lower bounding intractable distributions in graphical and neural models.
Cover for Practical Variational Inference for Neural Networks

Abstract

Variational methods have been previously explored as a tractable approximation to Bayesian inference for neural networks. However the approaches proposed so far have only been applicable to a few simple network architectures. This paper introduces an easy-to-implement stochastic variational method (or equivalently, minimum description length loss function) that can be applied to most neural networks. Along the way it revisits several common regularisers from a variational perspective. It also provides a simple pruning heuristic that can both drastically reduce the number of network weights and lead to improved generalisation. Experimental results are provided for a hierarchical multidimensional recurrent neural network applied to the TIMIT speech corpus.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Neural Networks
  • 3 Variational Inference
  • 4 Minimum Description Length
  • 5 Choice of Distributions
  • 5.1 Delta Posterior
  • 5.2 Gaussian Posterior
  • 6 Optimisation
  • 7 Pruning
  • 8 Experiments
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Minimum Description Length Formulation of Variational Inference for Neural Networks

    model/method

    Variational inference for a neural network parameterised by weights w={wi}i=1Ww = \{w_i\}_{i=1}^W on a dataset D={(x,y)}\mathcal{D} = \{(x, y)\} can be formulated as minimising a Minimum Description Length (MDL) loss function L(α,β,D)\mathcal{L}(\alpha, \beta, \mathcal{D}) that measures the total information (in nats) required to communicate the targets given the inputs:

    L(α,β,D)=LE(β,D)+LC(α,β)\mathcal{L}(\alpha, \beta, \mathcal{D}) = \mathcal{L}^E(\beta, \mathcal{D}) + \mathcal{L}^C(\alpha, \beta)

    where α\alpha parameterises the prior distribution over weights P(w∣α)P(w|\alpha), and β\beta parameterises a tractable variational posterior distribution Q(w∣β)Q(w|\beta).

    The error loss LE(β,D)\mathcal{L}^E(\beta, \mathcal{D}) represents the expected cost of transmitting the dataset targets using network outputs sampled from Q(w∣β)Q(w|\beta):

    LE(β,D)=⟨LN(w,D)⟩w∼Q(β)\mathcal{L}^E(\beta, \mathcal{D}) = \langle \mathcal{L}^N(w, \mathcal{D}) \rangle_{w \sim Q(\beta)}

    where LN(w,D)=−ln⁡Pr⁡(D∣w)=−∑(x,y)∈Dln⁡Pr⁡(y∣x,w)\mathcal{L}^N(w, \mathcal{D}) = -\ln \Pr(\mathcal{D}|w) = -\sum_{(x,y) \in \mathcal{D}} \ln \Pr(y|x, w) is the negative log-likelihood of the dataset.

    The complexity loss LC(α,β)\mathcal{L}^C(\alpha, \beta) represents the bits-back communication cost of transmitting the network parameters to a receiver who knows the prior P(w∣α)P(w|\alpha):

    LC(α,β)=DKL(Q(β)∥P(α))=∫Q(w∣β)ln⁡(Q(w∣β)P(w∣α))dw\mathcal{L}^C(\alpha, \beta) = D_{\text{KL}}(Q(\beta) \parallel P(\alpha)) = \int Q(w|\beta) \ln \left( \frac{Q(w|\beta)}{P(w|\alpha)} \right) dw

    Minimising L(α,β,D)\mathcal{L}(\alpha, \beta, \mathcal{D}) balances prediction accuracy against model complexity.

  2. Knowl 2 — Stochastic Gradient Estimation for Diagonal Gaussian Variational Posteriors

    equation

    For a neural network with WW parameters where the variational posterior is a diagonal Gaussian Q(w∣β)=∏i=1WN(wi;μi,σi2)Q(w|\beta) = \prod_{i=1}^W \mathcal{N}(w_i; \mu_i, \sigma_i^2) with variational parameters β={μi,σi2}i=1W\beta = \{\mu_i, \sigma_i^2\}_{i=1}^W, the exact derivatives of the error loss LE(β,D)=⟨LN(w,D)⟩w∼Q(β)\mathcal{L}^E(\beta, \mathcal{D}) = \langle \mathcal{L}^N(w, \mathcal{D}) \rangle_{w \sim Q(\beta)} are approximated via Monte Carlo sampling and the empirical Fisher information diagonal:

    ∂LE(β,D)∂μi=⟨∂LN(w,D)∂wi⟩w∼Q(β)≈1S∑k=1S∂LN(wk,D)∂wi\frac{\partial \mathcal{L}^E(\beta, \mathcal{D})}{\partial \mu_i} = \left\langle \frac{\partial \mathcal{L}^N(w, \mathcal{D})}{\partial w_i} \right\rangle_{w \sim Q(\beta)} \approx \frac{1}{S} \sum_{k=1}^S \frac{\partial \mathcal{L}^N(w^k, \mathcal{D})}{\partial w_i}

    ∂LE(β,D)∂σi2=12⟨∂2LN(w,D)∂wi2⟩w∼Q(β)≈12⟨(∂LN(w,D)∂wi)2⟩w∼Q(β)≈12S∑k=1S(∂LN(wk,D)∂wi)2\frac{\partial \mathcal{L}^E(\beta, \mathcal{D})}{\partial \sigma_i^2} = \frac{1}{2} \left\langle \frac{\partial^2 \mathcal{L}^N(w, \mathcal{D})}{\partial w_i^2} \right\rangle_{w \sim Q(\beta)} \approx \frac{1}{2} \left\langle \left( \frac{\partial \mathcal{L}^N(w, \mathcal{D})}{\partial w_i} \right)^2 \right\rangle_{w \sim Q(\beta)} \approx \frac{1}{2S} \sum_{k=1}^S \left( \frac{\partial \mathcal{L}^N(w^k, \mathcal{D})}{\partial w_i} \right)^2

    where {wk}k=1S\{w^k\}_{k=1}^S are SS independent parameter vectors sampled from Q(w∣β)Q(w|\beta), and LN(w,D)=−ln⁡Pr⁡(D∣w)\mathcal{L}^N(w, \mathcal{D}) = -\ln \Pr(\mathcal{D}|w). The second derivative approximation substitutes the negative diagonal of the empirical Fisher information matrix for the diagonal Hessian, which becomes exact when the model distribution Pr⁡(D∣w)\Pr(\mathcal{D}|w) matches the empirical distribution of D\mathcal{D}.

  3. Knowl 3 — Complexity Loss and Analytical Gradients for Gaussian Prior and Posterior

    equation

    When both the variational posterior Q(w∣β)=∏i=1WN(wi;μi,σi2)Q(w|\beta) = \prod_{i=1}^W \mathcal{N}(w_i; \mu_i, \sigma_i^2) and the prior P(w∣α)=∏i=1WN(wi;μ0,σ02)P(w|\alpha) = \prod_{i=1}^W \mathcal{N}(w_i; \mu_0, \sigma_0^2) are diagonal Gaussians over WW network weights, the complexity loss LC(α,β)=DKL(Q(β)∥P(α))\mathcal{L}^C(\alpha, \beta) = D_{\text{KL}}(Q(\beta) \parallel P(\alpha)) has the closed-form expression:

    LC(α,β)=∑i=1W[ln⁡σ0σi+12σ02((μi−μ0)2+σi2−σ02)]\mathcal{L}^C(\alpha, \beta) = \sum_{i=1}^W \left[ \ln \frac{\sigma_0}{\sigma_i} + \frac{1}{2\sigma_0^2} \left( (\mu_i - \mu_0)^2 + \sigma_i^2 - \sigma_0^2 \right) \right]

    The partial derivatives of the complexity loss with respect to the posterior parameters μi\mu_i and σi2\sigma_i^2 for weight ii are:

    ∂LC(α,β)∂μi=μi−μ0σ02\frac{\partial \mathcal{L}^C(\alpha, \beta)}{\partial \mu_i} = \frac{\mu_i - \mu_0}{\sigma_0^2}

    ∂LC(α,β)∂σi2=12(1σ02−1σi2)\frac{\partial \mathcal{L}^C(\alpha, \beta)}{\partial \sigma_i^2} = \frac{1}{2} \left( \frac{1}{\sigma_0^2} - \frac{1}{\sigma_i^2} \right)

  4. Knowl 4 — Closed-Form Optimal Prior Parameter Updates

    equation

    For a neural network with WW weights and variational parameters β\beta, the prior parameters α\alpha that minimize the free energy L(α,β,D)\mathcal{L}(\alpha, \beta, \mathcal{D}) can be updated in closed form at each training step.

    For a Gaussian prior α={μ0,σ02}\alpha = \{\mu_0, \sigma_0^2\} and a diagonal Gaussian posterior β={μi,σi2}i=1W\beta = \{\mu_i, \sigma_i^2\}_{i=1}^W, the optimal prior mean μ^0\hat{\mu}_0 and variance σ^02\hat{\sigma}_0^2 are:

    μ^0=1W∑i=1Wμi,σ^02=1W∑i=1W[σi2+(μi−μ^0)2]\hat{\mu}_0 = \frac{1}{W} \sum_{i=1}^W \mu_i, \qquad \hat{\sigma}_0^2 = \frac{1}{W} \sum_{i=1}^W \left[ \sigma_i^2 + (\mu_i - \hat{\mu}_0)^2 \right]

    For a delta posterior Q(w∣w∗)=δ(w−w∗)Q(w|w^*) = \delta(w - w^*) with weights w∗={wi}i=1Ww^* = \{w_i\}_{i=1}^W:

    1. Under a Gaussian prior α={μ0,σ02}\alpha = \{\mu_0, \sigma_0^2\}: μ^0=1W∑i=1Wwi,σ^02=1W∑i=1W(wi−μ^0)2\hat{\mu}_0 = \frac{1}{W} \sum_{i=1}^W w_i, \qquad \hat{\sigma}_0^2 = \frac{1}{W} \sum_{i=1}^W (w_i - \hat{\mu}_0)^2

    2. Under a Laplace prior P(w∣μ0,b)=∏i=1W12bexp⁡(−∣wi−μ0∣b)P(w|\mu_0, b) = \prod_{i=1}^W \frac{1}{2b} \exp\left(-\frac{|w_i - \mu_0|}{b}\right): μ^0=median({wi}i=1W),b^=1W∑i=1W∣wi−μ^0∣\hat{\mu}_0 = \text{median}(\{w_i\}_{i=1}^W), \qquad \hat{b} = \frac{1}{W} \sum_{i=1}^W |w_i - \hat{\mu}_0|

  5. Knowl 5 — Mini-Batch Online Loss Scaling for MDL Variational Training

    model/method

    When a dataset D\mathcal{D} is divided into BB equally sized batches {bj}j=1B\{b_j\}_{j=1}^B for online or mini-batch gradient optimization, the online variational loss L(α,β,bj)\mathcal{L}(\alpha, \beta, b_j) for batch bjb_j is defined as:

    L(α,β,bj)=1BLC(α,β)+LE(β,bj)\mathcal{L}(\alpha, \beta, b_j) = \frac{1}{B} \mathcal{L}^C(\alpha, \beta) + \mathcal{L}^E(\beta, b_j)

    where LC(α,β)=DKL(Q(β)∥P(α))\mathcal{L}^C(\alpha, \beta) = D_{\text{KL}}(Q(\beta) \parallel P(\alpha)) is the parameter complexity cost and LE(β,bj)=⟨−ln⁡Pr⁡(bj∣w)⟩w∼Q(β)\mathcal{L}^E(\beta, b_j) = \langle -\ln \Pr(b_j|w) \rangle_{w \sim Q(\beta)} is the batch error loss.

    The scaling factor 1/B1/B applied to the complexity loss accounts for the Minimum Description Length principle that the model parameters are transmitted only once for the entire dataset D\mathcal{D}, whereas the prediction targets must be transmitted for each batch.

  6. Knowl 6 — Signal-to-Noise Ratio Pruning Heuristic for Variational Neural Networks

    model/method

    For a neural network trained with a diagonal Gaussian posterior Q(w∣β)=∏i=1WN(wi;μi,σi2)Q(w|\beta) = \prod_{i=1}^W \mathcal{N}(w_i; \mu_i, \sigma_i^2), a weight wiw_i can be pruned (set permanently to zero) if its relative probability density at zero compared to its mode exceeds a threshold γ∈[0,1]\gamma \in [0, 1]:

    qi(0)qi(μi)=exp⁡(−μi22σi2)>γ  ⟺  ∣μiσi∣<λ\frac{q_i(0)}{q_i(\mu_i)} = \exp\left( -\frac{\mu_i^2}{2\sigma_i^2} \right) > \gamma \iff \left| \frac{\mu_i}{\sigma_i} \right| < \lambda

    where λ=−2ln⁡γ≥0\lambda = \sqrt{-2 \ln \gamma} \ge 0.

    Setting λ=0\lambda = 0 prunes no weights. A safe threshold rule of thumb is the point where setting a weight to zero makes it no less probable than an average sample drawn from qi(wi)q_i(w_i), which corresponds to:

    λ=2ln⁡2=ln⁡2≈0.83\lambda = \sqrt{2 \ln \sqrt{2}} = \sqrt{\ln 2} \approx 0.83

    Pruning eliminates parameters with low signal-to-noise ratio ∣μi∣σi\frac{|\mu_i|}{\sigma_i}. Subsequent retraining of the remaining unpruned parameters reduces gradient noise (since pruned weights are no longer sampled) without increasing model complexity.

  7. Knowl 7 — Variational Interpretation of Standard Regularisation and Weight Noise

    model/method

    Common neural network regularisation methods map directly to special cases of the variational Minimum Description Length framework:

    1. L2 Regularisation (Weight Decay): Equivalent to a delta posterior Q(w)=δ(w−w∗)Q(w) = \delta(w - w^*) paired with a fixed zero-mean Gaussian prior P(w)=∏i=1WN(0,σ2)P(w) = \prod_{i=1}^W \mathcal{N}(0, \sigma^2).
    2. L1 Regularisation: Equivalent to a delta posterior paired with a fixed zero-mean Laplace prior P(w)=∏i=1W12bexp⁡(−∣wi∣b)P(w) = \prod_{i=1}^W \frac{1}{2b} \exp\left(-\frac{|w_i|}{b}\right).
    3. Synaptic Weight Noise: Equivalent to a diagonal Gaussian posterior with fixed variance σi2=σnoise2\sigma_i^2 = \sigma_{\text{noise}}^2 and an improper uniform prior over weights, drawing S=1S=1 weight sample per input-target pair during gradient estimation.

    When a fixed-variance Gaussian posterior is paired with a proper adaptive Gaussian prior rather than a uniform prior, the description length becomes finite, enabling the model to achieve data compression and eliminating the requirement for early stopping.

  8. Knowl 8 — Experimental Setup for TIMIT Phoneme Recognition with Hierarchical MDRNN

    experimental setup

    Variational neural network methods were evaluated on phoneme recognition using the TIMIT acoustic-phonetic speech corpus, using spectrogram images as input and the standard 39-phoneme collapsed target set.

    • Dataset Partitioning: Core training set of 3,696 utterances and core test set of 192 utterances. For models requiring early stopping, a validation set of 184 utterances was held out from training, leaving 3,512 training utterances. Models with Gaussian posteriors and Gaussian priors were trained on all 3,696 utterances without early stopping.
    • Architecture: Hierarchical Multidimensional Recurrent Neural Network (MDRNN) with Long Short-Term Memory (LSTM) cells scanning the 2D spectrogram representations in both horizontal (time) and vertical (frequency) directions, followed by a Connectionist Temporal Classification (CTC) output layer with 40 units (39 phonemes plus blank). Subsampling window dimensions across three hierarchical stages were 2×42 \times 4, 2×42 \times 4, and 1×41 \times 4, yielding 15 layers, 1,306 units, and 139,536 weights.
    • Optimization: Online steepest descent with updates after every sequence, learning rate 10−410^{-4}, and momentum 0.90.9. For Gaussian posteriors, posterior standard deviations σi\sigma_i were initialised to 0.0750.075 and weight means μi\mu_i were drawn from N(0,0.12)\mathcal{N}(0, 0.1^2); one weight sample (S=1S=1) was drawn per sequence during training.
    • Inference: Maximum a posteriori (MAP) decoding using the posterior mean w∗=μw^* = \mu with CTC prefix search decoding at a probability threshold of 0.9950.995.
  9. Knowl 9 — Empirical Comparison of Prior and Posterior Formulations on TIMIT Phoneme Recognition

    data/table

    The table below compares phoneme error rate (PER, defined as total edit distance divided by target phoneme count ×100\times 100), convergence epochs, and target compression ratio (ratio of transmitted target description length relative to a uniform code of ≈5.3\approx 5.3 bits per phoneme) across different prior and posterior combinations on the TIMIT core test set.

    Name Posterior Prior Error (%) Epochs Ratio
    Adaptive L1 Delta Laplace 49.0 7 –
    Adaptive L2 Delta Gauss 35.1 421 –
    Adaptive mean L2 Delta Gauss σ2=0.1\sigma^2 = 0.1 28.0 53 –
    L2 Delta Gauss μ=0,σ2=0.1\mu = 0, \sigma^2 = 0.1 27.4 59 –
    Maximum likelihood Delta Uniform 27.1 44 –
    L1 Delta Laplace μ=0,b=1/12\mu = 0, b = 1/12 26.0 545 –
    Adaptive mean L1 Delta Laplace b=1/12b = 1/12 25.4 765 –
    Weight noise Gauss σi=0.075\sigma_i = 0.075 Uniform 25.4 220 –
    Adaptive prior weight noise Gauss σi=0.075\sigma_i = 0.075 Gauss 24.7 260 0.542
    Adaptive weight noise Gauss Gauss 23.8 384 0.286

    The fully adaptive diagonal Gaussian posterior with an adaptive Gaussian prior ('Adaptive weight noise') achieved the lowest error rate (23.8%) and highest data compression (ratio 0.286). Fully adaptive priors combined with delta posteriors performed poorly because prior variances collapsed to near-zero values (σ2≈0.003\sigma^2 \approx 0.003 for L2; b≈0.002b \approx 0.002 for L1).

  10. Knowl 10 — Effect of Signal-to-Noise Ratio Pruning and Retraining on Hierarchical MDRNN

    data/table

    Applying the signal-to-noise ratio pruning rule ∣μi/σi∣<λ|\mu_i / \sigma_i| < \lambda to the trained adaptive Gaussian posterior network followed by retraining yields the following results on the TIMIT core test set across various pruning thresholds λ\lambda:

    λ\lambda Weights Percent Initial error (%) Retrain error (%) Retrain Epochs Bits/weight
    0 139,536 100.0% 23.8 23.8 0 0.53
    0.01 107,974 77.4% 23.8 24.0 972 0.72
    0.05 63,079 45.2% 23.9 23.5 35 1.15
    0.1 52,984 37.9% 23.9 23.3 351 1.40
    0.2 43,182 30.9% 23.9 23.7 740 1.82
    0.5 31,120 22.3% 24.0 23.3 125 2.21
    1.0 22,806 16.3% 24.5 24.1 403 3.19
    2.0 16,029 11.5% 28.0 24.5 335 3.55

    Pruning up to 77.7%77.7\% of the weights (at λ=0.5\lambda = 0.5, retaining 31,120 weights) combined with retraining improved the test phoneme error rate from 23.8%23.8\% to 23.3%23.3\%. The initial error immediately after pruning remained stable up to λ≈0.5\lambda \approx 0.5 before increasing sharply beyond the theoretical safe threshold λ≈0.83\lambda \approx 0.83.

Coverage note — No substantial contributed material from the paper was omitted; qualitative visualizations of individual weight bit costs in 2D LSTM connections were subsumed under the quantitative pruning and complexity loss knowls.

References

  1. 1.D. Barber and C. M. Bishop. Ensemble learning in Bayesian neural networks., pages 215–237. Springer-Verlag, Berlin, 1998.
  2. 2.D. Barber and B. Schottky. Radial basis functions: A bayesian treatment. In NIPS, 1997.
  3. 3.G. E. Dahl, M. Ranzato, A. rahman Mohamed, and G. Hinton. Phone recognition with the mean-covariance restricted boltzmann machine. In J. Lafferty, C. K. I. Williams, J. Shawe-Taylor, R. Zemel, and A. Culotta, editors, Advances in Neural Information Processing Systems 23, pages 469–477. 2010.
  4. 4.DARPA-ISTO. The DARPA TIMIT Acoustic-Phonetic Continuous Speech Corpus (TIMIT), speech disc cd1-1.1 edition, 1990.
  5. 5.B. J. Frey. Graphical models for machine learning and digital communication. MIT Press, Cambridge, MA, USA, 1998.
  6. 6.K. fu Lee and H. wuen Hon. Speaker-independent phone recognition using hidden markov models. IEEE Transactions on Acoustics, Speech, and Signal Processing, 1989.
  7. 7.C. L. Giles and C. W. Omlin. Pruning recurrent neural networks for improved generalization performance. IEEE Transactions on Neural Networks, 5:848–851, 1994.
  8. 8.A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the International Conference on Machine Learning, ICML 2006, Pittsburgh, USA, 2006.
  9. 9.A. Graves and J. Schmidhuber. Offline handwriting recognition with multidimensional recurrent neural networks. In NIPS, pages 545–552, 2008.
  10. 10.G. E. Hinton and D. van Camp. Keeping the neural networks simple by minimizing the description length of the weights. In COLT, pages 5–13, 1993.
  11. 11.S. Hochreiter and J. Schmidhuber. Long Short-Term Memory. Neural Computation, 9(8):1735–1780, 1997.
  12. 12.A. Honkela and H. Valpola. Variational learning and bits-back coding: An information-theoretic view to bayesian learning. IEEE Transactions on Neural Networks, 15:800–810, 2004.
  13. 13.K.-C. Jim, C. Giles, and B. Horne. An analysis of noise in recurrent neural networks: convergence and generalization. Neural Networks, IEEE Transactions on, 7(6):1424 –1438, nov 1996.
  14. 14.N. D. Lawrence. Variational Inference in Probabilistic Models. PhD thesis, University of Cambridge, 2000.
  15. 15.Y. Le Cun, J. Denker, and S. Solla. Optimal brain damage. In D. S. Touretzky, editor, Advances in Neural Information Processing Systems, volume 2, pages 598–605. Morgan Kaufmann, San Mateo, CA, 1990.
  16. 16.D. J. C. MacKay. Probable networks and plausible predictions - a review of practical bayesian methods for supervised neural networks. Neural Computation, 1995.
  17. 17.S. J. Nowlan and G. E. Hinton. Simplifying neural networks by soft weight sharing. Neural Computation, 4:173–193, 1992.
  18. 18.M. Opper and C. Archambeau. The variational gaussian approximation revisited. Neural Computation, 21(3):786–792, 2009.
  19. 19.D. Plaut, S. Nowlan, and G. E. Hinton. Experiments on learning by back propagation. Technical Report CMU-CS-86-126, Department of Computer Science, Carnegie Mellon University, Pittsburgh, PA, 1986.
  20. 20.M. Riedmiller and T. Braun. A direst adaptive method for faster backpropagation learning: The rprop algorithm. In International Symposium on Neural Networks, 1993.
  21. 21.J. Rissanen. Modeling by shortest data description. Automatica, 14(5):465 – 471, 1978.
  22. 22.D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Learning representations by back-propagating errors, pages 696–699. MIT Press, Cambridge, MA, USA, 1988.
  23. 23.C. E. Shannon. A mathematical theory of communication. Bell system technical journal, 27, 1948.
  24. 24.P. Smolensky. Information processing in dynamical systems: foundations of harmony theory, pages 194–281. MIT Press, Cambridge, MA, USA, 1986.
  25. 25.C. S. Wallace. Classification by minimum-message-length inference. In Proceedings of the international conference on Advances in computing and information, ICCI'90, pages 72–81, New York, NY, USA, 1990. Springer-Verlag New York, Inc.
  26. 26.I. H. Witten, R. M. Neal, and J. G. Cleary. Arithmetic coding for data compression. Commun. ACM, 30:520–540, June 1987.

Citation

MLA
Graves, A. “Practical Variational Inference for Neural Networks”. Advances in Neural Information Processing Systems, vol. 24, 2011, https://proceedings.neurips.cc/paper_files/paper/2011/file/7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf.
APA
Graves, A. (2011). Practical Variational Inference for Neural Networks. Advances in Neural Information Processing Systems, 24. https://proceedings.neurips.cc/paper_files/paper/2011/file/7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf
Chicago
Graves, A. 2011. “Practical Variational Inference for Neural Networks”. Advances in Neural Information Processing Systems 24. https://proceedings.neurips.cc/paper_files/paper/2011/file/7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf.
Harvard
Graves, A. (2011) “Practical Variational Inference for Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2011/file/7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf.
Vancouver
1. Graves A (2011) Practical Variational Inference for Neural Networks. Advances in Neural Information Processing Systems 24:

BibTeX

@inproceedings{graves2011practical,
  title = {Practical Variational Inference for Neural Networks},
  author = {Graves, Alex},
  year = {2011},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {24},
  url = {https://proceedings.neurips.cc/paper_files/paper/2011/file/7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors