Variational Dropout and the Local Reparameterization Trick

Diederik P. KingmaTim SalimansMax Welling

article2015NeurIPS1,708 citations

Introduces the local reparameterization trick to drastically reduce gradient variance during variational inference in neural networks, establishing a Bayesian foundation for Gaussian dropout that allows dropout rates to be learned directly from data.

Listen

Deep neural networks are powerful tools for pattern recognition, but their immense flexibility frequently leads to overfitting, where models memorize spurious details in training data instead of learning generalizable rules. While standard regularization techniques such as dropout manage this problem empirically by injecting random noise, theoretical approaches like Bayesian inference remain attractive because they estimate uncertainty over model weights in a principled manner. However, existing stochastic gradient variational Bayes methods suffer from severe computational bottlenecks and high estimation variance when applied to network parameters, preventing them from matching the efficiency and performance of simpler heuristic techniques.

The article addresses this challenge by introducing the local reparameterization trick to dramatically reduce gradient variance during variational inference and establishing a formal Bayesian foundation for dropout. It evaluates whether this formulation can provide a faster, adaptive alternative—termed variational dropout—that learns optimal noise rates directly from data rather than relying on manually fixed settings.

To demonstrate this, the authors restructured the mathematical formulation of stochastic gradients so that global uncertainty over millions of network weights is converted into local uncertainty over neuron activations across data batches. This theoretical adjustment makes gradient variance scale inversely with batch size and ensures compatibility with high-performance, parallel matrix operations on graphics processing units. The authors validated their framework through empirical benchmarks on standard image classification datasets, including MNIST and CIFAR-10, across multiple fully connected and convolutional architectures.

The findings confirm three major advantages of this approach. First, the local reparameterization trick provides an over 200-fold speedup in training time compared to standard sampling approaches (reducing per-epoch runtime from 1,635 seconds to 7.4 seconds on a graphics processor) while reducing gradient variance by factors of two to four or more. Second, the article proves that common Gaussian dropout corresponds directly to variational inference under a scale-invariant log-uniform prior, resolving the long-standing gap between empirical dropout and Bayesian theory. Third, variational dropout matches or exceeds standard and Gaussian dropout in predictive accuracy, showing substantial performance gains in smaller networks by automatically learning lower dropout rates to prevent underfitting.

These results demonstrate that organizations can deploy rigorous Bayesian uncertainty estimation in deep learning without sacrificing training speed or computational efficiency. By enabling networks to learn their own regularizing noise levels per layer, neuron, or weight, engineering teams can eliminate time-consuming manual hyperparameter tuning. This lowers the computational cost and operational risk associated with deploying overparameterized models across diverse tasks.

Teams developing deep learning pipelines should adopt the local reparameterization formulation and evaluate variational dropout as an automated replacement for manual dropout tuning. To maintain training stability, practitioners should enforce the recommended constraint on noise variance parameters and consider slightly downscaling the regularizing divergence penalty to avoid underfitting. While the empirical evaluations rely on standard vision benchmarks and impose specific mathematical constraints on posterior distributions, the underlying framework provides high confidence for scaling principled Bayesian regularization in practical machine learning workflows.

arXiv: 1506.02557
  • Paper: Dropout: a simple way to prevent neural networks from overfitting, Nitish Srivastava et al. (2014). Introduces standard dropout and Gaussian dropout regularization, providing the primary foundation that Variational Dropout reinterprets and extends as variational inference.
  • Paper: Stochastic Backpropagation and Approximate Inference in Deep Generative Models, Danilo Jimenez Rezende et al. (2014). Establishes the stochastic gradient variational Bayes (SGVB) framework and reparameterization trick that the source adapts into a local reparameterization formulation.
  • Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Demonstrates backpropagation-compatible variational inference over neural network weights (Bayes by Backprop), setting up the parameter uncertainty problem that the source accelerates using local noise.
  • Paper: Practical Variational Inference for Neural Networks, Alex Graves (2011). Pioneers practical stochastic variational inference over neural network weights, providing early groundwork for learning posterior distributions over network parameters.
  • Paper: Stochastic variational inference, Matt Hoffman et al. (2012). Formulates stochastic variational inference using minibatch subsampling, foundational to scalable variational approximations.
  • Paper: Regularization of Neural Networks using DropConnect, Li Wan et al. (2013). Explores stochastic regularization applied directly to weights rather than activations, offering critical motivation for weight-space noise in neural networks.
  • Paper: Recurrent Neural Network Regularization, Wojciech Zaremba et al. (2014). Analyzes the challenges and initial heuristics of applying dropout to recurrent networks, which variational dropout later addresses systematically.
Cover for Variational Dropout and the Local Reparameterization Trick

Abstract

We investigate a local reparameterizaton technique for greatly reducing the variance of stochastic gradients for variational Bayesian inference (SGVB) of a posterior over model parameters, while retaining parallelizability. This local reparameterization translates uncertainty about global parameters into local noise that is independent across datapoints in the minibatch. Such parameterizations can be trivially parallelized and have variance that is inversely proportional to the minibatch size, generally leading to much faster convergence. Additionally, we explore a connection with dropout: Gaussian dropout objectives correspond to SGVB with local reparameterization, a scale-invariant prior and proportionally fixed posterior variance. Our method allows inference of more flexibly parameterized posteriors; specifically, we propose variational dropout, a generalization of Gaussian dropout where the dropout rates are learned, often leading to better models. The method is demonstrated through several experiments.

Table of Contents

  • 1 Introduction
  • 2 Efficient and Practical Bayesian Inference
  • 2.1 Stochastic Gradient Variational Bayes (SGVB)
  • 2.2 Variance of the SGVB estimator
  • 2.3 Local Reparameterization Trick
  • 3 Variational Dropout
  • 3.1 Variational dropout with independent weight noise
  • 3.2 Variational dropout with correlated weight noise
  • 3.3 Dropout’s scale-invariant prior and variational objective
  • 3.4 Adaptive regularization through optimizing the dropout rate
  • 4 Related Work
  • 5 Experiments
  • 6 Conclusion
  • References
  • A Floating-point numbers and compression
  • B Derivation of dropout’s implicit variational posterior
  • B.1 Gaussian dropout
  • B.2 Weight uncertainty
  • C Negative KL-divergence for the log-uniform prior
  • D Variance reduction of local parameterization compared to separate weight samples
  • E Variance of stochastic gradients for variational dropout with correlated weight noise
  • F Variance of SGVB estimator with minibatches of datapoints without replacement

Knowls

  1. Knowl 1 — Local Reparameterization Trick for Gaussian Distributed Neural Network Weights

    model/method

    In stochastic variational Bayesian inference for a neural network layer computing pre-activations B=AWB = AW, where A∈RM×KA \in \mathbb{R}^{M \times K} is an input matrix containing a minibatch of MM feature vectors of dimension KK, and W∈RK×LW \in \mathbb{R}^{K \times L} is a random weight matrix governed by a fully factorized Gaussian variational posterior qϕ(wi,j)=N(μi,j,σi,j2)q_\phi(w_{i,j}) = \mathcal{N}(\mu_{i,j}, \sigma_{i,j}^2), the local reparameterization trick translates parameter uncertainty into local activation noise by directly sampling the pre-activations BB.

    Conditioned on the input AA, the pre-activations bm,jb_{m,j} for example m∈{1,…,M}m \in \{1,\dots,M\} and output neuron j∈{1,…,L}j \in \{1,\dots,L\} follow independent Gaussian distributions: qϕ(bm,j∣A)=N(γm,j,δm,j)q_\phi(b_{m,j} \mid A) = \mathcal{N}(\gamma_{m,j}, \delta_{m,j}) where γm,j=∑i=1Kam,iμi,jandδm,j=∑i=1Kam,i2σi,j2\gamma_{m,j} = \sum_{i=1}^K a_{m,i}\mu_{i,j} \quad \text{and} \quad \delta_{m,j} = \sum_{i=1}^K a_{m,i}^2 \sigma_{i,j}^2 Sampling the pre-activations is performed elementwise via: bm,j=γm,j+δm,j ζm,j,with ζm,j∼N(0,1)b_{m,j} = \gamma_{m,j} + \sqrt{\delta_{m,j}}\,\zeta_{m,j}, \quad \text{with } \zeta_{m,j} \sim \mathcal{N}(0, 1) where ζ∈RM×L\zeta \in \mathbb{R}^{M \times L} is a matrix of standard normal random variables. This reduces the number of sampled random variables from M×K×LM \times K \times L (for per-example weight matrices) to M×LM \times L, and allows matrix multiplications to be computed using standard, parallelized BLAS operations.

  2. Knowl 2 — Minibatch Variance Scaling of Global versus Local Reparameterization in SGVB

    theoretical result

    For a dataset of NN observations and a minibatch of MM datapoints drawn with replacement, let Li=log⁡p(yi∣xi,w=f(ϵ,ϕ))L_i = \log p(y^i \mid x^i, w = f(\epsilon, \phi)) denote the log-likelihood contribution of the ii-th datapoint in stochastic gradient variational Bayes (SGVB), where ϵ\epsilon represents global model parameter noise. The variance of the Monte Carlo expected log-likelihood estimator LDSGVB(ϕ)=NM∑i=1MLiL_D^{\text{SGVB}}(\phi) = \frac{N}{M} \sum_{i=1}^M L_i is given by: Var[LDSGVB(ϕ)]=N2(1MVar[Li]+M−1MCov[Li,Lj])\text{Var}\left[ L_D^{\text{SGVB}}(\phi) \right] = N^2 \left( \frac{1}{M} \text{Var}[L_i] + \frac{M-1}{M} \text{Cov}[L_i, L_j] \right) where the variance and covariance are evaluated over both the data distribution and the noise distribution p(ϵ)p(\epsilon).

    When a single global weight matrix sample is shared across the minibatch, the covariance term Cov[Li,Lj]\text{Cov}[L_i, L_j] does not decrease with minibatch size MM and is positive on average (Exi,yi,xj,yj[Covϵ[Li,Lj]]=Varϵ[Ex,y[Li]]≥0\mathbb{E}_{x^i,y^i,x^j,y^j}[\text{Cov}_\epsilon[L_i, L_j]] = \text{Var}_\epsilon[\mathbb{E}_{x,y}[L_i]] \ge 0). As a result, the variance of the global estimator is dominated by parameter noise even for moderately large MM.

    Applying the local reparameterization trick makes the noise independent across individual datapoints in the minibatch, ensuring Cov[Li,Lj]=0\text{Cov}[L_i, L_j] = 0 for i≠ji \ne j. Consequently, the variance scales strictly inversely with the minibatch size: Var[LDlocal(ϕ)]=N2MVar[Li]\text{Var}\left[ L_D^{\text{local}}(\phi) \right] = \frac{N^2}{M} \text{Var}[L_i]

  3. Knowl 3 — Variational Dropout Formulations via Independent and Correlated Weight Noise

    model/method

    Variational dropout interprets continuous dropout as variational inference over neural network weights. For a linear layer with input minibatch A∈RM×KA \in \mathbb{R}^{M \times K} and pre-activations B∈RM×LB \in \mathbb{R}^{M \times L}, two distinct variational parameterizations are defined:

    1. Independent Weight Uncertainty (Type B / Post-linear Gaussian Dropout): The variational posterior over weights is fully factorized Gaussian: qϕ(wi,j)=N(θi,j,αi,jθi,j2)q_\phi(w_{i,j}) = \mathcal{N}\left(\theta_{i,j}, \alpha_{i,j} \theta_{i,j}^2\right) where θi,j\theta_{i,j} denotes the mean parameter and αi,j≥0\alpha_{i,j} \ge 0 parameterizes the variance relative to the squared mean. Under the local reparameterization trick, the pre-activations are distributed as: qϕ(bm,j∣A)=N(∑i=1Kam,iθi,j,  ∑i=1Kam,i2αi,jθi,j2)q_\phi(b_{m,j} \mid A) = \mathcal{N}\left(\sum_{i=1}^K a_{m,i}\theta_{i,j}, \; \sum_{i=1}^K a_{m,i}^2 \alpha_{i,j} \theta_{i,j}^2\right)

    2. Correlated Weight Uncertainty (Type A / Pre-linear Gaussian Dropout): The weight matrix rows wi∈R1×Lw_i \in \mathbb{R}^{1 \times L} are formed by multiplying non-stochastic parameters θi\theta_i by scalar stochastic scale variables sis_i: wi=siθi,with qϕ(si)=N(1,αi)w_i = s_i \theta_i, \quad \text{with } q_\phi(s_i) = \mathcal{N}(1, \alpha_i) This retains dependencies between the outputs of the layer across different dimensions and corresponds to multiplying the layer inputs by noise: B=(A∘ξ)θB = (A \circ \xi)\theta, where ξm,i∼N(1,αi)\xi_{m,i} \sim \mathcal{N}(1, \alpha_i).

  4. Knowl 4 — Characterization of the Scale-Invariant Log-Uniform Prior for Dropout Posteriors

    theoretical result

    For a dropout variational posterior qϕ(w)q_\phi(w) where weights decompose into a mean parameter θi\theta_i and a multiplicative noise variable ϵi\epsilon_i with Eqα[ϵi]=1\mathbb{E}_{q_\alpha}[\epsilon_i] = 1 (wi=θiϵiw_i = \theta_i \epsilon_i with ϵi∼qα(ϵi)\epsilon_i \sim q_\alpha(\epsilon_i)), the Kullback-Leibler divergence DKL(qϕ(w)∥p(w))D_{\text{KL}}(q_\phi(w) \parallel p(w)) is independent of θ\theta if and only if the prior p(w)p(w) is the scale-invariant log-uniform prior: p(log⁡∣wi∣)∝cwith p(sign(wi)=1)=p(sign(wi)=−1)=0.5p(\log |w_i|) \propto c \quad \text{with } p(\text{sign}(w_i) = 1) = p(\text{sign}(w_i) = -1) = 0.5 where cc is a normalization constant. Under this prior, the negative KL divergence separates into: −DKL(qϕ(wi) ∥ p(wi))=log⁡(c)+log⁡(0.5)+H(qα(ϵi))−Eqα[log⁡∣ϵi∣]-D_{\text{KL}}\left( q_\phi(w_i) \,\parallel\, p(w_i) \right) = \log(c) + \log(0.5) + \mathcal{H}(q_\alpha(\epsilon_i)) - \mathbb{E}_{q_\alpha}[\log |\epsilon_i|] where H(qα(ϵi))\mathcal{H}(q_\alpha(\epsilon_i)) is the differential entropy of the noise distribution. Because this term does not depend on θ\theta, optimizing the expected log-likelihood with respect to θ\theta is mathematically equivalent to maximizing the variational evidence lower bound (ELBO).

  5. Knowl 5 — Polynomial Approximation of the Negative KL Divergence for Gaussian Dropout

    equation

    For a factorized Gaussian dropout approximate posterior qα(ϵi)=N(1,α)q_\alpha(\epsilon_i) = \mathcal{N}(1, \alpha) and the improper scale-invariant log-uniform prior p(log⁡∣wi∣)∝cp(\log |w_i|) \propto c, the negative Kullback-Leibler divergence −DKL(qϕ(wi)∥p(wi))-D_{\text{KL}}(q_\phi(w_i) \parallel p(w_i)) cannot be evaluated analytically because of the expectation −Eϵ∼N(1,α)[log⁡∣ϵ∣]-\mathbb{E}_{\epsilon \sim \mathcal{N}(1, \alpha)}[\log |\epsilon|]. It is approximated with high precision across the domain of practical dropout rates by a third-order polynomial in α\alpha: −DKL(qϕ(wi) ∥ p(wi))≈C+0.5log⁡(α)+c1α+c2α2+c3α3-D_{\text{KL}}\left( q_\phi(w_i) \,\parallel\, p(w_i) \right) \approx C + 0.5 \log(\alpha) + c_1 \alpha + c_2 \alpha^2 + c_3 \alpha^3 where the constants are: c1=1.16145124,c2=−1.50204118,c3=0.58629921c_1 = 1.16145124, \quad c_2 = -1.50204118, \quad c_3 = 0.58629921 and CC is an arbitrary constant, conventionally chosen so that the divergence is zero at α=1\alpha = 1.

    Alternatively, using the property that −Eqα[log⁡∣ϵi∣]≥0-\mathbb{E}_{q_\alpha}[\log |\epsilon_i|] \ge 0 for all α\alpha, the KL divergence satisfies the lower bound: −DKL(qϕ(wi) ∥ p(wi))≥C+0.5log⁡(α)-D_{\text{KL}}\left( q_\phi(w_i) \,\parallel\, p(w_i) \right) \ge C + 0.5 \log(\alpha) Both the polynomial approximation and the lower bound become exact as log⁡(α)→−∞\log(\alpha) \to -\infty.

  6. Knowl 6 — Adaptive Variational Dropout Optimization with Variance Upper Bound Constraint

    model/method

    In variational dropout, the noise parameters α\alpha (which relate to dropout rates via p=α/(1+α)p = \alpha / (1 + \alpha)) are treated as variational parameters and optimized simultaneously with the mean weights θ\theta by maximizing the evidence lower bound: L(θ,α)=Eqα[LD(θ)]−∑iDKL(qαi(wi) ∥ p(wi))\mathcal{L}(\theta, \alpha) = \mathbb{E}_{q_\alpha}[\mathcal{L}_D(\theta)] - \sum_i D_{\text{KL}}\left( q_{\alpha_i}(w_i) \,\parallel\, p(w_i) \right) Optimization can be performed to learn separate dropout rates per layer, per neuron, or per individual weight parameter.

    To prevent gradients from exhibiting excessive variance and becoming trapped in poor local optima associated with large values of α\alpha, an upper bound constraint is enforced during training: α≤1\alpha \le 1 This constrains the posterior variance to not exceed the square of the posterior mean, corresponding to a maximum dropout rate of p=0.5p = 0.5.

    To prevent underfitting, the KL divergence term in the objective may additionally be downscaled by a constant factor (such as a factor of 3).

  7. Knowl 7 — Analytical Variance Difference Between Local and Sampled Weight Gradient Estimators

    theoretical result

    For a minibatch of size M=1M=1, input vector ama_m, pre-activation bm,jb_{m,j}, and Gaussian weight posterior parameters μi,j\mu_{i,j} and σi,j2\sigma_{i,j}^2, decomposing the variance of the gradient estimator ∂LDSGVB∂σi,j2\frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial \sigma_{i,j}^2} conditionally on bm,jb_{m,j} demonstrates the variance reduction of the local reparameterization trick over explicit weight sampling.

    When drawing separate random weight matrices using noise ϵi,j∼N(0,1)\epsilon_{i,j} \sim \mathcal{N}(0, 1): Ebm,j[Varqϕ,D[∂LDSGVB∂σi,j2 ∣ bm,j]]=Ebm,j[Varqϕ,D[∂LDSGVB∂bm,j ∣ bm,j]q]+Ebm,j[Eqϕ,D[(∂LDSGVB∂bm,j)2 ∣ bm,j]Varqϕ,D[ϵi,j2 ∣ bm,j]am,i24σi,j2]\mathbb{E}_{b_{m,j}}\left[ \text{Var}_{q_\phi, \mathcal{D}}\left[ \frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial \sigma_{i,j}^2} \,\Big|\, b_{m,j} \right] \right] = \mathbb{E}_{b_{m,j}}\left[ \text{Var}_{q_\phi, \mathcal{D}}\left[ \frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial b_{m,j}} \,\Big|\, b_{m,j} \right] q \right] + \mathbb{E}_{b_{m,j}}\left[ \mathbb{E}_{q_\phi, \mathcal{D}}\left[ \left(\frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial b_{m,j}}\right)^2 \,\Big|\, b_{m,j} \right] \text{Var}_{q_\phi, \mathcal{D}}\left[ \epsilon_{i,j}^2 \,\Big|\, b_{m,j} \right] \frac{a_{m,i}^2}{4\sigma_{i,j}^2} \right] where q=(bm,j−γm,j)2am,i44q = \frac{(b_{m,j} - \gamma_{m,j})^2 a_{m,i}^4}{4}.

    When using the local reparameterization trick with noise ζm,j∼N(0,1)\zeta_{m,j} \sim \mathcal{N}(0, 1): Ebm,j[Varqϕ,D[∂LDSGVB∂σi,j2 ∣ bm,j]]=Ebm,j[Varqϕ,D[∂LDSGVB∂bm,j ∣ bm,j]q]\mathbb{E}_{b_{m,j}}\left[ \text{Var}_{q_\phi, \mathcal{D}}\left[ \frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial \sigma_{i,j}^2} \,\Big|\, b_{m,j} \right] \right] = \mathbb{E}_{b_{m,j}}\left[ \text{Var}_{q_\phi, \mathcal{D}}\left[ \frac{\partial \mathcal{L}_D^{\text{SGVB}}}{\partial b_{m,j}} \,\Big|\, b_{m,j} \right] q \right] The second positive term in the weight sampling formulation vanishes under local reparameterization because the scalar variable ζm,j\zeta_{m,j} is uniquely determined given bm,jb_{m,j}, whereas the individual weight noise variables ϵi,j\epsilon_{i,j} are not uniquely determined.

  8. Knowl 8 — Empirical Gradient Variance Across SGVB Estimators on Fully Connected Networks

    data/table

    In a 3-hidden-layer fully connected ReLU neural network trained on MNIST with minibatch size M=1000M=1000 and independent weight noise variational dropout, the average empirical variance of stochastic gradient estimates across different estimator implementations is reported after 10 epochs (test error 3%) and 100 epochs (test error 1.3%):

    Stochastic Gradient Estimator Top layer (10 ep) Top layer (100 ep) Bottom layer (10 ep) Bottom layer (100 ep)
    Local reparameterization 7.8×1037.8 \times 10^3 1.2×1031.2 \times 10^3 1.9×1021.9 \times 10^2 1.1×1021.1 \times 10^2
    Weight sample per data point (slow) 1.4×1041.4 \times 10^4 2.6×1032.6 \times 10^3 4.3×1024.3 \times 10^2 2.5×1022.5 \times 10^2
    Weight sample per minibatch (standard) 4.9×1044.9 \times 10^4 4.3×1034.3 \times 10^3 8.5×1028.5 \times 10^2 3.3×1023.3 \times 10^2
    No dropout noise (minimal var.) 2.8×1032.8 \times 10^3 5.9×1015.9 \times 10^1 1.3×1021.3 \times 10^2 9.0×1009.0 \times 10^0

    The local reparameterization estimator yields the lowest gradient variance among all stochastic estimators across all layers and epochs. Sampling weights per data point reduces variance by approximately 2×2\times relative to sampling weights per minibatch early in optimization, and local reparameterization provides an additional factor-of-two variance reduction over per-data-point weight sampling due to drawing fewer random variables.

  9. Knowl 9 — Computational Speedup and Classification Performance of Variational Dropout

    empirical result

    Empirical benchmarks on MNIST and CIFAR-10 evaluate the wall-clock speed and generalization performance of variational dropout:

    1. Wall-Clock Speedup: On a GPU implementation of a fully connected neural network on MNIST, standard SGVB with per-datapoint weight sampling required 1635 seconds per epoch, whereas the local reparameterization trick required 7.4 seconds per epoch, demonstrating a speedup exceeding 200×200\times.

    2. Classification Performance:

    • On MNIST with 3-hidden-layer fully connected networks across layer widths ranging from 200 to 1200 units, adaptive variational dropout performs equal to or better than fixed-rate binary dropout and Gaussian dropout. Downscaling the KL divergence by a factor of 3 prevents underfitting and achieves the lowest error rates across model sizes.
    • On CIFAR-10 with a convolutional architecture consisting of two convolutional layers (32k32k and 64k64k feature maps, stride 2, softplus activations) followed by two fully connected layers (128k128k units), where k∈[1.0,3.0]k \in [1.0, 3.0] scales the model width, variational dropout achieves lower validation error than binary dropout and Gaussian dropout (reaching ≈23%\approx 23\% validation error at k=3.0k=3.0, compared to ≈31%\approx 31\% for binary and ≈36%\approx 36\% for Gaussian dropout).
    • On small networks, variational dropout automatically infers smaller dropout rates than on large networks, preventing the underfitting caused by fixed dropout hyperparameters.

Coverage note — None was omitted; all primary theoretical, methodological, and empirical contributions of the paper have been covered.

References

  1. 1.Ahn, S., Korattikara, A., and Welling, M. (2012). Bayesian posterior sampling via stochastic gradient Fisher scoring. arXiv preprint arXiv:1206.6380.
  2. 2.Ba, J. and Frey, B. (2013). Adaptive dropout for training deep neural networks. In Advances in Neural Information Processing Systems, pages 3084–3092.
  3. 3.Bayer, J., Karol, M., Korhammer, D., and Van der Smagt, P. (2015). Fast adaptive weight noise. arXiv preprint arXiv:1507.05331.
  4. 4.Bengio, Y. (2013). Estimating or propagating gradients through stochastic neurons. arXiv preprint arXiv:1305.2982.
  5. 5.Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010). Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy), volume 4.
  6. 6.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015). Weight uncertainty in neural networks. arXiv preprint arXiv:1505.05424.
  7. 7.Gal, Y. and Ghahramani, Z. (2015). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. arXiv preprint arXiv:1506.02142.
  8. 8.Graves, A. (2011). Practical variational inference for neural networks. In Advances in Neural Information Processing Systems, pages 2348–2356.
  9. 9.Hernández-Lobato, J. M. and Adams, R. P. (2015). Probabilistic backpropagation for scalable learning of Bayesian neural networks. arXiv preprint arXiv:1502.05336.
  10. 10.Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580.
  11. 11.Hinton, G. E. and Van Camp, D. (1993). Keeping the neural networks simple by minimizing the description length of the weights. In Proceedings of the sixth annual conference on Computational learning theory, pages 5–13. ACM.
  12. 12.Kingma, D. and Ba, J. (2015). Adam: A method for stochastic optimization. Proceedings of the International Conference on Learning Representations 2015.
  13. 13.Kingma, D. P. (2013). Fast gradient-based inference with continuous latent variable models in auxiliary form. arXiv preprint arXiv:1306.0733.
  14. 14.Kingma, D. P. and Welling, M. (2014). Auto-encoding variational Bayes. Proceedings of the 2nd International Conference on Learning Representations.
  15. 15.Maeda, S.-i. (2014). A Bayesian encourages dropout. arXiv preprint arXiv:1412.7003.
  16. 16.Neal, R. M. (1995). Bayesian learning for neural networks. PhD thesis, University of Toronto.
  17. 17.Rezende, D. J., Mohamed, S., and Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), pages 1278–1286.
  18. 18.Robbins, H. and Monro, S. (1951). A stochastic approximation method. The Annals of Mathematical Statistics, 22(3):400–407.
  19. 19.Salimans, T. and Knowles, D. A. (2013). Fixed-form variational posterior approximation through stochastic linear regression. Bayesian Analysis, 8(4).
  20. 20.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958.
  21. 21.Wan, L., Zeiler, M., Zhang, S., Cun, Y. L., and Fergus, R. (2013). Regularization of neural networks using dropconnect. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), pages 1058–1066.
  22. 22.Wang, S. and Manning, C. (2013). Fast dropout training. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), pages 118–126.
  23. 23.Welling, M. and Teh, Y. W. (2011). Bayesian learning via stochastic gradient Langevin dynamics. In Proceedings of the 28th International Conference on Machine Learning (ICML-11), pages 681–688.

Citation

MLA
Kingma, D. P., et al. “Variational Dropout and the Local Reparameterization Trick”. arXiv, 2015, http://arxiv.org/abs/1506.02557v2.
APA
Kingma, D. P., Salimans, T., & Welling, M. (2015). Variational Dropout and the Local Reparameterization Trick. arXiv. http://arxiv.org/abs/1506.02557v2
Chicago
Kingma, D. P., T. Salimans, and M. Welling. 2015. “Variational Dropout and the Local Reparameterization Trick”. arXiv. http://arxiv.org/abs/1506.02557v2.
Harvard
Kingma, D.P., Salimans, T. and Welling, M. (2015) “Variational Dropout and the Local Reparameterization Trick”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1506.02557v2.
Vancouver
1. Kingma DP, Salimans T, Welling M (2015) Variational Dropout and the Local Reparameterization Trick. arXiv

BibTeX

@article{kingma2015variational,
  title = {Variational Dropout and the Local Reparameterization Trick},
  author = {Kingma, Diederik P. and Salimans, Tim and Welling, Max},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1506.02557v2},
  eprint = {1506.02557}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors