The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

Chris J. MaddisonAndriy MnihYee Whye Teh

article2016ICLR3,005 citations

Proposes a continuous relaxation of discrete random variables that extends the reparameterization trick to discrete stochastic nodes, enabling efficient end-to-end gradient-based training of neural networks with discrete latent variables.

Listen

The paper introduces the Concrete distribution, a new family of continuous probability distributions on the simplex that serve as a practical relaxation of discrete random variables. The work addresses a core obstacle in training large neural networks that incorporate discrete stochastic nodes: standard automatic differentiation libraries cannot propagate low-variance gradients through discrete states, forcing practitioners to rely on high-variance score-function estimators or to forgo discrete variables altogether.

The authors set out to create a reparameterizable continuous surrogate that preserves the essential properties of any discrete distribution while allowing unbiased gradients with respect to a relaxed objective. Their approach begins with the Gumbel-Max trick for sampling discrete variables and replaces the non-differentiable argmax operation with a temperature-controlled softmax. The resulting Concrete random variable has a simple closed-form density, reduces exactly to the original discrete distribution in the zero-temperature limit, and can be sampled by adding fixed Gumbel noise to logits and applying a softmax. Experiments evaluated the method on density estimation and structured output prediction tasks using neural networks with hundreds of latent discrete nodes, trained on the MNIST and Omniglot datasets and compared against strong score-function baselines (VIMCO and NVIL).

The central empirical result is that Concrete relaxations produce competitive or superior test negative log-likelihoods on both tasks, often outperforming the baselines for non-linear models while requiring no custom gradient code. Linear models favored the score-function estimators, but the gap narrowed or reversed as model depth increased. Performance proved sensitive to the choice of temperature during training, with distinct temperatures for prior and posterior nodes yielding the best results; no annealing schedule was required. At test time the original discrete graph is evaluated, so the final model incurs no approximation error.

These findings matter because they let researchers incorporate discrete stochastic units—attractive for interpretability, sparsity, and structured reasoning—into end-to-end differentiable pipelines without sacrificing ease of implementation or incurring prohibitive variance. The method therefore broadens the set of architectures that can be trained at scale with off-the-shelf gradient descent.

The authors recommend replacing each discrete node with a Concrete node of fixed temperature during training, relaxing any log-probability terms that appear in the objective, and reverting to the discrete graph for evaluation. When further gains are desired, modest additional tuning of the two temperatures or exploration of temperature annealing schedules is likely to help. The main limitations are that the gradients remain biased with respect to the true discrete objective, that temperature selection remains an empirical hyper-parameter, and that the reported gains are demonstrated only on the two tasks and two datasets examined. Overall the evidence supports adoption of Concrete relaxations wherever discrete stochastic nodes are otherwise difficult to optimize.

arXiv: 1611.00712
  • Paper: An Introduction to Variational Autoencoders, Diederik P. Kingma et al. (2019). Reading the foundational VAE and reparameterization work first provides the core variational inference framework and notation that the Concrete distribution builds upon.
  • Paper: Categorical Reparameterization with Gumbel-Softmax, Eric Jang et al. (2017). Reviewing the Gumbel-Softmax method provides the primary alternative categorical relaxation technique, which makes understanding the Concrete distribution's formulation and design choices much clearer.
Cover for The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

Abstract

The reparameterization trick enables optimizing large scale stochastic computation graphs via gradient descent. The essence of the trick is to refactor each stochastic node into a differentiable function of its parameters and a random variable with fixed distribution. After refactoring, the gradients of the loss propagated by the chain rule through the graph are low variance unbiased estimators of the gradients of the expected loss. While many continuous random variables have such reparameterizations, discrete random variables lack useful reparameterizations due to the discontinuous nature of discrete states. In this work we introduce Concrete random variables---continuous relaxations of discrete random variables. The Concrete distribution is a new family of distributions with closed form densities and a simple reparameterization. Whenever a discrete stochastic node of a computation graph can be refactored into a one-hot bit representation that is treated continuously, Concrete stochastic nodes can be used with automatic differentiation to produce low-variance biased gradients of objectives (including objectives that depend on the log-probability of latent stochastic nodes) on the corresponding discrete graph. We demonstrate the effectiveness of Concrete relaxations on density estimation and structured prediction tasks using neural networks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Optimizing Stochastic Computation Graphs
  • 2.2 Score Function Estimators
  • 2.3 Reparameterization Trick
  • 2.4 Application: Variational Training of Latent Variable Models
  • 3 The Concrete Distribution
  • 3.1 Discrete Random Variables and the Gumbel-Max Trick
  • 3.2 Concrete Random Variables
  • 3.3 Concrete Relaxations
  • 4 Related Work
  • 5 Experiments
  • 5.1 Protocol
  • 5.2 Density Estimation
  • 5.3 Structured Output Prediction
  • 6 Conclusion
  • References
  • A Proof of Proposition
  • B The Binary Special Case
  • C Using Concrete Relaxations
  • C.1 The Basic Problem
  • C.2 What you might relax and why
  • C.3 Which random variable to treat as the stochastic node
  • C.3.1 nn-ary Concrete
  • C.3.2 Binary Concrete
  • C.4 Choosing the temperature
  • D Experimental Details
  • D.1 — vs ∼\sim
  • D.2 nn-ary layers
  • D.3 Bias Initialization
  • D.4 Centering
  • D.5 Hyperparameter Selection
  • E Extra Results
  • F Cheat Sheet

Knowls

  1. Knowl 1 — Concrete Distribution and Gumbel-Softmax Reparameterization

    definition

    Let α=(α1,…,αn)∈(0,∞)n\alpha = (\alpha_1, \ldots, \alpha_n) \in (0, \infty)^n be unnormalized location parameters and let λ∈(0,∞)\lambda \in (0, \infty) be the temperature parameter. A continuous random vector X=(X1,…,Xn)X = (X_1, \ldots, X_n) taking values on the open standard (n−1)(n-1)-simplex Δn−1={x∈(0,1)n∣∑k=1nxk=1}\Delta^{n-1} = \left\{x \in (0, 1)^n \mid \sum_{k=1}^n x_k = 1\right\} follows a Concrete distribution, denoted X∼Concrete(α,λ)X \sim \text{Concrete}(\alpha, \lambda), if its probability density function with respect to the Hausdorff measure on the simplex is:

    pα,λ(x)=(n−1)! λn−1∏k=1n(αkxk−λ−1∑i=1nαixi−λ)p_{\alpha, \lambda}(x) = (n - 1)! \, \lambda^{n-1} \prod_{k=1}^n \left( \frac{\alpha_k x_k^{-\lambda-1}}{\sum_{i=1}^n \alpha_i x_i^{-\lambda}} \right)

    A sample X∼Concrete(α,λ)X \sim \text{Concrete}(\alpha, \lambda) can be generated via the continuous, differentiable reparameterization:

    Xk=exp⁡((log⁡αk+Gk)/λ)∑i=1nexp⁡((log⁡αi+Gi)/λ),k∈{1,…,n}X_k = \frac{\exp\left((\log \alpha_k + G_k)/\lambda\right)}{\sum_{i=1}^n \exp\left((\log \alpha_i + G_i)/\lambda\right)}, \quad k \in \{1, \ldots, n\}

    where G1,…,GnG_1, \ldots, G_n are independent standard Gumbel random variables sampled via Gk=−log⁡(−log⁡Uk)G_k = -\log(-\log U_k) with Uk∼Uniform(0,1)U_k \sim \text{Uniform}(0, 1).

  2. Knowl 2 — Rounding, Zero-Temperature Limit, and Log-Convexity of Concrete Random Variables

    theoretical result

    Let X∼Concrete(α,λ)X \sim \text{Concrete}(\alpha, \lambda) be an nn-ary Concrete random variable on the simplex Δn−1\Delta^{n-1} with location parameters α∈(0,∞)n\alpha \in (0, \infty)^n and temperature λ∈(0,∞)\lambda \in (0, \infty). The distribution satisfies three fundamental structural properties:

    1. Rounding Property: Discretizing XX by selecting the index of its maximal element recovers the categorical distribution parameterized by α\alpha:
    P(Xk>Xi for all i≠k)=αk∑i=1nαi\mathbb{P}\left(X_k > X_i \text{ for all } i \ne k\right) = \frac{\alpha_k}{\sum_{i=1}^n \alpha_i}
    1. Zero-Temperature Limit: In the zero-temperature limit λ→0\lambda \to 0, XX converges almost surely to a one-hot discrete random variable:
    P(lim⁡λ→0Xk=1)=αk∑i=1nαi\mathbb{P}\left(\lim_{\lambda \to 0} X_k = 1\right) = \frac{\alpha_k}{\sum_{i=1}^n \alpha_i}
    1. Log-Convexity and Mode Absence: If λ≤(n−1)−1\lambda \le (n - 1)^{-1}, the density function pα,λ(x)p_{\alpha, \lambda}(x) is log-convex in xx. Consequently, under this temperature condition, the distribution has no local maxima in the interior of the simplex for any location parameter vector α\alpha.
  3. Knowl 3 — Variational Training of Latent Variable Models via Concrete Relaxations

    model/method

    In a variational autoencoder with discrete latent variables D∈{0,1}nD \in \{0, 1\}^n (where ∑k=1nDk=1\sum_{k=1}^n D_k = 1), prior Pa(d)P_a(d), likelihood pθ(x∣d)p_\theta(x \mid d), and approximate posterior Qα(d∣x)Q_\alpha(d \mid x), the evidence lower bound (ELBO) on log⁡p(x)\log p(x) is:

    L1(θ,a,α)=ED∼Qα(d∣x)[log⁡pθ(x∣D)+log⁡Pa(D)Qα(D∣x)]\mathcal{L}_1(\theta, a, \alpha) = \mathbb{E}_{D \sim Q_\alpha(d \mid x)} \left[ \log p_\theta(x \mid D) + \log \frac{P_a(D)}{Q_\alpha(D \mid x)} \right]

    To optimize this objective via standard automatic differentiation, the discrete posterior sampling is replaced with a Concrete relaxation Z∼Concrete(α(x),λ1)Z \sim \text{Concrete}(\alpha(x), \lambda_1) having density qα,λ1(z∣x)q_{\alpha, \lambda_1}(z \mid x), and the discrete prior is relaxed to a Concrete prior density pa,λ2(z)p_{a, \lambda_2}(z) with location aa and temperature λ2\lambda_2. The relaxed objective is:

    L1relax(θ,a,α)=EZ∼qα,λ1(z∣x)[log⁡pθ(x∣Z)+log⁡pa,λ2(Z)qα,λ1(Z∣x)]\mathcal{L}_1^{\text{relax}}(\theta, a, \alpha) = \mathbb{E}_{Z \sim q_{\alpha, \lambda_1}(z \mid x)} \left[ \log p_\theta(x \mid Z) + \log \frac{p_{a, \lambda_2}(Z)}{q_{\alpha, \lambda_1}(Z \mid x)} \right]

    This relaxed objective is a valid variational lower bound on the marginal likelihood log⁡∫pθ(x∣z)pa,λ2(z) dz\log \int p_\theta(x \mid z) p_{a, \lambda_2}(z) \, dz of the continuous relaxed model. Gradients with respect to all parameters (including α\alpha) are computed by backpropagating through the reparameterized samples ZZ. At test time, the model is evaluated using the original discrete graph.

  4. Knowl 4 — Binary Concrete Distribution

    definition

    For binary discrete states, the Concrete distribution is defined as a univariate continuous distribution over the open interval (0,1)(0, 1). A random variable X∈(0,1)X \in (0, 1) follows a Binary Concrete distribution, denoted X∼BinConcrete(α,λ)X \sim \text{BinConcrete}(\alpha, \lambda) with location parameter α∈(0,∞)\alpha \in (0, \infty) and temperature λ∈(0,∞)\lambda \in (0, \infty), if its density is:

    pα,λ(x)=λαx−λ−1(1−x)−λ−1(αx−λ+(1−x)−λ)2p_{\alpha,\lambda}(x) = \frac{\lambda \alpha x^{-\lambda-1}(1 - x)^{-\lambda-1}}{\left(\alpha x^{-\lambda} + (1 - x)^{-\lambda}\right)^2}

    It is reparameterized by passing perturbed logits through the logistic sigmoid function σ(t)=(1+exp⁡(−t))−1\sigma(t) = (1 + \exp(-t))^{-1}:

    X=σ(log⁡α+Lλ)=11+exp⁡(−log⁡α+Lλ)X = \sigma\left(\frac{\log \alpha + L}{\lambda}\right) = \frac{1}{1 + \exp\left(-\frac{\log \alpha + L}{\lambda}\right)}

    where L∼LogisticL \sim \text{Logistic} is a standard logistic random variable generated via L=log⁡U−log⁡(1−U)L = \log U - \log(1 - U) with U∼Uniform(0,1)U \sim \text{Uniform}(0, 1).

    The Binary Concrete distribution satisfies:

    • Rounding: P(X>0.5)=α1+α\mathbb{P}(X > 0.5) = \frac{\alpha}{1 + \alpha}
    • Zero-Temperature Limit: P(lim⁡λ→0X=1)=α1+α\mathbb{P}(\lim_{\lambda \to 0} X = 1) = \frac{\alpha}{1 + \alpha}
    • Log-Convexity: For λ≤1\lambda \le 1, pα,λ(x)p_{\alpha,\lambda}(x) is log-convex in xx on (0,1)(0, 1).
  5. Knowl 5 — ExpConcrete Distribution for Numerically Stable Variational Inference

    model/method

    To prevent numerical underflow when evaluating log-densities in variational bounds, Concrete random variables can be evaluated in log-space as ExpConcrete variables Y∈RnY \in \mathbb{R}^n, defined such that X=exp⁡(Y)∼Concrete(α,λ)X = \exp(Y) \sim \text{Concrete}(\alpha, \lambda) and LSEk=1n{Yk}=0\text{LSE}_{k=1}^n \{Y_k\} = 0, where LSEi=1n{vi}=log⁡∑i=1nexp⁡(vi)\text{LSE}_{i=1}^n \{v_i\} = \log \sum_{i=1}^n \exp(v_i).

    An ExpConcrete variable Y∼ExpConcrete(α,λ)Y \sim \text{ExpConcrete}(\alpha, \lambda) is sampled via:

    Yk=log⁡αk+Gkλ−LSEi=1n{log⁡αi+Giλ}Y_k = \frac{\log \alpha_k + G_k}{\lambda} - \text{LSE}_{i=1}^n \left\{ \frac{\log \alpha_i + G_i}{\lambda} \right\}

    where Gk∼Gumbel(0,1)G_k \sim \text{Gumbel}(0, 1) i.i.d. Its log-density log⁡κα,λ(y)\log \kappa_{\alpha, \lambda}(y) on {y∈(−∞,0)n∣LSEk=1n{yk}=0}\left\{y \in (-\infty, 0)^n \mid \text{LSE}_{k=1}^n \{y_k\} = 0\right\} is:

    log⁡κα,λ(y)=log⁡((n−1)!)+(n−1)log⁡λ+∑k=1n(log⁡αk−λyk)−n⋅LSEk=1n{log⁡αk−λyk}\log \kappa_{\alpha, \lambda}(y) = \log((n - 1)!) + (n - 1) \log \lambda + \sum_{k=1}^n (\log \alpha_k - \lambda y_k) - n \cdot \text{LSE}_{k=1}^n \{\log \alpha_k - \lambda y_k\}

    Because the exponential map Y↦XY \mapsto X is an invertible transformation, the Kullback-Leibler divergence between two ExpConcrete distributions is identical to the KL divergence between the corresponding Concrete distributions.

  6. Knowl 6 — Temperature Trade-Off and Integrality Gap in Concrete Optimization

    limitation

    Training discrete computation graphs with Concrete relaxations involves a fundamental trade-off governed by the temperature λ\lambda:

    • Excessive Temperature (λ\lambda too large): The relaxation allocates significant probability mass to the interior of the simplex rather than its vertices. Latent units can exploit this continuous interior capacity to communicate substantially more than log⁡2n\log_2 n bits of information. When the discrete graph is evaluated at test time, this causes a severe performance drop termed the integrality gap.
    • Low Temperature (λ→0\lambda \to 0): The sample distribution concentrates onto discrete states, but the softmax derivative approaches a step function, leading to vanishing or high-variance gradients.

    Empirically, using distinct fixed temperatures for the variational posterior (λ1\lambda_1) and prior (λ2\lambda_2) yielded optimal performance without annealing:

    • Binary (n=2n=2): posterior λ1=2/3\lambda_1 = 2/3, prior λ2=1/2\lambda_2 = 1/2
    • 4-ary (n=4n=4): posterior λ1=1\lambda_1 = 1, prior λ2=2/3\lambda_2 = 2/3
    • 8-ary (n=8n=8): posterior λ1=2/3\lambda_1 = 2/3, prior λ2=2/5\lambda_2 = 2/5

    While the theoretical guarantee against interior modes requires λ≤(n−1)−1\lambda \le (n - 1)^{-1}, higher temperatures are practically viable for larger nn because the random normalizer ∑k=1nexp⁡((log⁡αk+Gk)/λ)\sum_{k=1}^n \exp((\log \alpha_k + G_k)/\lambda) grows with nn, naturally concentrating mass toward vertices.

  7. Knowl 7 — Density Estimation Performance on MNIST and Omniglot

    empirical result

    Discrete variational autoencoders trained via Concrete relaxations were compared against score function estimators (NVIL for sample count m=1m=1, VIMCO for m∈{5,50}m \in \{5, 50\}) on binarized MNIST and Omniglot. Architectures had 1 or 2 latent stochastic layers with 200 units, using either linear (−-) or non-linear (∼\sim, two tanh⁡\tanh layers) conditional mappings. Performance was evaluated using test negative log-likelihood (NLL, lower is better) on the discrete graph via L50000\mathcal{L}_{50000}:

    MNIST Test NLL MNIST Train NLL Omniglot Test NLL Omniglot Train NLL
    Model mm Concrete VIMCO Concrete VIMCO Concrete VIMCO Concrete VIMCO
    (200H–784V) 1 107.3 104.4 107.5 104.2 118.7 115.7 117.0 112.2
    5 104.9 101.9 104.9 101.5 118.0 113.5 115.8 110.8
    50 104.3 98.8 104.2 98.3 118.9 113.0 115.8 110.0
    (200H–200H–784V) 1 102.1 92.9 102.3 91.7 116.3 109.2 114.4 104.8
    5 99.9 91.7 100.0 90.8 116.0 107.5 113.5 103.6
    50 99.5 90.7 99.4 89.7 117.0 108.1 113.9 103.6
    (200H∼\sim784V) 1 92.1 93.8 91.2 91.5 108.4 116.4 103.6 110.3
    5 89.5 91.4 88.1 88.6 107.5 118.2 101.4 102.3
    50 88.5 89.3 86.4 86.5 108.1 116.0 100.5 100.8
    (200H∼\sim200H∼\sim784V) 1 87.9 88.4 86.5 85.8 105.9 111.7 100.2 105.7
    5 86.3 86.4 84.1 82.5 105.8 108.2 98.6 101.1
    50 85.7 85.5 83.1 81.8 106.8 113.2 97.5 95.2

    VIMCO outperformed Concrete relaxations on linear generative models, whereas Concrete relaxations systematically outperformed VIMCO on non-linear models (e.g., Omniglot non-linear 1-layer test NLL of 108.1 for Concrete vs 116.0 for VIMCO at m=50m=50).

  8. Knowl 8 — Structured Output Prediction on MNIST with Concrete Units

    empirical result

    Structured prediction was evaluated by predicting the bottom half of an MNIST image (x1∈{0,1}392x_1 \in \{0, 1\}^{392}) given its top half (x2∈{0,1}392x_2 \in \{0, 1\}^{392}) using latent stochastic layers. Models were trained on the multi-sample objective LmSP(θ,ϕ)=EZi∼pϕ(z∣x2)[log⁡1m∑i=1mpθ(x1∣Zi)]\mathcal{L}_m^{\text{SP}}(\theta, \phi) = \mathbb{E}_{Z_i \sim p_\phi(z \mid x_2)} \left[ \log \frac{1}{m} \sum_{i=1}^m p_\theta(x_1 \mid Z_i) \right] and evaluated on L100SP\mathcal{L}_{100}^{\text{SP}}:

    Test NLL Train NLL
    Model mm Concrete VIMCO Concrete VIMCO
    (392V–240H–240H–392V) 1 58.5 61.4 54.2 59.3
    5 54.3 54.5 49.2 52.7
    50 53.4 51.8 48.2 49.6
    (392V–240H–240H–240H–392V) 1 56.3 59.7 51.6 58.4
    5 52.7 53.5 46.9 51.6
    50 52.0 50.2 45.9 47.9

    Concrete relaxations outperformed VIMCO for m∈{1,5}m \in \{1, 5\}. Furthermore, on the 3-layer architecture with m=1m=1, increasing unit arity improved performance under Concrete relaxation: 4-ary units achieved test/train NLL of 55.4 / 46.0, and 8-ary units achieved 54.7 / 44.8.

  9. Knowl 9 — Hypercube State Mapping for Multi-Category Discrete Layers

    model/method

    To construct discrete neural network layers of nn-ary units, states are mapped to the vertices of a hypercube {−1,1}log⁡2n\{-1, 1\}^{\log_2 n}. Let C∈{−1,1}log⁡2n×nC \in \{-1, 1\}^{\log_2 n \times n} denote the matrix whose columns enumerate the nn distinct corner vectors of the hypercube.

    • Discrete Layer Evaluation: For a discrete categorical variable D∼Discrete(α)D \sim \text{Discrete}(\alpha) represented as a one-hot column vector D∈{0,1}nD \in \{0, 1\}^n, the layer's output is Y=CD∈{−1,1}log⁡2nY = C D \in \{-1, 1\}^{\log_2 n}.
    • Concrete Continuous Relaxation: During training, DD is replaced by X∼Concrete(α,λ)X \sim \text{Concrete}(\alpha, \lambda), yielding relaxed continuous activations Y~=CX∈[−1,1]log⁡2n\tilde{Y} = C X \in [-1, 1]^{\log_2 n}.

    For binary units (n=2n = 2, log⁡2n=1\log_2 n = 1), this mapping simplifies to:

    Y~=2 σ(log⁡α+log⁡U−log⁡(1−U)λ)−1∈(−1,1)\tilde{Y} = 2 \, \sigma\left(\frac{\log \alpha + \log U - \log(1 - U)}{\lambda}\right) - 1 \in (-1, 1)

    where U∼Uniform(0,1)U \sim \text{Uniform}(0, 1) and σ(t)=(1+exp⁡(−t))−1\sigma(t) = (1 + \exp(-t))^{-1}.

Coverage note — Omitted the step-by-step change-of-variables Jacobian determinant derivation of Proposition 1 and background reviews on Dirichlet and Logistic-Normal distributions, as they serve solely as intermediate steps or prior literature.

References

  1. 1.Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah,  Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Watten-  berg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL http://tensorflow.org/. Software available from tensorflow.org.
  2. 2.J Aitchison. A general class of distributions on the simplex. Journal of the Royal Statistical Society. Series B (Methodological), pp. 136–146, 1985.
  3. 3.J Atchison and Sheng M Shen. Logistic-normal distributions: Some properties and uses. Biometrika, 67(2):261–272, 1980.
  4. 4.Yoshua Bengio, Nicholas Leonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  5. 5.David Blei and John Lafferty. Correlated topic models. 2006.
  6. 6.Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov. Importance weighted autoencoders. ICLR, 2016.
  7. 7.Robert J Connor and James E Mosimann. Concepts of independence for proportions with a generalization of the dirichlet distribution. Journal of the American Statistical Association, 64(325): 194–206, 1969.
  8. 8.Stefano Favaro, Georgia Hadjicharalambous, and Igor Prunster. On a class of distributions on the simplex. Journal of Statistical Planning and Inference, 141(9):2987 – 3004, 2011.
  9. 9.Brendan Frey. Continuous sigmoidal belief networks trained using slice sampling. In NIPS, 1997.
  10. 10.Michael C Fu. Gradient estimation. Handbooks in operations research and management science, 13:575–616, 2006.
  11. 11.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Aistats, volume 9, pp. 249–256, 2010.
  12. 12.Peter W Glynn. Likelihood ratio gradient estimation for stochastic systems. Communications of the ACM, 33(10):75–84, 1990.
  13. 13.Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka GrabskaBarwinska, Sergio G  omez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou,  et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538 (7626):471–476, 2016.
  14. 14.Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter. Variance reduction techniques for gradient estimates in reinforcement learning. JMLR, 5, 2004.
  15. 15.Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom. Learning to transduce with unbounded memory. In Advances in Neural Information Processing Systems, pp. 1828–1836, 2015.
  16. 16.Karol Gregor, Ivo Danihelka, Andriy Mnih, Charles Blundell, and Daan Wierstra. Deep autoregressive networks. arXiv preprint arXiv:1310.8499, 2013.
  17. 17.Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra. Draw: A recurrent neural network for image generation. arXiv preprint arXiv:1502.04623, 2015.
  18. 18.Shixiang Gu, Sergey Levine, Ilya Sutskever, and Andriy Mnih. MuProp: Unbiased backpropagation for stochastic neural networks. ICLR, 2016.
  19. 19.Emil Julius Gumbel. Statistical theory of extreme values and some practical applications: a series of lectures. Number 33. US Govt. Print. Office, 1954.
  20. 20.Tamir Hazan and Tommi Jaakkola. On the partition function and random maximum a-posteriori perturbations. In ICML, 2012.
  21. 21.Tamir Hazan, George Papandreou, and Daniel Tarlow. Perturbation, Optimization, and Statistics. MIT Press, 2016.
  22. 22.Matthew D Hoffman, David M Blei, Chong Wang, and John William Paisley. Stochastic variational inference. JMLR, 14(1):1303–1347, 2013.
  23. 23.E. Jang, S. Gu, and B. Poole. Categorical Reparameterization with Gumbel-Softmax. ArXiv e-prints, November 2016.
  24. 24.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  25. 25.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  26. 26.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. ICLR, 2014.
  27. 27.Tomas Ko ȇ cisk ȇ y, G  abor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, and  Karl Moritz Hermann. Semantic parsing with semi-supervised sequential autoencoders. In EMNLP, 2016.
  28. 28.R. Duncan Luce. Individual Choice Behavior: A Theoretical Analysis. New York: Wiley, 1959.
  29. 29.Chris J Maddison. A Poisson process model for Monte Carlo. In Tamir Hazan, George Papandreou, and Daniel Tarlow (eds.), Perturbation, Optimization, and Statistics, chapter 7. MIT Press, 2016.
  30. 30.Chris J Maddison, Daniel Tarlow, and Tom Minka. A∗ Sampling. In NIPS, 2014.
  31. 31.Andriy Mnih and Karol Gregor. Neural variational inference and learning in belief networks. In ICML, 2014.
  32. 32.Andriy Mnih and Danilo Jimenez Rezende. Variational inference for monte carlo objectives. In ICML, 2016.
  33. 33.Volodymyr Mnih, Nicolas Heess, Alex Graves, and koray kavukcuoglu. Recurrent Models of Visual Attention. In NIPS, 2014.
  34. 34.Christian A Naesseth, Francisco JR Ruiz, Scott W Linderman, and David M Blei. Rejection sampling variational inference. arXiv preprint arXiv:1610.05683, 2016.
  35. 35.John William Paisley, David M. Blei, and Michael I. Jordan. Variational bayesian inference with stochastic search. In ICML, 2012.
  36. 36.George Papandreou and Alan L Yuille. Perturb-and-map random fields: Using discrete optimization to learn and sample from energy models. In ICCV, 2011.
  37. 37.Tapani Raiko, Mathias Berglund, Guillaume Alain, and Laurent Dinh. Techniques for learning binary stochastic feedforward neural networks. arXiv preprint arXiv:1406.2989, 2014.
  38. 38.Rajesh Ranganath, Sean Gerrish, and David M. Blei. Black box variational inference. In AISTATS, 2014.
  39. 39.William S Rayens and Cidambi Srinivasan. Dependence properties of generalized liouville distributions on the simplex. Journal of the American Statistical Association, 89(428):1465–1470, 1994.
  40. 40.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In ICML, 2014.
  41. 41.Francisco JR Ruiz, Michalis K Titsias, and David M Blei. The generalized reparameterization gradient. arXiv preprint arXiv:1610.02287, 2016.
  42. 42.Ruslan Salakhutdinov and Geoffrey Hinton. Semantic hashing. International Journal of Approximate Reasoning, 50(7):969–978, 2009.
  43. 43.Ruslan Salakhutdinov and Iain Murray. On the quantitative analysis of deep belief networks. In ICML, 2008.
  44. 44.John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel. Gradient estimation using stochastic computation graphs. In NIPS, 2015.
  45. 45.Theano Development Team. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints, abs/1605.02688, May 2016. URL http://arxiv.org/abs/ 1605.02688.
  46. 46.Michalis Titsias and Miguel Lazaro-Gredilla. Doubly stochastic variational bayes for non-conjugate inference. In Tony Jebara and Eric P. Xing (eds.), ICML, 2014.
  47. 47.Michalis Titsias and Miguel Lazaro-Gredilla. Local expectation gradients for black box variational inference. In NIPS, 2015.
  48. 48.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3-4):229–256, 1992.
  49. 49.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015.
  50. 50.John I Yellott. The relationship between luce’s choice axiom, thurstone’s theory of comparative judgment, and the double exponential distribution. Journal of Mathematical Psychology, 15(2): 109–144, 1977.

Citation

MLA
Maddison, C. J., et al. “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”. arXiv, 2016, https://doi.org/10.48550/arxiv.1611.00712.
APA
Maddison, C. J., Mnih, A., & Teh, Y. W. (2016). The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables. arXiv. https://doi.org/10.48550/arxiv.1611.00712
Chicago
Maddison, C. J., A. Mnih, and Y. W. Teh. 2016. “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1611.00712.
Harvard
Maddison, C.J., Mnih, A. and Teh, Y.W. (2016) “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”. arXiv. Available at: https://doi.org/10.48550/arxiv.1611.00712.
Vancouver
1. Maddison CJ, Mnih A, Teh YW (2016) The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables. https://doi.org/10.48550/arxiv.1611.00712

BibTeX

@misc{https://doi.org/10.48550/arxiv.1611.00712,
  doi = {10.48550/ARXIV.1611.00712},
  url = {https://arxiv.org/abs/1611.00712},
  author = {Maddison, Chris J. and Mnih, Andriy and Teh, Yee Whye},
  keywords = {Machine Learning (cs.LG), Machine Learning (stat.ML), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables},
  publisher = {arXiv},
  year = {2016},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission