Normalizing Flows for Probabilistic Modeling and Inference

George PapamakariosEric NalisnickDanilo Jimenez RezendeShakir MohamedBalaji Lakshminarayanan

article2019JMLR2,507 citations

Unifies the foundational principles, expressive capacity, and computational trade-offs of normalizing flows to guide their application in generative modeling and approximate inference.

Listen

Normalizing flows address the long-standing need in statistics and machine learning for flexible probability distributions that can accurately capture complex, high-dimensional data. Simple base distributions such as Gaussians often fail to model real-world processes, and the article shows that repeated invertible transformations can produce distributions of arbitrary complexity while preserving tractable density evaluation and sampling.

The review sets out to synthesize the maturing literature on normalizing flows through the lens of probabilistic modeling and inference. It establishes core principles of flow design, examines expressive power and computational trade-offs, relates flows to more general probability transformations, and surveys practical uses in generative modeling, approximate inference, and supervised learning.

The authors review both finite compositions of simple maps and continuous-time formulations defined by ordinary differential equations. They demonstrate that autoregressive, linear, residual, and coupling-based constructions each offer distinct balances between flexibility and speed. Under mild regularity conditions, flows can represent any target density exactly, and the change-of-variables formula yields exact likelihoods when the transformation is invertible and differentiable.

Key results include the universality of triangular maps, efficient Jacobian-determinant calculations for structured transformations, and extensions to discrete variables and Riemannian manifolds via piecewise-invertible or embedding maps. Experiments and theoretical arguments show that flows improve density estimation on images and audio, stabilize variational inference, and support likelihood-free parameter estimation.

These capabilities matter because exact likelihoods and fast sampling enable better-specified models, more reliable uncertainty quantification, and scalable inference in scientific simulators. The work also highlights that flows subsume and extend autoregressive models while opening new routes for hybrid generative-discriminative learning.

Practitioners should adopt coupling-based flows when both sampling and density evaluation must be fast, and masked autoregressive flows when only one direction is required. Further gains are possible by interleaving linear flows for dimension mixing and by using continuous-time formulations when memory is limited. Additional research is needed on discrete and manifold-valued data, on finite-sample approximation bounds, and on automatic selection of flow depth and architecture.

The review draws on an extensive body of published work and provides consistent derivations, yet it necessarily omits the most recent empirical benchmarks. Readers should therefore treat reported performance numbers as indicative rather than definitive and should validate computational trade-offs on their own data and hardware.

arXiv: 1912.02762
  • Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). This seminal paper introduces normalizing flows to machine learning by parameterizing rich posterior distributions with sequences of invertible transformations.
  • Paper: NICE: Non-linear Independent Components Estimation, Laurent Dinh et al. (2014). This foundational work establishes coupling layers with triangular Jacobians to enable tractable exact likelihood evaluation and invertible transformations.
  • Paper: Density estimation using Real NVP, Laurent Dinh et al. (2016). This paper expands coupling-based normalizing flows using multi-scale architectures and affine transformations for high-dimensional density estimation.
  • Paper: Improving Variational Inference with Inverse Autoregressive Flow, Diederik P. Kingma et al. (2016). This work introduces inverse autoregressive flows, providing the theoretical and practical basis for autoregressive flow architectures covered in the review.
  • Paper: Neural Ordinary Differential Equations, Ricky T. Q. Chen et al. (2018). This foundational paper defines continuous-time normalizing flows parameterised by ordinary differential equations solved via adjoint sensitivity methods.
  • Paper: Glow: Generative Flow with Invertible 1x1 Convolutions, Diederik P. Kingma et al. (2018). This work introduces invertible 1x1 convolutions within coupling architectures, providing a core linear flow design surveyed in the review.
  • Paper: An Introduction to Variational Autoencoders, Diederik P. Kingma et al. (2019). This tutorial supplies the foundational theory of variational autoencoders and amortized variational inference that normalizing flows are frequently designed to enhance.
  • Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). This comprehensive review lays out the core principles and statistical foundations of variational inference necessary to understand normalizing flows' role in approximate inference.
Cover for Normalizing Flows for Probabilistic Modeling and Inference

Abstract

Normalizing flows provide a general mechanism for defining expressive probability distributions, only requiring the specification of a (usually simple) base distribution and a series of bijective transformations. There has been much recent work on normalizing flows, ranging from improving their expressive power to expanding their application. We believe the field has now matured and is in need of a unified perspective. In this review, we attempt to provide such a perspective by describing flows through the lens of probabilistic modeling and inference. We place special emphasis on the fundamental principles of flow design, and discuss foundational topics such as expressive power and computational trade-offs. We also broaden the conceptual framing of flows by relating them to more general probability transformations. Lastly, we summarize the use of flows for tasks such as generative modeling, approximate inference, and supervised learning.

Table of Contents

  • 1 Introduction
  • 2 Normalizing Flows
  • 2.1 Definition and Basics
  • 2.2 Expressive Power of Flow-Based Models
  • 2.3 Using Flows for Modeling and Inference
  • 2.3.1 Forward KL Divergence and Maximum Likelihood Estimation
  • 2.3.2 Reverse KL Divergence
  • 2.3.3 Relationship Between Forward and Reverse KL Divergence
  • 2.3.4 Alternative Divergences
  • 2.4 Brief Historical Overview
  • 3 Constructing Flows Part I: Finite Compositions
  • 3.1 Autoregressive Flows
  • 3.1.1 Implementing the Transformer
  • 3.1.2 Implementing the Conditioner
  • 3.1.3 Relationship with Autoregressive Models
  • 3.2 Linear Flows
  • 3.3 Residual Flows
  • 3.3.1 Contractive Residual Flows
  • 3.3.2 Residual Flows Based on the Matrix Determinant Lemma
  • 3.4 Practical Considerations when Combining Transformations
  • 4 Constructing Flows Part II: Continuous-Time Transformations
  • 4.1 Definition
  • 4.2 Solving and Optimizing Continuous-Time Flows
  • 4.2.1 Euler’s Method and Equivalence to Residual Flows
  • 4.2.2 The Adjoint Method
  • 5 Generalizations
  • 5.1 General Probability-Transformation Formula
  • 5.2 Piecewise-Invertible Transformations and Mixtures of Flows
  • 5.3 Flows for Discrete Random Variables
  • 5.4 Flows on Riemannian Manifolds
  • 5.5 Bypassing Topological Constraints
  • 5.6 Symmetric Densities and Equivariant Flows
  • 6 Applications
  • 6.1 Probabilistic Modeling
  • 6.1.1 Density Estimation
  • 6.1.2 Generation
  • 6.2 Inference
  • 6.2.1 Importance and Rejection Sampling
  • 6.2.2 Markov Chain Monte Carlo
  • 6.2.3 Variational Inference
  • 6.2.4 Likelihood-Free Inference
  • 6.3 Using Flows for Representation Learning
  • 6.3.1 Classification and Hybrid Modeling
  • 6.3.2 Reinforcement Learning
  • 7 Conclusions
  • A Proof of KL Dualities
  • B Constructing Linear Flows
  • References

Knowls

  1. Knowl 1 — Continuous Normalizing Flows and Change-of-Variables Density Formula

    definition

    A normalizing flow defines a probability distribution over a continuous random variable x∈RD\mathbf{x} \in \mathbb{R}^D as an invertible and differentiable mapping T:RD→RDT: \mathbb{R}^D \to \mathbb{R}^D (a diffeomorphism) applied to a base random variable u∼pu(u)\mathbf{u} \sim p_{\mathbf{u}}(\mathbf{u}):

    x=T(u)where u∼pu(u).\mathbf{x} = T(\mathbf{u}) \quad \text{where } \mathbf{u} \sim p_{\mathbf{u}}(\mathbf{u}).

    Because TT is bijective with a differentiable inverse T−1T^{-1}, the probability density function px(x)p_{\mathbf{x}}(\mathbf{x}) is determined exactly by the change-of-variables theorem:

    px(x)=pu(T−1(x))∣det⁡JT−1(x)∣=pu(u)∣det⁡JT(u)∣−1p_{\mathbf{x}}(\mathbf{x}) = p_{\mathbf{u}}(T^{-1}(\mathbf{x})) \left|\det J_{T^{-1}}(\mathbf{x})\right| = p_{\mathbf{u}}(\mathbf{u}) \left|\det J_T(\mathbf{u})\right|^{-1}

    where u=T−1(x)\mathbf{u} = T^{-1}(\mathbf{x}), and JT(u)∈RD×DJ_T(\mathbf{u}) \in \mathbb{R}^{D \times D} is the Jacobian matrix of partial derivatives [JT(u)]ij=∂Ti∂uj\left[J_T(\mathbf{u})\right]_{ij} = \frac{\partial T_i}{\partial u_j}.

    Diffeomorphic transformations are composable: for a sequence of KK transformations T=TK∘TK−1∘⋯∘T1T = T_K \circ T_{K-1} \circ \dots \circ T_1 where z0=u\mathbf{z}_0 = \mathbf{u} and zk=Tk(zk−1)\mathbf{z}_k = T_k(\mathbf{z}_{k-1}), the composite inverse is T−1=T1−1∘⋯∘TK−1T^{-1} = T_1^{-1} \circ \dots \circ T_K^{-1}, and the total log-absolute Jacobian determinant decomposes additively across intermediate states:

    log⁡∣det⁡JT(u)∣=∑k=1Klog⁡∣det⁡JTk(zk−1)∣.\log \left|\det J_T(\mathbf{u})\right| = \sum_{k=1}^K \log \left|\det J_{T_k}(\mathbf{z}_{k-1})\right|.

  2. Knowl 2 — Universal Representational Power of Normalizing Flows via Conditional CDFs

    theoretical result

    Let px(x)p_{\mathbf{x}}(\mathbf{x}) be an arbitrary target probability density on RD\mathbb{R}^D that is strictly positive everywhere (px(x)>0p_{\mathbf{x}}(\mathbf{x}) > 0 for all x∈RD\mathbf{x} \in \mathbb{R}^D) and whose conditional cumulative distribution functions (CDFs) Pr⁡(xi′≤xi∣x<i)\Pr(x'_i \le x_i \mid \mathbf{x}_{<i}) are differentiable with respect to (xi,x<i)(x_i, \mathbf{x}_{<i}).

    There exists a diffeomorphism F:RD→(0,1)DF: \mathbb{R}^D \to (0, 1)^D defined component-wise by conditional cumulative distribution functions:

    zi=Fi(xi,x<i)=∫−∞xipx(xi′∣x<i)dxi′=Pr⁡(xi′≤xi∣x<i)z_i = F_i(x_i, \mathbf{x}_{<i}) = \int_{-\infty}^{x_i} p_{\mathbf{x}}(x'_i \mid \mathbf{x}_{<i}) dx'_i = \Pr(x'_i \le x_i \mid \mathbf{x}_{<i})

    for i=1,…,Di = 1, \dots, D, where x<i=(x1,…,xi−1)\mathbf{x}_{<i} = (x_1, \dots, x_{i-1}). Because ∂Fi∂xj=0\frac{\partial F_i}{\partial x_j} = 0 for all j>ij > i, the Jacobian matrix JF(x)J_F(\mathbf{x}) is lower triangular. Its determinant equals the product of diagonal entries:

    det⁡JF(x)=∏i=1D∂Fi∂xi(xi,x<i)=∏i=1Dpx(xi∣x<i)=px(x).\det J_F(\mathbf{x}) = \prod_{i=1}^D \frac{\partial F_i}{\partial x_i}(x_i, \mathbf{x}_{<i}) = \prod_{i=1}^D p_{\mathbf{x}}(x_i \mid \mathbf{x}_{<i}) = p_{\mathbf{x}}(\mathbf{x}).

    The push-forward density of z=F(x)\mathbf{z} = F(\mathbf{x}) satisfies pz(z)=px(x)∣det⁡JF(x)∣−1=1p_{\mathbf{z}}(\mathbf{z}) = p_{\mathbf{x}}(\mathbf{x}) |\det J_F(\mathbf{x})|^{-1} = 1, making z\mathbf{z} uniformly distributed on the open hypercube (0,1)D(0, 1)^D.

    By constructing the analogous diffeomorphism G:RD→(0,1)DG: \mathbb{R}^D \to (0, 1)^D for any base distribution pu(u)p_{\mathbf{u}}(\mathbf{u}) satisfying the same regularity conditions, the composed mapping T=F−1∘GT = F^{-1} \circ G is a diffeomorphism transforming pu(u)p_{\mathbf{u}}(\mathbf{u}) exactly into px(x)p_{\mathbf{x}}(\mathbf{x}).

  3. Knowl 3 — Forward and Reverse Kullback-Leibler Duality in Normalizing Flows

    theoretical result

    Let px(x;θ)p_{\mathbf{x}}(\mathbf{x}; \boldsymbol{\theta}) be a flow model defined by a diffeomorphism T(⋅;ϕ)T(\cdot; \boldsymbol{\phi}) and base density pu(u;ψ)p_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\psi}), with parameters θ={ϕ,ψ}\boldsymbol{\theta} = \{\boldsymbol{\phi}, \boldsymbol{\psi}\}. Let px∗(x)p^*_{\mathbf{x}}(\mathbf{x}) be a target density on RD\mathbb{R}^D, and let pu∗(u;ϕ)p^*_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\phi}) denote the density induced on the base space by transforming samples from px∗(x)p^*_{\mathbf{x}}(\mathbf{x}) through the inverse mapping T−1(⋅;ϕ)T^{-1}(\cdot; \boldsymbol{\phi}).

    Under the change of variables x=T(u;ϕ)\mathbf{x} = T(\mathbf{u}; \boldsymbol{\phi}), the forward and reverse Kullback-Leibler (KL) divergences satisfy exact duality relationships:

    DKL[px∗(x)∥px(x;θ)]=DKL[pu∗(u;ϕ)∥pu(u;ψ)]D_{\text{KL}}[p^*_{\mathbf{x}}(\mathbf{x}) \parallel p_{\mathbf{x}}(\mathbf{x}; \boldsymbol{\theta})] = D_{\text{KL}}[p^*_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\phi}) \parallel p_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\psi})]

    DKL[px(x;θ)∥px∗(x)]=DKL[pu(u;ψ)∥pu∗(u;ϕ)].D_{\text{KL}}[p_{\mathbf{x}}(\mathbf{x}; \boldsymbol{\theta}) \parallel p^*_{\mathbf{x}}(\mathbf{x})] = D_{\text{KL}}[p_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\psi}) \parallel p^*_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\phi})].

    Consequently:

    1. Fitting the flow model px(x;θ)p_{\mathbf{x}}(\mathbf{x}; \boldsymbol{\theta}) to target data via forward KL divergence (maximum likelihood estimation) is mathematically equivalent to fitting the push-forward distribution pu∗(u;ϕ)p^*_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\phi}) to the base distribution pu(u;ψ)p_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\psi}) under the reverse KL divergence.
    2. Fitting the flow model via reverse KL divergence (as in variational inference or model distillation) is equivalent to fitting the base distribution pu(u;ψ)p_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\psi}) to pu∗(u;ϕ)p^*_{\mathbf{u}}(\mathbf{u}; \boldsymbol{\phi}) via maximum likelihood.
  4. Knowl 4 — Autoregressive Flows Structure and Computational Asymmetry

    model/method

    An autoregressive flow step transforms a vector z∈RD\mathbf{z} \in \mathbb{R}^D into z′∈RD\mathbf{z}' \in \mathbb{R}^D using elementwise scalar transformations:

    zi′=τ(zi;hi)where hi=ci(z<i)z'_i = \tau(z_i; \mathbf{h}_i) \quad \text{where } \mathbf{h}_i = c_i(\mathbf{z}_{<i})

    where τ:R→R\tau: \mathbb{R} \to \mathbb{R} is a strictly monotonic, invertible function called the transformer, and ci:Ri−1→RPc_i: \mathbb{R}^{i-1} \to \mathbb{R}^P is the ii-th conditioner that depends only on preceding components z<i=(z1,…,zi−1)\mathbf{z}_{<i} = (z_1, \dots, z_{i-1}).

    The Jacobian matrix Jfϕ(z)J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) is lower triangular with diagonal entries [Jfϕ(z)]ii=∂τ∂zi(zi;hi)[J_{f_{\boldsymbol{\phi}}}(\mathbf{z})]_{ii} = \frac{\partial \tau}{\partial z_i}(z_i; \mathbf{h}_i). The log-absolute Jacobian determinant is computed in O(D)\mathcal{O}(D) time:

    log⁡∣det⁡Jfϕ(z)∣=∑i=1Dlog⁡∣∂τ∂zi(zi;hi)∣.\log |\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z})| = \sum_{i=1}^D \log \left| \frac{\partial \tau}{\partial z_i}(z_i; \mathbf{h}_i) \right|.

    Autoregressive flows exhibit a fundamental computational asymmetry:

    • Forward transformation (→→): If the conditioners c1,…,cDc_1, \dots, c_D are computed jointly via a single masked neural network pass, all parameter vectors (h1,…,hD)(\mathbf{h}_1, \dots, \mathbf{h}_D) and all outputs z′\mathbf{z}' are evaluated in parallel in O(1)\mathcal{O}(1) network evaluations.
    • Inverse transformation (←←): Inverting requires computing zi=τ−1(zi′;hi)z_i = \tau^{-1}(z'_i; \mathbf{h}_i) sequentially for i=1,…,Di = 1, \dots, D, because hi=ci(z<i)\mathbf{h}_i = c_i(\mathbf{z}_{<i}) cannot be computed until all previous components z<i\mathbf{z}_{<i} are known. Exact inversion requires DD sequential neural network evaluations.
  5. Knowl 5 — Coupling Layers for Computationally Symmetric Normalizing Flows

    model/method

    A coupling layer partitions the input vector z∈RD\mathbf{z} \in \mathbb{R}^D into two subsets z=[z≤d,z>d]\mathbf{z} = [\mathbf{z}_{\le d}, \mathbf{z}_{>d}] for a split dimension 1≤d<D1 \le d < D (typically d=⌊D/2⌋d = \lfloor D/2 \rfloor). The first dd components are left unchanged, while the remaining D−dD - d components are transformed conditionally on z≤d\mathbf{z}_{\le d}:

    z≤d′=z≤d\mathbf{z}'_{\le d} = \mathbf{z}_{\le d} h>d=(hd+1,…,hD)=F(z≤d)\mathbf{h}_{>d} = (\mathbf{h}_{d+1}, \dots, \mathbf{h}_D) = F(\mathbf{z}_{\le d}) zi′=τ(zi;hi)for i=d+1,…,Dz'_i = \tau(z_i; \mathbf{h}_i) \quad \text{for } i = d+1, \dots, D

    where F:Rd→R(D−d)×PF: \mathbb{R}^d \to \mathbb{R}^{(D-d) \times P} is an arbitrary unconstrained neural network, and τ\tau is an invertible scalar transformer.

    The inverse mapping possesses the exact same computational structure and is computed in a single pass:

    z≤d=z≤d′\mathbf{z}_{\le d} = \mathbf{z}'_{\le d} h>d=F(z≤d)\mathbf{h}_{>d} = F(\mathbf{z}_{\le d}) zi=τ−1(zi′;hi)for i=d+1,…,D.z_i = \tau^{-1}(z'_i; \mathbf{h}_i) \quad \text{for } i = d+1, \dots, D.

    The Jacobian matrix has a block-triangular structure:

    Jfϕ(z)=[I0AD]J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) = \begin{bmatrix} \mathbf{I} & \mathbf{0} \\ \mathbf{A} & \mathbf{D} \end{bmatrix}

    where I\mathbf{I} is the d×dd \times d identity matrix, 0\mathbf{0} is the d×(D−d)d \times (D-d) zero matrix, A∈R(D−d)×d\mathbf{A} \in \mathbb{R}^{(D-d) \times d}, and D∈R(D−d)×(D−d)\mathbf{D} \in \mathbb{R}^{(D-d) \times (D-d)} is a diagonal matrix with diagonal entries Dii=∂τ∂zi(zi;hi)D_{ii} = \frac{\partial \tau}{\partial z_i}(z_i; \mathbf{h}_i). The determinant is evaluated in O(D)\mathcal{O}(D) operations as det⁡Jfϕ(z)=∏i=d+1D∂τ∂zi(zi;hi)\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) = \prod_{i=d+1}^D \frac{\partial \tau}{\partial z_i}(z_i; \mathbf{h}_i).

    A single coupling layer is not a universal approximator. Expressive models are formed by composing multiple coupling layers with permutations of coordinates between successive layers.

  6. Knowl 6 — Contractive Residual Flows and Power Series Log-Determinant Estimation

    model/method

    A contractive residual flow defines a transformation fϕ(z)=z+gϕ(z)f_{\boldsymbol{\phi}}(\mathbf{z}) = \mathbf{z} + g_{\boldsymbol{\phi}}(\mathbf{z}) where gϕ:RD→RDg_{\boldsymbol{\phi}}: \mathbb{R}^D \to \mathbb{R}^D is Lipschitz continuous with Lipschitz constant L<1L < 1 with respect to a norm δ\delta.

    By the Banach fixed-point theorem, the mapping F(z^)=z′−gϕ(z^)F(\hat{\mathbf{z}}) = \mathbf{z}' - g_{\boldsymbol{\phi}}(\hat{\mathbf{z}}) is a contraction with Lipschitz constant LL. The inverse z=fϕ−1(z′)\mathbf{z} = f_{\boldsymbol{\phi}}^{-1}(\mathbf{z}') is the unique fixed point z∗=F(z∗)\mathbf{z}^* = F(\mathbf{z}^*) obtained by the iteration:

    zk+1=z′−gϕ(zk)for k≥0\mathbf{z}_{k+1} = \mathbf{z}' - g_{\boldsymbol{\phi}}(\mathbf{z}_k) \quad \text{for } k \ge 0

    which converges exponentially from any z0\mathbf{z}_0 according to δ(zk,z∗)≤Lk1−Lδ(z0,z1)\delta(\mathbf{z}_k, \mathbf{z}^*) \le \frac{L^k}{1 - L} \delta(\mathbf{z}_0, \mathbf{z}_1).

    The log-absolute Jacobian determinant is expanded as a convergent matrix power series:

    log⁡∣det⁡Jfϕ(z)∣=log⁡∣det⁡(I+Jgϕ(z))∣=∑k=1∞(−1)k+1kTr⁡(Jgϕk(z)).\log |\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z})| = \log |\det(\mathbf{I} + J_{g_{\boldsymbol{\phi}}}(\mathbf{z}))| = \sum_{k=1}^\infty \frac{(-1)^{k+1}}{k} \operatorname{Tr}\left(J_{g_{\boldsymbol{\phi}}}^k(\mathbf{z})\right).

    The trace of the kk-th matrix power is estimated without forming the full D×DD \times D Jacobian using the Hutchinson trace estimator:

    Tr⁡(Jgϕk(z))≈v⊤Jgϕk(z)v\operatorname{Tr}\left(J_{g_{\boldsymbol{\phi}}}^k(\mathbf{z})\right) \approx \mathbf{v}^\top J_{g_{\boldsymbol{\phi}}}^k(\mathbf{z}) \mathbf{v}

    where v∈RD\mathbf{v} \in \mathbb{R}^D is a random vector with E[v]=0\mathbb{E}[\mathbf{v}] = \mathbf{0} and Cov⁡(v)=I\operatorname{Cov}(\mathbf{v}) = \mathbf{I}, computed with kk sequential vector-Jacobian products via automatic differentiation. The infinite power series is made unbiased and finite using Russian-roulette sampling.

  7. Knowl 7 — Residual Flows via the Matrix Determinant Lemma: Planar, Sylvester, and Radial

    model/method

    For an invertible matrix A∈RD×D\mathbf{A} \in \mathbb{R}^{D \times D} and low-rank factor matrices V,W∈RD×M\mathbf{V}, \mathbf{W} \in \mathbb{R}^{D \times M}, the matrix determinant lemma states:

    det⁡(A+VW⊤)=det⁡(I+W⊤A−1V)det⁡A.\det(\mathbf{A} + \mathbf{V}\mathbf{W}^\top) = \det(\mathbf{I} + \mathbf{W}^\top \mathbf{A}^{-1}\mathbf{V}) \det \mathbf{A}.

    Applying this identity yields residual flow architectures with O(D)\mathcal{O}(D) or O(M3+DM2)\mathcal{O}(M^3 + DM^2) Jacobian determinant computation:

    1. Planar Flow (M=1M=1): z′=z+vσ(w⊤z+b)\mathbf{z}' = \mathbf{z} + \mathbf{v}\sigma(\mathbf{w}^\top \mathbf{z} + b) with parameters v,w∈RD\mathbf{v}, \mathbf{w} \in \mathbb{R}^D, b∈Rb \in \mathbb{R}, and smooth activation σ:R→R\sigma: \mathbb{R} \to \mathbb{R}. The Jacobian determinant is: det⁡Jfϕ(z)=1+σ′(w⊤z+b)w⊤v.\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) = 1 + \sigma'(\mathbf{w}^\top \mathbf{z} + b) \mathbf{w}^\top \mathbf{v}. A sufficient condition for invertibility when σ′(x)>0\sigma'(x) > 0 and bounded is w⊤v>−1sup⁡xσ′(x)\mathbf{w}^\top \mathbf{v} > -\frac{1}{\sup_x \sigma'(x)}.

    2. Sylvester Flow (M≤DM \le D): z′=z+Vσ(W⊤z+b)\mathbf{z}' = \mathbf{z} + \mathbf{V}\sigma(\mathbf{W}^\top \mathbf{z} + \mathbf{b}) with V=QU\mathbf{V} = \mathbf{Q}\mathbf{U} and W=QL\mathbf{W} = \mathbf{Q}\mathbf{L}, where Q∈RD×M\mathbf{Q} \in \mathbb{R}^{D \times M} has orthonormal columns (Q⊤Q=I\mathbf{Q}^\top \mathbf{Q} = \mathbf{I}), U∈RM×M\mathbf{U} \in \mathbb{R}^{M \times M} is upper triangular, and L∈RM×M\mathbf{L} \in \mathbb{R}^{M \times M} is lower triangular. The Jacobian determinant simplifies to: det⁡Jfϕ(z)=∏i=1M(1+σ′(W⊤z+b)iLiiUii)\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) = \prod_{i=1}^M \left(1 + \sigma'(\mathbf{W}^\top \mathbf{z} + \mathbf{b})_i L_{ii} U_{ii}\right) which is invertible if LiiUii>−1sup⁡xσ′(x)L_{ii} U_{ii} > -\frac{1}{\sup_x \sigma'(x)} for all i∈{1,…,M}i \in \{1, \dots, M\}.

    3. Radial Flow: z′=z+βα+r(z)(z−z0)where r(z)=∥z−z0∥\mathbf{z}' = \mathbf{z} + \frac{\beta}{\alpha + r(\mathbf{z})}(\mathbf{z} - \mathbf{z}_0) \quad \text{where } r(\mathbf{z}) = \|\mathbf{z} - \mathbf{z}_0\| with parameters α>0\alpha > 0, β∈R\beta \in \mathbb{R}, and center z0∈RD\mathbf{z}_0 \in \mathbb{R}^D. The Jacobian determinant is: det⁡Jfϕ(z)=(1+αβ(α+r(z))2)(1+βα+r(z))D−1\det J_{f_{\boldsymbol{\phi}}}(\mathbf{z}) = \left(1 + \frac{\alpha \beta}{(\alpha + r(\mathbf{z}))^2}\right) \left(1 + \frac{\beta}{\alpha + r(\mathbf{z})}\right)^{D-1} and invertibility is guaranteed for β>−α\beta > -\alpha.

  8. Knowl 8 — Continuous-Time Normalizing Flows and Instantaneous Density Evolution

    model/method

    A continuous-time normalizing flow defines state trajectories zt∈RD\mathbf{z}_t \in \mathbb{R}^D over time t∈[t0,t1]t \in [t_0, t_1] via an ordinary differential equation (ODE) parameterizing infinitesimal dynamics:

    dztdt=gϕ(t,zt)\frac{d\mathbf{z}_t}{dt} = g_{\boldsymbol{\phi}}(t, \mathbf{z}_t)

    where zt0=u\mathbf{z}_{t_0} = \mathbf{u} and zt1=x\mathbf{z}_{t_1} = \mathbf{x}. The function gϕg_{\boldsymbol{\phi}} must be continuous in tt and uniformly Lipschitz continuous in zt\mathbf{z}_t.

    The forward transformation TT and inverse transformation T−1T^{-1} are computed by integration:

    x=u+∫t0t1gϕ(t,zt)dt,u=x−∫t0t1gϕ(t,zt)dt.\mathbf{x} = \mathbf{u} + \int_{t_0}^{t_1} g_{\boldsymbol{\phi}}(t, \mathbf{z}_t) dt, \qquad \mathbf{u} = \mathbf{x} - \int_{t_0}^{t_1} g_{\boldsymbol{\phi}}(t, \mathbf{z}_t) dt.

    Both directions require running the ODE solver over the interval [t0,t1][t_0, t_1] and thus share identical computational complexity.

    The instantaneous change in log density along the trajectory is governed by the trace of the Jacobian matrix of the vector field:

    dlog⁡p(zt)dt=−Tr⁡(Jgϕ(t,⋅)(zt))\frac{d \log p(\mathbf{z}_t)}{dt} = -\operatorname{Tr}\left(J_{g_{\boldsymbol{\phi}}(t, \cdot)}(\mathbf{z}_t)\right)

    yielding the total target log density:

    log⁡px(x)=log⁡pu(u)−∫t0t1Tr⁡(Jgϕ(t,⋅)(zt))dt.\log p_{\mathbf{x}}(\mathbf{x}) = \log p_{\mathbf{u}}(\mathbf{u}) - \int_{t_0}^{t_1} \operatorname{Tr}\left(J_{g_{\boldsymbol{\phi}}(t, \cdot)}(\mathbf{z}_t)\right) dt.

    Gradients of a scalar loss L(x;ϕ)\mathcal{L}(\mathbf{x}; \boldsymbol{\phi}) with respect to parameters ϕ\boldsymbol{\phi} can be computed in O(1)\mathcal{O}(1) memory via the adjoint sensitivity ODE without backpropagating through the internal steps of numerical ODE integrators:

    ddt(∂L∂zt)=−(∂L∂zt)⊤∂gϕ(t,zt)∂zt,∂L∂ϕ=∫t1t0∂L∂zt∂gϕ(t,zt)∂ϕdt.\frac{d}{dt} \left(\frac{\partial \mathcal{L}}{\partial \mathbf{z}_t}\right) = -\left(\frac{\partial \mathcal{L}}{\partial \mathbf{z}_t}\right)^\top \frac{\partial g_{\boldsymbol{\phi}}(t, \mathbf{z}_t)}{\partial \mathbf{z}_t}, \qquad \frac{\partial \mathcal{L}}{\partial \boldsymbol{\phi}} = \int_{t_1}^{t_0} \frac{\partial \mathcal{L}}{\partial \mathbf{z}_t} \frac{\partial g_{\boldsymbol{\phi}}(t, \mathbf{z}_t)}{\partial \boldsymbol{\phi}} dt.

  9. Knowl 9 — Equivariant Normalizing Flows for Symmetrically Invariant Target Densities

    theoretical result

    Let GG be a symmetry group where each element g∈Gg \in G has a matrix representation as an invertible linear transformation Rg∈RD×D\mathbf{R}_g \in \mathbb{R}^{D \times D} satisfying ∣det⁡Rg∣=1|\det \mathbf{R}_g| = 1. A density p(x)p(\mathbf{x}) is invariant under GG if p(Rgx)=p(x)p(\mathbf{R}_g \mathbf{x}) = p(\mathbf{x}) for all g∈Gg \in G and all x∈RD\mathbf{x} \in \mathbb{R}^D.

    Lemma (Equivariant Flows): If a base density pu(u)p_{\mathbf{u}}(\mathbf{u}) is invariant with respect to GG (i.e., pu(Rgu)=pu(u)p_{\mathbf{u}}(\mathbf{R}_g \mathbf{u}) = p_{\mathbf{u}}(\mathbf{u})) and the flow mapping T:RD→RDT: \mathbb{R}^D \to \mathbb{R}^D is GG-equivariant (i.e., T(Rgu)=RgT(u)T(\mathbf{R}_g \mathbf{u}) = \mathbf{R}_g T(\mathbf{u}) for all g∈Gg \in G), then the resulting model density px(x)p_{\mathbf{x}}(\mathbf{x}) is invariant with respect to GG.

    Lemma (Equivariance from Invariance): Let f:RD→Rf: \mathbb{R}^D \to \mathbb{R} be a scalar function invariant under GG (i.e., f(Rgu)=f(u)f(\mathbf{R}_g \mathbf{u}) = f(\mathbf{u})), and assume Rg\mathbf{R}_g is orthogonal (Rg⊤Rg=I\mathbf{R}_g^\top \mathbf{R}_g = \mathbf{I}) for all g∈Gg \in G. Then the gradient operator ∇uf(u)\nabla_{\mathbf{u}} f(\mathbf{u}) is GG-equivariant:

    ∇uf(Rgu)=Rg∇uf(u).\nabla_{\mathbf{u}} f(\mathbf{R}_g \mathbf{u}) = \mathbf{R}_g \nabla_{\mathbf{u}} f(\mathbf{u}).

  10. Knowl 10 — Linear Flows and Parameterization Constraints of Invertible Matrices

    theoretical result

    A linear flow layer performs an invertible matrix multiplication z′=Wz\mathbf{z}' = \mathbf{W}\mathbf{z} parameterized by an invertible matrix W∈RD×D\mathbf{W} \in \mathbb{R}^{D \times D}. Direct parameterization incurs O(D3)\mathcal{O}(D^3) inversion and determinant cost. Structured parameterizations reduce costs:

    1. PLU Parameterization: W=PLU\mathbf{W} = \mathbf{P}\mathbf{L}\mathbf{U} with permutation matrix P\mathbf{P}, lower-triangular L\mathbf{L}, and upper-triangular U\mathbf{U} having positive diagonal elements. The determinant is ∣det⁡W∣=∏i=1DLiiUii|\det \mathbf{W}| = \prod_{i=1}^D L_{ii} U_{ii} in O(D)\mathcal{O}(D) time, and the inverse linear system is solved via substitution in O(D2)\mathcal{O}(D^2) time.
    2. QR Parameterization: W=QR\mathbf{W} = \mathbf{Q}\mathbf{R} where Q\mathbf{Q} is orthogonal (Q−1=Q⊤\mathbf{Q}^{-1} = \mathbf{Q}^\top) and R\mathbf{R} is upper triangular with Rii>0R_{ii} > 0. The determinant is ∣det⁡W∣=∏i=1DRii|\det \mathbf{W}| = \prod_{i=1}^D R_{ii}, and inversion costs O(D2)\mathcal{O}(D^2).
    3. Orthogonal Maps: Parameterized using the matrix exponential Q=exp⁡(A)\mathbf{Q} = \exp(\mathbf{A}) or Cayley transform Q=(I+A)(I−A)−1\mathbf{Q} = (\mathbf{I}+\mathbf{A})(\mathbf{I}-\mathbf{A})^{-1} for skew-symmetric A=−A⊤\mathbf{A} = -\mathbf{A}^\top, or by products of K≤DK \le D Householder reflections Hk=I−2vkvk⊤∥vk∥2\mathbf{H}_k = \mathbf{I} - 2 \frac{\mathbf{v}_k \mathbf{v}_k^\top}{\|\mathbf{v}_k\|^2}.

    Topological Limitation: There is no continuous surjective function from RD2\mathbb{R}^{D^2} to the general linear group GL(D)GL(D) of all D×DD \times D invertible matrices. The set of invertible matrices contains two disconnected manifolds separated by the non-invertible boundary det⁡W=0\det \mathbf{W} = 0: one manifold with det⁡W>0\det \mathbf{W} > 0 and one with det⁡W<0\det \mathbf{W} < 0. Any continuous parameterization of W\mathbf{W} can only parameterize one connected component, fixing the sign of the determinant.

  11. Knowl 11 — Discrete Normalizing Flows and Representation Limitations of Factorized Bases

    theoretical result

    For discrete domains U=X\mathcal{U} = \mathcal{X}, normalizing flows define a bijection T:U→XT: \mathcal{U} \to \mathcal{X}. By conservation of counting measure, the probability mass function transforms without a Jacobian term:

    px(x)=pu(T−1(x)).p_{\mathbf{x}}(\mathbf{x}) = p_{\mathbf{u}}(T^{-1}(\mathbf{x})).

    Because TT is bijective, pxp_{\mathbf{x}} is strictly a permutation of the set of probability values assigned by pup_{\mathbf{u}}. Consequently, if the base distribution is uniform, the transformed distribution pxp_{\mathbf{x}} must also be uniform.

    Factorized Base Limitation: A discrete flow with a fully factorized base distribution pu(u)=∏i=1Dpu(ui)p_{\mathbf{u}}(\mathbf{u}) = \prod_{i=1}^D p_{\mathbf{u}}(u_i) cannot model arbitrary target distributions over X={0,…,K−1}D\mathcal{X} = \{0, \dots, K-1\}^D. Specifically, for a target joint distribution px(x1,x2)p_{\mathbf{x}}(x_1, x_2) represented as a probability matrix of rank R>1R > 1, permuting matrix elements cannot produce a rank-1 outer product matrix unless the target satisfies strict algebraic rank constraints.

    To represent general discrete distributions, discrete flows must either employ autoregressive base distributions or embed the variables into extended discrete spaces U′\mathcal{U}' of higher cardinality (space lifting).

  12. Knowl 12 — Normalizing Flows on Embedded Riemannian Manifolds

    model/method

    Let X\mathcal{X} be a DD-dimensional Riemannian manifold embedded in Euclidean space RM\mathbb{R}^M (M≥DM \ge D) traced by an injective immersion T:RD→XT: \mathbb{R}^D \to \mathcal{X}. The map TT induces a Riemannian metric tensor G(u)∈RD×DG(\mathbf{u}) \in \mathbb{R}^{D \times D} on the tangent space of X\mathcal{X} at x=T(u)\mathbf{x} = T(\mathbf{u}):

    G(u)=JT(u)⊤JT(u)G(\mathbf{u}) = J_T(\mathbf{u})^\top J_T(\mathbf{u})

    where JT(u)∈RM×DJ_T(\mathbf{u}) \in \mathbb{R}^{M \times D} is the Jacobian matrix of the embedding.

    The infinitesimal Riemannian volume measure on the manifold is dν(x)=det⁡G(u)dud\nu(\mathbf{x}) = \sqrt{\det G(\mathbf{u})} d\mathbf{u}. By conservation of probability measure, the density px(x)p_{\mathbf{x}}(\mathbf{x}) on the manifold is related to the density on RD\mathbb{R}^D by:

    px(x)=pu(T−1(x))[det⁡G(T−1(x))]−1/2p_{\mathbf{x}}(\mathbf{x}) = p_{\mathbf{u}}(T^{-1}(\mathbf{x})) \left[ \det G(T^{-1}(\mathbf{x})) \right]^{-1/2}

    where T−1:X→RDT^{-1}: \mathcal{X} \to \mathbb{R}^D is the inverse chart mapping.

    A global flow between manifolds is constructed by composing the inverse chart of a base manifold U\mathcal{U}, standard Euclidean flows on RD\mathbb{R}^D, and the forward chart of the target manifold X\mathcal{X}. This formulation requires U\mathcal{U} and X\mathcal{X} to be homeomorphic to RD\mathbb{R}^D. Applying this to manifolds with non-Euclidean topologies (such as spheres SDS^D or tori) necessarily creates coordinate singularities where the induced density becomes infinite.

  13. Knowl 13 — Conservation of Measure and Piecewise-Invertible Normalizing Flow Mixtures

    theoretical result

    The general probability transformation relation across arbitrary measure spaces (U,μ)(\mathcal{U}, \mu) and (X,ν)(\mathcal{X}, \nu) under a mapping T:U→XT: \mathcal{U} \to \mathcal{X} states that for all measurable sets ω⊆U\omega \subseteq \mathcal{U} with image γ={T(u)∣u∈ω}⊆X\gamma = \{T(\mathbf{u}) \mid \mathbf{u} \in \omega\} \subseteq \mathcal{X}:

    ∫u∈ωpu(u)dμ(u)=∫x∈γpx(x)dν(x).\int_{\mathbf{u} \in \omega} p_{\mathbf{u}}(\mathbf{u}) d\mu(\mathbf{u}) = \int_{\mathbf{x} \in \gamma} p_{\mathbf{x}}(\mathbf{x}) d\nu(\mathbf{x}).

    Applying this principle generalizes normalizing flows beyond standard diffeomorphisms:

    1. Many-to-One Piecewise Invertible Transformations: If U\mathcal{U} is partitioned into countable disjoint subsets {Ui}i∈I\{\mathcal{U}_i\}_{i \in \mathcal{I}} such that the restriction Ti=T∣Ui:Ui→XT_i = T|_{\mathcal{U}_i}: \mathcal{U}_i \to \mathcal{X} is a diffeomorphism onto X\mathcal{X}, the resulting density on X\mathcal{X} is an exact mixture of flows: px(x)=∑i∈Ipu(Ti−1(x))∣det⁡JTi−1(x)∣.p_{\mathbf{x}}(\mathbf{x}) = \sum_{i \in \mathcal{I}} p_{\mathbf{u}}(T_i^{-1}(\mathbf{x})) \left| \det J_{T_i^{-1}}(\mathbf{x}) \right|.

    2. One-to-Many Transformations (RAD / Extended Space): If X\mathcal{X} is partitioned into disjoint subsets {Xi}i∈I\{\mathcal{X}_i\}_{i \in \mathcal{I}} such that an inverse mapping R:X→UR: \mathcal{X} \to \mathcal{U} has restrictions R∣Xi=Ti−1:Xi→UR|_{\mathcal{X}_i} = T_i^{-1}: \mathcal{X}_i \to \mathcal{U} which are diffeomorphisms, and discrete branch indices i∈Ii \in \mathcal{I} are sampled according to p(i∣u)p(i \mid \mathbf{u}), the marginal density on X\mathcal{X} is: px(x)=pu(R(x))p(i(x)∣R(x))∣det⁡JR(x)∣p_{\mathbf{x}}(\mathbf{x}) = p_{\mathbf{u}}(R(\mathbf{x})) p(i(\mathbf{x}) \mid R(\mathbf{x})) \left| \det J_R(\mathbf{x}) \right| where i(x)i(\mathbf{x}) is the unique partition index containing x\mathbf{x}.

  14. Knowl 14 — Exact Sequential Inversion Procedure for Masked Autoregressive Flows

    algorithm

    Exact inversion of a masked autoregressive flow layer computes the input vector z∈RD\mathbf{z} \in \mathbb{R}^D given the output vector z′∈RD\mathbf{z}' \in \mathbb{R}^D, using a masked conditioner network c:RD→RD×Pc: \mathbb{R}^D \to \mathbb{R}^{D \times P} and a family of scalar invertible transformers τ−1(⋅;h)\tau^{-1}(\cdot; \mathbf{h}). Due to the autoregressive dependency, component ziz_i cannot be computed until the prefix z<i=(z1,…,zi−1)\mathbf{z}_{<i} = (z_1, \dots, z_{i-1}) is fully determined.

    Input: Transformed vector z′∈RD\mathbf{z}' \in \mathbb{R}^D, masked conditioner network cc, invertible transformer inverse τ−1\tau^{-1}
    Output: Original input vector z∈RD\mathbf{z} \in \mathbb{R}^D
    Initialize z∈RD\mathbf{z} \in \mathbb{R}^D to an arbitrary value (e.g., zeros or z′\mathbf{z}')
    for i=1i = 1 to DD do
        (h1,…,hD)=c(z)(\mathbf{h}_1, \dots, \mathbf{h}_D) = c(\mathbf{z})
        zi=τ−1(zi′;hi)z_i = \tau^{-1}(z'_i; \mathbf{h}_i)
    end for
    return z\mathbf{z}

    Correctness follows by induction: because the masked network cc ensures that hi\mathbf{h}_i depends only on z<i\mathbf{z}_{<i}, having correct values for indices 1,…,i−11, \dots, i-1 guarantees that the forward pass c(z)c(\mathbf{z}) computes the true conditioner output hi\mathbf{h}_i, which correctly determines ziz_i. The algorithm executes in O(D)\mathcal{O}(D) sequential forward passes of the conditioner network.

  15. Knowl 15 — Non-Affine Transformer Formulations: Combinations, Integrals, and Splines

    model/method

    To enhance expressivity beyond affine transformers τ(zi;hi)=αizi+βi\tau(z_i; \mathbf{h}_i) = \alpha_i z_i + \beta_i, autoregressive flows employ three principal classes of non-affine strictly monotonic scalar transformers:

    1. Combination-Based Transformers: Formed by conic combinations of strictly monotonic activation functions σ\sigma: τ(zi;hi)=wi0+∑k=1Kwikσ(αikzi+βik)with wik>0,αik>0\tau(z_i; \mathbf{h}_i) = w_{i0} + \sum_{k=1}^K w_{ik} \sigma(\alpha_{ik} z_i + \beta_{ik}) \quad \text{with } w_{ik} > 0, \alpha_{ik} > 0 implementing a strictly monotonic multi-layer perceptron. Derivatives are computed via backpropagation, but inversion lacks closed form and requires iterative root finding (e.g., bisection).

    2. Integration-Based Transformers (Sum-of-Squares): Defined via integrals of non-negative polynomial integrands: τ(zi;hi)=∫0zi∑k=1K(∑ℓ=0Lαikℓzℓ)2dz+βi\tau(z_i; \mathbf{h}_i) = \int_0^{z_i} \sum_{k=1}^K \left( \sum_{\ell=0}^L \alpha_{ik\ell} z^\ell \right)^2 dz + \beta_i with unconstrained coefficients αikℓ\alpha_{ik\ell}. The integral evaluates to an unconstrained positive polynomial of odd degree 2L+12L+1. Derivatives equal the positive integrand directly. Exact analytical inversion is possible only for L=0L=0 (affine) and L=1L=1 (cubic polynomial).

    3. Spline-Based Transformers: Piecewise rational or polynomial functions over KK segments defined by knots {(zik,zik′)}k=0K\{(z_{ik}, z'_{ik})\}_{k=0}^K. Types include monotonic linear, quadratic, cubic, linear-rational, and rational-quadratic splines. Splines provide exact closed-form analytic inversion and exact derivatives in O(log⁡K)\mathcal{O}(\log K) segment lookup time.

Coverage note — Omitted historical background narratives, standard textbook expositions of unconstrained baseline GAN/VAE literature, and specific empirical benchmark scores curated from external cited papers.

References

  1. 1.Justin Alsing, Benjamin D. Wandelt, and Stephen M. Feeney. Massive optimal data compression and density estimation for scalable, likelihood-free inference in cosmology. Monthly Notices of the Royal Astronomical Society, 477(3):2874–2885, 2018.
  2. 2.Lynton Ardizzone, Carsten Lüth, Jakob Kruse, Carsten Rother, and Ullrich Köthe. Guided image generation with conditional invertible neural networks. ArXiv preprint arXiv:1907.02392, 2019.
  3. 3.Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning, pages 214–223, 2017.
  4. 4.Matthias Bauer and Andriy Mnih. Resampled priors for variational autoencoders. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, pages 66–75, 2019.
  5. 5.Mark A. Beaumont. Approximate Bayesian computation in evolution and ecology. Annual Review of Ecology, Evolution, and Systematics, 41(1):379–406, 2010.
  6. 6.Mark A. Beaumont, Wenyang Zhang, and David J. Balding. Approximate Bayesian computation in population genetics. Genetics, 162:2025–2035, 2002.
  7. 7.Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen, David K. Duvenaud, and Jörn-Henrik Jacobsen. Invertible residual networks. In Proceedings of the 36th International Conference on Machine Learning, pages 573–582, 2019.
  8. 8.Yoshua Bengio and Samy Bengio. Modeling high-dimensional discrete data with multi-layer neural networks. In Advances in Neural Information Processing Systems, pages 400–406, 2000.
  9. 9.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. ArXiv Preprint arXiv:1308.3432, 2013.
  10. 10.Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In International Conference on Learning Representations, 2018.
  11. 11.David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe. Variational inference: A review for statisticians. Journal of the American Statistical Association, 112(518):859–877, 2017.
  12. 12.Vladimir I. Bogachev. Measure Theory. Springer Berlin Heidelberg, 2007.
  13. 13.Vladimir I. Bogachev, Alexander V. Kolesnikov, and Kirill V. Medvedev. Triangular transformations of measures. Sbornik: Mathematics, 196(3):309–335, 2005.
  14. 14.Johann Brehmer, Kyle Cranmer, Gilles Louppe, and Juan Pavez. Constraining effective field theories with machine learning. Physical Review Letters, 121(11):111801, 2018.
  15. 15.Richard L. Burden and J. Douglas Faires. Numerical Analysis. The Prindle, Weber and Schmidt Series in Mathematics. PWS-Kent Publishing Company, fourth edition, 1989.
  16. 16.Guillaume Carlier, Alfred Galichon, and Filippo Santambrogio. From Knothe’s transport to Brenier’s map and a continuation method for optimal transport. SIAM Journal on Mathematical Analysis, 41(6):2554–2576, 2010.
  17. 17.Ricky T. Q. Chen and David K. Duvenaud. Neural networks with cheap differential operators. In Advances in Neural Information Processing Systems, 2019.
  18. 18.Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, pages 6571–6583, 2018.
  19. 19.Ricky T. Q. Chen, Jens Behrmann, David K. Duvenaud, and Jörn-Henrik Jacobsen. Residual flows for invertible generative modeling. In Advances in Neural Information Processing Systems, 2019.
  20. 20.Scott Saobing Chen and Ramesh A. Gopinath. Gaussianization. In Advances in Neural Information Processing Systems, pages 423–429, 2000.
  21. 21.Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, pages 1724–1734, 2014.
  22. 22.Earl A. Coddington and Norman Levinson. Theory of ordinary differential equations. International Series in Pure and Applied Mathematics. McGraw-Hill, 1955.
  23. 23.Rob Cornish, Anthony L. Caterini, George Deligiannidis, and Arnaud Doucet. Localised generative flows. ArXiv Preprint arXiv:1909.13833, 2019.
  24. 24.Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based inference. Proceedings of the National Academy of Sciences, 2020. doi: 10.1073/pnas. 1912789117.
  25. 25.Ivo Danihelka, Balaji Lakshminarayanan, Benigno Uria, Daan Wierstra, and Peter Dayan. Comparison of maximum likelihood and GAN-based training of Real NVPs. ArXiv Preprint arXiv:1705.05263, 2017.
  26. 26.Nicola De Cao, Ivan Titov, and Wilker Aziz. Block neural autoregressive flow. In Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence, 2019.
  27. 27.Zhiwei Deng, Megha Nawhal, Lili Meng, and Greg Mori. Continuous graph flow. ArXiv Preprint arXiv:1908.02436, 2019.
  28. 28.Peter J. Diggle and Richard J. Gratton. Monte Carlo methods of inference for implicit statistical models. Journal of the Royal Statistical Society. Series B (Methodological), pages 193–227, 1984.
  29. 29.Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: Non-linear independent components estimation. ICLR Workshop Track, 2015.
  30. 30.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. In International Conference on Learning Representations, 2017.
  31. 31.Laurent Dinh, Jascha Sohl-Dickstein, Razvan Pascanu, and Hugo Larochelle. A RAD approach to deep mixture models. ICLR Workshop on Deep Generative Models for Highly Structured Data, 2019.
  32. 32.Hadi M. Dolatabadi, Sarah Erfani, and Christopher Leckie. Invertible generative modeling using linear rational splines. In Proceedings of the 23nd International Conference on Artificial Intelligence and Statistics, 2020.
  33. 33.Simon Duane, Anthony D. Kennedy, Brian J. Pendleton, and Duncan Roweth. Hybrid Monte Carlo. Physics Letters B, 195(2):216–222, 1987.
  34. 34.Emilien Dupont, Arnaud Doucet, and Yee Whye Teh. Augmented neural ODEs. In Advances in Neural Information Processing Systems, 2019.
  35. 35.Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Cubic-spline flows. ICML Workshop on Invertible Neural Networks and Normalizing Flows, 2019a.
  36. 36.Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Neural spline flows. In Advances in Neural Information Processing Systems, 2019b.
  37. 37.Gal Elidan. Copulas in machine learning. In Copulae in Mathematical and Quantitative Finance, pages 39–60, 2013.
  38. 38.Luca Falorsi, Pim de Haan, Tim R. Davidson, and Patrick Forré. Reparameterizing distributions on Lie groups. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, pages 3244–3253, 2019.
  39. 39.Brendan J Frey. Graphical models for machine learning and digital communication. MIT press, 1998.
  40. 40.Jerome H. Friedman. Exploratory projection pursuit. Journal of the American Statistical Association, 82(397):249–266, 1987.
  41. 41.Mevlana C. Gemici, Danilo Jimenez Rezende, and Shakir Mohamed. Normalizing flows on Riemannian manifolds. NeurIPS Workshop on Bayesian Deep Learning, 2016.
  42. 42.Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle. MADE: Masked autoencoder for distribution estimation. In Proceedings of the 32nd International Conference on Machine Learning, pages 881–889, 2015.
  43. 43.Adam Golinski, Mario Lezcano-Casado, and Tom Rainforth. Improving normalizing flows via better orthogonal parameterizations. ICML Workshop on Invertible Neural Networks and Normalizing Flows, 2019.
  44. 44.Aidan N. Gomez, Mengye Ren, Raquel Urtasun, and Roger B. Grosse. The reversible residual network: Backpropagation without storing activations. In Advances in Neural Information Processing Systems, pages 2214–2224, 2017.
  45. 45.Pedro J. Gonçalves, Jan-Matthis Lueckmann, Michael Deistler, Marcel Nonnenmacher, Kaan Öcal, Giacomo Bassetto, Chaitanya Chintaluri, William F. Podlaski, Sara A. Haddad, Tim P. Vogels, David S. Greenberg, and Jakob H. Macke. Training deep neural density estimators to identify mechanistic models of neural dynamics. Elife, 9, 2020. doi: 10.7554/eLife.56261.
  46. 46.Henry Gouk, Eibe Frank, Bernhard Pfahringer, and Michael Cree. Regularisation of neural networks by enforcing Lipschitz continuity. ArXiv Preprint arXiv:1804.04368, 2018.
  47. 47.Will Grathwohl, Ricky T. Q. Chen, Jesse Betterncourt, Ilya Sutskever, and David K. Duvenaud. FFJORD: Free-form continuous dynamics for scalable reversible generative models. In International Conference on Learning Representations, 2019.
  48. 48.Alex Graves. Generating sequences with recurrent neural networks. ArXiv Preprint arXiv:1308.0850, 2013.
  49. 49.David S. Greenberg, Marcel Nonnenmacher, and Jakob H. Macke. Automatic posterior transformation for likelihood-free inference. In Proceedings of the 36th International Conference on Machine Learning, pages 2404–2414, 2019.
  50. 50.Aditya Grover, Manik Dhar, and Stefano Ermon. Flow-GAN: Combining maximum likelihood and adversarial learning in generative models. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, 2018.
  51. 51.Tuomas Haarnoja, Kristian Hartikainen, Pieter Abbeel, and Sergey Levine. Latent space policies for hierarchical reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning, pages 1851–1860, 2018.
  52. 52.Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick. Unsupervised learning of syntactic structure with invertible neural projections. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1292–1302, 2018.
  53. 53.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  54. 54.Gustav Eje Henter, Simon Alexanderson, and Jonas Beskow. MoGlow: Probabilistic and controllable motion synthesis using normalising flows. ArXiv Preprint arXiv:1905.06598, 2019.
  55. 55.Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, and Pieter Abbeel. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In Proceedings of the 36th International Conference on Machine Learning, pages 2722–2730, 2019.
  56. 56.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  57. 57.Matthew Hoffman, Pavel Sountsov, Joshua V. Dillon, Ian Langmore, Dustin Tran, and Srinivas Vasudevan. NeuTra-lizing bad geometry in Hamiltonian Monte Carlo using neural transport. ArXiv Preprint arXiv:1903.03704, 2019.
  58. 58.Shion Honda, Hirotaka Akita, Katsuhiko Ishiguro, Toshiki Nakanishi, and Kenta Oono. Graph residual flow for molecular graph generation. ArXiv Preprint arXiv:1909.13521, 2019.
  59. 59.Emiel Hoogeboom, Jorn W. T. Peters, Rianne van den Berg, and Max Welling. Integer discrete flows and lossless compression. In Advances in Neural Information Processing Systems, 2019a.
  60. 60.Emiel Hoogeboom, Rianne Van Den Berg, and Max Welling. Emerging convolutions for generative normalizing flows. In Proceedings of the 36th International Conference on Machine Learning, pages 2771–2780, 2019b.
  61. 61.Chin-Wei Huang, David Krueger, Alexandre Lacoste, and Aaron Courville. Neural autoregressive flows. In Proceedings of the 35th International Conference on Machine Learning, pages 2078–2087, 2018.
  62. 62.Michael F. Hutchinson. A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines. Communications in Statistics—Simulation and Computation, 19 (2):433–450, 1990.
  63. 63.Aapo Hyvärinen and Petteri Pajunen. Nonlinear independent component analysis: Existence and uniqueness results. Neural Networks, 12(3):429–439, 1999.
  64. 64.David Inouye and Pradeep Ravikumar. Deep density destructors. In Proceedings of the 35th International Conference on Machine Learning, pages 2167–2175, 2018.
  65. 65.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on Machine Learning, pages 448–456, 2015.
  66. 66.Jörn-Henrik Jacobsen, Arnold W. M. Smeulders, and Edouard Oyallon. i-RevNet: Deep invertible networks. In International Conference on Learning Representations, 2018.
  67. 67.Jörn-Henrik Jacobsen, Jens Behrmann, Richard Zemel, and Matthias Bethge. Excessive invariance causes adversarial vulnerability. In International Conference on Learning Representations, 2019.
  68. 68.Priyank Jaini, Kira A. Selby, and Yaoliang Yu. Sum-of-squares polynomial flow. In Proceedings of the 36th International Conference on Machine Learning, pages 3009–3018, 2019.
  69. 69.Lifeng Jin, Finale Doshi-Velez, Timothy Miller, Lane Schwartz, and William Schuler. Unsupervised learning of PCFGs with normalizing flow. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2442–2452, 2019.
  70. 70.Richard M. Johnson. The minimal transformation to orthonormality. Psychometrika, 31: 61–66, 03 1966.
  71. 71.Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu. Video pixel networks. In Proceedings of the 34th International Conference on Machine Learning, pages 1771–1779, 2017.
  72. 72.Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon. FloWaveNet : A generative flow for raw audio. In Proceedings of the 36th International Conference on Machine Learning, pages 3370–3378, 2019.
  73. 73.Diederik P. Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1 × 1 convolutions. In Advances in Neural Information Processing Systems, pages 10215–10224, 2018.
  74. 74.Diederik P. Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations, 2014a.
  75. 75.Diederik P. Kingma and Max Welling. Efficient gradient-based inference through transformations between Bayes nets and neural nets. In Proceedings of the 31st International Conference on Machine Learning, pages 1782–1790, 2014b.
  76. 76.Diederik P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. Improved variational inference with inverse autoregressive flow. In Neural Information Processing Systems, pages 4743–4751, 2016.
  77. 77.Shoshichi Kobayashi and Katsumi Nomizu. Foundations of differential geometry, volume 1. Interscience Publishers, 1963.
  78. 78.Ivan Kobyzev, Simon Prince, and Marcus A. Brubaker. Normalizing flows: An introduction and review of current methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. doi: 10.1109/TPAMI.2020.2992934.
  79. 79.Jonas Köhler, Leon Klein, and Frank Noé. Equivariant flows: Sampling configurations for multi-body systems with symmetric energies. ArXiv Preprint arXiv:1910.00753, 2019.
  80. 80.Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Diederik P. Kingma. VideoFlow: A flow-based generative model for video. ICML Workshop on Invertible Neural Networks and Normalizing Flows, 2019.
  81. 81.Valero Laparra, Gustavo Camps-Valls, and Jesñs Malo. Iterative Gaussianization: From ICA to random rotations. IEEE Transactions on Neural Networks, 22(4):537–549, 2011.
  82. 82.Hugo Larochelle and Iain Murray. The neural autoregressive distribution estimator. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, pages 29–37, 2011.
  83. 83.Mario Lezcano-Casado and David Martínez-Rubio. Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group. In Proceedings of the 36th International Conference on Machine Learning, 2019.
  84. 84.Christos Louizos and Max Welling. Multiplicative normalizing flows for variational Bayesian neural networks. In Proceedings of the 34th International Conference on Machine Learning, pages 2218–2227, 2017.
  85. 85.Xuezhe Ma, Xiang Kong, Shanghang Zhang, and Eduard Hovy. MaCow: Masked convolutional generative flow. In Advances in Neural Information Processing Systems, 2019.
  86. 86.Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng. Rectifier nonlinearities improve neural network acoustic models. ICML Workshop on Deep Learning for Audio, Speech, and Language Processing, 2013.
  87. 87.Kaushalya Madhawa, Katushiko Ishiguro, Kosuke Nakago, and Motoki Abe. Graph-NVP: An invertible flow model for generating molecular graphs. ArXiv Preprint arXiv:1905.11600, 2019.
  88. 88.Murray Marshall. Positive polynomials and sums of squares. American Mathematical Society, 2008.
  89. 89.Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černocký, and Sanjeev Khudanpur. Recurrent neural network based language model. In Proceedings of the 11th Annual Conference of the International Speech Communication Association, 2010.
  90. 90.John W. Milnor and David W. Weaver. Topology from the differentiable viewpoint. Princeton University Press, 1997.
  91. 91.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018.
  92. 92.Shakir Mohamed and Balaji Lakshminarayanan. Learning in implicit generative models. NeurIPS Workshop on Adversarial Training, 2016.
  93. 93.Thomas Müller, Brian McWilliams, Fabrice Rousselle, Markus Gross, and Jan Novák. Neural importance sampling. ACM Transactions on Graphics, 38(5):145, 2019.
  94. 94.Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, Dilan Gorur, and Balaji Lakshminarayanan. Hybrid models with deep and invertible features. In Proceedings of the 36th International Conference on Machine Learning, pages 4723–4732, 2019.
  95. 95.Radford M. Neal. MCMC using Hamiltonian dynamics. Handbook of Markov Chain Monte Carlo, 54:113–162, 2010.
  96. 96.Frank Noé, Simon Olsson, Jonas Köhler, and Hao Wu. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365, 2019.
  97. 97.Junier Oliva, Avinava Dubey, Manzil Zaheer, Barnabas Poczos, Ruslan Salakhutdinov, Eric Xing, and Jeff Schneider. Transformation autoregressive networks. In Proceedings of the 35th International Conference on Machine Learning, pages 3898–3907, 2018.
  98. 98.George Papamakarios. Neural density estimation and likelihood-free inference. PhD thesis, University of Edinburgh, 2019. Available at https://arxiv.org/abs/1910.13233.
  99. 99.George Papamakarios, Theo Pavlakou, and Iain Murray. Masked autoregressive flow for density estimation. In Advances in Neural Information Processing Systems, pages 2338–2347, 2017.
  100. 100.George Papamakarios, David Sterratt, and Iain Murray. Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, pages 837–848, 2019.
  101. 101.Judea Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Elsevier Science, 1988. ISBN 9781558604797.
  102. 102.Lev Semenovich Pontryagin. Mathematical theory of optimal processes. Routledge, 1962.
  103. 103.Ryan Prenger, Rafael Valle, and Bryan Catanzaro. WaveGlow: A flow-based generative network for speech synthesis. In Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 3617–3621. IEEE, 2019.
  104. 104.Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, pages 1530–1538, 2015.
  105. 105.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the 31st International Conference on Machine Learning, pages 1278–1286, 2014.
  106. 106.Danilo Jimenez Rezende, Sébastien Racanière, Irina Higgins, and Peter Toth. Equivariant Hamiltonian flows. ArXiv Preprint arXiv:1909.13739, 2019.
  107. 107.Oren Rippel and Ryan Prescott Adams. High-dimensional probability estimation with deep density models. ArXiv Preprint arXiv:1302.5125, 2013.
  108. 108.Hannes Risken. Fokker–Planck equation. In The Fokker–Planck Equation: Methods of Solution and Applications, pages 63–95. Springer Berlin Heidelberg, 1996.
  109. 109.Murray Rosenblatt. Remarks on a multivariate transformation. The Annals of Mathematical Statistics, 23(3):470–472, 1952.
  110. 110.Walter Rudin. Principles of mathematical analysis. International Series in Pure and Applied Mathematics. McGraw-Hill, 1976.
  111. 111.Walter Rudin. Real and complex analysis. Tata McGraw-Hill Education, 2006.
  112. 112.Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma. PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications. In International Conference on Learning Representations, 2017.
  113. 113.Yannick Schroecker, Mel Vecerik, and Jon Scholz. Generative predecessor models for sample-efficient imitation learning. In International Conference on Learning Representations, 2019.
  114. 114.Ron Shepard, Scott R. Brozell, and Gergely Gidofalvi. The representation and parametrization of orthogonal matrices. The Journal of Physical Chemistry A, 119(28):7924–7939, 2015.
  115. 115.Abe Sklar. Fonctions de Répartition à N Dimensions et Leurs Marges. Université Paris, 1959.
  116. 116.Jiaming Song, Shengjia Zhao, and Stefano Ermon. A-NICE-MC: Adversarial training for MCMC. In Advances in Neural Information Processing Systems, volume 30, pages 5140–5150, 2017.
  117. 117.Yang Song, Chenlin Meng, and Stefano Ermon. MintNet: Building invertible neural networks with masked convolutions. In Advances in Neural Information Processing Systems, pages 11002–11012, 2019.
  118. 118.Endre Süli. Lecture notes on numerical solutions of ordinary differential equations, 2010.
  119. 119.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104–3112, 2014.
  120. 120.Esteban G. Tabak and Cristina V. Turner. A family of nonparametric density estimation algorithms. Communications on Pure and Applied Mathematics, 66(2):145–164, 2013.
  121. 121.Esteban G. Tabak and Eric Vanden-Eijnden. Density estimation by dual ascent of the log-likelihood. Communications in Mathematical Sciences, 8(1):217–233, 2010.
  122. 122.Lucas Theis and Matthias Bethge. Generative image modeling using spatial LSTMs. In Advances in Neural Information Processing Systems, pages 1927–1935, 2015.
  123. 123.Naftali Tishby and Noga Zaslavsky. Deep learning and the information bottleneck principle. In IEEE Information Theory Workshop, pages 1–5. IEEE, 2015.
  124. 124.Michalis K. Titsias. Learning model reparametrizations: Implicit variational inference by fitting MCMC distributions. ArXiv Preprint arXiv:1708.01529, 2017.
  125. 125.Jakub M. Tomczak and Max Welling. Improving variational auto-encoders using Householder flow. NeurIPS Workshop on Bayesian Deep Learning, 2016.
  126. 126.Dustin Tran, Keyon Vafa, Kumar Krishna Agrawal, Laurent Dinh, and Ben Poole. Discrete flows: Invertible generative models of discrete data. In Advances in Neural Information Processing Systems, 2019.
  127. 127.Benigno Uria, Iain Murray, and Hugo Larochelle. RNADE: The real-valued neural autoregressive density-estimator. In Advances in Neural Information Processing Systems, pages 2175–2183, 2013.
  128. 128.Benigno Uria, Iain Murray, and Hugo Larochelle. A deep and tractable density estimator. In Proceedings of the 31st International Conference on Machine Learning, pages 467–475, 2014.
  129. 129.Benigno Uria, Marc-Alexandre Côté, Karol Gregor, Iain Murray, and Hugo Larochelle. Neural autoregressive distribution estimation. Journal of Machine Learning Research, 17 (205):1–37, 2016.
  130. 130.Rianne van den Berg, Leonard Hasenclever, Jakub M. Tomczak, and Max Welling. Sylvester normalizing flows for variational inference. The 34th Conference on Uncertainty in Artificial Intelligence, 2018.
  131. 131.Rianne van den Berg, Alexey A. Gritsenko, Mostafa Dehghani, Casper Kaae Sønderby, and Tim Salimans. IDF++: Analyzing and improving integer discrete flows for lossless compression. ArXiv Preprint arXiv:2006.12459, 2020.
  132. 132.Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. ArXiv Preprint arXiv:1609.03499, 2016a.
  133. 133.Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In Proceedings of The 33rd International Conference on Machine Learning, pages 1747–1756, 2016b.
  134. 134.Aäron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with PixelCNN decoders. In Advances in Neural Information Processing Systems, pages 4797–4805, 2016c.
  135. 135.Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Dan Belov, and Demis Hassabis. Parallel WaveNet: Fast high-fidelity speech synthesis. In Proceedings of the 35th International Conference on Machine Learning, pages 3918–3926, 2018.
  136. 136.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017.
  137. 137.Cédric Villani. Optimal transport: Old and new, volume 338. Springer Science & Business Media, 2008.
  138. 138.Martin J. Wainwright and Michael I. Jordan. Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning, 1(1–2):1–305, 2008.
  139. 139.Prince Zizhuang Wang and William Yang Wang. Riemannian normalizing flow on variational Wasserstein autoencoder for text modeling. ArXiv Preprint arXiv:1904.02399, 2019.
  140. 140.Patrick Nadeem Ward, Ariella Smofsky, and Avishek Joey Bose. Improving exploration in soft-actor-critic with normalizing flows policies. ICML Workshop on Invertible Neural Networks and Normalizing Flows, 2019.
  141. 141.Antoine Wehenkel and Gilles Louppe. Unconstrained monotonic neural networks. In Advances in Neural Information Processing Systems, 2019.
  142. 142.Christina Winkler, Daniel E. Worrall, Emiel Hoogeboom, and Max Welling. Learning likelihoods with conditional normalizing flows. ArXiv Preprint arXiv:1912.00042, 2019.
  143. 143.Peter Wirnsberger, Andrew J. Ballard, George Papamakarios, Stuart Abercrombie, Sébastien Racanière, Alexander Pritzel, Danilo Jimenez Rezende, and Charles Blundell. Targeted free energy estimation via learned mappings. The Journal of Chemical Physics, 153(14):144112, 2020.
  144. 144.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge J. Belongie, and Bharath Hariharan. PointFlow: 3D point cloud generation with continuous normalizing flows. In Proceedings of the International Conference on Computer Vision, 2019.
  145. 145.Chunting Zhou, Xuezhe Ma, Di Wang, and Graham Neubig. Density matching for bilingual word embedding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, pages 1588–1598, 2019.
  146. 146.Zachary Ziegler and Alexander Rush. Latent normalizing flows for discrete sequences. In Proceedings of the 36th International Conference on Machine Learning, pages 7673–7682, 2019.

Citation

MLA
Papamakarios, G., et al. “Normalizing Flows for Probabilistic Modeling and Inference”. arXiv, 2019, https://doi.org/10.48550/arxiv.1912.02762.
APA
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., & Lakshminarayanan, B. (2019). Normalizing Flows for Probabilistic Modeling and Inference. arXiv. https://doi.org/10.48550/arxiv.1912.02762
Chicago
Papamakarios, G., E. Nalisnick, D. J. Rezende, S. Mohamed, and B. Lakshminarayanan. 2019. “Normalizing Flows for Probabilistic Modeling and Inference”. arXiv, ahead of print. https://doi.org/10.48550/arxiv.1912.02762.
Harvard
Papamakarios, G. et al. (2019) “Normalizing Flows for Probabilistic Modeling and Inference”, arXiv [Preprint]. Available at: https://doi.org/10.48550/arxiv.1912.02762.
Vancouver
1. Papamakarios G, Nalisnick E, Rezende DJ, Mohamed S, Lakshminarayanan B (2019) Normalizing Flows for Probabilistic Modeling and Inference. arXiv. https://doi.org/10.48550/arxiv.1912.02762

BibTeX

@article{https://doi.org/10.48550/arxiv.1912.02762,
  doi = {10.48550/ARXIV.1912.02762},
  url = {https://arxiv.org/abs/1912.02762},
  author = {Papamakarios, George and Nalisnick, Eric and Rezende, Danilo Jimenez and Mohamed, Shakir and Lakshminarayanan, Balaji},
  keywords = {Machine Learning (stat.ML), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Normalizing Flows for Probabilistic Modeling and Inference},
  publisher = {arXiv},
  year = {2019},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/