Normalizing Flows: An Introduction and Review of Current Methods

Ivan KobyzevSimon J.D. PrinceMarcus A. Brubaker

article2020TPAMI1,602 citations

Explains the mathematical foundations and architectural design principles of normalizing flows, offering a clear guide to how these models achieve exact probability density estimation and efficient sampling.

Listen

Generative modeling—the ability of artificial intelligence systems to learn and simulate complex, real-world probability distributions from unlabelled data—is critical for tasks such as anomaly detection, synthetic media creation, and scientific data summarization. Traditional generative architectures like Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) have achieved impressive qualitative results but suffer from key operational liabilities, including training instability, mode collapse, and an inability to compute the exact probability density of new observations. The article provides a comprehensive review of Normalizing Flows (NFs), an emerging class of generative models that mathematically transform simple, known base distributions into complex target densities using invertible and differentiable operations, thereby enabling exact, efficient density evaluation and direct data generation.

The article systematically analyzes the mathematical foundations, core design architectures, and experimental performance of normalizing flows across standard tabular benchmarks (such as power consumption and sensor measurements) and visual datasets (including MNIST, CIFAR-10, and ImageNet). The analyzed frameworks are categorized into several functional families: basic linear and elementwise transformations, structured coupling and autoregressive flows, invertible residual networks, and continuous formulations based on ordinary differential equations (ODEs) and stochastic differential equations (SDEs).

The review establishes several critical findings regarding model performance and trade-offs. First, universal flows—specifically those utilizing advanced coupling functions like monotonic splines, polynomials, and unconstrained neural networks—substantially outperform simple affine models on tabular density estimation. Second, on complex image benchmarks, architectural design and preprocessing choices are decisive; the Flow++ model achieved state-of-the-art results (reaching 3.08 bits per dimension on CIFAR-10) largely due to the implementation of learned variational dequantization rather than standard uniform noise addition. Third, continuous ordinary differential equation flows, such as FFJORD, demonstrate remarkable parameter efficiency, achieving competitive image-modeling performance with less than 2% of the parameters required by large discrete flow models like Glow. Finally, distinct architectures impose operational trade-offs: masked autoregressive models provide rapid density estimation during training but slow sequential sampling, whereas inverse autoregressive flows enable fast generation but costly density evaluation.

These findings have direct strategic implications for deploying generative machine learning systems. Normalizing flows eliminate the guesswork of approximate inference, significantly reducing the operational risks associated with training instability and unquantified uncertainty in safety-critical domains such as audio processing, medical imaging, and physics simulations. However, organizations must align their architectural selection with business requirements: systems requiring real-time data generation should avoid standard autoregressive flows in favor of inverse autoregressive or coupling models, while resource-constrained environments can leverage continuous ODE flows to minimize parameter footprints and hardware demands.

To advance the deployment of these technologies, engineering and research teams should focus on several immediate technical priorities. Practitioners should adopt expressive spline or neural-network coupling functions over basic affine layers and implement learned dequantization when handling discrete or ordinal data. For broader applicability, future research must address key theoretical and practical limitations, particularly adapting continuous flow theory to discrete domains like text processing, extending models to non-Euclidean geometries and Riemannian manifolds (for applications in robotics and physical sciences), and exploring alternative loss functions derived from optimal transport theory. While confidence in the evaluated continuous and tabular benchmarks is high, readers should exercise caution when applying current continuous flow methods to discrete datasets or manifold data, where standard Euclidean assumptions fail and robust solutions remain under active development.

arXiv: 1908.09257
  • Paper: Normalizing Flows for Probabilistic Modeling and Inference, George Papamakarios et al. (2019). This comprehensive review covers the foundational mathematics, core design principles, and taxonomy of invertible transformations that underlie normalizing flows.
  • Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). This seminal paper introduced normalizing flows for variational inference, formulating the change-of-variables framework using planar and radial transformations.
  • Paper: NICE: Non-linear Independent Components Estimation, Laurent Dinh et al. (2014). It introduces the coupling layer mechanism that enables tractable Jacobian determinants and invertible architectures central to modern flow methods.
  • Paper: Density estimation using Real NVP, Laurent Dinh et al. (2016). This foundational work establishes affine coupling layers and multi-scale architectures essential for scaling normalizing flows to complex, high-dimensional distributions.
  • Paper: Improving Variational Inference with Inverse Autoregressive Flow, Diederik P. Kingma et al. (2016). It introduces inverse autoregressive flows, providing a key autoregressive transformation technique heavily discussed throughout the flow literature.
  • Paper: Glow: Generative Flow with Invertible 1x1 Convolutions, Diederik P. Kingma et al. (2018). It develops the Glow architecture using invertible 1x1 convolutions, representing a primary milestone in generative flow design reviewed by the source.
  • Paper: Neural Ordinary Differential Equations, Ricky T. Q. Chen et al. (2018). This foundational text introduces continuous-time normalizing flows governed by neural ordinary differential equations.
  • Paper: An Introduction to Variational Autoencoders, Diederik P. Kingma et al. (2019). It provides the requisite probabilistic modeling context and variational lower-bound formulations where normalizing flows are frequently applied.
Cover for Normalizing Flows: An Introduction and Review of Current Methods

Abstract

Normalizing Flows are generative models which produce tractable distributions where both sampling and density evaluation can be efficient and exact. The goal of this survey article is to give a coherent and comprehensive review of the literature around the construction and use of Normalizing Flows for distribution learning. We aim to provide context and explanation of the models, review current state-of-the-art literature, and identify open questions and promising future directions.

Table of Contents

  • I Introduction
  • II Background
  • II-A Basics
  • II-A1 More formal construction
  • II-B Applications
  • II-B1 Density estimation and sampling
  • II-B2 Variational Inference
  • III Methods
  • III-A Elementwise Flows
  • III-B Linear Flows
  • III-B1 Diagonal
  • III-B2 Triangular
  • III-B3 Permutation and Orthogonal
  • III-B4 Factorizations
  • III-B5 Convolution
  • III-C Planar and Radial Flows
  • III-C1 Planar Flows
  • III-C2 Radial Flows
  • III-D Coupling and Autoregressive Flows
  • III-D1 Coupling Flows
  • III-D2 Autoregressive Flows
  • III-D3 Universality
  • III-D4 Coupling Functions
  • III-E Residual Flows
  • III-F Infinitesimal (Continuous) Flows
  • III-F1 ODE-based methods
  • III-F2 SDE-based methods (Langevin flows)
  • IV Datasets and performance
  • IV-A Tabular datasets
  • IV-B Image datasets
  • V Discussion and open problems
  • V-A Inductive biases
  • V-A1 Role of the base measure
  • V-A2 Form of diffeomorphisms
  • V-A3 Loss function
  • V-B Generalisation to non-Euclidean spaces
  • V-B1 Flows on manifolds.
  • V-B2 Discrete distributions
  • References

Knowls

  1. Knowl 1 — Change of Variables and Density Transformation in Normalizing Flows

    definition

    A Normalizing Flow transforms a simple base probability distribution into a complex target distribution via a sequence of invertible and differentiable transformations.

    Let Z∈RD\mathbf{Z} \in \mathbb{R}^D be a continuous random variable with a tractable base probability density function pZ:RD→Rp_Z : \mathbb{R}^D \to \mathbb{R}. Let g:RD→RDg : \mathbb{R}^D \to \mathbb{R}^D be a diffeomorphism (a bijection that is differentiable and has a differentiable inverse f=g−1f = g^{-1}). The pushforward probability density function pY:RD→Rp_Y : \mathbb{R}^D \to \mathbb{R} of the random variable Y=g(Z)\mathbf{Y} = g(\mathbf{Z}) is given by the change of variables formula:

    pY(y)=pZ(f(y))∣det⁡Df(y)∣=pZ(f(y))∣det⁡Dg(f(y))∣−1p_Y(\mathbf{y}) = p_Z(f(\mathbf{y})) |\det Df(\mathbf{y})| = p_Z(f(\mathbf{y})) |\det Dg(f(\mathbf{y}))|^{-1}

    where Df(y)=∂f∂yDf(\mathbf{y}) = \frac{\partial f}{\partial \mathbf{y}} is the Jacobian matrix of the inverse mapping ff (the normalizing direction), and Dg(z)=∂g∂zDg(\mathbf{z}) = \frac{\partial g}{\partial \mathbf{z}} is the Jacobian matrix of gg (the generative direction).

    When gg is constructed as a composition of NN bijective transformations g=gN∘gN−1∘⋯∘g1g = g_N \circ g_{N-1} \circ \dots \circ g_1 with inverse f=f1∘⋯∘fN−1∘fNf = f_1 \circ \dots \circ f_{N-1} \circ f_N, the overall Jacobian determinant factors as:

    det⁡Df(y)=∏i=1Ndet⁡Dfi(xi)\det Df(\mathbf{y}) = \prod_{i=1}^N \det Df_i(\mathbf{x}_i)

    where xi=gi∘⋯∘g1(z)=fi+1∘⋯∘fN(y)\mathbf{x}_i = g_i \circ \dots \circ g_1(\mathbf{z}) = f_{i+1} \circ \dots \circ f_N(\mathbf{y}) denotes the intermediate representation at step ii, with x0=z\mathbf{x}_0 = \mathbf{z} and xN=y\mathbf{x}_N = \mathbf{y}.

  2. Knowl 2 — Universality of Autoregressive Normalizing Flows

    theoretical result

    Autoregressive normalizing flows are universal density approximators under specific conditions on their coupling functions.

    Let μ\mu and ν\nu be absolutely continuous Borel probability measures on RD\mathbb{R}^D (or on [0,1]D[0, 1]^D). A transformation T=(T1,…,TD):RD→RDT = (T_1, \dots, T_D) : \mathbb{R}^D \to \mathbb{R}^D is triangular if each component TiT_i is a function of x1:i=(x1,…,xi)\mathbf{x}_{1:i} = (x_1, \dots, x_i) for i∈{1,…,D}i \in \{1, \dots, D\}. TT is an increasing triangular transformation if TiT_i is strictly increasing in xix_i for every ii.

    1. Existence and Uniqueness (Bogachev et al., 2005): There exists an increasing triangular transformation T:RD→RDT : \mathbb{R}^D \to \mathbb{R}^D such that ν=T∗μ\nu = T_*\mu (where T∗μT_*\mu denotes the pushforward measure μ(T−1(⋅))\mu(T^{-1}(\cdot))). This transformation TT is unique up to null sets of μ\mu.

    2. Weak Convergence of Pushforwards: If μ\mu is an absolutely continuous Borel probability measure on RD\mathbb{R}^D and {Tn}\{T_n\} is a sequence of measurable maps RD→RD\mathbb{R}^D \to \mathbb{R}^D that converges pointwise to a map TT, then the sequence of pushforward measures (Tn)∗μ(T_n)_*\mu converges weakly to T∗μT_*\mu.

    3. Universality Corollary: A class of autoregressive flows g(⋅;θ):RD→RDg(\cdot; \theta) : \mathbb{R}^D \to \mathbb{R}^D using scalar coupling functions h(⋅;θ):R→Rh(\cdot; \theta) : \mathbb{R} \to \mathbb{R} is universal (capable of approximating any target density arbitrarily well) if the family of coupling functions {h(⋅;θ)}\{h(\cdot; \theta)\} is dense in the set of all strictly monotone univariate functions under the pointwise convergence topology.

  3. Knowl 3 — Coupling Flow Architecture and Multi-Scale Transformations

    model/method

    A coupling flow partitions an input vector x∈RD\mathbf{x} \in \mathbb{R}^D into two disjoint subvectors xA∈Rd\mathbf{x}^A \in \mathbb{R}^d and xB∈RD−d\mathbf{x}^B \in \mathbb{R}^{D-d}. Given a parameterized bijection h(⋅;θ):Rd→Rdh(\cdot; \theta) : \mathbb{R}^d \to \mathbb{R}^d called the coupling function and an arbitrary conditioning function Θ:RD−d→Rp\Theta : \mathbb{R}^{D-d} \to \mathbb{R}^p, the coupling forward transformation g:RD→RDg : \mathbb{R}^D \to \mathbb{R}^D is:

    yA=h(xA;Θ(xB)),yB=xB\mathbf{y}^A = h(\mathbf{x}^A; \Theta(\mathbf{x}^B)), \quad \mathbf{y}^B = \mathbf{x}^B

    The inverse transformation is computed as:

    xA=h−1(yA;Θ(yB)),xB=yB\mathbf{x}^A = h^{-1}(\mathbf{y}^A; \Theta(\mathbf{y}^B)), \quad \mathbf{x}^B = \mathbf{y}^B

    The Jacobian matrix Dg(x)Dg(\mathbf{x}) is block lower-triangular:

    Dg(x)=[DxAh(xA;Θ(xB))DxBh(xA;Θ(xB))0ID−d]Dg(\mathbf{x}) = \begin{bmatrix} D_{\mathbf{x}^A} h(\mathbf{x}^A; \Theta(\mathbf{x}^B)) & D_{\mathbf{x}^B} h(\mathbf{x}^A; \Theta(\mathbf{x}^B)) \\ \mathbf{0} & \mathbf{I}_{D-d} \end{bmatrix}

    Hence, det⁡Dg(x)=det⁡DxAh(xA;Θ(xB))\det Dg(\mathbf{x}) = \det D_{\mathbf{x}^A} h(\mathbf{x}^A; \Theta(\mathbf{x}^B)). When hh applies a scalar bijection elementwise, DxAhD_{\mathbf{x}^A} h is diagonal, and the Jacobian determinant is the product of dd univariate derivatives:

    det⁡Dg(x)=∏i=1d∂hi(xiA;Θi(xB))∂xiA\det Dg(\mathbf{x}) = \prod_{i=1}^d \frac{\partial h_i(x_i^A; \Theta_i(\mathbf{x}^B))}{\partial x_i^A}

    In multi-scale architectures (such as RealNVP), dimensions are split and partially routed directly to the base distribution at intermediate steps, reducing computation on high-dimensional data.

  4. Knowl 4 — Direct Autoregressive Flows (MAF) and Inverse Autoregressive Flows (IAF)

    model/method

    Autoregressive flows enforce a triangular Jacobian by making each output coordinate depend on previous dimensions relative to a fixed ordering. Let h(⋅;θ):R→Rh(\cdot; \theta) : \mathbb{R} \to \mathbb{R} be a strictly monotone scalar coupling function and Θt\Theta_t be conditioning functions.

    1. Masked Autoregressive Flow (MAF / Direct Autoregressive Flow):

    yt=h(xt;Θt(x1:t−1))for t=1,…,Dy_t = h(x_t; \Theta_t(\mathbf{x}_{1:t-1})) \quad \text{for } t = 1, \dots, D

    where x1:t−1=(x1,…,xt−1)\mathbf{x}_{1:t-1} = (x_1, \dots, x_{t-1}) and Θ1\Theta_1 is a constant. The Jacobian is triangular with determinant:

    det⁡Dg(x)=∏t=1D∂yt∂xt\det Dg(\mathbf{x}) = \prod_{t=1}^D \frac{\partial y_t}{\partial x_t}

    Forward evaluation g(x)g(\mathbf{x}) (or computing likelihood f(y)f(\mathbf{y}) when MAF is defined in the normalizing direction) is parallelizable across all DD dimensions using masked neural networks (e.g., MADE). However, inversion requires sequential computation xt=h−1(yt;Θt(x1:t−1))x_t = h^{-1}(y_t; \Theta_t(\mathbf{x}_{1:t-1})) taking O(D)\mathcal{O}(D) sequential passes.

    1. Inverse Autoregressive Flow (IAF):

    yt=h(xt;Θt(y1:t−1))for t=1,…,Dy_t = h(x_t; \Theta_t(\mathbf{y}_{1:t-1})) \quad \text{for } t = 1, \dots, D

    In IAF, sampling g(z)g(\mathbf{z}) from base noise requires sequential evaluation across DD dimensions, whereas computing the inverse f(y)f(\mathbf{y}) and its log-determinant can be parallelized in a single pass. IAF is optimized for fast sampling (e.g., variational inference), whereas MAF is optimized for fast density estimation on data.

  5. Knowl 5 — Continuous Normalizing Flows and Neural Ordinary Differential Equations

    model/method

    Continuous Normalizing Flows specify the generative mapping y=g(z)\mathbf{y} = g(\mathbf{z}) as the time-1 map x(1)=Φ1(z)\mathbf{x}(1) = \Phi_1(\mathbf{z}) of an ordinary differential equation (ODE):

    ddtx(t)=F(x(t),θ(t)),x(0)=z,t∈[0,1]\frac{d}{dt}\mathbf{x}(t) = F(\mathbf{x}(t), \theta(t)), \quad \mathbf{x}(0) = \mathbf{z}, \quad t \in [0, 1]

    where F:RD×Θ→RDF : \mathbb{R}^D \times \Theta \to \mathbb{R}^D is Lipschitz continuous in x\mathbf{x} and continuous in tt.

    1. Instantaneous Change of Variables: The evolution of the log probability density under the continuous ODE dynamics satisfies:

    ddtlog⁡p(x(t))=−Tr⁡(∂F(x(t),θ(t))∂x(t))\frac{d}{dt} \log p(\mathbf{x}(t)) = -\operatorname{Tr}\left( \frac{\partial F(\mathbf{x}(t), \theta(t))}{\partial \mathbf{x}(t)} \right)

    This replaces matrix determinant calculation with a trace operation. In FFJORD, an unbiased stochastic estimate of the trace is computed using the Hutchinson estimator Tr⁡(A)=Ep(v)[vTAv]\operatorname{Tr}(A) = \mathbb{E}_{p(\mathbf{v})}[\mathbf{v}^T A \mathbf{v}], where E[v]=0\mathbb{E}[\mathbf{v}] = \mathbf{0} and Cov⁡(v)=I\operatorname{Cov}(\mathbf{v}) = \mathbf{I}.

    1. Adjoint Sensitivity Method: Gradients of a downstream loss L(x(1))\mathcal{L}(\mathbf{x}(1)) with respect to parameters θ(t)\theta(t) are computed backward in time via the adjoint state a(t)=∂L∂x(t)\mathbf{a}(t) = \frac{\partial \mathcal{L}}{\partial \mathbf{x}(t)}, which satisfies da(t)dt=−a(t)∂F(x(t),θ(t))∂x(t)\frac{d\mathbf{a}(t)}{dt} = -\mathbf{a}(t) \frac{\partial F(\mathbf{x}(t), \theta(t))}{\partial \mathbf{x}(t)}.

    2. Augmented Neural ODE (ANODE): Standard ODE flows can only learn orientation-preserving diffeomorphisms (where det⁡Dg>0\det Dg > 0). Augmented Neural ODEs concatenate pp auxiliary dimensions x^(t)∈Rp\hat{\mathbf{x}}(t) \in \mathbb{R}^p with initial condition x^(0)=0\hat{\mathbf{x}}(0) = \mathbf{0}, enabling representation of arbitrary diffeomorphisms as time-1 maps.

  6. Knowl 6 — Invertible Residual Networks and Unbiased Residual Flows

    model/method

    Residual flows construct invertible mappings via residual connections g(x)=x+F(x)g(\mathbf{x}) = \mathbf{x} + F(\mathbf{x}), where F:RD→RDF : \mathbb{R}^D \to \mathbb{R}^D is a parameterized neural network.

    1. Invertibility Condition: If the Lipschitz constant of the residual block satisfies Lip⁡(F)<1\operatorname{Lip}(F) < 1, the residual mapping g(x)g(\mathbf{x}) is guaranteed to be bijective. Inversion cannot be computed in closed form but is obtained numerically via fixed-point iteration x(k+1)=y−F(x(k))\mathbf{x}^{(k+1)} = \mathbf{y} - F(\mathbf{x}^{(k)}), which converges by the Banach fixed-point theorem.

    2. Log-Determinant Power Series: The Jacobian is Dg=I+DFDg = \mathbf{I} + DF. Under Lip⁡(F)<1\operatorname{Lip}(F) < 1, ln⁡∣det⁡(I+DF)∣=Tr⁡(ln⁡(I+DF))\ln |\det(\mathbf{I} + DF)| = \operatorname{Tr}(\ln(\mathbf{I} + DF)), which expands into an absolutely convergent series:

    Tr⁡(ln⁡(I+DF))=∑k=1∞(−1)k+1kTr⁡((DF)k)\operatorname{Tr}(\ln(\mathbf{I} + DF)) = \sum_{k=1}^\infty \frac{(-1)^{k+1}}{k} \operatorname{Tr}\left((DF)^k\right)

    1. Trace Estimation and Russian Roulette: In iResNet, the power series is truncated and traces are estimated using the Hutchinson estimator Tr⁡(A)=Ev[vTAv]\operatorname{Tr}(A) = \mathbb{E}_{\mathbf{v}}[\mathbf{v}^T A \mathbf{v}], resulting in a biased log-determinant estimate. Residual Flow introduces an unbiased estimator by applying a Russian roulette randomized stopping rule to evaluate the infinite series without truncation bias.
  7. Knowl 7 — Scalar Coupling Functions in Normalizing Flows

    model/method

    In coupling and autoregressive architectures, expressivity depends on the chosen scalar coupling function h(⋅;θ):R→Rh(\cdot; \theta) : \mathbb{R} \to \mathbb{R}:

    1. Affine Coupling:

    h(x;θ)=θ1x+θ2,θ1≠0,  θ2∈Rh(x; \theta) = \theta_1 x + \theta_2, \quad \theta_1 \neq 0, \; \theta_2 \in \mathbb{R}

    1. Continuous Mixture CDFs (Flow++):

    h(x;θ)=θ1F(x;π,μ,s)+θ2,F(x;π,μ,s)=σ−1(∑j=1Kπjσ(x−μjsj))h(x; \theta) = \theta_1 F(x; \boldsymbol{\pi}, \boldsymbol{\mu}, \mathbf{s}) + \theta_2, \quad F(x; \boldsymbol{\pi}, \boldsymbol{\mu}, \mathbf{s}) = \sigma^{-1}\left( \sum_{j=1}^K \pi_j \sigma\left( \frac{x - \mu_j}{s_j} \right) \right)

    where σ\sigma is the logistic sigmoid function, π∈RK\boldsymbol{\pi} \in \mathbb{R}^K is a probability vector, μ∈RK\boldsymbol{\mu} \in \mathbb{R}^K, and s∈R+K\mathbf{s} \in \mathbb{R}^K_+. Inversion is performed numerically with the bisection algorithm.

    1. Rational Quadratic Splines (RQ-NSF): h(x;θ)h(x; \theta) is defined as a monotone rational-quadratic spline on a compact interval [−B,B][-B, B] with KK bins parameterized by predicted knot coordinates and derivatives, and set to identity outside [−B,B][-B, B]. Inversion is computed analytically by solving a quadratic equation.

    2. Neural Autoregressive Functions (NAF & UMNN): NAF models h(x;θ)h(x; \theta) as an MLP with non-negative weights and strictly monotone activations. Unconstrained Monotonic Neural Networks (UMNN) model h(x;θ)=c+∫0xf(u;θ)duh(x; \theta) = c + \int_0^x f(u; \theta) du, where f(u;θ)>0f(u; \theta) > 0 is a strictly positive neural network evaluated via numerical integration.

    3. Sum-of-Squares (SOS) Polynomials:

    h(x;θ)=c+∫0x∑k=1K(∑l=0Laklul)2duh(x; \theta) = c + \int_0^x \sum_{k=1}^K \left( \sum_{l=0}^L a_{kl} u^l \right)^2 du

    where KK and LL are degree hyperparameters, guaranteeing strict monotonicity without weight constraints.

  8. Knowl 8 — Sylvester and Planar Normalizing Flows

    model/method

    Planar and Sylvester flows construct invertible transformations with linear-time or low-complexity Jacobian determinants:

    1. Planar Flow:

    g(x)=x+uh(wTx+b)g(\mathbf{x}) = \mathbf{x} + \mathbf{u} h(\mathbf{w}^T \mathbf{x} + b)

    where u,w∈RD\mathbf{u}, \mathbf{w} \in \mathbb{R}^D, b∈Rb \in \mathbb{R}, and h:R→Rh : \mathbb{R} \to \mathbb{R} is a smooth nonlinearity. Applying the matrix determinant lemma yields the Jacobian determinant in O(D)\mathcal{O}(D) operations:

    det⁡Dg(x)=det⁡(ID+uh′(wTx+b)wT)=1+h′(wTx+b)uTw\det Dg(\mathbf{x}) = \det\left(\mathbf{I}_D + \mathbf{u} h'(\mathbf{w}^T \mathbf{x} + b) \mathbf{w}^T\right) = 1 + h'(\mathbf{w}^T \mathbf{x} + b) \mathbf{u}^T \mathbf{w}

    Planar flows create a single-unit bottleneck along direction w\mathbf{w}.

    1. Sylvester Flow: Resolves the bottleneck by expanding the hidden dimension to M≤DM \le D:

    g(x)=x+Uh(WTx+b)g(\mathbf{x}) = \mathbf{x} + \mathbf{U} h(\mathbf{W}^T \mathbf{x} + \mathbf{b})

    where U,W∈RD×M\mathbf{U}, \mathbf{W} \in \mathbb{R}^{D \times M}, b∈RM\mathbf{b} \in \mathbb{R}^M, and h:RM→RMh : \mathbb{R}^M \to \mathbb{R}^M acts elementwise. By Sylvester's determinant identity, the Jacobian determinant is computed in O(M3+M2D)\mathcal{O}(M^3 + M^2 D) time via:

    det⁡Dg(x)=det⁡(IM+diag⁡(h′(WTx+b))WTU)\det Dg(\mathbf{x}) = \det\left(\mathbf{I}_M + \operatorname{diag}(h'(\mathbf{W}^T \mathbf{x} + \mathbf{b})) \mathbf{W}^T \mathbf{U}\right)

  9. Knowl 9 — Piecewise-Bijective Coupling (RAD)

    model/method

    Real and Discrete (RAD) flows allow normalizing flows to use piecewise-bijective, non-injective coupling transformations to model multimodal and topologically disconnected densities.

    The data domain is partitioned into KK disjoint subsets R=⨆k=1KAk\mathbb{R} = \bigsqcup_{k=1}^K A_k such that the restriction of h(⋅;θ)h(\cdot; \theta) to each subset AkA_k is strictly injective. A lookup function ϕ:R→{1,…,K}\phi : \mathbb{R} \to \{1, \dots, K\} maps y∈Aky \in A_k to index kk. The non-injective map is unfolded into y↦(h(y),ϕ(y))∈R×{1,…,K}y \mapsto (h(y), \phi(y)) \in \mathbb{R} \times \{1, \dots, K\}.

    The target probability density pY(y)p_Y(y) is evaluated by defining a joint density on the unfolded space pZ,[K](z,k)=p[K]∣Z(k∣z)pZ(z)p_{Z, [K]}(z, k) = p_{[K]|Z}(k|z) p_Z(z):

    pY(y)=pZ,[K](h(y),ϕ(y))∣Dh(y)∣=p[K]∣Z(ϕ(y)∣h(y))pZ(h(y))∣Dh(y)∣p_Y(y) = p_{Z, [K]}(h(y), \phi(y)) |Dh(y)| = p_{[K]|Z}(\phi(y) \mid h(y)) p_Z(h(y)) |Dh(y)|

    where p[K]∣Z(k∣z)p_{[K]|Z}(k \mid z) is parameterized by a trainable neural gating network. For generative sampling, a discrete component k∼p[K]∣Z(⋅∣z)k \sim p_{[K]|Z}(\cdot \mid z) is sampled from the gating network conditioned on z∼pZz \sim p_Z, and the corresponding local inverse branch (h∣Ak)−1(z)(h|_{A_k})^{-1}(z) is applied.

  10. Knowl 10 — Langevin Flows and Neural Stochastic Differential Equations

    model/method

    Langevin flows transform a data distribution pY(y)p_Y(\mathbf{y}) into a base distribution pZ(z)p_Z(\mathbf{z}) via a continuous Itô stochastic differential equation (SDE):

    dx(t)=b(x(t),t)dt+σ(x(t),t)dBt,t∈[0,T]d\mathbf{x}(t) = \mathbf{b}(\mathbf{x}(t), t) dt + \boldsymbol{\sigma}(\mathbf{x}(t), t) d\mathbf{B}_t, \quad t \in [0, T]

    where b(x,t)∈RD\mathbf{b}(\mathbf{x}, t) \in \mathbb{R}^D is the drift coefficient, σ(x,t)∈RD×D\boldsymbol{\sigma}(\mathbf{x}, t) \in \mathbb{R}^{D \times D} is the diffusion coefficient, and Bt\mathbf{B}_t is DD-dimensional standard Brownian motion.

    1. Forward Evolution: The time-dependent probability density p(x,t)p(\mathbf{x}, t) satisfies the Fokker-Planck (Kolmogorov forward) equation:

    ∂∂tp(x,t)=−∇x⋅(b(x,t)p(x,t))+∑i,j∂2∂xi∂xj(Dij(x,t)p(x,t))\frac{\partial}{\partial t} p(\mathbf{x}, t) = -\nabla_\mathbf{x} \cdot (\mathbf{b}(\mathbf{x}, t) p(\mathbf{x}, t)) + \sum_{i,j} \frac{\partial^2}{\partial x_i \partial x_j} \left( D_{ij}(\mathbf{x}, t) p(\mathbf{x}, t) \right)

    where D=12σσT\mathbf{D} = \frac{1}{2} \boldsymbol{\sigma} \boldsymbol{\sigma}^T and p(⋅,0)=pY(⋅)p(\cdot, 0) = p_Y(\cdot).

    1. Time-Reversed Evolution: The reverse evolution from t=Tt=T down to 00 satisfies Kolmogorov's backward equation:

    −∂∂tp(x,t)=b(x,t)⋅∇xp(x,t)+∑i,jDij(x,t)∂2p(x,t)∂xi∂xj-\frac{\partial}{\partial t} p(\mathbf{x}, t) = \mathbf{b}(\mathbf{x}, t) \cdot \nabla_\mathbf{x} p(\mathbf{x}, t) + \sum_{i,j} D_{ij}(\mathbf{x}, t) \frac{\partial^2 p(\mathbf{x}, t)}{\partial x_i \partial x_j}

    with terminal condition p(⋅,T)=pZ(⋅)p(\cdot, T) = p_Z(\cdot).

    1. Neural SDEs: Drift and diffusion are parameterized by neural networks and trained via adjoint SDE solvers or discretized MCMC diffusion paths matching forward and backward transition kernels.
  11. Knowl 11 — Density Estimation Benchmarks on Tabular Datasets

    data/table

    The table below compares the average test log-likelihood (in nats; higher is better) for unconditional density estimation across five standard preprocessed UCI and BSDS300 tabular datasets: POWER (D=6D=6), GAS (D=8D=8), HEPMASS (D=21D=21), MINIBOONE (D=43D=43), and BSDS300 (D=63D=63).

    Model POWER GAS HEPMASS MINIBOONE BSDS300
    MAF(5) 0.14±0.010.14 \pm 0.01 9.07±0.029.07 \pm 0.02 −17.70±0.02-17.70 \pm 0.02 −11.75±0.44-11.75 \pm 0.44 155.69±0.28155.69 \pm 0.28
    MAF(10) 0.24±0.010.24 \pm 0.01 10.08±0.0210.08 \pm 0.02 −17.73±0.02-17.73 \pm 0.02 −12.24±0.45-12.24 \pm 0.45 154.93±0.28154.93 \pm 0.28
    MAF MoG 0.30±0.010.30 \pm 0.01 9.59±0.029.59 \pm 0.02 −17.39±0.02-17.39 \pm 0.02 −11.68±0.44-11.68 \pm 0.44 156.36±0.28156.36 \pm 0.28
    RealNVP(5) −0.02±0.01-0.02 \pm 0.01 4.78±1.84.78 \pm 1.8 −19.62±0.02-19.62 \pm 0.02 −13.55±0.49-13.55 \pm 0.49 152.97±0.28152.97 \pm 0.28
    RealNVP(10) 0.17±0.010.17 \pm 0.01 8.33±0.148.33 \pm 0.14 −18.71±0.02-18.71 \pm 0.02 −13.84±0.52-13.84 \pm 0.52 153.28±1.78153.28 \pm 1.78
    Glow 0.170.17 8.158.15 −18.92-18.92 −11.35-11.35 155.07155.07
    FFJORD 0.460.46 8.598.59 −14.92-14.92 −10.43-10.43 157.40157.40
    NAF(5) 0.62±0.010.62 \pm 0.01 11.91±0.1311.91 \pm 0.13 −15.09±0.40-15.09 \pm 0.40 −8.86±0.15-8.86 \pm 0.15 157.73±0.04157.73 \pm 0.04
    NAF(10) 0.60±0.020.60 \pm 0.02 11.96±0.3311.96 \pm 0.33 −15.32±0.23-15.32 \pm 0.23 −9.01±0.01-9.01 \pm 0.01 157.43±0.30157.43 \pm 0.30
    UMNN 0.63±0.010.63 \pm 0.01 10.89±0.7010.89 \pm 0.70 −13.99±0.21-13.99 \pm 0.21 −9.67±0.13-9.67 \pm 0.13 157.98±0.01157.98 \pm 0.01
    SOS(7) 0.60±0.010.60 \pm 0.01 11.99±0.4111.99 \pm 0.41 −15.15±0.10-15.15 \pm 0.10 −8.90±0.11-8.90 \pm 0.11 157.48±0.41157.48 \pm 0.41
    Quadratic Spline (C) 0.64±0.010.64 \pm 0.01 12.80±0.0212.80 \pm 0.02 −15.35±0.02-15.35 \pm 0.02 −9.35±0.44-9.35 \pm 0.44 157.65±0.28157.65 \pm 0.28
    Quadratic Spline (AR) 0.66±0.010.66 \pm 0.01 12.91±0.0212.91 \pm 0.02 −14.67±0.03-14.67 \pm 0.03 −9.72±0.47-9.72 \pm 0.47 157.42±0.28157.42 \pm 0.28
    Cubic Spline 0.65±0.010.65 \pm 0.01 13.14±0.0213.14 \pm 0.02 −14.59±0.02-14.59 \pm 0.02 −9.06±0.48-9.06 \pm 0.48 157.24±0.07157.24 \pm 0.07
    RQ-NSF(C) 0.64±0.010.64 \pm 0.01 13.09±0.0213.09 \pm 0.02 −14.75±0.03-14.75 \pm 0.03 −9.67±0.47-9.67 \pm 0.47 157.54±0.28157.54 \pm 0.28
    RQ-NSF(AR) 0.66±0.010.66 \pm 0.01 13.09±0.0213.09 \pm 0.02 −14.01±0.03-14.01 \pm 0.03 −9.22±0.48-9.22 \pm 0.48 157.31±0.28157.31 \pm 0.28

    Across tabular benchmarks, universal flows equipped with non-linear coupling functions—such as Neural Autoregressive Flows (NAF), Sum-of-Squares polynomial flows (SOS), and Spline flows (Quadratic, Cubic, RQ-NSF)—consistently outperform standard affine coupling (RealNVP) and affine autoregressive (MAF) models.

  12. Knowl 12 — Generative Density Estimation Benchmarks on Image Datasets

    data/table

    The table below compares the average test negative log-likelihood (in bits per dimension; lower is better) for unconditional density estimation on standard image benchmark datasets: MNIST (D=784D=784), CIFAR-10 (D=3072D=3072), ImageNet 32×3232\times 32 (D=3072D=3072), and ImageNet 64×6464\times 64 (D=12288D=12288).

    Model MNIST CIFAR-10 ImageNet32 ImageNet64
    RealNVP 1.061.06 3.493.49 4.284.28 3.983.98
    Glow 1.051.05 3.353.35 4.094.09 3.813.81
    MAF 1.891.89 4.314.31 – –
    FFJORD 0.990.99 3.403.40 – –
    SOS 1.811.81 4.184.18 – –
    RQ-NSF(C) – 3.383.38 – 3.823.82
    UMNN 1.131.13 – – –
    iResNet 1.061.06 3.453.45 – –
    Residual Flow 0.970.97 3.283.28 4.014.01 3.763.76
    Flow++ – 3.083.08 3.863.86 3.693.69

    Flow++ achieves the best performance across CIFAR-10 and downsampled ImageNet benchmarks. This gain is primarily driven by its expressive logistic mixture coupling functions and the use of variational dequantization rather than standard uniform dequantization. Residual Flow also delivers competitive results, outperforming standard affine coupling baselines (RealNVP and Glow).

Coverage note — Domain-specific applied configurations in physics and audio processing, as well as brief introductory discussions on manifold Lie-group parameterizations, were omitted in favor of self-contained knowls covering all primary architectural paradigms, universality theory, coupling functions, and core benchmarks.

References

  1. 1.A. Abdelhamed, M. A. Brubaker, and M. S. Brown, ‘‘Noise flow: Noise modeling with conditional normalizing flows,’’ in Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 3165–3173.
  2. 2.J. Agnelli, M. Cadeiras, E. Tabak, T. Cristina, and E. Vanden-Eijnden, ‘‘Clustering and classification through normalizing flows in feature space,’’ Multiscale Modeling and Simulation, vol. 8, pp. 1784–1802, 2010.
  3. 3.J. Arango and A. Gómez, ‘‘Diffeomorphisms as time one maps,’’ Aequationes Math., vol. 64, pp. 304–314, 2002.
  4. 4.M. Arjovsky, S. Chintala, and L. Bottou, ‘‘Wasserstein Generative Adversarial Networks,’’ in ICML, 2017.
  5. 5.V. Arnold, Ordinary Differential Equations. The MIT Press, 1978.
  6. 6.A. Atanov, A. Volokhova, A. Ashukha, I. Sosnovik, and D. Vetrov, ‘‘Semi-Conditional Normalizing Flows for Semi-Supervised Learning,’’ in Workshop on Invertible Neural Nets and Normalizing Flows, ICML, 2019.
  7. 7.J. Behrmann, D. Duvenaud, and J.-H. Jacobsen, ‘‘Invertible residual networks,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2019.
  8. 8.Y. Bengio, N. Léonard, and A. Courville, ‘‘Estimating or propagating gradients through stochastic neurons for conditional computation,’’ arXiv preprint, arXiv:1308.3432, 2013.
  9. 9.V. Bogachev, A. Kolesnikov, and K. Medvedev, ‘‘Triangular transformations of measures,’’ Sbornik Math., vol. 196, no. 3-4, pp. 309–335, 2005.
  10. 10.A. J. Bose, A. Smofsky, R. Liao, P. Panangaden, and W. L. Hamilton, ‘‘Latent Variable Modelling with Hyperbolic Normalizing Flows,’’ arXiv preprint, arXiv:2002.06336, 2020.
  11. 11.S. R. Bowman, L. Vilnis, O. Vinyals, A. M. Dai, R. Józefowicz, and S. Bengio, ‘‘Generating sentences from a continuous space,’’ in CoNLL, 2015.
  12. 12.B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham, ‘‘Reversible Architectures for Arbitrarily Deep Residual Neural Networks,’’ in AAAI, 2018.
  13. 13.B. Chang, M. Chen, E. Haber, and E. H. Chi, ‘‘AntisymmetricRNN: A dynamical system view on recurrent neural networks,’’ in ICLR, 2019.
  14. 14.C. Chen, C. Li, L. Chen, W. Wang, Y. Pu, and L. Carin, ‘‘Continuous-Time Flows for Efficient Inference and Density Estimation,’’ in ICML, 2018.
  15. 15.R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud, ‘‘Neural ordinary differential equations,’’ Advances in Neural Information Processing Systems, 2018.
  16. 16.R. T. Q. Chen, J. Behrmann, D. Duvenaud, and J.-H. Jacobsen, ‘‘Residual Flows for Invertible Generative Modeling,’’ Advances in Neural Information Processing Systems, 2019.
  17. 17.A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, ‘‘Generative adversarial networks: An overview,’’ IEEE Signal Processing Magazine, vol. 35, pp. 53–65, 2018.
  18. 18.H. P. Das, P. Abbeel, and C. J. Spanos, ‘‘Dimensionality Reduction Flows,’’ arXiv preprint, arXiv:1908.01686, 2019.
  19. 19.L. Dinh, D. Krueger, and Y. Bengio, ‘‘NICE: Non-linear Independent Components Estimation,’’ in ICLR Workshop, 2015.
  20. 20.L. Dinh, J. Sohl-Dickstein, and S. Bengio, ‘‘Density Estimation using Real NVP,’’ in ICLR, 2017.
  21. 21.L. Dinh, J. Sohl-Dickstein, R. Pascanu, and H. Larochelle, ‘‘A RAD approach to deep mixture models,’’ in ICLR Workshop, 2019.
  22. 22.D. Dua and C. Graff, ‘‘UCI Machine Learning Repository,’’ 2017.
  23. 23.E. Dupont, A. Doucet, and Y. W. Teh, ‘‘Augmented Neural ODEs,’’ Advances in Neural Information Processing Systems, 2019.
  24. 24.C. Durkan, A. Bekasov, I. Murray, and G. Papamakarios, ‘‘Cubic-spline flows,’’ in Workshop on Invertible Neural Networks and Normalizing Flows, ICML, 2019.
  25. 25.——, ‘‘Neural Spline Flows,’’ Advances in Neural Information Processing Systems, 2019.
  26. 26.W. E, ‘‘A proposal on machine learning via dynamical systems,’’ Communications in Mathematics and Statistics, vol. 5, pp. 1–11, 2017.
  27. 27.P. Esling, N. Masuda, A. Bardet, R. Despres, and A. Chemla-Romeu-Santos, ‘‘Universal audio synthesizer control with normalizing flows,’’ arXiv preprint, arXiv:1907.00971, 2019.
  28. 28.L. Falorsi, P. de Haan, T. R. Davidson, and P. Forré, ‘‘Reparameterizing Distributions on Lie Groups,’’ arXiv preprint, arXiv:1903.02958, 2019.
  29. 29.C. Finlay, J.-H. Jacobsen, L. Nurbekyan, and A. M. Oberman, ‘‘How to train your neural ODE,’’ arXiv preprint, arXiv:2002.02798, 2020.
  30. 30.M. C. Gemici, D. Rezende, and S. Mohamed, ‘‘Normalizing Flows on Riemannian Manifolds,’’ arXiv preprint, arXiv:1611.02304, 2016.
  31. 31.M. Germain, K. Gregor, I. Murray, and H. Larochelle, ‘‘MADE: Masked Autoencoder for Distribution Estimation,’’ in ICML, 2015.
  32. 32.A. N. Gomez, M. Ren, R. Urtasun, and R. B. Grosse, ‘‘The Reversible Residual Network: Backpropagation Without Storing Activations,’’ Advances in Neural Information Processing Systems, 2017.
  33. 33.I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, ‘‘Generative Adversarial Nets,’’ Advances in Neural Information Processing Systems, 2014.
  34. 34.W. Grathwohl, R. T. Q Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud, ‘‘FFJORD: Free-form continuous dynamics for scalable reversible generative models,’’ in ICLR, 2019.
  35. 35.J. Gregory and R. Delbourgo, ‘‘Piecewise rational quadratic interpolation to monotonic data,’’ IMA Journal of Numerical Analysis, vol. 2, no. 2, pp. 123–130, 1982.
  36. 36.A. Grover, M. Dhar, and S. Ermon, ‘‘Flow-GAN: Combining Maximum Likelihood and Adversarial Learning in Generative Models,’’ in AAAI, 2018.
  37. 37.E. Haber, L. Ruthotto, and E. Holtham, ‘‘Learning across scales - a multiscale method for convolution neural networks,’’ in AAAI, 2018.
  38. 38.L. Hasenclever, J. M. Tomczak, R. Van Den Berg, and M. Welling, ‘‘Variational Inference with Orthogonal Normalizing Flows,’’ in Workshop on Bayesian Deep Learning, NIPS, 2017.
  39. 39.K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,’’ in ICCV, 2015.
  40. 40.——, ‘‘Deep Residual Learning for Image Recognition,’’ in CVPR, 2016.
  41. 41.J. Ho, X. Chen, A. Srinivas, Y. Duan, and P. Abbeel, ‘‘Flow++: Improving flow-based generative models with variational dequantization and architecture design,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2019.
  42. 42.E. Hoogeboom, R. V. D. Berg, and M. Welling, ‘‘Emerging Convolutions for Generative Normalizing Flows,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2019.
  43. 43.E. Hoogeboom, J. W. Peters, R. van den Berg, and M. Welling, ‘‘Integer discrete flows and lossless compression,’’ in NeurIPS, 2019.
  44. 44.E. Hoogeboom, T. S. Cohen, and J. M. Tomczak, ‘‘Learning discrete distributions by dequantization,’’ arXiv preprint, arXiv:2001.11235, 2020.
  45. 45.C.-W. Huang, D. Krueger, A. Lacoste, and A. Courville, ‘‘Neural Autoregressive Flows,’’ in ICML, 2018.
  46. 46.S. Ioffe and C. Szegedy, ‘‘Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,’’ in ICML, 2015.
  47. 47.J.-H. Jacobsen, A. W. Smeulders, and E. Oyallon, ‘‘i-RevNet: Deep Invertible Networks,’’ in ICLR, 2018.
  48. 48.P. Jaini, I. Kobyzev, M. Brubaker, and Y. Yu, ‘‘Tails of Triangular Flows,’’ arXiv preprint, arXiv:1907.04481, 2019.
  49. 49.P. Jaini, K. A. Selby, and Y. Yu, ‘‘Sum-of-squares polynomial flow,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 5 2019.
  50. 50.M. Jankowiak and F. Obermeyer, ‘‘Pathwise derivatives beyond the reparameterization trick,’’ in Proceedings of the 35th International Conference on Machine Learning, ICML, 2018.
  51. 51.G. Kanwar, M. S. Albergo, D. Boyda, K. Cranmer, D. C. Hackett, S. Racanière, D. J. Rezende, and P. E. Shanahan, ‘‘Equivariant flow-based sampling for lattice gauge theory,’’ arXiv preprint, arXiv:2003.06413, 2020.
  52. 52.A. Katok and B. Hasselblatt, Introduction to the modern theory of dynamical systems. Cambridge University Press, New York, 1995.
  53. 53.S. Kim, S. gil Lee, J. Song, J. Kim, and S. Yoon, ‘‘FloWaveNet: A Generative Flow for Raw Audio,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2018.
  54. 54.D. P. Kingma and M. Welling, ‘‘Auto-encoding variational bayes,’’ in Proceedings of the 2nd International Conference on Learning Representations, ICLR, 2014.
  55. 55.——, ‘‘An Introduction to Variational Autoencoders,’’ arXiv preprint, arXiv:1906.02691, 2019.
  56. 56.D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling, ‘‘Improved Variational Inference with Inverse Autoregressive Flow,’’ in NIPS, 2016.
  57. 57.D. P. Kingma and P. Dhariwal, ‘‘Glow: Generative flow with invertible 1x1 convolutions,’’ in Advances in Neural Information Processing Systems, 2018, pp. 10 215–10 224.
  58. 58.J. Köhler, L. Klein, and F. Noé, ‘‘Equivariant flows: sampling configurations for multi-body systems with symmetric energies,’’ in Workshop on Machine Learning and the Physical Sciences, NeurIPS, 2019.
  59. 59.D. Koller and N. Friedman, Probabilistic Graphical Models. Massachusetts: MIT Press, 2009.
  60. 60.M. Kumar, M. Babaeizadeh, D. Erhan, C. Finn, S. Levine, L. Dinh, and D. Kingma, ‘‘VideoFlow: A Flow-Based Generative Model for Video,’’ in Workshop on Invertible Neural Nets and Normalizing Flows, ICML, 2019.
  61. 61.P. M. Laurence, R. J. Pignol, and E. G. Tabak, ‘‘Constrained density estimation,’’ Proceedings of the 2011 Wolfgang Pauli Institute conference on energy and commodity trading, Springer Verlag, pp. 259–284, 2014.
  62. 62.X. Li, T.-K. L. Wong, R. T. Q. Chen, and D. Duvenaud, ‘‘Scalable Gradients for Stochastic Differential Equations,’’ arXiv preprint, arXiv:2001.01328, 2020.
  63. 63.A. Liutkus, U. Simsekli, S. Majewski, A. Durmus, and F.-R. Stöter, ‘‘Sliced-Wasserstein Flows: Nonparametric Generative Modeling via Optimal Transport and Diffusions,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2019.
  64. 64.A. L. Maas, A. Y. Hannun, and A. Y. Ng, ‘‘Rectifier Non-linearities Improve Neural Network Acoustic Models,’’ in ICML, 2013.
  65. 65.K. Madhawa, K. Ishiguro, K. Nakago, and M. Abe, ‘‘Graph-NVP: An Invertible Flow Model for Generating Molecular Graphs,’’ arXiv preprint, arXiv:1905.11600, 2019.
  66. 66.D. Martin, C. Fowlkes, D. Tal, and J. Malik, ‘‘A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,’’ in Proceedings of the 8th International Conference on Computer Vision, ICCV, 2001.
  67. 67.B. Mazoure, T. Doan, A. Durand, J. Pineau, and R. D. Hjelm, ‘‘Leveraging exploration in off-policy algorithms via normalizing flows,’’ in 3rd Conference on Robot Learning (CoRL 2019), 2019.
  68. 68.K. V. Medvedev, ‘‘Certain properties of triangular transformations of measures,’’ Theory Stoch. Process., vol. 14(30), pp. 95–99, 2008.
  69. 69.T. Müller, B. McWilliams, F. Rousselle, M. Gross, and J. Novak, ‘‘Neural Importance Sampling,’’ ACM Transactions on Graphics (TOG), vol. 38, 2018.
  70. 70.P. Nadeem Ward, A. Smofsky, and A. Joey Bose, ‘‘Improving exploration in soft-actor-critic with normalizing flows policies,’’ in Workshop on Invertible Neural Networks and Normalizing Flows, ICML, 2019.
  71. 71.F. Noé, S. Olsson, J. Köhler, and H. Wu, ‘‘Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning,’’ Science, vol. 365, 2019.
  72. 72.B. Oksendal, Stochastic Differential Equations (3rd Ed.): An Introduction with Applications. Berlin, Heidelberg: Springer-Verlag, 1992.
  73. 73.I. Ovinnikov, ‘‘Poincaré Wasserstein Autoencoder,’’ in Bayesian Deep Learning Workshop, NeurIPS, 2018.
  74. 74.G. Papamakarios, T. Pavlakou, and I. Murray, ‘‘Masked Autoregressive Flow for Density Estimation,’’ in NIPS, 2017.
  75. 75.G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mohamed, and B. Lakshminarayanan, ‘‘Normalizing Flows for Probabilistic Modeling and Inference,’’ arXiv preprint, arXiv:1912.02762, 2019.
  76. 76.S. Peluchetti and S. Favaro, ‘‘Neural Stochastic Differential Equations,’’ arXiv preprint, arXiv:1905.11065, 2019.
  77. 77.R. Prenger, R. Valle, and B. Catanzaro, ‘‘Waveglow: A flow-based generative network for speech synthesis,’’ in ICASSP, 2019.
  78. 78.D. J. Rezende and S. Mohamed, ‘‘Variational Inference with Normalizing Flows,’’ in ICML, 2015.
  79. 79.D. J. Rezende, G. Papamakarios, S. Racanière, M. S. Albergo, G. Kanwar, P. E. Shanahan, and K. Cranmer, ‘‘Normalizing Flows on Tori and Spheres,’’ arXiv preprint, arXiv:2002.02428, 2020.
  80. 80.O. Rippel and R. P. Adams, ‘‘High-dimensional probability estimation with deep density models,’’ arXiv preprint arXiv:1302.5125, 2013.
  81. 81.T. Salimans, A. Diederik, D. P. Kingma, and M. Welling, ‘‘Markov Chain Monte Carlo and Variational Inference: Bridging the Gap,’’ in ICML, 2015.
  82. 82.T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, ‘‘Improved Techniques for Training GANs,’’ in NIPS, 2016.
  83. 83.H. Salman, P. Yadollahpour, T. Fletcher, and N. Batmanghelich, ‘‘Deep diffeomorphic normalizing flows,’’ arXiv preprint, arXiv:1810.03256, 2018.
  84. 84.J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, ‘‘Deep unsupervised learning using nonequilibrium thermodynamics,’’ in Proceedings of the 32nd International Conference on Machine Learning, ICML, 2015.
  85. 85.A. Spantini, D. Bigoni, and Y. Marzouk, ‘‘Inference via low-dimensional couplings,’’ Journal of Machine Learning Research, vol. 19, 03 2017.
  86. 86.M. Spivak, Calculus on Manifolds: A Modern Approach to Classical Theorems of Advanced Calculus. Print, 1965.
  87. 87.J. Suykens, H. Verrelst, and J. Vandewalle, ‘‘On-Line Learning Fokker-Planck Machine,’’ Neural Processing Letters, vol. 7, pp. 81–89, 1998.
  88. 88.E. G. Tabak and C. V. Turner, ‘‘A Family of Nonparametric Density Estimation Algorithms,’’ Communications on Pure and Applied Mathematics, vol. 66, no. 2, pp. 145–164, 2013.
  89. 89.E. G. Tabak and E. Vanden-Eijnden, ‘‘Density Estimation by Dual Ascent of the Log-Likelihood,’’ Communications in Mathematical Sciences, vol. 8, no. 1, pp. 217–233, 2010.
  90. 90.I. O. Tolstikhin, O. Bousquet, S. Gelly, and B. Schölkopf, ‘‘Wasserstein Auto-Encoders,’’ in ICLR, 2018.
  91. 91.J. Tomczak and M. Welling, ‘‘Improving Variational Auto-Encoders using convex combination linear Inverse Autoregressive Flow,’’ Benelearn, 2017.
  92. 92.J. M. Tomczak and M. Welling, ‘‘Improving variational auto-encoders using householder flow,’’ arXiv preprint arXiv:1611.09630, 2016.
  93. 93.A. Touati, H. Satija, J. Romoff, J. Pineau, and P. Vincent, ‘‘Randomized value functions via multiplicative normalizing flows,’’ in UAI2019: Conference on Uncertainty in Artificial Intelligence, 2019.
  94. 94.D. Tran, K. Vafa, K. Agrawal, L. Dinh, and B. Poole, ‘‘Discrete Flows: Invertible Generative Models of Discrete Data,’’ in ICLR Workshop, 2019.
  95. 95.B. L. Trippe and R. E. Turner, ‘‘Conditional Density Estimation with Bayesian Normalising Flows,’’ in Workshop on Bayesian Deep Learning, NIPS, 2017.
  96. 96.B. Tzen and M. Raginsky, ‘‘Neural Stochastic Differential Equations: Deep Latent Gaussian Models in the Diffusion Limit,’’ arXiv preprint, arXiv:1905.09883, 2019.
  97. 97.R. van den Berg, L. Hasenclever, J. M. Tomczak, and M. Welling, ‘‘Sylvester normalizing flows for variational inference,’’ in Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, UAI, 2018.
  98. 98.A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, ‘‘Parallel wavenet: Fast high-fidelity speech synthesis,’’ in ICML, 2017.
  99. 99.C. Villani, Topics in optimal transportation (Graduate Studies in Mathematics 58). American Mathematical Society, Providence, RI, 2003.
  100. 100.K. Wang, C. Gou, Y. Duan, Y. Lin, X. Zheng, and F. yue Wang, ‘‘Generative adversarial networks: introduction and outlook,’’ IEEE/CAA Journal of Automatica Sinica, vol. 4, pp. 588–598, 2017.
  101. 101.P. Z. Wang and W. Y. Wang, ‘‘Riemannian Normalizing Flow on Variational Wasserstein Autoencoder for Text Modeling,’’ arXiv preprint, arXiv:1904.02399, 2019.
  102. 102.A. Wehenkel and G. Louppe, ‘‘Unconstrained Monotonic Neural Networks,’’ arXiv preprint, arXiv:1908.05164, 2019.
  103. 103.M. Welling and Y. W. Teh, ‘‘Bayesian Learning via Stochastic Gradient Langevin Dynamics,’’ in ICML, 2011.
  104. 104.P. Wirnsberger, A. Ballard, G. Papamakarios, S. Abercrombie, S. Racanière, A. Pritzel, D. Jimenez Rezende, and C. Blundell, ‘‘Targeted free energy estimation via learned mappings,’’ arXiv preprint, arXiv:2002.04913, 2020.
  105. 105.K. W. K. Wong, G. Contardo, and S. Ho, ‘‘Gravitational wave population inference with deep flow-based generative network,’’ arXiv preprint, arXiv:2002.09491, 2020.
  106. 106.H. Zhang, X. Gao, J. Unterman, and T. Arodz, ‘‘Approximation Capabilities of Neural Ordinary Differential Equations,’’ arXiv preprint, arXiv:1907.12998, 2019.
  107. 107.G. Zheng, Y. Yang, and J. Carbonell, ‘‘Convolutional Normalizing Flows,’’ in Workshop on Theoretical Foundations and Applications of Deep Generative Models, ICML, 2018.
  108. 108.Z. M. Ziegler and A. M. Rush, ‘‘Latent Normalizing Flows for Discrete Sequences,’’ in Proceedings of the 36th International Conference on Machine Learning, ICML, 2019.

Citation

MLA
Kobyzev, I., et al. “Normalizing Flows: An Introduction and Review of Current Methods”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 11, 2021, pp. 3964–79, https://doi.org/10.1109/TPAMI.2020.2992934.
APA
Kobyzev, I., Prince, S. J. D., & Brubaker, M. A. (2021). Normalizing Flows: An Introduction and Review of Current Methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11), 3964–3979. https://doi.org/10.1109/TPAMI.2020.2992934
Chicago
Kobyzev, I., S. J. D. Prince, and M. A. Brubaker. 2021. “Normalizing Flows: An Introduction and Review of Current Methods”. IEEE Transactions on Pattern Analysis and Machine Intelligence 43 (11): 3964–79. https://doi.org/10.1109/TPAMI.2020.2992934.
Harvard
Kobyzev, I., Prince, S.J.D. and Brubaker, M.A. (2021) “Normalizing Flows: An Introduction and Review of Current Methods”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11), pp. 3964–3979. Available at: https://doi.org/10.1109/TPAMI.2020.2992934.
Vancouver
1. Kobyzev I, Prince SJD, Brubaker MA (2021) Normalizing Flows: An Introduction and Review of Current Methods. IEEE Transactions on Pattern Analysis and Machine Intelligence 43:3964–3979

BibTeX

@article{Kobyzev_2021, title={Normalizing Flows: An Introduction and Review of Current Methods}, volume={43}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2020.2992934}, DOI={10.1109/tpami.2020.2992934}, number={11}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Kobyzev, Ivan and Prince, Simon J.D. and Brubaker, Marcus A.}, year={2021}, month=Nov, pages={3964–3979} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF