Normalizing Flows are Capable Generative Models

Shuangfei ZhaiRuixiang ZhangPreetum NakkiranDavid BerthelotJiatao GuHuangjie ZhengTianrong ChenMiguel ngel BautistaNavdeep JaitlyJoshua M. Susskind

article2025ICML115 citations

Demonstrates that normalizing flows can rival diffusion models in image generation quality while achieving state-of-the-art exact likelihood estimation through a scalable Transformer-based architecture paired with noise augmentation and guidance.

Listen

Generative artificial intelligence has largely been dominated by diffusion models and large language models, while normalizing flows—a foundational class of models known for exact likelihood computation and efficient mathematical invertibility—have lagged behind in perceptual generation quality and practical adoption. The article addresses this performance gap by investigating whether normalizing flows are fundamentally limited or merely constrained by suboptimal architectures and training procedures. Its primary objective is to introduce TARFLOW, a scalable transformer-based architecture combined with novel training and sampling techniques, to demonstrate that standalone normalizing flows can achieve state-of-the-art density estimation and image generation competitive with leading generative frameworks.

The authors evaluate TARFLOW across standard image benchmarks, including unconditional and class-conditional ImageNet (at 64x64 and 128x128 resolutions) and AFHQ (at 256x256 resolution). The approach replaces conventional masked multilayer perceptrons with causal vision transformer blocks that operate on image patches across alternating sequence directions. To bolster generative quality, the methodology introduces three core techniques: training with moderate Gaussian noise augmentation rather than narrow uniform noise, applying a post-training score-based denoising step via Tweedie's formula directly on the learned density, and adapting classifier-free guidance mechanisms for both conditional and unconditional generation.

The article reports several critical findings. First, TARFLOW sets a new state of the art in image density estimation on unconditional ImageNet 64x64, achieving a negative log-likelihood of 2.99 bits per dimension and breaking the sub-3.0 threshold for the first time. Second, in image generation quality, TARFLOW achieves competitive Fréchet Inception Distance (FID) scores—reaching 2.66 on conditional ImageNet 64x64 and 5.03 on ImageNet 128x128—outperforming traditional generative adversarial network (GAN) baselines and approaching diffusion model benchmarks. Third, ablation analyses show that noise augmentation paired with score-based denoising drastically reduces visual artifacts, while guidance reliably trades sample diversity for sharper class fidelity. Finally, depth ablations reveal that balancing the number of sequential flow blocks and layers per block yields optimal performance, confirming the model scales smoothly with compute.

These findings carry significant technical and operational implications. By proving that normalizing flows alone can generate high-fidelity images directly from continuous pixels without complex multi-stage tokenization or vector quantization, TARFLOW establishes an alternative, mathematically tractable paradigm for generative modeling. For decision-makers, this opens opportunities to leverage the exact likelihood tracking and deterministic objectives of flows without sacrificing visual quality. Although sampling currently requires sequential autoregressive generation—taking approximately two minutes for a batch of 32 images on an A100 GPU—the architecture provides a robust, modular baseline that can leverage existing transformer optimization ecosystems.

Organizations evaluating generative modeling pipelines should consider testing TARFLOW architectures where exact probability estimation, direct pixel generation, and training stability are critical requirements. Next steps should focus on exploring optimal guidance schedules, scaling model capacity to higher resolutions, and implementing engineering optimizations such as advanced caching and gradient checkpointing to reduce sampling latency and memory overhead. While the reported empirical confidence is high across standard benchmarks, stakeholders should note that sampling speed remains a key operational bottleneck compared to non-autoregressive alternatives before deploying the framework into real-time production workflows.

arXiv: 2412.06329
Cover for Normalizing Flows are Capable Generative Models

Abstract

Normalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In this work, we demonstrate that NFs are more powerful than previously believed. We present TARFLOW: a simple and scalable architecture that enables highly performant NF models. TARFLOW can be thought of as a Transformer-based variant of Masked Autoregressive Flows (MAFs): it consists of a stack of autoregressive Transformer blocks on image patches, alternating the autoregression direction between layers. TARFLOW is straightforward to train end-to-end, and capable of directly modeling and generating pixels. We also propose three key techniques to improve sample quality: Gaussian noise augmentation during training, a post training denoising procedure, and an effective guidance method for both class-conditional and unconditional settings. Putting these together, TARFLOW sets new state-of-the-art results on likelihood estimation for images, beating the previous best methods by a large margin, and generates samples with quality and diversity comparable to diffusion models, for the first time with a stand-alone NF model. We make our code available at https://github.com/apple/ml-tarflow.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Normalizing Flows
  • 2.2. Block Autoregressive Flows
  • 2.3. Transformer Autoregressive Flows
  • 2.4. Noise Augmented Training
  • 2.5. Score Based Denoising
  • 2.6. Guidance
  • 3. Experiments
  • 3.1. Likelihood
  • 3.2. Generation
  • 3.3. Ablation on Noise Augmentation and Denoising
  • 3.4. Ablation on Guidance
  • 3.5. Ablation on Model Scaling
  • 3.6. Comparison with VP and Channel Coupling
  • 4. Related Work
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Additional Related Work
  • B. Guidance
  • C. Experimental details
  • D. Inference Implementation
  • D.1. Visualizing Sample Trajectory
  • E. Additional samples

Knowls

  1. Knowl 1 — TARFLOW Architecture and Block Autoregressive Flow Formulation

    model/method

    TARFLOW defines a deep normalizing flow by composing TT block-autoregressive flow transformations on patchified continuous inputs. An input image x∈RC×H×Wx \in \mathbb{R}^{C \times H \times W} with channels CC, height HH, and width WW is converted into a sequence of non-overlapping patches of size S×SS \times S, yielding x∈RN×Dx \in \mathbb{R}^{N \times D} with sequence length N=HWS2N = \frac{HW}{S^2} and block dimension D=CS2D = C S^2.

    The full flow transformation maps xx to latent variable zT∈RN×Dz^T \in \mathbb{R}^{N \times D} via zT=(fT−1∘fT−2∘⋯∘f0)(x)z^T = (f^{T-1} \circ f^{T-2} \circ \dots \circ f^0)(x), initializing with z0=xz^0 = x. In each flow block t∈{0,…,T−1}t \in \{0, \dots, T-1\}:

    1. Sequence Permutation: The sequence is permuted as z~t=πt(zt)\tilde{z}^t = \pi^t(z^t), where π0\pi^0 is the identity map and πt(z)i=zN−1−i\pi^t(z)_i = z_{N-1-i} reverses the sequence for t≥1t \ge 1.

    2. Block Autoregressive Transformation: For patch token index i=0i = 0: z0t+1=z~0tz^{t+1}_0 = \tilde{z}^t_0 and for each i∈{1,…,N−1}i \in \{1, \dots, N-1\}: zit+1=(z~it−μit(z~<it))⊙exp⁡(−αit(z~<it))z^{t+1}_i = (\tilde{z}^t_i - \mu^t_i(\tilde{z}^t_{<i})) \odot \exp(-\alpha^t_i(\tilde{z}^t_{<i})) where ⊙\odot denotes element-wise multiplication, and μt,αt:RN×D→RN×D\mu^t, \alpha^t : \mathbb{R}^{N \times D} \to \mathbb{R}^{N \times D} are causal functions parameterized by a causal Vision Transformer (ViT) with attention masks.

    The analytical inverse mapping x=f−1(zT)x = f^{-1}(z^T) is computed by sequentially inverting each block tt from zt+1z^{t+1} to ztz^t: z~0t=z0t+1\tilde{z}^t_0 = z^{t+1}_0 z~it=zit+1⊙exp⁡(αit(z~<it))+μit(z~<it),i∈{1,…,N−1}\tilde{z}^t_i = z^{t+1}_i \odot \exp(\alpha^t_i(\tilde{z}^t_{<i})) + \mu^t_i(\tilde{z}^t_{<i}), \quad i \in \{1, \dots, N-1\} zt=(πt)−1(z~t)z^t = (\pi^t)^{-1}(\tilde{z}^t) iterating until x=z0x = z^0.

  2. Knowl 2 — Exact Likelihood Objective and Jacobian Determinant for TARFLOW

    equation

    Because the block permutation operation πt\pi^t preserves volume (zero log-determinant) and the autoregressive affine transformation has a lower-triangular Jacobian matrix, the Jacobian determinant depends strictly on the diagonal scale elements. For each flow layer ftf^t, the log-determinant evaluates to:

    log⁡∣det⁡(dft(zt)dzt)∣=−∑i=1N−1∑j=0D−1αit(z~<it)j\log \left| \det\left( \frac{d f^t(z^t)}{d z^t} \right) \right| = -\sum_{i=1}^{N-1} \sum_{j=0}^{D-1} \alpha^t_i(\tilde{z}^t_{<i})_j

    where NN is the patch sequence length, DD is the per-patch feature dimension, and αit(z~<it)j\alpha^t_i(\tilde{z}^t_{<i})_j is the jj-th element of the causal scale prediction vector for patch ii.

    Assuming a standard Gaussian prior distribution p0(z)=N(z;0,I)p_0(z) = \mathcal{N}(z; 0, I), maximum likelihood training minimizes the negative log-likelihood objective over data xx (omitting constant terms):

    min⁡f12∥zT∥22+∑t=0T−1∑i=1N−1∑j=0D−1αit(z~<it)j\min_f \frac{1}{2} \|z^T\|_2^2 + \sum_{t=0}^{T-1} \sum_{i=1}^{N-1} \sum_{j=0}^{D-1} \alpha^t_i(\tilde{z}^t_{<i})_j

    where zT=f(x)z^T = f(x) is the final latent vector produced by the stack of TT flow layers.

  3. Knowl 3 — Score-Based Post-Training Tweedie Denoising for Normalizing Flows

    model/method

    When a normalizing flow model pmodelp_{\text{model}} is trained on a Gaussian-noise-augmented data distribution q(y)=∫pdata(y−ϵ)N(ϵ;0,σ2I)dϵq(y) = \int p_{\text{data}}(y - \epsilon) \mathcal{N}(\epsilon; 0, \sigma^2 I) d\epsilon, raw generated samples y=f−1(z)y = f^{-1}(z) where z∼N(0,I)z \sim \mathcal{N}(0, I) contain residual noise of variance σ2\sigma^2.

    By Tweedie's formula, the conditional expectation of clean data x∼pdatax \sim p_{\text{data}} given noisy observation yy satisfies: E[x∣y]=y+σ2∇ylog⁡q(y)\mathbb{E}[x \mid y] = y + \sigma^2 \nabla_y \log q(y)

    Assuming the flow model is well-trained, its exact log-likelihood gradient ∇ylog⁡pmodel(y)\nabla_y \log p_{\text{model}}(y) serves as a direct estimator of the score function ∇ylog⁡q(y)\nabla_y \log q(y). Sampling proceeds without auxiliary networks as:

    1. Sample latent Gaussian noise z∼N(0,I)z \sim \mathcal{N}(0, I).
    2. Compute raw sample y=f−1(z)y = f^{-1}(z) via the inverse flow.
    3. Compute the score ∇ylog⁡pmodel(y)\nabla_y \log p_{\text{model}}(y) by differentiating the flow log-likelihood through a forward pass.
    4. Produce clean sample x^=y+σ2∇ylog⁡pmodel(y)\hat{x} = y + \sigma^2 \nabla_y \log p_{\text{model}}(y).

    This denoising step requires two forward model passes and eliminates sampling artifacts caused by training noise.

  4. Knowl 4 — Classifier-Free Guidance for Conditional Normalizing Flows

    model/method

    Guidance is incorporated into class-conditional normalizing flows during inverse autoregressive generation. During training, class conditioning labels cc are randomly dropped out with probability p=0.1p = 0.1 to simultaneously train class-conditional functions μit(⋅;c),αit(⋅;c)\mu^t_i(\cdot; c), \alpha^t_i(\cdot; c) and unconditional functions μit(⋅;∅),αit(⋅;∅)\mu^t_i(\cdot; \emptyset), \alpha^t_i(\cdot; \emptyset).

    During sampling, the causal shift and scale predictions in the reverse step of each flow layer tt are modified under guidance weight w≥0w \ge 0:

    μ~it(z~<it;c,w)=(1+w)μit(z~<it;c)−wμit(z~<it;∅)\tilde{\mu}^t_i(\tilde{z}^t_{<i}; c, w) = (1 + w) \mu^t_i(\tilde{z}^t_{<i}; c) - w \mu^t_i(\tilde{z}^t_{<i}; \emptyset) α~it(z~<it;c,w)=(1+w)αit(z~<it;c)−wαit(z~<it;∅)\tilde{\alpha}^t_i(\tilde{z}^t_{<i}; c, w) = (1 + w) \alpha^t_i(\tilde{z}^t_{<i}; c) - w \alpha^t_i(\tilde{z}^t_{<i}; \emptyset)

    The reverse generation update for token i>0i > 0 evaluates as: z~it=zit+1⊙exp⁡(α~it(z~<it;c,w))+μ~it(z~<it;c,w)\tilde{z}^t_i = z^{t+1}_i \odot \exp(\tilde{\alpha}^t_i(\tilde{z}^t_{<i}; c, w)) + \tilde{\mu}^t_i(\tilde{z}^t_{<i}; c, w)

    Positive guidance weights w>0w > 0 guide latent generation away from the unconditional distribution toward the specific class mode, trading sample diversity for visual fidelity and mode-seeking accuracy.

  5. Knowl 5 — Unconditional Flow Guidance via Attention Softmax Temperature Modulation

    model/method

    Guidance can be applied to unconditional normalizing flows by generating intentionally degraded predictions via attention temperature modulation. For every causal attention layer in flow block ftf^t, attention logits are scaled by a temperature hyperparameter τ≠1\tau \ne 1 before the Softmax function to produce degraded causal predictions μit(⋅;τ)\mu^t_i(\cdot; \tau) and αit(⋅;τ)\alpha^t_i(\cdot; \tau). Setting τ>1\tau > 1 or τ<1\tau < 1 overly smooths or sharpens attention distributions, degrading the Transformer's predictive quality.

    Guided unconditional generation combines nominal predictions (̂\tau = 1) and degraded predictions (̂\tau \ne 1) with guidance weight ww:

    μ~it(z~<it;τ,w)=(1+w)μit(z~<it;1)−wμit(z~<it;τ)\tilde{\mu}^t_i(\tilde{z}^t_{<i}; \tau, w) = (1 + w) \mu^t_i(\tilde{z}^t_{<i}; 1) - w \mu^t_i(\tilde{z}^t_{<i}; \tau) α~it(z~<it;τ,w)=(1+w)αit(z~<it;1)−wαit(z~<it;τ)\tilde{\alpha}^t_i(\tilde{z}^t_{<i}; \tau, w) = (1 + w) \alpha^t_i(\tilde{z}^t_{<i}; 1) - w \alpha^t_i(\tilde{z}^t_{<i}; \tau)

    Increasing ww or ∣τ−1∣|\tau - 1| strengthens guidance toward high-density image modes. Additionally, linearly scheduling the guidance weight across patch position ii via wi=iT−1ww_i = \frac{i}{T - 1} w achieves lower FID than a constant weight.

  6. Knowl 6 — Gaussian Noise Augmentation for Flow Training vs. Uniform Dequantization

    empirical result

    Standard normalizing flow training utilizes uniform dequantization noise U(0,bin)\mathcal{U}(0, \text{bin}) with bin size 1/1281/128 for 8-bit pixels normalized to [−1,1][-1, 1] (standard deviation ≈0.002\approx 0.002). While sufficient for likelihood evaluation, sampling from models trained with uniform noise suffers from severe numerical instability because forcing the model to map low-entropy discrete training data to an ambient Gaussian prior causes out-of-distribution generalization failure during inverse sampling.

    Training with isotropic Gaussian noise ϵ∼N(0,σ2I)\epsilon \sim \mathcal{N}(0, \sigma^2 I) at a moderate magnitude σ∈[0.05,0.15]\sigma \in [0.05, 0.15] (an order of magnitude larger than uniform dequantization noise) expands the support of the training data distribution across the continuous ambient space. Although raw generated samples retain visible noise, combining Gaussian noise augmentation during training with score-based Tweedie denoising during inference resolves numerical instability and produces high-fidelity images.

  7. Knowl 7 — Likelihood Estimation Performance on Unconditional ImageNet 64x64

    data/table

    Evaluation of test-set negative log-likelihood measured in bits per dimension (BPD, lower is better) on unconditional ImageNet 64×6464\times64. The TARFLOW configuration is specified as [P-Ch-T-K-pϵ][P\text{-}Ch\text{-}T\text{-}K\text{-}p_\epsilon] denoting patch size P=2P=2, channel dimension Ch=768Ch=768, T=8T=8 flow blocks, K=8K=8 layers per block, and uniform dequantization noise pϵ=U(0,1/128)p_\epsilon = \mathcal{U}(0, 1/128) in float32 precision.

    Model Type BPD ↓\downarrow
    Very Deep VAE (Child, 2021) VAE 3.52
    Glow (Kingma Dhariwal, 2018) Flow 3.81
    Flow++ (Ho et al., 2019) Flow 3.69
    PixelCNN (van den Oord et al., 2016a) AR 3.83
    SPN (Menick Kalchbrenner, 2019) AR 3.52
    Sparse Transformer (Child et al., 2019) AR 3.44
    Routing Transformer (Roy et al., 2021) AR 3.43
    Improved DDPM (Nichol Dhariwal, 2021) Diff/FM 3.54
    VDM (Kingma et al., 2021) Diff/FM 3.40
    Flow Matching (Lipman et al., 2023a) Diff/FM 3.31
    NFDM (Bartosh et al., 2024) Diff/FM 3.20
    TARFLOW [2-768-8-8-U(0,1128)\mathcal{U}(0, \frac{1}{128})] NF 2.99

    TARFLOW achieves the first sub-3.0 BPD likelihood result on ImageNet 64×6464\times64, outperforming prior normalizing flows, autoregressive models, VAEs, and diffusion/flow matching models.

  8. Knowl 8 — Generation Quality (FID) on ImageNet 64x64 and 128x128

    data/table

    Fréchet Inception Distance (FID, 50K samples, lower is better) on conditional and unconditional ImageNet benchmarks comparing TARFLOW to GANs, consistency models (CM), and diffusion/flow matching (Diff/FM) models. Configuration notation [P-Ch-T-K-pϵ][P\text{-}Ch\text{-}T\text{-}K\text{-}p_\epsilon] specifies patch size PP, channel dimension ChCh, flow blocks TT, layers per flow KK, and training noise pϵp_\epsilon.

    Task / Dataset Model Type FID ↓\downarrow
    Cond. ImageNet 64×6464\times64 EDM (Karras et al., 2022) Diff/FM 1.55
    ADM(dropout) (Dhariwal Nichol, 2021) Diff/FM 2.09
    iDDPM (Nichol Dhariwal, 2021) Diff/FM 2.92
    iCT-deep (Song Dhariwal, 2023) CM 3.25
    BigGAN (Brock et al., 2019) GAN 4.06
    CD(LPIPS) (Song et al., 2023) CM 4.70
    IC-GAN (Casanova et al., 2021) GAN 6.70
    TARFLOW [4-1024-8-8-N(0,0.052)\mathcal{N}(0, 0.05^2)] NF 3.99
    TARFLOW [2-768-8-8-N(0,0.052)\mathcal{N}(0, 0.05^2)] NF 2.90
    TARFLOW [2-1024-8-8-N(0,0.052)\mathcal{N}(0, 0.05^2)] NF 2.66
    Cond. ImageNet 128×128128\times128 Simple Diff (Hoogeboom et al., 2023) Diff/FM 1.94
    RIN (Jabri et al., 2023) Diff/FM 2.75
    ADM-G (Dhariwal Nichol, 2021) Diff/FM 2.97
    CDM (Ho et al., 2022) Diff/FM 3.52
    BigGAN-deep (Brock et al., 2019) GAN 5.70
    BigGAN (Brock et al., 2019) GAN 8.70
    TARFLOW [4-1024-8-8-N(0,0.052)\mathcal{N}(0, 0.05^2)] NF 5.29
    TARFLOW [4-1024-8-8-N(0,0.152)\mathcal{N}(0, 0.15^2)] NF 5.03
    Uncond. ImageNet 64×6464\times64 AGM (Chen et al., 2024) Diff/FM 10.07
    IC-GAN (Casanova et al., 2021) GAN 10.40
    MFM (Pooladian et al., 2023) Diff/FM 11.82
    FM (Lipman et al., 2023b) Diff/FM 13.93
    Self-sup GAN (Noroozi, 2020) GAN 19.20
    TARFLOW [2-768-8-8-N(0,0.052)\mathcal{N}(0, 0.05^2)] NF 18.42

    TARFLOW establishes competitive FID results for normalizing flows, outperforming strong GAN baselines on ImageNet 64×6464\times64 and 128×128128\times128 while approaching diffusion model performance.

  9. Knowl 9 — Ablation of Non-Volume Preservation and Causal Masking in Flow Design

    data/table

    Comparison of sample FID on class-conditional ImageNet 64×6464\times64 isolating two core architectural choices of TARFLOW: Non-Volume Preserving (NVP) scaling transformations and autoregressive causal attention masks. The Volume Preserving (VP) baseline sets αit(⋅)=0\alpha^t_i(\cdot) = 0 in the affine transformation, while the Channel Coupling baseline removes causal attention masks. Both baselines use identical compute budgets, σ=0.05\sigma = 0.05 Gaussian noise augmentation, and Tweedie score denoising.

    Guidance (ww) TARFLOW Volume Preserving (VP) Channel Coupling
    0 25.3 81.5 50.3
    2 5.7 51.0 20.4

    Removing either the learned scaling parameters or the causal autoregressive structure substantially degrades FID (from 5.7 to 51.0 for VP, and to 20.4 for channel coupling under guidance w=2w=2), showing both components are critical for generative performance in normalizing flows.

  10. Knowl 10 — Depth Allocation and Scaling Dynamics in Stacked Autoregressive Flows

    empirical result

    In TARFLOW, total depth is parameterized by the number of stacked flow blocks TT and the number of Transformer attention layers per block KK.

    1. Layer Allocation Trade-Off: Holding total network depth fixed at T×K=64T \times K = 64 on conditional ImageNet 64×6464\times64, loss and FID follow a symmetric U-shaped distribution across configurations {1×64,2×32,4×16,8×8,16×4,32×2,64×1}\{1\times64, 2\times32, 4\times16, 8\times8, 16\times4, 32\times2, 64\times1\}. The optimal trade-off occurs when capacity is balanced equally between the number of flows and the depth per flow (T=K=8T = K = 8).

    2. Failure of Single-Pass Autoregression: A single block autoregressive model (T=1,K=64T=1, K=64) fails completely to fit continuous image patch distributions, producing high training loss and an FID of 267 (equivalent to random noise). Stacking T≥2T \ge 2 blocks with alternating sequence permutations resolves this bottleneck.

    3. Monotonic Scaling: Increasing model depth (e.g., from 4×44\times4 to 8×48\times4 to 8×88\times8) consistently decreases negative log-likelihood loss and improves FID, demonstrating a strong positive correlation between likelihood optimization and image synthesis quality.

Coverage note — None was omitted; all primary contributions including the TARFLOW architecture, exact likelihood formulation, score-based denoising, conditional/unconditional guidance mechanisms, and experimental likelihood/FID benchmarks were included.

References

  1. 1.Bartosh, G., Vetrov, D., and Naesseth, C. A. Neural flow diffusion models: Learnable forward process for improved diffusion modelling. ArXiv preprint, abs/2404.12940, 2024. URL https://arxiv.org/abs/2404.12940.
  2. 2.Brock, A., Donahue, J., and Simonyan, K. Large scale GAN training for high fidelity natural image synthesis. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=B1xsqj09Fm.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  4. 4.Cao, N. D., Aziz, W., and Titov, I. Block neural autoregressive flow. In Globerson, A. and Silva, R. (eds.), Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, volume 115 of Proceedings of Machine Learning Research, pp. 1263–1273. AUAI Press, 2019. URL http://proceedings.mlr.press/v115/de-cao20a.html.
  5. 5.Casanova, A., Careil, M., Verbeek, J., Drozdzal, M., and Romero-Soriano, A. Instance-conditioned GAN. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 27517–27529, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/e7ac288b0f2d41445904d071ba37aaff-Abstract.html.
  6. 6.Chalvidal, M., Ricci, M., VanRullen, R., and Serre, T. Go with the flow: Adaptive control for neural odes. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=giit4HdDNa.
  7. 7.Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. Generative pretraining from pixels. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 1691–1703. PMLR, 2020. URL http://proceedings.mlr.press/v119/chen20s.html.
  8. 8.Chen, T., Gu, J., Dinh, L., Theodorou, E., Susskind, J. M., and Zhai, S. Generative modeling with phase stochastic bridge. In The Twelfth International Conference on Learning Representations, 2024.
  9. 9.Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. Neural ordinary differential equations. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pp. 6572–6583, 2018. URL https://proceedings.neurips.cc/paper/2018/hash/69386f6bb1dfed68692a24c8686939b9-Abstract.html.
  10. 10.Child, R. Very deep vaes generalize autoregressive models and can outperform them on images. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=RLRXCV6DbEJ.
  11. 11.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. ArXiv preprint, abs/1904.10509, 2019. URL https://arxiv.org/abs/1904.10509.
  12. 12.Choi, Y., Uh, Y., Yoo, J., and Ha, J. Stargan v2: Diverse image synthesis for multiple domains. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pp. 8185–8194. IEEE, 2020. doi: 10.1109/CVPR42600.2020.00821. URL https://doi.org/10.1109/CVPR42600.2020.00821.
  13. 13.Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pp. 248–255. IEEE Computer Society, 2009. doi: 10.1109/CVPR.2009.5206848. URL https://doi.org/10.1109/CVPR.2009.5206848.
  14. 14.Dhariwal, P. and Nichol, A. Q. Diffusion models beat gans on image synthesis. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 8780–8794, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/49ad23d1ec9fa4bd8d77d02681df5cfa-Abstract.html.
  15. 15.Dinh, L., Krueger, D., and Bengio, Y. Nice: Non-linear independent components estimation. International Conference on Learning Representations workshop Track, 2014.
  16. 16.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real NVP. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=HkpbnH9lx.
  17. 17.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  18. 18.Dupont, E., Doucet, A., and Teh, Y. W. Augmented neural odes. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 3134–3144, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/21be9a4bd4f81549a9d1d241981cec3c-Abstract.html.
  19. 19.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pp. 12873–12883. Computer Vision Foundation / IEEE, 2021. doi: 10.1109/CVPR46437.2021.01268. URL https://openaccess.thecvf.com/content/CVPR2021/html/Esser_Taming_Transformers_for_High-Resolution_Image_Synthesis_CVPR_2021_paper.html.
  20. 20.Germain, M., Gregor, K., Murray, I., and Larochelle, H. MADE: masked autoencoder for distribution estimation. In Bach, F. R. and Blei, D. M. (eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pp. 881–889. JMLR.org, 2015. URL http://proceedings.mlr.press/v37/germain15.html.
  21. 21.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y. Generative adversarial nets. In Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N. D., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pp. 2672–2680, 2014. URL https://proceedings.neurips.cc/paper/2014/hash/5ca3e9b122f61f8f06494c97b1afccf3-Abstract.html.
  22. 22.Grathwohl, W., Chen, R. T. Q., Bettencourt, J., Sutskever, I., and Duvenaud, D. FFJORD: free-form continuous dynamics for scalable reversible generative models. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=rJxgknCcK7.
  23. 23.Gu, J., Wang, Y., Zhang, Y., Zhang, Q., Zhang, D., Jaitly, N., Susskind, J., and Zhai, S. Dart: Denoising autoregressive transformer for scalable text-to-image generation. ArXiv preprint, abs/2410.08159, 2024. URL https://arxiv.org/abs/2410.08159.
  24. 24.Ho, J. and Salimans, T. Classifier-free diffusion guidance. ArXiv preprint, abs/2207.12598, 2022. URL https://arxiv.org/abs/2207.12598.
  25. 25.Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pp. 2722–2730. PMLR, 2019. URL http://proceedings.mlr.press/v97/ho19a.html.
  26. 26.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html.
  27. 27.Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T. Cascaded diffusion models for high fidelity image generation. J. Mach. Learn. Res., 23:47:1–47:33, 2022. URL http://jmlr.org/papers/v23/21-0635.html.
  28. 28.Hoogeboom, E., Heek, J., and Salimans, T. simple diffusion: End-to-end diffusion for high resolution images. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 13213–13232. PMLR, 2023. URL https://proceedings.mlr.press/v202/hoogeboom23a.html.
  29. 29.Huang, C., Krueger, D., Lacoste, A., and Courville, A. C. Neural autoregressive flows. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 2083–2092. PMLR, 2018. URL http://proceedings.mlr.press/v80/huang18d.html.
  30. 30.Hutchinson, M. F. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines. Communications in Statistics-Simulation and Computation, 18(3):1059–1076, 1989.
  31. 31.Jabri, A., Fleet, D. J., and Chen, T. Scalable adaptive computation for iterative generation. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 14569–14589. PMLR, 2023. URL https://proceedings.mlr.press/v202/jabri23a.html.
  32. 32.Kang, M., Zhu, J., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T. Scaling up gans for text-to-image synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pp. 10124–10134. IEEE, 2023. doi: 10.1109/CVPR52729.2023.00976. URL https://doi.org/10.1109/CVPR52729.2023.00976.
  33. 33.Karras, T., Laine, S., and Aila, T. A style-based generator architecture for generative adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pp. 4401–4410. Computer Vision Foundation / IEEE, 2019. doi: 10.1109/CVPR.2019.00453. URL http://openaccess.thecvf.com/content_CVPR_2019/html/Karras_A_Style-Based_Generator_Architecture_for_Generative_Adversarial_Networks_CVPR_2019_paper.html.
  34. 34.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022.
  35. 35.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. Advances in neural information processing systems, 34:21696–21707, 2021.
  36. 36.Kingma, D. P. and Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pp. 10236–10245, 2018. URL https://proceedings.neurips.cc/paper/2018/hash/d139db6a236200b21cc7f752979132d0-Abstract.html.
  37. 37.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In Bengio, Y. and LeCun, Y. (eds.), 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014. URL http://arxiv.org/abs/1312.6114.
  38. 38.Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. Improved variational inference with inverse autoregressive flow. Advances in neural information processing systems, 29, 2016.
  39. 39.Li, T., Tian, Y., Li, H., Deng, M., and He, K. Autoregressive image generation without vector quantization. ArXiv preprint, abs/2406.11838, 2024. URL https://arxiv.org/abs/2406.11838.
  40. 40.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023a. URL https://openreview.net/pdf?id=PqvMRDCJT9t.
  41. 41.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023b. URL https://openreview.net/pdf?id=PqvMRDCJT9t.
  42. 42.Liu, G., Chen, T., and Theodorou, E. A. Second-order neural ODE optimizer. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 25267–25279, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/d4c2e4a3297fe25a71d030b67eb83bfc-Abstract.html.
  43. 43.Menick, J. and Kalchbrenner, N. Generating high fidelity images with subscale pixel networks and multidimensional upscaling. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=HylzTiC5Km.
  44. 44.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 8162–8171. PMLR, 2021. URL http://proceedings.mlr.press/v139/nichol21a.html.
  45. 45.Noroozi, M. Self-labeled conditional gans. ArXiv preprint, abs/2012.02162, 2020. URL https://arxiv.org/abs/2012.02162.
  46. 46.Papamakarios, G., Murray, I., and Pavlakou, T. Masked autoregressive flow for density estimation. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 2338–2347, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/6c1da886822c67822bcf3679d04369fa-Abstract.html.
  47. 47.Patacchiola, M., Shysheya, A., Hofmann, K., and Turner, R. E. Transformer neural autoregressive flows. ArXiv preprint, abs/2401.01855, 2024. URL https://arxiv.org/abs/2401.01855.
  48. 48.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. In The Twelfth International Conference on Learning Representations, 2024.
  49. 49.Pooladian, A., Ben-Hamu, H., Domingo-Enrich, C., Amos, B., Lipman, Y., and Chen, R. T. Q. Multisample flow matching: Straightening flows with minibatch couplings. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 28100–28127. PMLR, 2023. URL https://proceedings.mlr.press/v202/pooladian23a.html.
  50. 50.Razavi, A., van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with VQ-VAE-2. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 14837–14847, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/5f8e2fa1718d1bbcadf1cd9c7a54fb8c-Abstract.html.
  51. 51.Rezende, D. J. and Mohamed, S. Variational inference with normalizing flows. In Bach, F. R. and Blei, D. M. (eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pp. 1530–1538. JMLR.org, 2015. URL http://proceedings.mlr.press/v37/rezende15.html.
  52. 52.Roy, A., Saffar, M., Vaswani, A., and Grangier, D. Efficient content-based sparse attention with routing transformers. Transactions of the Association for Computational Linguistics, 9:53–68, 2021. doi: 10.1162/tacl_a_00353. URL https://aclanthology.org/2021.tacl-1.4.
  53. 53.Sherstinsky, A. Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network. Physica D: Nonlinear Phenomena, 404:132306, 2020.
  54. 54.Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Bach, F. R. and Blei, D. M. (eds.), Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pp. 2256–2265. JMLR.org, 2015. URL http://proceedings.mlr.press/v37/sohl-dickstein15.html.
  55. 55.Song, Y. and Dhariwal, P. Improved techniques for training consistency models. In The Twelfth International Conference on Learning Representations, 2023.
  56. 56.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=PxTIG12RRHS.
  57. 57.Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. Consistency models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 32211–32252. PMLR, 2023. URL https://proceedings.mlr.press/v202/song23a.html.
  58. 58.Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z. Autoregressive model beats diffusion: Llama for scalable image generation. CoRR, abs/2406.06525, 2024. doi: 10.48550/ARXIV.2406.06525. URL https://doi.org/10.48550/arXiv.2406.06525.
  59. 59.Tabak, E. G. and Vanden-Eijnden, E. Density estimation by dual ascent of the log-likelihood. Communications in Mathematical Sciences, 8(1):217–233, 2010.
  60. 60.Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L. Visual autoregressive modeling: Scalable image generation via next-scale prediction. ArXiv preprint, abs/2404.02905, 2024. URL https://arxiv.org/abs/2404.02905.
  61. 61.Tschannen, M., Pinto, A. S., and Kolesnikov, A. Jetformer: An autoregressive generative model of raw images and text. ArXiv preprint, abs/2411.19722, 2024. URL https://arxiv.org/abs/2411.19722.
  62. 62.Tschannen, M., Eastwood, C., and Mentzer, F. Givt: Generative infinite-vocabulary transformers. In European Conference on Computer Vision, pp. 292–309. Springer, 2025.
  63. 63.van den Oord, A., Kalchbrenner, N., Espeholt, L., Kavukcuoglu, K., Vinyals, O., and Graves, A. Conditional image generation with pixelcnn decoders. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pp. 4790–4798, 2016a. URL https://proceedings.neurips.cc/paper/2016/hash/b1301141feffabac455e1f90a7de2054-Abstract.html.
  64. 64.van den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K. Pixel recurrent neural networks. In Balcan, M. and Weinberger, K. Q. (eds.), Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, volume 48 of JMLR Workshop and Conference Proceedings, pp. 1747–1756. JMLR.org, 2016b. URL http://proceedings.mlr.press/v48/oord16.html.
  65. 65.van den Oord, A., Vinyals, O., and Kavukcuoglu, K. Neural discrete representation learning. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 6306–6315, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/7a98af17e63a0ac09ce2e96d03992fbc-Abstract.html.
  66. 66.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need.(nips), 2017. Advances in neural information processing systems, 30, 2017.
  67. 67.Wiatrak, M., Albrecht, S. V., and Nystrom, A. Stabilizing generative adversarial networks: A survey. ArXiv preprint, abs/1910.00927, 2019. URL https://arxiv.org/abs/1910.00927.
  68. 68.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research, 2(3):5, 2022.
  69. 69.Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y. Open-sora: Democratizing efficient video production for all, 2024. URL https://github.com/hpcaitech/Open-Sora.
  70. 70.Zhuang, J., Dvornek, N. C., Tatikonda, S., and Duncan, J. S. MALI: A memory efficient and reverse accurate integrator for neural odes. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=blfSjHeFM_e.

Citation

MLA
Zhai, S., et al. “Normalizing Flows Are Capable Generative Models”. arXiv, 2024, http://arxiv.org/abs/2412.06329v3.
APA
Zhai, S., Zhang, R., Nakkiran, P., Berthelot, D., Gu, J., Zheng, H., Chen, T., Bautista, M. A., Jaitly, N., & Susskind, J. (2024). Normalizing Flows are Capable Generative Models. arXiv. http://arxiv.org/abs/2412.06329v3
Chicago
Zhai, S., R. Zhang, P. Nakkiran, et al. 2024. “Normalizing Flows Are Capable Generative Models”. arXiv. http://arxiv.org/abs/2412.06329v3.
Harvard
Zhai, S. et al. (2024) “Normalizing Flows are Capable Generative Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2412.06329v3.
Vancouver
1. Zhai S, Zhang R, Nakkiran P, Berthelot D, Gu J, Zheng H, Chen T, Bautista MA, Jaitly N, Susskind J (2024) Normalizing Flows are Capable Generative Models. arXiv

BibTeX

@article{zhai2024normalizing,
  title = {Normalizing Flows are Capable Generative Models},
  author = {Zhai, Shuangfei and Zhang, Ruixiang and Nakkiran, Preetum and Berthelot, David and Gu, Jiatao and Zheng, Huangjie and Chen, Tianrong and Bautista, Miguel Angel and Jaitly, Navdeep and Susskind, Josh},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2412.06329v3},
  eprint = {2412.06329}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/