Normalizing Flows are Capable Generative Models
Shuangfei ZhaiRuixiang ZhangPreetum NakkiranDavid BerthelotJiatao GuHuangjie ZhengTianrong ChenMiguel ngel BautistaNavdeep JaitlyJoshua M. Susskind
Demonstrates that normalizing flows can rival diffusion models in image generation quality while achieving state-of-the-art exact likelihood estimation through a scalable Transformer-based architecture paired with noise augmentation and guidance.
Generative artificial intelligence has largely been dominated by diffusion models and large language models, while normalizing flows—a foundational class of models known for exact likelihood computation and efficient mathematical invertibility—have lagged behind in perceptual generation quality and practical adoption. The article addresses this performance gap by investigating whether normalizing flows are fundamentally limited or merely constrained by suboptimal architectures and training procedures. Its primary objective is to introduce TARFLOW, a scalable transformer-based architecture combined with novel training and sampling techniques, to demonstrate that standalone normalizing flows can achieve state-of-the-art density estimation and image generation competitive with leading generative frameworks.
The authors evaluate TARFLOW across standard image benchmarks, including unconditional and class-conditional ImageNet (at 64x64 and 128x128 resolutions) and AFHQ (at 256x256 resolution). The approach replaces conventional masked multilayer perceptrons with causal vision transformer blocks that operate on image patches across alternating sequence directions. To bolster generative quality, the methodology introduces three core techniques: training with moderate Gaussian noise augmentation rather than narrow uniform noise, applying a post-training score-based denoising step via Tweedie's formula directly on the learned density, and adapting classifier-free guidance mechanisms for both conditional and unconditional generation.
The article reports several critical findings. First, TARFLOW sets a new state of the art in image density estimation on unconditional ImageNet 64x64, achieving a negative log-likelihood of 2.99 bits per dimension and breaking the sub-3.0 threshold for the first time. Second, in image generation quality, TARFLOW achieves competitive Fréchet Inception Distance (FID) scores—reaching 2.66 on conditional ImageNet 64x64 and 5.03 on ImageNet 128x128—outperforming traditional generative adversarial network (GAN) baselines and approaching diffusion model benchmarks. Third, ablation analyses show that noise augmentation paired with score-based denoising drastically reduces visual artifacts, while guidance reliably trades sample diversity for sharper class fidelity. Finally, depth ablations reveal that balancing the number of sequential flow blocks and layers per block yields optimal performance, confirming the model scales smoothly with compute.
These findings carry significant technical and operational implications. By proving that normalizing flows alone can generate high-fidelity images directly from continuous pixels without complex multi-stage tokenization or vector quantization, TARFLOW establishes an alternative, mathematically tractable paradigm for generative modeling. For decision-makers, this opens opportunities to leverage the exact likelihood tracking and deterministic objectives of flows without sacrificing visual quality. Although sampling currently requires sequential autoregressive generation—taking approximately two minutes for a batch of 32 images on an A100 GPU—the architecture provides a robust, modular baseline that can leverage existing transformer optimization ecosystems.
Organizations evaluating generative modeling pipelines should consider testing TARFLOW architectures where exact probability estimation, direct pixel generation, and training stability are critical requirements. Next steps should focus on exploring optimal guidance schedules, scaling model capacity to higher resolutions, and implementing engineering optimizations such as advanced caching and gradient checkpointing to reduce sampling latency and memory overhead. While the reported empirical confidence is high across standard benchmarks, stakeholders should note that sampling speed remains a key operational bottleneck compared to non-autoregressive alternatives before deploying the framework into real-time production workflows.
- Paper: Normalizing Flows: An Introduction and Review of Current Methods, Ivan Kobyzev et al. (2020). It provides a foundational overview of normalizing flows, invertible mappings, and exact density estimation principles upon which TARFLOW builds.
- Paper: Normalizing Flows for Probabilistic Modeling and Inference, George Papamakarios et al. (2019). It details the core mathematical framework of change-of-variables, coupling layers, and autoregressive flow transformations essential for understanding modern flow architectures.
- Paper: Density estimation using Real NVP, Laurent Dinh et al. (2016). It introduces the coupling-based multi-scale flow architecture that TARFLOW adapts and replaces with causal transformer blocks.
- Paper: Glow: Generative Flow with Invertible 1x1 Convolutions, Diederik P. Kingma et al. (2018). It establishes standard flow-based image synthesis techniques and negative log-likelihood benchmarks that TARFLOW directly surpasses.
- Paper: Image Transformer, Niki Parmar et al. (2018). It pioneers applying self-attention and transformer architectures to autoregressive generative modeling on visual patches and pixels.
- Paper: Improved Techniques for Training Score-Based Generative Models, Yang Song et al. (2020). It details noise-conditioned score techniques and Tweedie-based denoising methods that motivate TARFLOW's post-training artifact-reduction step.
- Paper: Variational Flow Maps: Make Some Noise for One-Step Conditional Generation, Abbas Mammadov et al. (2026). It builds beyond autoregressive flow generation by introducing variational flow maps that enable fast one-step conditional synthesis on high-resolution image benchmarks.
- Paper: Generative Modeling via Drifting, Mingyang Deng et al. (2026). It advances direct pixel-space generation beyond sequential flow trajectories by learning one-step drifting fields.
