Pixel Recurrent Neural Networks
Aaron van den OordNal KalchbrennerKoray Kavukcuoglu
Introduces two-dimensional recurrent neural network architectures with residual connections to sequentially model discrete pixel dependencies, setting a new standard for generative image modeling and log-likelihood performance on ImageNet.
Pixel Recurrent Neural Networks advance generative modeling of natural images by treating pixel prediction as a sequential task that captures full spatial and color dependencies. The work addresses the core difficulty of building models that are simultaneously expressive enough to represent complex image structure, tractable enough to compute exact likelihoods, and scalable to large datasets—an obstacle that has limited prior approaches such as latent-variable models and earlier autoregressive networks.
The authors set out to demonstrate that two-dimensional recurrent architectures can achieve substantially higher likelihoods on standard image benchmarks while also producing coherent samples. They constructed three main variants: Row LSTM and Diagonal BiLSTM networks that use specialized convolutional recurrences to scan images row-wise or diagonally, and a fully convolutional PixelCNN that trades unbounded context for greater parallelism during training. All models factor the joint distribution over raw RGB pixels into a product of conditional multinomial distributions implemented with softmax layers, enforce correct conditioning through masked convolutions, and employ residual connections to train networks up to twelve layers deep. Experiments covered MNIST, CIFAR-10, and ImageNet images resized to 32×32 and 64×64 pixels, with performance measured by negative log-likelihood after appropriate noise dequantization.
The strongest model, a twelve-layer Diagonal BiLSTM, reached 79.20 nats on MNIST and 3.00 bits per dimension on CIFAR-10, improving on previous published results by roughly 1–2 nats and 0.47 bits per dimension respectively. On ImageNet the same family of models established the first reported likelihood baselines at 3.86 bits per dimension for 32×32 images and 3.63 bits per dimension for 64×64 images. Samples drawn from the networks appear sharp and globally consistent, and image-completion examples show that the models fill occluded regions with diverse yet plausible content. Architectural ablations confirmed that the discrete softmax output outperforms continuous mixture-density alternatives and that residual connections enable effective scaling with depth.
These gains indicate that autoregressive recurrent models can now serve as practical density estimators for high-resolution images, supporting downstream uses such as lossless compression, inpainting, and conditional generation. Because likelihood improves steadily with added depth and data, further increases in model size and training compute are expected to yield additional gains. Generation remains inherently sequential, however, so practical deployment at high resolution will require either faster sampling techniques or hybrid architectures that combine the accuracy of PixelRNNs with the speed of parallel alternatives.
No sufficiently relevant recommendations were found.
- Paper: Video Pixel Networks, Nal Kalchbrenner et al. (2017). This work carries PixelRNN’s masked pixel-level prediction into video, extending its dependency modeling across both space and time.
- Paper: Image Transformer, Niki Parmar et al. (2018). Building on autoregressive image modeling, this paper replaces recurrent and convolutional context mechanisms with self-attention to capture broader visual dependencies more efficiently.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). This work advances PixelCNN-style autoregressive modeling by applying it to compressed, hierarchical image codes rather than generating every pixel directly.
- Paper: Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images, Rewon Child (2021). After seeing PixelRNN establish strong image likelihoods, this paper tests whether much deeper hierarchical VAEs can match or surpass autoregressive models while avoiding their slow sequential generation.
