keyword
PixelCNN Decoders
PixelCNN decoders are autoregressive neural network components based on the PixelCNN architecture that generate or reconstruct images conditioned on external inputs, such as class labels, descriptive tags, or latent representations from an encoder. Unlike conventional decoders that predict all pixel values simultaneously and independently, a PixelCNN decoder models the joint conditional probability distribution of image pixels sequentially, generating each pixel based on previously generated pixels as well as the provided conditioning signal. By employing masked or gated convolutional layers to enforce causal dependency, these decoders capture complex local textures and fine-grained spatial relationships, making them effective conditional image generators and expressive decoding mechanisms within architectures like autoencoders and variational autoencoders.
4 items

Image Super-Resolution via Iterative Refinement
Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, Mohammad Norouzi
Why you should read this
Proposes SR3, a conditional diffusion model that outperforms generative adversarial networks in high-magnification super-resolution by iteratively refining noisy inputs into photorealistic images.
We present SR3, an approach to image Super-Resolution via Repeated Refinement. SR3 adapts denoising diffusion probabilistic models to conditional image generation and performs super-resolution through a stochastic denoising process. Inference starts with pure Gaussian noise and iteratively refines the noisy output using a U-Net model trained on denoising at various noise levels. SR3 exhibits strong performance on super-resolution tasks at different magnification factors, on faces and natural images. We conduct human evaluation on a standard 8X face super-resolution task on CelebA-HQ, comparing with SOTA GAN methods. SR3 achieves a fool rate close to 50%, suggesting photo-realistic outputs, while GANs do not exceed a fool rate of 34%. We further show the effectiveness of SR3 in cascaded image generation, where generative models are chained with super-resolution models, yielding a competitive FID score of 11.3 on ImageNet.
Added
2026-09-14

Video Pixel Networks
Nal Kalchbrenner, Aäron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, Koray Kavukcuoglu
Why you should read this
Proposes a novel Video Pixel Network that achieves near-optimal video prediction performance, generates realistic video samples, and generalizes effectively to novel objects, representing a significant advancement in generative video modeling.
We propose a probabilistic video model, the Video Pixel Network (VPN), that estimates the discrete joint distribution of the raw pixel values in a video. The model and the neural architecture reflect the time, space and color structure of video tensors and encode it as a four-dimensional dependency chain. The VPN approaches the best possible performance on the Moving MNIST benchmark, a leap over the previous state of the art, and the generated videos show only minor deviations from the ground truth. The VPN also produces detailed samples on the action-conditional Robotic Pushing benchmark and generalizes to the motion of novel objects.
Added
2026-03-11


Variational Lossy Autoencoder
Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, Pieter Abbeel
Why you should read this
Reveals that the common failure of VAEs to use their latent codes when paired with powerful decoders isn't a bug but a controllable feature—by deliberately limiting what the decoder can model locally (like small texture patches), you can force the latent code to capture exactly the global structure you care about while achieving state-of-the-art density estimation.
Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good representation for 2D images might be one that describes only global structure and discards information about detailed texture. In this paper, we present a simple but principled method to learn such global representations by combining Variational Autoencoder (VAE) with neural autoregressive models such as RNN, MADE and PixelRNN/CNN. Our proposed VAE model allows us to have control over what the global latent code can learn and , by designing the architecture accordingly, we can force the global latent code to discard irrelevant information such as texture in 2D images, and hence the VAE only "autoencodes" data in a lossy fashion. In addition, by leveraging autoregressive models as both prior distribution p(z) and decoding distribution p(x|z), we can greatly improve generative modeling performance of VAEs, achieving new state-of-the-art results on MNIST, OMNIGLOT and Caltech-101 Silhouettes density estimation tasks.
Added
2026-02-21

Conditional Image Generation with PixelCNN Decoders
Aäron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Koray Kavukcuoglu, Oriol Vinyals, Alex Graves
Why you should read this
Presents masked-convolution autoregression for images with exact likelihood and straightforward conditioning, giving you a clean, reproducible AR baseline for density modeling and sampling.
This work explores conditional image generation with a new image density model based on the PixelCNN architecture. The model can be conditioned on any vector, including descriptive labels or tags, or latent embeddings created by other networks. When conditioned on class labels from the ImageNet database, the model is able to generate diverse, realistic scenes representing distinct animals, objects, landscapes and structures. When conditioned on an embedding produced by a convolutional network given a single image of an unseen face, it generates a variety of new portraits of the same person with different facial expressions, poses and lighting conditions. We also show that conditional PixelCNN can serve as a powerful decoder in an image autoencoder. Additionally, the gated convolutional layers in the proposed model improve the log-likelihood of PixelCNN to match the state-of-the-art performance of PixelRNN on ImageNet, with greatly reduced computational cost.
Added
2025-09-14
License
Published with permission
