Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks

Emily DentonSoumith ChintalaArthur SzlamRob Fergus

article2015NeurIPS2,352 citations

Proposes a multi-scale Laplacian pyramid architecture for generative adversarial networks that enables high-quality natural image synthesis through a coarse-to-fine generation cascade.

Listen

Generating realistic, high-resolution natural images has long been a difficult challenge in computer vision because images contain complex structures across many spatial scales. While deep learning has driven rapid progress in image classification, existing generative systems have struggled to produce high-fidelity full scenes without severe visual artifacts or training instabilities.

The article demonstrates a framework, known as the Laplacian Generative Adversarial Network (LAPGAN), designed to synthesize higher-quality images. It evaluates whether breaking image synthesis into a sequence of coarse-to-fine stages can significantly improve sample quality over standard generative adversarial methods.

To accomplish this, the authors combined generative adversarial networks (adversarial neural networks consisting of a generator competing against a discriminator) with a classic multi-scale image processing structure called a Laplacian pyramid. Instead of generating a full image in a single pass, the model starts by producing a tiny, low-resolution residual image and then progressively adds fine detail at each higher resolution level using a cascade of convolutional networks trained independently. The approach was tested across three image datasets: CIFAR10 (small object crops), STL (unlabeled natural images), and LSUN (a large database of approximately 10 million scene images), and evaluated through statistical likelihood estimations, visual inspection, and human perception trials.

The findings show substantial improvements in image quality. In human testing involving 15 volunteers and approximately 10,000 trials, samples from the class-conditional version of the model were mistaken for real images roughly 40% of the time, compared to 10% or less for standard generative adversarial networks. Statistically, the multi-scale approach yielded significantly higher log-likelihood estimates on benchmark datasets than standard baselines. Furthermore, the model successfully scaled to complex, multi-element scenes on the LSUN dataset without simply memorizing or copying training examples.

These results demonstrate that abandoning the attempt to enforce overall global fidelity in a single step and instead focusing on plausible, scale-by-scale refinements makes generative training far more robust. This architecture lowers the risk of neural network overfitting, provides an efficient feed-forward sampling process, and establishes a practical pathway for applying generative techniques to higher-resolution imaging.

For future work, the article suggests extending this multi-scale refinement framework to other complex signal domains that exhibit hierarchical structures, such as audio or video. Scaling the approach with larger datasets and deeper networks is recommended to further refine scene coherence and visual realism.

Key limitations include remaining visual artifacts, as human observers still identified real images correctly over 90% of the time, meaning generated samples are not yet indistinguishable from reality. In addition, evaluating model probability remains inherently challenging, relying on Parzen-window density approximations. Decision-makers should view this framework as a proven, high-performing foundation for image synthesis, while recognizing that full photorealism will require further scaling and refinement.

  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This foundational paper introduces the original Generative Adversarial Networks framework that the Laplacian pyramid architecture builds upon at each level.
Cover for Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks

Abstract

In this paper we introduce a generative parametric model capable of producing high quality samples of natural images. Our approach uses a cascade of convolutional networks within a Laplacian pyramid framework to generate images in a coarse-to-fine fashion. At each level of the pyramid, a separate generative convnet model is trained using the Generative Adversarial Nets (GAN) approach (Goodfellow et al.). Samples drawn from our model are of significantly higher quality than alternate approaches. In a quantitative assessment by human evaluators, our CIFAR10 samples were mistaken for real images around 40% of the time, compared to 10% for samples drawn from a GAN baseline model. We also show samples from models trained on the higher resolution images of the LSUN scene dataset.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 Approach
  • 2.1 Generative Adversarial Networks
  • 2.2 Laplacian Pyramid
  • 2.3 Laplacian Generative Adversarial Networks (LAPGAN)
  • 3 Model Architecture & Training
  • 3.1 CIFAR10 and STL
  • 3.2 LSUN
  • 4 Experiments
  • 4.1 Evaluation of Log-Likelihood
  • 4.2 Model Samples
  • 4.3 Human Evaluation of Samples
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — Laplacian Generative Adversarial Networks (LAPGAN) Framework

    model/method

    Laplacian Generative Adversarial Networks (LAPGAN) combine a linear Laplacian pyramid image representation with conditional Generative Adversarial Networks (CGAN) to generate high-resolution natural images in a coarse-to-fine hierarchy. A Laplacian pyramid decomposes an image II into a sequence of band-pass differential images [h0,h1,…,hK−1][h_0, h_1, \dots, h_{K-1}] and a low-frequency residual hK=IKh_K = I_K, where a downsampling operator d(⋅)d(\cdot) blurs and decimates an image by a factor of 2, and an upsampling operator u(⋅)u(\cdot) smooths and expands it by a factor of 2.

    In LAPGAN, the generation process is distributed across K+1K+1 levels:

    • At the coarsest scale KK, a standard unconditional GAN (GK,DKG_K, D_K) synthesizes the low-frequency residual image I~K=GK(zK)\tilde{I}_K = G_K(z_K), where zK∼pnoise(zK)z_K \sim p_{\text{noise}}(z_K) is a random noise vector.
    • At each finer scale k∈{K−1,…,0}k \in \{K-1, \dots, 0\}, a conditional GAN (Gk,DkG_k, D_k) generates the high-frequency band-pass coefficient image h~k=Gk(zk,u(I~k+1))\tilde{h}_k = G_k(z_k, u(\tilde{I}_{k+1})), conditioned on the upsampled image from the coarser level u(I~k+1)u(\tilde{I}_{k+1}) and a noise vector zkz_k.
    • The reconstructed image at level kk is computed recursively as: I~k=u(I~k+1)+h~k=u(I~k+1)+Gk(zk,u(I~k+1))\tilde{I}_k = u(\tilde{I}_{k+1}) + \tilde{h}_k = u(\tilde{I}_{k+1}) + G_k(z_k, u(\tilde{I}_{k+1})) yielding the full-resolution output I~0\tilde{I}_0.

    By breaking the image generation problem into successive scale-conditional residual generation tasks, LAPGAN avoids the need for a global discriminator over the entire cascade, which stabilizes adversarial training and mitigates sample memorization.

  2. Knowl 2 — LAPGAN Coarse-to-Fine Sampling Procedure

    algorithm

    Given a trained cascade of generative models {G0,G1,…,GK}\{G_0, G_1, \dots, G_K\}, where GKG_K generates the coarsest residual image and GkG_k (for k<Kk < K) generates high-frequency detail conditioned on an upsampled lower-resolution image, the image sampling procedure operates sequentially from scale KK down to scale 00.

    Input: Generative models G0,G1,…,GKG_0, G_1, \dots, G_K, upsampling operator u(⋅)u(\cdot), noise distributions pnoise(zk)p_{\text{noise}}(z_k) for k∈{0,…,K}k \in \{0, \dots, K\}
    Output: Full-resolution generated image I~0\tilde{I}_0
    Sample noise vector zK∼pnoise(zK)z_K \sim p_{\text{noise}}(z_K)
    Generate coarsest residual image I~K=GK(zK)\tilde{I}_K = G_K(z_K)
    for k=K−1k = K - 1 down to 00 do
        Upsample previous image: lk=u(I~k+1)l_k = u(\tilde{I}_{k+1})
        Sample noise vector zk∼pnoise(zk)z_k \sim p_{\text{noise}}(z_k)
        Generate high-frequency residual: h~k=Gk(zk,lk)\tilde{h}_k = G_k(z_k, l_k)
        Reconstruct image at scale kk: I~k=lk+h~k\tilde{I}_k = l_k + \tilde{h}_k
    return I~0\tilde{I}_0

    At the coarsest level, GKG_K maps a 100-dimensional uniform noise vector zK∼U[−1,1]z_K \sim \mathcal{U}[-1, 1] directly to a low-frequency image (e.g., 8×88 \times 8 or 4×44 \times 4). At each subsequent stage kk, the noise zk∼U[−1,1]z_k \sim \mathcal{U}[-1, 1] is concatenated as an additional spatial channel to the upsampled conditional image lk=u(I~k+1)l_k = u(\tilde{I}_{k+1}) and passed forward through convolutional layers to produce the difference image h~k\tilde{h}_k.

  3. Knowl 3 — Independent Scale-Wise Training of LAPGAN

    model/method

    In LAPGAN, the generative models {G0,…,GK}\{G_0, \dots, G_K\} and discriminative models {D0,…,DK}\{D_0, \dots, D_K\} at each scale of the pyramid are trained independently rather than via end-to-end backpropagation through the cascade.

    For a full-resolution training image I0=II_0 = I, a Gaussian pyramid [I0,I1,…,IK][I_0, I_1, \dots, I_K] is built via recursive downsampling Ik+1=d(Ik)I_{k+1} = d(I_k).

    1. Coarsest Level (k=Kk = K): The generator GKG_K and discriminator DKD_K are trained as a standard GAN on the smallest Gaussian residual images IKI_K.
    2. Intermediate Levels (k∈{0,…,K−1}k \in \{0, \dots, K-1\}):
      • A low-pass image is created by upsampling: lk=u(Ik+1)l_k = u(I_{k+1}).
      • The true high-frequency difference image is computed as hk=Ik−lkh_k = I_k - l_k.
      • The generator GkG_k takes noise zk∼U[−1,1]z_k \sim \mathcal{U}[-1, 1] and the conditioning low-pass image lkl_k to produce a synthetic high-pass image h~k=Gk(zk,lk)\tilde{h}_k = G_k(z_k, l_k).
      • The discriminator DkD_k receives either the real pair (hk,lk)(h_k, l_k) or the generated pair (h~k,lk)(\tilde{h}_k, l_k). The conditioning image lkl_k is explicitly added to hkh_k (or h~k\tilde{h}_k) prior to the first convolutional layer of DkD_k.
      • The parameters of GkG_k and DkD_k are optimized using the conditional GAN minimax objective: min⁡Gkmax⁡DkEIk∼pdata(Ik)[log⁡Dk(hk,lk)]+Ezk∼pnoise(zk),Ik+1∼pdata(Ik+1)[log⁡(1−Dk(Gk(zk,lk),lk))]\min_{G_k} \max_{D_k} \mathbb{E}_{I_k \sim p_{\text{data}}(I_k)} \left[ \log D_k(h_k, l_k) \right] + \mathbb{E}_{z_k \sim p_{\text{noise}}(z_k), I_{k+1} \sim p_{\text{data}}(I_{k+1})} \left[ \log(1 - D_k(G_k(z_k, l_k), l_k)) \right]

    This decoupled stage-wise training prevents high-capacity convolutional networks from memorizing training samples and eliminates the numerical instabilities of deep cascade optimization.

  4. Knowl 4 — Class-Conditional LAPGAN Architecture and Conditioning Mechanism

    model/method

    To incorporate class supervision into LAPGAN, both the generator GkG_k and the discriminator DkD_k at scale kk are conditioned on a one-hot class label vector c∈{0,1}Cc \in \{0, 1\}^C, where CC is the total number of classes.

    The class label vector cc is projected through a linear layer whose output matches the spatial dimensions of the layer's feature maps, reshaped into a single 2D feature map plane, and concatenated along the channel dimension with the first-layer feature maps of GkG_k and DkD_k.

    The resulting conditional objective function for each level kk is: min⁡Gkmax⁡DkE(hk,lk,c)∼pdata[log⁡Dk(hk,lk,c)]+Ezk∼pnoise,(lk,c)∼pdata[log⁡(1−Dk(Gk(zk,lk,c),lk,c))]\min_{G_k} \max_{D_k} \mathbb{E}_{(h_k, l_k, c) \sim p_{\text{data}}} \left[ \log D_k(h_k, l_k, c) \right] + \mathbb{E}_{z_k \sim p_{\text{noise}}, (l_k, c) \sim p_{\text{data}}} \left[ \log(1 - D_k(G_k(z_k, l_k, c), l_k, c)) \right]

    Class conditioning guides coarse-to-fine synthesis to produce objects belonging to a specific class with sharper boundaries and consistent category semantics.

  5. Knowl 5 — Multi-Scale Gaussian Parzen Window Log-Likelihood Estimation

    model/method

    Because GANs and LAPGAN do not define an explicit tractable probability density function, image log-likelihood is estimated using a multi-scale Gaussian Parzen window density estimator adapted to the Laplacian pyramid structure.

    Let an image I∈Rd2I \in \mathbb{R}^{d^2} be decomposed into a low-pass component l=d(I)l = d(I) and a high-pass component h=I−u(d(I))h = I - u(d(I)), where downsampling d(I)d(I) computes the mean over each disjoint 2×22 \times 2 pixel block, and uu removes the mean. Expressing hh in an orthonormal basis of the range of uu makes the linear mapping I↦(l,h)I \mapsto (l, h) unitary. The joint probability density factors as: p(I)=q0(l,h) q1(l)=q0(d(I),h(I)) q1(d(I))p(I) = q_0(l, h) \, q_1(l) = q_0(d(I), h(I)) \, q_1(d(I))

    For a KK-level pyramid, the density is estimated by accumulating Parzen window estimates across all levels: log⁡p(I)=log⁡qK(lK)+∑k=0K−1log⁡qk(lk,hk)\log p(I) = \log q_K(l_K) + \sum_{k=0}^{K-1} \log q_k(l_k, h_k) where:

    • qK(lK)∝∑i=1NKexp⁡(−∥lK−lK,i∥22σK2)q_K(l_K) \propto \sum_{i=1}^{N_K} \exp\left(-\frac{\|l_K - l_{K, i}\|^2}{2\sigma_K^2}\right) is the Parzen window density over the coarsest scale using generated residual samples {lK,i}\{l_{K, i}\}.
    • qk(lk,hk)∝∑i=1Nkexp⁡(−∥hk−hk,i(lk)∥22σk2)q_k(l_k, h_k) \propto \sum_{i=1}^{N_k} \exp\left(-\frac{\|h_k - h_{k, i}(l_k)\|^2}{2\sigma_k^2}\right) evaluates the true high-pass image hkh_k against high-pass samples generated conditioned on the true low-pass image lkl_k.
    • Parzen window bandwidths σk\sigma_k are tuned on a held-out validation set.
  6. Knowl 6 — Parzen-Window Log-Likelihood Evaluation on CIFAR-10 and STL

    data/table

    Log-likelihood estimates on held-out test sets evaluated using Gaussian Parzen window density estimation (with 50,000 generated samples) demonstrate that LAPGAN achieves substantially higher log-likelihood values than the standard unconditional GAN of Goodfellow et al. (2014) on both CIFAR-10 and STL-10 datasets (at 32×3232 \times 32 pixel resolution).

    Model CIFAR10 STL (@32×3232\times32)
    GAN −3617±353-3617 \pm 353 −3661±347-3661 \pm 347
    LAPGAN −1799±826-1799 \pm 826 −2906±728-2906 \pm 728

    The values report estimated test set log-likelihood ±\pm standard error. The LAPGAN model achieves a higher (less negative) log-likelihood by over 1800 nats on CIFAR-10 and over 750 nats on STL-10, indicating an improved density model over the single-stage GAN baseline.

  7. Knowl 7 — Human Subject Visual Realism Assessment on CIFAR-10

    empirical result

    In a quantitative psychophysical experiment evaluating image generation quality, 15 human subjects were presented with images shown for durations ranging from 50 ms to 2000 ms (followed by a gray mask) and tasked with classifying whether each image was real or generated. Across ∼10,000\sim 10,000 collected trials on CIFAR-10:

    • Samples from the Class-Conditional LAPGAN (CC-LAPGAN) were mistaken for real images approximately 40%40\% of the time.
    • Unconditional LAPGAN samples achieved a real-classification rate of approximately 30%–35%30\%\text{--}35\%.
    • Samples from the standard GAN baseline were mistaken for real images ≤10%\le 10\% of the time.
    • Real CIFAR-10 control images were correctly identified as real >90%> 90\% of the time.

    While still distinct from real images, CC-LAPGAN samples demonstrated an approximate fourfold improvement in fooling human evaluators over the standard single-stage GAN baseline.

  8. Knowl 8 — LAPGAN Neural Network Architectures and Optimization Setup

    experimental setup

    For CIFAR-10 (32×3232 \times 32 resolution) and STL (96×9696 \times 96 resolution) datasets, LAPGAN architectures and training settings are configured as follows:

    1. Initial Coarsest Scale (8×88 \times 8 resolution):

      • Fully connected networks for both GKG_K and DKD_K with 2 hidden layers and ReLU activations.
      • Discriminator DKD_K: 600 units per hidden layer with Dropout.
      • Generator GKG_K: 1200 units per hidden layer.
      • Latent noise zKz_K: 100-dimensional vector sampled from U[−1,1]\mathcal{U}[-1, 1].
    2. Subsequent Pyramid Scales:

      • CIFAR-10 levels: 8×8→14×14→28×288 \times 8 \to 14 \times 14 \to 28 \times 28 (trained on four 28×2828 \times 28 crops per 32×3232 \times 32 image).
      • STL levels: 8×8→16×16→32×32→64×64→96×968 \times 8 \to 16 \times 16 \to 32 \times 32 \to 64 \times 64 \to 96 \times 96 (4 conditional GAN stages).
      • At each conditional scale kk, GkG_k is a 3-layer convolutional network and DkD_k is a 2-layer convolutional network.
      • Noise zk∼U[−1,1]z_k \sim \mathcal{U}[-1, 1] is provided as a 4th input plane concatenated with the 3 RGB color channels of the upsampled low-pass image lkl_k.
    3. Optimization Hyperparameters:

      • Optimizer: Stochastic Gradient Descent (SGD).
      • Initial learning rate: 0.020.02, decayed at each epoch by a factor of (1+4×10−5)−1(1 + 4 \times 10^{-5})^{-1}.
      • Momentum: initialized at 0.50.5 and increased linearly by 0.00080.0008 per epoch up to a maximum of 0.80.8.
      • Model selection: Parzen-window estimator on validation set.
  9. Knowl 9 — Multi-Scale Scene Synthesis on the LSUN Dataset

    empirical result

    LAPGAN was scaled to synthesize 64×6464 \times 64 natural scene images across 10 categories of the LSUN dataset (including bedroom, tower, and church front) using a 5-scale pyramid (4×4→8×8→16×16→32×32→64×644 \times 4 \to 8 \times 8 \to 16 \times 16 \to 32 \times 32 \to 64 \times 64).

    • Architecture: Deep convolutional networks where GkG_k consists of 5 layers ({64, 368, 128, 224} feature maps with 7×77 \times 7 filters, ReLUs, batch normalization, Dropout, and linear output) and DkD_k consists of 3 hidden layers ({48, 448, 416} feature maps with sigmoid output).
    • Generation Behavior: When initialized with downsampled 4×44 \times 4 low-resolution seeds from held-out validation images, the model generates diverse, coherent 64×6464 \times 64 images capturing long-range structural dependencies and recomposing scene elements (e.g., spires, windows, furniture) into plausible scene arrangements without duplicating training examples.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.P. J. Burt, Edward, and E. H. Adelson. The laplacian pyramid as a compact image code. IEEE Transactions on Communications, 31:532–540, 1983.
  2. 2.J. S. De Bonet. Multiresolution sampling procedure for analysis and synthesis of texture images. In Proceedings of the 24th annual conference on Computer graphics and interactive techniques, pages 361–368. ACM Press/Addison-Wesley Publishing Co., 1997.
  3. 3.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. IEEE, 2009.
  4. 4.E. Denton, S. Chintala, A. Szlam, and R. Fergus. Deep generative image models using a laplacian pyramid of adversarial networks: Supplementary material. http://soumith.ch/eyescream.
  5. 5.A. Dosovitskiy, J. T. Springenberg, and T. Brox. Learning to generate chairs with convolutional neural networks. arXiv preprint arXiv:1411.5928, 2014.
  6. 6.A. A. Efros and T. K. Leung. Texture synthesis by non-parametric sampling. In ICCV, volume 2, pages 1033–1038. IEEE, 1999.
  7. 7.S. A. Eslami, N. Heess, C. K. Williams, and J. Winn. The shape boltzmann machine: a strong model of object shape. International Journal of Computer Vision, 107(2):155–176, 2014.
  8. 8.W. T. Freeman, T. R. Jones, and E. C. Pasztor. Example-based super-resolution. Computer Graphics and Applications, IEEE, 22(2):56–65, 2002.
  9. 9.J. Gauthier. Conditional generative adversarial nets for convolutional face generation. Class Project for Stanford CS231N: Convolutional Neural Networks for Visual Recognition, Winter semester 2014 2014.
  10. 10.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, pages 2672–2680. 2014.
  11. 11.K. Gregor, I. Danihelka, A. Graves, and D. Wierstra. DRAW: A recurrent neural network for image generation. CoRR, abs/1502.04623, 2015.
  12. 12.J. Hays and A. A. Efros. Scene completion using millions of photographs. ACM Transactions on Graphics (TOG), 26(3):4, 2007.
  13. 13.G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504–507, 2006.
  14. 14.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167v3, 2015.
  15. 15.D. P. Kingma and M. Welling. Auto-encoding variational bayes. ICLR, 2014.
  16. 16.A. Krizhevsky, G. E. Hinton, et al. Factored 3-way restricted boltzmann machines for modeling natural images. In AISTATS, pages 621–628, 2010.
  17. 17.M. Mirza and S. Osindero. Conditional generative adversarial nets. CoRR, abs/1411.1784, 2014.
  18. 18.B. A. Olshausen and D. J. Field. Sparse coding with an overcomplete basis set: A strategy employed by v1? Vision research, 37(23):3311–3325, 1997.
  19. 19.S. Osindero and G. E. Hinton. Modeling image patches with a directed hierarchy of markov random fields. In J. Platt, D. Koller, Y. Singer, and S. Roweis, editors, NIPS, pages 1121–1128. 2008.
  20. 20.J. Portilla and E. P. Simoncelli. A parametric texture model based on joint statistics of complex wavelet coefficients. International Journal of Computer Vision, 40(1):49–70, 2000.
  21. 21.M. Ranzato, V. Mnih, J. M. Susskind, and G. E. Hinton. Modeling natural images using gated MRFs. IEEE Transactions on Pattern Analysis & Machine Intelligence, (9):2206–2222, 2013.
  22. 22.D. J. Rezende, S. Mohamed, and D. Wierstra. Stochastic backpropagation and variational inference in deep latent gaussian models. arXiv preprint arXiv:1401.4082, 2014.
  23. 23.S. Roth and M. J. Black. Fields of experts: A framework for learning image priors. In In CVPR, pages 860–867, 2005.
  24. 24.R. Salakhutdinov and G. E. Hinton. Deep boltzmann machines. In AISTATS, pages 448–455, 2009.
  25. 25.E. P. Simoncelli, W. T. Freeman, E. H. Adelson, and D. J. Heeger. Shiftable multiscale transforms. Information Theory, IEEE Transactions on, 38(2):587–607, 1992.
  26. 26.J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. CoRR, abs/1503.03585, 2015.
  27. 27.L. Theis and M. Bethge. Generative image modeling using spatial LSTMs. Dec 2015.
  28. 28.P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol. Extracting and composing robust features with denoising autoencoders. In ICML, pages 1096–1103, 2008.
  29. 29.J. Wright, Y. Ma, J. Mairal, G. Sapiro, T. S. Huang, and S. Yan. Sparse representation for computer vision and pattern recognition. Proceedings of the IEEE, 98(6):1031–1044, 2010.
  30. 30.Y. Zhang, F. Yu, S. Song, P. Xu, A. Seff, and J. Xiao. Large-scale scene understanding challenge. In CVPR Workshop, 2015.
  31. 31.S. C. Zhu, Y. Wu, and D. Mumford. Filters, random fields and maximum entropy (frame): Towards a unified theory for texture modeling. International Journal of Computer Vision, 27(2):107–126, 1998.
  32. 32.D. Zoran and Y. Weiss. From learning models of natural image patches to whole image restoration. In ICCV, 2011.

Citation

MLA
Denton, E., et al. “Deep Generative Image Models Using a Laplacian Pyramid of Adversarial Networks”. arXiv, 2015, http://arxiv.org/abs/1506.05751v1.
APA
Denton, E., Chintala, S., Szlam, A., & Fergus, R. (2015). Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks. arXiv. http://arxiv.org/abs/1506.05751v1
Chicago
Denton, E., S. Chintala, A. Szlam, and R. Fergus. 2015. “Deep Generative Image Models Using a Laplacian Pyramid of Adversarial Networks”. arXiv. http://arxiv.org/abs/1506.05751v1.
Harvard
Denton, E. et al. (2015) “Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1506.05751v1.
Vancouver
1. Denton E, Chintala S, Szlam A, Fergus R (2015) Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks. arXiv

BibTeX

@article{denton2015deep,
  title = {Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks},
  author = {Denton, Emily and Chintala, Soumith and Szlam, Arthur and Fergus, Rob},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1506.05751v1},
  eprint = {1506.05751}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors