Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Zhengxue ChengHeming SunMasaru TakeuchiJiro Katto

article2020CVPR1,328 citations

Introduces a neural image compression model using discretized Gaussian mixture likelihoods and attention mechanisms to achieve rate-distortion performance that matches the Versatile Video Coding (VVC) standard in PSNR.

Listen

Efficient image compression is vital for modern digital storage and transmission, especially with the proliferation of high-resolution visual media. While deep learning methods have emerged as promising alternatives to traditional handcrafted formats, they have historically lagged behind the latest industry standards in objective image fidelity, largely due to inaccurate rate estimation in their underlying probability models.

The article demonstrates an advanced learned image compression framework that combines discretized Gaussian mixture likelihoods with a streamlined attention mechanism. The primary objective is to eliminate residual spatial redundancy in compressed representations and evaluate whether this design can match or exceed prevailing industry benchmarks in compression efficiency and reconstructed visual quality.

The researchers developed an autoencoder network utilizing residual blocks, sub-pixel convolution upsampling, and simplified attention modules that focus computational capacity on complex image regions. To accurately model compressed data, the system parameterizes latent representations using a combination of multiple Gaussian distributions rather than a single fixed distribution. The model was trained on 13,830 image patches from the ImageNet database and evaluated across standard benchmarks, including the 24-image Kodak dataset and the 41 high-resolution images of the CLIC professional validation set, against traditional codecs (JPEG, JPEG2000, HEVC, and the VVC test model VTM 5.2) as well as existing learned compression algorithms.

The evaluation yielded several key findings. First, the proposed method achieves state-of-the-art compression efficiency among learned codecs, outperforming prior deep learning approaches and established standards like HEVC, JPEG2000, and JPEG. Second, it is the first learned method to achieve objective reconstruction quality, measured by peak signal-to-noise ratio, comparable to the advanced VVC standard on the Kodak benchmark. Third, when optimized for structural similarity metrics, the model visibly outperforms classical standards by avoiding severe blocking artifacts and preserving fine structural textures at high compression ratios around 240:1. Finally, ablation experiments demonstrated that using three Gaussian components balances modeling performance and complexity, while the modified backbone reduces required bitrates by roughly 6% compared to standard architectures.

These findings show that flexible probability models significantly narrow the performance gap between artificial intelligence-driven compression and traditional handcrafted standards. For organizations handling large-scale image workflows, learned compression offers a viable path toward higher visual fidelity at reduced transmission bandwidth and storage footprint, while maintaining practical training runtimes via simplified attention blocks.

Stakeholders exploring modern compression infrastructure should consider piloting learned codecs in bandwidth-constrained visual services, particularly where perceptual quality is the priority. Future engineering work should focus on training networks on larger image patch sizes to improve performance on ultra-high-resolution content, as well as optimizing runtime decoding complexity for edge deployment.

Confidence in these findings is supported by extensive comparative benchmarks and consistent ablation metrics. However, decision-makers should note that the system slightly trailed VVC in objective signal-to-noise ratio on very high-resolution images, likely due to patch-size constraints during training, and requires input dimension padding to multiples of 64 during deployment.

  • Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This foundational paper establishes the joint autoregressive context model and hyperprior architecture that the source directly builds upon by upgrading the single Gaussian assumption to discretized Gaussian mixtures.
  • Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). This work introduces the variational scale hyperprior entropy model for learned image compression, providing the essential baseline framework extended by the source.
  • Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This seminal paper formulates end-to-end optimized rate-distortion autoencoders using generalized divisive normalization and uniform noise quantization, underpinning the modern neural image compression paradigm.
Cover for Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Abstract

Image compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. This project page is at this https URL this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Method
  • 3.1 Formulation of Learned Compression Models
  • 3.2 Discretized Gaussian Mixture Likelihoods
  • 3.3 Network Architecture
  • 4 Implementation Details
  • 5 Experiments
  • 5.1 Ablation Study
  • 5.2 Rate-distortion Performance
  • 5.3 Qualitative Results
  • 6 Conclusion
  • References
  • 7 Appendix
  • 7.1 Ablation Study on Network Architecture
  • 7.1.1 Backbone
  • 7.1.2 The Number of Mixtures KK
  • 7.2 Test Settings on Codecs
  • 7.2.1 Versatile Video Coding (VVC)
  • 7.2.2 Boundary Handling

Knowls

  1. Knowl 1 — Discretized Gaussian Mixture Likelihood Entropy Model

    model/method

    In learned transform image compression, the quantized latent representation y^\hat{\mathbf{y}} is modeled conditionally on hyperprior side information z^\hat{\mathbf{z}} and autoregressive context using a discretized Gaussian Mixture Model (GMM). For each spatial element y^i\hat{y}_i (where ii indexes the feature map location), the conditional probability distribution is parameterized as a mixture of KK Gaussian distributions:

    py^∣z^(y^∣z^)=∏ipy^i∣z^(y^i∣z^)p_{\hat{\mathbf{y}}|\hat{\mathbf{z}}}(\hat{\mathbf{y}}|\hat{\mathbf{z}}) = \prod_i p_{\hat{y}_i|\hat{\mathbf{z}}}(\hat{y}_i|\hat{\mathbf{z}})

    py^i∣z^(y^i∣z^)=(∑k=1Kwi(k)N(μi(k),σi2(k))∗U(−12,12))(y^i)=c(y^i+12)−c(y^i−12)p_{\hat{y}_i|\hat{\mathbf{z}}}(\hat{y}_i|\hat{\mathbf{z}}) = \left( \sum_{k=1}^K w_i^{(k)} \mathcal{N}\left(\mu_i^{(k)}, \sigma_i^{2(k)}\right) * \mathcal{U}\left(-\frac{1}{2}, \frac{1}{2}\right) \right)(\hat{y}_i) = c\left(\hat{y}_i + \frac{1}{2}\right) - c\left(\hat{y}_i - \frac{1}{2}\right)

    where k∈{1,…,K}k \in \{1, \dots, K\} indexes the mixture components, wi(k)w_i^{(k)} is the mixture weight satisfying ∑k=1Kwi(k)=1\sum_{k=1}^K w_i^{(k)} = 1 with wi(k)≥0w_i^{(k)} \ge 0, μi(k)\mu_i^{(k)} is the component mean, σi2(k)\sigma_i^{2(k)} is the component variance at location ii, and c(⋅)c(\cdot) is the cumulative distribution function of the continuous Gaussian mixture:

    c(v)=∑k=1Kwi(k)12[1+erf⁡(v−μi(k)2σi2(k))]c(v) = \sum_{k=1}^K w_i^{(k)} \frac{1}{2} \left[ 1 + \operatorname{erf}\left( \frac{v - \mu_i^{(k)}}{\sqrt{2 \sigma_i^{2(k)}}} \right) \right]

    To ensure numerical stability during training, the range of y^i\hat{y}_i is clamped to [−255,256][-255, 256]. At the boundaries, c(y^i−12)c\left(\hat{y}_i - \frac{1}{2}\right) is set to 00 when y^i=−255\hat{y}_i = -255 (representing c(−∞)=0c(-\infty) = 0), and c(y^i+12)c\left(\hat{y}_i + \frac{1}{2}\right) is set to 11 when y^i=256\hat{y}_i = 256 (representing c(+∞)=1c(+\infty) = 1). For a latent tensor with NN feature channels, the parameter prediction network outputs 3×N×K3 \times N \times K parameters per spatial location to specify weights, means, and variances.

  2. Knowl 2 — Rate-Distortion Optimization Objective for Learned Transform Coding

    equation

    End-to-end learned lossy image compression optimizes a Lagrangian rate-distortion loss balancing the transmission rate of quantized latents y^\hat{\mathbf{y}} and hyperprior side information z^\hat{\mathbf{z}} against reconstruction distortion:

    L=R(y^)+R(z^)+λ⋅D(x,x^)\mathcal{L} = R(\hat{\mathbf{y}}) + R(\hat{\mathbf{z}}) + \lambda \cdot D(\mathbf{x}, \hat{\mathbf{x}})

    L=Ex∼px[−log⁡2py^∣z^(y^∣z^)]+Ex∼px[−log⁡2pz^∣ψ(z^∣ψ)]+λ⋅D(x,x^)\mathcal{L} = \mathbb{E}_{\mathbf{x} \sim p_{\mathbf{x}}}\left[ -\log_2 p_{\hat{\mathbf{y}}|\hat{\mathbf{z}}}(\hat{\mathbf{y}}|\hat{\mathbf{z}}) \right] + \mathbb{E}_{\mathbf{x} \sim p_{\mathbf{x}}}\left[ -\log_2 p_{\hat{\mathbf{z}}|\boldsymbol{\psi}}(\hat{\mathbf{z}}|\boldsymbol{\psi}) \right] + \lambda \cdot D(\mathbf{x}, \hat{\mathbf{x}})

    where:

    • x\mathbf{x} is the uncompressed input image, and x^=gs(y^;θ)\hat{\mathbf{x}} = g_s(\hat{\mathbf{y}}; \boldsymbol{\theta}) is the reconstructed image generated by the synthesis transform gsg_s parameterized by θ\boldsymbol{\theta}.
    • y=ga(x;ϕ)\mathbf{y} = g_a(\mathbf{x}; \boldsymbol{\phi}) is the latent representation output by the analysis transform gag_a parameterized by ϕ\boldsymbol{\phi}.
    • z=ha(y;ϕh)\mathbf{z} = h_a(\mathbf{y}; \boldsymbol{\phi}_h) is the hyperprior side information produced by the hyper-analysis transform hah_a parameterized by ϕh\boldsymbol{\phi}_h.
    • y^\hat{\mathbf{y}} and z^\hat{\mathbf{z}} are the quantized representations. During training, non-differentiable quantization is approximated by adding uniform noise y~=y+n\tilde{\mathbf{y}} = \mathbf{y} + \mathbf{n} with n∼U(−12,12)\mathbf{n} \sim \mathcal{U}\left(-\frac{1}{2}, \frac{1}{2}\right) (and similarly for z~\tilde{\mathbf{z}}); during inference, standard scalar quantization y^=round⁡(y)\hat{\mathbf{y}} = \operatorname{round}(\mathbf{y}) is applied.
    • py^∣z^(y^∣z^)p_{\hat{\mathbf{y}}|\hat{\mathbf{z}}}(\hat{\mathbf{y}}|\hat{\mathbf{z}}) is the conditional entropy model for the main latent codes.
    • pz^∣ψ(z^∣ψ)=∏i(pzi∣ψ(ψ)∗U(−12,12))(z^i)p_{\hat{\mathbf{z}}|\boldsymbol{\psi}}(\hat{\mathbf{z}}|\boldsymbol{\psi}) = \prod_i \left( p_{z_i|\boldsymbol{\psi}}(\boldsymbol{\psi}) * \mathcal{U}\left(-\frac{1}{2}, \frac{1}{2}\right) \right)(\hat{z}_i) is a non-adaptive factorized prior parameterized by learnable scale parameters ψ\boldsymbol{\psi}.
    • D(x,x^)D(\mathbf{x}, \hat{\mathbf{x}}) denotes the distortion metric, implemented as Mean Squared Error (MSE) or multiscale structural similarity distortion D(x,x^)=1−MS-SSIM⁡(x,x^)D(\mathbf{x}, \hat{\mathbf{x}}) = 1 - \operatorname{MS-SSIM}(\mathbf{x}, \hat{\mathbf{x}}).
    • λ>0\lambda > 0 is the Lagrange multiplier governing the rate-distortion tradeoff.
  3. Knowl 3 — Simplified Attention Module for Learned Compression

    model/method

    To enhance network representation capacity for complex image regions while keeping training computation tractable, attention mechanisms can be integrated into the compression autoencoder. Standard non-local attention modules combine residual blocks with Non-Local Blocks (NLB) that compute pairwise spatial affinities; however, NLBs incur substantial memory and training time overhead.

    Because deep residual architectures already establish large effective receptive fields, the NLB is removed to create a simplified attention module comprising two parallel paths:

    1. A trunk branch composed of stacked residual blocks.
    2. A mask branch comprising residual blocks followed by a 1×11\times 1 convolution and a sigmoid activation function to generate spatial-channel gating weights in [0,1][0, 1].

    The outputs of the trunk and mask branches undergo elementwise multiplication, and the resulting feature map is added to the input tensor via a global residual connection.

    Attention Configuration Training Loss (after 16 epochs) Training Time (s/epoch)
    Full Attention (with Non-Local Block) 2.705 1119
    Simplified Attention (without Non-Local Block) 2.754 336
    Without Attention Module 3.026 216

    Omitting the NLB reduces training time from 11191119 to 336336 seconds per epoch while maintaining most of the loss reduction over the baseline without attention.

  4. Knowl 4 — Autoencoder Backbone with Residual Convolutions and Subpixel Upsampling

    model/method

    The autoencoder network architecture incorporates residual convolutions for transforms and subpixel convolutions for reconstruction:

    1. Analysis Transform (gag_a): Instead of single large-stride 5×55\times 5 or 9×99\times 9 convolutional layers, each downsampling stage utilizes four stacked 3×33\times 3 convolutions with a residual skip connection. This design provides an enlarged receptive field with fewer parameters and higher coding efficiency, achieving an approximate 6%6\% bitrate reduction over 5×55\times 5 convolution baselines at equal distortion.
    2. Synthesis Transform (gsg_s): Upsampling replaces transposed convolutions (deconvolutions) with subpixel convolutions (pixel shuffle operations) to better preserve fine spatial textures and avoid checkerboard artifacts.
    3. Joint Autoregressive and Hyperprior Entropy Architecture: A 5×55\times 5 masked convolution captures causal local spatial context within y^\hat{\mathbf{y}}, which is combined with hyperprior features decoded by hs(z^)h_s(\hat{\mathbf{z}}) and processed through 1×11\times 1 convolutions to predict the Gaussian mixture parameters (3×N×K3 \times N \times K parameters per location).
  5. Knowl 5 — Effect of the Number of Gaussian Mixture Components K

    empirical result

    Ablation evaluations comparing Gaussian mixture models with component counts K∈{2,3,4,5}K \in \{2, 3, 4, 5\} against a single Gaussian model (K=1K=1) for conditional entropy estimation show:

    • Discretized Gaussian mixture distributions (K≥2K \ge 2) systematically yield lower validation loss and improved rate-distortion curves compared to unimodal Gaussian priors (K=1K=1). The multi-modal formulation flexibly fits asymmetric, multi-peaked distributions occurring at high-entropy edges and complex textures.
    • Rate-distortion performance saturates when K≥3K \ge 3; moving from K=3K=3 to K=4K=4 or K=5K=5 produces minimal additional coding gain, identifying K=3K=3 as the optimal trade-off between coding efficiency and model complexity.
    • In smooth image regions (such as uniform sky), the mixture model collapses to a unimodal profile where the estimated means μi(k)\mu_i^{(k)} of all components converge to the same value.
  6. Knowl 6 — Training Setup and Operational Configurations

    experimental setup

    The training and evaluation pipeline uses the following specifications:

    • Dataset: 13,830 patches of size 256×256256 \times 256 cropped from the ImageNet database.
    • Optimization: Adam optimizer with batch size 88. Initial learning rate is 1×10−41 \times 10^{-4}, reduced to 1×10−51 \times 10^{-5} for the final 80,00080{,}000 iterations. Total training duration is 1×1061 \times 10^6 iterations per rate-distortion point λ\lambda.
    • Model Capacities and Tradeoff Weights:
      • For MSE distortion optimization, λ∈{0.0016,0.0032,0.0075,0.015,0.03,0.045}\lambda \in \{0.0016, 0.0032, 0.0075, 0.015, 0.03, 0.045\}. Main channel width is set to N=128N=128 for the three lowest bitrates and N=192N=192 for the three highest bitrates.
      • For MS-SSIM distortion optimization (D=1−MS-SSIM⁡D = 1 - \operatorname{MS-SSIM}), λ∈{3,12,40,120}\lambda \in \{3, 12, 40, 120\}. Main channel width is set to N=128N=128 for the two lower bitrates and N=192N=192 for the two higher bitrates.
    • Boundary Handling for Arbitrary Sizes: Input images are reflect-padded so that height and width are exact multiples of 64 (the minimum internal spatial resolution H64×W64\frac{H}{64} \times \frac{W}{64}). Original dimensions are encoded into the bitstream header, and decoded outputs are cropped back to original dimensions.
  7. Knowl 7 — Rate-Distortion Performance on Kodak and CLIC Benchmarks

    empirical result

    The learned compression framework using discretized Gaussian mixture likelihoods and simplified attention was evaluated against standard traditional codecs and prior learned codecs on the Kodak dataset (24 images, 768×512768 \times 512) and the CVPR CLIC Professional validation dataset (41 high-resolution images):

    • PSNR Performance on Kodak: When optimized for MSE, the model matches the coding efficiency of Versatile Video Coding (VVC test model VTM 5.2 intra profile with YUV444 format), achieving comparable PSNR across bitrates while outperforming HEVC/BPG, JPEG2000 (OpenJPEG), JPEG (PIL), and previous learning-based codecs (Minnen et al. 2018, Lee et al. 2019, Ballé et al. 2018).
    • PSNR Performance on CLIC: The model outperforms JPEG, JPEG2000, HEVC, and prior learned codecs across bitrates, performing close to VVC.
    • MS-SSIM Performance: When optimized for multiscale structural similarity (evaluated in dB=−10log⁡10(1−MS-SSIM⁡)\text{dB} = -10 \log_{10}(1 - \operatorname{MS-SSIM})), the method outperforms all tested classical standards (VVC, HEVC, JPEG2000, JPEG) and existing learned codecs on both Kodak and CLIC datasets. At low bitrates (sim0.10\\sim 0.10 bpp, compression ratios ≥200:1\ge 200:1), the MS-SSIM model preserves fine details (e.g., hair strands, brick patterns, flower petals) without the severe blocking or blur artifacts present in traditional codecs.

Coverage note — Omitted specific third-party command-line invocations for VVC/VTM 5.2 software execution and RGB-to-YUV color conversion definitions as they represent standard external benchmarking protocol rather than new methodological contributions.

References

  1. 1.G. K Wallace, “The JPEG still picture compression standard”, IEEE Trans. on Consumer Electronics, vol. 38, no. 1, pp. 43-59, Feb. 1991.
  2. 2.Majid Rabbani, Rajan Joshi, “An overview of the JPEG2000 still image compression standard”, ELSEVIER Signal Processing: Image Communication, vol. 17, no, 1, pp. 3-48, Jan. 2002.
  3. 3.G. J. Sullivan, J. Ohm, W. Han and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) Standard”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 12, pp. 1649-1668, Dec. 2012.
  4. 4.G. J. Sullivan and J. R. Ohm, “Versatile video coding Towards the next generation of video compression”, Picture Coding Symposium, Jun. 2018.
  5. 5.P. Vincent, H. Larochelle, Y. Bengio and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders”, Intl. conf. on Machine Learning (ICML), pp. 1096-1103, July 5-9. 2008.
  6. 6.Lucas Theis, Wenzhe Shi, Andrew Cunninghan and Ferenc Huszar, “Lossy Image Compression with Compressive Autoencoders”, Intl. Conf. on Learning Representations (ICLR), pp. 1-19, April 24-26, 2017.
  7. 7.J. Balle, Valero Laparra, Eero P. Simoncelli, “End-to-End Optimized Image Compression”, Intl. Conf. on Learning Representations (ICLR), pp. 1-27, April 24-26, 2017.
  8. 8.E. Agustsson, F. Mentzer, M. Tschannen, L. Cavigelli, R. Timofte, L. Benini, L. V. Gool, “Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations”, Neural Information Processing Systems (NIPS) 2017, arXiv:1704.00648v2.
  9. 9.G. Toderici, S. M.O’Malley, S. J. Hwang, et al., “Variable rate image compression with recurrent neural networks”, arXiv: 1511.06085, 2015.
  10. 10.G, Toderici, D. Vincent, N. Johnson, et al., “Full Resolution Image Compression with Recurrent Neural Networks”, IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1-9, July 21-26, 2017.
  11. 11.Nick Johnson, Damien Vincent, David Minnen, et al., “Improved Lossy Image Compression with Priming and Spatially Adaptive Bit Rates for Recurrent Networks”, arXiv:1703.10114, pp. 1-9, March 2017.
  12. 12.Ripple Oren, L. Bourdev, “Real Time Adaptive Image Compression”, Proc. of Machine Learning Research, Vol. 70, pp. 2922-2930, 2017.
  13. 13.S. Santurkar, D. Budden, N. Shavit, “Generative Compression”, Picture Coding Symposium, June 24-27, 2018.
  14. 14.E. Agustsson, M. Tschannen, F. Mentzer, R. Timofte, and L. V. Gool, “Generative Adversarial Networks for Extreme Learned Image Compression”, arXiv:1804.02958.
  15. 15.M. Li, W. Zuo, S. Gu, D. Zhao, D. Zhang, “Learning Convolutional Networks for Content-weighted Image Compression”, IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 17-22, 2018.
  16. 16.Z. Cheng, H. Sun, M. Takeuchi, J. Katto, “Deep Convolutional AutoEncoder-based Lossy Image Compression”, Picture Coding Symposium, pp. 1-5, June 24-27, 2018.
  17. 17.Z. Cheng, H. Sun, M. Takeuchi, J. Katto, “Learning Image and Video Compression through Spatial-Temporal Energy Compaction”, IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 16-20, 2019.
  18. 18.Z. Cheng, H. Sun, M. Takeuchi, J. Katto, “Deep Residual Learning for Image Compression”, CVPR Workshop, pp. 1-4, June 16-20, 2019.
  19. 19.Johannes Balle, D. Minnen, S. Singh, S. J. Hwang, N. Johnston, “Variational Image Compression with a Scale Hyperprior”, Intl. Conf. on Learning Representations (ICLR), pp. 1-23, 2018.
  20. 20.D. Minnen, J. Balle, G. Toderici, “Joint Autoregressive and Hierarchical Priors for Learned Image Compression”, arXiv.1809.02736.
  21. 21.J. Lee, S. Cho, S-K Beack, “Context-Adaptive Entropy Model for End-to-end optimized Image Compression”, Intl. Conf. on Learning Representations (ICLR) 2019.
  22. 22.F. Mentzer, E. Agustsson, M. Tschannen, R. Timofte, L. V. Gool, “Conditional Probability Models for Deep Image Compression”, IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 17-22, 2018.
  23. 23.Z. Cheng, H. Sun, M. Takeuchi, J. Katto, “Performance Comparison of Convolutional AutoEncoders, Generative Adversarial Networks and Super-Resolution for Image Compression”, CVPR Workshop and Challenge on Learned Image Compression (CLIC), pp. 1-4, June 17-22, 2018.
  24. 24.A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espehold, A. Graves, and K. Kavukcuoglu, “Conditional Image Generation with PixelCNN Decoders”, Advances in Neural Information Processing Systems (NIPS), 2016
  25. 25.T. Salimans, A. Karpathy, X. Chen, D. P. Kingma, “PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications”, Intl. Conf. on Learning Representations (ICLR), 2017.
  26. 26.F. Mentzer, E. Agustsson, M. Tschannen, R. Timofte, L. V. Gool, “Practical Full Resolution Learned Lossless Image Compression”, CVPR 2019.
  27. 27.Goyal Vivek K. “Theoretical Foundations of Transform Coding”, IEEE Signal Processing Magazine, Vol. 18, No. 5, Sep. 2001.
  28. 28.Rissanen Jorma and Glen G. Langdon Jr., “Universal modeling and coding”, IEEE Transactions on Information Theory, Vol. 27, No. 1, Jan. 1981.
  29. 29.Y.Zhang, K. Li, K. Li, B. Zhong, Y. Fu, “Residual Nonlocal Attention Networks for Image Restoration”, Intl. Conf. on Learning Representations (ICLR), pp. 1-18, 2019.
  30. 30.H. Liu, T. Chen, P. Guo, Q. Shen, X. Cao, Y. Wang, Z. Ma, “Non-local Attention Optimized Deep Image Compression”, arXiv.1904.09757.
  31. 31.J. Deng, W. Dong, R. Socher, L. Li, K. Li and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database”, IEEE Conf. on Computer Vision and Pattern Recognition, pp. 1-8, June 20-25, 2009.
  32. 32.D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization”, arXiv:1412.6980, pp.1-15, Dec. 2014.
  33. 33.Z. Wang, E. P. Simoncelli and A. C. Bovik, “Multiscale structural similarity for image quality assessment”, The 36th Asilomar Conference on Signals, Systems and Computers, Vol.2, pp. 1398-1402, Nov. 2013.
  34. 34.Kodak Lossless True Color Image Suite, Download from http://r0k.us/graphics/kodak/
  35. 35.Workshop and Challenge on Learned Image Compression, CVPR2018, http://www.compression.cc/challenge/
  36. 36.VVC Official Test Model VTM, https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/tree/VTM-5.2, accessed on July, 2019.
  37. 37.BPG Image Format, https://bellard.org/bpg/
  38. 38.JPEG2000 official software OpenJPEG, https://jpeg.org/jpeg2000/software.html
  39. 39.Python Imaging Library (PIL), https://pillow.readthedocs.io/en/5.1.x/index.html

Citation

MLA
Cheng, Z., et al. “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules”. arXiv, 2020, http://arxiv.org/abs/2001.01568v3.
APA
Cheng, Z., Sun, H., Takeuchi, M., & Katto, J. (2020). Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules. arXiv. http://arxiv.org/abs/2001.01568v3
Chicago
Cheng, Z., H. Sun, M. Takeuchi, and J. Katto. 2020. “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules”. arXiv. http://arxiv.org/abs/2001.01568v3.
Harvard
Cheng, Z. et al. (2020) “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2001.01568v3.
Vancouver
1. Cheng Z, Sun H, Takeuchi M, Katto J (2020) Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules. arXiv

BibTeX

@article{cheng2020learned,
  title = {Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules},
  author = {Cheng, Zhengxue and Sun, Heming and Takeuchi, Masaru and Katto, Jiro},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2001.01568v3},
  eprint = {2001.01568}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE