Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules
Zhengxue ChengHeming SunMasaru TakeuchiJiro Katto
Introduces a neural image compression model using discretized Gaussian mixture likelihoods and attention mechanisms to achieve rate-distortion performance that matches the Versatile Video Coding (VVC) standard in PSNR.
Efficient image compression is vital for modern digital storage and transmission, especially with the proliferation of high-resolution visual media. While deep learning methods have emerged as promising alternatives to traditional handcrafted formats, they have historically lagged behind the latest industry standards in objective image fidelity, largely due to inaccurate rate estimation in their underlying probability models.
The article demonstrates an advanced learned image compression framework that combines discretized Gaussian mixture likelihoods with a streamlined attention mechanism. The primary objective is to eliminate residual spatial redundancy in compressed representations and evaluate whether this design can match or exceed prevailing industry benchmarks in compression efficiency and reconstructed visual quality.
The researchers developed an autoencoder network utilizing residual blocks, sub-pixel convolution upsampling, and simplified attention modules that focus computational capacity on complex image regions. To accurately model compressed data, the system parameterizes latent representations using a combination of multiple Gaussian distributions rather than a single fixed distribution. The model was trained on 13,830 image patches from the ImageNet database and evaluated across standard benchmarks, including the 24-image Kodak dataset and the 41 high-resolution images of the CLIC professional validation set, against traditional codecs (JPEG, JPEG2000, HEVC, and the VVC test model VTM 5.2) as well as existing learned compression algorithms.
The evaluation yielded several key findings. First, the proposed method achieves state-of-the-art compression efficiency among learned codecs, outperforming prior deep learning approaches and established standards like HEVC, JPEG2000, and JPEG. Second, it is the first learned method to achieve objective reconstruction quality, measured by peak signal-to-noise ratio, comparable to the advanced VVC standard on the Kodak benchmark. Third, when optimized for structural similarity metrics, the model visibly outperforms classical standards by avoiding severe blocking artifacts and preserving fine structural textures at high compression ratios around 240:1. Finally, ablation experiments demonstrated that using three Gaussian components balances modeling performance and complexity, while the modified backbone reduces required bitrates by roughly 6% compared to standard architectures.
These findings show that flexible probability models significantly narrow the performance gap between artificial intelligence-driven compression and traditional handcrafted standards. For organizations handling large-scale image workflows, learned compression offers a viable path toward higher visual fidelity at reduced transmission bandwidth and storage footprint, while maintaining practical training runtimes via simplified attention blocks.
Stakeholders exploring modern compression infrastructure should consider piloting learned codecs in bandwidth-constrained visual services, particularly where perceptual quality is the priority. Future engineering work should focus on training networks on larger image patch sizes to improve performance on ultra-high-resolution content, as well as optimizing runtime decoding complexity for edge deployment.
Confidence in these findings is supported by extensive comparative benchmarks and consistent ablation metrics. However, decision-makers should note that the system slightly trailed VVC in objective signal-to-noise ratio on very high-resolution images, likely due to patch-size constraints during training, and requires input dimension padding to multiples of 64 during deployment.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This foundational paper establishes the joint autoregressive context model and hyperprior architecture that the source directly builds upon by upgrading the single Gaussian assumption to discretized Gaussian mixtures.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). This work introduces the variational scale hyperprior entropy model for learned image compression, providing the essential baseline framework extended by the source.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This seminal paper formulates end-to-end optimized rate-distortion autoencoders using generalized divisive normalization and uniform noise quantization, underpinning the modern neural image compression paradigm.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Latent Diffusion Models build upon high-fidelity learned latent autoencoders with perceptual and attention mechanisms to perform scalable generative image synthesis.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). This work pairs discrete latent autoencoding representations with autoregressive transformer modeling for high-resolution visual generation.
- Paper: Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images, Rewon Child (2021). Very Deep VAEs advance deep hierarchical latent density estimation beyond standard autoregressive priors for generative image modeling.
- Paper: NVAE: A Deep Hierarchical Variational Autoencoder, Arash Vahdat et al. (2020). NVAE designs advanced deep hierarchical architectures and residual distributions to optimize variational latent representation quality and likelihood estimation.
