Built independently by an author, for readers. Read the story and support ChapterPal

keyword

attention modules

Attention modules are specialized neural network components that dynamically assign varying levels of importance to different parts of an input signal, allowing models to focus computational processing on the most relevant information. Integrated into broader deep learning architectures, such as convolutional networks and transformers, these self-contained blocks compute attention weights across spatial dimensions, feature channels, or relational representations between distinct elements. By recalibrating intermediate feature maps according to contextual relevance, attention modules enhance the expressive power of neural networks, capture long-range dependencies, and improve overall feature selectivity without requiring additional supervision.

2 items

Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto

OrganizationsJapan Science and Technology AgencyWaseda University

Why you should read this

Introduces a neural image compression model using discretized Gaussian mixture likelihoods and attention mechanisms to achieve rate-distortion performance that matches the Versatile Video Coding (VVC) standard in PSNR.

Image compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. This project page is at this https URL this https URL

Added

2026-09-25