Built independently by an author, for readers. Read the story and support ChapterPal

keyword

learned image compression

Learned image compression is an image coding paradigm that uses machine learning models, primarily deep neural networks, to compress digital images for efficient storage and transmission. Unlike conventional compression algorithms that rely on hand-crafted mathematical transforms and fixed heuristics, learned image compression typically uses end-to-end trainable architectures, such as non-linear autoencoders. In this framework, an encoder maps input pixels into a compact latent representation, which is discretized through quantization and converted into a binary bitstream using a learned entropy model that estimates latent probability distributions. A decoder then reconstructs the original image from the decoded representation, with the entire network optimized through rate-distortion objectives to systematically balance compressed file size against reconstructed visual fidelity.

4 items

Frequency-Aware Transformer for Learned Image Compression

Frequency-Aware Transformer for Learned Image Compression

Han Li, Shaohui Li, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong

OrganizationsShanghai Jiao Tong UniversityTsinghua University

Why you should read this

Introduces a frequency-aware transformer architecture for learned image compression that captures directional image details through multiscale frequency decomposition, outperforming the VTM-12.1 standard codec by more than 13% in BD-rate across benchmark datasets.

Learned image compression (LIC) has gained traction as an effective solution for image storage and transmission in recent years. However, existing LIC methods are redundant in latent representation due to limitations in capturing anisotropic frequency components and preserving directional details. To overcome these challenges, we propose a novel frequency-aware transformer (FAT) block that for the first time achieves multiscale directional ananlysis for LIC. The FAT block comprises frequency-decomposition window attention (FDWA) modules to capture multiscale and directional frequency components of natural images. Additionally, we introduce frequency-modulation feed-forward network (FMFFN) to adaptively modulate different frequency components, improving rate-distortion performance. Furthermore, we present a transformer-based channel-wise autoregressive (T-CA) model that effectively exploits channel dependencies. Experiments show that our method achieves state-of-the-art rate-distortion performance compared to existing LIC methods, and evidently outperforms latest standardized codec VTM-12.1 by 14.5%, 15.1%, 13.0% in BD-rate on the Kodak, Tecnick, and CLIC datasets.

Added

2026-09-26

ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding

ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding

Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, Yan Wang

OrganizationsSenseTimeTsinghua University

Why you should read this

Presents ELIC, a learned image compression architecture that combines uneven space-channel contextual coding with efficient transform design to achieve state-of-the-art rate-distortion performance alongside fast inference, preview decoding, and progressive decoding.

Recently, learned image compression techniques have achieved remarkable performance, even surpassing the best manually designed lossy image coders. They are promising to be large-scale adopted. For the sake of practicality, a thorough investigation of the architecture design of learned image compression, regarding both compression performance and running speed, is essential. In this paper, we first propose uneven channel-conditional adaptive coding, motivated by the observation of energy compaction in learned image compression. Combining the proposed uneven grouping model with existing context models, we obtain a spatial-channel contextual adaptive model to improve the coding performance without damage to running speed. Then we study the structure of the main transform and propose an efficient model, ELIC, to achieve state-of-the-art speed and compression ability. With superior performance, the proposed model also supports extremely fast preview decoding and progressive decoding, which makes the coming application of learning-based image compression more promising.

Added

2026-09-26

Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules

Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto

OrganizationsJapan Science and Technology AgencyWaseda University

Why you should read this

Introduces a neural image compression model using discretized Gaussian mixture likelihoods and attention mechanisms to achieve rate-distortion performance that matches the Versatile Video Coding (VVC) standard in PSNR.

Image compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. This project page is at this https URL this https URL

Added

2026-09-25

Joint Autoregressive and Hierarchical Priors for Learned Image Compression

Joint Autoregressive and Hierarchical Priors for Learned Image Compression

David Minnen, Johannes Ballé, George Toderici

OrganizationsGoogle

Why you should read this

Presents a learned image compression architecture that couples autoregressive and hierarchical priors in the entropy model, establishing the first deep learning approach to outperform traditional BPG codecs across both PSNR and MS-SSIM rate-distortion metrics.

Recent models for learned image compression are based on autoencoders, learning approximately invertible mappings from pixels to a quantized latent representation. These are combined with an entropy model, a prior on the latent representation that can be used with standard arithmetic coding algorithms to yield a compressed bitstream. Recently, hierarchical entropy models have been introduced as a way to exploit more structure in the latents than simple fully factorized priors, improving compression performance while maintaining end-to-end optimization. Inspired by the success of autoregressive priors in probabilistic generative models, we examine autoregressive, hierarchical, as well as combined priors as alternatives, weighing their costs and benefits in the context of image compression. While it is well known that autoregressive models come with a significant computational penalty, we find that in terms of compression performance, autoregressive and hierarchical priors are complementary and, together, exploit the probabilistic structure in the latents better than all previous learned models. The combined model yields state-of-the-art rate--distortion performance, providing a 15.8% average reduction in file size over the previous state-of-the-art method based on deep learning, which corresponds to a 59.8% size reduction over JPEG, more than 35% reduction compared to WebP and JPEG2000, and bitstreams 8.4% smaller than BPG, the current state-of-the-art image codec. To the best of our knowledge, our model is the first learning-based method to outperform BPG on both PSNR and MS-SSIM distortion metrics.

Added

2026-09-24