topic
video coding
Video coding is the process of compressing and decompressing digital video data to enable efficient transmission, broadcasting, and storage across digital networks and multimedia systems. It reduces data volume while preserving perceptual visual quality by identifying and removing redundancies, including repetitive spatial details within individual frames and temporal similarities between consecutive frames. Standard video coding architectures employ a sequence of operations that typically includes intra-frame and inter-frame prediction, motion estimation and compensation, transform coding, quantization, and entropy coding. Widely deployed through international standards such as Advanced Video Coding and High Efficiency Video Coding, video coding balances compression efficiency, computational complexity, and error resilience to support applications ranging from low-latency teleconferencing to high-resolution on-demand streaming.
2 items

Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode Prediction
Zhihao Hu, Guo Lu, Jinyang Guo, Shan Liu, Wei Jiang, Dong Xu
Why you should read this
Presents a coarse-to-fine deep video compression framework that uses two-stage motion compensation alongside hyperprior-guided mode prediction to dynamically select block resolutions and skip residual coding without transmitting extra side information.
The previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep video compression framework for better motion compensation, in which we perform motion estimation, compression and compensation twice in a coarse to fine manner. Our C2F framework can achieve better motion compensation results without significantly increasing bit costs. Observing hyperprior information (i.e., the mean and variance values) from the hyperprior networks contains discriminant statistical information of different patches, we also propose two efficient hyperprior-guided mode prediction methods. Specifically, using hyperprior information as the input, we propose two mode prediction networks to respectively predict the optimal block resolutions for better motion coding and decide whether to skip residual information from each block for better residual coding without introducing additional bit cost while bringing negligible extra computation cost. Comprehensive experimental results demonstrate our proposed C2F video compression framework equipped with the new hyperprior-guided mode prediction methods achieves the state-of-the-art performance on HEVC, UVG and MCL-JCV datasets.
Added
2026-09-26

Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules
Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto
Why you should read this
Introduces a neural image compression model using discretized Gaussian mixture likelihoods and attention mechanisms to achieve rate-distortion performance that matches the Versatile Video Coding (VVC) standard in PSNR.
Image compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. This project page is at this https URL this https URL
Added
2026-09-25
