keyword
hierarchical priors
Hierarchical priors are multi-level probabilistic models used in learned data compression and generative modeling to estimate the probability distributions of primary latent representations by conditioning them on higher-level auxiliary latent variables. Instead of treating latent features as fully independent, a hierarchical prior extracts secondary latent features, commonly referred to as a hyperprior or side information, to capture complex statistical dependencies and redundancies such as localized scale and variance across the primary representations. This auxiliary information is compressed and transmitted to the decoder, where it is decoded to dynamically predict the parameters of the entropy model used for arithmetic decoding of the primary latents. By structuring the prior distribution into layered representations, hierarchical priors significantly improve entropy estimation accuracy and rate-distortion efficiency within end-to-end trainable compression architectures.
2 items

Efficient Hierarchical Entropy Model for Learned Point Cloud Compression
Rui Song, Chunyang Fu, Shan Liu, Ge Li
Why you should read this
Proposes an efficient octree-based entropy model using hierarchical attention and grouped contexts to achieve linear computational complexity and fast parallel decoding without sacrificing point cloud compression performance.
Learning an accurate entropy model is a fundamental way to remove the redundancy in point cloud compression. Recently, the octree-based auto-regressive entropy model which adopts the self-attention mechanism to explore dependencies in a large-scale context is proved to be promising. However, heavy global attention computations and auto-regressive contexts are inefficient for practical applications. To improve the efficiency of the attention model, we propose a hierarchical attention structure that has a linear complexity to the context scale and maintains the global receptive field. Furthermore, we present a grouped context structure to address the serial decoding issue caused by the auto-regression while preserving the compression performance. Experiments demonstrate that the proposed entropy model achieves superior rate-distortion performance and significant decoding latency reduction compared with the state-of-the-art large-scale auto-regressive entropy model.
Added
2026-09-26

Joint Autoregressive and Hierarchical Priors for Learned Image Compression
David Minnen, Johannes Ballé, George Toderici
Why you should read this
Presents a learned image compression architecture that couples autoregressive and hierarchical priors in the entropy model, establishing the first deep learning approach to outperform traditional BPG codecs across both PSNR and MS-SSIM rate-distortion metrics.
Recent models for learned image compression are based on autoencoders, learning approximately invertible mappings from pixels to a quantized latent representation. These are combined with an entropy model, a prior on the latent representation that can be used with standard arithmetic coding algorithms to yield a compressed bitstream. Recently, hierarchical entropy models have been introduced as a way to exploit more structure in the latents than simple fully factorized priors, improving compression performance while maintaining end-to-end optimization. Inspired by the success of autoregressive priors in probabilistic generative models, we examine autoregressive, hierarchical, as well as combined priors as alternatives, weighing their costs and benefits in the context of image compression. While it is well known that autoregressive models come with a significant computational penalty, we find that in terms of compression performance, autoregressive and hierarchical priors are complementary and, together, exploit the probabilistic structure in the latents better than all previous learned models. The combined model yields state-of-the-art rate--distortion performance, providing a 15.8% average reduction in file size over the previous state-of-the-art method based on deep learning, which corresponds to a 59.8% size reduction over JPEG, more than 35% reduction compared to WebP and JPEG2000, and bitstreams 8.4% smaller than BPG, the current state-of-the-art image codec. To the best of our knowledge, our model is the first learning-based method to outperform BPG on both PSNR and MS-SSIM distortion metrics.
Added
2026-09-24
