Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models
Matthew J. MuckleyAlaaeldin El-NoubyKaren UllrichHervé JégouJakob Verbeek
Introduces a non-binary adversarial discriminator conditioned on quantized local VQ-VAE representations that achieves state-of-the-art statistical fidelity and distortion trade-offs in neural image compression, matching HiFiC's FID using 30–40% fewer bits.
Digital image compression plays a vital role in managing data storage and network transmission demands. Conventional compression techniques and early deep learning systems optimize primarily for traditional mathematical distortion metrics, such as Peak Signal-to-Noise Ratio (PSNR). However, theoretical and empirical evidence shows that strictly minimizing distortion causes compressed images—especially at low bitrates—to become overly blurry and lose their natural statistical properties. While recent methods have incorporated generative adversarial networks to restore realistic detail, these previous frameworks adopted image-wide discriminator models from general image synthesis rather than tailoring them to the local patch-level detail synthesis required by compression.
The article aims to evaluate a new neural image compression framework called Mean-Scale Implicit Local Likelihood Model (MS-ILLM), demonstrating how conditioning adversarial training on quantized local image representations can significantly improve the realism of compressed images without degrading standard distortion metrics.
To achieve this, the authors designed a vector-quantized autoencoder that segments image features into discrete local neighborhood codes, pairing it with a specialized convolutional discriminator. Rather than classifying an entire image as simply real or fake, the discriminator predicts these local semantic labels while identifying compression artifacts. The system was trained on the large-scale OpenImages dataset and evaluated against traditional standard codecs as well as state-of-the-art neural compression models across benchmark datasets including CLIC2020, DIV2K, and Kodak across multiple compression bitrates.
The evaluation produced several key findings. First, the MS-ILLM framework achieved the same statistical realism—measured via Fréchet Inception Distance (FID)—as the prior state-of-the-art neural compressor while using 30% to 40% fewer bits on the CLIC2020 benchmark. Second, when compared at equal bitrate and traditional distortion levels, MS-ILLM consistently produced superior statistical fidelity and fewer unnatural visual artifacts than competing generative methods. Third, architectural ablations revealed that omitting standard normalization layers in the discriminator yielded the best performance and faster convergence, while varying codebook sizes had minimal sensitivity.
These results demonstrate that enforcing local statistical alignment in compression models resolves the trade-off between image sharpness and artifact generation much more effectively than prior global discriminators. For organizations handling large volumes of visual data, this offers substantial bandwidth and storage cost reductions without compromising visual texture quality. However, like other generative compression models, the system still degrades fine structured details such as small text compared to non-adversarial baselines.
Organizations evaluating neural compression should consider piloting local likelihood-based architectures where bandwidth optimization and visual texture are paramount. Before broad production deployment, teams must conduct comprehensive human perceptual studies and evaluate computational overhead. Further engineering work is recommended to quantize and miniaturize the neural network models to make them practical for real-time edge devices.
The findings are supported with high confidence through consistent results across multiple standard datasets and alternative feature-extraction metrics. Nonetheless, limitations remain: the deep learning models require substantial computing resources, and adversarial synthesis carries a potential risk of introducing demographic or visual biases, requiring careful auditing prior to commercial deployment.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This foundational end-to-end learned codec establishes the rate–distortion optimization framework that the source extends with a statistical-fidelity objective.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). Its learned hyperprior explains how neural codecs model local latent statistics, preparing you for the source’s conditioning on quantized local representations.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This codec combines hierarchical and autoregressive probability models, clarifying the learned likelihood machinery that the source seeks to complement.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). Its discretized Gaussian-mixture likelihoods provide useful context for the probability modeling in learned codecs that the source augments with a non-binary discriminator.
- Paper: Generating Diverse High-Fidelity Images with VQ-VAE-2, Ali Razavi et al. (2019). Its hierarchical VQ-VAE introduces discrete image codes, making the source’s use of quantized local representations easier to follow.
No sufficiently relevant recommendations were found.
