Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models

Matthew J. MuckleyAlaaeldin El-NoubyKaren UllrichHervé JégouJakob Verbeek

article2023ICML110 citations

Introduces a non-binary adversarial discriminator conditioned on quantized local VQ-VAE representations that achieves state-of-the-art statistical fidelity and distortion trade-offs in neural image compression, matching HiFiC's FID using 30–40% fewer bits.

Listen

Digital image compression plays a vital role in managing data storage and network transmission demands. Conventional compression techniques and early deep learning systems optimize primarily for traditional mathematical distortion metrics, such as Peak Signal-to-Noise Ratio (PSNR). However, theoretical and empirical evidence shows that strictly minimizing distortion causes compressed images—especially at low bitrates—to become overly blurry and lose their natural statistical properties. While recent methods have incorporated generative adversarial networks to restore realistic detail, these previous frameworks adopted image-wide discriminator models from general image synthesis rather than tailoring them to the local patch-level detail synthesis required by compression.

The article aims to evaluate a new neural image compression framework called Mean-Scale Implicit Local Likelihood Model (MS-ILLM), demonstrating how conditioning adversarial training on quantized local image representations can significantly improve the realism of compressed images without degrading standard distortion metrics.

To achieve this, the authors designed a vector-quantized autoencoder that segments image features into discrete local neighborhood codes, pairing it with a specialized convolutional discriminator. Rather than classifying an entire image as simply real or fake, the discriminator predicts these local semantic labels while identifying compression artifacts. The system was trained on the large-scale OpenImages dataset and evaluated against traditional standard codecs as well as state-of-the-art neural compression models across benchmark datasets including CLIC2020, DIV2K, and Kodak across multiple compression bitrates.

The evaluation produced several key findings. First, the MS-ILLM framework achieved the same statistical realism—measured via Fréchet Inception Distance (FID)—as the prior state-of-the-art neural compressor while using 30% to 40% fewer bits on the CLIC2020 benchmark. Second, when compared at equal bitrate and traditional distortion levels, MS-ILLM consistently produced superior statistical fidelity and fewer unnatural visual artifacts than competing generative methods. Third, architectural ablations revealed that omitting standard normalization layers in the discriminator yielded the best performance and faster convergence, while varying codebook sizes had minimal sensitivity.

These results demonstrate that enforcing local statistical alignment in compression models resolves the trade-off between image sharpness and artifact generation much more effectively than prior global discriminators. For organizations handling large volumes of visual data, this offers substantial bandwidth and storage cost reductions without compromising visual texture quality. However, like other generative compression models, the system still degrades fine structured details such as small text compared to non-adversarial baselines.

Organizations evaluating neural compression should consider piloting local likelihood-based architectures where bandwidth optimization and visual texture are paramount. Before broad production deployment, teams must conduct comprehensive human perceptual studies and evaluate computational overhead. Further engineering work is recommended to quantize and miniaturize the neural network models to make them practical for real-time edge devices.

The findings are supported with high confidence through consistent results across multiple standard datasets and alternative feature-extraction metrics. Nonetheless, limitations remain: the deep learning models require substantial computing resources, and adversarial synthesis carries a potential risk of introducing demographic or visual biases, requiring careful auditing prior to commercial deployment.

No sufficiently relevant recommendations were found.

Cover for Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models

Abstract

Lossy image compression aims to represent images in as few bits as possible while maintaining fidelity to the original. Theoretical results indicate that optimizing distortion metrics such as PSNR or MS-SSIM necessarily leads to a discrepancy in the statistics of original images from those of reconstructions, in particular at low bitrates, often manifested by the blurring of the compressed images. Previous work has leveraged adversarial discriminators to improve statistical fidelity. Yet these binary discriminators adopted from generative modeling tasks may not be ideal for image compression. In this paper, we introduce a non-binary discriminator that is conditioned on quantized local image representations obtained via VQ-VAE autoencoders. Our evaluations on the CLIC2020, DIV2K and Kodak datasets show that our discriminator is more effective for jointly optimizing distortion (e.g., PSNR) and statistical fidelity (e.g., FID) than the PatchGAN of the state-of-the-art HiFiC model. On CLIC2020, we obtain the same FID as HiFiC with 30-40% fewer bits.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Background
  • 3.1. Notation
  • 3.2. Rate-distortion-perception theory
  • 3.3. Lossy codec optimization
  • 3.4. Approximating the distributional divergence
  • 4. Method
  • 4.1. Autoencoder architecture
  • 4.2. Implicit local likelihood models
  • 4.3. Choice of labeling function
  • 4.4. Discriminator architecture
  • 4.5. Training
  • 5. Experiments
  • 5.1. Datasets and metrics
  • 5.2. Baseline models
  • 5.3. Main results
  • 5.4. Ablations
  • 6. Limitations
  • 7. Conclusion
  • References
  • A. Further training details and hyperparameters
  • A.1. Autoencoder pretraining
  • A.2. Labeler pretraining
  • A.3. Fine-tuning with GAN loss
  • B. Calculation of baseline metrics
  • C. Note on metrics computation: bitrate, FID, and KID
  • D. Further analysis of the rate-distortion-perception tradeoff
  • E. Additional qualitative examples
  • F. Experimental results on DIV2K and Kodak
  • G. Investigation of ImageNet feature alignment

Knowls

  1. Knowl 1 — Categorical adversarial losses enforce local reconstruction likelihoods

    model/method

    The implicit local likelihood model (ILLM) uses a discriminator that predicts a categorical label at each of W×HW\times H spatial locations, rather than making only a real-versus-fake decision for the whole image. For an image, the discriminator DϕD_\phi outputs probabilities in [0,1](C+1)×W×H[0,1]^{(C+1)\times W\times H}, normalized over the C+1C+1 classes at each location. The label map u(x)u(x) assigns original-image labels from classes 1,…,C1,\ldots,C; the one-hot map b0b_0 assigns class 00 everywhere to mark a reconstruction as fake. For source image x∼PXx\sim P_X, let x^\hat{x} be its codec reconstruction, and let ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle denote summation over classes and spatial locations, with the logarithm applied elementwise. The discriminator and reconstruction-generator losses are

    LD(ϕ)=−Ex∼PX⟨u(x),log⁡Dϕ(x)⟩−Ex∼PX⟨b0,log⁡Dϕ(x^)⟩,L_D(\phi)=-\mathbb{E}_{x\sim P_X}\langle u(x),\log D_\phi(x)\rangle-\mathbb{E}_{x\sim P_X}\langle b_0,\log D_\phi(\hat{x})\rangle, LG=−Ex∼PX⟨u(x),log⁡Dϕ(x^)⟩.L_G=-\mathbb{E}_{x\sim P_X}\langle u(x),\log D_\phi(\hat{x})\rangle.

    The discriminator learns the local labels of real images while assigning reconstructions to the fake class. The reconstruction loss instead asks the discriminator to assign each reconstructed image the labels of its corresponding source image. Thus local conditioning enters through the loss targets, not merely through an optional discriminator input.

  2. Knowl 2 — VQ-VAE code assignments provide the local labels

    model/method

    ILLM obtains its real-image label map from a pretrained vector-quantized variational autoencoder (VQ-VAE). At each latent location (i,j)(i,j), its encoder produces a channel vector e(i,j)(x)e^{(i,j)}(x); the codebook contains vectors m1,…,mCm_1,\ldots,m_C. The label at that location is the index of the nearest codebook vector:

    u^{(c,i,j)}(x)=\mathbf{1}\!\left[c=\arg\min_{q\in\{1,\ldots,C\}}\|e^{(i,j)}(x)-m_q\|_2^2\right],$$ where $c$ indexes a codebook label and $\mathbf{1}[\cdot]$ is the indicator. The resulting one-hot map has $C$ real-label channels; a separate channel is reserved for the fake class. The paper's default labeler uses a $32\times32$ latent grid for $256\times256$ images and a codebook of size $C=1024$. It is based on VQ-GAN, with ChannelNorm in place of GroupNorm and XCiT attention layers. The authors motivate this choice by noting that codebook Voronoi cells are convex, and that the labeler's continuous decoder maps connected latent neighborhoods to connected sets in image space; in their account, this makes label-space locality correspond to image-space locality. Labeler training uses a VGG-based LPIPS term, which the authors report was important for obtaining useful labels.
  3. Knowl 3 — MS-ILLM combines a mean-scale hyperprior codec with adversarial fidelity loss

    model/method

    MS-ILLM is a learned image codec formed by fine-tuning a Mean-Scale Hyperprior autoencoder with the ILLM adversarial loss. Its encoder fϕf_\phi maps an image xx to quantized latents yy, its entropy model gωg_\omega estimates the coded rate rϕ,ω(x)r_{\phi,\omega}(x), and its decoder hυh_\upsilon produces the reconstruction x^=hυ(y)\hat{x}=h_\upsilon(y). The latent prior uses a hyperprior to predict the means and scales of Gaussian distributions. The training objective is

    L=λr Ex∼PX[rϕ,ω(x)]+λρ Ex∼PX[ρ(x^,x)]+λdLG,L=\lambda_r\,\mathbb{E}_{x\sim P_X}[r_{\phi,\omega}(x)]+\lambda_\rho\,\mathbb{E}_{x\sim P_X}[\rho(\hat{x},x)]+\lambda_d L_G,

    where λr\lambda_r, λρ\lambda_\rho, and λd\lambda_d weight rate, reference distortion, and adversarial fidelity, respectively, and LGL_G is the ILLM reconstruction loss. The distortion used in the reported training recipe is ρ(x^,x)=150∥x^−x∥22+LPIPS⁡Alex(x^,x)\rho(\hat{x},x)=150\|\hat{x}-x\|_2^2+\operatorname{LPIPS}_{\mathrm{Alex}}(\hat{x},x) for images normalized to [0,1][0,1]. The discriminator is a residual U-Net with LeakyReLU activations, no convolutional normalization, and an output resolution matching the 32×3232\times32 label map. At adversarial fine-tuning, the encoder and bottleneck are frozen and the decoder is fine-tuned.

  4. Knowl 4 — Two-stage training and optimization settings

    experimental setup

    The authors first train the mean-scale hyperprior autoencoder without an adversarial discriminator for 1 million steps, using AdamW with peak learning rate 3×10−43\times10^{-4} and weight decay 5×10−55\times10^{-5}. The learning rate warms up linearly for 10,000 steps and then follows cosine decay. They use rate targeting and straight-through gradients for quantized latents so that the decoder receives quantized values during training. The VQ-VAE labeler is pretrained separately with the same general schedule; its distortion combines MSE weighted by 1.01.0 with VGG-based LPIPS.

    In the second stage, the encoder and bottleneck remain frozen while the decoder is fine-tuned with the rate-distortion-adversarial objective. The reported ILLM discriminator learning rate is 4×10−44\times10^{-4} and the generator/autoencoder learning rate is 1×10−41\times10^{-4}. Fine-tuning uses AdamW with betas (0.5,0.9)(0.5,0.9); pretraining uses the same optimizer with its other settings. The supplementary material lists eight pretraining target rates: 0.008750.00875, 0.01750.0175, 0.0350.035, 0.070.07, 0.140.14, 0.300.30, 0.450.45, and 0.90.9 bits per pixel.

  5. Knowl 5 — Evaluation datasets and metrics

    experimental setup

    All learned models are trained on the OpenImages V6 training split using full-resolution images, with random resized crops or crops to 256×256256\times256 and random horizontal flips. Evaluation uses the CLIC2020 test split, DIV2K validation split, and Kodak dataset. Reference-based measures are PSNR, MS-SSIM, LPIPS, and DISTS; distributional statistical-fidelity measures are FID and KID. FID and KID are computed from pretrained Inception V3 features, using the torch-fidelity implementation. The authors recompute baseline metrics through a common evaluation pipeline. They note that Kodak has only 24 images, too few for useful FID or KID estimates.

  6. Knowl 6 — ILLM improves the CLIC2020 rate–distortion–fidelity tradeoff

    empirical result

    On the CLIC2020 test set, MS-ILLM achieves better statistical-fidelity scores than HiFiC at comparable reference distortion. With discriminator weight λd=0.008\lambda_d=0.008, the authors match the PSNR of HiFiC's original reported model and obtain uniformly better FID and KID across the evaluated bitrates. With λd=0.0005\lambda_d=0.0005, they report matching HiFiC's FID/KID while obtaining higher PSNR. In comparisons across multiple MSE operating points against their own PatchGAN-trained baseline, MS-PatchGAN (HiFiC*), MS-ILLM has better FID at every investigated point. The abstract further reports achieving the same FID as HiFiC on CLIC2020 with 30–40% fewer bits.

  7. Knowl 7 — Results generalize across datasets and an alternative FID feature extractor

    empirical result

    On DIV2K validation, MS-ILLM matches HiFiC on reference-based metrics (MS-SSIM, PSNR, LPIPS, and DISTS) while outperforming it on FID and KID. On Kodak, where the dataset is too small for useful distributional metrics, the reported distortion curves show performance comparable to HiFiC across bitrates. The authors also recompute FID on CLIC2020 and DIV2K using features from a SwAV-trained ResNet-50 rather than Inception V3; MS-ILLM retains higher measured statistical fidelity across bitrates. They interpret this as evidence that the FID improvements are not solely due to alignment with ImageNet-class features. The paper notes that KID estimates on DIV2K were unstable at high bitrates, including negative estimates for MS-ILLM, so it omits that final KID point.

  8. Knowl 8 — Removing discriminator normalization improves the measured tradeoff

    data/table

    The authors compare no normalization with instance normalization in the ILLM U-Net at three approximate bitrates. FID is lower when no normalization is used at every listed rate, while PSNR is similar or slightly higher. The entries below report bitrate in bits per pixel, FID (lower is better), and PSNR in dB (higher is better).

    No normalization Instance normalization
    Bitrate FID PSNR (dB) FID PSNR (dB)
    ≈0.035\approx 0.035 6.27 27.9 8.28 28.0
    ≈0.121\approx 0.121 2.65 30.9 3.04 30.7
    ≈0.231\approx 0.231 1.77 33.0 2.21 32.8

    Spectral normalization was also tested, but the authors report that it took as much as 25 times more gradient steps to begin matching the performance of the other configurations. They report no major training-stability problems with no normalization.

  9. Knowl 9 — Label-grid and codebook-size ablations have modest effects

    data/table

    At approximately 0.1210.121 bits per pixel, the authors vary the VQ-VAE label map's spatial dimensions and codebook size. The measured FID changes modestly with spatial size and is nearly unchanged by codebook size; PSNR also changes only slightly. The table reports FID (lower is better) and PSNR in dB (higher is better).

    Spatial size Codebook size FID PSNR (dB)
    16×1616\times16 256 2.66 31.0
    16×1616\times16 1024 2.65 31.1
    32×3232\times32 256 2.57 30.9
    32×3232\times32 1024 2.55 30.8
  10. Knowl 10 — Limitations include deployment, bias, and imperfect quality surrogates

    limitation

    The authors caution that adversarial neural compression could produce demographic biases, including effects associated with race or gender, and recommend treating the results as research rather than production-ready without further assessment. The HiFiC-based autoencoder also requires substantial computation; practical deployment would require model miniaturization and heavy weight quantization. Better FID or KID does not establish better human preference. A qualitative CLIC2020 text example illustrates a further tradeoff: both HiFiC and MS-ILLM degrade textual quality compared with the non-adversarial model, although MS-ILLM is reported to make one word slightly more legible than HiFiC.

Coverage note — No substantial contribution was omitted; qualitative image panels and auxiliary learning-rate plots were left out because they illustrate, rather than add independent results to, the methods and findings captured here.

References

  1. 1.Agustsson, E. and Timofte, R. NTIRE 2017 challenge on single image super-resolution: Dataset and study. In CVPRW, 2017.
  2. 2.Agustsson, E., Tschannen, M., Mentzer, F., Timofte, R., and Gool, L. V. Generative adversarial networks for extreme learned image compression. In ICCV, 2019.
  3. 3.Agustsson, E., Minnen, D., Toderici, G., and Mentzer, F. Multi-realism image compression with a conditional generator. In CVPR, 2023.
  4. 4.Ahmed, N., Natarajan, T., and Rao, K. Discrete cosine transform. IEEE Trans. Comput., C-23(1):90–93, 1974.
  5. 5.Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K. Deep variational information bottleneck. In ICLR, 2017.
  6. 6.Ali, S. M. and Silvey, S. D. A general class of coefficients of divergence of one distribution from another. J R Stat Soc Series B Stat Methodol, 28(1):131–142, 1966.
  7. 7.Antonini, M., Barlaud, M., Mathieu, P., and Daubechies, I. Image coding using wavelet transform. IEEE TIP, 1(2):205–220, 1992.
  8. 8.Balle, J., Laparra, V., and Simoncelli, E. P. End-to-end optimized image compression. In ICLR, 2017.
  9. 9.Balle, J., Minnen, D., Singh, S., Hwang, S. J., and Johnston, N. Variational image compression with a scale hyperprior. In ICLR, 2018.
  10. 10.Balle, J., Hwang, S. J., and Agustsson, E. TensorFlow Compression: Learned data compression, 2022. URL http://github.com/tensorflow/compression.
  11. 11.Begaint, J., Racapé, F., Feltman, S., and Pushparaja, A. CompressAI: a PyTorch library and evaluation platform for end-to-end compression research. arXiv preprint, arXiv:2011.03029, 2020.
  12. 12.Bellard, F. BPG image format. URL https://bellard.org/bpg/.
  13. 13.Bhardwaj, S., Fischer, I., Balle, J., and Chinen, T. An unsupervised information-theoretic perceptual quality metric. In NeurIPS, 2020.
  14. 14.Binkowski, M., Sutherland, D. J., Arbel, M., and Gretton, A. Demystifying MMD GANs. ICLR, 2018.
  15. 15.Blau, Y. and Michaeli, T. The perception-distortion tradeoff. In CVPR, 2018.
  16. 16.Blau, Y. and Michaeli, T. Rethinking lossy compression: The rate-distortion-perception tradeoff. In ICML, 2019.
  17. 17.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. NeurIPS, 2020.
  18. 18.Cheng, Z., Sun, H., Takeuchi, M., and Katto, J. Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In CVPR, 2020.
  19. 19.Cover, T. and Thomas, J. Elements of information theory. John Wiley & Sons, 1991.
  20. 20.Csiszar, I. Information-type measures of difference of probability distributions and indirect observation. studia scientiarum Mathematicarum Hungarica, 2:229–318, 1967.
  21. 21.Ding, K., Ma, K., Wang, S., and Simoncelli, E. P. Image quality assessment: Unifying structure and texture similarity. IEEE TPAMI, 44(5):2567–2581, 2020.
  22. 22.Ding, K., Ma, K., Wang, S., and Simoncelli, E. P. Comparison of full-reference image quality models for optimization of image processing systems. IJCV, 129(4):1258–1281, 2021.
  23. 23.Dubois, Y., Bloem-Reddy, B., Ullrich, K., and Maddison, C. J. Lossy compression for lossless prediction. In NeurIPS, 2021.
  24. 24.El-Nouby, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., and Jegou, H. XCiT: Cross-covariance image transformers. In NeurIPS, 2021.
  25. 25.El-Nouby, A., Muckley, M. J., Ullrich, K., Laptev, I., Verbeek, J., and Jegou, H. Image compression with product quantized masked image modeling. Trans. Mach. Learn. Res., 2023.
  26. 26.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In CVPR, 2021.
  27. 27.Ghouse, N. F., Petersen, J., Wiggers, A., Xu, T., and Sautiere, G. A residual diffusion model for high perceptual quality codec augmentation. arXiv preprint, arXiv:2301.05489, 2023.
  28. 28.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. In NeurIPS, 2014.
  29. 29.He, D., Yang, Z., Peng, W., Ma, R., Qin, H., and Wang, Y. ELIC: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In CVPR, pp. 5718–5727, June 2022.
  30. 30.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. NeurIPS, 2017.
  31. 31.Krizhevsky, A., Sutskever, I., and Hinton, G. E. ImageNet classification with deep convolutional neural networks. In NeurIPS, 2012.
  32. 32.Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al. The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 128(7):1956–1981, 2020.
  33. 33.Kynkäänniemi, T., Karras, T., Aittala, M., Aila, T., and Lehtinen, J. The role of ImageNet classes in Fréchet inception distance. In ICLR, 2023.
  34. 34.Le Gall, D. and Tabatabai, A. Sub-band coding of digital images using symmetric short kernel filters and arithmetic coding techniques. In ICASSP, 1988.
  35. 35.Ledig, C., Theis, L., Huszár, F., Caballero, J., Cunningham, A., Acosta, A., Aitken, A., Tejani, A., Totz, J., Wang, Z., et al. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR, 2017.
  36. 36.Liese, F. and Vajda, I. f-divergences: Sufficiency, deficiency and testing of hypotheses. In Barnett, N. S. (ed.), Advances in Inequalities from Probability Theory & Statistics, pp. 113–173. Nova Publishers, 2008.
  37. 37.Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J. On the variance of the adaptive learning rate and beyond. In ICLR, 2020.
  38. 38.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. ICLR, 2019.
  39. 39.MacKay, D. J. Information Theory, Inference, and Learning Algorithms. Cambridge University Press, 2003.
  40. 40.Matsubara, Y., Yang, R., Levorato, M., and Mandt, S. Supervised compression for resource-constrained edge computing systems. In CVPR, 2022.
  41. 41.Mentzer, F., Toderici, G. D., Tschannen, M., and Agustsson, E. High-fidelity generative image compression. In NeurIPS, 2020.
  42. 42.Mentzer, F., Agustsson, E., Ballé, J., Minnen, D., Johnston, N., and Toderici, G. Neural video compression using GANs for detail synthesis and propagation. In ECCV, 2022.
  43. 43.Minnen, D., Ballé, J., and Toderici, G. D. Joint autoregressive and hierarchical priors for learned image compression. In NeurIPS, 2018.
  44. 44.Mirza, M. and Osindero, S. Conditional generative adversarial nets. arXiv preprint, arXiv:1411.1784, 2014.
  45. 45.Mittal, A., Soundararajan, R., and Bovik, A. C. Making a “completely blind” image quality analyzer. IEEE Sign. Process. Letters, 20(3):209–212, 2012.
  46. 46.Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. Spectral normalization for generative adversarial networks. In ICLR, 2018.
  47. 47.Nowozin, S., Cseke, B., and Tomioka, R. f-GAN: Training generative neural samplers using variational divergence minimization. NeurIPS, 2016.
  48. 48.Pasco, R. C. Source coding algorithms for fast data compression. PhD thesis, Stanford University CA, 1976.
  49. 49.Qian, J., Zhang, G., Chen, J., and Khisti, A. A rate-distortion-perception theory for binary sources. In International Zurich Seminar on Information and Communication, pp. 34–38, 2022.
  50. 50.Razavi, A., van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with VQ-VAE-2. In NeurIPS, 2019.
  51. 51.Rissanen, J. J. Generalized Kraft inequality and arithmetic coding. IBM J. Res. Dev., 20(3):198–203, 1976.
  52. 52.Ronneberger, O., Fischer, P., and Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  53. 53.Shannon, C. E. A mathematical theory of communication. Bell Syst. Tech. J., 27(3):379–423, 1948.
  54. 54.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. In ICLR, 2014.
  55. 55.Singh, S., Abu-El-Haija, S., Johnston, N., Ballé, J., Shrivastava, A., and Toderici, G. End-to-end learning of compressible features. In ICIP, 2020.
  56. 56.Snyder, H. L. Image quality: Measures and visual performance. In Flat-panel displays and CRTs, pp. 70–90. Springer, 1985.
  57. 57.Sushko, V., Schönfeld, E., Zhang, D., Gall, J., Schiele, B., and Khoreva, A. OASIS: Only adversarial supervision for semantic image synthesis. IJCV, 130(12):2903–2923, 2022.
  58. 58.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the Inception architecture for computer vision. In CVPR, 2016.
  59. 59.Theis, L., Shi, W., Cunningham, A., and Huszár, F. Lossy image compression with compressive autoencoders. In ICLR, 2017.
  60. 60.Theis, L., Salimans, T., Hoffman, M. D., and Mentzer, F. Lossy compression with Gaussian diffusion. arXiv preprint, arXiv:2206.08889, 2022.
  61. 61.Tishby, N., Pereira, F. C., and Bialek, W. The information bottleneck method. In Annu. Allerton Conf. Commun. Control Comput., pp. 368–377, 1999.
  62. 62.Toderici, G., Theis, L., Johnston, N., Agustsson, E., Mentzer, F., Ballé, J., Shi, W., and Timofte, R. CLIC 2020: Challenge on learned image compression, 2020.
  63. 63.Torfason, R., Mentzer, F., Agustsson, E., Tschannen, M., Timofte, R., and Van Gool, L. Towards image understanding from deep compression without decoding. In ICLR, 2018.
  64. 64.Tschannen, M., Agustsson, E., and Lucic, M. Deep generative models for distribution-preserving lossy compression. In NeurIPS, 2018.
  65. 65.Ulyanov, D., Vedaldi, A., and Lempitsky, V. Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis. In CVPR, 2017.
  66. 66.van den Oord, A., Vinyals, O., and Kavukcuoglu, k. Neural discrete representation learning. In NeurIPS, 2017.
  67. 67.Wang, Z., Simoncelli, E., and Bovik, A. Multiscale structural similarity for image quality assessment. In ACSSC, volume 2, pp. 1398–1402, 2003.
  68. 68.Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 13(4):600–612, 2004.
  69. 69.Xu, B., Wang, N., Chen, T., and Li, M. Empirical evaluation of rectified activations in convolutional network. arXiv preprint, arXiv:1505.00853, 2015.
  70. 70.Yan, Z., Wen, F., Ying, R., Ma, C., and Liu, P. On perceptual lossy compression: The cost of perceptual reconstruction and an optimal training framework. In ICML, 2021.
  71. 71.Yang, R. and Mandt, S. Lossy image compression with conditional diffusion models. arXiv preprint, arXiv:2209.06950, 2022.
  72. 72.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.

Citation

MLA
Muckley, M. J., et al. “Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 25426–43, https://proceedings.mlr.press/v202/muckley23a.html.
APA
Muckley, M. J., El-Nouby, A., Ullrich, K., Jegou, H., & Verbeek, J. (2023). Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models. International Conference on Machine Learning, 202, 25426–25443. https://proceedings.mlr.press/v202/muckley23a.html
Chicago
Muckley, M. J., A. El-Nouby, K. Ullrich, H. Jegou, and J. Verbeek. 2023. “Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models”. International Conference on Machine Learning 202: 25426–43. https://proceedings.mlr.press/v202/muckley23a.html.
Harvard
Muckley, M.J. et al. (2023) “Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models”, International Conference on Machine Learning. PMLR, pp. 25426–25443. Available at: https://proceedings.mlr.press/v202/muckley23a.html.
Vancouver
1. Muckley MJ, El-Nouby A, Ullrich K, Jegou H, Verbeek J (2023) Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models. In: International Conference on Machine Learning. PMLR, pp 25426–25443

BibTeX

@InProceedings{pmlr-v202-muckley23a,
  title = 	 {Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood Models},
  author =       {Muckley, Matthew J. and El-Nouby, Alaaeldin and Ullrich, Karen and Jegou, Herve and Verbeek, Jakob},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {25426--25443},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/muckley23a/muckley23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/muckley23a.html},
  abstract = 	 {Lossy image compression aims to represent images in as few bits as possible while maintaining fidelity to the original. Theoretical results indicate that optimizing distortion metrics such as PSNR or MS-SSIM necessarily leads to a discrepancy in the statistics of original images from those of reconstructions, in particular at low bitrates, often manifested by the blurring of the compressed images. Previous work has leveraged adversarial discriminators to improve statistical fidelity. Yet these binary discriminators adopted from generative modeling tasks may not be ideal for image compression. In this paper, we introduce a non-binary discriminator that is conditioned on quantized local image representations obtained via VQ-VAE autoencoders. Our evaluations on the CLIC2020, DIV2K and Kodak datasets show that our discriminator is more effective for jointly optimizing distortion (e.g., PSNR) and statistical fidelity (e.g., FID) than the PatchGAN of the state-of-the-art HiFiC model. On CLIC2020, we obtain the same FID as HiFiC with 30-40% fewer bits.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/