Built independently by an author, for readers. Read the story and support ChapterPal

keyword

lossy compression

Lossy compression is a data encoding method that reduces the size of a digital file by permanently discarding unnecessary or less perceptible information. Unlike lossless compression, the reconstructed file is an approximation of the original data rather than an exact duplicate. This approach deliberately trades reconstruction accuracy for substantially smaller file sizes and reduced transmission bandwidth, balancing the amount of data saved against the level of distortion introduced. Commonly applied to multimedia formats such as images, audio, and video, lossy compression often relies on quantization and mathematical transforms to remove high-frequency details or redundancies that human perception cannot easily detect.

3 items

End-to-end Optimized Image Compression

End-to-end Optimized Image Compression

Johannes Ballé, Valero Laparra, Eero P. Simoncelli

OrganizationsImage Processing LaboratoryNew York UniversityUniversitat de València

Why you should read this

Establishes an end-to-end framework for learned image compression by jointly optimizing nonlinear transforms and a continuous quantization proxy for rate-distortion performance, outperforming traditional codecs like JPEG and JPEG 2000.

We describe an image compression method, consisting of a nonlinear analysis transformation, a uniform quantizer, and a nonlinear synthesis transformation. The transforms are constructed in three successive stages of convolutional linear filters and nonlinear activation functions. Unlike most convolutional neural networks, the joint nonlinearity is chosen to implement a form of local gain control, inspired by those used to model biological neurons. Using a variant of stochastic gradient descent, we jointly optimize the entire model for rate-distortion performance over a database of training images, introducing a continuous proxy for the discontinuous loss function arising from the quantizer. Under certain conditions, the relaxed loss function may be interpreted as the log likelihood of a generative model, as implemented by a variational autoencoder. Unlike these models, however, the compression model must operate at any given point along the rate-distortion curve, as specified by a trade-off parameter. Across an independent set of test images, we find that the optimized method generally exhibits better rate-distortion performance than the standard JPEG and JPEG 2000 compression methods. More importantly, we observe a dramatic improvement in visual quality for all images at all bit rates, which is supported by objective quality estimates using MS-SSIM.

Added

2026-09-16

Generating Diverse High-Fidelity Images with VQ-VAE-2

Generating Diverse High-Fidelity Images with VQ-VAE-2

Ali Razavi, Aäron van den Oord, Oriol Vinyals

OrganizationsGoogle

Why you should read this

Demonstrates that hierarchical discrete VAEs can generate photorealistic, high-resolution images, rivaling the best adversarial models.

We explore the use of Vector Quantized Variational AutoEncoder (VQ-VAE) models for large scale image generation. To this end, we scale and enhance the autoregressive priors used in VQ-VAE to generate synthetic samples of much higher coherence and fidelity than possible before. We use simple feed-forward encoder and decoder networks, making our model an attractive candidate for applications where the encoding and/or decoding speed is critical. Additionally, VQ-VAE requires sampling an autoregressive model only in the compressed latent space, which is an order of magnitude faster than sampling in the pixel space, especially for large images. We demonstrate that a multi-scale hierarchical organization of VQ-VAE, augmented with powerful priors over the latent codes, is able to generate samples with quality that rivals that of state of the art Generative Adversarial Networks on multifaceted datasets such as ImageNet, while not suffering from GAN's known shortcomings such as mode collapse and lack of diversity.

Added

2026-03-09

Variational Lossy Autoencoder

Variational Lossy Autoencoder

Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, Pieter Abbeel

OrganizationsOpenAIUniversity of California Berkeley

Why you should read this

Reveals that the common failure of VAEs to use their latent codes when paired with powerful decoders isn't a bug but a controllable feature—by deliberately limiting what the decoder can model locally (like small texture patches), you can force the latent code to capture exactly the global structure you care about while achieving state-of-the-art density estimation.

Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good representation for 2D images might be one that describes only global structure and discards information about detailed texture. In this paper, we present a simple but principled method to learn such global representations by combining Variational Autoencoder (VAE) with neural autoregressive models such as RNN, MADE and PixelRNN/CNN. Our proposed VAE model allows us to have control over what the global latent code can learn and , by designing the architecture accordingly, we can force the global latent code to discard irrelevant information such as texture in 2D images, and hence the VAE only "autoencodes" data in a lossy fashion. In addition, by leveraging autoregressive models as both prior distribution p(z) and decoding distribution p(x|z), we can greatly improve generative modeling performance of VAEs, achieving new state-of-the-art results on MNIST, OMNIGLOT and Caltech-101 Silhouettes density estimation tasks.

Added

2026-02-21