keyword
bits-back coding
Bits-back coding is a lossless data compression method designed for latent variable models that achieves compression rates close to the theoretical variational bound on marginal likelihood. In this scheme, an encoder uses auxiliary information from an existing bitstream to pseudorandomly sample a latent variable from an approximate posterior distribution, and then encodes the data point conditioned on this latent state alongside the latent prior. The decoder reconstructs the data point, recovers the latent variable, and uses it to decode the initial auxiliary information, thereby retrieving the extra bits that were used during sampling. Because these auxiliary bits are recovered and can represent other useful data, the net transmission cost for the message is reduced by the entropy of the approximate posterior, allowing complex probabilistic models to compress data efficiently without requiring exact Bayesian inference.
2 items

Practical Variational Inference for Neural Networks
Alex Graves
Why you should read this
Proposes a stochastic variational inference method with a diagonal Gaussian posterior that scales practical Bayesian learning, regularisation, and weight pruning to general differentiable neural network architectures.
Variational methods have been previously explored as a tractable approximation to Bayesian inference for neural networks. However the approaches proposed so far have only been applicable to a few simple network architectures. This paper introduces an easy-to-implement stochastic variational method (or equivalently, minimum description length loss function) that can be applied to most neural networks. Along the way it revisits several common regularisers from a variational perspective. It also provides a simple pruning heuristic that can both drastically reduce the number of network weights and lead to improved generalisation. Experimental results are provided for a hierarchical multidimensional recurrent neural network applied to the TIMIT speech corpus.
Added
2026-09-18

Variational Lossy Autoencoder
Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, Pieter Abbeel
Why you should read this
Reveals that the common failure of VAEs to use their latent codes when paired with powerful decoders isn't a bug but a controllable feature—by deliberately limiting what the decoder can model locally (like small texture patches), you can force the latent code to capture exactly the global structure you care about while achieving state-of-the-art density estimation.
Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good representation for 2D images might be one that describes only global structure and discards information about detailed texture. In this paper, we present a simple but principled method to learn such global representations by combining Variational Autoencoder (VAE) with neural autoregressive models such as RNN, MADE and PixelRNN/CNN. Our proposed VAE model allows us to have control over what the global latent code can learn and , by designing the architecture accordingly, we can force the global latent code to discard irrelevant information such as texture in 2D images, and hence the VAE only "autoencodes" data in a lossy fashion. In addition, by leveraging autoregressive models as both prior distribution p(z) and decoding distribution p(x|z), we can greatly improve generative modeling performance of VAEs, achieving new state-of-the-art results on MNIST, OMNIGLOT and Caltech-101 Silhouettes density estimation tasks.
Added
2026-02-21
