Built independently by an author, for readers. Read the story and support ChapterPal

topic

arithmetic coding

Arithmetic coding is a lossless data compression method that encodes an entire message as a single fractional number, typically within the interval between zero and one. Rather than mapping each individual symbol to a separate sequence of bits, the algorithm begins with a baseline interval and iteratively narrows it into smaller subintervals whose sizes correspond to the probability of each successive symbol in the input sequence. By representing the full message within a progressively refined numerical range rather than assigning integer bit lengths to individual characters, arithmetic coding can effectively allocate fractional bits per symbol, allowing it to achieve compression ratios that closely approach the theoretical entropy limit of the source data.

2 items

Joint Autoregressive and Hierarchical Priors for Learned Image Compression

Joint Autoregressive and Hierarchical Priors for Learned Image Compression

David Minnen, Johannes Ballé, George Toderici

OrganizationsGoogle

Why you should read this

Presents a learned image compression architecture that couples autoregressive and hierarchical priors in the entropy model, establishing the first deep learning approach to outperform traditional BPG codecs across both PSNR and MS-SSIM rate-distortion metrics.

Recent models for learned image compression are based on autoencoders, learning approximately invertible mappings from pixels to a quantized latent representation. These are combined with an entropy model, a prior on the latent representation that can be used with standard arithmetic coding algorithms to yield a compressed bitstream. Recently, hierarchical entropy models have been introduced as a way to exploit more structure in the latents than simple fully factorized priors, improving compression performance while maintaining end-to-end optimization. Inspired by the success of autoregressive priors in probabilistic generative models, we examine autoregressive, hierarchical, as well as combined priors as alternatives, weighing their costs and benefits in the context of image compression. While it is well known that autoregressive models come with a significant computational penalty, we find that in terms of compression performance, autoregressive and hierarchical priors are complementary and, together, exploit the probabilistic structure in the latents better than all previous learned models. The combined model yields state-of-the-art rate--distortion performance, providing a 15.8% average reduction in file size over the previous state-of-the-art method based on deep learning, which corresponds to a 59.8% size reduction over JPEG, more than 35% reduction compared to WebP and JPEG2000, and bitstreams 8.4% smaller than BPG, the current state-of-the-art image codec. To the best of our knowledge, our model is the first learning-based method to outperform BPG on both PSNR and MS-SSIM distortion metrics.

Added

2026-09-24

Byte Latent Transformer: Patches Scale Better Than Tokens

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer

Why you should read this

Presents a novel byte-level LLM architecture, the Byte Latent Transformer, that for the first time matches tokenization-based LLM performance at scale, significantly improving inference efficiency and robustness by dynamically adapting compute allocation to data complexity.

We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency and robustness. BLT encodes bytes into dynamically sized patches, which serve as the primary units of computation. Patches are segmented based on the entropy of the next byte, allocating more compute and model capacity where increased data complexity demands it. We present the first FLOP controlled scaling study of byte-level models up to 8B parameters and 4T training bytes. Our results demonstrate the feasibility of scaling models trained on raw bytes without a fixed vocabulary. Both training and inference efficiency improve due to dynamically selecting long patches when data is predictable, along with qualitative improvements on reasoning and long tail generalization. Overall, for fixed inference costs, BLT shows significantly better scaling than tokenization-based models, by simultaneously growing both patch and model size.

Added

2026-01-07

Creative Commons License