keyword
transformer-based image compression
Transformer-based image compression is a machine learning approach for reducing the data size of digital images by utilizing transformer neural network architectures to encode, decode, and model visual information. Unlike traditional handcrafted codecs or convolutional neural network models that are constrained by local receptive fields, transformer-based methods leverage self-attention mechanisms to capture long-range spatial dependencies and global contextual relationships across an entire image. In an end-to-end learned compression framework, transformer modules are typically incorporated into the autoencoder transforms to extract compact latent representations or into the entropy modeling stage to accurately estimate probability distributions, thereby optimizing the trade-off between visual fidelity and the required bitrate.
1 item

