Frequency-Aware Transformer for Learned Image Compression
Han LiShaohui LiWenrui DaiChenglin LiJunni ZouHongkai Xiong
Introduces a frequency-aware transformer architecture for learned image compression that captures directional image details through multiscale frequency decomposition, outperforming the VTM-12.1 standard codec by more than 13% in BD-rate across benchmark datasets.
Modern digital workflows generate massive volumes of imagery, making efficient data compression critical for cutting bandwidth costs and storage demands. While neural network-based compression methods have advanced rapidly, existing solutions rely on uniform processing windows that fail to separate fine directional details from broader background structures, leading to redundant data encoding.
The article evaluates whether integrating multiscale directional frequency analysis into transformer architectures can optimize learned image compression. Specifically, it demonstrates a new framework that decomposes natural image features across orientations and adaptively compresses them.
The researchers developed the Frequency-aware Transformer-based learned Image Compression model, which introduces anisotropic window attention to process low, high, horizontal, and vertical image frequencies simultaneously. They also incorporated a frequency-modulation feed-forward network to adjust frequency weights dynamically and a transformer-based channel-wise autoregressive model to eliminate redundancy across data channels. The system was trained on standard public datasets totaling over a million images and evaluated across three established benchmarks: Kodak, Tecnick, and CLIC.
The evaluation yielded several key findings. First, the proposed framework achieved state-of-the-art compression efficiency, reducing average bitrate by 14.5% on Kodak, 15.1% on Tecnick, and 13.0% on CLIC compared to the latest standardized benchmark codec, VTM-12.1. Second, it surpassed existing deep-learning methods in preserving fine edge details and directional textures. Third, ablation testing confirmed that directional window attention accounts for the majority of the performance gains without incurring the computational penalties of larger uniform windows. Finally, the framework maintained practical inference speeds, achieving an encoding latency of 125 milliseconds and a decoding latency of 242 milliseconds per image.
These results demonstrate that frequency-aware attention enables higher visual fidelity at significantly lower file sizes than both conventional standards and earlier neural approaches. For organizations handling large-scale image storage and transmission, adopting these methods offers substantial reductions in bandwidth expenses and infrastructure loads without compromising visual quality.
Organizations should consider evaluating frequency-aware neural compression in pilot deployments for high-throughput image storage pipelines. For decision-makers planning implementations, trade-offs between hardware encoding speed and transmission bandwidth savings should be weighed. Future engineering should expand the underlying attention mechanisms to capture diagonal directions and adapt the framework to spatial-temporal video compression.
The conclusions are supported by rigorous benchmarking across standard datasets; however, evaluations were confined to static natural images on workstation-class graphics processors. Performance under edge-device hardware constraints or with non-standard visual data remains an area for cautious validation.
- Paper: ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive Coding, Dailan He et al. (2022). Introduces state-of-the-art space-channel contextual adaptive coding for learned image compression that motivates FAT's transformer-based channel-wise autoregressive design.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). Establishes joint autoregressive and hierarchical entropy priors for learned image compression upon which modern latent context modeling relies.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). Demonstrates the integration of attention modules with flexible likelihood models to capture spatial redundancy in learned image compression.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). Pioneers the variational scale hyperprior architecture that serves as the foundation for modern learned image compression autoencoders.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). Provides the seminal end-to-end rate-distortion optimization framework and continuous quantization relaxation used across neural compression models.
- Paper: Cross Aggregation Transformer for Image Restoration, Zheng Chen et al. (2022). Presents directional and rectangular window aggregation mechanisms in vision transformers that directly relate to directional frequency attention.
No sufficiently relevant recommendations were found.
