Frequency-Aware Transformer for Learned Image Compression

Han LiShaohui LiWenrui DaiChenglin LiJunni ZouHongkai Xiong

article2024ICLR87 citations

Introduces a frequency-aware transformer architecture for learned image compression that captures directional image details through multiscale frequency decomposition, outperforming the VTM-12.1 standard codec by more than 13% in BD-rate across benchmark datasets.

Listen

Modern digital workflows generate massive volumes of imagery, making efficient data compression critical for cutting bandwidth costs and storage demands. While neural network-based compression methods have advanced rapidly, existing solutions rely on uniform processing windows that fail to separate fine directional details from broader background structures, leading to redundant data encoding.

The article evaluates whether integrating multiscale directional frequency analysis into transformer architectures can optimize learned image compression. Specifically, it demonstrates a new framework that decomposes natural image features across orientations and adaptively compresses them.

The researchers developed the Frequency-aware Transformer-based learned Image Compression model, which introduces anisotropic window attention to process low, high, horizontal, and vertical image frequencies simultaneously. They also incorporated a frequency-modulation feed-forward network to adjust frequency weights dynamically and a transformer-based channel-wise autoregressive model to eliminate redundancy across data channels. The system was trained on standard public datasets totaling over a million images and evaluated across three established benchmarks: Kodak, Tecnick, and CLIC.

The evaluation yielded several key findings. First, the proposed framework achieved state-of-the-art compression efficiency, reducing average bitrate by 14.5% on Kodak, 15.1% on Tecnick, and 13.0% on CLIC compared to the latest standardized benchmark codec, VTM-12.1. Second, it surpassed existing deep-learning methods in preserving fine edge details and directional textures. Third, ablation testing confirmed that directional window attention accounts for the majority of the performance gains without incurring the computational penalties of larger uniform windows. Finally, the framework maintained practical inference speeds, achieving an encoding latency of 125 milliseconds and a decoding latency of 242 milliseconds per image.

These results demonstrate that frequency-aware attention enables higher visual fidelity at significantly lower file sizes than both conventional standards and earlier neural approaches. For organizations handling large-scale image storage and transmission, adopting these methods offers substantial reductions in bandwidth expenses and infrastructure loads without compromising visual quality.

Organizations should consider evaluating frequency-aware neural compression in pilot deployments for high-throughput image storage pipelines. For decision-makers planning implementations, trade-offs between hardware encoding speed and transmission bandwidth savings should be weighed. Future engineering should expand the underlying attention mechanisms to capture diagonal directions and adapt the framework to spatial-temporal video compression.

The conclusions are supported by rigorous benchmarking across standard datasets; however, evaluations were confined to static natural images on workstation-class graphics processors. Performance under edge-device hardware constraints or with non-standard visual data remains an area for cautious validation.

No sufficiently relevant recommendations were found.

Cover for Frequency-Aware Transformer for Learned Image Compression

Abstract

Learned image compression (LIC) has gained traction as an effective solution for image storage and transmission in recent years. However, existing LIC methods are redundant in latent representation due to limitations in capturing anisotropic frequency components and preserving directional details. To overcome these challenges, we propose a novel frequency-aware transformer (FAT) block that for the first time achieves multiscale directional ananlysis for LIC. The FAT block comprises frequency-decomposition window attention (FDWA) modules to capture multiscale and directional frequency components of natural images. Additionally, we introduce frequency-modulation feed-forward network (FMFFN) to adaptively modulate different frequency components, improving rate-distortion performance. Furthermore, we present a transformer-based channel-wise autoregressive (T-CA) model that effectively exploits channel dependencies. Experiments show that our method achieves state-of-the-art rate-distortion performance compared to existing LIC methods, and evidently outperforms latest standardized codec VTM-12.1 by 14.5%, 15.1%, 13.0% in BD-rate on the Kodak, Tecnick, and CLIC datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Transformer-based Image Compression
  • 2.2 Autoregressive Entropy Modeling
  • 2.3 Frequency Decomposition in Learned Image Compression
  • 3 Methods
  • 3.1 Overview
  • 3.2 Frequency-aware transformer block
  • 3.2.1 Frequency-Decomposed Window Attention
  • 3.2.2 Frequency-Modulation Feed-forward Network
  • 3.3 Transformer-based Channel-wise Autoregressive (T-CA) Entropy Model
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Rate-Distortion Performance
  • 4.3 Ablation Studies and Analysis
  • 5 Conclusion
  • References
  • A Model Architecture
  • A.1 Details of overall framework
  • A.2 Details of Shift-Window Operation
  • A.3 Details of Entropy Parameters Network
  • B Training
  • C Additional Experimental Results
  • C.1 R-D Performance on Tecnick and CLIC Professional Validation Datesets
  • C.2 Comparison on Coding Complexity
  • C.3 Comparison on Training Speed and GPU memory requirement
  • C.4 Visualization of channel attention weights in T-CA
  • C.4.1 Visual Quality Results
  • D Limitation and future work

Knowls

  1. Knowl 1 — FTIC Framework and Rate-Distortion Objective

    model/method

    The Frequency-aware Transformer-based learned Image Compression (FTIC) framework constructs nonlinear analysis and synthesis transforms using Frequency-Aware Transformer (FAT) blocks to capture multiscale and directional frequency components. Given an input raw image xx, the analysis transform ga(⋅)g_a(\cdot) extracts a latent representation y=ga(x)y = g_a(x). A uniform quantizer Q(⋅)Q(\cdot) discretizes yy into quantized latent y^=Q(y)\hat{y} = Q(y).

    To compress y^\hat{y}, a hyperprior path extracts hyper-latents z=ha(y)z = h_a(y), which are quantized to z^=Q(z)\hat{z} = Q(z), and decoded by the hyper-decoder to yield context ϕ=hs(z^)\phi = h_s(\hat{z}). The quantized latent y^\hat{y} is divided along the channel dimension into nsn_s slices {y^1,y^2,…,y^ns}\{\hat{y}_1, \hat{y}_2, \dots, \hat{y}_{n_s}\}. A Transformer-based Channel-wise Autoregressive (T-CA) entropy model predicts Gaussian distribution parameters (mean μi\mu_i, scale σi\sigma_i) and a latent quantization residual rir_i for each slice i∈{1,…,ns}i \in \{1, \dots, n_s\} conditioned on previously decoded slices y^<i\hat{y}_{<i} and hyperprior context ϕ\phi:

    ri,μi,σi=T-CA(ϕ,y^<i),1≤i≤nsr_i, \mu_i, \sigma_i = \text{T-CA}(\phi, \hat{y}_{<i}), \quad 1 \le i \le n_s

    A refined latent representation yˉ\bar{y} is formed by adding the accumulated residual r=Concat(r1,r2,…,rns)r = \text{Concat}(r_1, r_2, \dots, r_{n_s}) to y^\hat{y}:

    yˉ=y^+r\bar{y} = \hat{y} + r

    The reconstructed image x^\hat{x} is synthesized via x^=gs(yˉ)\hat{x} = g_s(\bar{y}). The model is optimized end-to-end via a rate-distortion Lagrangian loss:

    L=R(y^)+R(z^)+λ⋅D(x,x^)\mathcal{L} = R(\hat{y}) + R(\hat{z}) + \lambda \cdot D(x, \hat{x})

    where R(y^)R(\hat{y}) and R(z^)R(\hat{z}) denote the estimated bitrate for y^\hat{y} and z^\hat{z}, D(x,x^)D(x, \hat{x}) denotes distortion measured by Mean Squared Error (MSE) or Multiscale Structural Similarity (MS-SSIM), and λ\lambda controls the rate-distortion trade-off.

  2. Knowl 2 — Frequency-Decomposed Window Attention

    model/method

    Frequency-Decomposition Window Attention (FDWA) decomposes hidden representations into multiscale isotropic and anisotropic directional frequency components within a single attention layer. Standard self-attention acts as a low-pass filter, where smaller windows extract fine-grained high frequencies and larger windows extract low frequencies. Square windows, however, are isotropic and fail to capture directional orientations.

    Given an intermediate feature tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W} with CC channels and spatial dimensions H×WH \times W, the linear projection maps XX into KK attention heads. The KK heads are partitioned into four equal groups of K/4K/4 heads, each executing local window self-attention using a distinct window geometry (height ×\times width) based on a base scale parameter ss:

    1. Low-Frequency Window Attention (LL-WA): Window size 4s×4s4s \times 4s, capturing low-frequency isotropic features.
    2. High-Frequency Window Attention (HH-WA): Window size s×ss \times s, capturing fine-grained high-frequency isotropic details.
    3. Vertical Frequency Window Attention (HL-WA): Window size s×4ss \times 4s, capturing vertically oriented high frequencies.
    4. Horizontal Frequency Window Attention (LH-WA): Window size 4s×s4s \times s, capturing horizontally oriented high frequencies.

    For a specific window attention type (e.g., LH-WA), the feature XX is partitioned into M=H×W4s×sM = \frac{H \times W}{4s \times s} non-overlapping windows [X1,…,XM][X^1, \dots, X^M] with Xi∈RC×4s×sX^i \in \mathbb{R}^{C \times 4s \times s}. For head kk with head dimension dk=C/Kd_k = C/K and projection matrices WkQ,WkK,WkV∈RC×dkW_k^Q, W_k^K, W_k^V \in \mathbb{R}^{C \times d_k}:

    Yki=Attention(XiWkQ,XiWkK,XiWkV)Y_k^i = \text{Attention}(X^i W_k^Q, X^i W_k^K, X^i W_k^V)

    LH-WAk(X)=[Yk1,Yk2,…,YkM]\text{LH-WA}_k(X) = [Y_k^1, Y_k^2, \dots, Y_k^M]

    The outputs of all KK heads across the four groups are concatenated and linearly projected via WO∈RC×CW^O \in \mathbb{R}^{C \times C} to facilitate cross-frequency interaction:

    FDWA(X)=Concat[head1,…,headK]WO\text{FDWA}(X) = \text{Concat}[\text{head}_1, \dots, \text{head}_K] W^O

    headk={LL-WAk(X),1≤k≤K/4HH-WAk(X),K/4<k≤K/2HL-WAk(X),K/2<k≤3K/4LH-WAk(X),3K/4<k≤K\text{head}_k = \begin{cases} \text{LL-WA}_k(X), & 1 \le k \le K/4 \\ \text{HH-WA}_k(X), & K/4 < k \le K/2 \\ \text{HL-WA}_k(X), & K/2 < k \le 3K/4 \\ \text{LH-WA}_k(X), & 3K/4 < k \le K \end{cases}

    In consecutive FAT blocks, the first block applies regular FDWA and the second applies shifted-window FDWA, where windows are shifted by half their respective height and width dimensions.

  3. Knowl 3 — Frequency-Modulation Feed-Forward Network

    model/method

    The Frequency-Modulation Feed-Forward Network (FMFFN) adaptively modulates and filters frequency components produced by self-attention, dynamically amplifying or suppressing specific subbands according to target bitrates.

    Given an input feature representation XX, FMFFN first applies standard pointwise convolutions with a GELU activation to produce an expanded intermediate feature XffnX_{\text{ffn}}:

    Xffn=Conv1×1(GELU(Conv1×1(X)))X_{\text{ffn}} = \text{Conv}_{1\times 1}(\text{GELU}(\text{Conv}_{1\times 1}(X)))

    To make frequency modulation independent of varying input image resolutions, a spatial block partitioning operator B(⋅)B(\cdot) partitions XffnX_{\text{ffn}} into non-overlapping blocks of size 4s×4s4s \times 4s (matching the maximum window size in FDWA). A 2D Fast Fourier Transform F(⋅)\mathcal{F}(\cdot) converts each block into the frequency domain, where it undergoes element-wise multiplication ⊙\odot with a learnable filter tensor WW of fixed size 4s×4s4s \times 4s across channels:

    Xfm=F(B(Xffn))⊙WX_{\text{fm}} = \mathcal{F}(B(X_{\text{ffn}})) \odot W

    The modulated frequency representation XfmX_{\text{fm}} is converted back to the spatial domain using 2D Inverse Fast Fourier Transform F−1(⋅)\mathcal{F}^{-1}(\cdot), and non-overlapping blocks are merged back to the original feature map shape via B−1(⋅)B^{-1}(\cdot):

    Xout=B−1(F−1(Xfm))X_{\text{out}} = B^{-1}\left(\mathcal{F}^{-1}(X_{\text{fm}})\right)

  4. Knowl 4 — Transformer-Based Channel-Wise Autoregressive Entropy Model

    model/method

    The Transformer-based Channel-wise Autoregressive (T-CA) entropy model captures dynamic inter-slice and intra-slice channel correlations using causal masked channel self-attention rather than fixed-weight CNNs.

    The quantized latent representation y^∈RM×H×W\hat{y} \in \mathbb{R}^{M \times H \times W} is partitioned along channels into nsn_s slices {y^1,…,y^ns}\{\hat{y}_1, \dots, \hat{y}_{n_s}\}, where each slice contains Ms=M/nsM_s = M / n_s channels. Each slice is linearly expanded to r⋅Msr \cdot M_s channels (where r>1r > 1 is a projection ratio) using a 1×11 \times 1 group convolution with ng=nsn_g = n_s groups.

    The transformed slices are processed through LL consecutive transformer layers structured as follows:

    1. Masked-Slice Channel Attention: Self-attention is performed across the channel dimension (C=r⋅MC = r \cdot M). To enforce causal decoding, an upper-triangular slice-wise mask matrix Mmask∈RC×CM_{\text{mask}} \in \mathbb{R}^{C \times C} is applied to the channel attention maps, preventing tokens in subsequent slices from attending to future latent slices. A learnable pseudo start slice ysy_s provides context for the initial slice y^1\hat{y}_1.
    2. Group-wise Operations: Standard LayerNorm and linear projections are replaced with GroupNorm and 1×11 \times 1 group convolutions with ng=nsn_g = n_s groups, isolating channel mixing per slice and reducing computational parameter overhead.

    The final layer output yout∈R4M×H×Wy_{\text{out}} \in \mathbb{R}^{4M \times H \times W} is reshaped to RM×4×H×W\mathbb{R}^{M \times 4 \times H \times W} and concatenated with the reshaped hyperprior context ϕ∈RM×2×H×W\phi \in \mathbb{R}^{M \times 2 \times H \times W} to form yconcat∈R6M×H×Wy_{\text{concat}} \in \mathbb{R}^{6M \times H \times W}. This concatenated tensor is passed into an Entropy Parameters Network composed of three 3×33 \times 3 group convolutional layers (ng=nsn_g = n_s) with GELU activations to output the Gaussian distribution parameters μi,σi\mu_i, \sigma_i and quantization residual rir_i for each slice ii.

  5. Knowl 5 — Two-Stage Training Strategy for Autoregressive Learned Image Compression

    algorithm

    In single-Gaussian LIC models, encoding ⌈y−μ⌋\lceil y - \mu \rfloor rather than ⌈y⌋\lceil y \rfloor provides performance advantages. However, in autoregressive models, predicting μi\mu_i requires prior quantized latent slices y^<i\hat{y}_{<i}, leading to train-test mismatch if trained directly on centered representations. FTIC resolves this via a two-stage training strategy.

    Input: Training datasets Flickr2W and ImageNet-1k, Adam optimizer, loss function weights λ\lambda
    Output: Fully converged FTIC model parameters
    # Stage 1: Pretraining transforms with hyperprior only
    Initialize analysis transform gag_a, synthesis transform gsg_s, hyper-transforms ha,hsh_a, h_s
    Set learning rate η=10−4\eta = 10^{-4}, crop size =256×256= 256 \times 256, batch size =8= 8
    for step =1= 1 to 2,000,0002{,}000{,}000 do
        Sample batch xx
        Compute latents y=ga(x)y = g_a(x), z=ha(y)z = h_a(y), z^=Q(z)\hat{z} = Q(z), ϕ=hs(z^)\phi = h_s(\hat{z})
        Predict mean μ\mu and scale σ\sigma solely from hyperprior context ϕ\phi
        Encode symbols as ⌈y−μ⌋\lceil y - \mu \rfloor and restore as y^=⌈y−μ⌋+μ\hat{y} = \lceil y - \mu \rfloor + \mu
        Compute reconstructed image x^=gs(y^)\hat{x} = g_s(\hat{y})
        Update ga,gs,ha,hsg_a, g_s, h_a, h_s by minimizing R(y^)+R(z^)+λD(x,x^)R(\hat{y}) + R(\hat{z}) + \lambda D(x, \hat{x})
    end for
    # Stage 2: Joint autoregressive fine-tuning
    Load trained weights of ga,gs,ha,hsg_a, g_s, h_a, h_s from Stage 1
    Initialize T-CA entropy model randomly
    Set learning rate η=10−4\eta = 10^{-4}, crop size =256×256= 256 \times 256, batch size =8= 8
    for step =1= 1 to 1,000,0001{,}000{,}000 do
        Sample batch xx
        Compute latents y=ga(x)y = g_a(x), y^=Q(y)\hat{y} = Q(y), z=ha(y)z = h_a(y), z^=Q(z)\hat{z} = Q(z), ϕ=hs(z^)\phi = h_s(\hat{z})
        Predict slice parameters μi,σi,ri=T-CA(ϕ,y^<i)\mu_i, \sigma_i, r_i = \text{T-CA}(\phi, \hat{y}_{<i})
        Directly encode ⌈y⌋\lceil y \rfloor without mean subtraction; restore yˉ=y^+r\bar{y} = \hat{y} + r
        Compute reconstructed image x^=gs(yˉ)\hat{x} = g_s(\bar{y})
        Update all parameters by minimizing R(y^)+R(z^)+λD(x,x^)R(\hat{y}) + R(\hat{z}) + \lambda D(x, \hat{x})
    end for
    # Stage 2 Fine-tuning on larger crop size
    Set learning rate η=10−5\eta = 10^{-5}, crop size =384×384= 384 \times 384
    for step =1= 1 to 200,000200{,}000 do
        Execute joint autoregressive optimization step with direct ⌈y⌋\lceil y \rfloor encoding
    end for
  6. Knowl 6 — Rate-Distortion Compression Performance Across Standard Benchmarks

    empirical result

    The FTIC compression model was evaluated on three standard benchmark datasets: Kodak (24 images of resolution 768×512768 \times 512), Tecnick (100 images of resolution 1200×12001200 \times 1200), and the CLIC Professional Validation dataset (41 images up to 2K resolution). Performance was assessed using PSNR and MS-SSIM against bitrates measured in bits per pixel (BPP), with BD-rate computed relative to the standardized video codec VVC test model anchor VTM-12.1 (where negative BD-rate indicates bitrate savings at equal reconstruction quality).

    FTIC achieved state-of-the-art BD-rate improvements over VTM-12.1 across all evaluated datasets:

    • Kodak dataset: −14.5%-14.5\% BD-rate saving relative to VTM-12.1.
    • Tecnick dataset: −15.1%-15.1\% BD-rate saving relative to VTM-12.1.
    • CLIC Professional Validation dataset: −13.0%-13.0\% BD-rate saving relative to VTM-12.1.

    FTIC consistently outperformed prior learned image compression methods across all operational bitrates, including transformer-based architectures (Liu et al., 2023; Zou et al., 2022; Zhu et al., 2022) and CNN-based autoregressive models (He et al., 2022; Cheng et al., 2020; Minnen et al., 2018).

  7. Knowl 7 — Ablation Analysis of FAT Block Components

    data/table

    Ablation studies on the Kodak dataset evaluated the architectural contributions of the Frequency-Decomposition Window Attention (FDWA) and Frequency-Modulation Feed-Forward Network (FMFFN) against a baseline utilizing standard 4×44 \times 4 isotropic window self-attention with a conventional FFN. Computational complexity was quantified by Multiply-Accumulate Operations (GMACs) for the analysis transform ga(⋅)g_a(\cdot), overall parameter counts, and BD-rate savings relative to VTM-12.1.

    Methods Larger Window FDWA FMFFN GMACs #Params BD-rate
    Baseline 77.3 70.29M -9.3%
    Variant 1 ✓ 91.3 70.50M -11.2%
    Variant 2 ✓ 82.2 70.36M -13.3%
    Variant 3 (FAT Block) ✓ ✓ 83.1 70.97M -14.5%
    VTM-12.1 - - - - - 0%

    Expanding window size to 16×1616 \times 16 (Variant 1) improved BD-rate by 1.9%1.9\% over Baseline but increased GMACs by 18.1%18.1\%. Introducing FDWA (Variant 2) achieved a larger gain of 4.0%4.0\% BD-rate saving over Baseline with only a modest GMAC increase (82.2 vs. 77.3 GMACs), demonstrating that frequency-decomposed isotropic and anisotropic window geometry is more effective than scaling isotropic receptive fields. Adding FMFFN (Variant 3) yielded an additional 1.2%1.2\% BD-rate improvement (-14.5% total) with negligible parameter and computational increase.

  8. Knowl 8 — Parameter and Architectural Ablation of T-CA Entropy Model

    data/table

    The Transformer-based Channel-wise Autoregressive (T-CA) entropy model was evaluated against the CNN-based channel-wise autoregressive baseline CHARM (Minnen & Singh, 2020) and across various parameterizations of transformer layers LL and slice counts nsn_s. Models were evaluated on the Kodak dataset using pretrained transforms from mbt2018-mean and fine-tuned for 1M batches, with BPG codec used as the BD-rate anchor.

    Model #Params BD-rate
    CHARM (ns=5n_s=5) 18.3M -14.2%
    CHARM (ns=10n_s=10) 34.8M -15.9%
    T-CA (ns=5n_s=5) 30.4M -19.2%
    BPG - 0%
    LL nsn_s #Params BD-rate
    4 5 14.5M -17.6%
    8 5 22.4M -18.3%
    12 5 30.4M -19.2%
    16 5 38.4M -19.0%
    12 4 37.8M -19.0%
    12 5 30.4M -19.2%
    12 8 19.3M -18.0%
    12 10 15.6M -18.4%
    BPG - - 0%

    T-CA with 5 slices outperformed 10-slice CHARM by 3.3%3.3\% in BD-rate while requiring fewer parameters (30.4M vs 34.8M). Varying the number of transformer layers showed that performance peaked at L=12L=12 (-19.2%), with L=16L=16 providing no additional benefit. Unlike CHARM, where parameter counts grow linearly with slice count nsn_s, increasing nsn_s in T-CA reduces the parameter count of group convolutions and the overall entropy model due to shared transformer weights, with ns=5n_s=5 achieving optimal rate-distortion efficiency.

  9. Knowl 9 — Coding Complexity and Latency Benchmark

    data/table

    Computational complexity and inference latency were benchmarked on the Kodak dataset using an NVIDIA GeForce RTX 4090 GPU (24 GB memory). Inference latency represents model forward GPU runtime excluding CPU entropy arithmetic encoding/decoding. BD-rate is calculated against the VTM-12.1 anchor.

    Model GMACs Inference Latency (ms) #Params BD-rate
    Enc. Dec. Enc. Dec.
    Cheng et al. (2020) 154 229 >1000 >1000 26.60M 3.6%
    Minnen Singh (2020) 101 100 56 43 55.13M 1.1%
    Zhu et al. (2022) 116 116 110 99 32.71M -3.3%
    Zou et al. (2022) 143 161 97 101 99.58M -4.3%
    Liu et al. (2023) 317 453 255 322 76.57M -11.9%
    Ours (FTIC) 141 349 125 242 70.97M -14.5%
    VTM-12.1 - - - - - 0%

    While spatial autoregressive methods (Cheng et al., 2020) suffer from latencies exceeding 1000 ms, FTIC achieves the highest compression performance (-14.5% BD-rate) with an encoding latency of 125 ms and decoding latency of 242 ms, offering substantially lower computational complexity (141 Enc / 349 Dec GMACs vs 317 Enc / 453 Dec GMACs) and faster inference than prior SOTA transformer-CNN hybrid models (Liu et al., 2023).

  10. Knowl 10 — Limitations of Directional Decomposition and Scope in FTIC

    limitation

    The Frequency-Aware Transformer architecture exhibits two primary limitations:

    1. Restricted Directional Decompositions: The Frequency-Decomposition Window Attention (FDWA) module is restricted to two orthogonal directions—horizontal (4s×s4s \times s) and vertical (s×4ss \times 4s). It lacks support for diagonal, curvilinear, or arbitrary angular directional frequency subbands, which limits its ability to fully capture complex directional textures compared to general multiscale geometric analysis frameworks like scattering transforms.
    2. Scope Restricted to Still Image Compression: The FTIC formulation is designed exclusively for 2D spatial image compression and does not incorporate temporal frequency decomposition or inter-frame motion compensation mechanisms required for video compression.

Coverage note — None was omitted; all key contributions—including the FTIC overall architecture, FDWA, FMFFN, T-CA entropy model, two-stage training scheme, rate-distortion benchmarks, detailed ablations, and stated limitations—are fully covered.

References

  1. 1.Nicola Asuni and Andrea Giachetti. Testimages: a large-scale archive for testing visual devices and basic image processing algorithms. In STAG, pp. 63–70, 2014.
  2. 2.Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In Proceedings of the International Conference on Learning Representations, 2018.
  3. 3.Jean Bégaint, Fabien Racapé, Simon Feltman, and Akshay Pushparaja. Compressai: a pytorch library and evaluation platform for end-to-end compression research. arXiv preprint arXiv:2011.03029, 2020.
  4. 4.Gisle Bjontegaard. Calculation of average psnr differences between rd-curves. In VCEG-M33, 2001.
  5. 5.Emmanuel Jean Candès and David Leigh Donoho. Curvelets: A surprisingly effective nonadaptive representation for objects with edges. Vanderbilt Univ. Press, 2000.
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pp. 213–229. Springer, 2020.
  7. 7.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7939–7948, 2020.
  8. 8.CLIC. Workshop and challenge on learned image compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. pp. 248–255, 2009.
  10. 10.Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, and Hongkai Xiong. Motion-aware contrastive video representation learning via foregroundbackground merging. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9716–9726, 2022a.
  11. 11.Shuangrui Ding, Rui Qian, and Hongkai Xiong. Dual contrastive learning for spatio-temporal representation. In Proceedings of the 30th ACM international conference on multimedia, pp. 5649–5658, 2022b.
  12. 12.M.N. Do and M. Vetterli. The contourlet transform: an efficient directional multiresolution image representation. IEEE Transactions on Image Processing, 14(12):2091–2106, 2005.
  13. 13.David L. Donoho and Xiaoming Huo. Beamlets and multiscale image analysis. In Multiscale and Multiresolution Methods, pp. 149–196, Berlin, Heidelberg, 2002. Springer Berlin Heidelberg.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  15. 15.Haisheng Fu, Feng Liang, Jie Liang, Binglin Li, Guohe Zhang, and Jingning Han. Asymmetric learned image compression with multi-scale residual block, importance scaling, and postquantization filtering. IEEE Transactions on Circuits and Systems for Video Technology, 2023.
  16. 16.Ge Gao, Pei You, Rong Pan, Shunyuan Han, Yuanyuan Zhang, Yuchao Dai, and Hojae Lee. Neural image compression via attentional multi-scale back projection and frequency decomposition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14677–14686, 2021.
  17. 17.Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  18. 18.Eastman Kodak. Kodak lossless true color image suite (photocd pcd0992). 1993.
  19. 19.A Burakhan Koyuncu, Han Gao, and Eckehard Steinbach. Contextformer: A transformer with spatio-channel attention for context modeling in learned image compression. In European Conference on Computer Vision, 2022.
  20. 20.Han Li, Bowen Shi, Wenrui Dai, Hongwei Zheng, Botao Wang, Yu Sun, Min Guo, Chenglin Li, Junni Zou, and Hongkai Xiong. Pose-oriented transformer with uncertainty-guided refinement for 2d-to-3d human pose estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 1296–1304, 2023.
  21. 21.Jiaheng Liu, Guo Lu, Zhihao Hu, and Dong Xu. A unified end-to-end framework for efficient deep image compression. arXiv preprint arXiv:2002.03370, 2020.
  22. 22.Jinming Liu, Heming Sun, and Jiro Katto. Learned image compression with mixed transformercnn architectures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14388–14397, 2023.
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022, 2021.
  24. 24.Ming Lu, Peiyao Guo, Huiqing Shi, Chuntong Cao, and Zhan Ma. Transformer-based image compression. In 2022 Data Compression Conference (DCC), pp. 469–469. IEEE, 2022.
  25. 25.Haichuan Ma, Dong Liu, Ruiqin Xiong, and Feng Wu. iwave: Cnn-based wavelet-like transform for image compression. IEEE Transactions on Multimedia, 22(7):1667–1679, 2019.
  26. 26.Haichuan Ma, Dong Liu, Ning Yan, Houqiang Li, and Feng Wu. End-to-end optimized versatile image compression with wavelet-like transform. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(3):1247–1263, 2020.
  27. 27.Stephane G Mallat. A theory for multiresolution signal decomposition: the wavelet representation. IEEE transactions on pattern analysis and machine intelligence, 11(7):674–693, 1989.
  28. 28.Fabian Mentzer, George Toderici, David Minnen, Sung-Jin Hwang, Sergi Caelles, Mario Lucic, and Eirikur Agustsson. Vct: A video compression transformer. arXiv preprint arXiv:2206.07307, 2022.
  29. 29.David Minnen and Saurabh Singh. Channel-wise autoregressive entropy models for learned image compression. In IEEE International Conference on Image Processing (ICIP), pp. 3339–3343. IEEE, 2020.
  30. 30.David Minnen, Johannes Ballé, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems, 31, 2018.
  31. 31.Zizheng Pan, Jianfei Cai, and Bohan Zhuang. Fast vision transformers with hilo attention. Advances in Neural Information Processing Systems, 35:14541–14554, 2022.
  32. 32.Namuk Park and Songkuk Kim. How do vision transformers work? In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=IDwN6xjHnK8.
  33. 33.Badri Patro and Vijay Agneeswaran. Scattering vision transformer: Spectral mixing matters. Advances in Neural Information Processing Systems, 36, 2024.
  34. 34.Yichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan, and Rong Jin. Entroformer: A transformer-based entropy model for learned image compression. In International Conference on Learning Representations, 2022.
  35. 35.Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou. Global filter networks for image classification. Advances in neural information processing systems, 34:980–993, 2021.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  37. 37.Yueqi Xie, Ka Leong Cheng, and Qifeng Chen. Enhanced invertible encoding for learned image compression. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 162–170, 2021.
  38. 38.Ali Zafari, Atefeh Khoshkhahtinat, Piyush Mehta, Mohammad Saeed Ebrahimi Saadabadi, Mohammad Akyash, and Nasser M Nasrabadi. Frequency disentangled features in neural image compression. In 2023 IEEE International Conference on Image Processing (ICIP), pp. 2815–2819. IEEE, 2023.
  39. 39.Yinhao Zhu, Yang Yang, and Taco Cohen. Transformer-based transform coding. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=IDwN6xjHnK8.
  40. 40.Renjie Zou, Chunfeng Song, and Zhaoxiang Zhang. The devil is in the details: Window-based attention for image compression. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022.

Citation

MLA
Li, H., et al. “Frequency-Aware Transformer for Learned Image Compression”. arXiv, 2023, http://arxiv.org/abs/2310.16387v4.
APA
Li, H., Li, S., Dai, W., Li, C., Zou, J., & Xiong, H. (2023). Frequency-Aware Transformer for Learned Image Compression. arXiv. http://arxiv.org/abs/2310.16387v4
Chicago
Li, H., S. Li, W. Dai, C. Li, J. Zou, and H. Xiong. 2023. “Frequency-Aware Transformer for Learned Image Compression”. arXiv. http://arxiv.org/abs/2310.16387v4.
Harvard
Li, H. et al. (2023) “Frequency-Aware Transformer for Learned Image Compression”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.16387v4.
Vancouver
1. Li H, Li S, Dai W, Li C, Zou J, Xiong H (2023) Frequency-Aware Transformer for Learned Image Compression. arXiv

BibTeX

@article{li2023frequency,
  title = {Frequency-Aware Transformer for Learned Image Compression},
  author = {Li, Han and Li, Shaohui and Dai, Wenrui and Li, Chenglin and Zou, Junni and Xiong, Hongkai},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.16387v4},
  eprint = {2310.16387}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors