FFT-Based Dynamic Token Mixer for Vision

Yuki TatsunamiMasato Taki

article2024AAAI69 citations

Proposes an efficient FFT-based dynamic filter token-mixer within the MetaFormer framework to deliver global receptive field processing with substantially lower computational complexity and higher throughput than multi-head self-attention on high-resolution vision tasks.

Listen

Attention-based vision architectures deliver state-of-the-art accuracy in image recognition, but their computational complexity grows quadratically with input size. This creates severe performance bottlenecks and high memory consumption when handling high-resolution images or dense tasks like semantic segmentation. Fast Fourier Transform (FFT) based models offer a promising alternative by capturing global information at lower theoretical computational complexity, but prior versions relied on static, data-independent filters and lagged behind modern architecture standards.

The article demonstrates a novel dynamic token-mixing approach called the Dynamic Filter, along with two architecture families: DFFormer and the hybrid CDFFormer (combining standard convolutions with FFT blocks). The primary objective is to evaluate whether dynamically generating frequency-domain filters can match modern vision model accuracy while maintaining high processing speeds and low memory usage at high image resolutions.

The authors implemented their models within the modern MetaFormer framework and evaluated performance using standard computer vision benchmarks. Image classification was tested on the ImageNet-1K dataset (over 1.28 million training images), while semantic segmentation was evaluated on ADE20K. The models were benchmarked across standard configurations against leading convolutional, attention-based, and frequency-based vision models, with additional ablation and representational analyses to isolate architectural contributions.

The evaluation produced four key findings. First, DFFormer and CDFFormer established new performance highs among attention-free and FFT-based architectures, with DFFormer-B36 achieving an 84.8% top-1 accuracy on ImageNet-1K (outperforming prior FFT models by over 0.5%) and CDFFormer-B36 reaching 85.0%. Second, the architecture showed strong downstream effectiveness on ADE20K semantic segmentation, where DFFormer-S36 achieved 47.5% mean Intersection over Union (mIoU), outperforming comparable baselines by 5.5 points. Third, when scaling to high resolutions (from 256x256 up to 1024x1024 pixels), DFFormer and CDFFormer maintained high throughput and low peak memory, whereas attention-based models experienced severe throughput collapse and steep memory surges. Fourth, representational analysis revealed that unlike attention modules which act as both high- and low-pass filters, dynamic filters primarily suppress high frequencies to learn distinct low-frequency global representations.

These findings imply that dynamic FFT-based architectures are highly cost-effective alternatives to attention mechanisms for vision applications where latency, memory, and resolution are critical operational constraints. By avoiding quadratic computational scaling, these models allow real-time processing and dense prediction on resource-constrained hardware without sacrificing competitive accuracy.

Organizations developing high-resolution vision systems, such as medical imaging or autonomous segmentation pipelines, should consider adopting dynamic filter architectures to lower compute costs and memory footprints. Teams should benchmark DFFormer and CDFFormer against existing attention-based baselines in pilot deployments to determine the best trade-off between convolutional hybrid features and pure FFT blocks.

The primary limitation noted in the article is that at standard low resolutions (such as 224x224), implementation-level overheads mean FFT throughput can still trail optimized attention baselines on certain hardware. Because overall speed depends on specific hardware architectures and low-level library implementations, practitioners should validate inference throughput within their own deployment environments before broad production rollout.

Cover for FFT-Based Dynamic Token Mixer for Vision

Abstract

Multi-head-self-attention (MHSA)-equipped models have achieved notable performance in computer vision. Their computational complexity is proportional to quadratic numbers of pixels in input feature maps, resulting in slow processing, especially when dealing with high-resolution images. New types of token-mixer are proposed as an alternative to MHSA to circumvent this problem: an FFT-based token-mixer involves global operations similar to MHSA but with lower computational complexity. However, despite its attractive properties, the FFT-based token-mixer has not been carefully examined in terms of its compatibility with the rapidly evolving MetaFormer architecture. Here, we propose a novel token-mixer called Dynamic Filter and novel image recognition models, DFFormer and CDFFFormer, to close the gaps above. The results of image classification and downstream tasks, analysis, and visualization show that our models are helpful. Notably, their throughput and memory efficiency when dealing with high-resolution image recognition is remarkable. Our results indicate that Dynamic Filter is one of the token-mixer options that should be seriously considered. The code is available at https://github.com/okojoalg/dfformer

Table of Contents

  • Introduction
  • Related Work
  • Method
  • Preliminary: Global Filter
  • Dynamic Filter
  • DFFormer and CDFFormer
  • Experiments
  • Image Classification
  • Semantic Segmentation on ADE20K
  • Ablation Studies
  • Analysis
  • Advantages at Higher Resolutions
  • Representational Similarities
  • Analysis of Dynamic Filter Basis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Dynamic Filter Token Mixer

    model/method

    The Dynamic Filter is an FFT-based token-mixing module that dynamically determines spatial frequency filter weights conditioned on input content. For an input feature tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, a 2D real Fast Fourier Transform (2D-RFFT) F\mathcal{F} converts spatial features into the frequency domain CC×H×⌈W/2⌉\mathbb{C}^{C \times H \times \lceil W/2 \rceil}, where redundant Hermitian components are omitted.

    Instead of learning a static complex filter per channel, the Dynamic Filter maintains a shared basis of NN complex frequency filter matrices K={K1,…,KN}\mathbb{K} = \{K_1, \dots, K_N\}, where each Ki∈CH×⌈W/2⌉K_i \in \mathbb{C}^{H \times \lceil W/2 \rceil} (default N=4N = 4).

    The input tensor is spatially pooled via global average pooling:

    Xˉ=1HW∑h=1H∑w=1WX:,h,w∈RC\bar{X} = \frac{1}{HW} \sum_{h=1}^H \sum_{w=1}^W X_{:, h, w} \in \mathbb{R}^C

    A Multi-Layer Perceptron MM predicts dynamic weighting coefficients for each of the C′C' channels (where C′=2CC' = 2C in the expanded block configuration):

    (s1,…,sNC′)⊤=M(Xˉ)=W2 StarReLU(W1 LN(Xˉ))(s_1, \dots, s_{NC'})^\top = M(\bar{X}) = W_2 \, \text{StarReLU}(W_1 \, \text{LN}(\bar{X}))

    where LN\text{LN} denotes Layer Normalization, W1∈Rint(ρC)×CW_1 \in \mathbb{R}^{\text{int}(\rho C) \times C}, W2∈RNC′×int(ρC)W_2 \in \mathbb{R}^{NC' \times \text{int}(\rho C)} with intermediate dimension ratio ρ=0.25\rho = 0.25, and StarReLU(x)=x⋅max⁡(0,x)\text{StarReLU}(x) = x \cdot \max(0, x).

    The dynamic complex filter KM(X)∈CC′×H×⌈W/2⌉K_M(X) \in \mathbb{C}^{C' \times H \times \lceil W/2 \rceil} is constructed per channel c∈{1,…,C′}c \in \{1, \dots, C'\} via a channel-wise softmax combination over the NN basis filters:

    KM(X)c,:,::=∑i=1N(es(c−1)N+i∑n=1Nes(c−1)N+n)KiK_M(X)_{c, :, :} := \sum_{i=1}^N \left( \frac{e^{s_{(c-1)N + i}}}{\sum_{n=1}^N e^{s_{(c-1)N + n}}} \right) K_i

    For an intermediate transformed feature A(X)A(X) (with continuous real mapping AA), the Dynamic Filter transformation D(X)\mathcal{D}(X) is defined as:

    D(X)=F−1(KM(X)⊙F(A(X)))\mathcal{D}(X) = \mathcal{F}^{-1}(K_M(X) \odot \mathcal{F}(A(X)))

    where ⊙\odot denotes the element-wise Hadamard product in the complex domain, and F−1\mathcal{F}^{-1} is the 2D inverse real Fast Fourier Transform (2D-IRFFT).

  2. Knowl 2 — DFFormer and CDFFormer Architectural Family

    model/method

    DFFormer and CDFFormer are four-stage hierarchical vision architectures adhering to the MetaFormer macro-architecture. Each block transforms input X∈RC×H×WX \in \mathbb{R}^{C \times H \times W} according to an inverted bottleneck residual formulation:

    T(X)=X+Convpw2(L(StarReLU(Convpw1(LN(X)))))\mathcal{T}(X) = X + \text{Conv}_{pw2} \left( \mathcal{L}\left( \text{StarReLU}\left( \text{Conv}_{pw1}(\text{LN}(X)) \right) \right) \right)

    where LN\text{LN} is Layer Normalization, Convpw1\text{Conv}_{pw1} is a point-wise convolution projecting channels C→C′=2CC \to C'=2C, L(⋅)\mathcal{L}(\cdot) is a token mixer, and Convpw2\text{Conv}_{pw2} is a point-wise convolution projecting channels 2C→C2C \to C.

    In a DFFormer block (DF), the token mixer is L(Y)=F−1(KM(X)⊙F(Y))\mathcal{L}(Y) = \mathcal{F}^{-1}(K_M(X) \odot \mathcal{F}(Y)), applying a content-dependent dynamic Fourier filter. In a ConvFormer block (CF), L(Y)\mathcal{L}(Y) is depthwise separable convolution.

    Two architectural configurations are defined based on block composition:

    • DFFormer: Employs DFFormer blocks across all four stages [DF, DF, DF, DF].
    • CDFFormer: A hybrid architecture employing ConvFormer blocks in stages 1 and 2 and DFFormer blocks in stages 3 and 4 [CF, CF, DF, DF].

    Downsampling layers between stages use patch-embedding convolutions with kernel size KK, stride SS, and padding PP configured across stages as K=(7,3,3,3)K=(7, 3, 3, 3), S=(4,2,2,2)S=(4, 2, 2, 2), and P=(2,1,1,1)P=(2, 1, 1, 1). Model size variants are parameterized by block counts L=(L1,L2,L3,L4)L = (L_1, L_2, L_3, L_4) and stage channel dimensions C=(C1,C2,C3,C4)C = (C_1, C_2, C_3, C_4):

    • S18: L=(3,3,9,3)L = (3, 3, 9, 3), C=(64,128,320,512)C = (64, 128, 320, 512)
    • S36: L=(3,12,18,3)L = (3, 12, 18, 3), C=(64,128,320,512)C = (64, 128, 320, 512)
    • M36: L=(3,12,18,3)L = (3, 12, 18, 3), C=(96,192,384,576)C = (96, 192, 384, 576)
    • B36: L=(3,12,18,3)L = (3, 12, 18, 3), C=(128,256,512,768)C = (128, 256, 512, 768)

    The classifier head comprises Global Average Pooling (GAP), Layer Normalization, and an MLP classifier with Squared ReLU (extReLU2 ext{ReLU}^2) activation.

  3. Knowl 3 — ImageNet-1K Visual Classification Performance

    data/table

    ImageNet-1K classification performance of DFFormer and CDFFormer models compared against convolutional (C), attention-based (A), retention-based (R), MLP-based (M), FFT-based (F), and hybrid (CA, CF) architectures. All models were trained for 300 epochs at 224×224224 \times 224 resolution, and inference throughput was benchmarked on a single NVIDIA V100 GPU (16GB memory) at batch size 16.

    Model Type Params (M) FLOPs (G) Throughput (img/s) Top-1 (%)
    ConvNeXt-T C 29 4.5 1471 82.1
    ConvFormer-S18 C 27 3.9 756 83.0
    CSWin-T A 23 4.3 340 82.7
    MViTv2-T A 24 4.7 624 82.3
    DiNAT-T A 28 4.3 816 82.7
    DaViT-Tiny A 28 4.5 1121 82.8
    GCViT-T A 28 4.7 566 83.5
    MaxViT-T A 41 5.6 527 83.6
    RMT-S R 27 4.5 406 84.1
    AMixer-T M 26 4.5 724 82.0
    GFNet-H-S F 32 4.6 952 81.5
    DFFormer-S18 F 30 3.8 535 83.2
    CAFormer-S18 CA 26 4.1 741 83.6
    CDFFormer-S18 CF 30 3.9 567 83.1
    ConvNeXt-S C 50 8.7 864 83.1
    ConvFormer-S36 C 40 7.6 398 84.1
    MViTv2-S A 35 7.0 416 83.6
    DiNAT-S A 51 7.8 688 83.8
    DaViT-Small A 50 8.8 664 84.2
    GCViT-S A 51 8.5 478 84.3
    RMT-B R 54 9.7 264 85.0
    DynaMixer-S M 26 7.3 448 82.7
    AMixer-S M 46 9.0 378 83.5
    GFNet-H-B F 54 8.6 612 82.9
    DFFormer-S36 F 46 7.4 270 84.3
    CAFormer-S36 CA 39 8.0 382 84.5
    CDFFormer-S36 CF 45 7.5 319 84.2
    ConvNeXt-B C 89 15.4 687 83.8
    ConvFormer-M36 C 57 12.8 307 84.5
    MViTv2-B A 52 10.2 285 84.4
    GCViT-S2 A 68 10.7 415 84.8
    MaxViT-S A 69 11.7 449 84.5
    DynaMixer-M M 57 17.0 317 83.7
    AMixer-B M 83 16.0 325 84.0
    DFFormer-M36 F 65 12.5 210 84.6
    CAFormer-M36 CA 56 13.2 297 85.2
    CDFFormer-M36 CF 64 12.7 254 84.8
    ConvNeXt-L C 198 34.4 431 84.3
    ConvFormer-B36 C 100 22.6 235 84.8
    MViTv2-L A 218 42.1 128 85.3
    DiNAT-B A 90 13.7 499 84.4
    DaViT-Base A 88 15.5 528 84.6
    GCViT-B A 90 14.8 367 85.0
    MaxViT-B A 120 23.4 224 85.0
    RMT-L R 95 18.2 233 85.5
    DynaMixer-L M 97 27.4 216 84.3
    DFFormer-B36 F 115 22.1 161 84.8
    CAFormer-B36 CA 99 23.2 227 85.5
    CDFFormer-B36 CF 113 22.5 195 85.0

    DFFormer-S18 outperforms the static FFT baseline GFNet-H-S by 1.7% top-1 accuracy (83.2% vs. 81.5%) with fewer FLOPs (3.8G vs. 4.6G). DFFormer-B36 achieves 84.8% top-1 accuracy, and the hybrid CDFFormer-B36 reaches 85.0% top-1 accuracy, exceeding purely convolutional ConvNeXt and ConvFormer models while matching or approaching multi-head self-attention models.

  4. Knowl 4 — Resolution Scaling: Throughput and Memory Efficiency

    empirical result

    The theoretical computational complexity of Multi-Head Self-Attention (MHSA) is O(HWC2+(HW)2C)\mathcal{O}(HWC^2 + (HW)^2C), scaling quadratically with the spatial token resolution HWHW. In contrast, global and dynamic FFT-based filter operations have a theoretical complexity of O(HWC⌈log⁡2(HW)⌉+HWC)\mathcal{O}(HWC \lceil\log_2(HW)\rceil + HWC), scaling quasilinearly with spatial resolution.

    Empirical benchmarking on an NVIDIA V100 GPU with 16GB memory at batch size 4 across resolutions ranging from 256×256256 \times 256 to 1024×10241024 \times 1024 demonstrates the following behavior:

    1. Throughput Scaling: While attention-based hybrid models (CAFormer-B36) achieve higher throughput at standard resolution (224×224224 \times 224) due to hardware-optimized attention kernels, their throughput degrades rapidly as resolution scales, dropping precipitously from ≈102\approx 10^2 img/s at 2562256^2 to <101< 10^1 img/s at 102421024^2. Conversely, DFFormer-B36, CDFFormer-B36, and GFFormer-B36 maintain high throughput across all resolutions, closely following purely convolutional ConvFormer-B36.
    2. Peak Memory Scaling: CAFormer-B36 experiences exponential peak memory consumption exceeding 70007000 MB at 102421024^2 resolution. DFFormer-B36 and CDFFormer-B36 exhibit sublinear memory growth comparable to ConvFormer-B36, remaining below 35003500 MB at 102421024^2.

    These results establish that Dynamic Filter networks avoid the quadratic computational and memory bottlenecks of MHSA when handling high-resolution visual inputs.

  5. Knowl 5 — Semantic Segmentation Performance on ADE20K

    data/table

    Semantic segmentation performance on the ADE20K validation benchmark using Semantic FPN as the segmentation framework across different model backbones.

    Backbone Params (M) mIoU (%)
    ResNet-50 28.5 36.7
    PVT-Small 28.2 39.8
    PoolFormer-S24 23.2 40.3
    DFFormer-S18 31.7 45.1
    CDFFormer-S18 31.4 44.9
    ResNet-101 47.5 38.8
    ResNeXt-101-32x4d 47.1 39.7
    PVT-Medium 48.0 41.6
    PoolFormer-S36 34.6 42.0
    DFFormer-S36 47.2 47.5
    CDFFormer-S36 46.5 46.7
    PVT-Large 65.1 42.1
    PoolFormer-M36 59.8 42.4
    DFFormer-M36 66.4 47.6
    CDFFormer-M36 65.2 48.6

    DFFormer and CDFFormer backbones significantly outperform convolutional and pooling-based baselines across all parameter regimes. DFFormer-S36 achieves 47.5% mIoU, outperforming PoolFormer-S36 by 5.5 mIoU points. CDFFormer-M36 achieves the highest score of 48.6% mIoU, outperforming PVT-Large by 6.5 mIoU points.

  6. Knowl 6 — Dynamic Filter Design and Activation Ablation

    data/table

    Ablation experiments on ImageNet-1K using DFFormer-S18 evaluating filter parameters, mixer formulations, and activation functions. Inference throughput was measured on an NVIDIA V100 GPU (16GB memory) at batch size 16 with 224×224224 \times 224 inputs.

    Ablation Variant Params (M) FLOPs (G) Throughput (img/s) Top-1 Acc. (%)
    Baseline DFFormer-S18 30 3.8 535 83.2
    Filter N=4→2N = 4 \to 2 29 3.8 534 83.0
    ρ=0.25→0.125\rho = 0.25 \to 0.125 28 3.8 532 83.1
    GFFormer-S18 (Static Global Filter) 30 3.8 575 82.9
    DF →\to AFNO 30 3.8 389 82.6
    Activation StarReLU →\to GELU 30 3.8 672 82.7
    StarReLU →\to ReLU 30 3.8 640 82.5
    StarReLU, DF →\to ReLU, AFNO 30 3.8 444 82.0

    Key observations:

    1. Reducing the basis dimension NN from 4 to 2 or reducing the MLP expansion ratio ρ\rho from 0.25 to 0.125 lowers accuracy by 0.2% and 0.1%, respectively, with negligible throughput differences.
    2. Replacing dynamic filters with static global filters (GFFormer-S18) decreases accuracy by 0.3% (83.2% vs. 82.9%), demonstrating the advantage of input-dependent filter generation.
    3. Dynamic filters outperform non-separable Adaptive Fourier Neural Operators (AFNO) by 0.6% in accuracy and deliver 37.5% higher throughput (535 vs. 389 img/s).
    4. The performance advantage of the dynamic filter persists when controlling for activation functions: with ReLU activations, the Dynamic Filter baseline (82.5%) outperforms the AFNO variant (82.0%) by 0.5%.
  7. Knowl 7 — Representation Similarity and Frequency Filtering Properties

    empirical result

    Analysis of intermediate feature representations and frequency domain responses reveals significant mechanistic differences between dynamic filters and multi-head self-attention (MHSA):

    1. Representational Similarity via Linear CKA: Mini-batch linear Centered Kernel Alignment (CKA) between DFFormer-S18 and static global filter networks (GFFormer-S18) or convolutional networks (ConvFormer-S18) reveals high layer-wise similarity through stages 1 to 3, with minor divergence only in stage 4. In contrast, comparing attention-hybrid CAFormer-S18 and FFT-hybrid CDFFormer-S18 shows high representational similarity in stages 1 and 2 (where both use convolutions), but substantial divergence starting from stage 3 onward where MHSA and Dynamic Filters are introduced.
    2. Frequency Filtering Profile: Fourier analysis measuring the relative logarithmic amplitude ΔLog amplitude\Delta \text{Log amplitude} across normalized frequencies in stage 3 reveals that MHSA in CAFormer-S18 acts as a high-pass filter in hierarchical networks, preserving or enhancing higher spatial frequencies. In contrast, Dynamic Filters in CDFFormer-S18 strongly attenuate higher frequencies relative to zero frequency, functioning as low-pass and selective band-pass spatial filters.
  8. Knowl 8 — Frequency Domain Basis Diversity in Dynamic Filtering

    empirical result

    Static global filter architectures (such as GFNet) assign fixed complex filter matrices per channel, leading to noticeable filter redundancy across channels within the same layer. In the Dynamic Filter module of DFFormer-S18, each layer learns a compact set of N=4N=4 complex basis matrices Ki∈CH×⌈W/2⌉K_i \in \mathbb{C}^{H \times \lceil W/2 \rceil}.

    Visualizing the 2D frequency amplitude spectra of these basis matrices across network depths demonstrates:

    1. Reduced Redundancy: The learned basis elements within each layer exhibit distinct spatial frequency distributions without replicating redundant patterns.
    2. Multi-Spectral Filter Types: The learned basis elements within individual layers concurrently encompass low-pass (high central amplitudes), high-pass (attenuated central amplitudes and heightened perimeter amplitudes), and directional band-pass filter profiles.
    3. Adaptive First-Order Coupling: First-order linear combination via dynamic softmax weights allows the network to adaptively synthesize low-pass, high-pass, and band-pass filters per channel for each input instance.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016. Layer Normalization. In NeurIPS.
  2. 2.Bertasius, G.; Wang, H.; and Torresani, L. 2021. Is spacetime attention all you need for video understanding? In ICML.
  3. 3.Beyer, L.; Izmailov, P.; Kolesnikov, A.; Caron, M.; Kornblith, S.; Zhai, X.; Minderer, M.; Tschannen, M.; Alabdulmohsin, I.; and Pavetic, F. 2023. FlexiViT: One Model for All Patch Sizes. arXiv:2212.08013.
  4. 4.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-end object detection with transformers. In ECCV, 213–229. Springer.
  5. 5.Chen, C.-F.; Panda, R.; and Fan, Q. 2022. Regionvit: Regional-to-local attention for vision transformers. In ICLR.
  6. 6.Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; and Liu, Z. 2020. Dynamic convolution: Attention over convolution kernels. In CVPR, 11030–11039.
  7. 7.Chu, X.; Tian, Z.; Wang, Y.; Zhang, B.; Ren, H.; Wei, X.; Xia, H.; and Shen, C. 2021. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems, 34: 9355–9366.
  8. 8.Contributors, M. 2020. MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. https://github.com/open-mmlab/mmsegmentation.
  9. 9.Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020. RandAugment: Practical automated data augmentation with a reduced search space. In CVPRW, 702–703.
  10. 10.Ding, M.; Xiao, B.; Codella, N.; Luo, P.; Wang, J.; and Yuan, L. 2022a. Davit: Dual attention vision transformers. In ECCV, 74–92. Springer.
  11. 11.Ding, X.; Zhang, X.; Zhou, Y.; Han, J.; Ding, G.; and Sun, J. 2022b. Scaling Up Your Kernels to 31x31: Revisiting Large Kernel Design in CNNs. In CVPR.
  12. 12.Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2022. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In CVPR.
  13. 13.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR.
  14. 14.Fan, Q.; Huang, H.; Chen, M.; Liu, H.; and He, R. 2023. Rmt: Retentive networks meet vision transformers. arXiv:2309.11523.
  15. 15.Guibas, J.; Mardani, M.; Li, Z.; Tao, A.; Anandkumar, A.; and Catanzaro, B. 2022. Adaptive fourier neural operators: Efficient token mixers for transformers. In ICLR.
  16. 16.Guo, M.-H.; Cai, J.-X.; Liu, Z.-N.; Mu, T.-J.; Martin, R. R.; and Hu, S.-M. 2021. Pct: Point cloud transformer. Computational Visual Media, 7(2): 187–199.
  17. 17.Ha, D.; Dai, A.; and Le, Q. V. 2016. Hypernetworks. In ICLR.
  18. 18.Han, K.; Wang, Y.; Guo, J.; Tang, Y.; and Wu, E. 2022. Vision gnn: An image is worth graph of nodes. In NeurIPS.
  19. 19.Hassani, A.; and Shi, H. 2022. Dilated neighborhood attention transformer. arXiv:2209.15001.
  20. 20.Hatamizadeh, A.; Yin, H.; Heinrich, G.; Kautz, J.; and Molchanov, P. 2023. Global context vision transformers. In ICML, 12633–12646. PMLR.
  21. 21.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In CVPR, 770–778.
  22. 22.Hendrycks, D.; and Gimpel, K. 2023. Gaussian Error Linear Units (GELUs). arXiv:1606.08415.
  23. 23.Huang, G.; Sun, Y.; Liu, Z.; Sedra, D.; and Weinberger, K. Q. 2016. Deep Networks with Stochastic Depth. In ECCV, 646–661.
  24. 24.Jia, X.; De Brabandere, B.; Tuytelaars, T.; and Gool, L. V. 2016. Dynamic filter networks. In NeurIPS, volume 29.
  25. 25.Kirillov, A.; Girshick, R.; He, K.; and Dollár, P. 2019. Panoptic feature pyramid networks. In CVPR, 6399–6408.
  26. 26.Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS, volume 25, 1097–1105.
  27. 27.Lee-Thorp, J.; Ainslie, J.; Eckstein, I.; and Ontanon, S. 2022. Fnet: Mixing tokens with fourier transforms. In NAACL.
  28. 28.Li, Y.; Wu, C.-Y.; Fan, H.; Mangalam, K.; Xiong, B.; Malik, J.; and Feichtenhofer, C. 2022. MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. In CVPR, 4804–4814.
  29. 29.Liu, S.; Chen, T.; Chen, X.; Chen, X.; Xiao, Q.; Wu, B.; Pechenizkiy, M.; Mocanu, D.; and Wang, Z. 2023. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity. In ICLR.
  30. 30.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In ICCV.
  31. 31.Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022. A ConvNet for the 2020s. In CVPR.
  32. 32.Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. In ICLR.
  33. 33.Nair, V.; and Hinton, G. E. 2010. Rectified linear units improve restricted boltzmann machines. In ICML.
  34. 34.Neimark, D.; Bar, O.; Zohar, M.; and Asselmann, D. 2021. Video transformer network. In ICCV, 3163–3172.
  35. 35.Nguyen, T.; Raghu, M.; and Kornblith, S. 2021. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. In ICLR.
  36. 36.Park, N.; and Kim, S. 2022. How Do Vision Transformers Work? In ICLR.
  37. 37.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, volume 32.
  38. 38.Proakis, J. G. 2007. Digital signal processing: principles, algorithms, and applications, 4/E. Pearson Education India.
  39. 39.Rao, Y.; Zhao, W.; Zhou, J.; and Lu, J. 2022. AMixer: Adaptive Weight Mixing for Self-attention Free Vision Transformers. In ECCV, 50–67. Springer.
  40. 40.Rao, Y.; Zhao, W.; Zhu, Z.; Lu, J.; and Zhou, J. 2021. Global filter networks for image classification. In NeurIPS, volume 34.
  41. 41.Sevim, N.; Özyedek, E. O.; Şahinuç, F.; and Koç, A. 2023. Fast-FNet: Accelerating Transformer Encoder Models via Efficient Fourier Layers. arXiv:2209.12816.
  42. 42.So, D. R.; Manké, W.; Liu, H.; Dai, Z.; Shazeer, N.; and Lé, Q. V. 2021. Primer: Searching for efficient transformers for language modeling. In NeurIPS.
  43. 43.Subramanian, A. 2021. torch_cka. https://github.com/AntixK/PyTorch-Model-Compare.
  44. 44.Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016. Rethinking the Inception Architecture for Computer Vision. In CVPR, 2818–2826.
  45. 45.Tatsunami, Y.; and Taki, M. 2022. Sequencer: Deep LSTM for Image Classification. In NeurIPS.
  46. 46.Tay, Y.; Bahri, D.; Metzler, D.; Juan, D.-C.; Zhao, Z.; and Zheng, C. 2021. Synthesizer: Rethinking self-attention for transformer models. In ICML, 10183–10192. PMLR.
  47. 47.Tolstikhin, I. O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; et al. 2021. Mlp-mixer: An all-mlp architecture for vision. In NeurIPS, volume 34.
  48. 48.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021. Training data-efficient image transformers & distillation through attention. In ICML.
  49. 49.Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; and Li, Y. 2022. Maxvit: Multi-axis vision transformer. In ECCV, 459–479. Springer.
  50. 50.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In NeurIPS, volume 30.
  51. 51.Vo, X.-T.; Nguyen, D.-L.; Priadana, A.; and Jo, K.-H. 2023. Dynamic Circular Convolution for Image Classification. In International Workshop on Frontiers of Computer Vision.
  52. 52.Wang, W.; Xie, E.; Li, X.; Fan, D.-P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; and Shao, L. 2021. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions. In ICCV.
  53. 53.Wang, Z.; Jiang, W.; Zhu, Y. M.; Yuan, L.; Song, Y.; and Liu, W. 2022. Dynamixer: a vision MLP architecture with dynamic mixing. In ICML, 22691–22701. PMLR.
  54. 54.Wei, Y.; Liu, H.; Xie, T.; Ke, Q.; and Guo, Y. 2022. Spatialtemporal transformer for 3d point cloud sequences. In WACV, 1171–1180.
  55. 55.Wightman, R. 2019. PyTorch Image Models. https://github.com/rwightman/pytorch-image-models.
  56. 56.Xie, S.; Girshick, R.; Dollár, P.; Tu, Z.; and He, K. 2017. Aggregated residual transformations for deep neural networks. In CVPR, 1492–1500.
  57. 57.Yang, B.; Bender, G.; Le, Q. V.; and Ngiam, J. 2019. Condconv: Conditionally parameterized convolutions for efficient inference. In NeurIPS, volume 32.
  58. 58.Yu, W.; Luo, M.; Zhou, P.; Si, C.; Zhou, Y.; Wang, X.; Feng, J.; and Yan, S. 2022a. Metaformer is actually what you need for vision. In CVPR.
  59. 59.Yu, W.; Si, C.; Zhou, P.; Luo, M.; Zhou, Y.; Feng, J.; Yan, S.; and Wang, X. 2022b. MetaFormer Baselines for Vision. arXiv:2210.13452.
  60. 60.Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.-H.; Tay, F. E.; Feng, J.; and Yan, S. 2021. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In ICCV, 558–567.
  61. 61.Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. In ICCV, 6023–6032.
  62. 62.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR.
  63. 63.Zhang, P.; Dai, X.; Yang, J.; Xiao, B.; Yuan, L.; Zhang, L.; and Gao, J. 2021. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2998–3008.
  64. 64.Zhao, H.; Jiang, L.; Jia, J.; Torr, P. H.; and Koltun, V. 2021. Point transformer. In ICCV, 16259–16268.
  65. 65.Zhong, Z.; Zheng, L.; Kang, G.; Li, S.; and Yang, Y. 2020. Random erasing data augmentation. In AAAI, volume 34, 13001–13008.
  66. 66.Zhou, B.; Zhao, H.; Puig, X.; Fidler, S.; Barriuso, A.; and Torralba, A. 2017. Scene parsing through ade20k dataset. In CVPR, 633–641.
  67. 67.Zhou, D.; Shi, Y.; Kang, B.; Yu, W.; Jiang, Z.; Li, Y.; Jin, X.; Hou, Q.; and Feng, J. 2021. Refiner: Refining Self-attention for Vision Transformers. arXiv:2106.03714.

Citation

MLA
Tatsunami, Y., and M. Taki. “FFT-based Dynamic Token Mixer for Vision”. arXiv, 2023, http://arxiv.org/abs/2303.03932v2.
APA
Tatsunami, Y., & Taki, M. (2023). FFT-based Dynamic Token Mixer for Vision. arXiv. http://arxiv.org/abs/2303.03932v2
Chicago
Tatsunami, Y., and M. Taki. 2023. “FFT-based Dynamic Token Mixer for Vision”. arXiv. http://arxiv.org/abs/2303.03932v2.
Harvard
Tatsunami, Y. and Taki, M. (2023) “FFT-based Dynamic Token Mixer for Vision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.03932v2.
Vancouver
1. Tatsunami Y, Taki M (2023) FFT-based Dynamic Token Mixer for Vision. arXiv

BibTeX

@article{tatsunami2023fft,
  title = {FFT-based Dynamic Token Mixer for Vision},
  author = {Tatsunami, Yuki and Taki, Masato},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.03932v2},
  eprint = {2303.03932}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF