Kolmogorov-Arnold Transformer

Xingyi YangXinchao Wang

article2025ICLR153 citations

Introduces the Kolmogorov-Arnold Transformer, which scales Kolmogorov-Arnold Networks within deep learning architectures by using GPU-friendly rational basis functions, group-level parameter sharing, and variance-preserving initialization to outperform standard MLP-based transformers.

Listen

Modern artificial intelligence relies heavily on transformer architectures, which typically use multi-layer perceptron (MLP) modules to process information across feature channels. While effective, standard MLPs struggle to model complex functions efficiently. Recently, Kolmogorov-Arnold Networks (KANs) emerged as an expressive alternative with theoretical parameter efficiency. However, direct attempts to integrate KANs into large-scale vision models have consistently failed due to severe hardware bottlenecks, exponential parameter growth, and training instability that causes large models to crash.

The article demonstrates how to overcome these scaling barriers by introducing the Kolmogorov–Arnold Transformer (KAT). The primary objective is to redesign KAN layers to make them computationally efficient and stable enough to replace traditional MLPs across large-scale vision tasks.

To achieve this, the authors evaluated the structural bottlenecks of standard KANs and introduced Group-Rational KANs (GR-KAN). They replaced unoptimized spline functions with GPU-friendly rational functions accelerated via custom hardware instructions, shared activation functions across groups of neuron channels to prevent parameter explosion, and instituted a variance-preserving weight initialization scheme to ensure stable training. The proposed KAT models were benchmarked across standard computer vision tasks, including ImageNet-1K image classification (spanning 5.7M to 86.6M parameter configurations), MS-COCO object detection and instance segmentation, and ADE20K semantic segmentation.

The findings establish that KAT successfully scales and outperforms conventional transformer baselines with comparable parameter sizes and computational costs. On ImageNet-1K, KAT-Base achieved an 82.3% top-1 accuracy from scratch—outperforming standard Vision Transformers by 3.1% and DeiT by 0.5%—and reached 82.8% when initialized with pre-trained weights. Standard unadapted KAN models failed entirely at this scale, producing numerical errors. Furthermore, in object detection on the MS-COCO benchmark, KAT backbones delivered an improvement of up to 3.0 average precision points over standard vision transformer backbones with virtually negligible computational overhead. In semantic segmentation on ADE20K, KAT-Small improved mean intersection-over-union by 2.6% over DeiT-Small. Algorithmic optimizations using rational functions and nested polynomial calculations reduced operations by approximately 9.3 times compared to standard spline-based KAN configurations.

These results indicate that replacing traditional MLPs with group-rational activations significantly increases model capacity without requiring expanded network width or depth. For organizations deploying computer vision systems, KAT provides a path to enhanced accuracy without substantial increases in memory footprint or training compute. Additionally, the ability to transfer pre-trained weights from existing Vision Transformers directly into KAT reduces the financial and operational costs associated with training new architectures from scratch.

Based on these findings, teams maintaining vision transformer pipelines should consider piloting GR-KAN replacements, particularly for small- to medium-sized models where accuracy gains are highest relative to baseline costs. When integrating these layers, practitioners should utilize pre-trained transformer weights and dedicated low-level acceleration libraries to maximize performance and throughput. Future development should focus on extending this architecture to other domains, such as natural language processing and reinforcement learning, while exploring alternative mathematical bases like wavelets or Fourier transforms.

The authors note specific operational trade-offs and limitations. Although custom acceleration significantly improves efficiency, rational function evaluations still exhibit slightly lower processing throughput (approximately 13% lower batches per second on tested hardware) compared to simpler activation functions like ReLU. Moreover, hierarchical vision architectures (such as ConvNeXt) still outperform plain KAT models in semantic segmentation due to structural design advantages. While confidence in KAT's stability and accuracy across image recognition tasks is high, additional validation is required before deploying it in non-visual applications or strict low-latency inference environments.

  • Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). This foundational work introduces Kolmogorov-Arnold Networks (KANs) with learnable 1D edge activations, establishing the core architectural concept that the Kolmogorov-Arnold Transformer adapts to replace standard MLP layers.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This paper establishes the canonical Transformer architecture, providing the foundational multi-layer perceptron (MLP) and attention mechanics that Kolmogorov-Arnold Transformer seeks to enhance.
  • Paper: Understanding the difficulty of training deep feedforward neural networks, Xavier Glorot et al. (2010). This classic study demonstrates the principles of variance preservation during neural network weight initialization, directly motivating the variance-preserving initialization developed to stabilize deep KAN layers.
Cover for Kolmogorov-Arnold Transformer

Abstract

Transformers stand as the cornerstone of mordern deep learning. Traditionally, these models rely on multi-layer perceptron (MLP) layers to mix the information between channels. In this paper, we introduce the Kolmogorov-Arnold Transformer (KAT), a novel architecture that replaces MLP layers with Kolmogorov-Arnold Network (KAN) layers to enhance the expressiveness and performance of the model. Integrating KANs into transformers, however, is no easy feat, especially when scaled up. Specifically, we identify three key challenges: (C1) Base function. The standard B-spline function used in KANs is not optimized for parallel computing on modern hardware, resulting in slower inference speeds. (C2) Parameter and Computation Inefficiency. KAN requires a unique function for each input-output pair, making the computation extremely large. (C3) Weight initialization. The initialization of weights in KANs is particularly challenging due to their learnable activation functions, which are critical for achieving convergence in deep neural networks. To overcome the aforementioned challenges, we propose three key solutions: (S1) Rational basis. We replace B-spline functions with rational functions to improve compatibility with modern GPUs. By implementing this in CUDA, we achieve faster computations. (S2) Group KAN. We share the activation weights through a group of neurons, to reduce the computational load without sacrificing performance. (S3) Variance-preserving initialization. We carefully initialize the activation weights to make sure that the activation variance is maintained across layers. With these designs, KAT scales effectively and readily outperforms traditional MLP-based transformers.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 2.1 Kolmogorov-Arnold representation theorem
  • 2.2 Kolmogorov–Arnold Networks
  • 3 Why original KAN fails to scale?
  • 4 Kolmogorov–Arnold Transformer
  • 4.1 Overall Architecture
  • 4.2 Rational Base Functions
  • 4.3 Group KAN
  • 4.4 Variance-Preserving Initialization
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Image Recognition
  • 5.3 Object Detection and Instance Segmentation
  • 5.4 Semantic Segmentation
  • 5.5 Ablation Study and Analysis
  • 6 Conclusion and Future Work
  • References
  • 7 Derivation and Calculation of FLOPs
  • 7.1 Plain Computation
  • 7.2 Horner’s Method
  • 8 Hyper-parameters for KAT model

Knowls

  1. Knowl 1 — Kolmogorov–Arnold Transformer Architecture

    model/method

    The Kolmogorov–Arnold Transformer (KAT) modifies standard Vision Transformer architectures by replacing standard Multi-Layer Perceptron (MLP) channel-mixing blocks with two-layer Group-Rational Kolmogorov-Arnold Network (GR-KAN) blocks while retaining standard Multi-Head Self-Attention (MSA) modules.

    For an input sequence of tokens xℓ−1∈RT×Cx_{\ell-1} \in \mathbb{R}^{T \times C} at transformer layer ℓ∈{1,…,L}\ell \in \{1, \dots, L\}, KAT computes:

    x0(ℓ)=MSA(LN(xℓ−1))+xℓ−1x_0^{(\ell)} = \text{MSA}(\text{LN}(x_{\ell-1})) + x_{\ell-1}

    xℓ=GR-KAN(LN(x0(ℓ)))+x0(ℓ)x_\ell = \text{GR-KAN}(\text{LN}(x_0^{(\ell)})) + x_0^{(\ell)}

    where LN(⋅)\text{LN}(\cdot) denotes Layer Normalization, and GR-KAN(⋅)\text{GR-KAN}(\cdot) is a two-layer channel-mixing network employing group-shared rational basis activation functions followed by linear transformations.

  2. Knowl 2 — Group-Rational Kolmogorov–Arnold Network Layer

    model/method

    A Group-Rational Kolmogorov-Arnold Network (GR-KAN) layer maps an input vector x∈Rdinx \in \mathbb{R}^{d_{in}} to an output vector y∈Rdouty \in \mathbb{R}^{d_{out}} by grouping input channels to share learnable rational activation functions prior to a linear projection.

    The input channels dind_{in} are divided into gg disjoint groups, with each group containing dg=din/gd_g = d_{in}/g channels. For channel index i∈{1,…,din}i \in \{1, \dots, d_{in}\}, the group index is given by ⌊i/dg⌋\lfloor i / d_g \rfloor. The jj-th output component is computed as:

    yj=∑i=1dinwj,iF⌊i/dg⌋(xi)+bjy_j = \sum_{i=1}^{d_{in}} w_{j,i} F_{\lfloor i / d_g \rfloor}(x_i) + b_j

    where wj,i∈Rw_{j,i} \in \mathbb{R} is the learnable edge weight connecting input channel ii to output channel jj, bj∈Rb_j \in \mathbb{R} is a learnable bias, and Fk:R→RF_k: \mathbb{R} \to \mathbb{R} is the rational activation function shared across input channels in group k∈{0,…,g−1}k \in \{0, \dots, g-1\}.

    In matrix notation, GR-KAN is expressed as:

    GR-KAN(x)=WF(x)+b=linear(group_rational(x))\text{GR-KAN}(x) = W F(x) + b = \text{linear}(\text{group\_rational}(x))

    where W∈Rdout×dinW \in \mathbb{R}^{d_{out} \times d_{in}} and F(x)=[F⌊1/dg⌋(x1),…,F⌊din/dg⌋(xdin)]⊤∈RdinF(x) = [F_{\lfloor 1 / d_g \rfloor}(x_1), \dots, F_{\lfloor d_{in} / d_g \rfloor}(x_{d_{in}})]^\top \in \mathbb{R}^{d_{in}}.

  3. Knowl 3 — Safe Padé Activation Unit Base Function

    equation

    To prevent numerical divergence caused by zero-crossing denominators in standard rational functions, GR-KAN parameterizes each univariate base activation function as a Safe Padé Activation Unit (PAU) of polynomial degrees mm (numerator) and nn (denominator):

    F(x)=∑i=0maixi1+∣∑j=1nbjxj∣=P(x)Q(x)F(x) = \frac{\sum_{i=0}^m a_i x^i}{1 + \left| \sum_{j=1}^n b_j x^j \right|} = \frac{P(x)}{Q(x)}

    where P(x)=a0+a1x+⋯+amxmP(x) = a_0 + a_1 x + \dots + a_m x^m and Q(x)=1+∣A(x)∣Q(x) = 1 + |A(x)| with A(x)=b1x+⋯+bnxnA(x) = b_1 x + \dots + b_n x^n. The parameters {ai}i=0m\{a_i\}_{i=0}^m and {bj}j=1n\{b_j\}_{j=1}^n are learnable coefficients optimized via backpropagation. In KAT, default polynomial orders are set to m=5m=5 and n=4n=4. The denominator coefficients {bj}\{b_j\} are shared across all gg channel groups, while the numerator coefficients {ai}\{a_i\} are learned independently per group.

  4. Knowl 4 — Variance-Preserving Weight Initialization for GR-KAN

    algorithm

    To maintain signal variance Var[y]=Var[x]\text{Var}[y] = \text{Var}[x] across stacked layers, GR-KAN initializes rational coefficients {ai,bj}\{a_i, b_j\} and linear weights W∈Rdout×dinW \in \mathbb{R}^{d_{out} \times d_{in}} sequentially:

    Input: Input dimension dind_{in}, target base activation shape (e.g., Identity, GELU, Swish)
    Output: Initialized rational coefficients a,ba, b, linear weights WW, and biases bbiasb_{bias}
    1. Fit coefficients a=(a0,…,am)a = (a_0, \dots, a_m) and b=(b1,…,bn)b = (b_1, \dots, b_n) such that F(x)F(x) minimizes regression error against the target activation function.
    2. Numerically evaluate the activation gain α\alpha assuming standard normal input x∼N(0,1)x \sim \mathcal{N}(0, 1):
       α=Var[x]E[F(x)2]=1∫−∞+∞F(x)212πe−x2/2dx\alpha = \frac{\text{Var}[x]}{\mathbb{E}[F(x)^2]} = \frac{1}{\int_{-\infty}^{+\infty} F(x)^2 \frac{1}{\sqrt{2\pi}} e^{-x^2 / 2} dx}
    3. Initialize each linear weight entry wj,i∼N(0,αdin)w_{j,i} \sim \mathcal{N}\left(0, \frac{\alpha}{d_{in}}\right).
    4. Initialize each bias entry bj=0b_j = 0.
    5. return a,b,W,bbiasa, b, W, b_{bias}

    Numerical gain values α\alpha for common target functions are:

    • Identity: α=1.0000\alpha = 1.0000
    • ReLU: α=2.0000\alpha = 2.0000
    • GELU: α≈2.3568\alpha \approx 2.3568
    • Swish / SiLU: α≈2.8178\alpha \approx 2.8178
    • GEGLU: α≈0.7112\alpha \approx 0.7112
    • SwishGLU: α≈0.8434\alpha \approx 0.8434
  5. Knowl 5 — Parameter and FLOP Complexity Comparison of GR-KAN vs. KAN and MLP

    data/table

    Compared to standard KANs parameterized with B-splines, GR-KAN eliminates the multiplicative scaling of parameter count and computational complexity with spline order KK and grid interval count GG, reducing them to a constant overhead relative to standard MLPs.

    Model No. Parameters FLOPs (per sample)
    MLP din×dout+doutd_{in} \times d_{out} + d_{out} Func FLOPs×dout+2(din×dout)\text{Func FLOPs} \times d_{out} + 2(d_{in} \times d_{out})
    KAN din×dout×(G+K+3)+doutd_{in} \times d_{out} \times (G+K+3) + d_{out} Func FLOPs×din+(din×dout)×[9K(G+1.5K)+2G−2.5K+3]\text{Func FLOPs} \times d_{in} + (d_{in} \times d_{out}) \times [9K(G+1.5K) + 2G - 2.5K + 3]
    GR-KAN (Ours) din×dout+dout+(m+n×g)d_{in} \times d_{out} + d_{out} + (m + n \times g) (2m+2n+3)×din+2(din×dout)(2m + 2n + 3) \times d_{in} + 2(d_{in} \times d_{out})

    For a single scalar activation evaluation with m=5,n=4m=5, n=4 using Horner's method, the rational activation requires 21 FLOPs (9 multiplications, 10 additions, 1 absolute value, 1 division), compared to 46 FLOPs for unnested evaluation and 204 FLOPs for a B-spline with G=3,K=3G=3, K=3 (a 9.3×\times reduction in activation FLOPs).

  6. Knowl 6 — Analytical Gradients and Horner Evaluation for Rational Base Functions

    equation

    In the custom CUDA implementation of the rational base function F(x)=P(x)Q(x)=P(x)1+∣A(x)∣F(x) = \frac{P(x)}{Q(x)} = \frac{P(x)}{1 + |A(x)|}, polynomial evaluation is performed using Horner's nested rule:

    P(x)=a0+x(a1+x(a2+⋯+x(am−1+xam)… ))P(x) = a_0 + x(a_1 + x(a_2 + \dots + x(a_{m-1} + x a_m)\dots))

    which evaluates a degree-mm polynomial with exactly mm multiplications and mm additions.

    The analytical partial derivatives used for exact backward gradient computation are:

    ∂F∂ai=xiQ(x),i∈{0,…,m}\frac{\partial F}{\partial a_i} = \frac{x^i}{Q(x)}, \quad i \in \{0, \dots, m\}

    ∂F∂bj=−xjA(x)∣A(x)∣P(x)Q(x)2,j∈{1,…,n}\frac{\partial F}{\partial b_j} = - x^j \frac{A(x)}{|A(x)|} \frac{P(x)}{Q(x)^2}, \quad j \in \{1, \dots, n\}

    ∂F∂x=1Q(x)∂P(x)∂x−P(x)Q(x)2∂Q(x)∂x\frac{\partial F}{\partial x} = \frac{1}{Q(x)} \frac{\partial P(x)}{\partial x} - \frac{P(x)}{Q(x)^2} \frac{\partial Q(x)}{\partial x}

    where ∂P(x)∂x=∑i=1miaixi−1\frac{\partial P(x)}{\partial x} = \sum_{i=1}^m i a_i x^{i-1} and ∂Q(x)∂x=A(x)∣A(x)∣∑j=1njbjxj−1\frac{\partial Q(x)}{\partial x} = \frac{A(x)}{|A(x)|} \sum_{j=1}^n j b_j x^{j-1}.

  7. Knowl 7 — Pretrained Weight Transfer from Vision Transformers to KAT

    model/method

    KAT models can directly transfer parameters from a pretrained Vision Transformer (ViT) without architecture modification outside the channel mixer:

    1. Linear Layer Cloning: The linear weights and biases of the two fully connected (FC) layers in the ViT MLP are copied directly into the corresponding linear matrices of the two GR-KAN layers.
    2. First Rational Layer Initialization: The rational activation function of the first GR-KAN layer is initialized to fit an Identity function (F1(x)≈xF_1(x) \approx x).
    3. Second Rational Layer Initialization: The rational activation function of the second GR-KAN layer is initialized to fit the non-linear activation of the original ViT MLP (such as GELU or Swish).
    4. Transformer Block Cloning: Patch embedding, positional embedding, Layer Normalization, and Multi-Head Self-Attention layers are loaded directly without modification.
  8. Knowl 8 — ImageNet-1K Image Classification Performance

    data/table

    Models were evaluated on ImageNet-1K image classification at 224×224224 \times 224 resolution using a 300-epoch training schedule with AdamW (batch size 1024). Models marked with an asterisk (∗*) were initialized from pretrained ViT weights rather than trained from scratch.

    Model Channel Mixer #Param. FLOPs IN-1k Top-1 (%)
    ViT-Ti/16 MLP 5.7M 1.08G 72.7
    DeiT-T MLP 5.7M 1.08G 72.2
    ViT-T + KAN KAN 12.8M 1.78G 64.9
    KAT-T GR-KAN 5.7M 1.13G 74.6
    KAT-T* GR-KAN 5.7M 1.13G 75.7
    ViT-S/16 MLP 22.1M 4.25G 78.8
    DeiT-S MLP 22.1M 4.25G 79.8
    ViT-S + KAN KAN 50.4M 7.05G 62.9
    KAT-S GR-KAN 22.1M 4.35G 81.2
    KAT-S* GR-KAN 22.1M 4.35G 82.0
    ViT-B/16 MLP 86.6M 16.87G 79.1
    DeiT-B MLP 86.6M 16.87G 81.8
    ViT-B + KAN KAN 199.8M 28.04G NaN
    KAT-B GR-KAN 86.6M 17.06G 82.3
    KAT-B* GR-KAN 86.6M 17.06G 82.8

    KAT outperforms standard MLP mixers across all model capacities (e.g., KAT-S achieves 81.2% Top-1 vs. 79.8% for DeiT-S). Vanilla ViT+KAN exhibits severe degradation (62.9% on Small) or fails to converge (NaN on Base).

  9. Knowl 9 — Object Detection and Instance Segmentation Performance on MS-COCO

    data/table

    Object detection and instance segmentation performance evaluated using Mask R-CNN backbones within the ViTDet framework on MS-COCO 2017 with a 3×3\times schedule (36 epochs), input image size 800×1333800 \times 1333, batch size 16, and FP16 precision on 4 NVIDIA H100 GPUs.

    Backbone #Param. FLOPs APbox\text{AP}^{\text{box}} AP50box\text{AP}^{\text{box}}_{50} AP75box\text{AP}^{\text{box}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} AP75mask\text{AP}^{\text{mask}}_{75}
    PVT-Small 44.1M - 43.0 65.3 46.9 39.9 62.5 42.8
    Swin-T 48M 267G 46.0 68.1 50.3 41.6 65.1 44.9
    ConvNeXt-T 48M 262G 46.2 67.9 50.8 41.7 65.0 44.9
    ViT-S 43.8M 423G 44.0 66.9 47.8 39.9 63.4 42.2
    ViTDet-S 44.5M 423G 44.5 66.9 48.4 40.1 63.6 42.5
    KATDet-S 44.5M 424G 47.5 69.0 51.2 41.5 65.7 44.0
    ViT-B 113.6M 767G 45.8 68.2 50.1 41.3 65.1 44.4
    ViTDet-B 113.6M 767G 46.3 68.6 50.5 41.6 65.3 44.5
    KATDet-B 113.7M 770G 47.7 69.1 51.6 41.6 65.9 44.3

    KATDet-S achieves a +3.0 APbox\text{AP}^{\text{box}} and +1.4 APmask\text{AP}^{\text{mask}} improvement over ViTDet-S with 1 GFLOP computational overhead.

  10. Knowl 10 — Semantic Segmentation Performance on ADE20K

    data/table

    Semantic segmentation results using UperNet with different backbones initialized with ImageNet pre-trained weights on the ADE20K validation set (150 categories, 512×512512 \times 512 crop training, 160,000 iterations, evaluation MACs measured at 512×2048512 \times 2048).

    Backbone #Param. FLOPs mIoU (%)
    Swin-T 60M 945G 45.8
    ConvNeXt-T 60M 939G 46.7
    DeiT-S 57M 1217G 43.5
    KAT-S 57M 1219G 46.1
    Swin-B 121M 1188G 49.5
    ConvNeXt-B 122M 1170G 49.6
    DeiT-B 142M 2007G 47.2
    KAT-B 142M 2011G 47.4

    KAT-S improves mIoU by 2.6 percentage points over DeiT-S (46.1% vs. 43.5%) with virtually identical parameter count (57M) and FLOPs (1219G vs. 1217G).

  11. Knowl 11 — Ablation of Activation Function Variants and Channel Mixer Initializations

    data/table

    Ablation experiments comparing activation functions in the channel mixer of ViT-Ti/16 and initializations of the two GR-KAN layers on ImageNet-1K.

    Activation Name Learnable? IN-1k Top-1 (%)
    GELU (Default) No 72.7
    ReLU No 72.8
    SiLU No 69.8
    PReLU Yes 73.2
    PAU Yes 73.6
    KAT-T (GR-KAN) Yes 74.6
    Layer 1 Rational Init Layer 2 Rational Init IN-1k Top-1 (%)
    Identity Identity 69.7
    Swish Swish 74.4
    Identity GELU 74.5
    Identity Swish 74.6

    Throughput and peak memory measured on an NVIDIA A5000 GPU with input shape [64,1000,512][64, 1000, 512] shows that while KAT-T achieves 2313 batch/s compared to 2643 batch/s for GELU (a ~12.5% throughput reduction), peak memory usage remains unchanged at 1380 MB.

  12. Knowl 12 — Limitations of Kolmogorov–Arnold Transformers

    limitation

    The authors identify three main limitations of the KAT architecture:

    1. Inference Throughput: Despite custom CUDA implementations, rational function evaluations remain slower than simple elementwise activations like ReLU and GELU due to higher per-element arithmetic complexity.
    2. Gradient Stability: Higher-order polynomial gradients with respect to coefficients ama_m and bnb_n depend strongly on powers of the input magnitude, creating potential gradient instability during deep backpropagation if inputs are poorly scaled.
    3. Hybrid Representation: GR-KAN relies on group-shared activation functions followed by a full linear transformation matrix, functioning as a hybrid between MLPs and KANs rather than a pure edge-wise KAN with distinct functions per input-output edge pair.

Coverage note — None was omitted; all key architectural components, theoretical derivations, algorithms, benchmark tables, ablation analyses, and stated limitations are represented.

References

  1. 1.Alireza Afzal Aghaei. rkan: Rational kolmogorov-arnold networks. arXiv preprint arXiv:2406.14495, 2024.
  2. 2.Zavareh Bozorgasl and Hao Chen. Wav-kan: Wavelet kolmogorov-arnold networks. arXiv preprint arXiv:2405.12832, 2024.
  3. 3.Zavareh Bozorgasl and Hao Chen. Wav-kan: Wavelet kolmogorov-arnold networks, 2024.
  4. 4.Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs, Yoni Kasten, and Shira Kritchman. Frequency bias in neural networks for input of non-uniform density. In International Conference on Machine Learning, pages 685–694. PMLR, 2020.
  5. 5.George A Baker Jr and John L Gammel. The padé approximant. Journal of Mathematical Analysis and Applications, 2(1):21–30, 1961.
  6. 6.Nicolas Boullé, Yuji Nakatsukasa, and Alex Townsend. Rational neural networks. Advances in neural information processing systems, 33:14243–14253, 2020.
  7. 7.C de Boor. Subroutine package for calculating with b-splines, 1971.
  8. 8.Alexander Dylan Bodner, Antonio Santiago Tepsich, Jack Natan Spolski, and Santiago Pourteau. Convolutional kolmogorov-arnold networks. arXiv preprint arXiv:2406.13155, 2024.
  9. 9.Ziwen Chen, Gundavarapu, and WU DI. Vision-kan: Exploring the possibility of kan replacing mlp in vision transformer. https://github.com/chenziwenhaoshuai/Vision-KAN.git, 2024.
  10. 10.Minjong Cheon. Demonstrating the efficacy of kolmogorov-arnold networks in vision tasks. arXiv preprint arXiv:2406.14916, 2024.
  11. 11.Minjong Cheon. Kolmogorov-arnold network for satellite image classification in remote sensing. arXiv preprint arXiv:2406.00600, 2024.
  12. 12.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, et al. Mmdetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  13. 13.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 702–703, 2020.
  14. 14.Yifei Chen, Zhu Zhu, Shenghao Zhu, Linwei Qiu, Binfeng Zou, Fan Jia, Yunpeng Zhu, Chenyan Zhang, Zhaojie Fang, Feiwei Qin, et al. Sckansformer: Fine-grained classification of bone marrow cells via kansformer backbone and hierarchical attention mechanisms. arXiv preprint arXiv:2406.09931, 2024.
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  16. 16.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural networks, 107:3–11, 2018.
  17. 17.Kunihiko Fukushima. Visual feature extraction by a multilayered network of analog threshold elements. IEEE Transactions on Systems Science and Cybernetics, 5(4):322–333, 1969.
  18. 18.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256. JMLR Workshop and Conference Proceedings, 2010.
  19. 19.William J Gordon and Richard F Riesenfeld. B-spline curves and surfaces. In Computer aided geometric design, pages 95–126. Elsevier, 1974.
  20. 20.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  22. 22.Robert Hecht-Nielsen. Kolmogorov’s mapping neural network existence theorem. In Proceedings of the international conference on Neural Networks, volume 3, pages 11–14. IEEE press New York, NY, USA, 1987.
  23. 23.WG Horner. A new method of solving numerical equations of all orders, by continuous approximation. In Abstracts of the Papers Printed in the Philosophical Transactions of the Royal Society of London, volume 2, pages 117–117. JSTOR, 1815.
  24. 24.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 646–661. Springer, 2016.
  25. 25.Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks, 2(5):359–366, 1989.
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015.
  27. 27.Yann LeCun, Yoshua Bengio, et al. Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks, 3361(10):1995, 1995.
  28. 28.Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1(4):541–551, 1989.
  29. 29.Yann LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller. Efficient backprop. In Neural networks: Tricks of the trade, pages 9–50. Springer, 2002.
  30. 30.Henry Leung and Simon Haykin. Rational function neural network. Neural Computation, 5(6):928–938, 1993.
  31. 31.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019.
  32. 32.Ziyao Li. Kolmogorov-arnold networks are radial basis function networks. ArXiv, abs/2405.06721, 2024.
  33. 33.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.
  34. 34.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  35. 35.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. In European conference on computer vision, pages 280–296. Springer, 2022.
  36. 36.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022.
  37. 37.Ziming Liu, Pingchuan Ma, Yixuan Wang, Wojciech Matusik, and Max Tegmark. Kan 2.0: Kolmogorov-arnold networks meet science. arXiv preprint arXiv:2408.10205, 2024.
  38. 38.Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljačić, Thomas Y Hou, and Max Tegmark. Kan: Kolmogorov-arnold networks. arXiv preprint arXiv:2404.19756, 2024.
  39. 39.Alejandro Molina, Patrick Schramowski, and Kristian Kersting. Padé activation units: End-to-end learning of flexible activation functions in deep networks. In International Conference on Learning Representations, 2020.
  40. 40.John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue, 6(2):40–53, 2008.
  41. 41.Gist Noesis. Fourierkan, 2024.
  42. 42.Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. In International conference on machine learning, pages 5301–5310. PMLR, 2019.
  43. 43.Basri Ronen, David Jacobs, Yoni Kasten, and Shira Kritchman. The convergence rate of neural networks for learned functions of different frequencies. Advances in Neural Information Processing Systems, 32, 2019.
  44. 44.Daniel Ruijters and Philippe Thévenaz. Gpu prefilter for accurate cubic b-spline interpolation. The Computer Journal, 55(1):15–20, 2012.
  45. 45.Daniel Ruijters, Bart M ter Haar Romeny, and Paul Suetens. Efficient gpu-based texture interpolation using uniform b-splines. Journal of Graphics Tools, 13(4):61–69, 2008.
  46. 46.Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017.
  47. 47.Christian Sigg and Markus Hadwiger. Fast third-order texture filtering. GPU gems, 2:313–329, 2005.
  48. 48.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  49. 49.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  50. 50.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International conference on machine learning, pages 10347–10357. PMLR, 2021.
  51. 51.Matus Telgarsky. Neural networks and rational functions. In International Conference on Machine Learning, pages 3387–3393. PMLR, 2017.
  52. 52.Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. Mlp-mixer: An all-mlp architecture for vision. Advances in neural information processing systems, 34:24261–24272, 2021.
  53. 53.Asher Trockman and J Zico Kolter. Mimetic initialization of self-attention layers. In International Conference on Machine Learning, pages 34456–34468. PMLR, 2023.
  54. 54.Michael Unser, Akram Aldroubi, and Murray Eden. B-spline signal processing. i. theory. IEEE transactions on signal processing, 41(2):821–833, 1993.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.
  56. 56.Joseph Leonard Walsh. Interpolation and approximation by rational functions in the complex domain, volume 20. American Mathematical Soc., 1935.
  57. 57.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European conference on computer vision (ECCV), pages 3–19, 2018.
  58. 58.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European conference on computer vision (ECCV), pages 418–434, 2018.
  59. 59.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  60. 60.Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10819–10829, 2022.
  61. 61.Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. Metaformer baselines for vision. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  62. 62.Runpeng Yu, Weihao Yu, and Xinchao Wang. Kan or mlp: A fairer comparison. arXiv preprint arXiv:2407.16674, 2024.
  63. 63.Xingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye, and Xinchao Wang. Deep model reassembly. Advances in neural information processing systems, 35:25739–25753, 2022.
  64. 64.Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018.
  65. 65.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 13001–13008, 2020.
  66. 66.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017.

Citation

MLA
Yang, X., and X. Wang. “Kolmogorov-Arnold Transformer”. arXiv, 2024, http://arxiv.org/abs/2409.10594v1.
APA
Yang, X., & Wang, X. (2024). Kolmogorov-Arnold Transformer. arXiv. http://arxiv.org/abs/2409.10594v1
Chicago
Yang, X., and X. Wang. 2024. “Kolmogorov-Arnold Transformer”. arXiv. http://arxiv.org/abs/2409.10594v1.
Harvard
Yang, X. and Wang, X. (2024) “Kolmogorov-Arnold Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2409.10594v1.
Vancouver
1. Yang X, Wang X (2024) Kolmogorov-Arnold Transformer. arXiv

BibTeX

@article{yang2024kolmogorov,
  title = {Kolmogorov-Arnold Transformer},
  author = {Yang, Xingyi and Wang, Xinchao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2409.10594v1},
  eprint = {2409.10594}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors