Kolmogorov-Arnold Transformer
Xingyi YangXinchao Wang
Introduces the Kolmogorov-Arnold Transformer, which scales Kolmogorov-Arnold Networks within deep learning architectures by using GPU-friendly rational basis functions, group-level parameter sharing, and variance-preserving initialization to outperform standard MLP-based transformers.
Modern artificial intelligence relies heavily on transformer architectures, which typically use multi-layer perceptron (MLP) modules to process information across feature channels. While effective, standard MLPs struggle to model complex functions efficiently. Recently, Kolmogorov-Arnold Networks (KANs) emerged as an expressive alternative with theoretical parameter efficiency. However, direct attempts to integrate KANs into large-scale vision models have consistently failed due to severe hardware bottlenecks, exponential parameter growth, and training instability that causes large models to crash.
The article demonstrates how to overcome these scaling barriers by introducing the Kolmogorov–Arnold Transformer (KAT). The primary objective is to redesign KAN layers to make them computationally efficient and stable enough to replace traditional MLPs across large-scale vision tasks.
To achieve this, the authors evaluated the structural bottlenecks of standard KANs and introduced Group-Rational KANs (GR-KAN). They replaced unoptimized spline functions with GPU-friendly rational functions accelerated via custom hardware instructions, shared activation functions across groups of neuron channels to prevent parameter explosion, and instituted a variance-preserving weight initialization scheme to ensure stable training. The proposed KAT models were benchmarked across standard computer vision tasks, including ImageNet-1K image classification (spanning 5.7M to 86.6M parameter configurations), MS-COCO object detection and instance segmentation, and ADE20K semantic segmentation.
The findings establish that KAT successfully scales and outperforms conventional transformer baselines with comparable parameter sizes and computational costs. On ImageNet-1K, KAT-Base achieved an 82.3% top-1 accuracy from scratch—outperforming standard Vision Transformers by 3.1% and DeiT by 0.5%—and reached 82.8% when initialized with pre-trained weights. Standard unadapted KAN models failed entirely at this scale, producing numerical errors. Furthermore, in object detection on the MS-COCO benchmark, KAT backbones delivered an improvement of up to 3.0 average precision points over standard vision transformer backbones with virtually negligible computational overhead. In semantic segmentation on ADE20K, KAT-Small improved mean intersection-over-union by 2.6% over DeiT-Small. Algorithmic optimizations using rational functions and nested polynomial calculations reduced operations by approximately 9.3 times compared to standard spline-based KAN configurations.
These results indicate that replacing traditional MLPs with group-rational activations significantly increases model capacity without requiring expanded network width or depth. For organizations deploying computer vision systems, KAT provides a path to enhanced accuracy without substantial increases in memory footprint or training compute. Additionally, the ability to transfer pre-trained weights from existing Vision Transformers directly into KAT reduces the financial and operational costs associated with training new architectures from scratch.
Based on these findings, teams maintaining vision transformer pipelines should consider piloting GR-KAN replacements, particularly for small- to medium-sized models where accuracy gains are highest relative to baseline costs. When integrating these layers, practitioners should utilize pre-trained transformer weights and dedicated low-level acceleration libraries to maximize performance and throughput. Future development should focus on extending this architecture to other domains, such as natural language processing and reinforcement learning, while exploring alternative mathematical bases like wavelets or Fourier transforms.
The authors note specific operational trade-offs and limitations. Although custom acceleration significantly improves efficiency, rational function evaluations still exhibit slightly lower processing throughput (approximately 13% lower batches per second on tested hardware) compared to simpler activation functions like ReLU. Moreover, hierarchical vision architectures (such as ConvNeXt) still outperform plain KAT models in semantic segmentation due to structural design advantages. While confidence in KAT's stability and accuracy across image recognition tasks is high, additional validation is required before deploying it in non-visual applications or strict low-latency inference environments.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). This foundational work introduces Kolmogorov-Arnold Networks (KANs) with learnable 1D edge activations, establishing the core architectural concept that the Kolmogorov-Arnold Transformer adapts to replace standard MLP layers.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This paper establishes the canonical Transformer architecture, providing the foundational multi-layer perceptron (MLP) and attention mechanics that Kolmogorov-Arnold Transformer seeks to enhance.
- Paper: Understanding the difficulty of training deep feedforward neural networks, Xavier Glorot et al. (2010). This classic study demonstrates the principles of variance preservation during neural network weight initialization, directly motivating the variance-preserving initialization developed to stabilize deep KAN layers.
- Paper: SineKAN: Kolmogorov-Arnold Networks Using Sinusoidal Activation Functions, Eric A. F. Reinhardt et al. (2025). This paper explores substituting standard B-spline KAN activations with sinusoidal functions, presenting an alternative solution to the efficiency and scaling challenges of KAN architectures analyzed in KAT.
- Paper: U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation, Chenxin Li et al. (2025). This work applies Kolmogorov-Arnold Network layers within tokenized encoder-decoder backbones for medical image segmentation and diffusion models, offering a complementary application of KANs in vision and generative domains.
