SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks
Xinyu ShiZecheng HaoZhaofei Yu
Presents SpikingResformer, a hybrid multi-stage architecture powered by Dual Spike Self-Attention that eliminates floating-point matrix operations while achieving a state-of-the-art 79.40% top-1 accuracy on ImageNet with superior energy efficiency.
Spiking neural networks (SNNs) are energy-efficient, brain-inspired models well-suited for neuromorphic hardware, yet their accuracy on complex computer vision tasks has historically lagged behind conventional artificial neural networks. Recent efforts have attempted to integrate high-performing vision transformer architectures into SNNs. However, existing spiking transformers struggle to extract local image features effectively because they rely on shallow convolutional modules and lack appropriate mathematical scaling methods to handle multi-scale inputs.
The article designs and evaluates SpikingResformer, a novel architecture that combines a multi-stage residual network backbone with a new spike-driven attention mechanism called Dual Spike Self-Attention (DSSA). The main objective is to eliminate architectural bottlenecks, improve accuracy, and reduce computational energy and parameter overhead on large-scale visual recognition tasks.
The researchers evaluated this framework through extensive direct-training experiments on the standard ImageNet benchmark, ablation studies on ImageNet100, and transfer learning evaluations across static and event-based image datasets. The proposed DSSA mechanism replaces floating-point operations and softmax functions with dual spike transformations, paired with statistically derived scaling factors to support varying feature map resolutions without gradient vanishing.
The findings show that SpikingResformer sets a new state of the art in SNN performance, achieving up to 79.40% top-1 accuracy on ImageNet with only 4 time-steps. Compared to prior spiking transformers, it consistently achieves higher accuracy with significantly fewer parameters and lower energy consumption. For example, the tiny variant delivers a 2.06% accuracy gain while saving 5.67 million parameters and roughly 37% of inference energy compared to baseline spiking transformers. Ablation tests confirmed that the multi-stage architecture, the group-wise convolutional feed-forward network, and the novel scaling factors are all essential for model convergence and top performance. Furthermore, transfer learning from pre-trained ImageNet weights achieved top results on static benchmarks (97.40% on CIFAR10 and 85.98% on CIFAR100) and static-derived event data.
These results demonstrate that energy-efficient spiking models can achieve competitive visual recognition accuracy without relying on costly floating-point computations, significantly lowering the barrier for deploying high-performance vision models on low-power neuromorphic edge devices. However, the evaluation also revealed a key limitation: while transfer learning works well on static datasets, it transfers poorly to real-world event datasets with rich temporal dynamics, such as DVSGesture, where accuracy dropped 5.9% behind direct training methods due to the absence of temporal features in pre-training. Moving forward, organizations exploring low-power neuromorphic vision should consider adopting multi-stage spiking transformer architectures for static imagery, while prioritizing further research into temporal pre-training strategies to unlock similar performance gains on dynamic event streams.
- Paper: Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting, Shikuang Deng et al. (2022). Understanding direct training techniques and surrogate gradients for deep spiking neural networks is essential for following the optimization of spiking vision architectures.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). Learn how convolutional operations can be systematically integrated into multi-stage vision transformer blocks to improve local feature extraction and efficiency.
- Paper: PVT v2: Improved baselines with Pyramid Vision Transformer, Wenhai Wang et al. (2021). Explore hierarchical multi-stage vision transformer designs that combine convolutional patch embeddings with efficient spatial reduction attention.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). Discover how classic ResNet-style multi-stage architectures can be modernized with transformer-like design principles to maintain strong local feature extraction.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Examine the formulation and scaling of linear attention mechanisms that eliminate quadratic complexity, which underlies efficient spiking self-attention designs.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). Gain foundational insights into the complementary representational differences and local-versus-global feature trade-offs between ResNets and Vision Transformers.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Review the canonical Vision Transformer architecture that SpikingResformer adapts into the energy-efficient spiking domain.
No sufficiently relevant recommendations were found.
