SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks

Xinyu ShiZecheng HaoZhaofei Yu

article2024CVPR78 citations

Presents SpikingResformer, a hybrid multi-stage architecture powered by Dual Spike Self-Attention that eliminates floating-point matrix operations while achieving a state-of-the-art 79.40% top-1 accuracy on ImageNet with superior energy efficiency.

Listen

Spiking neural networks (SNNs) are energy-efficient, brain-inspired models well-suited for neuromorphic hardware, yet their accuracy on complex computer vision tasks has historically lagged behind conventional artificial neural networks. Recent efforts have attempted to integrate high-performing vision transformer architectures into SNNs. However, existing spiking transformers struggle to extract local image features effectively because they rely on shallow convolutional modules and lack appropriate mathematical scaling methods to handle multi-scale inputs.

The article designs and evaluates SpikingResformer, a novel architecture that combines a multi-stage residual network backbone with a new spike-driven attention mechanism called Dual Spike Self-Attention (DSSA). The main objective is to eliminate architectural bottlenecks, improve accuracy, and reduce computational energy and parameter overhead on large-scale visual recognition tasks.

The researchers evaluated this framework through extensive direct-training experiments on the standard ImageNet benchmark, ablation studies on ImageNet100, and transfer learning evaluations across static and event-based image datasets. The proposed DSSA mechanism replaces floating-point operations and softmax functions with dual spike transformations, paired with statistically derived scaling factors to support varying feature map resolutions without gradient vanishing.

The findings show that SpikingResformer sets a new state of the art in SNN performance, achieving up to 79.40% top-1 accuracy on ImageNet with only 4 time-steps. Compared to prior spiking transformers, it consistently achieves higher accuracy with significantly fewer parameters and lower energy consumption. For example, the tiny variant delivers a 2.06% accuracy gain while saving 5.67 million parameters and roughly 37% of inference energy compared to baseline spiking transformers. Ablation tests confirmed that the multi-stage architecture, the group-wise convolutional feed-forward network, and the novel scaling factors are all essential for model convergence and top performance. Furthermore, transfer learning from pre-trained ImageNet weights achieved top results on static benchmarks (97.40% on CIFAR10 and 85.98% on CIFAR100) and static-derived event data.

These results demonstrate that energy-efficient spiking models can achieve competitive visual recognition accuracy without relying on costly floating-point computations, significantly lowering the barrier for deploying high-performance vision models on low-power neuromorphic edge devices. However, the evaluation also revealed a key limitation: while transfer learning works well on static datasets, it transfers poorly to real-world event datasets with rich temporal dynamics, such as DVSGesture, where accuracy dropped 5.9% behind direct training methods due to the absence of temporal features in pre-training. Moving forward, organizations exploring low-power neuromorphic vision should consider adopting multi-stage spiking transformer architectures for static imagery, while prioritizing further research into temporal pre-training strategies to unlock similar performance gains on dynamic event streams.

No sufficiently relevant recommendations were found.

Cover for SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks

Abstract

The remarkable success of Vision Transformers in Artificial Neural Networks (ANNs) has led to a growing interest in incorporating the self-attention mechanism and transformer-based architecture into Spiking Neural Networks (SNNs). While existing methods propose spiking self-attention mechanisms that are compatible with SNNs, they lack reasonable scaling methods, and the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting local features. To address these challenges, we propose a novel spiking self-attention mechanism named Dual Spike Self-Attention (DSSA) with a reasonable scaling method. Based on DSSA, we propose a novel spiking Vision Transformer architecture called SpikingResformer, which combines the ResNet-based multi-stage architecture with our proposed DSSA to improve both performance and energy efficiency while reducing parameters. Experimental results show that SpikingResformer achieves higher accuracy with fewer parameters and lower energy consumption than other spiking Vision Transformer counterparts. Notably, our SpikingResformer-L achieves 79.40% top-1 accuracy on ImageNet with 4 time-steps, which is the state-of-the-art result in the SNN field. Codes are available at https://github.com/xyshi2000/SpikingResformer

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminary
  • 4. Dual Spike Self-Attention
  • 4.1. Vanilla Self-Attention
  • 4.2. Dual Spike Self-Attention
  • 4.3. Scaling Factors in DSSA
  • 4.4. Spike-driven Characteristic of DSSA
  • 5. SpikingResformer
  • 5.1. Overall Architecture
  • 5.2. Spiking Resformer Block
  • 6. Experiments
  • 6.1. ImageNet Classification
  • 6.2. Ablation Study
  • 6.3. Transfer Learning
  • 7. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Dual Spike Transformation

    model/method

    Dual Spike Transformation (DST) is a matrix transformation designed for Spiking Neural Networks (SNNs) that computes self-attention projections using only binary spike inputs, thereby avoiding floating-point matrix multiplications.

    Given two binary spike tensors X∈{0,1}T×p×m\mathbf{X} \in \{0, 1\}^{T \times p \times m} and Y∈{0,1}T×m×q\mathbf{Y} \in \{0, 1\}^{T \times m \times q}, where TT is the number of time steps and p,m,qp, m, q are arbitrary dimensions, and a generalized linear transformation f(⋅)f(\cdot) on Y\mathbf{Y} parameterized by weight matrix W∈Rq×q\mathbf{W} \in \mathbb{R}^{q \times q}, standard DST is defined as:

    DST(X,Y;f(⋅))=Xf(Y)=XYW\mathrm{DST}(\mathbf{X}, \mathbf{Y}; f(\cdot)) = \mathbf{X} f(\mathbf{Y}) = \mathbf{X}\mathbf{Y}\mathbf{W}

    Similarly, for Y∈{0,1}T×q×m\mathbf{Y} \in \{0, 1\}^{T \times q \times m} and W∈Rm×m\mathbf{W} \in \mathbb{R}^{m \times m}, transposed DST (DSTT\mathrm{DST}_T) is defined as:

    DSTT(X,Y;f(⋅))=Xf(Y)T=XWTYT\mathrm{DST}_T(\mathbf{X}, \mathbf{Y}; f(\cdot)) = \mathbf{X} f(\mathbf{Y})^{\mathrm{T}} = \mathbf{X}\mathbf{W}^{\mathrm{T}}\mathbf{Y}^{\mathrm{T}}

    Because both X\mathbf{X} and Y\mathbf{Y} are binary spike matrices, the matrix multiplications evaluate to selective summations of weight parameters, making the operation fully spike-driven and compatible with neuromorphic hardware.

  2. Knowl 2 — Dual Spike Self-Attention Mechanism

    model/method

    Dual Spike Self-Attention (DSSA) is a spiking self-attention mechanism that replaces floating-point matrix multiplications and softmax normalizations with Dual Spike Transformations (DST) and spiking neuron activations.

    Given an input spike tensor X∈{0,1}T×HW×d\mathbf{X} \in \{0, 1\}^{T \times HW \times d}, where TT denotes time steps, HWHW is the spatial sequence length (height ×\times width), and dd is the embedding dimension, the spiking attention map is computed as:

    AttnMap(X)=SN(DSTT(X,X;f(⋅))⋅c1)\mathrm{AttnMap}(\mathbf{X}) = \mathrm{SN}\left(\mathrm{DST}_T(\mathbf{X}, \mathbf{X}; f(\cdot)) \cdot c_1\right)

    f(X)=BN(Convp(X))f(\mathbf{X}) = \mathrm{BN}(\mathrm{Conv}_p(\mathbf{X}))

    where Convp(⋅)\mathrm{Conv}_p(\cdot) denotes a p×pp \times p convolution with a stride of pp used to downsample the spatial dimension, BN(⋅)\mathrm{BN}(\cdot) denotes batch normalization, SN(⋅)\mathrm{SN}(\cdot) denotes a spiking neuron activation layer (such as Leaky Integrate-and-Fire), and c1c_1 is a scaling factor. The resulting attention map AttnMap(X)∈{0,1}T×HW×(HW/p2)\mathrm{AttnMap}(\mathbf{X}) \in \{0, 1\}^{T \times HW \times (HW/p^2)} is binary-valued, where each spike sij∈{0,1}s_{ij} \in \{0, 1\} indicates whether patch ii attends to downsampled patch jj.

    The final DSSA output is obtained by applying DST between the spiking attention map and the input spikes:

    DSSA(X)=SN(DST(AttnMap(X),X;f(⋅))⋅c2)\mathrm{DSSA}(\mathbf{X}) = \mathrm{SN}\left(\mathrm{DST}(\mathrm{AttnMap}(\mathbf{X}), \mathbf{X}; f(\cdot)) \cdot c_2\right)

    where c2c_2 is the second scaling factor.

  3. Knowl 3 — Statistical Scaling Factors for Dual Spike Transformation

    theoretical result

    In Dual Spike Self-Attention (DSSA), variance scaling is required to prevent surrogate gradients from vanishing across multi-scale feature maps. Unlike vanilla self-attention where inputs are assumed to have zero mean and unit variance, the input tensors in DST are binary spikes governed by Bernoulli distributions.

    Let X∈{0,1}T×p×m\mathbf{X} \in \{0, 1\}^{T \times p \times m} and Y∈{0,1}T×m×q\mathbf{Y} \in \{0, 1\}^{T \times m \times q} be independent random spike tensors whose individual elements follow a Bernoulli distribution xix,jx[t]∼Bernoulli(fx)x_{i_x, j_x}[t] \sim \mathrm{Bernoulli}(f_x) with mean firing rate fxf_x. Let f(⋅)f(\cdot) be a linear transformation with weight matrix W∈Rq×q\mathbf{W} \in \mathbb{R}^{q \times q} such that the output f(Y)f(\mathbf{Y}) has mean 0 and variance 1. Then the output current I=DST(X,Y;f(⋅))\mathbf{I} = \mathrm{DST}(\mathbf{X}, \mathbf{Y}; f(\cdot)) satisfies:

    E[IiI,jI[t]]=0,Var(IiI,jI[t])=fxm\mathbb{E}[I_{i_I, j_I}[t]] = 0, \quad \mathrm{Var}(I_{i_I, j_I}[t]) = f_x m

    Similarly, for I=DSTT(X,Y;f(⋅))\mathbf{I} = \mathrm{DST}_T(\mathbf{X}, \mathbf{Y}; f(\cdot)) with Y∈{0,1}T×q×m\mathbf{Y} \in \{0, 1\}^{T \times q \times m} and W∈Rm×m\mathbf{W} \in \mathbb{R}^{m \times m}, E[IiI,jI[t]]=0\mathbb{E}[I_{i_I, j_I}[t]] = 0 and Var(IiI,jI[t])=fxm\mathrm{Var}(I_{i_I, j_I}[t]) = f_x m.

    To normalize the variance of the input current to 1 before entering the spiking neuron layers, the scaling factors c1c_1 and c2c_2 in DSSA are set to:

    c1=1fXdc_1 = \frac{1}{\sqrt{f_X d}}

    c2=1fAttnHWp2c_2 = \frac{1}{\sqrt{f_{\mathrm{Attn}} \frac{HW}{p^2}}}

    where fXf_X and fAttnf_{\mathrm{Attn}} denote the average firing rates of the input spike tensor X\mathbf{X} and the attention map AttnMap(X)\mathrm{AttnMap}(\mathbf{X}), dd is the embedding dimension, and HW/p2HW/p^2 is the reduced spatial dimension from the p×pp \times p convolution with stride pp.

  4. Knowl 4 — Spike-Driven Formulation of Dual Spike Transformation

    theoretical result

    A spiking neural network is spike-driven if the input current Ii[t]I_i[t] to postsynaptic neuron ii at time step tt is computed strictly as a sparse summation of synaptic weights triggered by incoming spikes:

    Ii[t]=∑jwi,jsj[t]=∑j,sj[t]≠0wi,jI_i[t] = \sum_{j} w_{i,j} s_j[t] = \sum_{j, s_j[t] \neq 0} w_{i,j}

    where sj[t]∈{0,1}s_j[t] \in \{0, 1\} is the binary spike output of presynaptic neuron jj and wi,jw_{i,j} is the synaptic weight.

    Dual Spike Transformation satisfies this definition through the logical conjunction of dual spikes. For I=DSTT(X,X;f(⋅))=XWTYT\mathbf{I} = \mathrm{DST}_T(\mathbf{X}, \mathbf{X}; f(\cdot)) = \mathbf{X}\mathbf{W}^{\mathrm{T}}\mathbf{Y}^{\mathrm{T}}, the input current is:

    Ii,j[t]=∑k=1m∑l=1mxi,k[t]wl,kyj,l[t]=∑k,l(xi,k[t]∧yj,l[t])≠0wl,kI_{i,j}[t] = \sum_{k=1}^m \sum_{l=1}^m x_{i,k}[t] w_{l,k} y_{j,l}[t] = \sum_{\substack{k,l \\ (x_{i,k}[t] \land y_{j,l}[t]) \neq 0}} w_{l,k}

    For I=DST(AttnMap(X),X;f(⋅))=XYW\mathbf{I} = \mathrm{DST}(\mathrm{AttnMap}(\mathbf{X}), \mathbf{X}; f(\cdot)) = \mathbf{X}\mathbf{Y}\mathbf{W}, the input current is:

    Ii,j[t]=∑k=1m∑l=1qxi,k[t]yk,l[t]wl,j=∑k,l(xi,k[t]∧yk,l[t])≠0wl,jI_{i,j}[t] = \sum_{k=1}^m \sum_{l=1}^q x_{i,k}[t] y_{k,l}[t] w_{l,j} = \sum_{\substack{k,l \\ (x_{i,k}[t] \land y_{k,l}[t]) \neq 0}} w_{l,j}

    In both equations, synaptic accumulation occurs only when the logical AND of the two spikes is active (x[t]∧y[t]=1x[t] \land y[t] = 1). Thus, the computation requires only synaptic additions and no floating-point multiplications.

  5. Knowl 5 — Group-Wise Spiking Feed-Forward Network

    model/method

    The Group-Wise Spiking Feed-Forward Network (GWSFFN) extends standard spiking feed-forward networks by incorporating a 3×33 \times 3 group-wise convolution with a residual shortcut between two point-wise linear layers, enhancing local feature extraction while constraining computational cost.

    Let X\mathbf{X} be the intermediate feature tensor. GWSFFN is formulated as:

    FFLi(X)=BN(Conv1(SN(X))),i∈{1,2}\mathrm{FFL}_i(\mathbf{X}) = \mathrm{BN}(\mathrm{Conv}_1(\mathrm{SN}(\mathbf{X}))), \quad i \in \{1, 2\}

    GWL(X)=BN(GWConv(SN(X)))+X\mathrm{GWL}(\mathbf{X}) = \mathrm{BN}(\mathrm{GWConv}(\mathrm{SN}(\mathbf{X}))) + \mathbf{X}

    GWSFFN(X)=FFL2(GWL(FFL1(X)))\mathrm{GWSFFN}(\mathbf{X}) = \mathrm{FFL}_2(\mathrm{GWL}(\mathrm{FFL}_1(\mathbf{X})))

    where Conv1(⋅)\mathrm{Conv}_1(\cdot) is a 1×11 \times 1 point-wise convolution (linear projection), SN(⋅)\mathrm{SN}(\cdot) denotes spiking neuron activation, BN(⋅)\mathrm{BN}(\cdot) is batch normalization, and GWConv(⋅)\mathrm{GWConv}(\cdot) is a 3×33 \times 3 group-wise convolution with group size G=64G = 64 channels per group. The first feed-forward layer FFL1\mathrm{FFL}_1 expands the channel dimension by an expansion ratio of R=4R = 4, and FFL2\mathrm{FFL}_2 reduces it back to the input dimension.

  6. Knowl 6 — SpikingResformer Architecture and Block Structure

    model/method

    SpikingResformer is a spiking Vision Transformer architecture that integrates a ResNet-style multi-stage convolutional hierarchy with spiking self-attention.

    The overall pipeline consists of:

    1. Stem: A 7×77 \times 7 convolution with stride 2 and batch normalization followed by 3×33 \times 3 max pooling with stride 2, extracting localized features and reducing a 224×224224 \times 224 input to a resolution of 56×5656 \times 56.
    2. Multi-Stage Backbone: Three hierarchical stages operating at resolutions 56×5656 \times 56, 28×2828 \times 28, and 14×1414 \times 14. A downsampling layer (3×33 \times 3 convolution with stride 2) is applied before Stage 2 and Stage 3, halving spatial resolution and doubling channel dimension.
    3. Classifier: A global average pooling layer followed by a fully connected linear classification head.

    Inside each stage, multiple Spiking Resformer blocks are stacked sequentially. Each block comprises a Multi-Head Dual Spike Self-Attention (MHDSSA) module and a Group-Wise Spiking Feed-Forward Network (GWSFFN) with residual identity connections:

    Yi=MHDSSA(Xi)+Xi\mathbf{Y}_i = \mathrm{MHDSSA}(\mathbf{X}_i) + \mathbf{X}_i

    Xi+1=GWSFFN(Yi)+Yi\mathbf{X}_{i+1} = \mathrm{GWSFFN}(\mathbf{Y}_i) + \mathbf{Y}_i

    where Xi\mathbf{X}_i is the input to the ii-th block, and MHDSSA\mathrm{MHDSSA} is defined across hh attention heads as:

    MHDSSA(X)=BN(Conv1([DSSAk(SN(X))]k=1h))\mathrm{MHDSSA}(\mathbf{X}) = \mathrm{BN}\left(\mathrm{Conv}_1\left([\mathrm{DSSA}_k(\mathrm{SN}(\mathbf{X}))]_{k=1}^h\right)\right)

    with [… ][\dots] denoting concatenation across heads, Conv1\mathrm{Conv}_1 denoting a 1×11 \times 1 point-wise projection, and DSSAk\mathrm{DSSA}_k denoting the single-head Dual Spike Self-Attention module.

  7. Knowl 7 — Architectural Configurations of SpikingResformer Variants

    data/table

    SpikingResformer is instantiated in four model scales: Tiny (Ti), Small (S), Medium (M), and Large (L). Across all variants, the stem consists of a Conv 7×77 \times 7 (stride 2) and Maxpooling 3×33 \times 3 (stride 2). GWSFFN uses expansion ratio Ri=4R_i = 4 and group size Gi=64G_i = 64 across all stages i∈{1,2,3}i \in \{1, 2, 3\}. Downsampling layers between stages use Conv 3×33 \times 3 with stride 2. In stage ii, DiD_i is embedding dimension, HiH_i is head count, and pip_i is the convolution patch size in DST.

    Stage Output Size Layer / Param SpikingResformer-Ti SpikingResformer-S SpikingResformer-M SpikingResformer-L
    Stem 56×5656 \times 56 Stem Conv 7×77 \times 7, stride 2, Maxpooling 3×33 \times 3, stride 2
    Stage 1 56×5656 \times 56 MHDSSA + GWSFFN [D1=64H1=1,p1=4]×1\begin{bmatrix} D_1=64 \\ H_1=1, p_1=4 \end{bmatrix} \times 1 [D1=64H1=1,p1=4]×1\begin{bmatrix} D_1=64 \\ H_1=1, p_1=4 \end{bmatrix} \times 1 [D1=64H1=1,p1=4]×1\begin{bmatrix} D_1=64 \\ H_1=1, p_1=4 \end{bmatrix} \times 1 [D1=128H1=1,p1=4]×1\begin{bmatrix} D_1=128 \\ H_1=1, p_1=4 \end{bmatrix} \times 1
    Stage 2 28×2828 \times 28 Downsample Conv 3×3,1923 \times 3, 192, stride 2 Conv 3×3,2563 \times 3, 256, stride 2 Conv 3×3,3843 \times 3, 384, stride 2 Conv 3×3,5123 \times 3, 512, stride 2
    MHDSSA + GWSFFN [D2=192H2=3,p2=2]×2\begin{bmatrix} D_2=192 \\ H_2=3, p_2=2 \end{bmatrix} \times 2 [D2=256H2=4,p2=2]×2\begin{bmatrix} D_2=256 \\ H_2=4, p_2=2 \end{bmatrix} \times 2 [D2=384H2=6,p2=2]×2\begin{bmatrix} D_2=384 \\ H_2=6, p_2=2 \end{bmatrix} \times 2 [D2=512H2=8,p2=2]×2\begin{bmatrix} D_2=512 \\ H_2=8, p_2=2 \end{bmatrix} \times 2
    Stage 3 14×1414 \times 14 Downsample Conv 3×3,3843 \times 3, 384, stride 2 Conv 3×3,5123 \times 3, 512, stride 2 Conv 3×3,7683 \times 3, 768, stride 2 Conv 3×3,10243 \times 3, 1024, stride 2
    MHDSSA + GWSFFN [D3=384H3=6,p3=1]×3\begin{bmatrix} D_3=384 \\ H_3=6, p_3=1 \end{bmatrix} \times 3 [D3=512H3=8,p3=1]×3\begin{bmatrix} D_3=512 \\ H_3=8, p_3=1 \end{bmatrix} \times 3 [D3=768H3=12,p3=1]×3\begin{bmatrix} D_3=768 \\ H_3=12, p_3=1 \end{bmatrix} \times 3 [D3=1024H3=16,p3=1]×3\begin{bmatrix} D_3=1024 \\ H_3=16, p_3=1 \end{bmatrix} \times 3
    Classifier 1×11 \times 1 Linear 1000-FC
  8. Knowl 8 — ImageNet Classification Performance and Energy Efficiency of SpikingResformer

    data/table

    SpikingResformer models were evaluated on the ImageNet validation set at time steps T=4T=4 using default 224×224224 \times 224 input resolution (and 288×288288 \times 288 denoted by †\dagger). Energy is estimated from Synaptic Operations (SOPs).

    Method Type Architecture TT Param (M) SOPs (G) Energy (mJ) Top-1 Acc. (%)
    Spiking ResNet ANN-to-SNN ResNet-34 350 21.79 65.28 59.30 71.61
    Spiking ResNet ANN-to-SNN ResNet-50 350 25.56 78.29 70.93 72.75
    STBP-tdBN Direct Training Spiking ResNet-34 6 21.79 6.50 6.39 63.72
    TET Direct Training Spiking ResNet-34 6 21.79 - - 64.79
    TET Direct Training SEW ResNet-34 4 21.79 - - 68.00
    SEW ResNet Direct Training SEW ResNet-34 4 21.79 3.88 4.04 67.04
    SEW ResNet Direct Training SEW ResNet-50 4 25.56 4.83 4.89 67.78
    SEW ResNet Direct Training SEW ResNet-101 4 44.55 9.30 8.91 68.76
    SEW ResNet Direct Training SEW ResNet-152 4 60.19 13.72 12.89 69.26
    Spikformer Direct Training Spikformer-8-384 4 16.81 6.82 7.73 70.24
    Spikformer Direct Training Spikformer-8-512 4 29.68 11.09 11.58 73.38
    Spikformer Direct Training Spikformer-8-768 4 66.34 22.09 21.48 74.81
    Spikingformer Direct Training Spikingformer-8-384 4 16.81 - 4.69 72.45
    Spikingformer Direct Training Spikingformer-8-512 4 29.68 - 7.46 74.79
    Spikingformer Direct Training Spikingformer-8-768 4 66.34 - 13.68 75.85
    Spike-driven Transformer Direct Training Spike-driven Transformer-8-384 4 16.81 - 3.90 72.28
    Spike-driven Transformer Direct Training Spike-driven Transformer-8-512 4 29.68 - 4.50 74.57
    Spike-driven Transformer Direct Training Spike-driven Transformer-8-768 4 66.34 - 6.09 76.32 / 77.07†77.07^\dagger
    SpikingResformer (Ours) Direct Training SpikingResformer-Ti 4 11.14 2.73 / 4.71†4.71^\dagger 2.46 / 4.24†4.24^\dagger 74.34 / 75.57†75.57^\dagger
    SpikingResformer (Ours) Direct Training SpikingResformer-S 4 17.76 3.74 / 6.40†6.40^\dagger 3.37 / 5.76†5.76^\dagger 75.95 / 76.90†76.90^\dagger
    SpikingResformer (Ours) Direct Training SpikingResformer-M 4 35.52 6.07 / 10.24†10.24^\dagger 5.46 / 9.22†9.22^\dagger 77.24 / 78.06†78.06^\dagger
    SpikingResformer (Ours) Direct Training SpikingResformer-L 4 60.38 9.74 / 16.40†16.40^\dagger 8.76 / 14.76†14.76^\dagger 78.77 / 79.40†79.40^\dagger

    SpikingResformer outperforms all prior SNNs. SpikingResformer-Ti (11.14M params, 2.46 mJ) achieves 74.34% accuracy, exceeding Spike-driven Transformer-8-384 by 2.06% while using 5.67M fewer parameters and saving 1.44 mJ of energy. SpikingResformer-L achieves 79.40% top-1 accuracy at 288×288288 \times 288 resolution.

  9. Knowl 9 — Component Ablation Study of SpikingResformer on ImageNet-100

    data/table

    Ablation experiments were conducted on the ImageNet-100 dataset to isolate the contributions of the multi-stage architecture, the group-wise convolution in GWSFFN, the p×pp \times p downsampling convolution, and the variance scaling factor in DSSA. All variant architectures were configured to have parameter counts comparable to SpikingResformer-S.

    Model SOPs (G) Energy (mJ) Acc. (%)
    SpikingResformer-S 2.43 2.18 88.06
    w/o multi-stage architecture 1.84 1.66 85.32
    w/o group-wise convolution 2.37 2.13 84.64
    w/o DSSA (replaced with SSA / SDSA) - - not converge
    w/o p×pp \times p convolution 4.34 3.91 86.14
    w/o scaling (or with 1/d1/\sqrt{d}) - - not converge

    The ablation reveals:

    1. The multi-stage architecture improves accuracy by 2.74% over single-scale Spikingformer-style backbones.
    2. The 3×33 \times 3 group-wise convolution in GWSFFN improves accuracy by 3.42% compared to point-wise feed-forward layers.
    3. Replacing DSSA with existing spiking self-attention mechanisms (SSA from Spikformer or SDSA from Spike-driven Transformer) causes training divergence on multi-scale inputs due to improper or missing variance scaling.
    4. Replacing p×pp \times p convolutions with 1×11 \times 1 convolutions in DST reduces accuracy by 1.92% while increasing energy from 2.18 mJ to 3.91 mJ.
    5. Omitting variance scaling or using vanilla 1/d1/\sqrt{d} scaling prevents model convergence.
  10. Knowl 10 — Transfer Learning Performance and Neuromorphic Domain Gap of SpikingResformer

    empirical result

    Fine-tuning ImageNet pre-trained SpikingResformer models on downstream static and neuromorphic datasets yields the following top-1 classification accuracies:

    Method Type CIFAR10 CIFAR100 CIFAR10-DVS DVSGesture
    TT Acc. TT Acc. TT Acc. TT Acc.
    STBP-tdBN Direct Training 6 93.16% - - 10 67.8% 40 96.87%
    PLIF Direct Training 8 93.50% - - 20 74.8% 20 97.57%
    Dspike Direct Training 6 94.25% 6 74.24% 10 75.4% - -
    Spikformer Direct Training 4 95.19% 4 77.86% 16 80.6% 16 97.9%
    Spikingformer Direct Training 4 95.61% 4 79.09% 16 81.3% 16 98.3%
    Spike-driven Transformer Direct Training 4 95.6% 4 78.4% 16 80.0% 16 99.3%
    Spikformer Transfer Learning 4 97.03% 4 83.83% - - - -
    SpikingResformer (Ours) Transfer Learning 4 97.40% 4 85.98% 10 84.8% 10 93.4%

    Transfer learning on static datasets achieves 97.40% on CIFAR-10 and 85.98% on CIFAR-100, outperforming Spikformer transfer learning by 0.37% and 2.15%, respectively. On CIFAR10-DVS, transfer learning achieves 84.8% accuracy (outperforming direct training by 3.5%).

    However, transfer learning from static image pre-training on DVSGesture achieves only 93.4%, falling behind direct training (99.3%). This occurs because CIFAR10-DVS is converted from static images and lacks native temporal dynamics, enabling static pre-trained weights to transfer effectively, whereas DVSGesture is recorded with dynamic vision sensors containing rich temporal information that static image pre-training cannot capture.

Coverage note — None was omitted; all primary architectural definitions, theoretical scaling factor formulations, spike-driven proofs, ImageNet comparisons, ablation studies, and transfer learning evaluations are included.

References

  1. 1.Arnon Amir, Brian Taba, David Berg, Timothy Melano, Jeffrey McKinstry, Carmelo Di Nolfo, Tapan Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, et al. A low power, fully event-based gesture recognition system. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7243–7252, 2017.
  2. 2.Tong Bu, Jianhao Ding, Zhaofei Yu, and Tiejun Huang. Optimized potential initialization for low-latency spiking neural networks. In In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11–20, 2022.
  3. 3.Tong Bu, Wei Fang, Jianhao Ding, PengLin Dai, Zhaofei Yu, and Tiejun Huang. Optimal ANN-SNN conversion for high-accuracy and ultra-low-latency spiking neural networks. In Proceedings of the International Conference on Learning Representations, pages 1–19, 2023.
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  5. 5.Shikuang Deng and Shi Gu. Optimal conversion of conventional artificial neural networks to spiking neural networks. In Proceedings of the International Conference on Learning Representations, pages 1–14, 2021.
  6. 6.Shikuang Deng, Yuhang Li, Shanghang Zhang, and Shi Gu. Temporal efficient training of spiking neural network via gradient re-weighting. In Proceedings of the International Conference on Learning Representations, pages 1–15, 2022.
  7. 7.Jianhao Ding, Zhaofei Yu, Yonghong Tian, and Tiejun Huang. Optimal ANN-SNN conversion for fast and accurate inference in deep spiking neural networks. In Proceedings of the International Joint Conference on Artificial Intelligence, pages 2328–2336, 2021.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations, pages 1–22, 2021.
  9. 9.Wei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang, Timothee Masquelier, and Yonghong Tian. Deep residual learning in spiking neural networks. In Advances in Neural Information Processing Systems, pages 21056–21069, 2021.
  10. 10.Wei Fang, Zhaofei Yu, Yanqi Chen, Timothee Masquelier, Tiejun Huang, and Yonghong Tian. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2661–2671, 2021.
  11. 11.Steve B Furber, Francesco Galluppi, Steve Temple, and Luis A Plana. The spiNNaker project. Proceedings of the IEEE, 102(5):652–665, 2014.
  12. 12.Yufei Guo, Yuhan Zhang, Yuanpei Chen, Weihang Peng, Xiaode Liu, Liwen Zhang, Xuhui Huang, and Zhe Ma. Membrane potential batch normalization for spiking neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19420–19430, 2023.
  13. 13.Zecheng Hao, Tong Bu, Jianhao Ding, Tiejun Huang, and Zhaofei Yu. Reducing ANN-SNN conversion error through residual membrane potential. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11–21, 2023.
  14. 14.Yangfan Hu, Huajin Tang, and Gang Pan. Spiking deep residual networks. IEEE Transactions on Neural Networks and Learning Systems, 34(8):5200–5205, 2021.
  15. 15.Yangfan Hu, Qian Zheng, Xudong Jiang, and Gang Pan. Fast-SNN: Fast spiking neural network by converting quantized ann. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12):14546–14562, 2023.
  16. 16.Haiyan Jiang, Srinivas Anumasa, Giulia De Masi, Huan Xiong, and Bin Gu. A unified optimization framework of ANN-SNN conversion: Towards optimal mapping from activation values to firing rates. In Proceedings of the International Conference on Machine Learning, pages 14945–14974, 2023.
  17. 17.Seijoon Kim, Seongsik Park, Byunggook Na, and Sungroh Yoon. Spiking-YOLO: Spiking neural network for energy-efficient object detection. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11270–11277, 2020.
  18. 18.Paul Kirkland, Gaetano Di Caterina, John Soraghan, and George Matich. SpikeSEG: Spiking segmentation via STDP saliency mapping. In Proceedings of the International Joint Conference on Neural Networks, pages 1–8, 2020.
  19. 19.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  20. 20.Yuxiang Lan, Yachao Zhang, Xu Ma, Yanyun Qu, and Yun Fu. Efficient converted spiking neural network for 3d and 2d classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9211–9220, 2023.
  21. 21.Chankyu Lee, Syed Shakib Sarwar, Priyadarshini Panda, Gopalakrishnan Srinivasan, and Kaushik Roy. Enabling spike-based backpropagation for training deep neural network architectures. Frontiers in Neuroscience, 14:119, 2020.
  22. 22.Chen Li, Edward Jones, and Steve Furber. Unleashing the potential of spiking neural networks by dynamic confidence. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13350–13360, 2023.
  23. 23.Hongmin Li, Hanchao Liu, Xiangyang Ji, Guoqi Li, and Luping Shi. CIFAR10-DVS: an event-stream dataset for object classification. Frontiers in Neuroscience, 11:309, 2017.
  24. 24.Yuhang Li, Shikuang Deng, Xin Dong, Ruihao Gong, and Shi Gu. A free lunch from ANN: Towards efficient, accurate spiking neural networks calibration. In Proceedings of the International Conference on Machine Learning, pages 6316–6325, 2021.
  25. 25.Yuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng, Yongqing Hai, and Shi Gu. Differentiable spike: Rethinking gradient-descent for training spiking neural networks. Advances in Neural Information Processing Systems, 34: 23426–23439, 2021.
  26. 26.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.
  27. 27.Wolfgang Maass. Networks of spiking neurons: the third generation of neural network models. Neural networks, 10 (9):1659–1671, 1997.
  28. 28.Qingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang, Zhouchen Lin, and Zhi-Quan Luo. Towards memory-and time-efficient backpropagation for training spiking neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6166–6176, 2023.
  29. 29.Paul A Merolla, John V Arthur, Rodrigo Alvarez-Icaza, Andrew S Cassidy, Jun Sawada, Filipp Akopyan, Bryan L Jackson, Nabil Imam, Chen Guo, Yutaka Nakamura, et al. A million spiking-neuron integrated circuit with a scalable communication network and interface. Science, 345(6197):668–673, 2014.
  30. 30.Emre O Neftci, Hesham Mostafa, and Friedemann Zenke. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36 (6):51–63, 2019.
  31. 31.Kinjal Patel, Eric Hunsberger, Sean Batir, and Chris Eliasmith. A spiking neural network for image segmentation. arXiv preprint arXiv:2106.08921, pages 1–25, 2021.
  32. 32.Jing Pei, Lei Deng, Sen Song, Mingguo Zhao, Youhui Zhang, Shuang Wu, Guanrui Wang, Zhe Zou, Zhenzhi Wu, Wei He, et al. Towards artificial general intelligence with hybrid tianjic chip architecture. Nature, 572(7767):106–111, 2019.
  33. 33.Catherine D Schuman, Thomas E Potok, Robert M Patton, J Douglas Birdwell, Mark E Dean, Garrett S Rose, and James S Plank. A survey of neuromorphic computing and neural networks in hardware. arXiv preprint arXiv:1705.06963, pages 1–88, 2017.
  34. 34.Abhronil Sengupta, Yuting Ye, Robert Wang, Chiao Liu, and Kaushik Roy. Going deeper in spiking neural networks: VGG and residual architectures. Frontiers in Neuroscience, 13:95, 2019.
  35. 35.Qiaoyi Su, Yuhong Chou, Yifan Hu, Jianing Li, Shijie Mei, Ziyang Zhang, and Guoqi Li. Deep directly-trained spiking neural networks for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6555–6565, 2023.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30:1–11, 2017.
  37. 37.Xiao Wang, Zongzhen Wu, Yao Rong, Lin Zhu, Bo Jiang, Jin Tang, and Yonghong Tian. SSTFormer: Bridging spiking neural network and memory support transformer for frame-event based recognition. arXiv preprint arXiv:2308.04369, pages 1–12, 2023.
  38. 38.Wenjie Wei, Malu Zhang, Hong Qu, Ammar Belatreche, Jian Zhang, and Hong Chen. Temporal-coded spiking neural networks with dynamic firing threshold: Learning with event-driven backpropagation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10552–10562, 2023.
  39. 39.Jibin Wu, Chenglin Xu, Xiao Han, Daquan Zhou, Malu Zhang, Haizhou Li, and Kay Chen Tan. Progressive tandem learning for pattern recognition with deep spiking neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7824–7840, 2021.
  40. 40.Qi Xu, Yaxin Li, Jiangrong Shen, Jian K Liu, Huajin Tang, and Gang Pan. Constructing deep spiking neural networks from artificial neural networks with knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7886–7895, 2023.
  41. 41.Man Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan, Yonghong Tian, Bo Xu, and Guoqi Li. Spike-driven transformer. In Advances in neural information processing systems, pages 1–20, 2023.
  42. 42.Friedemann Zenke and Tim P Vogels. The remarkable robustness of surrogate gradient learning for instilling complex function in spiking neural networks. Neural computation, 33 (4):899–925, 2021.
  43. 43.Hanle Zheng, Yujie Wu, Lei Deng, Yifan Hu, and Guoqi Li. Going deeper with directly-trained larger spiking neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11062–11070, 2021.
  44. 44.Chenlin Zhou, Liutao Yu, Zhaokun Zhou, Han Zhang, Zhengyu Ma, Huihui Zhou, and Yonghong Tian. Spikingformer: Spike-driven residual learning for transformer-based spiking neural network. arXiv preprint arXiv:2304.11954, pages 1–16, 2023.
  45. 45.Zhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang, Shuicheng Yan, Yonghong Tian, and Li Yuan. Spikformer: When spiking neural network meets transformer. In Proceedings of the International Conference on Learning Representations, pages 1–17, 2023.
  46. 46.Yaoyu Zhu, Zhaofei Yu, Wei Fang, Xiaodong Xie, Tiejun Huang, and Timothee Masquelier. Training spiking neural networks with event-driven backpropagation. In Advances in Neural Information Processing Systems, pages 30528–30541, 2022.
  47. 47.Yaoyu Zhu, Wei Fang, Xiaodong Xie, Tiejun Huang, and Zhaofei Yu. Exploring loss functions for time-based training strategy in spiking neural networks. In Advances in Neural Information Processing Systems, pages 65366–65379, 2023.
  48. 48.Shihao Zou, Yuxuan Mu, Xinxin Zuo, Sen Wang, and Li Cheng. Event-based human pose tracking by spiking spatiotemporal transformer. arXiv preprint arXiv:2303.09681, pages 1–12, 2023.

Citation

MLA
Shi, X., et al. “SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks”. arXiv, 2024, http://arxiv.org/abs/2403.14302v2.
APA
Shi, X., Hao, Z., & Yu, Z. (2024). SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks. arXiv. http://arxiv.org/abs/2403.14302v2
Chicago
Shi, X., Z. Hao, and Z. Yu. 2024. “SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks”. arXiv. http://arxiv.org/abs/2403.14302v2.
Harvard
Shi, X., Hao, Z. and Yu, Z. (2024) “SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.14302v2.
Vancouver
1. Shi X, Hao Z, Yu Z (2024) SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks. arXiv

BibTeX

@article{shi2024spikingresformer,
  title = {SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks},
  author = {Shi, Xinyu and Hao, Zecheng and Yu, Zhaofei},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.14302v2},
  eprint = {2403.14302}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE