Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation

Qingyan MengMingqing XiaoShen YanYisen WangZhouchen LinZhi-Quan Luo

article2022CVPR183 citations

Presents the Differentiation on Spike Representation method to train spiking neural networks via sub-differentiable mapping of firing rates, eliminating expensive backpropagation through time to achieve ANN-competitive accuracy at low latency across static and neuromorphic benchmarks.

Listen

Spiking neural networks (SNNs) are brain-inspired artificial intelligence models that process information using discrete pulses or spikes. When deployed on specialized neuromorphic hardware, they offer substantial energy efficiency compared to conventional deep artificial neural networks (ANNs). However, training high-performing SNNs remains difficult because spike generation is non-differentiable, preventing standard gradient-based optimization. Existing training approaches either require hundreds of simulation time steps to match standard model accuracy, causing high latency, or rely on computationally heavy surrogate gradient methods that struggle to scale and match top-tier performance.

The article introduces and evaluates the Differentiation on Spike Representation (DSR) method, a training technique designed to deliver high accuracy with low latency across both standard image datasets and neuromorphic event-based streams.

The proposed approach encodes temporal spike trains into rate-based continuous representations and mathematically maps the forward network operations to sub-differentiable functions. Gradients are then propagated directly through these mappings across network layers, completely bypassing temporal backpropagation. To counter the representation error that arises when simulating with very few time steps, the authors introduce two mechanisms: dynamically training the firing thresholds of neurons with regularized optimization, and introducing a firing-control hyperparameter to cut quantization error.

The evaluation demonstrates that DSR achieves state-of-the-art accuracy across several benchmark datasets while maintaining short simulation latencies of 5 to 20 time steps. On CIFAR-100, the method achieves 78.50% accuracy, surpassing prior spiking approaches by 5% to 10% and matching standard ANN baselines. On CIFAR-10, it reaches 95.40% accuracy at 20 time steps and maintains over 94.4% accuracy down to an ultra-low latency of 5 time steps. On neuromorphic DVS-CIFAR10 data, it attains 77.27% accuracy, outperforming competing methods. Furthermore, the approach scales successfully to deep residual networks exceeding 100 layers without performance degradation.

These results show that spiking neural networks can achieve standard deep learning accuracy without incurring large computational and memory overheads during training or long inference delays. By enabling faster, low-power processing, the method improves the viability of deploying energy-efficient neuromorphic systems in edge computing and real-time sensory applications.

Organizations developing low-power AI systems should consider adopting representation-based differentiation frameworks for neuromorphic applications and evaluating these models on target hardware platforms. Further research and hardware-in-the-loop validation are recommended to assess real-world energy savings and performance trade-offs under extreme latency conditions below five time steps, where representation errors may increase.

arXiv: 2205.00459
Cover for Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation

Abstract

Spiking Neural Network (SNN) is a promising energy-efficient AI model when implemented on neuromorphic hardware. However, it is a challenge to efficiently train SNNs due to their non-differentiability. Most existing methods either suffer from high latency (i.e., long simulation time steps), or cannot achieve as high performance as Artificial Neural Networks (ANNs). In this paper, we propose the Differentiation on Spike Representation (DSR) method, which could achieve high performance that is competitive to ANNs yet with low latency. First, we encode the spike trains into spike representation using (weighted) firing rate coding. Based on the spike representation, we systematically derive that the spiking dynamics with common neural models can be represented as some sub-differentiable mapping. With this viewpoint, our proposed DSR method trains SNNs through gradients of the mapping and avoids the common non-differentiability problem in SNN training. Then we analyze the error when representing the specific mapping with the forward computation of the SNN. To reduce such error, we propose to train the spike threshold in each layer, and to introduce a new hyperparameter for the neural models. With these components, the DSR method can achieve state-of-the-art SNN performance with low latency on both static and neuromorphic datasets, including CIFAR-10, CIFAR-100, ImageNet, and DVS-CIFAR10.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Differentiation on Spike Representation (DSR) Method
  • 3.1. Spiking Neural Models
  • 3.2. Forward Pass
  • 3.3. Spike Representation
  • 3.3.1 Spike Representation for the LIF Model
  • 3.3.2 Spike Representation for the IF Model
  • 3.4. Differentiation on Spike Representation
  • 4. Reducing Representation Error
  • 5. Experiments
  • 5.1. Comparison to the State-of-the-Art
  • 5.2. Model Validation and Ablation Study
  • Effectiveness of the Proposed Method with Low Latency
  • 6. Conclusion and Discussions
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — Differentiation on Spike Representation Framework

    model/method

    The Differentiation on Spike Representation (DSR) framework trains feedforward Spiking Neural Networks (SNNs) by decoupling the temporal forward simulation of spiking neurons from spatial gradient backpropagation.

    For an LL-layer feedforward SNN, let si=(si[1],…,si[N])\mathbf{s}^i = (\mathbf{s}^i[1], \dots, \mathbf{s}^i[N]) denote the output spike trains of layer ii across NN discrete time steps, and let Wi\mathbf{W}^i be the synaptic weight matrix connecting layer i−1i-1 to layer ii. A spike representation operator oi=r(si)\mathbf{o}^i = r(\mathbf{s}^i) transforms spike trains into continuous-valued representations (such as firing rates). The layer-to-layer forward transformation over time is mapped to a continuous sub-differentiable activation mapping gWig_{\mathbf{W}^i}:

    oi=r(si)≈gWi(oi−1)=clamp⁡(Wioi−1,0,bi),i=1,2,…,L,\mathbf{o}^i = r(\mathbf{s}^i) \approx g_{\mathbf{W}^i}(\mathbf{o}^{i-1}) = \operatorname{clamp}\left(\mathbf{W}^i \mathbf{o}^{i-1}, 0, b_i\right), \quad i=1,2,\dots,L,

    where bi=Vthib_i = V_{\mathrm{th}}^i for Integrate-and-Fire (IF) neurons and bi=VthiΔtb_i = \frac{V_{\mathrm{th}}^i}{\Delta t} for Leaky Integrate-and-Fire (LIF) neurons, with VthiV_{\mathrm{th}}^i denoting the layer spike threshold, Δt\Delta t the discrete simulation step size, and clamp⁡(x,a,b)≜max⁡(a,min⁡(x,b))\operatorname{clamp}(x, a, b) \triangleq \max(a, \min(x, b)).

    Given a scalar loss function ℓ\ell computed on the final representation oL=r(sL)\mathbf{o}^L = r(\mathbf{s}^L), gradients with respect to parameters Wi\mathbf{W}^i are computed across spatial layers via standard backpropagation:

    ∂ℓ∂Wi=∂ℓ∂oi∂oi∂Wi,∂ℓ∂oi=∂ℓ∂oi+1∂oi+1∂oi.\frac{\partial \ell}{\partial \mathbf{W}^i} = \frac{\partial \ell}{\partial \mathbf{o}^i} \frac{\partial \mathbf{o}^i}{\partial \mathbf{W}^i}, \quad \frac{\partial \ell}{\partial \mathbf{o}^i} = \frac{\partial \ell}{\partial \mathbf{o}^{i+1}} \frac{\partial \mathbf{o}^{i+1}}{\partial \mathbf{o}^i}.

    This formulation eliminates the need to backpropagate through time (BPTT) and avoids calculating gradients through discontinuous Heaviside step functions.

  2. Knowl 2 — Sub-Differentiable Mapping Convergence for LIF Neurons

    theoretical result

    Consider an LL-layer feedforward SNN composed of Leaky Integrate-and-Fire (LIF) neurons whose discrete membrane potential dynamics are given by:

    Vi[n]=exp⁡(−Δtτi)Vi[n−1]+(1−exp⁡(−Δtτi))ΔtVthi−1Wisi−1[n]−Vthisi[n],\mathbf{V}^i[n] = \exp\left(-\frac{\Delta t}{\tau^i}\right) \mathbf{V}^i[n-1] + \left(1 - \exp\left(-\frac{\Delta t}{\tau^i}\right)\right)\Delta t V_{\mathrm{th}}^{i-1}\mathbf{W}^i \mathbf{s}^{i-1}[n] - V_{\mathrm{th}}^i \mathbf{s}^i[n],

    where i∈{1,…,L}i \in \{1, \dots, L\} indexes the layer, Vi[n]\mathbf{V}^i[n] is the membrane potential vector at time step n∈{1,…,N}n \in \{1, \dots, N\}, τi>Δt>0\tau^i > \Delta t > 0 is the membrane time constant, Vthi>0V_{\mathrm{th}}^i > 0 is the spike threshold, Wi\mathbf{W}^i is the synaptic weight matrix, si[n]∈{0,1}dim⁡\mathbf{s}^i[n] \in \{0, 1\}^{\dim} is the output spike vector, and Vi[0]=0\mathbf{V}^i[0] = \mathbf{0}.

    Let λi=exp⁡(−Δt/τi)\lambda_i = \exp(-\Delta t / \tau^i). Define the weighted average spike representation as:

    a^0[N]=∑n=1Nλ1N−ns0[n]∑n=1Nλ1N−nΔt,a^i[N]=Vthi∑n=1NλiN−nsi[n]∑n=1NλiN−nΔt,∀i=1,…,L.\hat{\mathbf{a}}^0[N] = \frac{\sum_{n=1}^N \lambda_1^{N-n}\mathbf{s}^0[n]}{\sum_{n=1}^N \lambda_1^{N-n}\Delta t}, \quad \hat{\mathbf{a}}^i[N] = \frac{V_{\mathrm{th}}^i \sum_{n=1}^N \lambda_i^{N-n}\mathbf{s}^i[n]}{\sum_{n=1}^N \lambda_i^{N-n}\Delta t}, \quad \forall i=1, \dots, L.

    Define the recursive sub-differentiable mappings:

    zi=clamp⁡(1τiWizi−1,0,VthiΔt),i=1,…,L,\mathbf{z}^i = \operatorname{clamp}\left(\frac{1}{\tau^i}\mathbf{W}^i \mathbf{z}^{i-1}, 0, \frac{V_{\mathrm{th}}^i}{\Delta t}\right), \quad i=1, \dots, L,

    where clamp⁡(x,a,b)≜max⁡(a,min⁡(x,b))\operatorname{clamp}(x, a, b) \triangleq \max(a, \min(x, b)).

    If lim⁡N→∞a^i[N]=zi\lim_{N \to \infty} \hat{\mathbf{a}}^i[N] = \mathbf{z}^i holds for all i=0,1,…,L−1i = 0, 1, \dots, L-1, then as simulation latency N→∞N \to \infty, the weighted firing rate representation a^i+1[N]\hat{\mathbf{a}}^{i+1}[N] asymptotically converges to the clamped mapping zi+1\mathbf{z}^{i+1}.

  3. Knowl 3 — Sub-Differentiable Mapping Convergence for IF Neurons

    theoretical result

    Consider an LL-layer feedforward SNN composed of Integrate-and-Fire (IF) neurons governed by the discrete dynamics:

    Vi[n]=Vi[n−1]+Vthi−1Wisi−1[n]−Vthisi[n],\mathbf{V}^i[n] = \mathbf{V}^i[n-1] + V_{\mathrm{th}}^{i-1}\mathbf{W}^i \mathbf{s}^{i-1}[n] - V_{\mathrm{th}}^i \mathbf{s}^i[n],

    where i∈{1,…,L}i \in \{1, \dots, L\} is the layer index, Vi[n]\mathbf{V}^i[n] is the membrane potential at time step n∈{1,…,N}n \in \{1, \dots, N\}, Wi\mathbf{W}^i is the synaptic weight matrix, Vthi>0V_{\mathrm{th}}^i > 0 is the layer threshold, si[n]∈{0,1}dim⁡\mathbf{s}^i[n] \in \{0, 1\}^{\dim} are the generated binary spikes, and Vi[0]=0\mathbf{V}^i[0] = \mathbf{0}.

    Define the scaled firing rate representations as:

    a0[N]=1N∑n=1Ns0[n],ai[N]=1N∑n=1NVthisi[n],∀i=1,…,L.\mathbf{a}^0[N] = \frac{1}{N} \sum_{n=1}^N \mathbf{s}^0[n], \quad \mathbf{a}^i[N] = \frac{1}{N}\sum_{n=1}^N V_{\mathrm{th}}^i \mathbf{s}^i[n], \quad \forall i=1, \dots, L.

    Define the corresponding sub-differentiable mappings:

    zi=clamp⁡(Wizi−1,0,Vthi),i=1,…,L,\mathbf{z}^i = \operatorname{clamp}\left(\mathbf{W}^i \mathbf{z}^{i-1}, 0, V_{\mathrm{th}}^i\right), \quad i=1, \dots, L,

    where clamp⁡(x,a,b)≜max⁡(a,min⁡(x,b))\operatorname{clamp}(x, a, b) \triangleq \max(a, \min(x, b)).

    If the input representation satisfies lim⁡N→∞a0[N]=z0\lim_{N \to \infty} \mathbf{a}^0[N] = \mathbf{z}^0, then for all layers i=1,…,Li = 1, \dots, L:

    lim⁡N→∞ai[N]=zi.\lim_{N \to \infty} \mathbf{a}^i[N] = \mathbf{z}^i.
  4. Knowl 4 — Modified Firing Mechanism to Reduce Quantization Error

    model/method

    In low-latency SNNs with a small time step horizon NN, the finite number of discrete spike levels introduces a quantization error eqe_q. For standard IF neurons receiving a constant input current I∗I^*, the scaled firing rate is quantized as:

    a[N]=VthN⋅clamp⁡(⌊NI∗Vth⌋,0,N),a[N] = \frac{V_{\mathrm{th}}}{N} \cdot \operatorname{clamp}\left(\left\lfloor \frac{N I^*}{V_{\mathrm{th}}} \right\rfloor, 0, N\right),

    which implements a floor rounding operation ⌊⋅⌋\lfloor \cdot \rfloor.

    To reduce this quantization error, a firing hyperparameter α∈[0,1]\alpha \in [0, 1] is introduced into the spike generation mechanism:

    s[n]=H(U[n]−αVth),s[n] = H(U[n] - \alpha V_{\mathrm{th}}),

    where HH is the Heaviside step function and U[n]U[n] is the pre-reset membrane potential.

    Setting α=0.5\alpha = 0.5 for IF neurons converts the scaled firing rate into nearest-integer rounding:

    a[N]=VthN⋅clamp⁡([NI∗Vth],0,N),a[N] = \frac{V_{\mathrm{th}}}{N} \cdot \operatorname{clamp}\left(\left[ \frac{N I^*}{V_{\mathrm{th}}} \right], 0, N\right),

    where [⋅][\cdot] is the rounding operator. This halves the maximum absolute quantization error and minimizes the average absolute quantization error. For LIF neurons, the parameter α∈[0,1]\alpha \in [0, 1] is selected according to latency NN to minimize the average absolute quantization error.

  5. Knowl 5 — Trainable Spike Thresholds with L2 Regularization

    model/method

    To balance the trade-off between neuron approximation capacity (promoted by higher thresholds VthV_{\mathrm{th}}) and quantization error minimization (promoted by lower thresholds VthV_{\mathrm{th}}), each layer's spike threshold VthiV_{\mathrm{th}}^i is treated as a learnable parameter optimized by backpropagation alongside synaptic weights.

    An L2 regularization penalty on the thresholds is added to the training loss function. For an IF neuron with steady scaled firing rate a∗a^* under average input current I∗I^*, the derivative of the steady firing rate with respect to threshold VthV_{\mathrm{th}} is:

    ∂a∗∂Vth={1,if I∗>Vth,0,otherwise.\frac{\partial a^*}{\partial V_{\mathrm{th}}} = \begin{cases} 1, & \text{if } I^* > V_{\mathrm{th}}, \\ 0, & \text{otherwise}. \end{cases}

    The gradient of the training loss with respect to VthV_{\mathrm{th}} is computed via the chain rule. During mini-batch stochastic gradient descent, the gradient for each threshold is scaled based on the mini-batch size and the neural model (IF or LIF).

  6. Knowl 6 — Spike Representation Definitions for IF and LIF Models

    definition

    The spike representation operator r(⋅)r(\cdot) transforms an individual neuron's binary spike train s=(s[1],s[2],…,s[N])∈{0,1}N\mathbf{s} = (s[1], s[2], \dots, s[N]) \in \{0, 1\}^N over latency NN into a continuous scalar representation:

    1. Integrate-and-Fire (IF) Model: The spike representation is defined as the scaled firing rate:
    r(s)=1N∑n=1NVths[n],r(\mathbf{s}) = \frac{1}{N} \sum_{n=1}^N V_{\mathrm{th}} s[n],

    where VthV_{\mathrm{th}} is the spike threshold.

    1. Leaky Integrate-and-Fire (LIF) Model: The spike representation is defined as the scaled weighted firing rate:
    r(s)=Vth∑n=1NλN−ns[n]∑n=1NλN−nΔt,r(\mathbf{s}) = \frac{V_{\mathrm{th}} \sum_{n=1}^N \lambda^{N-n} s[n]}{\sum_{n=1}^N \lambda^{N-n} \Delta t},

    where λ=exp⁡(−Δt/τ)\lambda = \exp(-\Delta t / \tau), τ\tau is the membrane time constant, Δt\Delta t is the discrete time step size, and VthV_{\mathrm{th}} is the spike threshold.

  7. Knowl 7 — Image Classification Performance on Benchmark Datasets

    data/table

    The table below compares the Differentiation on Spike Representation (DSR) method with artificial neural networks (ANN), ANN-to-SNN conversion methods, and direct training methods across CIFAR-10, CIFAR-100, ImageNet, and neuromorphic DVS-CIFAR10. Results on CIFAR and DVS-CIFAR10 report the mean and standard deviation over 3 runs.

    Method Network Neural Model Time Steps Accuracy
    CIFAR-10
    ANN PreAct-ResNet-18 / / 95.41%
    ANN-to-SNN ResNet-20 IF 128 93.56%
    ANN-to-SNN VGG-16 IF 2048 93.63%
    ANN-to-SNN VGG-like IF 600 94.20%
    Tandem Learning CIFARNet IF 8 90.98%
    ASF-BP VGG-7 IF 400 91.35%
    STBP CIFARNet LIF 12 90.53%
    IDE CIFARNet-F LIF 100 92.52% ±\pm 0.17%
    STBP-tdBN ResNet-19 LIF 6 93.16%
    TSSL-BP CIFARNet LIF w/ syn. 5 91.41%
    DSR (ours) PreAct-ResNet-18 IF 20 95.24% ±\pm 0.17%
    DSR (ours) PreAct-ResNet-18 LIF 20 95.40% ±\pm 0.15%
    CIFAR-100
    ANN PreAct-ResNet-18 / / 78.12%
    ANN-to-SNN ResNet-20 IF 400–600 69.82%
    ANN-to-SNN VGG-16 IF 768 70.09%
    ANN-to-SNN VGG-like IF 300 71.84%
    Hybrid Training VGG-11 LIF 125 67.84%
    DIET-SNN VGG-16 LIF 5 69.67%
    IDE CIFARNet-F LIF 100 73.07% ±\pm 0.21%
    DSR (ours) PreAct-ResNet-18 IF 20 78.20% ±\pm 0.13%
    DSR (ours) PreAct-ResNet-18 LIF 20 78.50% ±\pm 0.12%
    ImageNet
    ANN PreAct-ResNet-18 / / 70.79%
    ANN-to-SNN ResNet-34 IF 2000 65.47%
    ANN-to-SNN ResNet-34 IF 4096 69.89%
    Hybrid training ResNet-34 LIF 250 61.48%
    STBP-tdBN ResNet-34 LIF 6 63.72%
    SEW ResNet SEW ResNet-34 LIF 4 67.04%
    SEW ResNet SEW ResNet-18 LIF 4 63.18%
    DSR (ours) PreAct-ResNet-18 IF 50 67.74%
    DVS-CIFAR10
    ASF-BP VGG-7 IF 50 62.50%
    Tandem Learning 7-layer CNN IF 20 65.59%
    STBP 7-layer CNN LIF 40 60.50%
    STBP-tdBN ResNet-19 LIF 10 67.80%
    Fang et al. 7-layer CNN LIF 20 74.80%
    DSR (ours) VGG-11 IF 20 75.03% ±\pm 0.39%
    DSR (ours) VGG-11 LIF 20 77.27% ±\pm 0.24%

    The DSR method matches or outperforms corresponding ANN models on CIFAR-10 (95.40% vs. 95.41%) and CIFAR-100 (78.50% vs. 78.12%) with 20 simulation time steps, outperforming prior SNN direct training and conversion methods. On ImageNet with PreAct-ResNet-18 at 50 time steps, DSR achieves 67.74%, exceeding directly trained SEW ResNet-18 (63.18% at 4 steps) and SEW ResNet-34 (67.04% at 4 steps). On neuromorphic DVS-CIFAR10 with VGG-11 at 20 steps, DSR achieves 77.27% (LIF).

  8. Knowl 8 — Low-Latency Robustness and Deep Architecture Scaling

    empirical result

    The DSR method shows consistent performance when trained across small latencies and deep architectures on CIFAR-10:

    1. Ultra-Low Latency: When trained from scratch using PreAct-ResNet-18 across time step settings N∈{20,15,10,5}N \in \{20, 15, 10, 5\} (averaged over 3 runs):

      • IF model: 95.24%95.24\% at N=20N=20, 94.85%94.85\% at N=15N=15, 94.69%94.69\% at N=10N=10, and 94.48%94.48\% at N=5N=5.
      • LIF model: 95.40%95.40\% at N=20N=20, 95.39%95.39\% at N=15N=15, 94.47%94.47\% at N=10N=10, and 94.42%94.42\% at N=5N=5. Reducing simulation latency from N=20N=20 down to N=5N=5 results in less than a 1%1\% accuracy decrease.
    2. Deep Architecture Scaling: When evaluated using narrow PreAct-ResNet architectures of depth 20, 32, 44, 56, and 110 at latency N=20N=20:

      • IF model: depth 20 yields 92.67%92.67\%, depth 32 yields 93.73%93.73\%, depth 44 yields 93.74%93.74\%, depth 56 yields 94.15%94.15\%, and depth 110 yields 94.60%94.60\%.
      • LIF model: depth 20 yields 92.82%92.82\%, depth 32 yields 93.74%93.74\%, depth 44 yields 93.99%93.99\%, depth 56 yields 94.03%94.03\%, and depth 110 yields 94.61%94.61\%. Classification accuracy monotonically improves with network depth up to 110 layers, indicating that error accumulation across layers remains controlled.
  9. Knowl 9 — Ablation Study of Threshold Training and Firing Mechanism Modification

    data/table

    An ablation study on CIFAR-10 with PreAct-ResNet-18 using IF neurons at N=20N=20 time steps assesses the individual contributions of the modified firing mechanism (hyperparameter α=0.5\alpha = 0.5, denoted 'F') and trainable spike thresholds (denoted 'T') over 3 experimental runs.

    Setting Accuracy
    DSR (with F T), initial Vth=6V_{\mathrm{th}} = 6 95.24% ±\pm 0.17%
    DSR without F (only T), initial Vth=6V_{\mathrm{th}} = 6 92.88% ±\pm 0.25%
    DSR without T (only F), fixed Vth=6V_{\mathrm{th}} = 6 90.45% ±\pm 1.84%
    DSR without T (only F), fixed Vth=2V_{\mathrm{th}} = 2 90.47% ±\pm 0.12%
    DSR without F and T (baseline), fixed Vth=6V_{\mathrm{th}} = 6 92.59% ±\pm 0.81%

    The results show that combining threshold training and firing mechanism modification yields the highest performance (95.24%), outperforming individual and unassisted variants. Threshold training also provides optimization stability, reducing the standard deviation from ±1.84%\pm 1.84\% to ±0.17%\pm 0.17\% when Vth=6V_{\mathrm{th}} = 6.

  10. Knowl 10 — Performance Limitation Under Extreme Low Latency

    limitation

    The Differentiation on Spike Representation (DSR) method relies on the accuracy of the spike representation operator r(s)r(\mathbf{s}) to estimate the sub-differentiable mapping gWg_{\mathbf{W}}. When the simulation latency is extremely short (e.g., N=2N = 2 or N=3N = 3 time steps), the quantization error and the deviation error between the actual firing dynamics and the continuous sub-differentiable activation mapping increase significantly. This estimation discrepancy causes performance degradation under extreme low-latency regimes.

Coverage note — No substantial contributed material was omitted. The mathematical proofs and derivations of the asymptotic convergence for Propositions 1 and 2 were omitted in accordance with the rule excluding intermediate proofs.

References

  1. 1.Guillaume Bellec, Darjan Salaj, Anand Subramoney, Robert Legenstein, and Wolfgang Maass. Long short-term memory and learning-to-learn in networks of spiking neurons. In NeurIPS, 2018.
  2. 2.Sander M Bohte, Joost N Kok, and Johannes A La Poutré. Spikeprop: backpropagation for networks of spiking neurons. In ESANN, 2000.
  3. 3.Anthony N Burkitt. A review of the integrate-and-fire neuron model: I. homogeneous synaptic input. Biological cybernetics, 95(1):1–19, 2006.
  4. 4.Yongqiang Cao, Yang Chen, and Deepak Khosla. Spiking deep convolutional neural networks for energy-efficient object recognition. IJCV, 113(1):54–66, 2015.
  5. 5.Natalia Caporale and Yang Dan. Spike timing–dependent plasticity: a hebbian learning rule. Annu. Rev. Neurosci., 31:25–46, 2008.
  6. 6.Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. Pact: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085, 2018.
  7. 7.Mike Davies, Narayan Srinivasa, Tsung-Han Lin, Gautham Chinya, Yongqiang Cao, Sri Harsha Choday, Georgios Dimou, Prasad Joshi, Nabil Imam, Shweta Jain, et al. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro, 38(1):82–99, 2018.
  8. 8.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. IEEE, 2009.
  9. 9.Lei Deng, Yujie Wu, Xing Hu, Ling Liang, Yufei Ding, Guoqi Li, Guangshe Zhao, Peng Li, and Yuan Xie. Rethinking the performance comparison between snns and anns. Neural Networks, 121:294–307, 2020.
  10. 10.Shikuang Deng and Shi Gu. Optimal conversion of conventional artificial neural networks to spiking neural networks. In ICLR, 2021.
  11. 11.Peter U Diehl, Daniel Neil, Jonathan Binas, Matthew Cook, Shih-Chii Liu, and Michael Pfeiffer. Fast-classifying, highaccuracy spiking deep networks through weight and threshold balancing. In IJCNN, 2015.
  12. 12.Jianhao Ding, Zhaofei Yu, Yonghong Tian, and Tiejun Huang. Optimal ann-snn conversion for fast and accurate inference in deep spiking neural networks. In IJCAI, 2021.
  13. 13.Steve K Esser, Rathinakumar Appuswamy, Paul Merolla, John V Arthur, and Dharmendra S Modha. Backpropagation for energy-efficient neuromorphic computing. In NeurIPS, 2015.
  14. 14.Steven K. Esser, Paul A. Merolla, John V. Arthur, Andrew S. Cassidy, Rathinakumar Appuswamy, Alexander Andreopoulos, David J. Berg, Jeffrey L. McKinstry, Timothy Melano, Davis R. Barch, Carmelo di Nolfo, Pallab Datta, Arnon Amir, Brian Taba, Myron D. Flickner, and Dharmendra S. Modha. Convolutional networks for fast, energy-efficient neuromorphic computing. PNAS, 113(41):11441–11446, 2016.
  15. 15.Wei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang, Timothée Masquelier, and Yonghong Tian. Deep residual learning in spiking neural networks. In NeurIPS, 2021.
  16. 16.Wei Fang, Zhaofei Yu, Yanqi Chen, Timothee Masquelier, Tiejun Huang, and Yonghong Tian. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In ICCV, 2021.
  17. 17.Bing Han and Kaushik Roy. Deep spiking neural network: Energy efficiency through time based coding. In ECCV, 2020.
  18. 18.Bing Han, Gopalakrishnan Srinivasan, and Kaushik Roy. RMP-SNN: residual membrane potential neuron for enabling deeper high-accuracy and low-latency spiking neural network. In CVPR, 2020.
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In ECCV, 2016.
  20. 20.Donald Olding Hebb. The organisation of behaviour: a neuropsychological theory. Science Editions New York, 1949.
  21. 21.Yangfan Hu, Huajin Tang, Yueming Wang, and Gang Pan. Spiking deep residual network. arXiv preprint arXiv:1805.01352, 2018.
  22. 22.Dongsung Huh and Terrence J. Sejnowski. Gradient descent for spiking neural networks. In NeurIPS, 2018.
  23. 23.Eric Hunsberger and Chris Eliasmith. Spiking deep networks with lif neurons. arXiv preprint arXiv:1510.08829, 2015.
  24. 24.Saeed Reza Kheradpisheh, Mohammad Ganjtabesh, Simon J Thorpe, and Timothee Masquelier. Stdp-based spiking deep convolutional neural networks for object recognition. Neural Networks, 99:56–67, 2018.
  25. 25.Sei Joon Kim, Seongsik Park, Byunggook Na, and Sungroh Yoon. Spiking-yolo: Spiking neural network for energyefficient object detection. In AAAI, 2020.
  26. 26.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  27. 27.Chankyu Lee, Syed Shakib Sarwar, Priyadarshini Panda, Gopalakrishnan Srinivasan, and Kaushik Roy. Enabling spike-based backpropagation for training deep neural network architectures. Frontiers in Neuroscience, 14:119, 2020.
  28. 28.Robert Legenstein, Dejan Pecevski, and Wolfgang Maass. A learning theory for reward-modulated spike-timingdependent plasticity with application to biofeedback. PLoS computational biology, 4(10):e1000180, 2008.
  29. 29.Hongmin Li, Hanchao Liu, Xiangyang Ji, Guoqi Li, and Luping Shi. Cifar10-dvs: an event-stream dataset for object classification. Frontiers in neuroscience, 11:309, 2017.
  30. 30.Yuhang Li, Shikuang Deng, Xin Dong, Ruihao Gong, and Shi Gu. A free lunch from ann: Towards efficient, accurate spiking neural networks calibration. In ICML, 2021.
  31. 31.Paul A Merolla, John V Arthur, Rodrigo Alvarez-Icaza, Andrew S Cassidy, Jun Sawada, Filipp Akopyan, Bryan L Jackson, Nabil Imam, Chen Guo, Yutaka Nakamura, et al. A million spiking-neuron integrated circuit with a scalable communication network and interface. Science, 345(6197):668–673, 2014.
  32. 32.Hesham Mostafa. Supervised learning based on temporal coding in spiking neural networks. IEEE transactions on neural networks and learning systems, 29(7):3227–3235, 2017.
  33. 33.Emre O Neftci, Hesham Mostafa, and Friedemann Zenke. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6):51–63, 2019.
  34. 34.Stefano Panzeri and Simon R Schultz. A unified approach to the study of temporal, correlational, and rate coding. Neural Computation, 13(6):1311–1349, 2001.
  35. 35.Jing Pei, Lei Deng, Sen Song, Mingguo Zhao, Youhui Zhang, Shuang Wu, Guanrui Wang, Zhe Zou, Zhenzhi Wu, Wei He, et al. Towards artificial general intelligence with hybrid tianjic chip architecture. Nature, 572(7767):106–111, 2019.
  36. 36.Nitin Rathi and Kaushik Roy. Diet-snn: Direct input encoding with leakage and threshold optimization in deep spiking neural networks. arXiv preprint arXiv:2008.03658, 2020.
  37. 37.Nitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, and Kaushik Roy. Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation. In ICLR, 2020.
  38. 38.Bodo Rueckauer, Iulia-Alexandra Lungu, Yuhuang Hu, Michael Pfeiffer, and Shih-Chii Liu. Conversion of continuous-valued deep networks to efficient event-driven networks for image classification. Frontiers in neuroscience, 11:682, 2017.
  39. 39.Abhronil Sengupta, Yuting Ye, Robert Wang, Chiao Liu, and Kaushik Roy. Going deeper in spiking neural networks: Vgg and residual architectures. Frontiers in neuroscience, 13:95, 2019.
  40. 40.Sumit Bam Shrestha and Garrick Orchard. SLAYER: spike layer error reassignment in time. In NeurIPS, 2018.
  41. 41.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2014.
  42. 42.Christoph Stöckl and Wolfgang Maass. Optimized spiking neurons can classify images with high accuracy through temporal coding with two spikes. Nature Machine Intelligence, pages 1–9, 2021.
  43. 43.Amirhossein Tavanaei, Masoud Ghodrati, Saeed Reza Kheradpisheh, Timothee Masquelier, and Anthony Maida. Deep learning in spiking neural networks. Neural Networks, 111:47–63, 2019.
  44. 44.Hao Wu, Yueyi Zhang, Wenming Weng, Yongting Zhang, Zhiwei Xiong, Zheng-Jun Zha, Xiaoyan Sun, and Feng Wu. Training spiking neural networks with accumulated spiking flow. In AAAI, 2021.
  45. 45.Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li, Haizhou Li, and Kay Chen Tan. A tandem learning rule for effective training and rapid inference of deep spiking neural networks. TNNLS, 2021.
  46. 46.Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, and Luping Shi. Spatio-temporal backpropagation for training highperformance spiking neural networks. Frontiers in neuroscience, 12:331, 2018.
  47. 47.Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, Yuan Xie, and Luping Shi. Direct training for spiking neural networks: Faster, larger, better. In AAAI, 2019.
  48. 48.Timo C Wunderlich and Christian Pehle. Event-based backpropagation can compute exact gradients for spiking neural networks. Scientific Reports, 11(1):1–17, 2021.
  49. 49.Mingqing Xiao, Qingyan Meng, Zongpeng Zhang, Yisen Wang, and Zhouchen Lin. Training feedback spiking neural networks by implicit differentiation on the equilibrium state. In NeurIPS, 2021.
  50. 50.Zhanglu Yan, Jun Zhou, and Weng-Fai Wong. Near lossless transfer learning for spiking neural networks. In AAAI, 2021.
  51. 51.Yukun Yang, Wenrui Zhang, and Peng Li. Backpropagated neighborhood aggregation for accurate training of spiking neural networks. In ICML, 2021.
  52. 52.Wenrui Zhang and Peng Li. Temporal spike sequence learning via backpropagation for deep spiking neural networks. In NeurIPS, 2020.
  53. 53.Hanle Zheng, Yujie Wu, Lei Deng, Yifan Hu, and Guoqi Li. Going deeper with directly-trained larger spiking neural networks. In AAAI, 2021.
  54. 54.Shibo Zhou, Xiaohua Li, Ying Chen, Sanjeev T Chandrasekaran, and Arindam Sanyal. Temporal-coded deep spiking neural network with easy training and robust performance. In AAAI, 2021.

Citation

MLA
Meng, Q., et al. “Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation”. arXiv, 2022, http://arxiv.org/abs/2205.00459v2.
APA
Meng, Q., Xiao, M., Yan, S., Wang, Y., Lin, Z., & Luo, Z.-Q. (2022). Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation. arXiv. http://arxiv.org/abs/2205.00459v2
Chicago
Meng, Q., M. Xiao, S. Yan, Y. Wang, Z. Lin, and Z.-Q. Luo. 2022. “Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation”. arXiv. http://arxiv.org/abs/2205.00459v2.
Harvard
Meng, Q. et al. (2022) “Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.00459v2.
Vancouver
1. Meng Q, Xiao M, Yan S, Wang Y, Lin Z, Luo Z-Q (2022) Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation. arXiv

BibTeX

@article{meng2022training,
  title = {Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation},
  author = {Meng, Qingyan and Xiao, Mingqing and Yan, Shen and Wang, Yisen and Lin, Zhouchen and Luo, Zhi-Quan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.00459v2},
  eprint = {2205.00459}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE