Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks

Yujie WuLei DengGuoqi LiJun ZhuLuping Shi

article2017Frontiers in Neuroscience1,403 citations

Proposes a spatio-temporal backpropagation framework with surrogate gradient approximation for iterative leaky integrate-and-fire models, overcoming the non-differentiability of spikes to achieve high-performance direct supervised training of spiking neural networks on both static and neuromorphic benchmarks.

Listen

Spiking neural networks are energy-efficient, brain-inspired models well-suited for processing event-driven, time-dependent data on specialized hardware. However, training them directly using gradient-based learning has historically been difficult due to the discontinuous, non-differentiable nature of spiking events and the common neglect of timing dynamics during optimization. Existing approaches typically convert pre-trained conventional models or rely on complex mathematical workarounds and regularizations that limit overall performance.

The article develops and evaluates a spatio-temporal backpropagation training algorithm that combines spatial layer-to-layer error propagation with temporal timing dynamics. To enable gradient-based optimization, the authors establish an iterative leaky integrate-and-fire neuronal model and introduce continuous approximation curves to handle the non-differentiable spiking threshold.

The evaluation tests both fully connected and convolutional spiking architectures across static visual datasets (MNIST handwritten digits and a custom pedestrian detection dataset) and dynamic event-based sensor data (N-MNIST). The authors also assess model sensitivity to various derivative approximation curves, curve widths, and the removal of the temporal domain component.

The findings show that the proposed method establishes state-of-the-art accuracy across all tested benchmarks for spiking architectures. The algorithm achieved 98.89% accuracy on static MNIST with a fully connected network and 99.42% with a convolutional architecture. On dynamic N-MNIST data, the model reached 98.78% accuracy, outperforming existing spiking networks as well as standard non-spiking deep neural networks. Ablation experiments demonstrated that including temporal dynamics prevents performance degradation and stabilizes training, while the specific shape of the derivative approximation curve matters far less than choosing an appropriate curve steepness.

These results indicate that directly training spiking networks across both spatial and temporal dimensions yields high accuracy without requiring specialized tricks such as weight normalization or complex reset rules. This significantly lowers algorithmic complexity, making the approach appealing for practical applications on energy-efficient neuromorphic hardware and edge devices.

Organizations developing brain-inspired AI and low-power sensory hardware should adopt spatio-temporal gradient frameworks for direct model training rather than relying on cumbersome translation pipelines. Future efforts should prioritize accelerating software simulation routines—which remain considerably slower than conventional network training—and testing the algorithm on richer temporal benchmarks such as speech datasets and larger-scale computer vision tasks.

Confidence in these findings is high for standard image classification and dynamic vision benchmarks, but caution is warranted when scaling to larger, more complex datasets like CIFAR-10, where high software simulation runtimes and lack of advanced optimization techniques currently limit absolute accuracy.

  • Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). Introduces straight-through and continuous gradient estimators for backpropagating through non-differentiable hard thresholds, establishing the mathematical prerequisite for surrogate gradient approximations in spiking neurons.
  • Paper: Learning representations by back-propagating errors, David E. Rumelhart et al. (1986). Presents the foundational backpropagation algorithm for computing parameter gradients across network layers, which the source extends to the temporal domain and non-smooth dynamics of spiking networks.
  • Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Establishes the backpropagation-through-time framework for recurrent architectures and discrete time-unrolled computational graphs, directly informing the iterative state modeling in spatio-temporal backpropagation.
  • Paper: On the difficulty of training recurrent neural networks, Razvan Pascanu et al. (2012). Analyzes the mechanics and mathematical conditions of vanishing and exploding gradients in time-unrolled networks, providing essential context for the stability challenges of temporal credit assignment.
Cover for Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks

Abstract

Compared with artificial neural networks (ANNs), spiking neural networks (SNNs) are promising to explore the brain-like behaviors since the spikes could encode more spatio-temporal information. Although pre-training from ANN or direct training based on backpropagation (BP) makes the supervised training of SNNs possible, these methods only exploit the networks' spatial domain information which leads to the performance bottleneck and requires many complicated training skills. Another fundamental issue is that the spike activity is naturally non-differentiable which causes great difficulties in training SNNs. To this end, we build an iterative LIF model that is more friendly for gradient descent training. By simultaneously considering the layer-by-layer spatial domain (SD) and the timing-dependent temporal domain (TD) in the training phase, as well as an approximated derivative for the spike activity, we propose a spatio-temporal backpropagation (STBP) training framework without using any complicated technology. We achieve the best performance of multi-layered perceptron (MLP) compared with existing state-of-the-art algorithms over the static MNIST and the dynamic N-MNIST dataset as well as a custom object detection dataset. This work provides a new perspective to explore the high-performance SNNs for future brain-like computing paradigm with rich spatio-temporal dynamics.

Table of Contents

  • I Introduction
  • II Method and Material
  • II-A Iterative Leaky Integrate-and-Fire Model in Spiking Neural Networks
  • II-B Spatio-Temporal Backpropagation Training
  • II-C Derivative Approximation of the Non-differentiable Spike Activity
  • III Results
  • III-A Parameter Initialization
  • III-B Dataset Experiments
  • III-B1 Spatio-temporal fully connected neural network
  • III-B2 Spatio-temporal convolution neural network
  • III-C Performance Analysis
  • III-C1 The Impact of Derivative Approximation Curves
  • III-C2 The Impact of Temporal Domain
  • IV Conclusion
  • References

Knowls

  1. Knowl 1 — Iterative Leaky Integrate-and-Fire Neuronal Model for Spatio-Temporal Backpropagation

    model/method

    The continuous-time Leaky Integrate-and-Fire (LIF) neuron model, governed by τdu(t)dt=−u(t)+I(t)\tau \frac{du(t)}{dt} = -u(t) + I(t) (where u(t)u(t) is the membrane potential, τ\tau is the decay time constant, and I(t)I(t) is the pre-synaptic input), is discretized into an iterative update rule suitable for backpropagation across discrete time steps t∈{1,…,T}t \in \{1, \dots, T\} and network layers n∈{1,…,N}n \in \{1, \dots, N\}:

    xit+1,n=∑j=1l(n−1)wijnojt+1,n−1x_i^{t+1, n} = \sum_{j=1}^{l(n-1)} w_{ij}^n o_j^{t+1, n-1}

    uit+1,n=uit,nf(oit,n)+xit+1,n+binu_i^{t+1, n} = u_i^{t, n} f(o_i^{t, n}) + x_i^{t+1, n} + b_i^n

    oit+1,n=g(uit+1,n)o_i^{t+1, n} = g(u_i^{t+1, n})

    where:

    • l(n−1)l(n-1) is the number of neurons in layer n−1n-1.
    • wijnw_{ij}^n is the synaptic weight from neuron jj in layer n−1n-1 to neuron ii in layer nn.
    • ojt+1,n−1∈{0,1}o_j^{t+1, n-1} \in \{0, 1\} is the binary spike output of neuron jj in layer n−1n-1 at time step t+1t+1.
    • xit+1,nx_i^{t+1, n} denotes the aggregated pre-synaptic spatial input to neuron ii in layer nn.
    • uit+1,nu_i^{t+1, n} is the membrane potential of neuron ii in layer nn at time step t+1t+1.
    • binb_i^n is a learnable bias parameter mimicking a threshold adjustment.
    • f(x)=τe−x/τf(x) = \tau e^{-x/\tau} is a temporal decay gate. For small positive τ\tau, it is approximated as: f(oit,n)≈{τ,oit,n=00,oit,n=1f(o_i^{t, n}) \approx \begin{cases} \tau, & o_i^{t, n} = 0 \\ 0, & o_i^{t, n} = 1 \end{cases} causing the membrane potential to reset toward 0 upon firing (oit,n=1o_i^{t, n} = 1) and to decay by factor τ\tau when no spike occurs (oit,n=0o_i^{t, n} = 0).
    • g(u)g(u) is the spike generation step function: g(u)={1,u≥Vth0,u<Vthg(u) = \begin{cases} 1, & u \ge V_{\text{th}} \\ 0, & u < V_{\text{th}} \end{cases} where VthV_{\text{th}} is the firing threshold.
  2. Knowl 2 — Spatio-Temporal Backpropagation (STBP) Training Algorithm

    algorithm

    The STBP algorithm trains spiking neural networks by propagating error gradients through both the spatial domain (layer-by-layer connections) and the temporal domain (iterative intra-neuron dynamics) over a full simulation window TT. It minimizes the mean squared error loss over SS training samples:

    L=12S∑s=1S∥ys−1T∑t=1Tost,N∥22L = \frac{1}{2S} \sum_{s=1}^S \left\| y_s - \frac{1}{T} \sum_{t=1}^T o_s^{t, N} \right\|_2^2

    where ysy_s is the target label vector of sample ss, and ost,No_s^{t, N} is the output spike vector of the final layer NN at time step tt.

    Input: Training batch of samples and targets, network depth NN, total time steps TT, weights WnW^n, biases bnb^n, threshold VthV_{\text{th}}, decay factor τ\tau, surrogate derivative function h(u)h(u).
    Output: Parameter updates ΔWn\Delta W^n and Δbn\Delta b^n for all layers n=1,…,Nn = 1, \dots, N.
    for each training batch do
        // Forward Pass
        for t=1t = 1 to TT do
            for n=1n = 1 to NN do
                xit,n=∑j=1l(n−1)wijnojt,n−1x_i^{t, n} = \sum_{j=1}^{l(n-1)} w_{ij}^n o_j^{t, n-1}
                uit,n=uit−1,nf(oit−1,n)+xit,n+binu_i^{t, n} = u_i^{t-1, n} f(o_i^{t-1, n}) + x_i^{t, n} + b_i^n
                oit,n=g(uit,n)o_i^{t, n} = g(u_i^{t, n})
            end for
        end for
        // Backward Pass
        for t=Tt = T down to 1 do
            for n=Nn = N down to 1 do
                if t==Tt == T and n==Nn == N then
                    ∂L∂oiT,N=−1TS(yi−1T∑k=1Toik,N)\frac{\partial L}{\partial o_i^{T, N}} = -\frac{1}{TS} (y_i - \frac{1}{T} \sum_{k=1}^T o_i^{k, N})
                    ∂L∂uiT,N=∂L∂oiT,Nh(uiT,N)\frac{\partial L}{\partial u_i^{T, N}} = \frac{\partial L}{\partial o_i^{T, N}} h(u_i^{T, N})
                else if t==Tt == T and n<Nn < N then
                    ∂L∂oiT,n=∑j=1l(n+1)∂L∂ojT,n+1h(ujT,n+1)wjin+1\frac{\partial L}{\partial o_i^{T, n}} = \sum_{j=1}^{l(n+1)} \frac{\partial L}{\partial o_j^{T, n+1}} h(u_j^{T, n+1}) w_{ji}^{n+1}
                    ∂L∂uiT,n=∂L∂oiT,nh(uiT,n)\frac{\partial L}{\partial u_i^{T, n}} = \frac{\partial L}{\partial o_i^{T, n}} h(u_i^{T, n})
                else if t<Tt < T and n==Nn == N then
                    ∂L∂oit,N=∂L∂oit+1,Nh(uit+1,N)uit,N∂f∂oit,N−1TS(yi−1T∑k=1Toik,N)\frac{\partial L}{\partial o_i^{t, N}} = \frac{\partial L}{\partial o_i^{t+1, N}} h(u_i^{t+1, N}) u_i^{t, N} \frac{\partial f}{\partial o_i^{t, N}} - \frac{1}{TS} (y_i - \frac{1}{T} \sum_{k=1}^T o_i^{k, N})
                    ∂L∂uit,N=∂L∂oit+1,Nh(uit+1,N)f(oit,N)\frac{\partial L}{\partial u_i^{t, N}} = \frac{\partial L}{\partial o_i^{t+1, N}} h(u_i^{t+1, N}) f(o_i^{t, N})
                else if t<Tt < T and n<Nn < N then
                    ∂L∂oit,n=∑j=1l(n+1)∂L∂ojt,n+1h(ujt,n+1)wjin+1+∂L∂oit+1,nh(uit+1,n)uit,n∂f∂oit,n\frac{\partial L}{\partial o_i^{t, n}} = \sum_{j=1}^{l(n+1)} \frac{\partial L}{\partial o_j^{t, n+1}} h(u_j^{t, n+1}) w_{ji}^{n+1} + \frac{\partial L}{\partial o_i^{t+1, n}} h(u_i^{t+1, n}) u_i^{t, n} \frac{\partial f}{\partial o_i^{t, n}}
                    ∂L∂uit,n=∂L∂oit,nh(uit,n)+∂L∂oit+1,nh(uit+1,n)f(oit,n)\frac{\partial L}{\partial u_i^{t, n}} = \frac{\partial L}{\partial o_i^{t, n}} h(u_i^{t, n}) + \frac{\partial L}{\partial o_i^{t+1, n}} h(u_i^{t+1, n}) f(o_i^{t, n})
                end if
            end for
        end for
        // Gradient Accumulation and Update
        for n=1n = 1 to NN do
            ∂L∂bn=∑t=1T∂L∂ut,n\frac{\partial L}{\partial b^n} = \sum_{t=1}^T \frac{\partial L}{\partial u^{t, n}}
            ∂L∂Wn=∑t=1T∂L∂ut,n(ot,n−1)T\frac{\partial L}{\partial W^n} = \sum_{t=1}^T \frac{\partial L}{\partial u^{t, n}} (o^{t, n-1})^T
            Update WnW^n and bnb^n via Adam optimizer
        end for
    end for
  3. Knowl 3 — Surrogate Derivative Approximations for Non-Differentiable Spike Generation

    model/method

    Because the step activation function g(u)g(u) has a Dirac delta derivative δ(u−Vth)\delta(u - V_{\text{th}}) (zero everywhere except infinity at threshold), standard gradient descent fails due to vanishing or exploding gradients. To enable backpropagation, ∂g∂u\frac{\partial g}{\partial u} is approximated by a smooth function hi(u)h_i(u) parametrized by steepness/width ai>0a_i > 0 such that ∫−∞∞hi(u)du=1\int_{-\infty}^{\infty} h_i(u) du = 1 and lim⁡ai→0+hi(u)=dgdu\lim_{a_i \to 0^+} h_i(u) = \frac{dg}{du}:

    1. Rectangular function derivative: h1(u)=1a1sign(∣u−Vth∣<a12)h_1(u) = \frac{1}{a_1} \text{sign}\left(|u - V_{\text{th}}| < \frac{a_1}{2}\right)
    2. Polynomial function derivative: h2(u)=(a22−a24∣u−Vth∣)sign(2a2−∣u−Vth∣)h_2(u) = \left(\frac{\sqrt{a_2}}{2} - \frac{a_2}{4}|u - V_{\text{th}}|\right) \text{sign}\left(\frac{2}{\sqrt{a_2}} - |u - V_{\text{th}}|\right)
    3. Sigmoid derivative: h3(u)=1a3e(Vth−u)/a3(1+e(Vth−u)/a3)2h_3(u) = \frac{1}{a_3} \frac{e^{(V_{\text{th}}-u)/a_3}}{\left(1 + e^{(V_{\text{th}}-u)/a_3}\right)^2}
    4. Gaussian cumulative distribution function derivative: h4(u)=12πa4e−(u−Vth)22a4h_4(u) = \frac{1}{\sqrt{2\pi a_4}} e^{-\frac{(u - V_{\text{th}})^2}{2a_4}}
  4. Knowl 4 — Weight Initialization and Normalization Strategy for SNNs

    model/method

    To prevent premature saturation or silence across network layers without relying on complex runtime training heuristics (e.g., error normalization, weight/threshold balancing, or fixed-amount proportional resets), synaptic weights are initialized from a standard uniform distribution:

    W∼U[−1,1]W \sim U[-1, 1]

    and subsequently normalized per post-synaptic neuron ii in layer nn by the Euclidean norm of incoming weights:

    wijn=wijn∑j=1l(n−1)(wijn)2,i=1,…,l(n)w_{ij}^n = \frac{w_{ij}^n}{\sqrt{\sum_{j=1}^{l(n-1)} (w_{ij}^n)^2}}, \quad i = 1, \dots, l(n)

    where l(n−1)l(n-1) is the fan-in dimension from the previous layer. The firing threshold VthV_{\text{th}} is kept constant across all neurons in a layer, leaving weight magnitudes to govern firing balance.

  5. Knowl 5 — Spiking Convolutional Neural Network Formulation

    model/method

    The STBP framework extends to Spiking Convolutional Neural Networks (CNNs) through two structural adaptations:

    • In convolutional layers, the matrix-vector dot product ∑jwijnojt,n−1\sum_j w_{ij}^n o_j^{t, n-1} is replaced with a 2D spatial convolution between the learnable kernel tensor and the previous layer's binary spike feature maps, with the resulting pre-synaptic potential updating the iterative LIF neuron states.
    • In pooling layers, standard max pooling is replaced by average pooling, because binary spike activations ({0,1}\{0, 1\}) cannot preserve rate-graded temporal intensity under max operations, whereas average pooling maintains signal continuity across time steps.
  6. Knowl 6 — Benchmark Accuracy Comparison on Static MNIST

    data/table

    On the static MNIST benchmark (28x28 grayscale digits converted to spike trains over T=30 msT=30\text{ ms} with dt=1 msdt=1\text{ ms} via independent Bernoulli sampling), STBP achieves higher accuracy than competing spiking networks without using auxiliary regularization tricks.

    Model Network Structure Training Skills Accuracy (%)
    Spiking RBM (STDP) (Neftci et al., 2013) 784-500-40 None 93.16
    Spiking RBM (pre-training) (Peter et al., 2013) 784-500-500-10 None 97.48
    Spiking MLP (pre-training) (Diehl et al., 2015) 784-1200-1200-10 Weight normalization 98.64
    Spiking MLP (pre-training) (Hunsberger and Eliasmith, 2015) 784-500-200-10 None 98.37
    Spiking MLP (BP) (O'Connor and Welling, 2016) 784-200-200-10 None 97.66
    Spiking MLP (STDP) (Diehl and Cook, 2015) 784-6400 None 95.00
    Spiking MLP (BP) (Lee et al., 2016) 784-800-10 Error norm. / param. reg. 98.71
    Spiking MLP (STBP) 784-800-10 None 98.89
    Spiking CNN (pre-training) (Esser et al., 2016) 28×\times28×\times1-12C5-P2-64C5-P2-10 - 99.12
    Spiking CNN (BP) (Lee et al., 2016) 28×\times28×\times1-20C5-P2-50C5-P2-200-10 - 99.31
    Spiking CNN (STBP) 28×\times28×\times1-15C5-P2-40C5-P2-300-10 None 99.42

    The Spiking MLP trained via STBP achieves 98.89% test accuracy on MNIST, exceeding prior direct BP and converted models. The Spiking CNN with STBP achieves 99.42% accuracy with a lighter convolutional structure than competing direct-BP spiking CNN architectures.

  7. Knowl 7 — Benchmark Accuracy Comparison on Dynamic N-MNIST

    data/table

    The N-MNIST dataset converts static MNIST digits into dynamic event streams using a Dynamic Vision Sensor (DVS) undergoing 3 triangular saccades, yielding 34×3434 \times 34 spatial resolution with separate ON and OFF event polarity channels (34×34×234 \times 34 \times 2). STBP directly processes this event-stream data without frame-based accumulation preprocessing.

    Model Network Structure Training Skills Accuracy (%)
    Non-spiking CNN (BP) (Neil et al., 2016) - None 95.30
    Non-spiking CNN (BP) (Neil and Liu, 2016) - None 98.30
    Non-spiking MLP (BP) (Lee et al., 2016) 34×\times34×\times2-800-10 None 97.80
    LSTM (BPTT) (Neil et al., 2016) - Batch normalization 97.05
    Phased-LSTM (BPTT) (Neil et al., 2016) - None 97.38
    Spiking CNN (pre-training) (Neil and Liu, 2016) - None 95.72
    Spiking MLP (BP) (Lee et al., 2016) 34×\times34×\times2-800-10 Error norm. / param. reg. 98.74
    Spiking MLP (BP) (Cohen et al., 2016) 34×\times34×\times2-10000-10 None 92.87
    Spiking MLP (STBP) 34×\times34×\times2-800-10 None 98.78

    The Spiking MLP trained with STBP achieves 98.78% accuracy on N-MNIST, outperforming all non-spiking RNN/CNN architectures (e.g., Phased-LSTM at 97.38%) and prior direct-BP spiking networks (98.74%), without requiring error normalization or parameter regularization.

  8. Knowl 8 — Classification Performance on Pedestrian Object Detection Dataset

    data/table

    On a binary pedestrian object detection dataset composed of 1,509 training and 631 testing 28×2828 \times 28 grayscale image patches, spiking networks trained with STBP achieve classification accuracy comparable to equivalent non-spiking artificial neural networks. Due to higher mean input firing rate, the firing threshold is set to Vth=2.0V_{\text{th}} = 2.0.

    Model Network Structure Mean Accuracy (%) Accuracy Interval (%)
    Non-spiking MLP (BP) 784-400-10 98.31 [97.62, 98.57]
    Spiking MLP (STBP) 784-400-10 98.34 [97.94, 98.57]
    Non-spiking CNN (BP) 28×\times28×\times1-6C3-300-10 98.57 [98.57, 98.57]
    Spiking CNN (STBP) 28×\times28×\times1-6C3-300-10 98.59 [98.26, 98.89]

    Results are averaged across epochs 201 to 210. The Spiking MLP with STBP achieves 98.34% mean accuracy versus 98.31% for the non-spiking MLP, and the Spiking CNN with STBP achieves 98.59% mean accuracy versus 98.57% for the non-spiking CNN.

  9. Knowl 9 — Sensitivity of STBP to Surrogate Gradient Curve Shapes and Peak Widths

    empirical result

    Empirical evaluation on a 784-400-10 Spiking MLP trained on MNIST reveals that model performance is largely insensitive to the exact functional shape of the surrogate derivative (h1h_1: rectangular, h2h_2: polynomial, h3h_3: sigmoid derivative, h4h_4: Gaussian derivative), but sensitive to the peak width/steepness coefficient aia_i:

    • All four curve families with ai=1.0a_i = 1.0 exhibit virtually identical convergence trajectories and reach approximately 98.4%98.4\% accuracy after 200 epochs.
    • Sweeping the rectangular width parameter a1∈{0.1,1.0,2.5,5.0,7.5,10}a_1 \in \{0.1, 1.0, 2.5, 5.0, 7.5, 10\} shows that moderate values (a1∈[0.5,5.0]a_1 \in [0.5, 5.0]) yield rapid and stable convergence. Extremely small values (a1=0.1a_1 = 0.1, severe gradient sparsity) or large values (a1=10a_1 = 10, overly blurred gradient localization) degrade accuracy and slow convergence.
  10. Knowl 10 — Ablation of Temporal Domain Error Propagation (STBP vs. SDBP)

    empirical result

    Ablation experiments comparing Spatio-Temporal Backpropagation (STBP, propagating errors across both layers and time steps) against Spatial-Domain-only Backpropagation (SDBP, dropping intra-neuron temporal error feedback) on a 784-400-10 Spiking MLP demonstrate the necessity of temporal domain backpropagation:

    Model Dataset Mean Accuracy (%) Accuracy Interval (%)
    Spiking MLP (SDBP) Object Detection 97.11 [96.04, 97.78]
    Spiking MLP (SDBP) MNIST 98.29 [98.23, 98.39]
    Spiking MLP (STBP) Object Detection 98.32 [97.94, 98.57]
    Spiking MLP (STBP) MNIST 98.48 [98.42, 98.51]

    Removing the temporal backpropagation term (SDBP) causes a 1.21% drop in mean accuracy on the pedestrian detection dataset (97.11% vs. 98.32%) and a 0.19% drop on MNIST (98.29% vs. 98.48%), alongside higher performance variance across training epochs.

  11. Knowl 11 — CIFAR-10 Performance and Computational Bottlenecks of Direct SNN Training

    limitation

    Direct training of a spiking CNN (28×28×3→20C5→P2→30C5→P2→256→1028 \times 28 \times 3 \to 20\text{C}5 \to \text{P}2 \to 30\text{C}5 \to \text{P}2 \to 256 \to 10) with STBP on CIFAR-10 without data augmentation or regularizations (e.g., batch normalization, weight decay) achieves 50.7% testing accuracy after 100 epochs, compared to 52.9% for an equivalent non-spiking CNN.

    Direct supervised training of SNNs on larger datasets faces two major bottlenecks:

    1. Simulating the temporal kinetic states and threshold comparisons of SNNs over multiple discrete time steps on conventional CPU/GPU software requires 10×10\times to 100×100\times the runtime of an equivalent ANN.
    2. Optimization difficulty is elevated due to the discontinuous, non-differentiable spiking activity and multi-step spatio-temporal error surfaces on complex multi-class image tasks.

Coverage note — None was omitted; all key contributions—including the iterative LIF formulation, the STBP algorithm, surrogate derivative approximations, parameter initialization, convolutional SNN extensions, experimental benchmarks (MNIST, N-MNIST, pedestrian detection, CIFAR-10), and ablation studies—are covered.

References

  1. 1.Allen, J. N., Abdel-Aty-Zohdy, H. S., and Ewing, R. L. (2009). ‘‘Cognitive processing using spiking neural networks,’’ in IEEE 2009 National Aerospace and Electronics Conference (Dayton, OH), 56–64.
  2. 2.Bengio, Y., Mesnard, T., Fischer, A., Zhang, S., and Wu, Y. (2015). An objective function for stdp. Comput. Sci. preprint arXiv.
  3. 3.Benjamin, B. V., Gao, P., Mcquinn, E., Choudhary, S., Chandrasekaran, A. R., Bussat, J. M., et al. (2014). Neurogrid: a mixed-analog-digital multichip system for large-scale neural simulations. Proc. IEEE 102, 699–716. doi: 10.1109/JPROC.2014.2313565
  4. 4.Bohte, S. M., Kok, J. N., and Poutr, J. A. L. (2000). ‘‘Spikeprop: backpropagation for networks of spiking neurons,’’ in Esann 2000, European Symposium on Artificial Neural Networks, Bruges, Belgium, April 26-28, 2000, Proceedings (Bruges), 419–424.
  5. 5.Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2015). Gated feedback recurrent neural networks. Comput. Sci. 2067–2075. preprint arXiv.
  6. 6.Cohen, G. K., Orchard, G., Leng, S. H., Tapson, J., Benosman, R. B., and Schaik, A. V. (2016). Skimming digits: neuromorphic classification of spike-encoded images. Front. Neurosci. 10:184. doi: 10.3389/fnins.2016.00184
  7. 7.Davis, K. H., Biddulph, R., and Balashek, S. (1952). Automatic recognition of spoken digits. J. Acoust. Soc. Am. 24:637. doi: 10.1121/1.19 06946
  8. 8.Diehl, P. U., and Cook, M. (2015). Unsupervised learning of digit recognition using spike-timing-dependent plasticity. Front. Comput. Neurosci. 9:99. doi: 10.3389/fncom.2015.00099
  9. 9.Diehl, P. U., Neil, D., Binas, J., and Cook, M. (2015). ‘‘Fast-classifying, high-accuracy spiking deep networks through weight and threshold balancing’’ in International Joint Conference on Neural Networks (Killarney), 1–8.
  10. 10.Esser, S. K., Merolla, P. A., Arthur, J. V., Cassidy, A. S., Appuswamy, R., Andreopoulos, A., et al. (2016). Convolutional networks for fast, energy-efficient neuromorphic computing. Proc. Natl. Acad. Sci. U.S.A. 113, 11441– 11446. doi: 10.1073/pnas.1604850113
  11. 11.Furber, S. B., Galluppi, F., Temple, S., and Plana, L. A. (2014). The spinnaker project. Proc. IEEE 102, 652–665. doi: 10.1109/JPROC.2014. 2304638
  12. 12.Garofolo, J. S., Lamel, L. F., Fisher, W. M., Fiscus, J. G., Pallett, D. S., Dahlgren, N. L. (1993). DARPA TIMIT Acoustic-Phonetic Continous Speech Corpus CD-ROM. NIST Speech Disc 1-1.1. NASA STI/Recon Technical Report n 93.
  13. 13.Gers, F. A., Schmidhuber, J., and Cummins, F. (1999). Learning to forget: continual prediction with lstm. Neural Comput. 12, 2451–2571. doi: 10.1162/089976600300015015
  14. 14.Gtig, R., and Sompolinsky, H. (2006). The tempotron: a neuron that learns spike timing-based decisions. Nat. Neurosci. 9:420–428. doi: 10.1038/nn1643
  15. 15.Hochreiter, S., and Schmidhuber, J. (1997). Long short-term memory. Neural Comput. 9, 1735–1780. doi: 10.1162/neco.1997.9.8.1735
  16. 16.Hunsberger, E., and Eliasmith, C. (2015). Spiking deep networks with lif neurons. Comput. Sci. preprint arXiv.
  17. 17.Hwu, T., Isbell, J., Oros, N., and Krichmar, J. (2016). A self-driving robot using deep convolutional neural networks on neuromorphic hardware. arXiv.org.
  18. 18.Kasabov, N., and Capecci, E. (2015). Spiking neural network methodology for modelling, classification and understanding of eeg spatio-temporal data measuring cognitive processes. Infor. Sci. 294, 565–575. doi: 10.1016/j.ins.2014.06.028
  19. 19.Kheradpisheh, S. R., Ganjtabesh, M., and Masquelier, T. (2016). Bio-inspired unsupervised learning of visual features leads to robust invariant object recognition. Neurocomputing 205, 382–392. doi: 10.1016/j.neucom.2016.04.029
  20. 20.Kingma, D., and Ba, J. (2015). ‘‘Adam: A method for stochastic optimization,’’ in International Conference on Learning Representations (ICLR).
  21. 21.Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998). Gradient-based learning applied to document recognition. Proc. IEEE 86, 2278–2324. doi: 10.1109/5.726791
  22. 22.Lee, J. H., Delbruck, T., and Pfeiffer, M. (2016). Training deep spiking neural networks using backpropagation. Front. Neurosci. 10:508. doi: 10.3389/fnins.2016.00508
  23. 23.Lichtsteiner, P., Posch, C., and Delbruck, T. (2007). A 128x128 120db 15us latency asynchronous temporal contrast vision sensor. IEEE J. Solid State Circ. 43, 566–576. doi: 10.1109/JSSC.2007.914337
  24. 24.Mckennoch, S., Liu, D., and Bushnell, L. G. (2006). ‘‘Fast modifications of the spikeprop algorithm,’’ in IEEE International Joint Conference on Neural Network Proceedings (Vancouver, BC), 3970–3977.
  25. 25.Merolla, P. A., Arthur, J. V., Alvarezicaza, R., Cassidy, A. S., Sawada, J., Akopyan, F., et al. (2014). Artificial brains. A million spiking-neuron integrated circuit with a scalable communication network and interface. Science 345, 668–673. doi: 10.1126/science.1254642
  26. 26.Mozafari, M., Kheradpisheh, S. R., Masquelier, T., Nowzaridalini, A., and Ganjtabesh, M. (2017). First-spike based visual categorization using reward-modulated stdp. preprint arXiv.
  27. 27.Neftci, E., Das, S., Pedroni, B., Kreutzdelgado, K., and Cauwenberghs, G. (2013). Event-driven contrastive divergence for spiking neuromorphic systems. Front. Neurosci. 7:272. doi: 10.3389/fnins.2013.00272
  28. 28.Neil, D., and Liu, S. C. (2016). ‘‘Effective sensor fusion with event-based sensors and deep network architectures,’’ in IEEE International Symposium on Circuits and Systems, ed O. Rene Levesque (Montreal, QC).
  29. 29.Neil, D., Pfeiffer, M., and Liu, S. C. (2016). Phased lstm: accelerating recurrent network training for long or event-based sequences. arXiv.org.
  30. 30.O’Connor, P., and Welling, M. (2016). Deep spiking networks. arXiv.org.
  31. 31.Orchard, G., Jayawant, A., Cohen, G. K., and Thakor, N. (2015). Converting static image datasets to spiking neuromorphic datasets using saccades. Front. Neurosci. 9:437. doi: 10.3389/fnins.2015.00437
  32. 32.Peter, O., Daniel, N., Liu, S. C., Tobi, D., and Michael, P. (2013). Real-time classification and sensor fusion with a spiking deep belief network. Front. Neurosci. 7:178. doi: 10.3389/fnins.2013.00178
  33. 33.Ponulak, F., (2005). ReSuMe-New Supervised Learning Method for Spiking Neural Networks[J]. Institute of Control and Information Engineering, Poznan University of Technology.
  34. 34.Ponulak, F., and Kasiski, A. (2010). Supervised learning in spiking neural networks with resume: sequence learning, classification, and spike shifting. Neural Comput. 22, 467–510. doi: 10.1162/neco.2009.11-08-901
  35. 35.Querlioz, D., Bichler, O., Dollfus, P., and Gamrat, C. (2013). Immunity to device variations in a spiking neural network with memristive nanodevices. IEEE Trans. Nanotechnol. 12, 288–295. doi: 10.1109/TNANO.2013.2250995
  36. 36.Schrauwen, B., and Campenhout, J. V. (2004). Extending spikeprop. arXiv.org 1, 471–475.
  37. 37.Simard, P. Y., Steinkraus, D., and Platt, J. C. (2003). ‘‘Best practices for convolutional neural networks applied to visual document analysis,’’ in International Conference on Document Analysis and Recognition (Edinburgh), 958.
  38. 38.Tavanaei, A., and Maida, A. S. (2017). Bio-inspired spiking convolutional neural network using layer-wise sparse coding and stdp learning. preprint arXiv.
  39. 39.Urbanczik, R., and Senn, W. (2009). A gradient learning rule for the tempotron. Neural Comput. 21, 340–352. doi: 10.1162/neco.2008.09-07-605
  40. 40.Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P. A. (2008). ‘‘Extracting and composing robust features with denoising autoencoders,’’ in International Conference on Machine Learning (Helsinki), 1096–1103.
  41. 41.Werbos, P. J. (1990). Backpropagation through time: what it does and how to do it. Proc. IEEE 78, 1550–1560. doi: 10.1109/5.58337
  42. 42.Zhang, B., Shi, L., and Song, S. (2016). Creating more intelligent robots through brain-inspired computing. Science 354:1445. doi: 10.1126/science.354.6318.1445-b
  43. 43.Zhang, X., Xu, Z., Henriquez, C., and Ferrari, S. (2013). ‘‘Spike-based indirect training of a spiking neural network-controlled virtual insect,’’ in Decision and Control (CDC), 2013 IEEE 52nd Annual Conference on (Florence: IEEE), 6798–6805.

Citation

MLA
Wu, Y., et al. “Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks”. Frontiers in Neuroscience, vol. 12, 2018, https://doi.org/10.3389/fnins.2018.00331.
APA
Wu, Y., Deng, L., Li, G., Zhu, J., & Shi, L. (2018). Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks. Frontiers in Neuroscience, 12. https://doi.org/10.3389/fnins.2018.00331
Chicago
Wu, Y., L. Deng, G. Li, J. Zhu, and L. Shi. 2018. “Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks”. Frontiers in Neuroscience 12. https://doi.org/10.3389/fnins.2018.00331.
Harvard
Wu, Y. et al. (2018) “Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks”, Frontiers in Neuroscience, 12. Available at: https://doi.org/10.3389/fnins.2018.00331.
Vancouver
1. Wu Y, Deng L, Li G, Zhu J, Shi L (2018) Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks. Frontiers in Neuroscience. https://doi.org/10.3389/fnins.2018.00331

BibTeX

@article{Wu_2018, title={Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks}, volume={12}, ISSN={1662-453X}, url={http://dx.doi.org/10.3389/fnins.2018.00331}, DOI={10.3389/fnins.2018.00331}, journal={Frontiers in Neuroscience}, publisher={Frontiers Media SA}, author={Wu, Yujie and Deng, Lei and Li, Guoqi and Zhu, Jun and Shi, Luping}, year={2018}, month=May }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF