Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks
Yujie WuLei DengGuoqi LiJun ZhuLuping Shi
Proposes a spatio-temporal backpropagation framework with surrogate gradient approximation for iterative leaky integrate-and-fire models, overcoming the non-differentiability of spikes to achieve high-performance direct supervised training of spiking neural networks on both static and neuromorphic benchmarks.
Spiking neural networks are energy-efficient, brain-inspired models well-suited for processing event-driven, time-dependent data on specialized hardware. However, training them directly using gradient-based learning has historically been difficult due to the discontinuous, non-differentiable nature of spiking events and the common neglect of timing dynamics during optimization. Existing approaches typically convert pre-trained conventional models or rely on complex mathematical workarounds and regularizations that limit overall performance.
The article develops and evaluates a spatio-temporal backpropagation training algorithm that combines spatial layer-to-layer error propagation with temporal timing dynamics. To enable gradient-based optimization, the authors establish an iterative leaky integrate-and-fire neuronal model and introduce continuous approximation curves to handle the non-differentiable spiking threshold.
The evaluation tests both fully connected and convolutional spiking architectures across static visual datasets (MNIST handwritten digits and a custom pedestrian detection dataset) and dynamic event-based sensor data (N-MNIST). The authors also assess model sensitivity to various derivative approximation curves, curve widths, and the removal of the temporal domain component.
The findings show that the proposed method establishes state-of-the-art accuracy across all tested benchmarks for spiking architectures. The algorithm achieved 98.89% accuracy on static MNIST with a fully connected network and 99.42% with a convolutional architecture. On dynamic N-MNIST data, the model reached 98.78% accuracy, outperforming existing spiking networks as well as standard non-spiking deep neural networks. Ablation experiments demonstrated that including temporal dynamics prevents performance degradation and stabilizes training, while the specific shape of the derivative approximation curve matters far less than choosing an appropriate curve steepness.
These results indicate that directly training spiking networks across both spatial and temporal dimensions yields high accuracy without requiring specialized tricks such as weight normalization or complex reset rules. This significantly lowers algorithmic complexity, making the approach appealing for practical applications on energy-efficient neuromorphic hardware and edge devices.
Organizations developing brain-inspired AI and low-power sensory hardware should adopt spatio-temporal gradient frameworks for direct model training rather than relying on cumbersome translation pipelines. Future efforts should prioritize accelerating software simulation routines—which remain considerably slower than conventional network training—and testing the algorithm on richer temporal benchmarks such as speech datasets and larger-scale computer vision tasks.
Confidence in these findings is high for standard image classification and dynamic vision benchmarks, but caution is warranted when scaling to larger, more complex datasets like CIFAR-10, where high software simulation runtimes and lack of advanced optimization techniques currently limit absolute accuracy.
- Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). Introduces straight-through and continuous gradient estimators for backpropagating through non-differentiable hard thresholds, establishing the mathematical prerequisite for surrogate gradient approximations in spiking neurons.
- Paper: Learning representations by back-propagating errors, David E. Rumelhart et al. (1986). Presents the foundational backpropagation algorithm for computing parameter gradients across network layers, which the source extends to the temporal domain and non-smooth dynamics of spiking networks.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Establishes the backpropagation-through-time framework for recurrent architectures and discrete time-unrolled computational graphs, directly informing the iterative state modeling in spatio-temporal backpropagation.
- Paper: On the difficulty of training recurrent neural networks, Razvan Pascanu et al. (2012). Analyzes the mechanics and mathematical conditions of vanishing and exploding gradients in time-unrolled networks, providing essential context for the stability challenges of temporal credit assignment.
- Paper: Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-based optimization to spiking neural networks, Emre O. Neftci et al. (2019). Synthesizes and formalizes surrogate gradient optimization across recurrent and deep spiking neural network architectures, generalizing the continuous derivative approximations introduced in STBP.
- Paper: Deep Learning in Spiking Neural Networks, Amirhossein Tavanaei et al. (2018). Reviews modern deep spiking neural networks and places direct spatio-temporal gradient descent methods into the broader landscape of conversion techniques and biologically plausible learning rules.
