Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting
Shikuang DengYuhang LiShanghang ZhangShi Gu
Proposes a temporal efficient training method that re-weights surrogate gradients in spiking neural networks to converge to flatter minima, substantially boosting generalization performance across standard vision and neuromorphic datasets.
Brain-inspired spiking neural networks (SNNs) offer major advantages for energy-efficient computing and fast processing on specialized neuromorphic hardware. However, training deep spiking networks directly from scratch has remained a critical bottleneck. Because spiking neurons rely on non-differentiable binary activations, standard gradient-based optimization cannot be directly applied. Conventional direct training methods use surrogate gradients to approximate derivatives, but this creates a fundamental mismatch between the loss landscape and gradient estimates. As a result, standard training easily gets trapped in sharp local minima, yielding poor generalizability on test data and high computational training costs.
The main objective of the article is to demonstrate that optimizing network outputs at every individual time step, rather than only evaluating the time-averaged final output, fundamentally improves generalization and training efficiency in deep spiking neural networks. To address the issue, the authors introduce the Temporal Efficient Training (TET) framework and a complementary Time Inheritance Training (TIT) acceleration scheme, evaluating their mathematical foundations and empirical performance across diverse image and neuromorphic datasets.
The authors designed a re-weighted loss function that constrains the pre-synaptic output distribution at each discrete time step alongside a regularizing mean squared error term. They evaluated the method using standard convolutional and residual network architectures—including ResNet-19, ResNet-34, and custom architectures—tested on standard static image benchmarks (CIFAR-10, CIFAR-100, and ImageNet) as well as the challenging neuromorphic dataset DVS-CIFAR10. They also carried out loss landscape visualizations, convergence proofs, and ablation studies comparing standard training against the new framework.
The analysis produced several vital findings. First, the proposed training method consistently outperformed existing state-of-the-art spiking models across all evaluated benchmarks. On the challenging DVS-CIFAR10 neuromorphic dataset, the method achieved an 83.17% top-1 accuracy, representing an improvement of over 10% compared to prior published results. Second, on CIFAR-100, the method improved accuracy by more than 3% across various simulation lengths, coming within 0.63% of equivalent artificial neural network baselines. Third, loss landscape analysis confirmed that the approach reliably guides optimization toward flatter, highly generalizable minima while standard training stalls in sharp valleys. Finally, the time inheritance scheme cut total training time roughly in half by first training on short simulation lengths before fine-tuning on longer time horizons.
These findings indicate that directly optimizing temporal dynamics resolves the long-standing generalization deficit of spiking networks without modifying their energy-efficient inference mechanisms. By converging to flatter loss regions, spiking models achieve performance parity with conventional neural networks while maintaining low computational and energy footprints. The halved training timeline substantially mitigates the development expense and carbon footprint associated with training temporal models.
Organizations developing low-power artificial intelligence, edge devices, or neuromorphic applications should adopt per-step temporal training and inheritance strategies to accelerate development cycles and boost model accuracy. Teams deploying these pipelines should tune the regularization strength carefully based on task type, using moderate penalties for static datasets and smaller penalties for sparse neuromorphic streams. Further work is recommended to validate the approach across larger-scale neuromorphic video streams and explore automatic hyperparameter scheduling during temporal scaling.
The reported evidence provides high confidence in the method's effectiveness across standard image classification benchmarks. Nevertheless, practitioners should exercise caution regarding specific limitations: neuromorphic datasets exhibit high noise and sparsity, which can disrupt early time-step training if regularization hyperparameters are miscalibrated. Additionally, while the training efficiency gains are substantial, the training phase still incurs recurrent memory overhead that requires careful hardware resource management.
- Paper: Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-based optimization to spiking neural networks, Emre O. Neftci et al. (2019). This paper establishes the foundational framework of surrogate gradient learning in spiking neural networks to overcome the non-differentiability barrier, providing essential background for the optimization issues addressed by TET.
- Paper: Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks, Yujie Wu et al. (2017). It introduces spatio-temporal backpropagation through iterative leaky integrate-and-fire dynamics with continuous derivative approximations, forming the direct training baseline that TET seeks to improve.
- Paper: Deep Learning in Spiking Neural Networks, Amirhossein Tavanaei et al. (2018). This comprehensive survey outlines the key paradigms and challenges in direct training and gradient approximation for deep spiking neural networks.
- Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). It provides fundamental principles for estimating and propagating gradients through hard non-differentiable thresholds and stochastic binary units, underpinning modern surrogate gradient techniques.
- Paper: Averaging Weights Leads to Wider Optima and Better Generalization, Pavel Izmailov et al. (2018). This work establishes the connection between converging to wider loss-landscape optima and achieving superior generalization, which motivates TET's gradient re-weighting objective.
No sufficiently relevant recommendations were found.
