RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks
Yufei GuoXinyi TongYuanpei ChenLiwen ZhangXiaode LiuZhe MaXuhui Huang
Proposes a novel distribution loss that penalizes membrane potential shifts and reduces quantization errors during training, enabling direct optimization of deeper and more accurate spiking neural networks with fewer timesteps and zero inference overhead.
Modern artificial intelligence requires substantial computing power and energy, creating severe bottlenecks for edge devices and latency-critical systems. Brain-inspired Spiking Neural Networks offer an energy-efficient alternative by communicating through sparse, binary spike events rather than high-precision continuous values. However, directly training deep spiking models using gradient descent remains difficult because continuous neural voltages, known as membrane potentials, shift undesirably during information processing. These distribution shifts trigger three primary failures: degeneration, where neurons emit identical signals and destroy information flow; saturation, where gradients drop to zero and halt optimization; and gradient mismatch, where mathematical approximations used during back-propagation accumulate substantial estimation errors.
The article develops and evaluates a targeted training framework called RecDis-SNN to rectify these internal voltage distributions. Rather than adding complex model parameters or runtime normalization layers, the approach introduces an explicit regularization loss function during training to penalize degeneration, saturation, and gradient mismatch simultaneously.
The researchers assessed the method through comprehensive classification experiments across standard visual benchmarks, including CIFAR-10, CIFAR-100, and ImageNet, as well as a neuromorphic event-stream dataset, DVS-CIFAR10. They tested multiple deep network backbones, such as CIFARNet, VGG-16, ResNet-19, and ResNet-34, analyzing classification accuracy, latency requirements across discrete timesteps, training computational overhead, and internal quantization error.
The evaluation revealed several critical findings. First, the proposed regularizer prevented optimization breakdown and allowed randomly initialized networks to converge successfully, boosting standalone accuracy on CIFAR-10 by up to 3.11 percentage points over unregularized baselines. Second, the framework achieved state-of-the-art accuracy across all evaluated benchmarks while reducing latency; on CIFAR-10, it attained 95.55% accuracy within only 6 timesteps, and on ImageNet, it achieved 67.33% accuracy. Third, on neuromorphic event-stream data, the approach demonstrated substantial gains, outperforming prior state-of-the-art models on DVS-CIFAR10 by 4.62 to 6.80 percentage points across matching architectures. Fourth, the loss function reduced average internal quantization error by roughly 15% to 26% compared to threshold-dependent batch normalization because it naturally shapes internal voltages into a bimodal distribution. Finally, these improvements incurred only a modest training time overhead of under 17% and introduced zero extra computational operations during operational deployment.
These results establish that internal voltage rectification resolves fundamental optimization barriers in deep spiking models. By addressing gradient estimation mismatch and quantization error in addition to standard variance shifts, the framework enables high-accuracy neuromorphic models that require fewer time steps to process information. This significantly lowers runtime energy consumption and computational latency without introducing inference penalties.
Engineering teams deploying neural networks to energy-constrained hardware should consider integrating membrane potential distribution loss directly into their training pipelines. The method can serve as a standalone regularizer or be combined with existing batch normalization techniques for further performance gains. For future research, teams should explore applying this distribution control framework to other complex vision and temporal tasks beyond standard image classification.
The findings are supported by consistent multi-trial experiments across varied benchmark datasets and architectures. Nevertheless, potential users should note that the mathematical loss formulation assumes internal neuron voltages approximate a Gaussian distribution, and experimental evaluations were centered on visual classification tasks. Organizations should conduct targeted validation on their specific domain-specific tasks and neuromorphic hardware platforms prior to full-scale production deployment.
- Paper: Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-based optimization to spiking neural networks, Emre O. Neftci et al. (2019). This paper establishes the foundational surrogate gradient learning framework for overcoming non-differentiable spiking dynamics in SNNs, whose estimation errors RecDis-SNN directly regularizes.
- Paper: Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks, Yujie Wu et al. (2017). It introduces spatio-temporal backpropagation and continuous threshold approximations for directly training deep spiking neural networks, providing the baseline direct-training formulation that RecDis-SNN optimizes.
- Paper: Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift, Sergey Ioffe et al. (2015). It introduces batch normalization to control internal activation distribution shifts, serving as the benchmark normalization technique against which RecDis-SNN's distribution loss is compared.
- Paper: Deep Learning in Spiking Neural Networks, Amirhossein Tavanaei et al. (2018). This survey provides an essential overview of direct supervised training challenges and conversion methods in deep spiking neural networks.
- Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). It introduces foundational techniques for estimating and propagating gradients through hard thresholds and stochastic neurons, which underpin gradient estimation methods in SNNs.
- Paper: Adaptive Smoothing Gradient Learning for Spiking Neural Networks, Ziming Wang et al. (2023). This work directly extends the line of research on mitigating gradient estimation mismatch in SNNs by adaptively tuning surrogate relaxation degrees throughout training.
- Paper: Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting, Shikuang Deng et al. (2022). It presents a complementary temporal re-weighting optimization approach to improve gradient estimation and landscape convergence in directly trained deep SNNs.
- Paper: Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation, Qingyan Meng et al. (2022). This paper offers an alternative low-latency direct training paradigm by differentiating directly over continuous spike rate representations to bypass temporal surrogate backpropagation errors.
- Paper: SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks, Xinyu Shi et al. (2024). It builds upon deep spiking architectures and low-latency direct training advances to scale SNNs effectively to hybrid vision transformer backbones.
- Paper: Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies, Wei Fang et al. (2023). It addresses the latency and temporal dynamics bottlenecks of spiking neurons by eliminating the reset mechanism to enable parallel state computation across time steps.
- Paper: ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural Networks, Jiangrong Shen et al. (2023). This work explores structural plasticity and dynamic sparse connectivity during the direct training of deep SNNs to reduce hardware and training overhead.
