Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies

Wei FangZhaofei YuZhaokun ZhouDing ChenYanqi ChenZhengyu MaTimothée MasquelierYonghong Tian

article2023NeurIPS102 citations

Proposes a parallelizable formulation of spiking neurons by removing iterative reset dynamics, enabling significantly faster simulation speeds and improved long-term dependency learning across static and sequential benchmarks.

Listen

Spiking neural networks offer extreme energy efficiency for artificial intelligence by mimicking biological brain communication using discrete, binary spikes. Despite their low power potential on specialized hardware, standard spiking models are slow to train and deploy because their traditional charge-fire-reset computing mechanism operates serially over time. This step-by-step dependency limits the ability to exploit modern parallel computing hardware and hinders the network's capacity to learn long-term temporal relationships.

The article demonstrates that eliminating the reset step allows the internal states of spiking neurons to be computed in parallel across time without degrading performance. Based on this insight, the authors introduce the Parallel Spiking Neuron along with two variants—a masked version for real-time, step-by-step processing and a sliding version with shared parameters across time for variable-length inputs.

To validate this framework, the authors conducted simulation speed benchmarks on graphics processors and evaluated classification accuracy across sequential, static, and neuromorphic image datasets using standard benchmarking environments.

The findings establish that parallel spiking neurons run substantially faster during both training and inference compared to traditional models, cutting computational complexity across time-steps from linear to logarithmic or matrix-parallel forms. In sequential image recognition tasks, the proposed models achieved superior accuracy, reaching 88.45% on sequential CIFAR-10 compared to 83.66% for the best-performing traditional gated spiking neuron. On large-scale static benchmarks such as ImageNet, integrating parallel spiking neurons increased model accuracy by over 3 percentage points compared to standard implementations. Furthermore, on neuromorphic vision tasks, the sliding model reached 82.30% accuracy with only 4 time-steps, marking the first work to exceed 80% accuracy with such low latency.

These results indicate that deep spiking networks can achieve higher accuracy and significantly reduced training times while adding less than 0.003% additional parameters and maintaining lower memory overhead. This shifts spiking neural networks from computationally prohibitive serial simulations to scalable, parallelized deep learning pipelines suitable for commercial graphics hardware.

Organizations developing low-power artificial intelligence solutions should consider adopting parallel spiking architectures to accelerate model training and improve temporal processing accuracy. For streaming or variable-length inputs, adopting the sliding or masked configurations provides the optimal balance between latency and accuracy. Further research should evaluate these parallel formulations directly on physical neuromorphic hardware to confirm that energy efficiency gains translate seamlessly from simulation to production chips.

The primary limitation of the base parallel neuron is increased latency, as it requires the full sequence of inputs before producing outputs; however, the masked and sliding variants effectively resolve this trade-off for sequential workflows. Firing rates are slightly higher when reset is omitted, but the resulting impact on power consumption remains minor compared to the substantial gains in execution speed and predictive accuracy.

arXiv: 2304.12760
Cover for Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies

Abstract

Vanilla spiking neurons in Spiking Neural Networks (SNNs) use charge-fire-reset neuronal dynamics, which can only be simulated serially and can hardly learn long-time dependencies. We find that when removing reset, the neuronal dynamics can be reformulated in a non-iterative form and parallelized. By rewriting neuronal dynamics without reset to a general formulation, we propose the Parallel Spiking Neuron (PSN), which generates hidden states that are independent of their predecessors, resulting in parallelizable neuronal dynamics and extremely high simulation speed. The weights of inputs in the PSN are fully connected, which maximizes the utilization of temporal information. To avoid the use of future inputs for step-by-step inference, the weights of the PSN can be masked, resulting in the masked PSN. By sharing weights across time-steps based on the masked PSN, the sliding PSN is proposed to handle sequences of varying lengths. We evaluate the PSN family on simulation speed and temporal/static data classification, and the results show the overwhelming advantage of the PSN family in efficiency and accuracy. To the best of our knowledge, this is the first study about parallelizing spiking neurons and can be a cornerstone for the spiking deep learning research. Our codes are available at https://github.com/fangwei123456/Parallel-Spiking-Neuron.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Deep Learning for Spiking Neural Networks
  • 2.2 Improvement of Spiking Neurons
  • 2.3 Improvement of Training Methods and Network Structure
  • 2.4 Acceleration for Sequence Processing
  • 3 Methods
  • 3.1 Vanilla Spiking Neurons W/WO Reset
  • 3.2 Parallel Spiking Neuron
  • 3.3 k-Order Masked Parallel Spiking Neuron
  • 3.4 k-Order Sliding Parallel Spiking Neuron
  • 3.5 Summary of the PSN Family
  • 4 Experiments
  • 4.1 Simulation Speed Benchmark
  • 4.2 Learning Long-term Dependencies
  • 4.3 Static and Neuromorphic Data Classification
  • 5 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Parallel Spiking Neuron Formulation

    model/method

    In standard spiking neurons, charging, firing, and resetting create sequential temporal dependencies (H[t]=f(V[t−1],X[t])H[t] = f(V[t-1], X[t]), S[t]=Θ(H[t]−Vth)S[t] = \Theta(H[t] - V_{\text{th}}), and V[t]=g(H[t],S[t])V[t] = g(H[t], S[t])). When the reset operation is omitted, the membrane potential accumulation simplifies to a non-iterative linear combination over time-steps:

    H[t]=∑i=0T−1Wt,i⋅X[i]H[t] = \sum_{i=0}^{T-1} W_{t,i} \cdot X[i]

    The Parallel Spiking Neuron (PSN) generalizes this into fully learnable matrix operations:

    H=WXH = W X

    S=Θ(H−B)S = \Theta(H - B)

    where X∈RT×NX \in \mathbb{R}^{T \times N} is the input current tensor over TT time-steps and batch size NN, W∈RT×TW \in \mathbb{R}^{T \times T} is a learnable temporal weight matrix, B∈RTB \in \mathbb{R}^{T} is a learnable threshold vector, H∈RT×NH \in \mathbb{R}^{T \times N} represents the hidden membrane potential state, Θ(⋅)\Theta(\cdot) is the Heaviside step function (Θ(x)=1\Theta(x) = 1 for x≥0x \ge 0, and 00 otherwise), and S∈{0,1}T×NS \in \{0, 1\}^{T \times N} is the output binary spike tensor.

    Because H[t]H[t] connects directly to all time-step inputs X[i]X[i] rather than relying on iterative recurrence from H[t−1]H[t-1], the PSN acts as a TT-order neuron and enables fully parallelized forward and backward execution on parallel hardware.

  2. Knowl 2 — k-Order Masked Parallel Spiking Neuron

    model/method

    Unconstrained Parallel Spiking Neurons require all future inputs X[i]X[i] (i≥ti \ge t) to evaluate H[t]H[t], which imposes a latency of TT time-steps. To enable causal, step-by-step inference where H[t]H[t] depends only on current and past inputs within a temporal window of size kk, the kk-order masked PSN applies an element-wise band mask MkM_k to the weight matrix WW:

    H=(W⊙Mk)XH = (W \odot M_k)X

    S=Θ(H−B)S = \Theta(H - B)

    where W∈RT×TW \in \mathbb{R}^{T \times T}, B∈RTB \in \mathbb{R}^T, X∈RT×NX \in \mathbb{R}^{T \times N}, S∈{0,1}T×NS \in \{0, 1\}^{T \times N}, and ⊙\odot represents the Hadamard (element-wise) product. The binary mask matrix Mk∈{0,1}T×TM_k \in \{0, 1\}^{T \times T} is defined as:

    Mk[i][j]={1,j≤i≤j+k−10,otherwiseM_k[i][j] = \begin{cases} 1, & j \le i \le j + k - 1 \\ 0, & \text{otherwise} \end{cases}

    Assuming zero-padding for non-existent past inputs (X[i]=0X[i] = 0 for all i<0i < 0), H[t]H[t] is determined exclusively by the recent inputs {X[t−k+1],…,X[t]}\{X[t - k + 1], \dots, X[t]\}, permitting output spike S[t]S[t] to be generated as soon as X[t]X[t] arrives.

  3. Knowl 3 — Progressive Masking Training Schedule for Masked PSN

    algorithm

    To improve optimization stability when training a kk-order masked Parallel Spiking Neuron, a progressive masking schedule smoothly transitions the effective mask from an all-ones matrix to the target causal band mask MkM_k.

    Input: Total training epochs EE, current epoch index e∈{0,…,E−1}e \in \{0, \dots, E-1\}, target causal mask Mk∈{0,1}T×TM_k \in \{0, 1\}^{T \times T}, all-ones matrix 1∈RT×T\mathbf{1} \in \mathbb{R}^{T \times T}
    Output: Interpolated mask matrix Mk(λ)∈[0,1]T×TM_k(\lambda) \in [0, 1]^{T \times T}
    λ←min⁡(1,8⋅eE−1)\lambda \leftarrow \min\left(1, \frac{8 \cdot e}{E - 1}\right)
    Mk(λ)←λ⋅Mk+(1−λ)⋅1M_k(\lambda) \leftarrow \lambda \cdot M_k + (1 - \lambda) \cdot \mathbf{1}
    return Mk(λ)M_k(\lambda)

    During early epochs, λ≈0\lambda \approx 0 allows gradients to flow across all time-steps to establish suitable initial parameters before the network is constrained to strict causality as λ\lambda reaches 11 (at 12.5%12.5\% of total training epochs).

  4. Knowl 4 — k-Order Sliding Parallel Spiking Neuron

    model/method

    To handle input sequences of arbitrary or variable length TT, the kk-order sliding Parallel Spiking Neuron (SPSN) shares learnable temporal weights across time-steps rather than using time-wise parameters. The neuronal dynamics follow a 1D convolution over input history:

    H[t]=∑i=0k−1Wi⋅X[t−k+1+i]H[t] = \sum_{i=0}^{k-1} W_i \cdot X[t - k + 1 + i]

    S[t]=Θ(H[t]−Vth)S[t] = \Theta(H[t] - V_{\text{th}})

    where W=[W0,W1,…,Wk−1]∈RkW = [W_0, W_1, \dots, W_{k-1}] \in \mathbb{R}^k is a learnable kernel vector, Vth∈RV_{\text{th}} \in \mathbb{R} is a scalar learnable firing threshold, and X[j]=0X[j] = 0 for all j<0j < 0.

    For an input sequence X∈RT×NX \in \mathbb{R}^{T \times N} of known length TT, this 1D convolution is computed as a matrix-matrix multiplication H=AXH = AX, where the matrix A∈RT×TA \in \mathbb{R}^{T \times T} is defined as:

    A[i][j]={Wk−1−i+j,i+1−k≤j≤i0,otherwiseA[i][j] = \begin{cases} W_{k - 1 - i + j}, & i + 1 - k \le j \le i \\ 0, & \text{otherwise} \end{cases}

    Evaluating sliding dynamics via matrix-matrix multiplication yields faster execution on GPUs than standard 1D convolution primitives.

  5. Knowl 5 — Parameter Complexity of the PSN Family

    theoretical result

    For a sequence length TT and history order kk, the number of learnable parameters nparamn_{\text{param}} in each variant of the Parallel Spiking Neuron family is:

    • Parallel Spiking Neuron (PSN): nparam=T2+Tn_{\text{param}} = T^2 + T (T×TT \times T weights in WW plus TT threshold values in BB).
    • kk-Order Masked PSN: nparam=(2T+1−k)k2+Tn_{\text{param}} = \frac{(2T + 1 - k)k}{2} + T (active masked weights in WW plus TT threshold values in BB).
    • kk-Order Sliding PSN (SPSN): nparam=k+1n_{\text{param}} = k + 1 (kk kernel weights in WW plus 1 scalar threshold VthV_{\text{th}}).

    In deep Spiking Neural Networks trained with small temporal steps (e.g., T=4T=4), this parameter overhead is negligible: adding PSN to Spiking ResNet-18 adds 340 parameters (+0.00291%+0.00291\%) and to VGG-11 adds 200 parameters (+0.00015%+0.00015\%). Furthermore, PSN maintains only one hidden state tensor HH in memory during backpropagation, compared to two states (HH and VV) in standard charge-fire-reset neurons.

  6. Knowl 6 — Simulation Speedup of Parallel Spiking Neurons over LIF Neurons

    empirical result

    Execution time benchmarks comparing the Parallel Spiking Neuron (PSN) against the standard Leaky Integrate-and-Fire (LIF) neuron across neuron counts N∈{28,212,216,220}N \in \{2^8, 2^{12}, 2^{16}, 2^{20}\} and time-steps T∈{2,4,8,16,32,64}T \in \{2, 4, 8, 16, 32, 64\} demonstrate substantial simulation acceleration:

    • Inference (Forward): Even when LIF charging, firing, and resetting equations across all time-steps are compiled into a single fused GPU kernel using PyTorch Just-In-Time (JIT), the ratio of execution times tLIF/tPSNt_{\text{LIF}} / t_{\text{PSN}} reaches up to ∼23×\sim 23\times at T=64T=64.
    • Training (Forward + Backward): Because surrogate gradient functions in the backward pass prevent full-loop JIT kernel fusion across time-steps (requiring separate kernel launches per time-step for LIF), the ratio tLIF/tPSNt_{\text{LIF}} / t_{\text{PSN}} exceeds 30×30\times at T=64T=64 for N∈{28,212}N \in \{2^8, 2^{12}\}.

    PSN maintains identical implementations for inference and training based on cuBLAS/MKL matrix multiplication, bypassing sequential time-stepping entirely.

  7. Knowl 7 — Sequential CIFAR Classification and Ablation of Neuronal Reset

    data/table

    To assess long-term dependency modeling, spiking neurons were evaluated on sequential CIFAR-10 and CIFAR-100 tasks, where images are processed column-by-column across T=32T=32 time-steps using 1D convolutional networks.

    Dataset PSN Masked PSN Sliding PSN GLIF KLIF PLIF LIF LIF wo reset
    Sequential CIFAR-10 88.45% 85.81% 86.70% 83.66% 83.26% 83.49% 81.50% 79.50%
    Sequential CIFAR-100 62.21% 60.69% 62.11% 58.92% 57.37% 57.55% 55.45% 53.33%

    The accuracy ranking is PSN>Sliding PSN>Masked PSN>GLIF>PLIF>KLIF>LIF>LIF without reset\text{PSN} > \text{Sliding PSN} > \text{Masked PSN} > \text{GLIF} > \text{PLIF} > \text{KLIF} > \text{LIF} > \text{LIF without reset}. Simply deleting the reset mechanism from vanilla LIF reduces accuracy by 2.00%2.00\% on CIFAR-10 and 2.12%2.12\% on CIFAR-100. However, the generalized multi-order weighting in PSN, masked PSN, and sliding PSN surpasses all conventional spiking neurons. Moreover, removing reset does not cause uninterrupted spiking: layerwise firing rates remain controlled between 0.050.05 and 0.350.35.

  8. Knowl 8 — Effect of Temporal Window Order on Sliding and Masked PSN

    empirical result

    Evaluating sequential CIFAR-100 accuracy across varying dependency orders k∈[1,32]k \in [1, 32] for T=32T=32 inputs shows:

    • Performance increases sharply as order kk expands from 11 to 44 (accuracy improves from ∼51%\sim 51\% to over 59%59\%).
    • Increasing kk beyond 4 yields marginal improvements, peaking at k=20k=20 for masked PSN and k=31k=31 for sliding PSN (62.11%62.11\%).
    • The sliding PSN accuracy curve stays consistently above that of masked PSN and exceeds the full T=32T=32 PSN model at k=21k=21 and k=31k=31. This demonstrates that parameter sharing across time-steps introduces an inductive bias that improves generalization on temporal sequence processing compared to time-varying weight matrices.
  9. Knowl 9 — Static and Neuromorphic Dataset Classification Benchmarks

    data/table

    The PSN family was tested on static image recognition (CIFAR-10, ImageNet) and neuromorphic event-stream classification (CIFAR10-DVS).

    Dataset Method / Neuron Architecture Time-steps (TT) Accuracy (%)
    CIFAR-10 Dspike Modified ResNet-18 6 94.25
    TET ResNet-19 6 94.50
    TDBN ResNet-19 6 93.16
    TEBN ResNet-19 6 94.71
    PLIF PLIF Net 8 93.50
    KLIF Modified PLIF Net 10 92.52
    GLIF ResNet-19 6 95.03
    PSN Modified PLIF Net 4 95.32
    ImageNet Dspike ResNet-34 6 68.19
    Dspike VGG-16 5 71.24
    TET SEW ResNet-34 4 68.00
    TDBN ResNet-34 (double channels) 6 67.05
    TEBN SEW ResNet-34 4 68.28
    GLIF ResNet-34 4 67.52
    SEW ResNet SEW ResNet-18 4 63.18
    SEW ResNet SEW ResNet-34 4 67.04
    PSN SEW ResNet-18 4 67.63
    PSN SEW ResNet-34 4 70.54
    CIFAR10-DVS Dspike ResNet-18 10 75.40
    TET VGG 10 83.17
    TDBN ResNet-19 10 67.80
    TEBN VGG 10 84.90
    PLIF PLIF Net 20 74.80
    KLIF Modified PLIF Net 15 70.90
    GLIF Wide 7B Net 16 78.10
    SEW ResNet Wide 7B Net 16 74.40
    Sliding PSN (k=2k=2) VGG 4 82.30
    Sliding PSN (k=2k=2) VGG 8 85.30
    Sliding PSN (k=2k=2) VGG 10 85.90

    On CIFAR-10, PSN achieves 95.32%95.32\% with only T=4T=4. On ImageNet, PSN achieves 70.54%70.54\% on SEW ResNet-34, improving by >3.5%>3.5\% over the IF neuron baseline (67.04%67.04\%). On CIFAR10-DVS, the 2-order sliding PSN achieves 82.30%82.30\% at T=4T=4, 85.30%85.30\% at T=8T=8, and 85.90%85.90\% at T=10T=10.

Coverage note — Specific hyperparameter configurations and training implementation details deferred to the supplementary materials in the paper are omitted.

References

  1. 1.Wolfgang Maass. Networks of spiking neurons: the third generation of neural network models. Neural Networks, 10(9):1659–1671, 1997.
  2. 2.Paul A Merolla, John V Arthur, Rodrigo Alvarez-Icaza, Andrew S Cassidy, Jun Sawada, Filipp Akopyan, Bryan L Jackson, Nabil Imam, Chen Guo, Yutaka Nakamura, et al. A million spiking-neuron integrated circuit with a scalable communication network and interface. Science, 345(6197):668–673, 2014.
  3. 3.Mike Davies, Narayan Srinivasa, Tsung-Han Lin, Gautham Chinya, Yongqiang Cao, Sri Harsha Choday, Georgios Dimou, Prasad Joshi, Nabil Imam, Shweta Jain, Yuyun Liao, Chit-Kwan Lin, Andrew Lines, Ruokun Liu, Deepak Mathaikutty, Steven McCoy, Arnab Paul, Jonathan Tse, Guruguhanathan Venkataramanan, Yi-Hsin Weng, Andreas Wild, Yoonseok Yang, and Hong Wang. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro, 38(1):82–99, 2018.
  4. 4.Jing Pei, Lei Deng, Sen Song, Mingguo Zhao, Youhui Zhang, Shuang Wu, Guanrui Wang, Zhe Zou, Zhenzhi Wu, Wei He, et al. Towards artificial general intelligence with hybrid tianjic chip architecture. Nature, 572(7767):106–111, 2019.
  5. 5.Chris Eliasmith, Terrence C. Stewart, Xuan Choo, Trevor Bekolay, Travis DeWolf, Yichuan Tang, and Daniel Rasmussen. A large-scale model of the functioning brain. Science, 338(6111):1202–1205, 2012.
  6. 6.Marcel Stimberg, Romain Brette, and Dan FM Goodman. Brian 2, an intuitive and efficient neural simulator. eLife, 8:e47314, 2019.
  7. 7.Emre O Neftci, Hesham Mostafa, and Friedemann Zenke. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6):51–63, 2019.
  8. 8.Amirhossein Tavanaei, Masoud Ghodrati, Saeed Reza Kheradpisheh, Timothée Masquelier, and Anthony Maida. Deep learning in spiking neural networks. Neural Networks, 111:47–63, 2019.
  9. 9.Yufei Guo, Xuhui Huang, and Zhe Ma. Direct learning-based deep spiking neural networks: a review. Frontiers in Neuroscience, 17:1209795, 2023.
  10. 10.Yongqiang Cao, Yang Chen, and Deepak Khosla. Spiking deep convolutional neural networks for energy-efficient object recognition. International Journal of Computer Vision, 113(1):54–66, 2015.
  11. 11.Bodo Rueckauer, Iulia-Alexandra Lungu, Yuhuang Hu, Michael Pfeiffer, and Shih-Chii Liu. Conversion of continuous-valued deep networks to efficient event-driven networks for image classification. Frontiers in Neuroscience, 11:682, 2017.
  12. 12.Saeed Reza Kheradpisheh and Timothée Masquelier. Temporal backpropagation for spiking neural networks with one spike per neuron. International Journal of Neural Systems, 30(06):2050027, 2020.
  13. 13.Kaushik Roy, Akhilesh Jaiswal, and Priyadarshini Panda. Towards spike-based machine intelligence with neuromorphic computing. Nature, 575(7784):607–617, 2019.
  14. 14.Rui Yuan, Qingxi Duan, Pek Jun Tiw, Ge Li, Zhuojian Xiao, Zhaokun Jing, Ke Yang, Chang Liu, Chen Ge, Ru Huang, et al. A calibratable sensory neuron based on epitaxial vo2 for spike-based neuromorphic multisensory system. Nature Communications, 13(1):1–12, 2022.
  15. 15.Wei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier, Tiejun Huang, and Yonghong Tian. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 2661–2671, 2021.
  16. 16.Xingting Yao, Fanrong Li, Zitao Mo, and Jian Cheng. GLIF: A unified gated leaky integrate-and-fire neuron for spiking neural networks. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems (NeurIPS), 2022.
  17. 17.Ilyass Hammouamri, Timothée Masquelier, and Dennis George Wilson. Mitigating catastrophic forgetting in spiking neural networks through threshold modulation. Transactions on Machine Learning Research, 2022.
  18. 18.Chunming Jiang and Yilei Zhang. Klif: An optimized spiking neuron unit for tuning surrogate gradient slope and membrane potential, 2023.
  19. 19.Lang Feng, Qianhui Liu, Huajin Tang, De Ma, and Gang Pan. Multi-level firing with spiking ds-resnet: Enabling better and deeper directly-trained spiking neural networks. In Lud De Raedt, editor, International Joint Conferences on Artificial Intelligence Organization (IJCAI), pages 2471–2477, 7 2022. Main Track.
  20. 20.Wachirawit Ponghiran and Kaushik Roy. Spiking neural networks with improved inherent recurrence dynamics for sequential learning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 36, pages 8001–8008, 2022.
  21. 21.Yufei Guo, Xinyi Tong, Yuanpei Chen, Liwen Zhang, Xiaode Liu, Zhe Ma, and Xuhui Huang. Recdis-snn: Rectifying membrane potential distribution for directly training spiking neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 326–335, June 2022.
  22. 22.Bing Han, Gopalakrishnan Srinivasan, and Kaushik Roy. Rmp-snn: Residual membrane potential neuron for enabling deeper high-accuracy and low-latency spiking neural network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13558–13567, 2020.
  23. 23.Jianhao Ding, Zhaofei Yu, Yonghong Tian, and Tiejun Huang. Optimal ann-snn conversion for fast and accurate inference in deep spiking neural networks. In Zhi-Hua Zhou, editor, Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 2328–2336. International Joint Conferences on Artificial Intelligence Organization, 8 2021. Main Track.
  24. 24.Tong Bu, Wei Fang, Jianhao Ding, PengLin Dai, Zhaofei Yu, and Tiejun Huang. Optimal ann-snn conversion for high-accuracy and ultra-low-latency spiking neural networks. In International Conference on Learning Representations (ICLR), 2021.
  25. 25.Shikuang Deng and Shi Gu. Optimal conversion of conventional artificial neural networks to spiking neural networks. In International Conference on Learning Representations (ICLR), 2021.
  26. 26.Zecheng Hao, Tong Bu, Jianhao Ding, Tiejun Huang, and Zhaofei Yu. Reducing ann-snn conversion error through residual membrane potential. Proceedings of the AAAI Conference on Artificial Intelligence, 37(1):11–21, Jun. 2023.
  27. 27.Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, and Luping Shi. Spatio-temporal backpropagation for training high-performance spiking neural networks. Frontiers in Neuroscience, 12, 2018.
  28. 28.Sumit Bam Shrestha and Garrick Orchard. Slayer: Spike layer error reassignment in time. In Advances in Neural Information Processing Systems (NeurIPS), pages 1419–1428, 2018.
  29. 29.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  30. 30.Hanle Zheng, Yujie Wu, Lei Deng, Yifan Hu, and Guoqi Li. Going deeper with directly-trained larger spiking neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 35, pages 11062–11070, 2021.
  31. 31.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), pages 448–456. PMLR, 2015.
  32. 32.Youngeun Kim and Priyadarshini Panda. Revisiting batch normalization for training low-latency deep spiking neural networks from scratch. Frontiers in Neuroscience, 15, 2021.
  33. 33.Chaoteng Duan, Jianhao Ding, Shiyan Chen, Zhaofei Yu, and Tiejun Huang. Temporal effective batch normalization in spiking neural networks. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems (NeurIPS), 2022.
  34. 34.Jacques Kaiser, Hesham Mostafa, and Emre Neftci. Synaptic plasticity dynamics for deep continuous local learning (decolle). Frontiers in Neuroscience, 14, 2020.
  35. 35.Mingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He, and Zhouchen Lin. Online training through time for spiking neural networks. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems (NeurIPS), 2022.
  36. 36.Nicolas Perez-Nieves and Dan F. M. Goodman. Sparse spiking gradient descent. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems (NeurIPS), 2021.
  37. 37.Man Yao, Guangshe Zhao, Hengyu Zhang, Yifan Hu, Lei Deng, Yonghong Tian, Bo Xu, and Guoqi Li. Attention spiking neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–18, 2023.
  38. 38.Yuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng, Yongqing Hai, and Shi Gu. Differentiable spike: Rethinking gradient-descent for training spiking neural networks. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems (NeurIPS), 2021.
  39. 39.Shikuang Deng, Yuhang Li, Shanghang Zhang, and Shi Gu. Temporal efficient training of spiking neural network via gradient re-weighting. In International Conference on Learning Representations (ICLR), 2022.
  40. 40.Wei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang, Timothée Masquelier, and Yonghong Tian. Deep residual learning in spiking neural networks. Advances in Neural Information Processing Systems (NeurIPS), 34, 2021.
  41. 41.Zhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang, Shuicheng YAN, Yonghong Tian, and Li Yuan. Spikformer: When spiking neural network meets transformer. In International Conference on Learning Representations (ICLR), 2023.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017.
  43. 43.Rajat Raina, Anand Madhavan, and Andrew Y Ng. Large-scale deep unsupervised learning using graphics processors. In International Conference on Machine Learning (ICML), pages 873–880, 2009.
  44. 44.Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. Neural machine translation in linear time. arXiv preprint arXiv:1610.10099, 2016.
  45. 45.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. Convolutional sequence to sequence learning. In International Conference on Machine Learning (ICML), pages 1243–1252. PMLR, 2017.
  46. 46.Eric Martin and Chris Cundy. Parallelizing linear recurrent neural nets over sequence length. In International Conference on Learning Representations (ICLR), 2018.
  47. 47.Mark Harris, Shubhabrata Sengupta, and John D Owens. Parallel prefix sum (scan) with cuda. GPU gems, 3(39):851–876, 2007.
  48. 48.Eimantas Ledinauskas, Julius Ruseckas, Alfonsas Juršėnas, and Giedrius Buračas. Training Deep Spiking Neural Networks. arXiv preprint arXiv:2006.04436, 2020.
  49. 49.Seongsik Park, Seijoon Kim, Byunggook Na, and Sungroh Yoon. T2fsnn: deep spiking neural networks with time-to-first-spike coding. In 2020 57th ACM/IEEE Design Automation Conference, pages 1–6. IEEE, 2020.
  50. 50.Dongwoo Lew, Kyungchul Lee, and Jongsun Park. A time-to-first-spike coding and conversion aware training for energy-efficient deep spiking neural network processor design. In Proceedings of the 59th ACM/IEEE Design Automation Conference, pages 265–270, 2022.
  51. 51.Lina Bonilla, Jacques Gautrais, Simon Thorpe, and Timothée Masquelier. Analyzing time-to-first-spike coding schemes. Frontiers in Neuroscience, 16, 2022.
  52. 52.Wei Fang, Yanqi Chen, Jianhao Ding, Zhaofei Yu, Timothée Masquelier, Ding Chen, Liwei Huang, Huihui Zhou, Guoqi Li, and Yonghong Tian. Spikingjelly: An open-source machine learning infrastructure platform for spike-based intelligence. Science Advances, 9(40):eadi1480, 2023.
  53. 53.Friedemann Zenke and Tim P Vogels. The remarkable robustness of surrogate gradient learning for instilling complex function in spiking neural networks. BioRxiv, 2020.
  54. 54.Hongmin Li, Hanchao Liu, Xiangyang Ji, Guoqi Li, and Luping Shi. Cifar10-dvs: An event-stream dataset for object classification. Frontiers in Neuroscience, 11, 2017.
  55. 55.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015.
  56. 56.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022.
  57. 57.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  58. 58.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  59. 59.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 6023–6032, 2019.
  60. 60.Samuel G Müller and Frank Hutter. Trivialaugment: Tuning-free yet state-of-the-art data augmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 774–782, 2021.
  61. 61.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence (AAAI), volume 34, pages 13001–13008, 2020.
  62. 62.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, 2016.
  63. 63.Qingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang, Zhouchen Lin, and Zhi-Quan Luo. Training high-performance low-latency spiking neural networks by differentiation on spike representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12444–12453, 2022.
  64. 64.Nitin Rathi and Kaushik Roy. Diet-snn: A low-latency spiking neural network with direct input encoding and leakage and threshold optimization. IEEE Transactions on Neural Networks and Learning Systems, 2021.
  65. 65.Nitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, and Kaushik Roy. Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation. In International Conference on Learning Representations (ICLR), 2020.

Citation

MLA
Fang, W., et al. “Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 53674–87, https://proceedings.neurips.cc/paper_files/paper/2023/file/a834ac3dfdb90da54292c2c932c997cc-Paper-Conference.pdf.
APA
Fang, W., Yu, Z., Zhou, Z., Chen, D., Chen, Y., Ma, Z., Masquelier, T., & Tian, Y. (2023). Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies. Advances in Neural Information Processing Systems, 36, 53674–53687. https://proceedings.neurips.cc/paper_files/paper/2023/file/a834ac3dfdb90da54292c2c932c997cc-Paper-Conference.pdf
Chicago
Fang, W., Z. Yu, Z. Zhou, et al. 2023. “Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies”. Advances in Neural Information Processing Systems 36: 53674–87. https://proceedings.neurips.cc/paper_files/paper/2023/file/a834ac3dfdb90da54292c2c932c997cc-Paper-Conference.pdf.
Harvard
Fang, W. et al. (2023) “Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 53674–53687. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/a834ac3dfdb90da54292c2c932c997cc-Paper-Conference.pdf.
Vancouver
1. Fang W, Yu Z, Zhou Z, Chen D, Chen Y, Ma Z, Masquelier T, Tian Y (2023) Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 53674–53687

BibTeX

@inproceedings{fang2023parallel,
  title = {Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies},
  author = {Fang, Wei and Yu, Zhaofei and Zhou, Zhaokun and Chen, Ding and Chen, Yanqi and Ma, Zhengyu and Masquelier, Timothée and Tian, Yonghong},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {53674-53687},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/a834ac3dfdb90da54292c2c932c997cc-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors