Adaptive Smoothing Gradient Learning for Spiking Neural Networks

Ziming WangRunhao JiangShuang LianRui YanHuajin Tang

article2023ICML67 citations

Proposes an adaptive smoothing gradient learning method that eliminates gradient mismatch in spiking neural networks by incorporating learnable relaxation factors with random spike noise, achieving state-of-the-art accuracy across vision and audio tasks.

Listen

Artificial intelligence applications running on edge devices face severe energy constraints. Spiking neural networks offer an appealing solution because they process information using sparse binary pulses, drastically lowering energy consumption on specialized hardware. However, training these networks is difficult because binary pulses create a mathematical discontinuity that prevents standard optimization algorithms from computing exact gradients. While current methods approximate these gradients using fixed smoothing curves, this approximation creates a persistent mismatch between estimated gradients and actual network behavior, leading to unstable training, degraded accuracy, and severe sensitivity to manual tuning.

The article demonstrates a novel training framework called adaptive smoothing gradient learning to eliminate this gradient mismatch and automate hyperparameter tuning. The method introduces a dual-mode training approach where standard continuous activations are injected with random binary pulse noise. By isolating pulse activations from backward error calculation and making the smoothing factor a learnable parameter, the network naturally adapts its gradient estimates during training. Over successive optimization steps, the hybrid architecture progressively converges into a pure, energy-efficient spiking neural network without requiring manual intervention.

The authors validated this approach across diverse benchmark datasets spanning static images, neuromorphic event streams, human speech, and musical instruments. Key findings show that the proposed method consistently achieves state-of-the-art performance across all tested domains. On static image recognition benchmarks, the model reached 95.35% accuracy on CIFAR-10 and 77.74% on CIFAR-100 within only four time steps, while consuming only 8.96% of the energy required by an equivalent non-spiking network. On dynamic sensory tasks, it achieved 84.50% accuracy on DVS-CIFAR10 and 97.90% on gesture recognition. Crucially, the approach eliminated sensitivity to initial smoothness settings: while conventional methods suffered catastrophic performance drops of up to 60 percentage points when smoothness factors were misconfigured, the adaptive method maintained stable, high accuracy.

These results demonstrate that organizations can achieve deep learning accuracy at a fraction of the computational power cost, reducing operational expenses and latency for edge intelligence without incurring expensive hyperparameter search costs. Decision-makers should consider piloting this adaptive training methodology for battery-powered, real-time, or edge-deployed vision and audio systems. Future work should focus on validating this technique on larger-scale foundational models and testing its deployment across diverse physical neuromorphic hardware architectures.

Wang et al (2023).pdf
Cover for Adaptive Smoothing Gradient Learning for Spiking Neural Networks

Abstract

Spiking neural networks (SNNs) with biologically inspired spatio-temporal dynamics demonstrate superior energy efficiency on neuromorphic architectures. Error backpropagation in SNNs is prohibited by the all-or-none nature of spikes. The existing solution circumvents this problem by a relaxation on the gradient calculation using a continuous function with a constant relaxation degree, so-called surrogate gradient learning. Nevertheless, such a solution introduces additional smoothing error on spike firing which leads to the gradients being estimated inaccurately. Thus, how to adaptively adjust the relaxation degree and eliminate smoothing error progressively is crucial. Here, we propose a methodology such that training a prototype neural network will evolve into training an SNN gradually by fusing the learnable relaxation degree into the network with random spike noise. In this way, the network learns adaptively the accurate gradients of loss landscape in SNNs. The theoretical analysis further shows optimization on such a noisy network could be evolved into optimization on the embedded SNN with shared weights progressively. Moreover, the experiments on static images, dynamic event streams, speech, and instrumental sounds show the proposed method achieves state-of-the-art performance across all the datasets with remarkable robustness on different relaxation degrees.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Preliminary
  • 4. Method
  • 4.1. Spike-based Backpropagation
  • 4.2. Design of ASGL
  • 4.3. Theoretical Analysis
  • 5. Experiments
  • 5.1. Ablation Study
  • 5.2. Performance on Static Images
  • 5.3. Performance on Spatio-temporal Patterns.
  • 5.4. Effect of Noise Probability
  • 5.5. Network Evolution
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Detailed Derivation and Neuron Dynamics
  • A.1. Spike-based Backpropagation incorporated with Adaptive Width
  • A.2. Theory Analysis for Error Decomposition
  • A.3. C-LIF Model
  • A.4. Reccurent Connections
  • B.3. Ablation Study on Image Reconstruction
  • B.4. Experiments on MedlyDB, DVS128 Gesture
  • B.5. Ablation Study on ResNet-19
  • B.6. Width Update
  • B.7. Effect of fixed α
  • B.8. Effect of Adaptive α ( α ↛ 0 )
  • B.9. Results without Auto-Augment and Cutout

Knowls

  1. Knowl 1 — Dual-Mode Forward Activation with Detached Spike Noise Injection

    model/method

    Adaptive Smoothing Gradient Learning (ASGL) resolves gradient mismatching in Spiking Neural Networks (SNNs) by replacing the non-differentiable Heaviside step function Θ(x)\Theta(x) during training with a hybrid activation H^α(x)\hat{H}_\alpha(x) that couples continuous analog activation Hα(x)H_\alpha(x) with binary spike noise. Here, Hα(x)=∫−∞xhα(u) duH_\alpha(x) = \int_{-\infty}^x h_\alpha(u) \, du is the antiderivative of the chosen surrogate gradient function hα(x)h_\alpha(x). For the rectangular surrogate gradient hα(x)=1αsign⁡(∣x∣<α2)h_\alpha(x) = \frac{1}{\alpha} \operatorname{sign}(|x| < \frac{\alpha}{2}), the corresponding continuous activation is the shifted clipping function:

    Hα(x)=clip⁡(1αx+12,0,1)H_\alpha(x) = \operatorname{clip}\left(\frac{1}{\alpha}x + \frac{1}{2}, 0, 1\right)

    During training forward propagation, given the normalized membrane potential difference u^l[t]=ul[t]−ϑ\hat{u}^l[t] = u^l[t] - \vartheta (where ul[t]u^l[t] is the membrane potential at layer ll and time step tt, and ϑ\vartheta is the firing threshold), the layer activation H^α(u^l[t])\hat{H}_\alpha(\hat{u}^l[t]) is computed as:

    H^α(u^l[t])=Hα(u^l[t])+m⊙Φ(Θ(u^l[t])−Hα(u^l[t]))\hat{H}_\alpha(\hat{u}^l[t]) = H_\alpha(\hat{u}^l[t]) + m \odot \Phi\left(\Theta(\hat{u}^l[t]) - H_\alpha(\hat{u}^l[t])\right)

    where m∼Bernoulli⁡(p)m \sim \operatorname{Bernoulli}(p) is an independent binary random mask with spike noise probability p∈(0,1]p \in (0, 1], ⊙\odot represents element-wise multiplication, and Φ\Phi is an identity mapping in forward propagation that detaches gradients during backward propagation (such that ∂Φ(x)∂x=0\frac{\partial \Phi(x)}{\partial x} = 0). By detaching the gradient of the binary spikes, backpropagation flows entirely through the continuous function Hα(x)H_\alpha(x), eliminating surrogate gradient mismatching while exposing the network to spike-induced perturbations. During validation/inference, the deterministic binary activation sl[t]=Θ(u^l[t])s^l[t] = \Theta(\hat{u}^l[t]) is used exclusively.

  2. Knowl 2 — ASGL Dual-Mode Spike Activation Procedure

    algorithm

    The core activation algorithm in ASGL conditionally switches between stochastic noise injection during training and discrete spike firing during inference.

    Input: Normalized membrane potential u^l[t]=ul[t]−ϑ\hat{u}^l[t] = u^l[t] - \vartheta
    Input: Mode indicator TT (T=trueT = \text{true} for training, T=falseT = \text{false} for validation)
    Input: Learnable width parameter α\alpha, noise probability pp
    Output: Output activation sl[t]s^l[t]
    if TT is true then
        Sample random mask m∼Bernoulli⁡(p)m \sim \operatorname{Bernoulli}(p)
        Compute continuous activation Hα(u^l[t])=clip⁡(1αu^l[t]+12,0,1)H_\alpha(\hat{u}^l[t]) = \operatorname{clip}(\frac{1}{\alpha}\hat{u}^l[t] + \frac{1}{2}, 0, 1)
        Compute binary spike Θ(u^l[t])\Theta(\hat{u}^l[t])
        sl[t]=Hα(u^l[t])+m⊙Φ(Θ(u^l[t])−Hα(u^l[t]))s^l[t] = H_\alpha(\hat{u}^l[t]) + m \odot \Phi(\Theta(\hat{u}^l[t]) - H_\alpha(\hat{u}^l[t]))
    else
        sl[t]=Θ(u^l[t])s^l[t] = \Theta(\hat{u}^l[t])
    end if
    return sl[t]s^l[t]

    Here, Φ\Phi detaches gradients from the spike discrepancy Θ(u^l[t])−Hα(u^l[t])\Theta(\hat{u}^l[t]) - H_\alpha(\hat{u}^l[t]), allowing exact gradient updates with respect to HαH_\alpha and the learnable width α\alpha across spatial and temporal dimensions.

  3. Knowl 3 — Spatial-Temporal Gradient Formulation for Learnable Smoothing Width Parameter

    equation

    In ASGL, the surrogate function width parameter αl\alpha^l at layer ll is treated as a learnable parameter optimized end-to-end via spatio-temporal backpropagation. For a network with NN discrete time steps, target label yy, predicted softmax output y^=softmax⁡(cˉL)\hat{y} = \operatorname{softmax}(\bar{c}^L) where cˉL=1N∑t=1NcL[t]\bar{c}^L = \frac{1}{N} \sum_{t=1}^N c^L[t] is the time-averaged postsynaptic current at the final layer LL, the loss gradient with respect to αl\alpha^l is:

    ∇αl=−yT−y^TN∑t∗=1N∂cL[t∗]∂αl\nabla \alpha^l = -\frac{y^T - \hat{y}^T}{N} \sum_{t^*=1}^N \frac{\partial c^L[t^*]}{\partial \alpha^l}

    where for any target time step t∗t^*, the temporal partial derivative is decomposed as:

    ∂cL[t∗]∂αl=∑t=1t∗(δl+1[t]Wl+1−γδl[t+1]diag⁡(ul[t]))∂sl[t]∂αl[t]\frac{\partial c^L[t^*]}{\partial \alpha^l} = \sum_{t=1}^{t^*} \left( \delta^{l+1}[t] W^{l+1} - \gamma \delta^l[t+1] \operatorname{diag}(u^l[t]) \right) \frac{\partial s^l[t]}{\partial \alpha^l[t]}

    In this expression, δl[t]=∂cL[t∗]∂ul[t]\delta^l[t] = \frac{\partial c^L[t^*]}{\partial u^l[t]} is the credit assigned to the membrane potential ul[t]u^l[t], Wl+1W^{l+1} is the forward weight matrix to layer l+1l+1, γ=1−1/τm\gamma = 1 - 1/\tau_m is the membrane leak factor, and the derivative of the shifted clipping activation HαH_\alpha with respect to α\alpha is:

    ∂sl[t]∂αl[t]=∂Hα(u^l[t])∂α={0,if ∣u^l[t]∣>12α−diag⁡(1α2⊙u^l[t]),otherwise\frac{\partial s^l[t]}{\partial \alpha^l[t]} = \frac{\partial H_\alpha(\hat{u}^l[t])}{\partial \alpha} = \begin{cases} 0, & \text{if } |\hat{u}^l[t]| > \frac{1}{2}\alpha \\ -\operatorname{diag}\left(\frac{1}{\alpha^2} \odot \hat{u}^l[t]\right), & \text{otherwise} \end{cases}

  4. Knowl 4 — Equivalence of ASGL Optimization to Regularized SNN Loss Minimization

    theoretical result

    Let FnoiseF_{\text{noise}} denote the training network utilizing mixed analog activations Hα(x)H_\alpha(x) with random binary spike masking, and let FsnnF_{\text{snn}} denote the target embedded SNN using purely discrete spike activations Θ(x)\Theta(x). Let ℓnoise(F,s)≜Em^[ℓ(Fnoise(s))]=Em^[ℓ(Fsnn(s,m^))]\ell_{\text{noise}}(F, s) \triangleq \mathbb{E}_{\hat{m}}[\ell(F_{\text{noise}}(s))] = \mathbb{E}_{\hat{m}}[\ell(F_{\text{snn}}(s, \hat{m}))] represent the expected loss of the hybrid network over the normalized random mask m^=m/p\hat{m} = m/p, where pp is the spike noise probability.

    Via second-order Taylor expansion around the perturbation Δ=(1−m^l)⊙(Hα(u^l)−Θ(u^l))\Delta = (1 - \hat{m}^l) \odot (H_\alpha(\hat{u}^l) - \Theta(\hat{u}^l)), the expected noisy loss ℓnoise(F,s)\ell_{\text{noise}}(F, s) is approximated by the target SNN loss ℓsnn(F,s)\ell_{\text{snn}}(F, s) regularized by the layer-wise distance between continuous activations Hα(u^l)H_\alpha(\hat{u}^l) and binary activations Θ(u^l)\Theta(\hat{u}^l):

    ℓnoise(F,s)≈ℓsnn(F,s)+1−p2p∑l=1L⟨Cl,diag⁡(Hα(u^l)−Θ(u^l))⊙2⟩\ell_{\text{noise}}(F, s) \approx \ell_{\text{snn}}(F, s) + \frac{1-p}{2p} \sum_{l=1}^L \left\langle C^l, \operatorname{diag}\left(H_\alpha(\hat{u}^l) - \Theta(\hat{u}^l)\right)^{\odot 2} \right\rangle

    where Cl=D2(ℓ∘Em^[Gl])[sl]C^l = D^2(\ell \circ \mathbb{E}_{\hat{m}}[G^l])[s^l] represents the Hessian (second derivative) of the loss function ℓ\ell with respect to the ll-th layer spike activation sls^l through the subnetwork GlG^l (which employs mixed activations after layer ll and fully discrete spikes in the preceding ll layers). Under alternating optimization of weights WW and learnable width α\alpha, minimizing ℓnoise\ell_{\text{noise}} dynamically contracts the activation discrepancy ∥Hα(u^l)−Θ(u^l)∥2\|H_\alpha(\hat{u}^l) - \Theta(\hat{u}^l)\|^2, driving the dynamics of the hybrid network to converge into those of the embedded SNN.

  5. Knowl 5 — Robustness of ASGL to Smoothing Width Initialization

    data/table

    Standard surrogate gradient (SG) learning exhibits extreme sensitivity to the surrogate width parameter α\alpha, where poor choices cause severe gradient vanishing (dead neurons) or excessive smoothing errors. In contrast, ASGL incorporates a learnable α\alpha, making network accuracy robust across a wide range of initial widths α\alpha.

    Width (α\alpha) CIFAR-10 (ResNet-19, N=3N=3) CIFAR-100 (ResNet-18, N=3N=3)
    SG Acc. (%) ASGL Acc. (%) SG Acc. (%) ASGL Acc. (%)
    0.5 93.19 94.11 75.76 76.54
    1.0 93.78 94.30 65.19 76.09
    2.5 90.68 94.09 15.12 76.18
    5.0 62.34 93.61 8.04 76.68
    10.0 30.85 93.53 6.14 76.00

    While standard SG performance collapses completely on CIFAR-100 when α\alpha increases from 0.50.5 to 10.010.0 (falling from 75.76%75.76\% to 6.14%6.14\%, a 69.62%69.62\% drop), ASGL maintains stable performance across all initializations (76.00%76.00\% to 76.68%76.68\%).

  6. Knowl 6 — Direct Training Performance on Static Image Benchmarks

    data/table

    ASGL achieves state-of-the-art accuracy when directly training spiking neural networks on CIFAR-10, CIFAR-100, and Tiny-ImageNet under ultra-low latency (N=2,4,8N = 2, 4, 8 time steps) without requiring pretrained ANN initialization or Time Inheritance Training.

    Dataset Method Architecture Time steps (NN) Accuracy (%)
    CIFAR-10 TET ResNet-19 4 94.44±0.0894.44 \pm 0.08
    TET ResNet-19 2 94.16±0.0394.16 \pm 0.03
    SpikeDHS SpikeDHS-CLA (n4s1) 6 94.68±0.0594.68 \pm 0.05
    GLIF ResNet-18 4 94.67±0.0594.67 \pm 0.05
    GLIF ResNet-18 2 94.15±0.0494.15 \pm 0.04
    ASGL (Ours) CifarNet 4 94.74±0.1094.74 \pm 0.10
    ASGL (Ours) CifarNet 2 93.80±0.1193.80 \pm 0.11
    ASGL (Ours) ResNet-18 4 95.35±0.25\mathbf{95.35 \pm 0.25}
    ASGL (Ours) ResNet-18 2 95.27±0.06\mathbf{95.27 \pm 0.06}
    CIFAR-100 TET ResNet-19 4 74.47±0.1574.47 \pm 0.15
    TET ResNet-19 2 72.87±0.1072.87 \pm 0.10
    Dspike ResNet-18 4 73.35±0.1473.35 \pm 0.14
    SpikeDHS SpikeDHS-CLA (n4s1) 6 76.03±0.2076.03 \pm 0.20
    ASGL (Ours) CifarNet 4 74.59±0.0774.59 \pm 0.07
    ASGL (Ours) ResNet-18 4 77.74±0.07\mathbf{77.74 \pm 0.07}
    ASGL (Ours) ResNet-18 2 76.59±0.05\mathbf{76.59 \pm 0.05}
    Tiny-ImageNet Online LTL VGG-13 16 54.82
    Offline LTL VGG-13 16 55.37
    ASGL (Ours) VGG-13 8 56.81\mathbf{56.81}
    ASGL (Ours) VGG-13 4 56.57\mathbf{56.57}

    On CIFAR-10 ResNet-18, ASGL reaches 95.27%95.27\% accuracy with only 2 time steps. On CIFAR-100 ResNet-18, ASGL obtains 77.74%77.74\% accuracy at 4 time steps. On Tiny-ImageNet (200 classes, 64×6464\times 64 resolution), ASGL achieves 56.57%56.57\% at N=4N=4 and 56.81%56.81\% at N=8N=8, outperforming previous tandem learning approaches that required 16 time steps.

  7. Knowl 7 — Performance on Neuromorphic and Audio Spatio-Temporal Datasets

    data/table

    ASGL was evaluated across diverse spatio-temporal benchmarks including event-camera streams (DVS-CIFAR10, DVS128 Gesture) and audio/speech recognition (Spiking Heidelberg Digits SHD, MedleyDB music instruments).

    Dataset Method Architecture Time steps (NN) Metric / Acc. (%)
    DVS-CIFAR10 PLIF-SNN CifarNet-C 20 74.80
    Dspike ResNet-18 10 75.45
    TET VGGSNN 10 83.17±0.1583.17 \pm 0.15
    ASGL (Ours) VGGSNN 10 84.50±0.08\mathbf{84.50 \pm 0.08}
    DVS128 Gesture STBP-tdBN SNN (ResNet17) 60 96.87
    PLIF-SNN SNN (8 layers) 60 97.57
    PointNet-like ANN DNN 60 95.32
    RG-CNN DNN 60 97.20
    ASGL (Ours) SNN (8 layers) 60 97.90\mathbf{97.90}
    SHD SG (Cramer et al.) 700-240-20 SNN 250 83.2±1.383.2 \pm 1.3
    EventProp 700-240-20 SNN 250 84.8±1.584.8 \pm 1.5
    ASGL (Ours) 700-240-20 SNN 250 86.9±1.0\mathbf{86.9 \pm 1.0}
    MedleyDB STCA 384-700-10 SNN 500 F1: 97.25
    CNN (direct BP) CNN (76.9w) 500 F1: 97.51
    ASGL (Ours) 384-700-10 SNN 500 F1: 98.59

    On DVS-CIFAR10, ASGL achieves 84.50%84.50\%, improving over TET by 1.33%1.33\%. On DVS128 Gesture, ASGL reaches 97.90%97.90\%, exceeding specialized event DNNs. On SHD, ASGL reaches 86.9%86.9\%, outperforming exact-gradient EventProp by 2.1%2.1\%. On MedleyDB, ASGL achieves an F1-score of 98.59%98.59\%, outperforming both surrogate gradient SNNs and spectrogram CNNs.

  8. Knowl 8 — Tradeoff Between Dead Neurons and Smoothing Error in Surrogate Width Optimization

    empirical result

    Analyzing spiking neurons across the width parameter α\alpha reveals an inherent tension:

    1. Dead Neuron Proportion: As width α\alpha decreases toward zero, the derivative hα(x)h_\alpha(x) becomes non-zero over an increasingly narrow band. Neurons whose membrane potential falls outside [−α/2,α/2][-\alpha/2, \alpha/2] receive zero gradient, causing the proportion of inactive/dead neurons to rise sharply.
    2. Smoothing Error: Conversely, as α\alpha increases, the continuous surrogate activation deviates further from the true step function Θ(x)\Theta(x). Measuring the mismatch via the metric err⁡=1−cos⁡(al,sl)\operatorname{err} = 1 - \cos(a^l, s^l) (where ala^l and sls^l are analog and binary activations respectively) shows that smoothing error grows monotonically with larger α\alpha.

    Because the optimal trade-off point between gradient propagation (preventing dead neurons) and gradient accuracy (minimizing smoothing error) varies across layers and training epochs, fixed-width surrogate gradients fail. Learnable α\alpha in ASGL resolves this trade-off adaptively: α\alpha does not simply shrink to zero during training, but stabilizes at layer-specific non-zero values (e.g., higher α\alpha in final layers and lower α\alpha in intermediate layers), achieving over 9%9\% higher cosine similarity between hybrid and target SNN activations than training with a fixed width.

  9. Knowl 9 — Energy Consumption and Synaptic Operations Efficiency of ASGL SNNs

    empirical result

    The computational overhead of SNNs is evaluated by counting total Synaptic Operations (SOP) across all TT time steps and L−1L-1 synaptic layers:

    NAC=∑t=1T∑l=1L−1∑i=1Nlfilsil[t]N_{\text{AC}} = \sum_{t=1}^T \sum_{l=1}^{L-1} \sum_{i=1}^{N_l} f_i^l s_i^l[t]

    where filf_i^l is the fan-out (outgoing connection count) of neuron ii at layer ll, NlN_l is the number of neurons in layer ll, and sil[t]∈{0,1}s_i^l[t] \in \{0, 1\} is the binary spike event. In contrast, an equivalent ANN requires fixed Multiply-Accumulate (MAC) operations:

    NMAC=∑l=1L−1∑i=1NlfilN_{\text{MAC}} = \sum_{l=1}^{L-1} \sum_{i=1}^{N_l} f_i^l

    Using standard 45nm CMOS hardware energy estimates (32-bit floating-point AC at 0.9 pJ0.9\text{ pJ} per operation and 32-bit floating-point MAC at 4.6 pJ4.6\text{ pJ} per operation, with MAC used for the direct-current first layer in SNNs), a spiking ResNet-18 trained with ASGL under 2 time steps achieves 94.11%94.11\% accuracy on CIFAR-10 while consuming only 8.96%8.96\% of the energy of an equivalent ANN architecture.

  10. Knowl 10 — Effective Soft-Reset Dynamics Induced by Dual-Mode Forwarding

    model/method

    In standard LIF neurons with hard reset, the membrane potential resets to zero when a spike is emitted: ul[t]=γul[t−1]⊙(1−sl[t−1])+cl[t]u^l[t] = \gamma u^l[t-1] \odot (1 - s^l[t-1]) + c^l[t]. During ASGL training, replacing the discrete spike sl[t−1]=Θ(u^l[t−1])s^l[t-1] = \Theta(\hat{u}^l[t-1]) with the hybrid activation H^α(u^l[t−1])\hat{H}_\alpha(\hat{u}^l[t-1]) naturally transforms the hard reset into an expected soft reset:

    E[ul[t]]=γul[t−1]⊙(1−Hα(u^l[t−1]))+cl[t]\mathbb{E}[u^l[t]] = \gamma u^l[t-1] \odot \left(1 - H_\alpha(\hat{u}^l[t-1])\right) + c^l[t]

    Under this formulation, Hα(u[t])H_\alpha(u[t]) acts as the expected firing probability within the discrete time step. Rather than forcing the post-spike voltage immediately to zero, residual membrane potential exceeding the threshold is retained with probability 1−Hα(u[t])1 - H_\alpha(u[t]), smoothly interpolating temporal reset behavior during training without modifying the deterministic hard reset used in deployment.

Coverage note — None was omitted; all key theoretical derivations, algorithmic components, static/spatio-temporal benchmark results, ablation studies, and architectural variants (LIF, C-LIF, Recurrent SNN) are comprehensively covered.

References

  1. 1.Akopyan, F., Sawada, J., Cassidy, A., Alvarez-Icaza, R., Arthur, J., Merolla, P., Imam, N., Nakamura, Y., Datta, P., Nam, G.-J., et al. Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip. IEEE transactions on computer-aided design of integrated circuits and systems, 34(10):1537–1557, 2015.
  2. 2.Amir, A., Taba, B., Berg, D., Melano, T., McKinstry, J., Di Nolfo, C., Nayak, T., Andreopoulos, A., Garreau, G., Mendoza, M., et al. A low power, fully event-based gesture recognition system. In CVPR, pp. 7243–7252, 2017.
  3. 3.Bi, Y., Chadha, A., Abbas, A., Bourtsoulatze, E., and Andreopoulos, Y. Graph-based spatio-temporal feature learning for neuromorphic vision sensing. IEEE Transactions on Image Processing, 29:9084–9098, 2020.
  4. 4.Bittner, R. M., Salamon, J., Tierney, M., Mauch, M., Cannam, C., and Bello, J. P. Medleydb: A multitrack dataset for annotation-intensive mir research. In ISMIR, volume 14, pp. 155–160, 2014.
  5. 5.Bohte, S. M., Kok, J. N., and La Poutre, H. Error-backpropagation in temporally encoded networks of spiking neurons. Neurocomputing, 48(1-4):17–37, 2002.
  6. 6.Bu, T., Fang, W., Ding, J., Dai, P., Yu, Z., and Huang, T. Optimal ann-snn conversion for high-accuracy and ultra-low-latency spiking neural networks. In International Conference on Learning Representations, 2021.
  7. 7.Cramer, B., Stradmann, Y., Schemmel, J., and Zenke, F. The heidelberg spiking data sets for the systematic evaluation of spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems, 2020.
  8. 8.Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. Autoaugment: Learning augmentation policies from data. arXiv preprint arXiv:1805.09501, 2018.
  9. 9.Davies, M., Srinivasa, N., Lin, T.-H., Chinya, G., Cao, Y., Choday, S. H., Dimou, G., Joshi, P., Imam, N., Jain, S., et al. Loihi: A neuromorphic manycore processor with on-chip learning. Ieee Micro, 38(1):82–99, 2018.
  10. 10.Deng, S. and Gu, S. Optimal conversion of conventional artificial neural networks to spiking neural networks. ICLR, 2021.
  11. 11.Deng, S., Li, Y., Zhang, S., and Gu, S. Temporal efficient training of spiking neural network via gradient reweighting. arXiv preprint arXiv:2202.11946, 2022.
  12. 12.DeVries, T. and Taylor, G. W. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
  13. 13.Fang, W., Yu, Z., Chen, Y., Huang, T., Masquelier, T., and Tian, Y. Deep residual learning in spiking neural networks. Advances in Neural Information Processing Systems, 34:21056–21069, 2021a.
  14. 14.Fang, W., Yu, Z., Chen, Y., Masquelier, T., Huang, T., and Tian, Y. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In ICCV, pp. 2661–2671, 2021b.
  15. 15.Garg, I., Chowdhury, S. S., and Roy, K. Dct-snn: Using dct to distribute spatial information over time for learning low-latency spiking neural networks. arXiv preprint arXiv:2010.01795, 2020.
  16. 16.Gu, P., Xiao, R., Pan, G., and Tang, H. Stca: Spatio-temporal credit assignment with delayed feedback in deep spiking neural networks. In IJCAI, pp. 1366–1372, 2019.
  17. 17.Guo, Y., Chen, Y., Zhang, L., Liu, X., Wang, Y., Huang, X., and Ma, Z. Im-loss: Information maximization loss for spiking neural networks. In Advances in Neural Information Processing Systems, 2022.
  18. 18.Gütig, R. Spiking neurons can discover predictive features by aggregate-label learning. Science, 351(6277), 2016.
  19. 19.Hagenaars, J., Paredes-Vallés, F., and De Croon, G. Self-supervised learning of event-based optical flow with spiking neural networks. Advances in Neural Information Processing Systems, 34:7167–7179, 2021.
  20. 20.Han, S., Pool, J., Tran, J., and Dally, W. J. Learning both weights and connections for efficient neural networks. CoRR, abs/1506.02626, 2015.
  21. 21.He, W., Wu, Y., Deng, L., Li, G., Wang, H., Tian, Y., Ding, W., Wang, W., and Xie, Y. Comparing snns and rnns on neuromorphic vision datasets: similarities and differences. Neural Networks, 132:108–120, 2020.
  22. 22.Kaiser, J., Mostafa, H., and Neftci, E. Synaptic plasticity dynamics for deep continuous local learning (decolle). Frontiers in Neuroscience, 14:424, 2020.
  23. 23.Kim, J., Kim, K., and Kim, J.-J. Unifying activation-and timing-based learning rules for spiking neural networks. NeurIPS, 33:19534–19544, 2020.
  24. 24.Kugele, A., Pfeil, T., Pfeiffer, M., and Chicca, E. Efficient processing of spatio-temporal data streams with spiking neural networks. Frontiers in Neuroscience, 14: 439, 2020.
  25. 25.Kundu, S., Datta, G., Pedram, M., and Beerel, P. A. Spike-thrift: Towards energy-efficient deep spiking neural networks by limiting spiking activity via attention-guided compression. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3953–3962, 2021.
  26. 26.Leng, L., Che, K., Zhang, K., Zhang, J., Meng, Q., Cheng, J., Guo, Q., and Liao, J. Differentiable hierarchical and surrogate gradient search for spiking neural networks. In Advances in Neural Information Processing Systems, 2022.
  27. 27.Lewicki, M. S. Efficient coding of natural sounds. Nature neuroscience, 5(4):356–363, 2002.
  28. 28.Li, H., Liu, H., Ji, X., Li, G., and Shi, L. Cifar10-dvs: an event-stream dataset for object classification. Frontiers in neuroscience, 11:309, 2017.
  29. 29.Li, Y., Deng, S., Dong, X., Gong, R., and Gu, S. A free lunch from ANN: towards efficient, accurate spiking neural networks calibration. In ICML, volume 139, pp. 6316–6325, 2021a.
  30. 30.Li, Y., Guo, Y., Zhang, S., Deng, S., Hai, Y., and Gu, S. Differentiable spike: Rethinking gradient-descent for training spiking neural networks. NeurIPS, 34, 2021b.
  31. 31.Mostafa, H. Supervised learning based on temporal coding in spiking neural networks. IEEE transactions on neural networks and learning systems, 29(7):3227–3235, 2017.
  32. 32.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pp. 7197–7206. PMLR, 2020.
  33. 33.Neftci, E. O., Mostafa, H., and Zenke, F. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6):51–63, 2019.
  34. 34.Nowotny, T., Turner, J. P., and Knight, J. C. Loss shaping enhances exact gradient learning with eventprop in spiking neural networks. arXiv preprint arXiv:2212.01232, 2022.
  35. 35.Pei, J., Deng, L., Song, S., Zhao, M., Zhang, Y., Wu, S., Wang, G., Zou, Z., Wu, Z., He, W., et al. Towards artificial general intelligence with hybrid tianjic chip architecture. Nature, 572(7767):106–111, 2019.
  36. 36.Perez-Nieves, N. and Goodman, D. Sparse spiking gradient descent. Advances in Neural Information Processing Systems, 34, 2021.
  37. 37.Perez-Nieves, N., Leung, V. C., Dragotti, P. L., and Goodman, D. F. Neural heterogeneity promotes robust learning. Nature communications, 12(1):1–9, 2021.
  38. 38.Pons, J., Slizovskaia, O., Gong, R., Gómez, E., and Serra, X. Timbre analysis of music audio signals with convolutional neural networks. In 2017 25th European Signal Processing Conference (EUSIPCO), pp. 2744–2748. IEEE, 2017.
  39. 39.Rathi, N. and Roy, K. DIET-SNN: direct input encoding with leakage and threshold optimization in deep spiking neural networks. CoRR, abs/2008.03658, 2020.
  40. 40.Rathi, N., Srinivasan, G., Panda, P., and Roy, K. Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation. In ICLR 2020,. OpenReview.net, 2020.
  41. 41.Roy, K., Jaiswal, A., and Panda, P. Towards spike-based machine intelligence with neuromorphic computing. Nature, 575(7784):607–617, 2019.
  42. 42.Samadi, A., Lillicrap, T. P., and Tweed, D. B. Deep learning with dynamic spiking neurons and fixed feedback weights. Neural computation, 29(3):578–602, 2017.
  43. 43.Severa, W., Vineyard, C. M., Dellana, R., Verzi, S. J., and Aimone, J. B. Training deep neural networks for binary communication with the whetstone method. Nature Machine Intelligence, 1(2):86–94, 2019.
  44. 44.Shrestha, S. B. and Orchard, G. Slayer: Spike layer error reassignment in time. arXiv preprint arXiv:1810.08646, 2018.
  45. 45.Wang, Q., Zhang, Y., Yuan, J., and Lu, Y. Space-time event clouds for gesture recognition: From rgb cameras to event cameras. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1826–1835. IEEE, 2019.
  46. 46.Wang, Z., Lian, S., Zhang, Y., Cui, X., Yan, R., and Tang, H. Towards lossless ann-snn conversion under ultra-low latency with dual-phase optimization. arXiv preprint arXiv:2205.07473, 2022.
  47. 47.Wei, C., Kakade, S., and Ma, T. The implicit and explicit regularization effects of dropout. In International Conference on Machine Learning, pp. 10181–10192. ICML, 2020.
  48. 48.Wu, J., Chua, Y., Zhang, M., Li, G., Li, H., and Tan, K. C. A tandem learning rule for effective training and rapid inference of deep spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems, pp. 1–15, 2021a.
  49. 49.Wu, Y., Deng, L., Li, G., Zhu, J., and Shi, L. Spatio-temporal backpropagation for training high-performance spiking neural networks. Frontiers in neuroscience, 12: 331, 2018.
  50. 50.Wu, Y., Deng, L., Li, G., Zhu, J., Xie, Y., and Shi, L. Direct training for spiking neural networks: Faster, larger, better. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 1311–1318, 2019.
  51. 51.Wu, Z., Zhang, H., Lin, Y., Li, G., Wang, M., and Tang, Y. Liaf-net: Leaky integrate and analog fire network for lightweight and efficient spatiotemporal information processing. IEEE Transactions on Neural Networks and Learning Systems, 2021b.
  52. 52.Wunderlich, T. C. and Pehle, C. Event-based backpropagation can compute exact gradients for spiking neural networks. Scientific Reports, 11(1):1–17, 2021.
  53. 53.Yang, Q., Wu, J., Zhang, M., Chua, Y., Wang, X., and Li, H. Training spiking neural networks with local tandem learning. NeurIPS, 2022.
  54. 54.Yang, Y., Zhang, W., and Li, P. Backpropagated neighborhood aggregation for accurate training of spiking neural networks. In ICML, pp. 11852–11862, 2021.
  55. 55.Yao, M., Gao, H., Zhao, G., Wang, D., Lin, Y., Yang, Z., and Li, G. Temporal-wise attention spiking neural networks for event streams classification. In ICCV, pp. 10221–10230, 2021.
  56. 56.Yao, X., Li, F., Mo, Z., and Cheng, J. Glif: A unified gated leaky integrate-and-fire neuron for spiking neural networks. arXiv preprint arXiv:2210.13768, 2022.
  57. 57.Yin, B., Corradi, F., and Bohte, S. M. Effective and efficient computation with multiple-timescale spiking recurrent neural networks. In International Conference on Neuromorphic Systems 2020, pp. 1–8, 2020.
  58. 58.Zenke, F. and Vogels, T. P. The remarkable robustness of surrogate gradient learning for instilling complex function in spiking neural networks. Neural Computation, 33(4): 899–925, 2021.
  59. 59.Zhang, W. and Li, P. Temporal spike sequence learning via backpropagation for deep spiking neural networks. In NeurIPS, 2020.
  60. 60.Zheng, H., Wu, Y., Deng, L., Hu, Y., and Li, G. Going deeper with directly-trained larger spiking neural networks. In AAAI 2021, pp. 11062–11070, 2021.

Citation

MLA
Wang, Z., et al. “Adaptive Smoothing Gradient Learning for Spiking Neural Networks”. International Conference on Machine Learning, vol. 202, 2023, pp. 35798–816, https://proceedings.mlr.press/v202/wang23j.html.
APA
Wang, Z., Jiang, R., Lian, S., Yan, R., & Tang, H. (2023). Adaptive Smoothing Gradient Learning for Spiking Neural Networks. International Conference on Machine Learning, 202, 35798–35816. https://proceedings.mlr.press/v202/wang23j.html
Chicago
Wang, Z., R. Jiang, S. Lian, R. Yan, and H. Tang. 2023. “Adaptive Smoothing Gradient Learning for Spiking Neural Networks”. International Conference on Machine Learning 202: 35798–816. https://proceedings.mlr.press/v202/wang23j.html.
Harvard
Wang, Z. et al. (2023) “Adaptive Smoothing Gradient Learning for Spiking Neural Networks”, International Conference on Machine Learning. PMLR, pp. 35798–35816. Available at: https://proceedings.mlr.press/v202/wang23j.html.
Vancouver
1. Wang Z, Jiang R, Lian S, Yan R, Tang H (2023) Adaptive Smoothing Gradient Learning for Spiking Neural Networks. In: International Conference on Machine Learning. PMLR, pp 35798–35816

BibTeX

@InProceedings{pmlr-v202-wang23j,
  title = 	 {Adaptive Smoothing Gradient Learning for Spiking Neural Networks},
  author =       {Wang, Ziming and Jiang, Runhao and Lian, Shuang and Yan, Rui and Tang, Huajin},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {35798--35816},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/wang23j/wang23j.pdf},
  url = 	 {https://proceedings.mlr.press/v202/wang23j.html},
  abstract = 	 {Spiking neural networks (SNNs) with biologically inspired spatio-temporal dynamics demonstrate superior energy efficiency on neuromorphic architectures. Error backpropagation in SNNs is prohibited by the all-or-none nature of spikes. The existing solution circumvents this problem by a relaxation on the gradient calculation using a continuous function with a constant relaxation de- gree, so-called surrogate gradient learning. Nevertheless, such a solution introduces additional smoothing error on spike firing which leads to the gradients being estimated inaccurately. Thus, how to adaptively adjust the relaxation degree and eliminate smoothing error progressively is crucial. Here, we propose a methodology such that training a prototype neural network will evolve into training an SNN gradually by fusing the learnable relaxation degree into the network with random spike noise. In this way, the network learns adaptively the accurate gradients of loss landscape in SNNs. The theoretical analysis further shows optimization on such a noisy network could be evolved into optimization on the embedded SNN with shared weights progressively. Moreover, The experiments on static images, dynamic event streams, speech, and instrumental sounds show the proposed method achieves state-of-the-art performance across all the datasets with remarkable robustness on different relaxation degrees.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/