SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation

Malyaban BalAbhronil Sengupta

article2024AAAI77 citations

Proposes a scalable spiking language model trained via equilibrium-state implicit differentiation and BERT knowledge distillation, bypassing non-differentiability challenges without surrogate gradients to deliver energy-efficient NLP performance on the GLUE benchmark.

Listen

Modern large language models require substantial computational power and energy to train and run, creating growing environmental, infrastructure, and financial burdens. While brain-inspired spiking neural networks communicate via discrete electrical pulses to achieve high energy efficiency on specialized neuromorphic hardware, scaling them to complex natural language processing tasks has been historically blocked by high training memory requirements and non-differentiable spiking dynamics.

The article demonstrates a fully operational, energy-efficient spiking language model called SpikingBERT, evaluating its performance across diverse language understanding tasks and establishing a scalable training framework.

To overcome traditional training barriers, the researchers modeled the spiking network as a dynamic system that converges to a steady equilibrium rate of firing over time. This design allows the model to compute training gradients implicitly at equilibrium, completely eliminating the need to store massive computational histories in memory or rely on approximate surrogate gradients. The authors designed a specialized spiking attention mechanism that approximates conventional attention while relying on low-energy accumulation operations rather than power-intensive multiplications. Furthermore, they introduced a two-stage knowledge distillation method, transferring learned representations from a full-sized standard BERT model into a compact four-layer spiking model using equilibrium firing states.

The findings establish that SpikingBERT achieves competitive accuracy compared to standard and efficient non-spiking language models across seven language benchmark tasks, reaching 88.19% accuracy on sentiment analysis and 86.82% on question-pair classification. Crucially, eliminating the proposed knowledge distillation process led to an immediate 4% to 5% drop in accuracy across all evaluated benchmarks. On specialized hardware architectures, the spike-based design utilizes accumulation operations that consume approximately 0.9 picojoules per operation—making them over five times more energy-efficient than standard matrix operations consuming 4.6 picojoules. Analysis of trade-offs demonstrated that setting an operating threshold of 16 convergence time steps delivers nearly double the energy efficiency of an equivalent non-spiking model while sacrificing only 2% in task accuracy.

These results confirm that brain-inspired spiking models can effectively perform complex natural language understanding tasks at significantly lower energy footprints. This technological shift opens up viable deployment pathways for sophisticated language processing on resource-constrained edge devices and mobile systems, mitigating the operational costs and energy consumption typical of artificial intelligence systems.

Organizations planning low-power natural language processing deployments should consider developing and piloting spiking transformer architectures on specialized neuromorphic hardware processors. For near-term implementation, engineering teams should leverage the multi-stage knowledge distillation approach and adjust neuron threshold voltages and convergence steps to balance instantaneous power against latency and accuracy requirements.

A slight performance gap remains between SpikingBERT and fully sized, fine-tuned standard language models. Additionally, real-world energy savings depend on deploying models directly onto neuromorphic hardware rather than general-purpose graphic processing units. Nevertheless, the theoretical stability proofs and consistent benchmark validations provide strong confidence in the viability of the training methodology.

Cover for SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation

Abstract

Large language Models (LLMs), though growing exceedingly powerful, comprises of orders of magnitude less neurons and synapses than the human brain. However, it requires significantly more power/energy to operate. In this work, we propose a novel bio-inspired spiking language model (LM) which aims to reduce the computational cost of conventional LMs by drawing motivation from the synaptic information flow in the brain. In this paper, we demonstrate a framework that leverages the average spiking rate of neurons at equilibrium to train a neuromorphic spiking LM using implicit differentiation technique, thereby overcoming the non-differentiability problem of spiking neural network (SNN) based algorithms without using any type of surrogate gradient. The steady-state convergence of the spiking neurons also allows us to design a spiking attention mechanism, which is critical in developing a scalable spiking LM. Moreover, the convergence of average spiking rate of neurons at equilibrium is utilized to develop a novel ANN-SNN knowledge distillation based technique wherein we use a pre-trained BERT model as “teacher” to train our “student” spiking architecture. While the primary architecture proposed in this paper is motivated by BERT, the technique can be potentially extended to different kinds of LLMs. Our work is the first one to demonstrate the performance of an operational spiking LM architecture on multiple different tasks in the GLUE benchmark. Our implementation source code is available at https://github.com/NeuroCompLab-psu/SpikingBERT.

Table of Contents

  • Introduction
  • Related Works
  • Methods
  • Spiking Neural Networks
  • Implicit Modeling
  • Architecture
  • Spiking Attention Mechanism
  • ANN-SNN KD using Equilibrium States
  • Experimentation
  • Datasets
  • Baselines & SpikingBERT Settings
  • Analysis of Power & Energy Efficiency
  • Conclusion and Future Works
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — SpikingBERT Architecture and LIF Neuron Dynamics

    model/method

    SpikingBERT is an encoder-only spiking language model composed of a stack of NN Spiking Encoder (SE) blocks. Token embeddings are converted into spike trains through an initial Leaky Integrate-and-Fire (LIF) neuron layer. Within each SE block, communication occurs exclusively via discrete binary spikes across a Spiking Attention module, fully connected intermediate layers (Intermediate Layer-1 and Intermediate Layer-2), Layer Normalization, and residual skip connections.

    The membrane potential dynamics of discrete-time LIF neuron ii at time step tt are governed by:

    ui[t+δ]=γui[t]+∑j(wijsj[t])+biu_i[t + \delta] = \gamma u_i[t] + \sum_j (w_{ij} s_j[t]) + b_i

    si[t+1]=S(ui[t+δ])s_i[t + 1] = S(u_i[t + \delta])

    ui[t+1]=ui[t+δ]−Vthsi[t+1]u_i[t + 1] = u_i[t + \delta] - V_{th} s_i[t + 1]

    where ui[t]u_i[t] represents the membrane potential of neuron ii at time tt, γ≤1\gamma \le 1 is the leak factor (gamma<1\\gamma < 1 for LIF neurons, γ=1\gamma = 1 for Integrate-and-Fire (IF) neurons), wijw_{ij} is the synaptic weight connecting pre-synaptic neuron jj to post-synaptic neuron ii, sj[t]∈{0,1}s_j[t] \in \{0, 1\} is the binary spike emitted by pre-synaptic neuron jj at time tt, bib_i is a learnable bias, VthV_{th} is the firing threshold, and S(⋅)S(\cdot) is the non-differentiable step function defined as S(x)=1S(x) = 1 if x>Vthx > V_{th} and 00 otherwise. An optional feedback connection FF can feed spikes from the final SE layer output s(N,out)[t]s_{(N,out)}[t] back into the first SE layer input:

    u1[t+1]=γu1[t]+Fs(N,out)[t]+W0(x)+b1−Vths1[t+1]u_1[t + 1] = \gamma u_1[t] + F s_{(N,out)}[t] + W_0(x) + b_1 - V_{th} s_1[t + 1]

    where W0(x)W_0(x) is the input embedding mapping token sequence xx of length NsN_s to dimension RNs×Demb\mathbb{R}^{N_s \times D_{emb}}.

  2. Knowl 2 — Equilibrium State Formulation for Spiking Rate in Spiking Networks

    theoretical result

    For a spiking layer ii with synaptic operation W(i−1)W_{(i-1)}, the Average Spiking Rate (ASR) at time step tt is defined as the exponentially weighted average over past spikes:

    ai[t]=∑τ=1tγt−τsi[τ]∑τ=1tγt−τa_i[t] = \frac{\sum_{\tau=1}^t \gamma^{t-\tau} s_i[\tau]}{\sum_{\tau=1}^t \gamma^{t-\tau}}

    Assuming initial conditions ui[0]=0u_i[0] = 0 and si[0]=0s_i[0] = 0, applying a weighted temporal average to the discrete-time LIF update yields:

    ai[t+1]=1Vth(W(i−1)ai−1[t+1]+bi−ui[t+1]∑τ=0tγτ)a_i[t + 1] = \frac{1}{V_{th}} \left( W_{(i-1)} a_{i-1}[t + 1] + b_i - \frac{u_i[t + 1]}{\sum_{\tau=0}^t \gamma^\tau} \right)

    When the input sequence converges to a steady state xˉ[t]→x∗\bar{x}[t] \to x^*, the membrane potential remains bounded while the denominator ∑τ=0tγτ\sum_{\tau=0}^t \gamma^\tau accumulates. Consequently, the layer-wise ASR converges to an equilibrium fixed point ai∗a_i^* satisfying:

    ai∗=σ(1Vth(W(i−1)ai−1∗+bi))a_i^* = \sigma\left( \frac{1}{V_{th}} (W_{(i-1)} a_{i-1}^* + b_i) \right)

    where σ(z)=max⁡(0,min⁡(1,z))\sigma(z) = \max(0, \min(1, z)) is the clipping function bounding firing rates to the interval [0,1][0, 1]. For an MM-layer network with feedback FF, the equilibrium state of the first layer satisfies the fixed-point equation a1∗=l1(lM∘⋯∘l2(a1∗),x∗)a_1^* = l_1(l_M \circ \dots \circ l_2(a_1^*), x^*), with l1(a,x)=σ(1Vth(Fa+W0(x)+b1))l_1(a, x) = \sigma\left(\frac{1}{V_{th}}(F a + W_0(x) + b_1)\right).

  3. Knowl 3 — Training Spiking Language Models via Implicit Differentiation

    model/method

    Rather than unrolling the network across TconvT_{conv} operational time steps using backpropagation through time (BPTT) with surrogate gradients, SpikingBERT is trained by differentiating the dynamical system at its equilibrium state.

    Let the converged steady-state ASR across network layers be represented as the fixed-point equation z∗=fθ(z∗)z^* = f_\theta(z^*), where θ\theta denotes the model parameters and z∗=zTconvz^* = z_{T_{conv}} is the state after TconvT_{conv} forward simulation steps. Defining the root-finding equation gθ(z)=fθ(z)−z=0g_\theta(z) = f_\theta(z) - z = 0, the gradient of a scalar loss function L(z∗)\mathcal{L}(z^*) with respect to θ\theta is derived via the implicit function theorem:

    ∂L(z∗)∂θ=−∂L(z∗)∂z∗(Jgθ∣z∗)−1∂fθ(z∗)∂θ\frac{\partial \mathcal{L}(z^*)}{\partial \theta} = - \frac{\partial \mathcal{L}(z^*)}{\partial z^*} \left( J_{g_\theta}\big|_{z^*} \right)^{-1} \frac{\partial f_\theta(z^*)}{\partial \theta}

    where Jgθ∣z∗=∂gθ(z)∂z∣z∗J_{g_\theta}\big|_{z^*} = \frac{\partial g_\theta(z)}{\partial z}\big|_{z^*} is the Jacobian of gθg_\theta evaluated at equilibrium z∗z^*. This computes exact parameter gradients directly through the surrogate equilibrium network without storing intermediate unrolled temporal activations, eliminating BPTT memory overhead.

  4. Knowl 4 — Spiking Attention Mechanism and Equilibrium Equivalence

    model/method

    The Spiking Attention module processes binary input spike sequences Sx(t)∈{0,1}Ns×DembS_x(t) \in \{0, 1\}^{N_s \times D_{emb}} at time step tt using the following operations:

    Attn(Sx(t),SK(t),SV(t))=π(s⋅Q(Sx(t))(SK(t))T)⋅SV(t)\text{Attn}(S_x(t), S_K(t), S_V(t)) = \pi\left( s \cdot Q(S_x(t)) (S_K(t))^T \right) \cdot S_V(t)

    where Q(Sx(t))=WQSx(t)Q(S_x(t)) = W_Q S_x(t) computes the Query matrix using real weights WQW_Q, while Key spikes SK(t)S_K(t) and Value spikes SV(t)S_V(t) are obtained by passing linear transformations WKSx(t)W_K S_x(t) and WVSx(t)W_V S_x(t) through discrete LIF neuron layers. Here, π(⋅)\pi(\cdot) is a normalization function (e.g., softmax or identity), and s=1dks = \frac{1}{\sqrt{d_k}} is the scaling factor with dkd_k being the Key dimension.

    Because SK(t)S_K(t) and SV(t)S_V(t) consist purely of discrete binary spikes {0,1}\{0, 1\}, the matrix multiplications require only O(n3)\mathcal{O}(n^3) accumulative (ACC) operations rather than O(n3)\mathcal{O}(n^3) floating-point multiply-accumulate (MAC) operations.

    At equilibrium, the surrogate steady-state ASR a(attn)∗a_{(attn)}^* of the attention module satisfies:

    a(attn)∗=σ(1Vth(Attn(ax∗,ak∗,av∗)+b(attn)))a_{(attn)}^* = \sigma\left( \frac{1}{V_{th}} \left( \text{Attn}(a_x^*, a_k^*, a_v^*) + b_{(attn)} \right) \right)

    where ax∗,ak∗,av∗a_x^*, a_k^*, a_v^* denote the equilibrium ASRs of the Query, Key, and Value inputs, respectively.

  5. Knowl 5 — Multi-Stage ANN-to-SNN Knowledge Distillation via Equilibrium States

    model/method

    SpikingBERT transfers representations from a pre-trained non-spiking BERT teacher to the spiking student LM by aligning the student's equilibrium ASRs with the teacher's continuous activations across two stages: general domain distillation on Wikipedia text, followed by task-specific distillation.

    Distillation utilizes three objectives evaluated at the steady-state equilibrium:

    1. Intermediate Layer Distillation: The MSE loss between the linear-projected equilibrium ASR of the ii-th Spiking Encoder output ahi∗a_{h_i}^* and the activation of the mapped teacher layer Tf(hi)T_{f(h_i)}:

    Lhi=MSE(ahi∗WTd,  Tf(hi))\mathcal{L}_{h_i} = \text{MSE}\left( a_{h_i}^* W_{Td},\; T_{f(h_i)} \right)

    where WTdW_{Td} is a dimension-aligning linear projection, and the layer mapping function is f(hi)=hp⋅i′f(h_i) = h'_{p \cdot i} with p=Tenc/Sencp = T_{enc} / S_{enc} (TencT_{enc} and SencS_{enc} being the number of teacher and student encoder layers; p=12/4=3p = 12/4 = 3).

    1. Attention and Embedding Distillation: MSE losses between the student's equilibrium attention maps and teacher attention maps, and between student embedding ASR and teacher embedding activations.

    2. Prediction Layer Distillation: Cross-entropy loss between student logits derived from the final encoder equilibrium ASR apred∗a_{pred}^* and teacher soft logits TpredT_{pred} scaled by temperature t′t':

    Lpred=CE(c(apred∗)t′,  Tpredt′)\mathcal{L}_{pred} = \text{CE}\left( \frac{c(a_{pred}^*)}{t'},\; \frac{T_{pred}}{t'} \right)

    where cc is a linear classification projection.

  6. Knowl 6 — Normalized Operations and SNN Energy Efficiency Factor Formulation

    equation

    In 45 nm45\text{ nm} CMOS technology, spike-based accumulative operations (ACC) consume 0.9 pJ0.9\text{ pJ}, whereas multiply-accumulate operations (MAC) in standard non-spiking ANNs consume 4.6 pJ4.6\text{ pJ} (a factor of 5.1×5.1\times higher).

    The total normalized operations (Norm#OPS\text{Norm\#OPS}) of a spiking neural network relative to an equivalent iso-architecture ANN are defined as:

    Norm#OPS=∑iIFRi⋅Layer#OPSi+1∑iLayer#OPSi\text{Norm\#OPS} = \frac{\sum_i \text{IFR}_i \cdot \text{Layer\#OPS}_{i+1}}{\sum_i \text{Layer\#OPS}_i}

    where IFRi\text{IFR}_i is the inter-spike firing rate (total spikes over inference time steps divided by the total number of neurons) of layer ii, and Layer#OPSi\text{Layer\#OPS}_i is the theoretical operation count of layer ii.

    The overall energy-efficiency factor ee, defined as the ratio of energy consumed by an iso-architecture ANN to the proposed SNN, is:

    e=(15.1⋅Norm#OPS)−1e = \left( \frac{1}{5.1} \cdot \text{Norm\#OPS} \right)^{-1}

  7. Knowl 7 — Evaluation of SpikingBERT on the GLUE Benchmark

    data/table

    SpikingBERT4 (4 Spiking Encoder blocks, 50M parameters) was evaluated across seven tasks from the General Language Understanding Evaluation (GLUE) benchmark against baseline sequence models and compact non-spiking BERT variants. Tasks include Quora Question Pair (QQP), Multi-Genre Natural Language Inference (MNLI-m), Stanford Sentiment Treebank (SST-2), Question-answering NLI (QNLI), Recognizing Textual Entailment (RTE), Microsoft Research Paraphrase Corpus (MRPC), and Semantic Textual Similarity Benchmark (STS-B).

    Model QQP MNLI-m SST-2 QNLI RTE MRPC STS-B
    CBoW 75.00 57.10 79.50 62.50 71.90 75.00/83.70 70.60/71.10
    BiLSTM 85.30 66.70 87.50 77.00 58.50 77.90/85.10 71.60/72.00
    BiLSTM + Attn, CoVe 83.50 67.90 89.20 72.50 58.10 72.80/82.40 59.40/58.00
    GenSen 82.60 71.40 87.20 62.50 78.40 80.40/86.20 81.30/81.80
    BERT5\text{BERT}_5 + PF 84.10 67.70 81.60 80.90 62.80 78.60/- -/81.10
    NAS-BERT5\text{NAS-BERT}_5 + PF 85.70 74.20 84.90 83.90 67.00 80.00/- -/82.80
    NAS-BERT5\text{NAS-BERT}_5 + KD 85.80 74.40 87.30 84.90 66.60 79.60/- -/83.00
    NAS-BERT10\text{NAS-BERT}_{10} + PF 88.40 76.00 88.60 86.30 68.70 81.50/- -/84.30
    BERTTINY\text{BERT}_{\text{TINY}} Adam 81.09 65.36 80.11 77.85 - 69.90/- 64.39/-
    BERTMINI\text{BERT}_{\text{MINI}} Adam 86.45 73.30 85.46 83.85 - 76.57/- 82.09/-
    SpikingBERT4_4 86.82 78.10 88.19 85.20 66.06 79.17/85.15 82.20/81.90
    TinyBERT4\text{TinyBERT}_4 (no DA) 88.50 80.60 90.50 87.00 68.20 82.40/- 86.20/85.70

    Metric for QQP, MNLI-m, SST-2, QNLI, and RTE is accuracy (%); MRPC reports accuracy/F1 score; STS-B reports Pearson/Spearman correlation. SpikingBERT4 outperforms non-spiking baselines of comparable depth and width (e.g., BERTTINY, BERTMini, BiLSTM) across multiple GLUE tasks while operating via sparse spiking additions.

  8. Knowl 8 — Hyperparameter Specifications for SpikingBERT Training

    experimental setup

    The architectural and training hyperparameters for SpikingBERT4 across general and task-based knowledge distillation stages are specified below:

    Hyper-parameter Explored Range Optimal Value
    TconvT_{conv}: General KD (5–150) 80
    TconvT_{conv}: Task-based IKD (5–150) 80
    VthV_{th} (Threshold Voltage) (0.25–5.0) 1.0
    γ\gamma (Leak term) (0.8–1.0) 0.99 (LIF); 1 (IF)
    t′t' (Temperature) (0.1–10.0) 1.0
    Batch Size: General KD (8–256) 128
    Batch Size: Task-based IKD (8–128) [16, 32]
    Epochs: General KD – 5
    Epochs: Task-based IKD – 20

    All models use a maximum sequence length of 128 tokens, an embedding dimension Demb=768D_{emb} = 768, an intermediate feed-forward layer (IL-2) size of 3072, and 4 Spiking Encoder (SE) blocks. Distillation omitting internal-layer KD causes a 4% to 5% accuracy drop across all GLUE datasets.

  9. Knowl 9 — Energy-Accuracy Tradeoff and Firing Threshold Dynamics in SpikingBERT

    empirical result

    Empirical analysis on the SST-2 dataset demonstrates two operational trade-offs:

    1. Latency-Energy Tradeoff: Decreasing the forward convergence horizon from Tconv=80T_{conv} = 80 to Tconv=16T_{conv} = 16 reduces accuracy by only 2% (from ≈88.2%\approx 88.2\% to ≈86.2%\approx 86.2\%) while doubling the energy efficiency factor ee to approximately 2×2\times that of an iso-architecture non-spiking ANN.

    2. Firing Threshold (VthV_{th}) Tuning: Increasing VthV_{th} from 1.01.0 to 2.0 V2.0\text{ V} systematically reduces the mean Average Spiking Rate (ASR) per neuron across all internal sub-layers of SpikingBERT4. Although higher VthV_{th} extends the number of timesteps required for equilibrium convergence, it lowers instantaneous power consumption, facilitating deployment on power-constrained neuromorphic edge devices.

  10. Knowl 10 — Performance Gap and Architecture Constraints of SpikingBERT

    limitation

    SpikingBERT exhibits two primary limitations:

    1. Performance Gap: A residual performance gap remains between the distilled SpikingBERT4 and the fully fine-tuned 12-layer non-spiking teacher BERTBASE\text{BERT}_{\text{BASE}} across GLUE benchmark tasks (e.g., MNLI-m score of 78.10% versus TinyBERT4\text{TinyBERT}_4's 80.60%).

    2. Ineffectiveness of Recurrent Feedback: Unlike vision-based equilibrium SNNs where recurrent feedback connections consistently enhance performance, adding a feedback connection FF from the final SE layer to the first SE layer does not provide measurable accuracy gains on GLUE benchmark tasks.

Coverage note — None was omitted; all key contributions including neuron dynamics, equilibrium rate equations, implicit differentiation formulation, spiking attention, multi-stage KD, energy/OPS models, GLUE experimental results, hyperparameter configurations, and empirical tradeoff analyses are covered.

References

  1. 1.Alawad, M.; Yoon, H.-J.; and Tourassi, G. 2017. Energy efficient stochastic-based deep spiking neural networks for sparse datasets. In 2017 IEEE International Conference on Big Data (Big Data), 311–318. IEEE.
  2. 2.Amir, A.; Taba, B.; Berg, D.; Melano, T.; McKinstry, J.; Di Nolfo, C.; Nayak, T.; Andreopoulos, A.; Garreau, G.; Mendoza, M.; et al. 2017. A low power, fully event-based gesture recognition system. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7243–7252.
  3. 3.Bai, S.; Kolter, J. Z.; and Koltun, V. 2019. Deep equilibrium models. Advances in Neural Information Processing Systems, 32.
  4. 4.Bal, M.; and Sengupta, A. 2022. Sequence Learning using Equilibrium Propagation. arXiv preprint arXiv:2209.09626.
  5. 5.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  6. 6.Davies, M.; Wild, A.; Orchard, G.; Sandamirskaya, Y.; Guerra, G. A. F.; Joshi, P.; Plank, P.; and Risbud, S. R. 2021. Advancing neuromorphic computing with loihi: A survey of results and outlook. Proceedings of the IEEE, 109(5): 911–934.
  7. 7.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee.
  8. 8.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  9. 9.Frantar, E.; Kurtic, E.; and Alistarh, D. 2021. M-FAC: Efficient matrix-free approximations of second-order information. Advances in Neural Information Processing Systems, 34: 14873–14886.
  10. 10.Ghosh-Dastidar, S.; and Adeli, H. 2009. Spiking neural networks. International journal of neural systems, 19(04): 295–308.
  11. 11.Han, S.; Pool, J.; Tran, J.; and Dally, W. J. 2015. Learning both Weights and Connections for Efficient Neural Networks. arXiv:1506.02626.
  12. 12.Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  13. 13.Hong, D.; Shen, J.; Qi, Y.; and Wang, Y. 2023. LaSNN: Layer-wise ANN-to-SNN Distillation for Effective and Efficient Training in Deep Spiking Neural Networks. arXiv preprint arXiv:2304.09101.
  14. 14.Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2019. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351.
  15. 15.Kim, S.; Gholami, A.; Yao, Z.; Mahoney, M. W.; and Keutzer, K. 2021. I-bert: Integer-only bert quantization. In International conference on machine learning, 5506–5518. PMLR.
  16. 16.Krizhevsky, A.; Nair, V.; and Hinton, G. 2009. Cifar-10 and cifar-100 datasets. URl: https://www. cs. toronto. edu/kriz/cifar. html, 6(1): 1.
  17. 17.Kubilius, J.; Schrimpf, M.; Kar, K.; Rajalingham, R.; Hong, H.; Majaj, N.; Issa, E.; Bashivan, P.; Prescott-Roy, J.; Schmidt, K.; et al. 2019. Brain-like object recognition with high-performing shallow recurrent ANNs. Advances in neural information processing systems, 32.
  18. 18.Kurtic, E.; Campos, D.; Nguyen, T.; Frantar, E.; Kurtz, M.; Fineran, B.; Goin, M.; and Alistarh, D. 2022. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259.
  19. 19.Lee, C.; Sarwar, S. S.; Panda, P.; Srinivasan, G.; and Roy, K. 2020. Enabling spike-based backpropagation for training deep neural network architectures. Frontiers in neuroscience, 119.
  20. 20.Lu, S.; and Sengupta, A. 2020. Exploring the connection between binary and spiking neural networks. Frontiers in neuroscience, 14: 535.
  21. 21.Neftci, E. O.; Mostafa, H.; and Zenke, F. 2019. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Processing Magazine, 36(6): 51–63.
  22. 22.Orchard, G.; Jayawant, A.; Cohen, G. K.; and Thakor, N. 2015. Converting static image datasets to spiking neuromorphic datasets using saccades. Frontiers in neuroscience, 9: 437.
  23. 23.Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018. Improving language understanding by generative pre-training.
  24. 24.Scellier, B.; and Bengio, Y. 2017. Equilibrium propagation: Bridging the gap between energy-based models and backpropagation. Frontiers in computational neuroscience, 11: 24.
  25. 25.Sengupta, A.; Ye, Y.; Wang, R.; Liu, C.; and Roy, K. 2019. Going deeper in spiking neural networks: VGG and residual architectures. Frontiers in neuroscience, 13: 95.
  26. 26.Takuya, S.; Zhang, R.; and Nakashima, Y. 2021. Training low-latency spiking neural network through knowledge distillation. In 2021 IEEE Symposium in Low-Power and High-Speed Chips (COOL CHIPS), 1–3. IEEE.
  27. 27.Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019. Distilling task-specific knowledge from bert into simple neural networks. arXiv preprint arXiv:1903.12136.
  28. 28.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. arXiv:1706.03762.
  29. 29.Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  30. 30.Wei, J.; Tay, Y.; Bommasani, R.; Raffel, C.; Zoph, B.; Borgeaud, S.; Yogatama, D.; Bosma, M.; Zhou, D.; Metzler, D.; Chi, E. H.; Hashimoto, T.; Vinyals, O.; Liang, P.; Dean, J.; and Fedus, W. 2022. Emergent Abilities of Large Language Models. arXiv:2206.07682.
  31. 31.Xiao, M.; Meng, Q.; Zhang, Z.; Wang, Y.; and Lin, Z. 2021. Training feedback spiking neural networks by implicit differentiation on the equilibrium state. Advances in Neural Information Processing Systems, 34: 14516–14528.
  32. 32.Xu, J.; Tan, X.; Luo, R.; Song, K.; Li, J.; Qin, T.; and Liu, T.-Y. 2021. NAS-BERT: task-agnostic and adaptive-size BERT compression with neural architecture search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 1933–1943.
  33. 33.Xu, Q.; Li, Y.; Shen, J.; Liu, J. K.; Tang, H.; and Pan, G. 2023. Constructing deep spiking neural networks from artificial neural networks with knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7886–7895.
  34. 34.Zhou, Z.; Zhu, Y.; He, C.; Wang, Y.; Yan, S.; Tian, Y.; and Yuan, L. 2022. Spikformer: When spiking neural network meets transformer. arXiv preprint arXiv:2209.15425.
  35. 35.Zhu, R.-J.; Zhao, Q.; and Eshraghian, J. K. 2023. Spikegpt: Generative pre-trained language model with spiking neural networks. arXiv preprint arXiv:2302.13939.

Citation

MLA
Bal, M., and A. Sengupta. “SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 10, 2024, pp. 10998–1006, https://doi.org/10.1609/AAAI.V38I10.28975.
APA
Bal, M., & Sengupta, A. (2024). SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation. Proceedings of the AAAI Conference on Artificial Intelligence, 38(10), 10998–11006. https://doi.org/10.1609/AAAI.V38I10.28975
Chicago
Bal, M., and A. Sengupta. 2024. “SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation”. Proceedings of the AAAI Conference on Artificial Intelligence 38 (10): 10998–11006. https://doi.org/10.1609/AAAI.V38I10.28975.
Harvard
Bal, M. and Sengupta, A. (2024) “SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation”, Proceedings of the AAAI Conference on Artificial Intelligence, 38(10), pp. 10998–11006. Available at: https://doi.org/10.1609/AAAI.V38I10.28975.
Vancouver
1. Bal M, Sengupta A (2024) SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation. Proceedings of the AAAI Conference on Artificial Intelligence 38:10998–11006

BibTeX

@article{Bal_2024, title={SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation}, volume={38}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V38I10.28975}, DOI={10.1609/aaai.v38i10.28975}, number={10}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Bal, Malyaban and Sengupta, Abhronil}, year={2024}, month=Mar, pages={10998–11006} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF