Gaussian Error Linear Units (GELUs)

Dan HendrycksKevin Gimpel

article2016arXiv7,346 citations

Introduces the Gaussian Error Linear Unit (GELU), a smooth activation function that weights inputs by their magnitude rather than gating them by sign, establishing a high-performing alternative to ReLU across vision, language, and speech tasks.

Listen

Deep neural network performance depends heavily on the choice of activation function, which determines how individual artificial neurons transform incoming signals to model complex patterns. Historically, network designs have relied on deterministic gating mechanisms like the Rectified Linear Unit (ReLU) or Exponential Linear Unit (ELU), while treating stochastic regularization techniques, such as dropout, as entirely separate architectural decisions. This separation leaves open the question of whether neural network activations can be made more effective by mathematically integrating probabilistic regularizers directly into the activation function itself.

The article sets out to introduce and evaluate the Gaussian Error Linear Unit (GELU), a novel nonlinearity designed to merge the probabilistic benefits of dropout-style regularizers with standard neuron activation. It aims to demonstrate that GELU outperforms standard ReLUs and ELUs across diverse core machine learning domains, including computer vision, natural language processing, and speech recognition.

The authors conducted a comprehensive set of empirical benchmark experiments comparing standard GELUs against ReLUs and ELUs across multiple architectures without introducing extra hyperparameters. The evaluation spanned image classification and self-supervised autoencoding on MNIST, part-of-speech tagging on social media text from Twitter, phone frame classification on the TIMIT acoustic speech dataset, and image classification on CIFAR-10 and CIFAR-100 using shallow convolutional and deep wide residual networks. Models were trained across varying learning rates and tested with momentum-based optimizers and dropout settings to ensure robust, fair comparisons.

The empirical findings consistently favored the new activation function. On deep image classification using a 40-layer Wide Residual Network on CIFAR-100, GELU achieved an error rate of 20.74%, outperforming ReLU at 21.77% and ELU at 22.98%. In standard CIFAR-10 image classification, GELU achieved a median error rate of 7.89%, surpassing ReLU (8.16%) and ELU (8.41%). In MNIST autoencoding, GELU accommodated multiple learning rates smoothly and yielded significantly lower reconstruction error than both alternatives. On natural language part-of-speech tagging and TIMIT speech recognition, GELU similarly secured the lowest median test errors (12.57% and 29.3%, respectively). Additionally, GELU exhibited superior resilience when artificial noise was added to input data during image classification.

These results indicate that continuous curvature and probabilistic weighting allow neural networks to fit complex functions and converge faster than traditional threshold-based activations. In practical terms, adopting GELU provides measurable accuracy gains without requiring additional hyperparameter tuning, thereby improving model performance without inflating computational engineering complexity. Because GELU acts as a smooth approximation to standard gating functions, it serves as an effective drop-in replacement across modern architectures.

Organizations developing or deploying deep learning systems should consider adopting GELU or its fast analytic approximations as the default nonlinearity in new and existing architectures. When deploying GELU, engineering teams should pair it with momentum-based optimization algorithms and use accurate approximations to ensure numerical stability and computational efficiency.

While the results demonstrate consistent advantages across multiple benchmarks, the experiments in the article rely on standard academic datasets and specific baseline architectures. Stakeholders should validate performance on proprietary internal datasets and evaluate runtime inference overhead when selecting between exact mathematical formulations and fast mathematical approximations.

  • Paper: GLU Variants Improve Transformer, Noam Shazeer. This work extends GELU into gated linear unit formulations (such as GEGLU) within Transformer feed-forward networks, demonstrating substantial downstream empirical gains.
  • Paper: Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning, Stefan Elfwing et al. (2017). This paper introduces and evaluates sigmoid-weighted linear units (SiLU/Swish), offering a closely related self-gated continuous nonlinearity evaluated across reinforcement learning domains.
  • Paper: Low-dimensional topology of deep neural networks, Junyu Ren et al. (2026). This theoretical study demonstrates why non-monotonic activations like GELU possess superior topological expressivity over monotonic units in unlinking data manifolds.
  • Paper: Self-Normalizing Neural Networks, Günter Klambauer et al. (2017). This paper follows the trajectory of modern activation design by introducing Scaled Exponential Linear Units (SELU) that theoretically induce self-normalizing dynamics across deep networks.
Cover for Gaussian Error Linear Units (GELUs)

Abstract

We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function. The GELU activation function is xΦ(x)x\Phi(x), where Φ(x)\Phi(x) the standard Gaussian cumulative distribution function. The GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs (x1x>0x\mathbf{1}_{x>0}). We perform an empirical evaluation of the GELU nonlinearity against the ReLU and ELU activations and find performance improvements across all considered computer vision, natural language processing, and speech tasks.

Table of Contents

  • 1 Introduction
  • 2 GELU Formulation
  • 3 GELU Experiments
  • 3.1 MNIST Classification
  • 3.2 MNIST Autoencoder
  • 3.3 Twitter POS Tagging
  • 3.4 TIMIT Frame Classification
  • 3.5 CIFAR-10/100 Classification
  • 4 Discussion
  • 5 Conclusion
  • References
  • A Neural network architecture for CIFAR-10 experiments
  • B History of the GELU and SiLU

Knowls

  1. Knowl 1 — Gaussian Error Linear Unit (GELU) Definition and Probabilistic Formulation

    definition

    The Gaussian Error Linear Unit (GELU) is a neural network activation function that weights an input x∈Rx \in \mathbb{R} by the probability that a standard normal variable X∼N(0,1)X \sim \mathcal{N}(0, 1) is less than or equal to xx:

    GELU(x)=x⋅Φ(x)=x⋅P(X≤x)=x⋅12[1+erf(x2)]\text{GELU}(x) = x \cdot \Phi(x) = x \cdot P(X \le x) = x \cdot \frac{1}{2} \left[ 1 + \text{erf}\left( \frac{x}{\sqrt{2}} \right) \right]

    where Φ(x)\Phi(x) is the cumulative distribution function (CDF) of the standard normal distribution N(0,1)\mathcal{N}(0, 1) and erf(⋅)\text{erf}(\cdot) is the standard error function:

    erf(z)=2π∫0ze−t2 dt\text{erf}(z) = \frac{2}{\sqrt{\pi}} \int_0^z e^{-t^2} \, dt

    Probabilistically, the GELU is the expected value of a stochastic gating operation. If an input xx is multiplied by a stochastic binary mask m∼Bernoulli(Φ(x))m \sim \text{Bernoulli}(\Phi(x)), then xx is retained (scaled by 11) with probability Φ(x)\Phi(x) and zeroed with probability 1−Φ(x)1 - \Phi(x). The expected transformation of this stochastic regularizer on xx yields:

    E[m⋅x]=Φ(x)⋅x+(1−Φ(x))⋅0=xΦ(x)\mathbb{E}[m \cdot x] = \Phi(x) \cdot x + (1 - \Phi(x)) \cdot 0 = x \Phi(x)

    Rather than deterministically gating inputs by sign as done in the Rectified Linear Unit (ReLU(x)=x1(x>0)\text{ReLU}(x) = x \mathbf{1}(x > 0)), the GELU weights inputs continuously according to their relative magnitude under a standard Gaussian prior.

  2. Knowl 2 — Fast Approximations for the GELU Activation Function

    equation

    Evaluating the exact Gaussian Error Linear Unit GELU(x)=xΦ(x)\text{GELU}(x) = x \Phi(x) requires computing the error function erf(x/2)\text{erf}(x / \sqrt{2}). For computationally efficient feedforward evaluation, the GELU can be approximated using either a hyperbolic tangent approximation or a scaled logistic sigmoid approximation:

    1. Hyperbolic tangent approximation: GELU(x)≈0.5x(1+tanh⁡[2π(x+0.044715x3)])\text{GELU}(x) \approx 0.5x \left( 1 + \tanh\left[ \sqrt{\frac{2}{\pi}} \left( x + 0.044715 x^3 \right) \right] \right)

    2. Scaled sigmoid approximation: GELU(x)≈x⋅σ(1.702x)=x1+e−1.702x\text{GELU}(x) \approx x \cdot \sigma(1.702 x) = \frac{x}{1 + e^{-1.702 x}}

    where σ(z)=1/(1+e−z)\sigma(z) = 1 / (1 + e^{-z}) is the standard logistic function. Both formulations provide smooth, fast approximations to xΦ(x)x\Phi(x) without evaluating error integrals.

  3. Knowl 3 — Sigmoid Linear Unit (SiLU) Definition

    definition

    The Sigmoid Linear Unit (SiLU) is an activation function formed by multiplying an input x∈Rx \in \mathbb{R} by the cumulative distribution function of the standard logistic distribution:

    SiLU(x)=x⋅σ(x)=x1+e−x\text{SiLU}(x) = x \cdot \sigma(x) = \frac{x}{1 + e^{-x}}

    where σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}} is the standard logistic sigmoid function.

    Like the GELU, the SiLU can be interpreted as the expectation of a stochastic binary gating variable m∼Bernoulli(σ(x))m \sim \text{Bernoulli}(\sigma(x)) acting on input xx. While SiLU yields competitive performance over standard ReLUs and ELUs across several tasks, it generally performs slightly below the standard normal GELU xΦ(x)x \Phi(x).

  4. Knowl 4 — Mathematical Properties and Relations of GELU to ReLU and ELU

    theoretical result

    The Gaussian Error Linear Unit GELU(x)=xΦ(x)\text{GELU}(x) = x\Phi(x) possesses several structural and asymptotic properties relative to standard activation functions:

    • Non-monotonicity and Curvature: Unlike the Rectified Linear Unit (ReLU(x)=max⁡(x,0)\text{ReLU}(x) = \max(x, 0)) and Exponential Linear Unit (ELU(x)=x\text{ELU}(x) = x if x>0x > 0, α(ex−1)\alpha(e^x - 1) if x≤0x \le 0), which are monotonic and linear for x>0x > 0, the GELU is non-convex and non-monotonic. It attains negative values for negative inputs (with a global minimum near x≈−0.7517x \approx -0.7517) and exhibits nonzero curvature across all points.
    • Asymptotic Equivalence to ReLU: In the limits, lim⁡x→∞GELU(x)=x\lim_{x \to \infty} \text{GELU}(x) = x and lim⁡x→−∞GELU(x)=0\lim_{x \to -\infty} \text{GELU}(x) = 0. For a parameterized Gaussian CDF Φμ,σ(x)=P(X≤x)\Phi_{\mu, \sigma}(x) = P(X \le x) where X∼N(μ,σ2)X \sim \mathcal{N}(\mu, \sigma^2), setting μ=0\mu = 0 and taking the limit σ→0\sigma \to 0 recovers the ReLU exactly: lim⁡σ→0xΦ0,σ(x)=x1(x>0)\lim_{\sigma \to 0} x \Phi_{0, \sigma}(x) = x \mathbf{1}(x > 0).
    • Cauchy CDF Relation to ELU: When replacing the Gaussian CDF with the CDF of a standard Cauchy distribution C∼Cauchy(0,1)C \sim \text{Cauchy}(0, 1), the resulting activation xP(C≤x)x P(C \le x) is asymptotically equivalent to an ELU with α=1/π\alpha = 1/\pi for negative values, and equivalent to xP(C≤x)−1/πx P(C \le x) - 1/\pi for positive values.
  5. Knowl 5 — Empirical Performance on CIFAR-10 and CIFAR-100 Image Classification

    empirical result

    The GELU activation function outperforms ReLU and ELU activations across shallow and deep convolutional architectures on CIFAR image classification benchmarks:

    • CIFAR-10 with a 9-layer CNN: On an unaugmented CIFAR-10 dataset using a 9-layer convolutional architecture trained with Adam for 200 epochs (batch normalization applied, initial learning rate tuned over {10−3,10−4,10−5}\{10^{-3}, 10^{-4}, 10^{-5}\} and linearly decayed to zero over the final 100 epochs), the median test error rates across three runs are:

      • GELU: 7.89%7.89\%
      • ReLU: 8.16%8.16\%
      • ELU: 8.41%8.41\%
    • CIFAR-100 with Wide Residual Networks (WideResNet 40-4): On CIFAR-100 using a 40-layer Wide Residual Network with a widening factor of 4 (using a Conv-Activation-Conv-Activation-BatchNorm residual block structure, dropout keep probability 0.70.7, trained for 50 epochs with Nesterov momentum and cosine annealing learning rate schedule T0=50,η=0.1T_0 = 50, \eta = 0.1), the median test error rates across three runs are:

      • GELU: 20.74%20.74\%
      • ReLU: 21.77%21.77\% (the baseline 40-4 WideResNet with ReLU without the end-of-block BatchNorm achieves 22.89%22.89\%)
      • ELU: 22.98%22.98\%

    Across both benchmarks, GELU achieves faster training convergence and higher final test accuracy.

  6. Knowl 6 — Empirical Performance and Robustness on MNIST Classification

    empirical result

    On MNIST handwritten digit classification, fully connected networks utilizing GELU activations achieve lower log loss and higher robustness to input noise compared to ReLU and ELU activations:

    • Network Setup: Fully connected 8-layer networks, each hidden layer having 128 neurons, trained with the Adam optimizer for 50 epochs with batch size 128 and unit-norm row weight initialization. Learning rates were tuned over {10−3,10−4,10−5}\{10^{-3}, 10^{-4}, 10^{-5}\}.
    • Convergence and Dropout Compatibility: Across five runs, GELU achieves the lowest median training log loss both without dropout and with a dropout rate of 0.5 (keep rate 0.5), demonstrating that the deterministic expected gating of GELU interfaces effectively with stochastic dropout regularization.
    • Noise Robustness: When adding additive uniform noise Unif[−a,a]\text{Unif}[-a, a] to input images at increasing noise strengths a∈[0,3]a \in [0, 3], networks trained with GELU retain higher test set accuracy and lower test set log loss across all noise magnitudes compared to networks trained with ReLU or ELU.
  7. Knowl 7 — Empirical Performance on MNIST Deep Autoencoding

    empirical result

    In self-supervised deep autoencoding on MNIST, GELU activations achieve lower reconstruction error than ReLU and ELU activations across different learning rates:

    • Architecture and Setup: A 7-layer deep autoencoder with layer widths 1000-500-250-30-250-500-1000 trained on MNIST images with the Adam optimizer, batch size 64, using mean squared error (MSE) reconstruction loss.
    • Results: Across three runs, GELU achieves lower median test set reconstruction error across 250 training epochs than both ReLU and ELU when trained with learning rates of 10−310^{-3} and 10−410^{-4}. At a higher learning rate of 0.010.01, ELU diverged, while GELU and ReLU converged poorly; across stable learning rate regimes, GELU demonstrates higher optimization speed and lower final reconstruction error.
  8. Knowl 8 — Empirical Performance on NLP POS Tagging and Speech Recognition

    empirical result

    The GELU nonlinearity achieves lower error rates than ReLU and ELU on sequence modeling tasks in natural language processing and speech recognition:

    • Twitter Part-of-Speech Tagging: Evaluated on a 25-tag Twitter POS dataset (1,000 training, 327 validation, and 500 test tweets) using a 2-layer network (256 neurons per layer, dropout keep probability 0.8) fed with concatenated pretrained word vectors (for the target word and left/right neighbors) trained on 56 million tweets. With Adam optimization tuned over learning rates {10−3,10−4,10−5}\{10^{-3}, 10^{-4}, 10^{-5}\} across 5 runs, the median test errors are:

      • GELU: 12.57%12.57\%
      • ReLU: 12.67%12.67\%
      • ELU: 12.91%12.91\%
    • TIMIT Phone Recognition: Evaluated on TIMIT frame classification (680 speakers, 39 output phone classes) using a 5-layer fully connected network (2048 neurons per layer, dropout rate 0.5) taking 11 audio frames (26 MFCC, energy, and derivative features per frame). Optimized with Adam tuned over {10−3,10−4,10−5}\{10^{-3}, 10^{-4}, 10^{-5}\} across 5 runs, the median test errors evaluated at minimum validation error are:

      • GELU: 29.3%29.3\%
      • ReLU: 29.5%29.5\%
      • ELU: 29.6%29.6\%
  9. Knowl 9 — Convolutional Network Architecture for CIFAR-10 Benchmarking

    data/table

    The convolutional architecture used for evaluating activations on CIFAR-10 without data augmentation consists of 9 convolutional layers with batch normalization and interleaved dropout and pooling stages:

    Layer Type Channels Spatial Dimension (x×yx \times y)
    Raw RGB input 3 32×3232 \times 32
    ZCA whitening 3 32×3232 \times 32
    Gaussian noise (σ=0.15\sigma = 0.15) 3 32×3232 \times 32
    3×33 \times 3 conv + activation 96 32×3232 \times 32
    3×33 \times 3 conv + activation 96 32×3232 \times 32
    3×33 \times 3 conv + activation 96 32×3232 \times 32
    2×22 \times 2 max pool, stride 2 96 16×1616 \times 16
    Dropout (p=0.5p = 0.5) 96 16×1616 \times 16
    3×33 \times 3 conv + activation 192 16×1616 \times 16
    3×33 \times 3 conv + activation 192 16×1616 \times 16
    3×33 \times 3 conv + activation 192 16×1616 \times 16
    2×22 \times 2 max pool, stride 2 192 8×88 \times 8
    Dropout (p=0.5p = 0.5) 192 8×88 \times 8
    3×33 \times 3 conv + activation 192 6×66 \times 6
    1×11 \times 1 conv + activation 192 6×66 \times 6
    1×11 \times 1 conv + activation 192 6×66 \times 6
    Global average pool 192 1×11 \times 1
    Softmax output 10 1×11 \times 1

    This network structure allows comparing GELU, ReLU, and ELU activations under identical training conditions (Adam optimizer for 200 epochs with learning rate linear decay after epoch 100).

Coverage note — Appendix B, which documents the historical priority narrative regarding the naming of SiLU and Swish, was omitted as it represents historical commentary rather than a technical or scientific contribution.

References

  1. 1.Jimmy Ba and Brendan Frey. Adaptive dropout for training deep neural networks. In Neural Information Processing Systems, 2013.
  2. 2.Philip Bachman, Ouais Alsharif, and Doina Precup. Learning with pseudo-ensembles. In Neural Information Processing Systems, 2014.
  3. 3.Amit Choudhury. A simple approximation to the area under standard normal curve. In Mathematics and Statistics, 2014.
  4. 4.Djork-Arne Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network learning by exponential linear units (ELUs). In International Conference on Learning Representations, 2016.
  5. 5.Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu. Natural neural networks. In arXiv, 2015.
  6. 6.Kevin Gimpel, Nathan Schneider, Brendan O′Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith. Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments. Association for Computational Linguistics (ACL), 2011.
  7. 7.Dan Hendrycks and Kevin Gimpel. Adjusting for dropout variance in batch normalization and weight initialization. In arXiv, 2016.
  8. 8.John Hopfield. Neural networks and physical systems with emergent collective computational abilities. In Proceedings of the National Academy of Sciences of the USA, 1982.
  9. 9.Diederik Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. International Conference for Learning Representations, 2015.
  10. 10.Alex Krizhevsky. Learning Multiple Layers of Features from Tiny Images, 2009.
  11. 11.David Krueger, Tegan Maharaj, Janos Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke1, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron Courville, and Chris Pal. Zoneout: Regularizing RNNs by randomly preserving hidden activations. In Neural Information Processing Systems, 2016.
  12. 12.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with restarts. arXiv, 2016.
  13. 13.Andrew L. Maas, Awni Y. Hannun, , and Andrew Y. Ng. Rectifier nonlinearities improve neural network acoustic models. In International Conference on Machine Learning, 2013.
  14. 14.Warren S. McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. In Bulletin of Mathematical Biophysics, 1943.
  15. 15.Dmytro Mishkin and Jiri Matas. All you need is a good init. In International Conference on Learning Representations, 2016.
  16. 16.Abdelrahman Mohamed, George E. Dahl, and Geoffrey E. Hinton. Acoustic modeling using deep belief networks. In IEEE Transactions on Audio, Speech, and Language Processing, 2012.
  17. 17.Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted boltzmann machines. In International Conference on Machine Learning, 2010.
  18. 18.Olutobi Owoputi, Brendan O’Connor, Chris Dyer, Kevin Gimpel, Nathan Schneider, and Noah A. Smith. Improved part-of-speech tagging for online conversational text with word clusters. In North American Chapter of the Association for Computational Linguistics (NAACL), 2013.
  19. 19.Tim Salimans and Diederik P. Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Neural Information Processing Systems, 2016.
  20. 20.Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. In International Conference on Learning Representations, 2014.
  21. 21.Anish Shah, Sameer Shinde, Eashan Kadam, Hena Shah, and Sandip Shingade. Deep residual networks with exponential linear unit. In Vision Net, 2016.
  22. 22.Nitish Srivastava. Improving neural networks with dropout. In University of Toronto, 2013.
  23. 23.Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. In Journal of Machine Learning Research, 2014.
  24. 24.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. British Machine Vision Conference, 2016.

Citation

MLA
Hendrycks, D., and K. Gimpel. “Gaussian Error Linear Units (GELUs)”. arXiv, 2016, http://arxiv.org/abs/1606.08415v5.
APA
Hendrycks, D., & Gimpel, K. (2016). Gaussian Error Linear Units (GELUs). arXiv. http://arxiv.org/abs/1606.08415v5
Chicago
Hendrycks, D., and K. Gimpel. 2016. “Gaussian Error Linear Units (GELUs)”. arXiv. http://arxiv.org/abs/1606.08415v5.
Harvard
Hendrycks, D. and Gimpel, K. (2016) “Gaussian Error Linear Units (GELUs)”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1606.08415v5.
Vancouver
1. Hendrycks D, Gimpel K (2016) Gaussian Error Linear Units (GELUs). arXiv

BibTeX

@article{hendrycks2016gaussian,
  title = {Gaussian Error Linear Units (GELUs)},
  author = {Hendrycks, Dan and Gimpel, Kevin},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1606.08415v5},
  eprint = {1606.08415}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors