Gaussian Error Linear Units (GELUs)
Dan HendrycksKevin Gimpel
Introduces the Gaussian Error Linear Unit (GELU), a smooth activation function that weights inputs by their magnitude rather than gating them by sign, establishing a high-performing alternative to ReLU across vision, language, and speech tasks.
Deep neural network performance depends heavily on the choice of activation function, which determines how individual artificial neurons transform incoming signals to model complex patterns. Historically, network designs have relied on deterministic gating mechanisms like the Rectified Linear Unit (ReLU) or Exponential Linear Unit (ELU), while treating stochastic regularization techniques, such as dropout, as entirely separate architectural decisions. This separation leaves open the question of whether neural network activations can be made more effective by mathematically integrating probabilistic regularizers directly into the activation function itself.
The article sets out to introduce and evaluate the Gaussian Error Linear Unit (GELU), a novel nonlinearity designed to merge the probabilistic benefits of dropout-style regularizers with standard neuron activation. It aims to demonstrate that GELU outperforms standard ReLUs and ELUs across diverse core machine learning domains, including computer vision, natural language processing, and speech recognition.
The authors conducted a comprehensive set of empirical benchmark experiments comparing standard GELUs against ReLUs and ELUs across multiple architectures without introducing extra hyperparameters. The evaluation spanned image classification and self-supervised autoencoding on MNIST, part-of-speech tagging on social media text from Twitter, phone frame classification on the TIMIT acoustic speech dataset, and image classification on CIFAR-10 and CIFAR-100 using shallow convolutional and deep wide residual networks. Models were trained across varying learning rates and tested with momentum-based optimizers and dropout settings to ensure robust, fair comparisons.
The empirical findings consistently favored the new activation function. On deep image classification using a 40-layer Wide Residual Network on CIFAR-100, GELU achieved an error rate of 20.74%, outperforming ReLU at 21.77% and ELU at 22.98%. In standard CIFAR-10 image classification, GELU achieved a median error rate of 7.89%, surpassing ReLU (8.16%) and ELU (8.41%). In MNIST autoencoding, GELU accommodated multiple learning rates smoothly and yielded significantly lower reconstruction error than both alternatives. On natural language part-of-speech tagging and TIMIT speech recognition, GELU similarly secured the lowest median test errors (12.57% and 29.3%, respectively). Additionally, GELU exhibited superior resilience when artificial noise was added to input data during image classification.
These results indicate that continuous curvature and probabilistic weighting allow neural networks to fit complex functions and converge faster than traditional threshold-based activations. In practical terms, adopting GELU provides measurable accuracy gains without requiring additional hyperparameter tuning, thereby improving model performance without inflating computational engineering complexity. Because GELU acts as a smooth approximation to standard gating functions, it serves as an effective drop-in replacement across modern architectures.
Organizations developing or deploying deep learning systems should consider adopting GELU or its fast analytic approximations as the default nonlinearity in new and existing architectures. When deploying GELU, engineering teams should pair it with momentum-based optimization algorithms and use accurate approximations to ensure numerical stability and computational efficiency.
While the results demonstrate consistent advantages across multiple benchmarks, the experiments in the article rely on standard academic datasets and specific baseline architectures. Stakeholders should validate performance on proprietary internal datasets and evaluate runtime inference overhead when selecting between exact mathematical formulations and fast mathematical approximations.
- Paper: Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs), Djork-Arné Clevert et al. (2015). This paper introduces the Exponential Linear Unit (ELU), establishing the core comparative baseline and theoretical motivation for smooth, negative-valued linear units that GELU aims to improve upon.
- Paper: Deep Sparse Rectifier Neural Networks, Xavier Glorot et al. (2011). This work establishes the foundational advantages of non-saturating rectified linear activations in deep networks, providing the essential context for GELU's probabilistic gating formulation.
- Paper: Empirical Evaluation of Rectified Activations in Convolutional Network, Bing Xu et al. (2015). This empirical study examines non-zero slope variants of rectified activations, directly contextualizing why smooth non-zero regimes in negative inputs benefit neural representations.
- Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). This paper explores combining stochastic gating and smooth gradient pathways, which conceptually underpins GELU's view of deterministic activation as expected stochastic regularization.
- Paper: Maxout Networks, Ian J. Goodfellow et al. (2013). This work presents an alternative activation unit designed specifically to interact synergistically with stochastic dropout regularization.
- Paper: GLU Variants Improve Transformer, Noam Shazeer. This work extends GELU into gated linear unit formulations (such as GEGLU) within Transformer feed-forward networks, demonstrating substantial downstream empirical gains.
- Paper: Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning, Stefan Elfwing et al. (2017). This paper introduces and evaluates sigmoid-weighted linear units (SiLU/Swish), offering a closely related self-gated continuous nonlinearity evaluated across reinforcement learning domains.
- Paper: Low-dimensional topology of deep neural networks, Junyu Ren et al. (2026). This theoretical study demonstrates why non-monotonic activations like GELU possess superior topological expressivity over monotonic units in unlinking data manifolds.
- Paper: Self-Normalizing Neural Networks, Günter Klambauer et al. (2017). This paper follows the trajectory of modern activation design by introducing Scaled Exponential Linear Units (SELU) that theoretically induce self-normalizing dynamics across deep networks.
