Maxout Networks
Ian J. GoodfellowDavid Warde-FarleyMehdi MirzaAaron CourvilleYoshua Bengio
Introduces maxout networks, demonstrating how learning piecewise linear activation functions via a maximum operation over affine feature maps improves optimization and ensemble averaging when combined with dropout.
Deep neural networks are central to modern computer vision and pattern recognition, but training them effectively remains challenging. While dropout—a technique that randomly deactivates network units during training to simulate averaging an ensemble of sub-models—has become widely used, standard network architectures are not specifically designed to maximize its benefits. Conventional activation functions frequently suffer from mathematical inaccuracies during fast model averaging or become permanently inactive during training, which limits the depth and accuracy of the resulting models.
To address this limitation, the article introduces and evaluates "maxout," a simple feed-forward neural network unit specifically constructed to improve both model averaging accuracy and optimization speed when paired with dropout. Rather than treating dropout as an afterthought for arbitrary architectures, the article establishes the theoretical basis of maxout and tests whether designing networks specifically for dropout can achieve superior classification performance across diverse visual recognition benchmarks.
The authors proved mathematically that a maxout network with just two units is a universal approximator capable of modeling any continuous function. They then conducted extensive empirical experiments by pairing maxout layers with dropout across four major benchmark datasets: MNIST handwritten digits, CIFAR-10, CIFAR-100 object images, and Street View House Numbers (SVHN). These models were compared directly against established architectures, including conventional networks using rectified linear units and hyperbolic tangent activations, across various model sizes, depths, and optimization stress tests.
The findings demonstrate significant, consistent improvements across all benchmarks. First, maxout established new state-of-the-art accuracy on all four datasets, cutting test classification error to 0.45% on MNIST, 9.38% on CIFAR-10 with data augmentation (and 11.68% without augmentation, beating the previous 14.98% baseline), 38.57% on CIFAR-100, and 2.47% on SVHN. Second, mathematical and empirical tracking showed that dropout's fast prediction averaging is substantially more accurate in maxout networks than in traditional curved activation functions. Third, maxout dramatically improved optimization stability; unlike standard rectified units that frequently saturate at zero and permanently block gradient flow (saturating up to 60% of the time under dropout), maxout units never saturate and fully utilized over 99.9% of their learned filters. Finally, maxout maintained healthy gradient variance to the lowest network layers—producing 3.4 times greater first-layer gradient variance than rectified units—which enabled effective training of deeper and narrower networks.
These results imply that co-designing network activation functions alongside regularization methods like dropout yields substantial gains in performance and training efficiency. By eliminating saturated, dead neurons and retaining gradient flow throughout the architecture, engineering teams can train deeper models with lower risk of optimization failure and without requiring complex unsupervised pretraining. Moreover, because maxout achieves high representational power with fewer output units via cross-channel pooling, it offers an efficient pathway to improve classification accuracy without ballooning parameter counts to the extent required by standard rectifier networks.
Based on these findings, teams developing vision and pattern recognition pipelines should adopt maxout activation functions when regularizing models with dropout. Future engineering efforts should explore applying maxout to other complex domains beyond computer vision and design further neural architectures that explicitly optimize inexpensive model averaging techniques.
Confidence in these findings is high given consistent state-of-the-art results across four diverse, standard benchmarks and thorough optimization stress testing. However, decision-makers should note that the evaluation was bounded by image classification datasets, and certain configurations (such as hyperparameter settings on CIFAR-100) were transferred directly from CIFAR-10 rather than extensively cross-validated due to computational time limits. Additional tuning may be required when adapting maxout to non-visual data types or new domain constraints.
- Paper: Improving neural networks by preventing co-adaptation of feature detectors, Geoffrey E. Hinton et al. (2012). Read this foundational paper on dropout first to understand the core regularization mechanism that maxout networks were specifically designed to optimize and complement.
- Paper: Understanding the difficulty of training deep feedforward neural networks, Xavier Glorot et al. (2010). Reviewing this analysis of training deep neural networks provides essential context on initialization and activation saturation before exploring maxout units.
- Paper: Regularization of Neural Networks using DropConnect, Li Wan et al. (2013). This paper extends the regularization principles introduced by dropout and maxout by randomly dropping network weights instead of activations.
- Paper: Dropout: a simple way to prevent neural networks from overfitting, Nitish Srivastava et al. (2014). This definitive study builds directly upon maxout and dropout to demonstrate comprehensive state-of-the-art performance across a wide array of benchmark datasets.
