Maxout Networks

Ian J. GoodfellowDavid Warde-FarleyMehdi MirzaAaron CourvilleYoshua Bengio

article2013ICML2,270 citations

Introduces maxout networks, demonstrating how learning piecewise linear activation functions via a maximum operation over affine feature maps improves optimization and ensemble averaging when combined with dropout.

Listen

Deep neural networks are central to modern computer vision and pattern recognition, but training them effectively remains challenging. While dropout—a technique that randomly deactivates network units during training to simulate averaging an ensemble of sub-models—has become widely used, standard network architectures are not specifically designed to maximize its benefits. Conventional activation functions frequently suffer from mathematical inaccuracies during fast model averaging or become permanently inactive during training, which limits the depth and accuracy of the resulting models.

To address this limitation, the article introduces and evaluates "maxout," a simple feed-forward neural network unit specifically constructed to improve both model averaging accuracy and optimization speed when paired with dropout. Rather than treating dropout as an afterthought for arbitrary architectures, the article establishes the theoretical basis of maxout and tests whether designing networks specifically for dropout can achieve superior classification performance across diverse visual recognition benchmarks.

The authors proved mathematically that a maxout network with just two units is a universal approximator capable of modeling any continuous function. They then conducted extensive empirical experiments by pairing maxout layers with dropout across four major benchmark datasets: MNIST handwritten digits, CIFAR-10, CIFAR-100 object images, and Street View House Numbers (SVHN). These models were compared directly against established architectures, including conventional networks using rectified linear units and hyperbolic tangent activations, across various model sizes, depths, and optimization stress tests.

The findings demonstrate significant, consistent improvements across all benchmarks. First, maxout established new state-of-the-art accuracy on all four datasets, cutting test classification error to 0.45% on MNIST, 9.38% on CIFAR-10 with data augmentation (and 11.68% without augmentation, beating the previous 14.98% baseline), 38.57% on CIFAR-100, and 2.47% on SVHN. Second, mathematical and empirical tracking showed that dropout's fast prediction averaging is substantially more accurate in maxout networks than in traditional curved activation functions. Third, maxout dramatically improved optimization stability; unlike standard rectified units that frequently saturate at zero and permanently block gradient flow (saturating up to 60% of the time under dropout), maxout units never saturate and fully utilized over 99.9% of their learned filters. Finally, maxout maintained healthy gradient variance to the lowest network layers—producing 3.4 times greater first-layer gradient variance than rectified units—which enabled effective training of deeper and narrower networks.

These results imply that co-designing network activation functions alongside regularization methods like dropout yields substantial gains in performance and training efficiency. By eliminating saturated, dead neurons and retaining gradient flow throughout the architecture, engineering teams can train deeper models with lower risk of optimization failure and without requiring complex unsupervised pretraining. Moreover, because maxout achieves high representational power with fewer output units via cross-channel pooling, it offers an efficient pathway to improve classification accuracy without ballooning parameter counts to the extent required by standard rectifier networks.

Based on these findings, teams developing vision and pattern recognition pipelines should adopt maxout activation functions when regularizing models with dropout. Future engineering efforts should explore applying maxout to other complex domains beyond computer vision and design further neural architectures that explicitly optimize inexpensive model averaging techniques.

Confidence in these findings is high given consistent state-of-the-art results across four diverse, standard benchmarks and thorough optimization stress testing. However, decision-makers should note that the evaluation was bounded by image classification datasets, and certain configurations (such as hyperparameter settings on CIFAR-100) were transferred directly from CIFAR-10 rather than extensively cross-validated due to computational time limits. Additional tuning may be required when adapting maxout to non-visual data types or new domain constraints.

Cover for Maxout Networks

Abstract

We consider the problem of designing models to leverage a recently introduced approximate model averaging technique called dropout. We define a simple new model called maxout (so named because its output is the max of a set of inputs, and because it is a natural companion to dropout) designed to both facilitate optimization by dropout and improve the accuracy of dropout's fast approximate model averaging technique. We empirically verify that the model successfully accomplishes both of these tasks. We use maxout and dropout to demonstrate state of the art classification performance on four benchmark datasets: MNIST, CIFAR-10, CIFAR-100, and SVHN.

Table of Contents

  • 1 Introduction
  • 2 Review of dropout
  • 3 Description of maxout
  • 4 Maxout is a universal approximator
  • 5 Benchmark results
  • 5.1 MNIST
  • 5.2 CIFAR-10
  • 5.3 CIFAR-100
  • 5.4 Street View House Numbers
  • 6 Comparison to rectifiers
  • 7 Model averaging
  • 8 Optimization
  • 8.1 Optimization experiments
  • 8.2 Saturation
  • 8.3 Lower layer gradients and bagging
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — Maxout Unit and Hidden Layer Formulation

    model/method

    A maxout hidden layer is a feedforward layer defined by taking the maximum over a set of affine feature responses. Given an input vector x∈Rdx \in \mathbb{R}^d (which may be the network input vv or the hidden state vector of a preceding layer), the activation of the ii-th maxout unit hi(x)∈Rh_i(x) \in \mathbb{R} in a layer of mm units is defined as:

    hi(x)=max⁡j∈{1,…,k}zijh_i(x) = \max_{j \in \{1, \dots, k\}} z_{ij}

    where each affine component zij∈Rz_{ij} \in \mathbb{R} is computed as:

    zij=xTW…ij+bijz_{ij} = x^T W_{\dots ij} + b_{ij}

    Here, W∈Rd×m×kW \in \mathbb{R}^{d \times m \times k} is a learned weight tensor, b∈Rm×kb \in \mathbb{R}^{m \times k} is a learned bias matrix, and W…ij∈RdW_{\dots ij} \in \mathbb{R}^d denotes the parameter vector associated with the jj-th affine component of unit ii. The hyperparameter k∈N+k \in \mathbb{N}^+ denotes the number of affine pieces per unit.

    In convolutional architectures, a maxout feature map is formed by computing kk distinct spatial affine feature maps and taking the pointwise maximum across channels at each spatial coordinate.

    When trained with dropout, binary mask variables μ\mu are multiplied elementwise directly onto the input xx before matrix multiplication with WW; individual affine inputs zijz_{ij} to the max⁡\max operator are not dropped independently. A single maxout unit parameterized by kk affine functions acts as a continuous piecewise linear (PWL) approximator to an arbitrary convex function.

  2. Knowl 2 — Universal Approximation Theorem for Two-Unit Maxout Networks

    theoretical result

    Any continuous function f:C→Rf: C \to \mathbb{R} defined on a compact domain C⊂RnC \subset \mathbb{R}^n can be approximated arbitrarily well by a feedforward network containing just two maxout hidden units, provided that the number of affine components kk per unit is allowed to be arbitrarily large.

    The result is established from two foundational properties:

    1. Continuous Piecewise Linear (PWL) Representation: For any continuous PWL function g:Rn→Rg: \mathbb{R}^n \to \mathbb{R} comprising kk locally affine regions, there exist parameter vectors [W1j,b1j][W_{1j}, b_{1j}] and [W2j,b2j][W_{2j}, b_{2j}] for j∈{1,…,k}j \in \{1, \dots, k\} such that g(v)g(v) is expressed as the difference of two convex continuous PWL functions: g(v)=h1(v)−h2(v)=max⁡j∈{1,…,k}(vTW1j+b1j)−max⁡j∈{1,…,k}(vTW2j+b2j)g(v) = h_1(v) - h_2(v) = \max_{j \in \{1, \dots, k\}} (v^T W_{1j} + b_{1j}) - \max_{j \in \{1, \dots, k\}} (v^T W_{2j} + b_{2j})

    2. Stone-Weierstrass Approximation: By the Stone-Weierstrass theorem, any continuous function ff on a compact set C⊂RnC \subset \mathbb{R}^n can be uniformly approximated to within an arbitrary tolerance ϵ>0\epsilon > 0 by a continuous PWL function gg, such that sup⁡v∈C∣f(v)−g(v)∣<ϵ\sup_{v \in C} |f(v) - g(v)| < \epsilon.

    A network with two maxout hidden units h1(v)h_1(v) and h2(v)h_2(v) connected to a linear output layer with weights W1=1W_1 = 1 and W2=−1W_2 = -1 computes g(v)=h1(v)−h2(v)g(v) = h_1(v) - h_2(v). As ϵ→0\epsilon \to 0, k→∞k \to \infty, allowing the two-unit maxout network to approximate any continuous function on CC.

  3. Knowl 3 — Model Averaging with Dropout in Locally Linear Maxout Networks

    theoretical result

    For a single-layer softmax model p(y∣v;θ)=softmax(vTW+b)p(y \mid v; \theta) = \text{softmax}(v^T W + b), computing the geometric mean over all 2∣v∣2^{|v|} sub-models induced by binary dropout masks μ\mu and renormalizing produces the exact predictive distribution:

    pensemble(y∣v)=softmax(vTW2+b)p_{\text{ensemble}}(y \mid v) = \text{softmax}\left(v^T \frac{W}{2} + b\right)

    This exactness of the weight halving rule (W/2W/2) extends to multi-layer feedforward networks that are entirely linear.

    Maxout networks are locally linear almost everywhere. Dropout training encourages maxout units to form large linear regions around training examples because each sub-model must produce consistent predictions regardless of which inputs are masked out. Consequently, changing the dropout mask μ\mu rarely changes the index j∗=arg⁡max⁡jzijj^* = \arg\max_j z_{ij} of the active affine filter in each maxout unit for a given training point. Because the network remains locally linear across the set of inputs visited under dropout sampling, the W/2W/2 weight-scaling heuristic provides a close mathematical approximation to the geometric mean ensemble prediction.

    Empirical evaluation confirms that as the number of sampled dropout sub-models increases, the geometric mean prediction converges toward the W/2W/2 prediction, achieving a significantly smaller Kullback-Leibler (KL) divergence in maxout networks than in hyperbolic tangent (tanh⁡\tanh) networks.

  4. Knowl 4 — Filter Utilization and Optimization Advantage of Maxout over Rectifiers under Dropout

    empirical result

    Maxout units max⁡jzij\max_{j} z_{ij} differ from max pooling over rectified linear units (ReLUs) max⁡(0,zi1,…,zik)\max(0, z_{i1}, \dots, z_{ik}) solely by omitting the constant 0 in the maximization. Under dropout training with large learning rates, this distinction is critical for parameter optimization:

    1. Unit Saturation and Reactivation: In standard ReLU networks trained with dropout, units transition from active to zero-activation more frequently than the reverse, causing unit saturation at 0 to rise to approximately 60%60\% (compared to less than 5%5\% under standard SGD). Because the constant 0 produces zero gradient, inactive ReLU units become locked and cannot easily receive updates. In maxout units, activations transition between positive and negative values at equal rates, and gradient continuously flows through whichever affine component achieves the maximum, keeping all parameters steerable.

    2. Filter Inactivation on MNIST: In a 2-hidden-layer MLP (1200 filters per layer pooled in groups of k=5k=5) trained with dropout on MNIST, including a constant 0 in the pooling caused 17.6%17.6\% of first-layer filters and 39.2%39.2\% of second-layer filters to go completely unused (never achieving the maximum on any training example), worsening validation error from 1.04%1.04\% to over 1.2%1.2\%. Maxout utilized 2398 of the 2400 total filters (99.9%99.9\%), tuning virtually every parameter.

    3. Optimization Stress Test on SVHN: On a 600,000-example SVHN dataset with a small 2-layer convolutional network (k=2k=2, 16 kernels), rectified units stalled at a training error of 7.3%7.3\%, whereas maxout achieved 5.1%5.1\% training error.

  5. Knowl 5 — Propagation of Dropout-Induced Gradient Variance Across Network Depth

    empirical result

    For dropout training to function as an ensemble bagging procedure rather than standard stochastic gradient descent (SGD), the backpropagated gradient for a fixed data point must exhibit high variance with respect to the sampled dropout mask μ\mu.

    In experiments comparing two-hidden-layer MLPs on MNIST trained with dropout:

    • At the output layer, the variance of the gradient across different dropout masks for fixed data was 1.4×1.4\times larger for maxout networks than for rectified linear networks.
    • At the first hidden layer (closest to the input), the gradient variance across dropout masks was 3.4×3.4\times larger for maxout networks than for rectifier networks.

    Because rectifier networks lose gradient flow through saturated (zeroed) units, the dropout stochasticity attenuates toward the bottom of the network, causing lower layers to behave like standard SGD. Maxout maintains gradient flow through every unit, propagating dropout variance down to the lowest layers and ensuring parameter sharing under bagging operates throughout the entire architecture.

  6. Knowl 6 — Depth Scaling in Narrow Maxout Networks vs. Pooled Rectifiers

    empirical result

    When evaluating network optimization across depths from 1 to 7 layers on MNIST using narrow layers (80 units per layer, pooling factor k=5k=5):

    • Maxout networks maintain stable optimization, with training and test classification errors degrading gracefully as depth increases up to 7 layers (test error remaining below 3%3\%).
    • Pooled rectifier networks exhibit severe optimization failure as depth increases, showing noticeable error increases at 6 layers and a dramatic degradation at 7 layers (test error escalating above 14%14\% and training error rising sharply).

    The continuous non-zero gradient flow of maxout units mitigates the vanishing gradient and unit death phenomena that impede deep narrow ReLU architectures under dropout.

  7. Knowl 7 — Benchmark Classification Error on Permutation Invariant and General MNIST

    data/table

    Maxout networks trained with dropout set the state of the art on both the permutation invariant and general MNIST classification benchmarks without unsupervised pretraining.

    Method Test Error
    Permutation Invariant MNIST
    Rectifier MLP + dropout (Srivastava, 2013) 1.05%
    DBM (Salakhutdinov Hinton, 2009) 0.95%
    Maxout MLP + dropout 0.94%
    MP-DBM (Goodfellow et al., 2013) 0.91%
    Deep Convex Network (Yu Deng, 2011) 0.83%
    Manifold Tangent Classifier (Rifai et al., 2011) 0.81%
    DBM + dropout (Hinton et al., 2012) 0.79%
    General MNIST (No Data Augmentation)
    2-layer CNN + 2-layer NN (Jarrett et al., 2009) 0.53%
    Stochastic pooling (Zeiler Fergus, 2013) 0.47%
    Conv. maxout + dropout 0.45%

    On permutation invariant MNIST, a fully connected model with two maxout layers regularized with dropout and max-norm weight constraints achieved 0.94%0.94\% test error, outperforming all purely supervised models. On general MNIST (without data augmentation), a network with three convolutional maxout layers followed by spatial pooling and a maxout dense layer achieved 0.45%0.45\% test error.

  8. Knowl 8 — Benchmark Classification Error on CIFAR-10 and CIFAR-100

    data/table

    Maxout convolutional networks regularized with dropout established state-of-the-art error rates on CIFAR-10 and CIFAR-100 using global contrast normalization (GCN) and ZCA whitening.

    Method Test Error
    CIFAR-10
    Stochastic pooling (Zeiler Fergus, 2013) 15.13%
    CNN + Spearmint (Snoek et al., 2012) 14.98%
    Conv. maxout + dropout (no aug) 11.68%
    CNN + Spearmint + data augmentation (Snoek et al., 2012) 9.50%
    Conv. maxout + dropout + data augmentation 9.38%
    CIFAR-100
    Learned pooling (Malinowski Fritz, 2013) 43.71%
    Stochastic pooling (Zeiler Fergus, 2013) 42.51%
    Conv. maxout + dropout 38.57%

    The architecture consisted of three convolutional maxout layers, one fully connected maxout layer, and a softmax output layer. On CIFAR-10 without data augmentation, maxout achieved 11.68%11.68\% error (13.2%13.2\% without validation set retraining), improving prior methods by over 3 percentage points. With horizontal flips and translations, it achieved 9.38%9.38\%. On CIFAR-100, transferring CIFAR-10 hyperparameters yielded 38.57%38.57\% test error (41.48%41.48\% without validation set retraining).

  9. Knowl 9 — Benchmark Classification Error on Street View House Numbers (SVHN)

    data/table

    On the Street View House Numbers (SVHN) digit classification task (format 2, 32×3232 \times 32 cropped color digit images), a network of three convolutional maxout layers, one dense maxout layer, and a softmax layer trained with dropout achieved state-of-the-art performance.

    Method Test Error
    Sermanet et al. (2012a) 4.90%
    Stochastic pooling (Zeiler Fergus, 2013) 2.80%
    Rectifiers + dropout (Srivastava, 2013) 2.78%
    Rectifiers + dropout + synthetic translation (Srivastava, 2013) 2.68%
    Conv. maxout + dropout 2.47%

    Using local contrast normalization (LCN) preprocessing and training on the combined train and extra sets without retraining on the validation set, convolutional maxout achieved a test error rate of 2.47%2.47\%, improving upon both stochastic pooling and rectified linear networks with synthetic translations.

  10. Knowl 10 — Cross-Channel Pooling Parameter Efficiency in Maxout vs. Rectifiers

    empirical result

    Controlled cross-validation experiments comparing maxout and rectified linear networks on CIFAR-10 under identical preprocessing (global contrast normalization and ZCA whitening) and hyperparameter schedules revealed distinct structural behaviors:

    1. Cross-Channel Pooling Impact: Applying cross-channel pooling to rectifier networks does not yield substantial performance gains; rectifier performance is primarily governed by the number of output units rather than the total number of intermediate filters.
    2. Parameter Scaling: In maxout networks, performance scales with the total number of affine filters pooled within each unit. For a rectifier network to match the generalization performance of a maxout network, it must omit cross-channel pooling and use kk times as many hidden units. Because increasing layer ii output units by a factor of kk expands the input dimensionality to layer i+1i+1 by kk, the equivalent rectifier network requires approximately kk times more parameters and substantially greater memory and runtime than the maxout model.

Coverage note — No substantial contributed material was omitted; all core theoretical definitions, the universal approximation theorem, dropout model averaging analysis, optimization investigations, and benchmark empirical results across MNIST, CIFAR-10, CIFAR-100, and SVHN are included.

References

  1. 1.Bastien, Frederic, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian, Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua. Theano: new features and speed improvements. Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop, 2012.
  2. 2.Bergstra, James, Breuleux, Olivier, Bastien, Frederic, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua. Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy), June 2010. Oral Presentation.
  3. 3.Breiman, Leo. Bagging predictors. Machine Learning, 24 (2):123–140, 1994.
  4. 4.Ciresan, D. C., Meier, U., Gambardella, L. M., and Schmidhuber, J. Deep big simple neural nets for handwritten digit recognition. Neural Computation, 22:1–14, 2010.
  5. 5.Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua. Deep sparse rectifier neural networks. In JMLR W&CP: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS 2011), April 2011.
  6. 6.Goodfellow, Ian J., Courville, Aaron, and Bengio, Yoshua. Joint training of deep Boltzmann machines for classification. In International Conference on Learning Representations: Workshops Track, 2013.
  7. 7.Hahnloser, Richard H. R. On the piecewise analysis of networks of linear threshold neurons. Neural Networks, 11(4):691–697, 1998.
  8. 8.Hinton, Geoffrey E., Srivastava, Nitish, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan. Improving neural networks by preventing co-adaptation of feature detectors. Technical report, arXiv:1207.0580, 2012.
  9. 9.Jarrett, Kevin, Kavukcuoglu, Koray, Ranzato, Marc’Aurelio, and LeCun, Yann. What is the best multi-stage architecture for object recognition? In Proc. International Conference on Computer Vision (ICCV’09), pp. 2146–2153. IEEE, 2009.
  10. 10.Krizhevsky, Alex and Hinton, Geoffrey. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  11. 11.Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 (NIPS’2012). 2012.
  12. 12.LeCun, Yann, Bottou, Leon, Bengio, Yoshua, and Haffner, Patrick. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, November 1998.
  13. 13.Malinowski, Mateusz and Fritz, Mario. Learnable pooling regions for image classification. In International Conference on Learning Representations: Workshop track, 2013.
  14. 14.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. Deep Learning and Unsupervised Feature Learning Workshop, NIPS, 2011.
  15. 15.Rifai, Salah, Dauphin, Yann, Vincent, Pascal, Bengio, Yoshua, and Muller, Xavier. The manifold tangent classifier. In NIPS’2011, 2011. Student paper award.
  16. 16.Salakhutdinov, R. and Hinton, G.E. Deep Boltzmann machines. In Proceedings of the Twelfth International Conference on Artificial Intelligence and Statistics (AISTATS 2009), volume 8, 2009.
  17. 17.Salinas, E. and Abbott, L. F. A model of multiplicative neural responses in parietal cortex. Proc Natl Acad Sci U S A, 93(21):11956–11961, October 1996.
  18. 18.Sermanet, Pierre, Chintala, Soumith, and LeCun, Yann. Convolutional neural networks applied to house numbers digit classification. CoRR, abs/1204.3968, 2012a.
  19. 19.Sermanet, Pierre, Chintala, Soumith, and LeCun, Yann. Convolutional neural networks applied to house numbers digit classification. In International Conference on Pattern Recognition (ICPR 2012), 2012b.
  20. 20.Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan Prescott. Practical bayesian optimization of machine learning algorithms. In Neural Information Processing Systems, 2012.
  21. 21.Srebro, Nathan and Shraibman, Adi. Rank, trace-norm and max-norm. In Proceedings of the 18th Annual Conference on Learning Theory, pp. 545–560. Springer-Verlag, 2005.
  22. 22.Srivastava, Nitish. Improving neural networks with dropout. Master’s thesis, U. Toronto, 2013.
  23. 23.Wang, Shuning. General constructive representations for continuous piecewise-linear functions. IEEE Trans. Circuits Systems, 51(9):1889–1896, 2004.
  24. 24.Yu, Dong and Deng, Li. Deep convex net: A scalable architecture for speech pattern classification. In INTERSPEECH, pp. 2285–2288, 2011.
  25. 25.Zeiler, Matthew D. and Fergus, Rob. Stochastic pooling for regularization of deep convolutional neural networks. In International Conference on Learning Representations, 2013.

Citation

MLA
Goodfellow, I. J., et al. “Maxout Networks”. JMLR WCP 28 (3): 1319-1327, 2013, 2013, http://arxiv.org/abs/1302.4389v4.
APA
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., & Bengio, Y. (2013). Maxout Networks. JMLR WCP 28 (3): 1319-1327, 2013. http://arxiv.org/abs/1302.4389v4
Chicago
Goodfellow, I. J., D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio. 2013. “Maxout Networks”. JMLR WCP 28 (3): 1319-1327, 2013. http://arxiv.org/abs/1302.4389v4.
Harvard
Goodfellow, I.J. et al. (2013) “Maxout Networks”, JMLR WCP 28 (3): 1319-1327, 2013 [Preprint]. Available at: http://arxiv.org/abs/1302.4389v4.
Vancouver
1. Goodfellow IJ, Warde-Farley D, Mirza M, Courville A, Bengio Y (2013) Maxout Networks. JMLR WCP 28 (3): 1319-1327, 2013

BibTeX

@article{goodfellow2013maxout,
  title = {Maxout Networks},
  author = {Goodfellow, Ian J. and Warde-Farley, David and Mirza, Mehdi and Courville, Aaron and Bengio, Yoshua},
  year = {2013},
  journal = {JMLR WCP 28 (3): 1319-1327, 2013},
  url = {http://arxiv.org/abs/1302.4389v4},
  eprint = {1302.4389}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission