ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation

Adam PaszkeAbhishek ChaurasiaSangpil KimEugenio Culurciello

article2016arXiv2,400 citations

Proposes ENet, a lightweight deep neural network architecture that enables real-time semantic segmentation on resource-constrained embedded devices by reducing computational cost and parameter size by over an order of magnitude without sacrificing accuracy.

Listen

Real-time visual scene understanding—labeling every pixel in an image with its corresponding object class—is critical for autonomous vehicles, augmented reality headsets, and mobile robotics. However, existing high-performing deep learning models are computationally prohibitive. Because they rely on massive network architectures that require billions of operations, they cannot achieve the low-latency processing necessary for battery-powered or embedded mobile devices.

This article introduces and evaluates ENet (Efficient Neural Network), an extremely compact deep neural network architecture designed from the ground up for low latency and high accuracy in real-time visual segmentation tasks.

The authors conducted empirical performance and accuracy evaluations across three established benchmark datasets: Cityscapes and CamVid for autonomous driving environments, and SUN RGB-D for indoor scenes. ENet's efficiency and latency were measured against the standard baseline model, SegNet, on both an embedded mobile board (NVIDIA Jetson TX1) and a high-end desktop graphics processor (NVIDIA Titan X). The architecture achieves efficiency through several deliberate design choices, such as compressing resolution in the earliest layers, using an asymmetric encoder-decoder structure with a minimal decoder, and employing factorized and dilated convolutions to preserve contextual information without adding computational bloat.

The evaluation produced several notable findings. First, ENet reduces hardware requirements dramatically: it uses 75 times fewer floating-point operations (3.83 GFLOPs versus 286.03 GFLOPs) and 79 times fewer parameters (0.37 million versus 29.46 million), yielding a compact 0.7 MB footprint that fits entirely into fast on-chip processor memory. Second, ENet is up to 18 to 20 times faster than the baseline on embedded hardware, delivering 14.6 frames per second at 640x360 resolution where the baseline achieved less than 1 frame per second. Third, despite its small size, ENet matched or surpassed baseline accuracy on road scenes, achieving higher intersection-over-union scores on Cityscapes (58.3% versus 56.1%) and outperforming prior models on difficult, smaller object classes such as signs, pedestrians, and cyclists.

These results demonstrate that organizations deploying computer vision do not need to choose between accuracy and resource efficiency. ENet eliminates the need for expensive, specialized model-compression workflows or heavy supplementary post-processing algorithms. For embedded systems, it enables real-time perception on low-cost, low-power hardware, directly reducing hardware bills of materials and energy consumption. On high-end data center hardware, its high frame rate (over 135 frames per second) offers substantial cost and runtime reductions for large-scale video processing.

Organizations developing edge-device vision applications should evaluate ENet as a primary architecture or baseline. Furthermore, software teams should investigate software optimization techniques such as kernel fusion in underlying machine learning libraries. Combining multiple mathematical operations into single execution steps would alleviate memory transaction bottlenecks and further enhance execution speed.

The findings are subject to a few boundaries. ENet's indoor segmentation performance on the SUN RGB-D dataset lagged behind the baseline in global accuracy (59.5% versus 70.3%), indicating that highly diverse, cluttered indoor scenes remain more challenging for compact networks than structured road environments. Additionally, because the architecture breaks operations down into numerous small calculations, software overhead from graphics processor function calls currently accounts for a notable fraction of runtime. Nevertheless, there is high confidence that ENet provides a viable, state-of-the-art solution for real-time mobile scene segmentation.

arXiv: 1606.02147e-lab/ENet-training
Cover for ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation

Abstract

The ability to perform pixel-wise semantic segmentation in real-time is of paramount importance in mobile applications. Recent deep neural networks aimed at this task have the disadvantage of requiring a large number of floating point operations and have long run-times that hinder their usability. In this paper, we propose a novel deep neural network architecture named ENet (efficient neural network), created specifically for tasks requiring low latency operation. ENet is up to 18×\times faster, requires 75×\times less FLOPs, has 79×\times less parameters, and provides similar or better accuracy to existing models. We have tested it on CamVid, Cityscapes and SUN datasets and report on comparisons with existing state-of-the-art methods, and the trade-offs between accuracy and processing time of a network. We present performance measurements of the proposed architecture on embedded systems and suggest possible software improvements that could make ENet even faster.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Network architecture
  • 4 Design choices
  • 5 Results
  • 5.1 Performance Analysis
  • 5.2 Benchmarks
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — ENet Architecture and Bottleneck Module Structure

    model/method

    ENet is an efficient, compact encoder-decoder convolutional neural network designed for real-time pixel-wise semantic segmentation. The overall architecture is organized into five sequential stages followed by a final full convolution:

    • Initial Stage: A single initial block outputting 16 feature maps.
    • Stage 1 (Encoder): 1 downsampling bottleneck module followed by 4 regular bottleneck modules (64 feature maps).
    • Stage 2 (Encoder): 1 downsampling bottleneck module followed by 8 bottleneck modules that interleave regular, dilated (2,4,8,162, 4, 8, 16), and asymmetric (5×15\times 1 and 1×51\times 5) convolutions (128 feature maps).
    • Stage 3 (Encoder): An exact repetition of Stage 2's structure except that the initial downsampling block is omitted (128 feature maps).
    • Stage 4 (Decoder): 1 upsampling bottleneck module followed by 2 regular bottleneck modules (64 feature maps).
    • Stage 5 (Decoder): 1 upsampling bottleneck module followed by 1 regular bottleneck module (16 feature maps).
    • Fullconv: A single full deconvolution layer that maps 16 feature maps to CC output channels corresponding to the target semantic classes.

    Each bottleneck module contains a main branch (residual connection) and an extension branch containing three convolutional layers:

    1. A 1×11\times 1 projection convolution that reduces channel dimensionality.
    2. A main convolutional layer (3×33\times 3 standard convolution, 3×33\times 3 dilated convolution, 3×33\times 3 deconvolution/full convolution, or decomposed asymmetric 5×15\times 1 and 1×51\times 5 convolutions).
    3. A 1×11\times 1 expansion convolution that restores the channel dimensionality.

    Batch Normalization and Parametric Rectified Linear Units (PReLU) are placed between all convolutions. Spatial Dropout is applied at the end of the extension branch before merging back with the main branch via element-wise addition. Bias terms are omitted in all convolutional projections to reduce CUDA kernel launch overhead and unnecessary memory round-trips.

  2. Knowl 2 — ENet Initial Block and Resolution Scaling Modules

    model/method

    ENet employs specialized modules for early feature extraction and spatial resolution scaling:

    • Initial Block: To reduce the heavy computational cost of processing high-resolution inputs without creating a representational bottleneck, the input image (3 channels) is processed simultaneously by two parallel operations: a 3×33\times 3 convolution with stride 2 (13 filters) and a non-overlapping 2×22\times 2 max pooling operation with stride 2. The resulting feature representations are concatenated along the channel dimension to produce 16 feature maps at half the spatial resolution.

    • Downsampling Bottleneck Module: When reducing spatial resolution in the encoder, the main branch applies non-overlapping 2×22\times 2 max pooling with stride 2, and the resulting activations are zero-padded to match the channel count of the extension branch. The max pooling indices are saved for subsequent unpooling in the decoder. In the extension branch, the initial 1×11\times 1 projection is replaced by a 2×22\times 2 convolution with stride 2 in both dimensions, ensuring all input pixels are sampled without discarding features.

    • Upsampling Bottleneck Module: In the decoder, max pooling in the main branch is replaced by max unpooling using the saved indices from the corresponding downsampling stage, and padding is replaced by a spatial convolution without bias. The extension branch uses a 3×33\times 3 full convolution (fractionally strided deconvolution). The final upsampling module of the network does not use pooling indices because the initial block operated on 3 RGB channels whereas the final output requires CC class channels.

  3. Knowl 3 — Hardware Efficiency and Inference Speed Comparison

    data/table

    ENet achieves real-time inference on embedded hardware (NVIDIA Jetson TX1) and high-throughput execution on desktop GPUs (NVIDIA Titan X), outperforming SegNet by large margins across latency, operations, and memory footprint.

    Model GFLOPs (3×640×3603\times 640\times 360) Parameters Model Size (FP16) TX1 FPS (640×360640\times 360) Titan X FPS (640×360640\times 360) Titan X FPS (1920×10801920\times 1080)
    SegNet 286.03 29.46M 56.2 MB 0.8 14.6 1.6
    ENet 3.83 0.37M 0.7 MB 14.6 135.4 21.6

    ENet requires 75×75\times fewer FLOPs, has 79×79\times fewer parameters, and achieves an FP16 footprint of 0.7 MB (enabling the entire model to fit within embedded on-chip memory). On an NVIDIA Jetson TX1, ENet achieves 14.6 frames per second (fps) at 640×360640\times 360 resolution (surpassing the 10 fps real-time requirement for automotive scene parsing, where SegNet achieves 0.8 fps). On an NVIDIA Titan X, ENet processes 640×360640\times 360 inputs at 135.4 fps and full HD 1920×10801920\times 1080 inputs at 21.6 fps.

  4. Knowl 4 — Semantic Segmentation Benchmarks on Cityscapes, CamVid, and SUN RGB-D

    empirical result

    ENet was evaluated against SegNet across three semantic segmentation benchmarks covering driving and indoor scenes:

    1. Cityscapes Test Set (19 classes, urban road scenes):

      • ENet achieves 58.3% Class IoU, 34.4% Class iIoU (instance-level IoU), 80.4% Category IoU, and 64.0% Category iIoU.
      • In comparison, SegNet achieves 56.1% Class IoU, 34.2% Class iIoU, 79.8% Category IoU, and 66.4% Category iIoU.
    2. CamVid Test Set (11 classes, downsampled to 480×360480\times 360):

      • ENet achieves 68.3% Class Average accuracy and 51.3% Class IoU, compared to SegNet's 65.2% Class Average and 55.6% Class IoU.
      • ENet substantially outperforms SegNet on smaller object categories: Sign (51.0% vs 20.5%), Pedestrian (67.2% vs 57.1%), Pole (35.4% vs 27.5%), and Bicyclist (34.1% vs 30.7%).
    3. SUN RGB-D Test Set (37 indoor classes, RGB data only):

      • ENet achieves 59.5% Global Average accuracy, 32.6% Class Average accuracy, and 19.7% Mean IoU, while SegNet achieves 70.3% Global Average, 35.6% Class Average, and 26.3% Mean IoU.
  5. Knowl 5 — Bounded Class Weighting Function for Semantic Imbalance

    equation

    To handle severe class imbalance in pixel-wise cross-entropy training without allowing class weights to diverge towards infinity for rare classes, the class weight wclassw_{\text{class}} for a class with empirical pixel probability pclass∈(0,1]p_{\text{class}} \in (0, 1] is computed as:

    wclass=1ln⁡(c+pclass)w_{\text{class}} = \frac{1}{\ln(c + p_{\text{class}})}

    where c>1c > 1 is a hyperparameter set to c=1.02c = 1.02. This bounds all class weights strictly within the interval [1.0,50.0][1.0, 50.0] as pclass→0p_{\text{class}} \to 0, providing stable gradients compared to unbounded inverse class probability weighting.

  6. Knowl 6 — Context Aggregation via Factorized and Dilated Convolutions

    model/method

    To enlarge the receptive field without aggressive downsampling (which causes spatial edge degradation and mandates expensive upsampling decoders), ENet combines two convolutional strategies:

    • Asymmetric Convolutions: A 5×55\times 5 2D spatial convolution is decomposed into a sequence of two 1D operations: a 5×15\times 1 convolution followed by a 1×51\times 5 convolution. This factorization reduces the parameter and computational cost to that of a single 3×33\times 3 convolution while expanding the receptive field and adding intermediate non-linearities.
    • Dilated Convolutions: Standard 3×33\times 3 convolutions inside the bottleneck extension branches are replaced with dilated convolutions with dilation factors r∈{2,4,8,16}r \in \{2, 4, 8, 16\}. Interleaving dilated bottleneck modules with regular and asymmetric bottlenecks in Stages 2 and 3 provides a 4 percentage point gain in Class IoU on Cityscapes without adding parameters or arithmetic operations.
  7. Knowl 7 — Asymmetric Encoder-Decoder Network Design

    model/method

    In contrast to symmetric architectures (such as SegNet) where the decoder is a mirror image of the encoder, ENet uses an asymmetric structure consisting of a large encoder (Stages 1–3) and a small, lightweight decoder (Stages 4–5):

    • The encoder operates on downsampled feature representations to extract semantic features, aggregate spatial context, and perform categorization, similar to standard image classification backbones.
    • The decoder functions solely to upsample the encoder's feature representations and fine-tune spatial boundary details. Reducing decoder depth and width drastically lowers latency and parameter count.
  8. Knowl 8 — Two-Stage Training Protocol for Semantic Segmentation

    algorithm

    ENet is trained using the Adam optimization algorithm in a decoupled two-stage process:

    Input: Training dataset with images and ground truth segmentation masks
    Output: Trained ENet model weights
    Hyperparameters: Initial learning rate α=5×10−4\alpha = 5 \times 10^{-4}, weight decay λ=2×10−4\lambda = 2 \times 10^{-4}, batch size B=10B = 10, class weight parameter c=1.02c = 1.02
    Stage 1: Train Encoder Only
    Initialize ENet Encoder (Initial block, Stages 1, 2, and 3)
    Compute class weights: wclass=1/ln⁡(c+pclass)w_{\text{class}} = 1 / \ln(c + p_{\text{class}})
    Downsample ground truth masks to match the spatial resolution of Stage 3 output
    Train the encoder using Adam with learning rate α\alpha, weight decay λ\lambda, and batch size BB to predict downsampled class masks
    Stage 2: Train Full Encoder-Decoder
    Append the Decoder (Stages 4, 5, and fullconv) to the trained Encoder
    Train the complete ENet network end-to-end using full-resolution ground truth masks with Adam (α=5×10−4\alpha = 5 \times 10^{-4}, λ=2×10−4\lambda = 2 \times 10^{-4}, B=10B = 10)
    return Complete trained ENet model
  9. Knowl 9 — Spatial Dropout Regularization in Bottleneck Modules

    model/method

    To prevent overfitting on small semantic segmentation datasets (typically ∼103\sim 10^3 images), ENet applies Spatial Dropout rather than standard dropout, L2L_2 weight decay, or stochastic depth. Spatial Dropout drops entire 2D feature map channels rather than isolated individual activations. It is inserted at the end of the convolutional extension branch in every bottleneck block, immediately before the element-wise addition with the residual main branch. The drop probability is set to p=0.01p = 0.01 in the early layers (before bottleneck 2.0) and increased to p=0.1p = 0.1 for all subsequent bottleneck layers.

  10. Knowl 10 — Kernel Launch and Memory Bandwidth Overhead in Layer Factorization

    limitation

    Factorizing convolutional layers into smaller operations (1×11\times 1 projections, 1×n1\times n and n×1n\times 1 asymmetric convolutions) substantially reduces theoretical FLOP counts and parameters, but significantly increases the number of individual GPU kernel calls. For lightweight point-wise operations, GPU kernel launch latency and global memory round-trips dominate computation time because kernels do not retain intermediate activations in hardware registers across calls. In ENet, isolated PReLU activation kernels consume over 25%25\% of total inference latency. Realizing the full speedup potential of such compact architectures requires software-level kernel fusion in deep learning GPU backends.

Coverage note — None was omitted; all key contributions including network architecture, design principles, mathematical formulations, training algorithms, benchmark comparisons, and system-level performance analyses are covered.

References

  1. 1.Y. LeCun and Y. Bengio, “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks, pp. 255–258, 1998.
  2. 2.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25, 2012, pp. 1097–1105.
  3. 3.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  4. 4.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1–9.
  5. 5.J. Shotton, J. Winn, C. Rother, and A. Criminisi, “Textonboost for image understanding: Multi-class object recognition and segmentation by jointly modeling texture, layout, and context,” Int. Journal of Computer Vision (IJCV), January 2009.
  6. 6.F. Perronnin, Y. Liu, J. Sánchez, and H. Poirier, “Large-scale image retrieval with compressed fisher vectors,” in Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, 2010, pp. 3384–3391.
  7. 7.K. E. A. van de Sande, J. R. R. Uijlings, T. Gevers, and A. W. M. Smeulders, “mathbf{S}egmentation as selective search for object recognition,” in IEEE International Conference on Computer Vision, 2011.
  8. 8.C. Farabet, C. Couprie, L. Najman, and Y. LeCun, “mathbf{L}earning hierarchical features for scene labeling,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, pp. 1915–1929, Aug 2013.
  9. 9.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “mathbf{S}emantic image segmentation with deep convolutional nets and fully connected crfs,” arXiv preprint arXiv:1412.7062, 2014.
  10. 10.V. Badrinarayanan, A. Handa, and R. Cipolla, “mathbf{S}egnet: A deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling,” arXiv preprint arXiv:1505.07293, 2015.
  11. 11.V. Badrinarayanan, A. Kendall, and R. Cipolla, “mathbf{S}egnet: A deep convolutional encoder-decoder architecture for image segmentation,” arXiv preprint arXiv:1511.00561, 2015.
  12. 12.J. Long, E. Shelhamer, and T. Darrell, “mathbf{F}ully convolutional networks for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431–3440.
  13. 13.K. Simonyan and A. Zisserman, “mathbf{V}ery deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  14. 14.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “mathbf{T}he cityscapes dataset for semantic urban scene understanding,” in Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  15. 15.G. J. Brostow, J. Shotton, J. Fauqueur, and R. Cipolla, “mathbf{S}egmentation and recognition using structure from motion point clouds,” in ECCV (1), 2008, pp. 44–57.
  16. 16.S. Song, S. P. Lichtenberg, and J. Xiao, “mathbf{S}un rgb-d: A rgb-d scene understanding benchmark suite,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 567–576.
  17. 17.M. A. Ranzato, F. J. Huang, Y.-L. Boureau, and Y. LeCun, “mathbf{U}nsupervised learning of invariant feature hierarchies with applications to object recognition,” in Computer Vision and Pattern Recognition, 2007. CVPR’07. IEEE Conference on, 2007, pp. 1–8.
  18. 18.J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “mathbf{M}ultimodal deep learning,” in Proceedings of the 28th international conference on machine learning (ICML-11), 2011, pp. 689–696.
  19. 19.H. Noh, S. Hong, and B. Han, “mathbf{L}earning deconvolution network for semantic segmentation,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1520–1528.
  20. 20.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr, “mathbf{C}onditional random fields as recurrent neural networks,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1529–1537.
  21. 21.D. Eigen and R. Fergus, “mathbf{P}redicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2650–2658.
  22. 22.S. Hong, H. Noh, and B. Han, “mathbf{D}ecoupled deep neural network for semi-supervised semantic segmentation,” in Advances in Neural Information Processing Systems, 2015, pp. 1495–1503.
  23. 23.P. Sturgess, K. Alahari, L. Ladicky, and P. H. Torr, “mathbf{C}ombining appearance and structure from motion features for road scene understanding,” in BMVC 2012-23rd British Machine Vision Conference, 2009.
  24. 24.K. He, X. Zhang, S. Ren, and J. Sun, “mathbf{D}eep residual learning for image recognition,” arXiv preprint arXiv:1512.03385, 2015.
  25. 25.S. Ioffe and C. Szegedy, “mathbf{B}atch normalization: Accelerating deep network training by reducing internal covariate shift,” arXiv preprint arXiv:1502.03167, 2015.
  26. 26.K. He, X. Zhang, S. Ren, and J. Sun, “mathbf{D}elving deep into rectifiers: Surpassing human-level performance on imagenet classification,” pp. 1026–1034, 2015.
  27. 27.J. Tompson, R. Goroshin, A. Jain, Y. LeCun, and C. Bregler, “mathbf{E}fficient object localization using convolutional networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 648–656.
  28. 28.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “mathbf{R}ethinking the inception architecture for computer vision,” arXiv preprint arXiv:1512.00567, 2015.
  29. 29.S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer, “mathbf{c}udnn: Efficient primitives for deep learning,” arXiv preprint arXiv:1410.0759, 2014.
  30. 30.F. Yu and V. Koltun, “mathbf{M}ulti-scale context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.
  31. 31.K. He, X. Zhang, S. Ren, and J. Sun, “mathbf{I}dentity mappings in deep residual networks,” arXiv preprint arXiv:1603.05027, 2016.
  32. 32.J. Jin, A. Dundar, and E. Culurciello, “mathbf{F}lattened convolutional neural networks for feedforward acceleration,” arXiv preprint arXiv:1412.5474, 2014.
  33. 33.G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger, “mathbf{D}eep networks with stochastic depth,” arXiv preprint arXiv:1603.09382, 2016.
  34. 34.S. Han, H. Mao, and W. J. Dally, “mathbf{D}eep compression: Compressing deep neural network with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149, 2015.
  35. 35.D. Kingma and J. Ba, “mathbf{A}dam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014.

Citation

MLA
Paszke, A., et al. “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation”. arXiv, 2016, https://doi.org/10.48550/arxiv.1606.02147.
APA
Paszke, A., Chaurasia, A., Kim, S., & Culurciello, E. (2016). ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation. arXiv. https://doi.org/10.48550/arxiv.1606.02147
Chicago
Paszke, A., A. Chaurasia, S. Kim, and E. Culurciello. 2016. “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1606.02147.
Harvard
Paszke, A. et al. (2016) “ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation”. arXiv. Available at: https://doi.org/10.48550/arxiv.1606.02147.
Vancouver
1. Paszke A, Chaurasia A, Kim S, Culurciello E (2016) ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation. https://doi.org/10.48550/arxiv.1606.02147

BibTeX

@misc{https://doi.org/10.48550/arxiv.1606.02147,
  doi = {10.48550/ARXIV.1606.02147},
  url = {https://arxiv.org/abs/1606.02147},
  author = {Paszke, Adam and Chaurasia, Abhishek and Kim, Sangpil and Culurciello, Eugenio},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation},
  publisher = {arXiv},
  year = {2016},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors