LinkNet: Exploiting encoder representations for efficient semantic segmentation

Abhishek ChaurasiaEugenio Culurciello

article2017Visual Communications and Image Processing1,746 citations

Introduces LinkNet, an efficient neural network architecture that directly connects encoder feature maps to decoder stages, enabling real-time semantic segmentation on embedded hardware with only 11.5 million parameters.

Listen

Real-time visual scene parsing is essential for time-critical automated systems, such as autonomous vehicles and augmented reality platforms. These applications require pixel-level scene understanding, known as semantic segmentation, to accurately identify objects and drivable surfaces. However, most existing deep learning models are computationally massive, requiring millions of parameters and heavy computational budgets. This makes them too slow and power-hungry to operate in real time on mobile or embedded devices.

The article introduces and evaluates LinkNet, an efficient neural network architecture designed to perform fast and accurate semantic segmentation. The primary objective is to demonstrate that directly connecting encoder representations to the corresponding decoder stages can significantly lower computational demands and parameter size while preserving or improving segmentation accuracy.

The authors designed a compact network leveraging a lightweight ResNet18 encoder and paired it with a custom decoder. To test its effectiveness, they evaluated the model on two benchmark driving datasets, Cityscapes and CamVid. They measured segmentation accuracy using intersection-over-union metrics and evaluated processing speed and operational efficiency across high-end desktop graphics processors and embedded system hardware.

The findings show that LinkNet achieves high accuracy while drastically cutting resource requirements. First, the model achieves a 76.4% class intersection-over-union score on Cityscapes and 68.3% on CamVid, outperforming larger baseline networks like SegNet and Deep-Lab. Second, the architecture utilizes only 11.5 million parameters and 21.2 billion floating-point operations for standard resolution inputs, which is over 60% fewer parameters and over 90% fewer operations than SegNet. Third, the model achieves real-time inference speeds of up to 65.8 frames per second on a desktop graphics processor and maintains functional processing speeds on embedded hardware where several bulkier alternatives fail to run.

These results indicate that systems no longer need to compromise between high accuracy and computational speed. By directly reusing encoder features in the decoder, networks avoid wasting parameters and processing cycles to reconstruct lost spatial details. For organizations deploying computer vision, this design lowers hardware and energy costs, decreases latency, and enables high-quality visual perception on resource-constrained embedded modules and edge devices.

Engineering and deployment teams should consider adopting bypass-linked encoder-decoder designs when building real-time vision pipelines for embedded systems. Furthermore, high-throughput data centers can leverage these efficiencies to process large volumes of imagery with reduced computing infrastructure. Future development should evaluate LinkNet across broader weather and environmental conditions to establish robust confidence before deploying in safety-critical production settings.

Cover for LinkNet: Exploiting encoder representations for efficient semantic segmentation

Abstract

Pixel-wise semantic segmentation for visual scene understanding not only needs to be accurate, but also efficient in order to find any use in real-time application. Existing algorithms even though are accurate but they do not focus on utilizing the parameters of neural network efficiently. As a result they are huge in terms of parameters and number of operations; hence slow too. In this paper, we propose a novel deep neural network architecture which allows it to learn without any significant increase in number of parameters. Our network uses only 11.5 million parameters and 21.2 GFLOPs for processing an image of resolution 3x640x360. It gives state-of-the-art performance on CamVid and comparable results on Cityscapes dataset. We also compare our networks processing time on NVIDIA GPU and embedded system device with existing state-of-the-art architectures for different image resolutions.

Table of Contents

  • I Introduction
  • II Related work
  • III Network architecture
  • IV Results
  • IV-A Performance Analysis
  • IV-B Benchmarks
  • V Conclusion
  • References

Knowls

  1. Knowl 1 — LinkNet Overall Architecture and Bypass Connection Mechanism

    model/method

    LinkNet is an encoder-decoder deep neural network designed for real-time semantic segmentation. It employs a pre-trained ResNet-18 as its encoder and pairs it with a lightweight decoder. The key architectural mechanism is direct bypass connections that route spatial representations from the encoder directly to the corresponding decoder stage via element-wise addition.

    The overall architecture operates as follows:

    1. Initial Block: The input image is processed by a 7×77 \times 7 convolution layer with stride 2, followed by batch normalization, ReLU activation, and a 3×33 \times 3 max-pooling operation with stride 2. This downsamples the input by a spatial factor of 4 and outputs 64 feature maps.

    2. Encoder-Decoder Pathway: The network has four sequential encoder blocks (Encoder Block 1 to 4) matched symmetrically with four decoder blocks (Decoder Block 4 down to Decoder Block 1). The input to Encoder Block ii is added directly via an additive skip connection to the output of Decoder Block ii before that representation is passed to Decoder Block i−1i-1: Input(Decoder Block i−1)=Output(Decoder Block i)+Input(Encoder Block i)\text{Input}(\text{Decoder Block } i-1) = \text{Output}(\text{Decoder Block } i) + \text{Input}(\text{Encoder Block } i) This bypass mechanism preserves fine spatial details lost during strided operations and pooling in the encoder, eliminating the need to store pooling indices or relearn lost spatial features with heavy decoder parameterization.

    3. Final Classification Head: The output of Decoder Block 1 (64 channels) is upsampled to the full image resolution and mapped to NN semantic class scores through the following sequence:

    • Transposed convolution full-conv(3×3,(64,32),∗2)\text{full-conv}(3 \times 3, (64, 32), *2) with upsampling factor 2
    • Standard convolution conv(3×3,(32,32))\text{conv}(3 \times 3, (32, 32))
    • Transposed convolution full-conv(2×2,(32,N),∗2)\text{full-conv}(2 \times 2, (32, N), *2) with upsampling factor 2

    All convolution and transposed convolution layers in the architecture are followed by Batch Normalization and ReLU non-linearities.

  2. Knowl 2 — LinkNet Decoder Block Module

    model/method

    Each decoder block in LinkNet upsamples feature maps spatially by a factor of 2 while reducing the channel dimension from mm input channels to nn output channels. To minimize parameter count and computational cost, the decoder block applies a channel-reduction bottleneck prior to upsampling.

    For a decoder block with mm input channels and nn output channels, the internal operations proceed sequentially:

    1. 1×11 \times 1 Convolution (Channel Reduction): conv(1×1,(m,m/4))\text{conv}(1 \times 1, (m, m/4)) reduces the channel dimensionality by a factor of 4.
    2. 3×33 \times 3 Full Convolution (Upsampling): full-conv(3×3,(m/4,m/4),∗2)\text{full-conv}(3 \times 3, (m/4, m/4), *2) applies a transposed convolution with a 3×33 \times 3 kernel and stride 2 to double the spatial resolution while maintaining m/4m/4 feature channels.
    3. 1×11 \times 1 Convolution (Channel Expansion/Projection): conv(1×1,(m/4,n))\text{conv}(1 \times 1, (m/4, n)) projects the intermediate representation to the target nn output channels.

    Each convolution and transposed convolution is followed by Batch Normalization and a ReLU activation function.

    The channel parameters (m,n)(m, n) for the four decoder blocks are:

    • Decoder Block 4: m=512→n=256m = 512 \to n = 256
    • Decoder Block 3: m=256→n=128m = 256 \to n = 128
    • Decoder Block 2: m=128→n=64m = 128 \to n = 64
    • Decoder Block 1: m=64→n=64m = 64 \to n = 64
  3. Knowl 3 — LinkNet Encoder Block Module

    model/method

    The encoder in LinkNet is structured into four residual blocks based on ResNet-18. Each encoder block takes mm input channels and outputs nn channels.

    An encoder block consists of two stacked residual sub-modules:

    1. First Sub-Module (Downsampling Residual):
    • Main path: conv(3×3,(m,n),/2)\text{conv}(3 \times 3, (m, n), /2) with stride 2 for spatial downsampling, followed by conv(3×3,(n,n))\text{conv}(3 \times 3, (n, n)) with stride 1.
    • Shortcut path: A projection connection with stride 2 matching spatial and channel dimensions (m→n)(m \to n).
    • The outputs of the main path and shortcut path are summed via element-wise addition.
    1. Second Sub-Module (Identity Residual):
    • Main path: conv(3×3,(n,n))\text{conv}(3 \times 3, (n, n)) with stride 1, followed by another conv(3×3,(n,n))\text{conv}(3 \times 3, (n, n)) with stride 1.
    • Shortcut path: An identity skip connection.
    • The outputs are summed via element-wise addition.

    Every convolution operation is followed by Batch Normalization and a ReLU non-linearity.

    The channel parameters (m,n)(m, n) across the four encoder blocks are:

    • Encoder Block 1: m=64→n=64m = 64 \to n = 64 (operates with stride 1, without spatial downsampling)
    • Encoder Block 2: m=64→n=128m = 64 \to n = 128 (downsamples spatial dimensions by factor 2)
    • Encoder Block 3: m=128→n=256m = 128 \to n = 256 (downsamples spatial dimensions by factor 2)
    • Encoder Block 4: m=256→n=512m = 256 \to n = 512 (downsamples spatial dimensions by factor 2)
  4. Knowl 4 — Inverse Logarithmic Class Weighting Formulation

    equation

    To counteract heavy class imbalance in semantic segmentation datasets without causing unbounded gradient values for extremely rare classes, class-specific loss weights are assigned using a bounded inverse logarithmic function:

    wclass=1ln⁡(c+pclass)w_{\text{class}} = \frac{1}{\ln(c + p_{\text{class}})}

    where:

    • pclass∈[0,1]p_{\text{class}} \in [0, 1] denotes the probability (frequency) of a given class occurring across all pixels in the training dataset.
    • c=1.02c = 1.02 is a constant offset parameter that prevents the denominator from approaching zero for rare classes (pclass→0p_{\text{class}} \to 0), capping the maximum class weight at ≈1ln⁡(1.02)≈50.5\approx \frac{1}{\ln(1.02)} \approx 50.5.
    • wclassw_{\text{class}} is the scalar multiplying factor applied to the cross-entropy loss for pixels belonging to that class.
  5. Knowl 5 — Computational Complexity and Hardware Inference Benchmark

    data/table

    LinkNet was benchmarked against SegNet and ENet to assess parameter count, theoretical floating-point operations (GFLOPs), and empirical inference frame rates across multiple resolutions (W×HW \times H) on both an embedded GPU module (NVIDIA Jetson TX1) and a desktop workstation GPU (NVIDIA Titan X).

    Model GFLOPs (640×360640 \times 360) Parameters Model size (fp16)
    SegNet 286.0 29.5M 56.2 MB
    ENet 3.8 0.4M 0.7 MB
    LinkNet 21.2 11.5M 22.0 MB
    NVIDIA TX1 NVIDIA Titan X
    Model 480×320480 \times 320 640×360640 \times 360 1280×7201280 \times 720 640×360640 \times 360 1280×7201280 \times 720 1920×10801920 \times 1080
    ms fps ms fps ms fps ms fps ms fps ms fps
    SegNet 757 1.3 1251 0.8 - - 69 14.6 289 3.5 637 1.6
    ENet 47 21.1 69 14.6 262 3.8 7 135.4 21 46.8 46 21.6
    LinkNet 108 9.3 134 7.8 501 2.0 15 65.8 53 18.7 117 8.5

    Note: '-' indicates the model could not process the resolution on the embedded device due to hardware resource constraints. LinkNet processes 640×360640 \times 360 frames at 65.8 fps on Titan X and 7.8 fps on Jetson TX1, achieving real-time capabilities with over 13×13\times fewer GFLOPs than SegNet.

  6. Knowl 6 — Cityscapes Validation Results and Bypass Connection Ablation

    data/table

    LinkNet was evaluated on the Cityscapes validation set across 19 semantic classes using Class Intersection over Union (Class IoU) and Class Instance-level Intersection over Union (Class iIoU). To evaluate the contribution of the encoder-to-decoder bypass connections, LinkNet was benchmarked both with and without these connections.

    Model Class IoU (%) Class iIoU (%)
    SegNet* 56.1 34.2
    ENet* 58.3 34.4
    Dilation10 68.7 -
    Deep-Lab CRF (VGG16) 65.9 -
    Deep-Lab CRF (ResNet101) 71.4 42.6
    LinkNet without bypass 72.6 51.4
    LinkNet 76.4 58.6

    Note: * denotes results evaluated on the Cityscapes test set, whereas the others are on the validation set.

    The bypass connections provide an absolute improvement of +3.8%+3.8\% in Class IoU (from 72.6%72.6\% to 76.4%76.4\%) and +7.2%+7.2\% in Class iIoU (from 51.4%51.4\% to 58.6%58.6\%), demonstrating that routing encoder representations directly to decoder outputs substantially enhances segmentation accuracy without adding parameters.

  7. Knowl 7 — CamVid Test Set Benchmark Results

    data/table

    LinkNet was evaluated on the CamVid automotive dataset test set across 11 semantic classes (unlabeled pixels were ignored). Performance is compared against baseline and state-of-the-art models in terms of per-class IoU, mean Class IoU, and Class iIoU.

    Model Building Tree Sky Car Sign Road Pedestrian Fence Pole Sidewalk Bicyclist IoU iIoU
    SegNet 88.8 87.3 92.4 82.1 20.5 97.2 57.1 49.3 27.5 84.4 30.7 65.2 55.6
    ENet 74.7 77.8 95.1 82.4 51.0 95.1 67.2 51.7 35.4 86.7 34.1 68.3 51.3
    Dilation8 82.6 76.2 89.9 84.0 46.9 92.2 56.3 35.8 23.4 75.3 55.5 65.3 -
    LinkNet w/o bypass 84.6 87.4 88.8 72.6 37.1 95.3 61.2 56.0 33.1 88.3 24.4 66.3 52.7
    LinkNet 88.8 85.3 92.8 77.6 41.7 96.8 57.0 57.8 37.8 88.4 27.2 68.3 55.8

    LinkNet achieves a mean IoU of 68.3%68.3\% and an instance-level IoU of 55.8%55.8\%, outperforming LinkNet without bypass connections (66.3%66.3\% IoU, 52.7%52.7\% iIoU) and achieving higher iIoU than all benchmarked models.

  8. Knowl 8 — LinkNet Training Protocol and Implementation Hyperparameters

    experimental setup

    LinkNet was implemented and trained under the following experimental configuration:

    • Framework & Optimization: Implemented in Torch7 and optimized using RMSProp across four NVIDIA Titan X GPUs.
    • Loss Weighting: Custom inverse logarithmic class weighting (wclass=1ln⁡(1.02+pclass)w_{\text{class}} = \frac{1}{\ln(1.02 + p_{\text{class}})}) applied to cross-entropy loss to handle class imbalance.
    • Cityscapes Dataset Setup:
      • Dataset size: 5,000 fine-annotated images (2,975 train, 500 validation, 1,525 test) across 19 semantic classes.
      • Training input resolution: 1024×5121024 \times 512.
      • Batch size: 10.
      • Initial learning rate: 5×10−45 \times 10^{-4}.
    • CamVid Dataset Setup:
      • Dataset size: 701 total images (367 train, 101 validation, 233 test) labeled with 11 target semantic classes (12th unlabeled class omitted during training).
      • Training input resolution: Original 960×720960 \times 720 downsampled by a factor of 1.25 to 768×576768 \times 576.
      • Batch size: 8.

Coverage note — None was omitted; all key architectural components, equations, performance tables, ablations, and experimental parameters have been extracted as knowls.

References

  1. 1.Y. LeCun and Y. Bengio, “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks, pp. 255–258, 1998.
  2. 2.Y. LeCun, L. Bottou, G. B. Orr, and K. R. Muller, ¨ Neural Networks: Tricks of the Trade. Berlin, Heidelberg: Springer Berlin Heidelberg, 1998, ch. Efficient BackProp, pp. 9–50.
  3. 3.M. A. Ranzato, F. J. Huang, Y.-L. Boureau, and Y. LeCun, “Unsupervised learning of invariant feature hierarchies with applications to object recognition,” in Computer Vision and Pattern Recognition, 2007. CVPR’07. IEEE Conference on, 2007, pp. 1–8.
  4. 4.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25, 2012, pp. 1097–1105.
  5. 5.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  6. 6.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” arXiv preprint arXiv:1512.03385, 2015.
  7. 7.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1–9.
  8. 8.C. Szegedy, S. Ioffe, and V. Vanhoucke, “Inception-v4, inception-resnet and the impact of residual connections on learning,” arXiv preprint arXiv:1602.07261, 2016.
  9. 9.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” arXiv preprint arXiv:1312.6229, 2013.
  10. 10.J. Tompson, R. Goroshin, A. Jain, Y. LeCun, and C. Bregler, “Efficient object localization using convolutional networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 648–656.
  11. 11.P. Sturgess, K. Alahari, L. Ladicky, and P. H. Torr, “Combining appearance and structure from motion features for road scene understanding,” in BMVC 2012-23rd British Machine Vision Conference, 2009.
  12. 12.D. Eigen and R. Fergus, “Linearizing depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2650–2658.
  13. 13.X. Ren, L. Bo, and D. Fox, “Linear-(d) scene labeling: Features and algorithms,” in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, 2012, pp. 2759–2766.
  14. 14.C. Farabet, C. Couprie, L. Najman, and Y. LeCun, “Linear hierarchical features for scene labeling,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, pp. 1915–1929, Aug 2013.
  15. 15.J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Linear deep learning,” in Proceedings of the 28th international conference on machine learning (ICML-11), 2011, pp. 689–696.
  16. 16.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “Linear only look once: Unified, real-time object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779–788.
  17. 17.S. Ren, K. He, R. Girshick, and J. Sun, “Linear r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems, 2015, pp. 91–99.
  18. 18.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Linear single shot multibox detector,” in European Conference on Computer Vision. Springer, 2016, pp. 21–37.
  19. 19.A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “Linear A deep neural network architecture for real-time semantic segmentation,” arXiv preprint arXiv:1606.02147, 2016.
  20. 20.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “Linear cityscapes dataset for semantic urban scene understanding,” in Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  21. 21.G. J. Brostow, J. Shotton, J. Fauqueur, and R. Cipolla, “Linear and recognition using structure from motion point clouds,” in ECCV (1), 2008, pp. 44–57.
  22. 22.V. Badrinarayanan, A. Handa, and R. Cipolla, “Linear A deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling,” arXiv preprint arXiv:1505.07293, 2015.
  23. 23.H. Noh, S. Hong, and B. Han, “Linear deconvolution network for semantic segmentation,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1520–1528.
  24. 24.V. Badrinarayanan, A. Kendall, and R. Cipolla, “Linear A deep convolutional encoder-decoder architecture for image segmentation,” arXiv preprint arXiv:1511.00561, 2015.
  25. 25.J. Long, E. Shelhamer, and T. Darrell, “Linear convolutional networks for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431–3440.
  26. 26.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Linear image segmentation with deep convolutional nets and fully connected crfs,” arXiv preprint arXiv:1412.7062, 2014.
  27. 27.F. Yu and V. Koltun, “Linear context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.
  28. 28.F. Visin, M. Ciccone, A. Romero, K. Kastner, K. Cho, Y. Bengio, M. Matteucci, and A. Courville, “Linear A recurrent neural network-based model for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016, pp. 41–48.
  29. 29.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr, “Linear random fields as recurrent neural networks,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1529–1537.
  30. 30.S. Han, H. Mao, and W. J. Dally, “Linear compression: Compressing deep neural network with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149, 2015.
  31. 31.V. Nair and G. E. Hinton, “Linear linear units improve restricted boltzmann machines,” in Proceedings of the 27th international conference on machine learning (ICML-10), 2010, pp. 807–814.
  32. 32.S. Ioffe and C. Szegedy, “Linear normalization: Accelerating deep network training by reducing internal covariate shift,” arXiv preprint arXiv:1502.03167, 2015.
  33. 33.R. Collobert, K. Kavukcuoglu, and C. Farabet, “Linear A matlab-like environment for machine learning,” in BigLearn, NIPS Workshop, 2011.
  34. 34.F. Yu and V. Koltun, “Linear context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.
  35. 35.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Linear Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” arXiv preprint arXiv:1606.00915, 2016.

Citation

MLA
Chaurasia, A., and E. Culurciello. “LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation”. 2017 IEEE Visual Communications and Image Processing (VCIP), 2017, pp. 1–4, https://doi.org/10.1109/VCIP.2017.8305148.
APA
Chaurasia, A., & Culurciello, E. (2017). LinkNet: Exploiting encoder representations for efficient semantic segmentation. 2017 IEEE Visual Communications and Image Processing (VCIP), 1–4. https://doi.org/10.1109/VCIP.2017.8305148
Chicago
Chaurasia, A., and E. Culurciello. 2017. “LinkNet: Exploiting Encoder Representations for Efficient Semantic Segmentation”. 2017 IEEE Visual Communications and Image Processing (VCIP), 1–4. https://doi.org/10.1109/VCIP.2017.8305148.
Harvard
Chaurasia, A. and Culurciello, E. (2017) “LinkNet: Exploiting encoder representations for efficient semantic segmentation”, 2017 IEEE Visual Communications and Image Processing (VCIP). IEEE, pp. 1–4. Available at: https://doi.org/10.1109/VCIP.2017.8305148.
Vancouver
1. Chaurasia A, Culurciello E (2017) LinkNet: Exploiting encoder representations for efficient semantic segmentation. In: 2017 IEEE Visual Communications and Image Processing (VCIP). IEEE, pp 1–4

BibTeX

@inproceedings{Chaurasia_2017, title={LinkNet: Exploiting encoder representations for efficient semantic segmentation}, url={http://dx.doi.org/10.1109/VCIP.2017.8305148}, DOI={10.1109/vcip.2017.8305148}, booktitle={2017 IEEE Visual Communications and Image Processing (VCIP)}, publisher={IEEE}, author={Chaurasia, Abhishek and Culurciello, Eugenio}, year={2017}, month=Dec, pages={1–4} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF