A Review on Deep Learning Techniques Applied to Semantic Segmentation

Alberto Garcia-GarciaSergio Orts-EscolanoSergiu OpreaVictor Villena-MartinezJose Garcia-Rodriguez

article2017arXiv1,398 citations

Surveys deep learning methods for semantic segmentation by systematically categorizing leading architectures, comparing quantitative benchmark performance across standard datasets, and identifying key future research directions for computer vision applications.

Listen

Semantic segmentation—the automated assignment of a category label to every individual pixel or point in visual data—is essential for critical technologies such as autonomous driving, robotics, augmented reality, and indoor navigation. While deep learning methods have rapidly replaced traditional approaches by automatically learning powerful visual representations, the rapid influx of research has created a fragmented landscape lacking unified performance comparisons and standard baselines.

The article provides a systematic review of deep learning techniques applied to semantic segmentation. It evaluates the architectural evolution, performance metrics, and key design trade-offs across 27 deep learning methods and 28 standard benchmark datasets spanning two-dimensional (2D), depth-augmented (2.5D), volumetric (3D), and video formats.

The review traces architectural progress from the foundational Fully Convolutional Network through specialized encoder-decoder structures, dilated convolutions, recurrent neural networks, and graphical models. The analysis synthesizes quantitative results primarily using Mean Intersection over Union (a standard overlap metric for segmentation accuracy), while also assessing practical operational constraints such as execution time and memory footprint.

The findings show that deep learning significantly outperforms traditional methods, with DeepLab achieving the highest accuracy across standard 2D benchmarks (such as a 79.70% score on PASCAL VOC-2012 and 70.40% on Cityscapes). For sequential and depth-augmented data, recurrent architectures like DAG-RNN and LSTM-CF dominate, reaching up to 91.60% on CamVid and 58.50% on SUN3D, respectively. For raw 3D point cloud data, PointNet demonstrated baseline viability by achieving 83.70% on ShapeNet Part without requiring voxel discretization. However, most leading architectures exhibit severe computational latency—often taking between 100 and over 500 milliseconds per low-resolution image—falling far short of the real-time threshold of at least 25 frames per second needed for camera streams.

These findings indicate that while visual accuracy has reached practical utility, current models pose substantial deployment risks for safety-critical and resource-constrained environments like self-driving vehicles or mobile robots. Deploying high-accuracy models directly onto embedded hardware without architectural changes risks hardware bottlenecks, delayed decision-making, and high power consumption.

To move forward, engineering teams should evaluate lightweight architectures such as ENet or adopt network pruning and compression techniques to balance accuracy against memory and latency limits. In video processing, practitioners should consider adaptive update schedules (like Clockwork networks) to reduce redundant computations across frames. Future research must also focus on creating standardized, real-world 3D datasets and developing graph-based convolutions to handle spatial and temporal coherence without artificial data discretization.

Readers should note that comparisons across the literature remain limited by severe reporting gaps: very few studies report memory footprint or execution runtime, and several models rely on non-standard datasets or omit implementation details. While confidence is high regarding relative 2D accuracy rankings on major benchmarks, caution is necessary when evaluating hardware suitability and performance in complex 3D or video applications.

arXiv: 1704.06857
Cover for A Review on Deep Learning Techniques Applied to Semantic Segmentation

Abstract

Image semantic segmentation is more and more being of interest for computer vision and machine learning researchers. Many applications on the rise need accurate and efficient segmentation mechanisms: autonomous driving, indoor navigation, and even virtual or augmented reality systems to name a few. This demand coincides with the rise of deep learning approaches in almost every field or application target related to computer vision, including semantic segmentation or scene understanding. This paper provides a review on deep learning methods for semantic segmentation applied to various application areas. Firstly, we describe the terminology of this field as well as mandatory background concepts. Next, the main datasets and challenges are exposed to help researchers decide which are the ones that best suit their needs and their targets. Then, existing methods are reviewed, highlighting their contributions and their significance in the field. Finally, quantitative results are given for the described methods and the datasets in which they were evaluated, following up with a discussion of the results. At last, we point out a set of promising future works and draw our own conclusions about the state of the art of semantic segmentation using deep learning techniques.

Table of Contents

  • I Introduction
  • II Terminology and Background Concepts
  • II-A Common Deep Network Architectures
  • II-A1 AlexNet
  • II-A2
  • II-A3 GoogLeNet
  • II-A4 ResNet
  • II-A5 ReNet
  • II-B Transfer Learning
  • II-C Data Preprocessing and Augmentation
  • III Datasets and Challenges
  • III-A 2D Datasets
  • III-B Datasets
  • III-C Datasets
  • IV Methods
  • IV-A Decoder Variants
  • IV-B Integrating Context Knowledge
  • IV-B1
  • IV-B2 Dilated Convolutions
  • IV-B3 Multi-scale Prediction
  • IV-B4 Feature Fusion
  • IV-B5
  • IV-C Instance Segmentation
  • IV-D Data
  • IV-E Data
  • IV-F Video Sequences
  • V Discussion
  • V-A Evaluation Metrics
  • V-A1 Execution Time
  • V-A2 Memory Footprint
  • V-A3 Accuracy
  • V-B Results
  • V-B1
  • V-B2
  • V-B3
  • V-B4 Sequences
  • V-C Summary
  • V-D Future Research Directions
  • VI Conclusion
  • References

Knowls

  1. Knowl 1 — Problem Formulation and Mathematical Framework of Semantic Segmentation

    definition

    Semantic segmentation is formulated as a dense per-pixel labeling problem where an input spatial or volumetric domain is mapped to a discrete set of semantic categories.

    Let X={x1,x2,…,xN}\mathcal{X} = \{x_1, x_2, \dots, x_N\} denote a set of NN discrete elements or random variables representing spatial coordinates (for instance, the pixels of a 2D image of dimension W×H=NW \times H = N, or volumetric voxels/points in 3D data). Let L={l0,l1,l2,…,lk}\mathcal{L} = \{l_0, l_1, l_2, \dots, l_k\} denote the discrete label space comprising kk distinct target object or material categories along with a void or background class l0l_0.

    The objective of semantic segmentation is to determine an optimal assignment f:X→Lf: \mathcal{X} \to \mathcal{L} that allocates a class state li∈Ll_i \in \mathcal{L} to every element xj∈Xx_j \in \mathcal{X}, thereby producing a dense, fine-grained semantic partition of the input domain.

  2. Knowl 2 — Mathematical Formulations of Standard Evaluation Metrics for Semantic Segmentation

    equation

    Semantic segmentation models are evaluated on their per-pixel classification accuracy and boundary overlap using four standard metrics derived from a contingency matrix. Let k+1k + 1 denote the total number of classes (classes l0,…,lkl_0, \dots, l_k, including background/void), and let pijp_{ij} represent the total number of pixels belonging to ground-truth class ii that are predicted as class jj. Under this notation, piip_{ii} denotes the true positive count for class ii, pijp_{ij} (with i≠ji \neq j) denotes false positives for class jj arising from class ii, and pjip_{ji} (with j≠ij \neq i) denotes false negatives for class ii misclassified as class jj.

    1. Pixel Accuracy (PA) computes the global ratio of correctly classified pixels to the total number of pixels:

    PA=∑i=0kpii∑i=0k∑j=0kpijPA = \frac{\sum_{i=0}^{k} p_{ii}}{\sum_{i=0}^{k} \sum_{j=0}^{k} p_{ij}}

    1. Mean Pixel Accuracy (MPA) computes the proportion of correct pixels on a per-class basis and averages across all classes:

    MPA=1k+1∑i=0kpii∑j=0kpijMPA = \frac{1}{k + 1} \sum_{i=0}^{k} \frac{p_{ii}}{\sum_{j=0}^{k} p_{ij}}

    1. Mean Intersection over Union (MIoU) computes the ratio between the intersection (true positives) and the union (sum of true positives, false positives, and false negatives) across all classes:

    MIoU=1k+1∑i=0kpii∑j=0kpij+∑j=0kpji−piiMIoU = \frac{1}{k + 1} \sum_{i=0}^{k} \frac{p_{ii}}{\sum_{j=0}^{k} p_{ij} + \sum_{j=0}^{k} p_{ji} - p_{ii}}

    1. Frequency Weighted Intersection over Union (FWIoU) weights each category's Intersection over Union according to its frequency of occurrence in the ground truth:

    FWIoU=1∑i=0k∑j=0kpij∑i=0k[(∑j=0kpij)pii∑j=0kpij+∑j=0kpji−pii]FWIoU = \frac{1}{\sum_{i=0}^{k} \sum_{j=0}^{k} p_{ij}} \sum_{i=0}^{k} \left[ \frac{\left( \sum_{j=0}^{k} p_{ij} \right) p_{ii}}{\sum_{j=0}^{k} p_{ij} + \sum_{j=0}^{k} p_{ji} - p_{ii}} \right]

  3. Knowl 3 — Taxonomy of Deep Learning Architectures for Semantic Segmentation

    model/method

    Deep learning architectures for semantic segmentation are structured around resolving the limitations of canonical Fully Convolutional Networks (FCNs), which transform image classification backbones (such as AlexNet, VGG-16, GoogLeNet, and ResNet) into dense predictors by replacing fully connected layers with convolutions and using fractionally strided convolutions (deconvolutions) for upsampling. Segmentation architectures are categorized across six primary paradigms:

    1. Decoder Variants: Contrasting upsampling mechanisms. FCNs employ learnable deconvolution filters with skip-connections from encoder feature maps. In contrast, SegNet and Bayesian SegNet store max-pooling indices from encoder pooling layers to guide unpooling in the decoder, followed by trainable convolutional filter banks to recover dense resolution without storing large encoder feature maps.

    2. Context Integration Mechanisms:

      • Conditional Random Fields (CRFs): Capture long-range spatial dependencies and recover boundary details lost to CNN pooling invariance. DeepLab employs a fully connected pairwise CRF as an offline post-processing step, whereas CRFasRNN unrolls CRF mean-field inference iterations as recurrent layers for end-to-end training.
      • Dilated Convolutions: Generalize convolutional filters with a dilation factor ℓ\ell, expanding receptive field size exponentially without downsampling resolution or introducing additional parameters (e.g., Dilation network, DeepLab, ENet).
      • Multi-Scale Prediction: Processes images through multiple parallel branches operating on different scales (e.g., multi-scale VGG variants, coarse-to-fine sequential networks) to improve scale invariance.
      • Feature Fusion: Early fusion unpools and concatenates global features with local features (ParseNet), whereas late fusion combines predictions or representations across different layers (SharpMask, FCN skip connections).
      • Recurrent Neural Networks (RNNs): Employs multidimensional sequence modeling across vertical and horizontal directions (e.g., ReSeg using Gated Recurrent Units, 2D-LSTM, Directed Acyclic Graph RNNs [DAG-RNN]) to capture global topological and long-range spatial context.
    3. Instance Segmentation: Segmenting separate instances within classes using region-proposal refinement pipelines (Simultaneous Detection and Segmentation [SDS], DeepMask, SharpMask, and MultiPathNet).

    4. RGB-D Multimodal Architectures: Incorporating geometric depth cues encoded as three-channel representations (such as Horizontal Disparity, Height above ground, and surface normal Angle with gravity [HHA]) or cross-modal recurrent fusion (LSTM-CF) and multi-view SLAM warping.

    5. 3D Volumetric and Point Cloud Segmentation: Processing spatial data through voxelized 3D occupancy grids with 3D CNNs, or directly consuming raw, unordered point clouds using Multi-Layer Perceptrons and permutation-invariant symmetric pooling functions (PointNet).

    6. Video Sequence Segmentation: Leveraging temporal continuity and feature velocity across frames to update deeper semantic layers less frequently than shallow layers (Clockwork FCN), or learning spatiotemporal features via 3D convolutions across multi-frame clips (C3D, voxel-to-voxel 3DCNNs).

  4. Knowl 4 — Comprehensive Dataset Taxonomy for 2D, 2.5D, and 3D Semantic Segmentation

    data/table

    Semantic segmentation benchmarks span 2D RGB imagery, 2.5D RGB-D data, and 3D volumetric or point cloud data across generic, urban/driving, indoor, and specialized domains.

    Dataset Domain Data Type Classes Train Val Test
    PASCAL VOC 2012 Generic 2D 21 1,464 1,449 Private
    PASCAL-Context Generic 2D 540 (59) 10,103 – 9,637
    PASCAL-Part Generic-Part 2D 20 10,103 – 9,637
    SBD Generic 2D 21 8,498 2,857 –
    Microsoft COCO Generic 2D >80 82,783 40,504 81,434
    SYNTHIA Urban (Driving) 2D (Synthetic) 11 13,407 – –
    Cityscapes (fine) Urban 2D 30 (8) 2,975 500 1,525
    Cityscapes (coarse) Urban 2D 30 (8) 22,973 500 –
    CamVid Urban (Driving) 2D 32 701 – –
    CamVid-Sturgess Urban (Driving) 2D 11 367 100 233
    KITTI-Layout Urban (Driving) 2D 3 323 – –
    KITTI-Ros Urban (Driving) 2D 11 170 – 46
    KITTI-Zhang Urban (Driving) 2D/3D 10 140 – 112
    Stanford Background Outdoor 2D 8 725 – –
    SiftFlow Outdoor 2D 33 2,688 – –
    Youtube-Objects Objects 2D (Video) 10 10,167 – –
    Adobe Portrait Portrait 2D 2 1,500 300 –
    MINC Materials 2D 23 7,061 2,500 5,000
    DAVIS Generic (Video) 2D 4 4,219 2,023 2,180
    NYUDv2 Indoor 2.5D 40 795 654 –
    SUN3D Indoor 2.5D (Video) – 19,640 – –
    SUNRGBD Indoor 2.5D 37 2,666 2,619 5,050
    RGB-D Object Household 2.5D 51 207,920 – –
    ShapeNet Part 3D Object/Part 3D (Synthetic) 16 / 50 31,963 – –
    Stanford 2D-3D-S Indoor 2D/2.5D/3D 13 70,469 – –
    3D Mesh 3D Object/Part 3D (Synthetic) 19 380 – –
    Sydney Urban Objects Urban (Objects) 3D 26 41 – –
    Point Cloud Benchmark Urban/Nature 3D 8 15 – 15

    The benchmark analysis demonstrates that 2D datasets dominate the literature in volume and standardization, with PASCAL VOC and COCO providing standard evaluation protocols. In contrast, 2.5D and 3D datasets are constrained by smaller scales, reliance on synthetic data, or fragmented class partitions across research studies.

  5. Knowl 5 — Empirical Accuracy Benchmark of Deep Segmentation Methods Across 2D RGB Datasets

    empirical result

    Quantitative performance across standard 2D RGB segmentation benchmarks indicates distinct architectural advantages depending on dataset characteristics:

    Dataset Method Accuracy (MIoU %)
    PASCAL VOC 2012 DeepLab 79.70
    Dilation 75.30
    CRFasRNN 74.70
    ParseNet 69.80
    FCN-8s 67.20
    Multi-scale-CNN-Eigen 62.60
    Bayesian SegNet 60.50
    PASCAL-Context DeepLab 45.70
    CRFasRNN 39.28
    FCN-8s 39.10
    PASCAL-Person-Part DeepLab 64.94
    CamVid DAG-RNN 91.60
    Bayesian SegNet 63.10
    SegNet 60.10
    ReSeg 58.80
    ENet 55.60
    Cityscapes DeepLab 70.40
    Dilation10 67.10
    FCN-8s 65.30
    CRFasRNN 62.50
    ENet 58.30
    Stanford Background rCNN 80.20
    2D-LSTM 78.56
    SiftFlow DAG-RNN 85.30
    rCNN 77.70
    2D-LSTM 70.11

    DeepLab achieves the highest accuracy on broad object segmentation benchmarks (PASCAL VOC 2012 with 79.70% MIoU, PASCAL-Context with 45.70% MIoU, and Cityscapes with 70.40% MIoU) due to dilated convolutions combined with fully connected CRFs. Conversely, recurrent graphical architectures such as DAG-RNN dominate structured outdoor scene parsing datasets, reaching 91.60% MIoU on CamVid and 85.30% MIoU on SiftFlow by modeling long-range contextual graph dependencies.

  6. Knowl 6 — Empirical Accuracy Benchmark for 2.5D RGB-D, 3D Point Cloud, and Video Segmentation

    empirical result

    Empirical evaluations for multimodal, 3D spatial, and temporal video segmentation show the performance of specialized deep architectures:

    Modality / Task Dataset Method Accuracy (MIoU %)
    2.5D (RGB-D) SUN-RGB-D LSTM-CF 48.10
    NYUDv2 LSTM-CF 49.40
    SUN3D LSTM-CF 58.50
    3D Spatial ShapeNet Part PointNet 83.70
    Stanford 2D-3D-S PointNet 47.71
    Video Sequences Cityscapes Clockwork Convnet 64.40
    Youtube-Objects Clockwork Convnet 68.50

    In 2.5D RGB-D scene parsing, LSTM-CF achieves 48.10% MIoU on SUN-RGB-D, 49.40% MIoU on NYUDv2, and 58.50% on SUN3D by combining multiscale photometric features with vertical and horizontal context fusion over depth and RGB channels.

    In 3D segmentation, PointNet operates directly on raw, unordered 3D point sets to obtain 83.70% MIoU on ShapeNet Part and 47.71% MIoU on the large-scale Stanford 2D-3D-S indoor benchmark without voxel quantization.

    In video sequence segmentation, Clockwork Convnet attains 64.40% MIoU on Cityscapes sequences and 68.50% MIoU on YouTube-Objects by exploiting temporal feature stability to reduce redundant computations across consecutive frames.

  7. Knowl 7 — Mathematical and Structural Properties of Dilated (À Trous) Convolutions

    model/method

    Dilated convolutions (also known as à trous convolutions) expand the receptive field of standard convolutional layers without reducing the spatial resolution of intermediate feature maps or adding parameters.

    For a 1D discrete input signal x[i]x[i] and a convolutional filter w[k]w[k] of length KK, the dilated convolution with dilation rate ℓ∈Z+\ell \in \mathbb{Z}^+ is defined as:

    y[i]=∑k=1Kx[i+ℓ⋅k]w[k]y[i] = \sum_{k=1}^{K} x[i + \ell \cdot k] w[k]

    In 2D spatial feature processing, a convolution with dilation rate ℓ\ell matches kernel weights to elements spaced ℓ\ell steps apart, effectively interleaving ℓ−1\ell - 1 zeros between consecutive filter taps without storing zero parameters.

    Stacking dilated convolutional layers with increasing dilation factors ℓ1,ℓ2,…,ℓm\ell_1, \ell_2, \dots, \ell_m expands the effective receptive field exponentially across depth, while preserving linear growth in parameters and avoiding pooling-induced spatial resolution loss.

  8. Knowl 8 — End-to-End Dense Labeling via Recurrent Formulation of Conditional Random Fields

    model/method

    To overcome the spatial oversmoothing and loss of boundary localization caused by successive pooling layers in FCNs, dense Conditional Random Fields (CRFs) can be structured as deep network layers rather than detached post-processing modules.

    In a dense pairwise CRF, the energy function over an image of pixel variables X={X1,…,XN}\mathbf{X} = \{X_1, \dots, X_N\} taking labels from L\mathcal{L} is defined as:

    E(x)=∑iψu(xi)+∑i<jψp(xi,xj)E(\mathbf{x}) = \sum_{i} \psi_u(x_i) + \sum_{i < j} \psi_p(x_i, x_j)

    where ψu(xi)=−log⁡P(li)\psi_u(x_i) = -\log P(l_i) denotes the unary potential provided by the deep network's pixel classifier scores, and ψp(xi,xj)\psi_p(x_i, x_j) denotes the pairwise potential modeling spatial and bilateral color interactions between all pixel pairs (i,j)(i, j) using Gaussian kernels:

    ψp(xi,xj)=μ(xi,xj)∑m=1Mw(m)k(m)(fi,fj)\psi_p(x_i, x_j) = \mu(x_i, x_j) \sum_{m=1}^{M} w^{(m)} k^{(m)}(\mathbf{f}_i, \mathbf{f}_j)

    where μ\mu is a label compatibility function, w(m)w^{(m)} are linear combination weights, and fi,fj\mathbf{f}_i, \mathbf{f}_j are feature vectors (such as spatial pixel coordinates and RGB color values).

    CRFasRNN reformulates the mean-field approximate inference algorithm—comprising message passing, compatibility transformation, unary potential addition, and softmax normalization—as a sequence of recurrent operations. This unrolling allows error gradients from downstream segmentation losses to backpropagate through the CRF inference steps, enabling end-to-end joint optimization of CNN feature extractors and CRF potentials.

  9. Knowl 9 — Evaluation and Reproducibility Limitations in Deep Semantic Segmentation Literature

    limitation

    A critical assessment of the deep semantic segmentation literature identifies three systematic methodological deficiencies:

    1. Omission of Computational Efficiency Metrics: The overwhelming majority of literature exclusively reports accuracy metrics (primarily MIoU and PA) while omitting execution time (inference latency), memory footprint (peak and average GPU/RAM consumption), and model parameter counts.

    2. Hardware and Environment Under-Specification: Where runtime metrics are provided, papers frequently fail to document hardware specifications (GPU/CPU models), batch sizes, input resolutions, and backend implementation environments, preventing reproducible benchmarking.

    3. Evaluation Protocol Inconsistencies: A substantial proportion of published methods evaluate on non-standard dataset partitions, employ private validation splits without public test submissions, or withhold implementation source code and pre-trained weights, creating barriers to objective cross-architecture comparison.

  10. Knowl 10 — Open Challenges and Future Directions in Deep Semantic Segmentation

    theoretical result

    Synthesizing the state of the art in deep semantic segmentation highlights six open technical challenges and research trajectories:

    1. Real-Time Inference vs. Accuracy Trade-Off: Existing high-accuracy architectures exhibit latencies incompatible with real-time requirements (≥25\ge 25 frames per second); for example, FCN-8s requires ≈100\approx 100 ms per low-resolution PASCAL VOC frame, while CRFasRNN requires >500> 500 ms. Specialized lightweight designs and structural pruning are required to bridge this gap.

    2. Graph Convolutions for 3D Point Clouds: Standard 3D CNNs rely on dense voxel grids that cause spatial discretization loss, quantization errors, and cubic memory growth (O(N3)O(N^3)). Formulating point cloud segmentation via Graph Convolutional Networks (GCNs) preserves raw geometric coordinates and continuous spatial cues.

    3. 3D and Temporal Dataset Scarcity: While 2D RGB segmentation benefits from extensive datasets, 3D point cloud and video sequence domains suffer from a deficit of large-scale, annotated real-world datasets, which restricts end-to-end spatiotemporal training.

    4. Temporal Coherency in Video Streams: Applying single-frame segmentation models independently across frames introduces high-frequency temporal flickering artifacts. Future video architectures require explicit temporal smoothing and inter-frame consistency constraints.

    5. Model Compression and Pruning for Embedded Platforms: Deployment in robotics and autonomous driving demands aggressive network compression, weight pruning, and quantization to fit tight on-chip memory bounds without sacrificing segmentation fidelity.

    6. Multi-View Integration: Multi-view 2.5D/3D segmentation remains largely confined to single-object settings, requiring general frameworks capable of multi-view geometric fusion across large-scale dynamic environments.

Coverage note — None was omitted; all contributed taxonomies, mathematical formulations, evaluation metric definitions, tabular performance benchmarks across 2D/2.5D/3D/video modalities, methodological critiques, and future research directions from the survey paper were fully captured.

References

  1. 1.A. Ess, T. Müller, H. Grabner, and L. J. Van Gool, “Segmentation-based urban traffic scene understanding.” in BMVC, vol. 1, 2009, p. 2.
  2. 2.A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in 2012 IEEE Conference on Computer Vision and Pattern Recognition, June 2012, pp. 3354–3361.
  3. 3.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3213–3223.
  4. 4.M. Oberweger, P. Wohlhart, and V. Lepetit, “Hands deep in deep learning for hand pose estimation,” arXiv preprint arXiv:1502.06807, 2015.
  5. 5.Y. Yoon, H.-G. Jeon, D. Yoo, J.-Y. Lee, and I. So Kweon, “ Learning a deep convolutional network for light-field image super-resolution,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 2015, pp. 24–32.
  6. 6.J. Wan, D. Wang, S. C. H. Hoi, P. Wu, J. Zhu, Y. Zhang, and J. Li, “Deep learning for content-based image retrieval: A comprehensive study,” in Proceedings of the 22nd ACM international conference on Multimedia. ACM, 2014, pp. 157–166.
  7. 7.F. Ning, D. Delhomme, Y. LeCun, F. Piano, L. Bottou, and P. E. Barbano, “ Toward automatic phenotyping of developing embryos from videos,” IEEE Transactions on Image Processing, vol. 14, no. 9, pp. 1360–1371, 2005.
  8. 8.D. Ciresan, A. Giusti, L. M. Gambardella, and J. Schmidhuber, “Deep neural networks segment neuronal membranes in electron microscopy images,” in Advances in neural information processing systems, 2012, pp. 2843–2851.
  9. 9.C. Farabet, C. Couprie, L. Najman, and Y. LeCun, “ Learning hierarchical features for scene labeling,” IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 8, pp. 1915–1929, 2013.
  10. 10.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik, “ Simultaneous detection and segmentation,” in European Conference on Computer Vision. Springer, 2014, pp. 297–312.
  11. 11.S. Gupta, R. Girshick, P. Arbelaez, and J. Malik, “ Learning rich features from rgb-d images for object detection and segmentation,” in European Conference on Computer Vision. Springer, 2014, pp. 345–360.
  12. 12.H. Zhu, F. Meng, J. Cai, and S. Lu, “ Beyond pixels: A comprehensive survey from bottom-up to semantic image segmentation and cosegmentation,” Journal of Visual Communication and Image Representation, vol. 34, pp. 12 – 27, 2016. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S1047320315002035
  13. 13.M. Thoma, “ A survey of semantic segmentation,” CoRR, vol. abs/1602.06541, 2016. [Online]. Available: http://arxiv.org/abs/1602.06541
  14. 14.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems, 2012, pp. 1097–1105.
  15. 15.K. Simonyan and A. Zisserman, “ Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  16. 16.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “ Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1–9.
  17. 17.K. He, X. Zhang, S. Ren, and J. Sun, “ Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778.
  18. 18.A. Graves, S. Fernández, and J. Schmidhuber, “ Multi-dimensional recurrent neural networks,” CoRR, vol. abs/0705.2011, 2007. [Online]. Available: http://arxiv.org/abs/0705.2011
  19. 19.F. Visin, K. Kastner, K. Cho, M. Matteucci, A. C. Courville, and Y. Bengio, “ Renet: A recurrent neural network based alternative to convolutional networks,” CoRR, vol. abs/1505.00393, 2015. [Online]. Available: http://arxiv.org/abs/1505.00393
  20. 20.A. Ahmed, K. Yu, W. Xu, Y. Gong, and E. Xing, “ Training hierarchical feed-forward visual recognition models using transfer learning from pseudo-tasks,” in European Conference on Computer Vision. Springer, 2008, pp. 69–82.
  21. 21.M. Oquab, L. Bottou, I. Laptev, and J. Sivic, “ Learning and transferring mid-level image representations using convolutional neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 1717–1724.
  22. 22.J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “ How transferable are features in deep neural networks?” in Advances in neural information processing systems, 2014, pp. 3320–3328.
  23. 23.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. IEEE, 2009, pp. 248–255.
  24. 24.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “ Imagenet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015.
  25. 25.S. C. Wong, A. Gatt, V. Stamatescu, and M. D. McDonnell, “Understanding data augmentation for classification: when to warp?” CoRR, vol. abs/1609.08764, 2016. [Online]. Available: http://arxiv.org/abs/1609.08764
  26. 26.X. Shen, A. Hertzmann, J. Jia, S. Paris, B. Price, E. Shechtman, and I. Sachs, “Automatic portrait segmentation for image stylization,” in Computer Graphics Forum, vol. 35, no. 2. Wiley Online Library, 2016, pp. 93–102.
  27. 27.M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “ The pascal visual object classes challenge: A retrospective,” International Journal of Computer Vision, vol. 111, no. 1, pp. 98–136, Jan. 2015.
  28. 28.R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille, “ The role of context for object detection and semantic segmentation in the wild,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
  29. 29.X. Chen, R. Mottaghi, X. Liu, S. Fidler, R. Urtasun, and A. Yuille, “Detect what you can: Detecting and representing objects using holistic models and body parts,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
  30. 30.B. Hariharan, P. Arbelaez, L. Bourdev, S. Maji, and J. Malik, “ Semantic contours from inverse detectors,” in 2011 International Conference on Computer Vision. IEEE, 2011, pp. 991–998.
  31. 31.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “ Microsoft coco: Common objects in context,” in European Conference on Computer Vision. Springer, 2014, pp. 740–755.
  32. 32.G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, “The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3234–3243.
  33. 33.M. Cordts, M. Omran, S. Ramos, T. Scharwächter, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “ The cityscapes dataset,” in CVPR Workshop on The Future of Datasets in Vision, 2015.
  34. 34.G. J. Brostow, J. Fauqueur, and R. Cipolla, “ Semantic object classes in video: A high-definition ground truth database,” Pattern Recognition Letters, vol. 30, no. 2, pp. 88–97, 2009.
  35. 35.P. Sturgess, K. Alahari, L. Ladicky, and P. H. Torr, “ Combining appearance and structure from motion features for road scene understanding,” in BMVC 2012-23rd British Machine Vision Conference. BMVA, 2009.
  36. 36.J. M. Alvarez, T. Gevers, Y. LeCun, and A. M. Lopez, “ Road scene segmentation from a single image,” in European Conference on Computer Vision. Springer, 2012, pp. 376–389.
  37. 37.G. Ros and J. M. Alvarez, “ Unsupervised image transformation for outdoor semantic labelling,” in Intelligent Vehicles Symposium (IV), 2015 IEEE. IEEE, 2015, pp. 537–542.
  38. 38.G. Ros, S. Ramos, M. Granados, A. Bakhtiary, D. Vazquez, and A. M. Lopez, “ Vision-based offline-online perception paradigm for autonomous driving,” in Applications of Computer Vision (WACV), 2015 IEEE Winter Conference on. IEEE, 2015, pp. 231–238.
  39. 39.R. Zhang, S. A. Candra, K. Vetter, and A. Zakhor, “ Sensor fusion for semantic segmentation of urban scenes,” in Robotics and Automation (ICRA), 2015 IEEE International Conference on. IEEE, 2015, pp. 1850–1857.
  40. 40.S. Gould, R. Fulton, and D. Koller, “ Decomposing a scene into geometric and semantically consistent regions,” in Computer Vision, 2009 IEEE 12th International Conference on. IEEE, 2009, pp. 1–8.
  41. 41.C. Liu, J. Yuen, and A. Torralba, “ Nonparametric scene parsing: Label transfer via dense scene alignment,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. IEEE, 2009, pp. 1972–1979.
  42. 42.S. D. Jain and K. Grauman, “ Supervoxel-consistent foreground propagation in video,” in European Conference on Computer Vision. Springer, 2014, pp. 656–671.
  43. 43.S. Bell, P. Upchurch, N. Snavely, and K. Bala, “ Material recognition in the wild with the materials in context database,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3479–3487.
  44. 44.F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “ A benchmark dataset and evaluation methodology for video object segmentation,” in Computer Vision and Pattern Recognition, 2016.
  45. 45.J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbelaez, A. Sorkine-Hornung, and L. Van Gool, “ The 2017 davis challenge on video object segmentation,” arXiv:1704.00675, 2017.
  46. 46.N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “ Indoor segmentation and support inference from rgbd images,” in European Conference on Computer Vision. Springer, 2012, pp. 746–760.
  47. 47.J. Xiao, A. Owens, and A. Torralba, “ Sun3d: A database of big spaces reconstructed using sfm and object labels,” in 2013 IEEE International Conference on Computer Vision, Dec 2013, pp. 1625–1632.
  48. 48.S. Song, S. P. Lichtenberg, and J. Xiao, “ Sun rgb-d: A rgb-d scene understanding benchmark suite,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 567–576.
  49. 49.K. Lai, L. Bo, X. Ren, and D. Fox, “ A large-scale hierarchical multi-view rgb-d object dataset,” in Robotics and Automation (ICRA), 2011 IEEE International Conference on. IEEE, 2011, pp. 1817–1824.
  50. 50.L. Yi, V. G. Kim, D. Ceylan, I.-C. Shen, M. Yan, H. Su, C. Lu, Q. Huang, A. Sheffer, and L. Guibas, “ A scalable active framework for region annotation in 3d shape collections,” SIGGRAPH Asia, 2016.
  51. 51.I. Armeni, A. Sax, A. R. Zamir, and S. Savarese, “ Joint 2D-3D-Semantic Data for Indoor Scene Understanding,” ArXiv e-prints, Feb. 2017.
  52. 52.X. Chen, A. Golovinskiy, and T. Funkhouser, “ A benchmark for 3D mesh segmentation,” ACM Transactions on Graphics (Proc. SIGGRAPH), vol. 28, no. 3, Aug. 2009.
  53. 53.A. Quadros, J. Underwood, and B. Douillard, “ An occlusion-aware feature for range images,” in Robotics and Automation, 2012. ICRA’12. IEEE International Conference on. IEEE, May 14-18 2012.
  54. 54.T. Hackel, J. D. Wegner, and K. Schindler, “ Contour detection in unstructured 3d point clouds,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 1610–1618.
  55. 55.G. J. Brostow, J. Shotton, J. Fauqueur, and R. Cipolla, “ Segmentation and recognition using structure from motion point clouds,” in European Conference on Computer Vision. Springer, 2008, pp. 44–57.
  56. 56.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “ Vision meets robotics: The kitti dataset,” The International Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013.
  57. 57.A. Prest, C. Leistner, J. Civera, C. Schmid, and V. Ferrari, “ Learning object class detectors from weakly annotated video,” in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on. IEEE, 2012, pp. 3282–3289.
  58. 58.S. Bell, P. Upchurch, N. Snavely, and K. Bala, “ OpenSurfaces: A richly annotated catalog of surface appearance,” ACM Trans. on Graphics (SIGGRAPH), vol. 32, no. 4, 2013.
  59. 59.B. C. Russell, A. Torralba, K. P. Murphy, and W. T. Freeman, “ Labelme: a database and web-based tool for image annotation,” International journal of computer vision, vol. 77, no. 1, pp. 157–173, 2008.
  60. 60.S. Gupta, P. Arbelaez, and J. Malik, “ Perceptual organization and recognition of indoor scenes from rgb-d images,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2013, pp. 564–571.
  61. 61.A. Janoch, S. Karayev, Y. Jia, J. T. Barron, M. Fritz, K. Saenko, and T. Darrell, A Category-Level 3D Object Dataset: Putting the Kinect to Work. London: Springer London, 2013, pp. 141–165. [Online]. Available: http://dx.doi.org/10.1007/978-1-4471-4640-7 8
  62. 62.A. Richtsfeld, “The object segmentation database (osd),” 2012.
  63. 63.A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su et al., “ Shapenet: An information-rich 3d model repository,” arXiv preprint arXiv:1512.03012, 2015.
  64. 64.I. Armeni, O. Sener, A. R. Zamir, H. Jiang, I. Brilakis, M. Fischer, and S. Savarese, “3d semantic parsing of large-scale indoor spaces,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 1534–1543.
  65. 65.J. Long, E. Shelhamer, and T. Darrell, “ Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431–3440.
  66. 66.V. Badrinarayanan, A. Kendall, and R. Cipolla, “ Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” arXiv preprint arXiv:1511.00561, 2015.
  67. 67.A. Kendall, V. Badrinarayanan, and R. Cipolla, “ Bayesian segnet: Model uncertainty in deep convolutional encoder-decoder architectures for scene understanding,” arXiv preprint arXiv:1511.02680, 2015.
  68. 68.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “ Semantic image segmentation with deep convolutional nets and fully connected crfs,” arXiv preprint arXiv:1412.7062, 2014.
  69. 69.——, “ Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” arXiv preprint arXiv:1606.00915, 2016.
  70. 70.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr, “ Conditional random fields as recurrent neural networks,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1529–1537.
  71. 71.F. Yu and V. Koltun, “ Multi-scale context aggregation by dilated convolutions,” arXiv preprint arXiv:1511.07122, 2015.
  72. 72.A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ Enet: A deep neural network architecture for real-time semantic segmentation,” arXiv preprint arXiv:1606.02147, 2016.
  73. 73.A. Raj, D. Maturana, and S. Scherer, “ Multi-scale convolutional architecture for semantic segmentation,” 2015.
  74. 74.D. Eigen and R. Fergus, “ Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2650–2658.
  75. 75.A. Roy and S. Todorovic, “ A multi-scale cnn for affordance segmentation in rgb images,” in European Conference on Computer Vision. Springer, 2016, pp. 186–201.
  76. 76.X. Bian, S. N. Lim, and N. Zhou, “ Multiscale fully convolutional network with application to industrial inspection,” in Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on. IEEE, 2016, pp. 1–8.
  77. 77.W. Liu, A. Rabinovich, and A. C. Berg, “ Parsenet: Looking wider to see better,” arXiv preprint arXiv:1506.04579, 2015.
  78. 78.F. Visin, M. Ciccone, A. Romero, K. Kastner, K. Cho, Y. Bengio, M. Matteucci, and A. Courville, “ Reseg: A recurrent neural network-based model for semantic segmentation,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2016.
  79. 79.Z. Li, Y. Gan, X. Liang, Y. Yu, H. Cheng, and L. Lin, LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling. Cham: Springer International Publishing, 2016, pp. 541–557. [Online]. Available: http://dx.doi.org/10.1007/978-3-319-46475-6 34
  80. 80.W. Byeon, T. M. Breuel, F. Raue, and M. Liwicki, “ Scene labeling with lstm recurrent neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3547–3555.
  81. 81.P. H. Pinheiro and R. Collobert, “ Recurrent convolutional neural networks for scene labeling.” in ICML, 2014, pp. 82–90.
  82. 82.B. Shuai, Z. Zuo, G. Wang, and B. Wang, “ Dag-recurrent neural networks for scene labeling,” CoRR, vol. abs/1509.00552, 2015. [Online]. Available: http://arxiv.org/abs/1509.00552
  83. 83.P. O. Pinheiro, R. Collobert, and P. Dollar, “ Learning to segment object candidates,” in Advances in Neural Information Processing Systems, 2015, pp. 1990–1998.
  84. 84.P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollár, “ Learning to refine object segments,” in European Conference on Computer Vision. Springer, 2016, pp. 75–91.
  85. 85.S. Zagoruyko, A. Lerer, T.-Y. Lin, P. O. Pinheiro, S. Gross, S. Chintala, and P. Dollár, “ A multipath network for object detection,” arXiv preprint arXiv:1604.02135, 2016.
  86. 86.J. Huang and S. You, “ Point cloud labeling using 3d convolutional neural network,” in Proc. of the International Conf. on Pattern Recognition (ICPR), vol. 2, 2016.
  87. 87.C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “ Pointnet: Deep learning on point sets for 3d classification and segmentation,” arXiv preprint arXiv:1612.00593, 2016.
  88. 88.E. Shelhamer, K. Rakelly, J. Hoffman, and T. Darrell, “ Clockwork convnets for video semantic segmentation,” in Computer Vision–ECCV 2016 Workshops. Springer, 2016, pp. 852–868.
  89. 89.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “ Deep end2end voxel2voxel prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016, pp. 17–24.
  90. 90.M. D. Zeiler, G. W. Taylor, and R. Fergus, “ Adaptive deconvolutional networks for mid and high level feature learning,” in Computer Vision (ICCV), 2011 IEEE International Conference on. IEEE, 2011, pp. 2018–2025.
  91. 91.M. D. Zeiler and R. Fergus, “ Visualizing and understanding convolutional networks,” in European conference on computer vision. Springer, 2014, pp. 818–833.
  92. 92.C. Rother, V. Kolmogorov, and A. Blake, “ Grabcut: Interactive foreground extraction using iterated graph cuts,” in ACM transactions on graphics (TOG), vol. 23, no. 3. ACM, 2004, pp. 309–314.
  93. 93.J. Shotton, J. Winn, C. Rother, and A. Criminisi, “ Textonboost for image understanding: Multi-class object recognition and segmentation by jointly modeling texture, layout, and context,” International Journal of Computer Vision, vol. 81, no. 1, pp. 2–23, 2009.
  94. 94.V. Koltun, “ Efficient inference in fully connected crfs with gaussian edge potentials,” Adv. Neural Inf. Process. Syst, vol. 2, no. 3, p. 4, 2011.
  95. 95.P. Krähenbühl and V. Koltun, “ Parameter learning and convergent inference for dense random fields.” in ICML (3), 2013, pp. 513–521.
  96. 96.S. Zhou, J.-N. Wu, Y. Wu, and X. Zhou, “ Exploiting local structures with the kronecker layer in convolutional networks,” arXiv preprint arXiv:1512.09194, 2015.
  97. 97.S. Hochreiter and J. Schmidhuber, “ Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  98. 98.K. Cho, B. Van Merrienboer, D. Bahdanau, and Y. Bengio, “ On the properties of neural machine translation: Encoder-decoder approaches,” arXiv preprint arXiv:1409.1259, 2014.
  99. 99.Z. Li, Y. Gan, X. Liang, Y. Yu, H. Cheng, and L. Lin, “ RGB-D scene labeling with long short-term memorized fusion model,” CoRR, vol. abs/1604.05000, 2016. [Online]. Available: http://arxiv.org/abs/1604.05000
  100. 100.G. Li and Y. Yu, “ Deep contrast learning for salient object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 478–487.
  101. 101.P. Arbelaez, J. Pont-Tuset, J. T. Barron, F. Marques, and J. Malik, “ Multiscale combinatorial grouping,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 328–335.
  102. 102.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “ Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
  103. 103.A. Zeng, K. Yu, S. Song, D. Suo, E. W. Jr., A. Rodriguez, and J. Xiao, “ Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge,” CoRR, vol. abs/1609.09475, 2016. [Online]. Available: http://arxiv.org/abs/1609.09475
  104. 104.L. Ma, J. Stuckler, C. Kerl, and D. Cremers, “ Multi-view deep learning for consistent semantic mapping with rgb-d cameras,” in arXiv:1703.08866, Mar 2017.
  105. 105.C. Hazirbas, L. Ma, C. Domokos, and D. Cremers, “ Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture,” in Proc. ACCV, vol. 2, 2016.
  106. 106.H. Zhang, K. Jiang, Y. Zhang, Q. Li, C. Xia, and X. Chen, “ Discriminative feature learning for video semantic segmentation,” in Virtual Reality and Visualization (ICVRV), 2014 International Conference on. IEEE, 2014, pp. 321–326.
  107. 107.Y. Boykov, O. Veksler, and R. Zabih, “ Fast approximate energy minimization via graph cuts,” IEEE Transactions on pattern analysis and machine intelligence, vol. 23, no. 11, pp. 1222–1239, 2001.
  108. 108.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “ Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 4489–4497.
  109. 109.M. Henaff, J. Bruna, and Y. LeCun, “ Deep convolutional networks on graph-structured data,” arXiv preprint arXiv:1506.05163, 2015.
  110. 110.T. N. Kipf and M. Welling, “ Semi-supervised classification with graph convolutional networks,” arXiv preprint arXiv:1609.02907, 2016.
  111. 111.M. Niepert, M. Ahmed, and K. Kutzkov, “ Learning convolutional neural networks for graphs,” in Proceedings of the 33rd annual international conference on machine learning. ACM, 2016.
  112. 112.S. Anwar, K. Hwang, and W. Sung, “ Structured pruning of deep convolutional neural networks,” arXiv preprint arXiv:1512.08571, 2015.
  113. 113.S. Han, H. Mao, and W. J. Dally, “ Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149, 2015.
  114. 114.P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz, “ Pruning convolutional neural networks for resource efficient transfer learning,” arXiv preprint arXiv:1611.06440, 2016.

Citation

MLA
Garcia-Garcia, A., et al. “A Review on Deep Learning Techniques Applied to Semantic Segmentation”. arXiv, 2017, http://arxiv.org/abs/1704.06857v1.
APA
Garcia-Garcia, A., Orts-Escolano, S., Oprea, S., Villena-Martinez, V., & Garcia-Rodriguez, J. (2017). A Review on Deep Learning Techniques Applied to Semantic Segmentation. arXiv. http://arxiv.org/abs/1704.06857v1
Chicago
Garcia-Garcia, A., S. Orts-Escolano, S. Oprea, V. Villena-Martinez, and J. Garcia-Rodriguez. 2017. “A Review on Deep Learning Techniques Applied to Semantic Segmentation”. arXiv. http://arxiv.org/abs/1704.06857v1.
Harvard
Garcia-Garcia, A. et al. (2017) “A Review on Deep Learning Techniques Applied to Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1704.06857v1.
Vancouver
1. Garcia-Garcia A, Orts-Escolano S, Oprea S, Villena-Martinez V, Garcia-Rodriguez J (2017) A Review on Deep Learning Techniques Applied to Semantic Segmentation. arXiv

BibTeX

@article{garciagarcia2017review,
  title = {A Review on Deep Learning Techniques Applied to Semantic Segmentation},
  author = {Garcia-Garcia, Alberto and Orts-Escolano, Sergio and Oprea, Sergiu and Villena-Martinez, Victor and Garcia-Rodriguez, Jose},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1704.06857v1},
  eprint = {1704.06857}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors