ICNet for Real-Time Semantic Segmentation on High-Resolution Images

Hengshuang ZhaoXiaojuan QiXiaoyong ShenJianping ShiJiaya Jia

article2017ECCV1,624 citations

Introduces an image cascade network that combines multi-resolution branches with cascade feature fusion to deliver real-time, high-accuracy semantic segmentation on high-resolution images.

Listen

Real-time computer vision is essential for autonomous vehicles, robotics, and mobile devices, where systems must quickly classify every pixel in an image to navigate and interact with environments safely. While modern deep learning models achieve high accuracy, they require heavy computational power and typically take roughly one second to process a single high-resolution image on standard hardware. Conversely, existing lightweight alternatives run quickly but suffer substantial drops in accuracy, creating a critical bottleneck for time-sensitive, safety-critical applications.

The article aims to evaluate and demonstrate the Image Cascade Network (ICNet), a new framework designed to achieve real-time semantic segmentation on high-resolution images while maintaining competitive prediction quality. To establish credibility, the authors tested the system across three benchmark datasets—Cityscapes (street scenes at 1024x2048 resolution), CamVid (720x960), and COCO-Stuff (640x640)—measuring processing speed in frames per second and accuracy using the standard mean intersection-over-union metric on a single graphics processing unit.

The findings show that standard speedup tactics—such as downsampling input images, reducing internal feature maps, or pruning model parameters—either fail to achieve real-time speeds or severely degrade segmentation details. In contrast, the cascading multi-resolution approach processes low-resolution inputs through a deeper network to extract broad context, then uses lightweight layers and a custom fusion unit to restore fine details at higher resolutions. On the challenging Cityscapes dataset, the system achieved a 5-fold speedup (processing 1024x2048 images at over 30 frames per second, or 33 milliseconds per frame) and reduced memory usage by over 5-fold compared to baseline architectures. It reached 69.5% accuracy (and 70.6% with supplemental training data), outperforming prior real-time models by roughly 10 percentage points and performing comparably to several computationally intensive models.

These results demonstrate that high-resolution visual perception does not require costly multi-GPU setups to achieve real-time performance, significantly lowering hardware expenses, energy demands, and latency risks in deployment. The framework provides a practical design template for engineering teams developing embedded vision systems. Organizations seeking to deploy efficient vision models should consider adopting multi-resolution cascading architectures and evaluating the open-source code base for their specific edge-device constraints. However, decision-makers should note that the system still trails top-tier, non-real-time models by roughly 10 percentage points in overall accuracy, meaning applications with zero tolerance for boundary errors should conduct targeted validation on fine-grained objects before full deployment.

Cover for ICNet for Real-Time Semantic Segmentation on High-Resolution Images

Abstract

We focus on the challenging task of real-time semantic segmentation in this paper. It finds many practical applications and yet is with fundamental difficulty of reducing a large portion of computation for pixel-wise label inference. We propose an image cascade network (ICNet) that incorporates multi-resolution branches under proper label guidance to address this challenge. We provide in-depth analysis of our framework and introduce the cascade feature fusion unit to quickly achieve high-quality segmentation. Our system yields real-time inference on a single GPU card with decent quality results evaluated on challenging datasets like Cityscapes, CamVid and COCO-Stuff.

Table of Contents

  • 1 Introduction
  • Status of Fast Semantic Segmentation
  • Our Focus and Contributions
  • 2 Related Work
  • High Quality Semantic Segmentation
  • High Efficiency Semantic Segmentation
  • Video Semantic Segmentation
  • 3 Image Cascade Network
  • 3.1 Speed Analysis
  • 3.2 Network Architecture
  • 3.3 Cascade Feature Fusion
  • 3.4 Cascade Label Guidance
  • 4 Structure Comparison and Analysis
  • 5 Experimental Evaluation
  • 5.1 Implementation Details
  • 5.2 Cityscapes
  • Intuitive Speedup
  • Cascade Branches
  • Cascade Structure
  • Methods Comparison
  • Visual Improvement
  • Quantitative Analysis
  • 5.3 CamVid
  • 5.4 COCO-Stuff
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Image Cascade Network (ICNet) Architecture

    model/method

    The Image Cascade Network (ICNet) is a real-time semantic segmentation architecture designed for high-resolution images (such as 1024×20481024 \times 2048 pixels). It processes an input image across three multi-resolution cascade branches:

    1. Low-Resolution Branch: The full-resolution input image is downsampled by a factor of 4 (1/41/4 spatial resolution) and passed into a deep segmentation network (such as PSPNet50 with downsampling rate 8), resulting in a coarse feature map at 1/321/32 resolution relative to the original image. This branch extracts global context and high-level semantic layout at low computational cost.
    2. Medium-Resolution Branch: The input image is downsampled by a factor of 2 (1/21/2 spatial resolution) and passed through a lightweight convolutional network (sharing the first 17 convolutional layers with the low-resolution branch), yielding a feature map at 1/161/16 resolution. This representation is combined with the low-resolution branch output via a Cascade Feature Fusion (CFF) unit.
    3. High-Resolution Branch: The full-resolution input image (1×1\times) is fed into a shallow, lightweight convolutional network with downsampling rate 8, producing an output feature map at 1/81/8 resolution. It is merged with the fused medium/low-resolution feature map using a second CFF unit.

    During testing, a final 4×4\times bilinear upsampling is applied to the output of the high-resolution branch to produce the full-resolution semantic segmentation prediction map.

  2. Knowl 2 — Cascade Feature Fusion (CFF) Unit

    model/method

    The Cascade Feature Fusion (CFF) unit combines feature maps from adjacent cascade branches of differing resolutions to refine coarse semantic predictions:

    • Inputs: A lower-resolution feature map F1∈RC1×H1×W1F_1 \in \mathbb{R}^{C_1 \times H_1 \times W_1} from a lower branch, a higher-resolution feature map F2∈RC2×H2×W2F_2 \in \mathbb{R}^{C_2 \times H_2 \times W_2} from an upper branch (where H2=2H1H_2 = 2H_1 and W2=2W1W_2 = 2W_1), and a corresponding downsampled ground-truth label map of dimension 1×H2×W21 \times H_2 \times W_2.
    • Lower-Resolution Processing: F1F_1 is upsampled by a factor of 2 via bilinear interpolation to match the spatial dimensions H2×W2H_2 \times W_2, followed by a dilated convolution with kernel size C3×3×3C_3 \times 3 \times 3 and dilation rate 2, and batch normalization. Dilated convolution achieves an effective 7×77 \times 7 receptive field with fewer parameters and operations than a 7×77 \times 7 deconvolution.
    • Higher-Resolution Processing: F2F_2 undergoes a C3×1×1C_3 \times 1 \times 1 projection convolution and batch normalization to align its channel depth to C3C_3.
    • Fusion: The normalized features are combined via element-wise addition and a rectified linear unit (ReLU) activation to produce the fused feature map:

    F2′=ReLU(BN(DilatedConv3×3,d=2(2×F1))+BN(Conv1×1(F2)))∈RC3×H2×W2F_2' = \text{ReLU}\left( \text{BN}(\text{DilatedConv}_{3\times 3, d=2}(2\times F_1)) + \text{BN}(\text{Conv}_{1\times 1}(F_2)) \right) \in \mathbb{R}^{C_3 \times H_2 \times W_2}

    • Auxiliary Guidance: An auxiliary classifier with a 1×11 \times 1 convolution is appended to the upsampled F1F_1 feature map to compute an auxiliary loss against the ground-truth label map during training.
  3. Knowl 3 — Cascade Label Guidance Loss Formulation

    equation

    The Cascade Label Guidance (CLG) training objective optimizes the multi-resolution branches simultaneously using ground-truth segmentation masks downsampled to the corresponding feature map scales (1/161/16, 1/81/8, and 1/41/4 of the original resolution):

    L=−∑t=1Tλt1YtXt∑y=1Yt∑x=1Xtlog⁡exp⁡(Fn^,y,xt)∑n=1Nexp⁡(Fn,y,xt)\mathcal{L} = -\sum_{t=1}^T \lambda_t \frac{1}{Y_t X_t} \sum_{y=1}^{Y_t} \sum_{x=1}^{X_t} \log \frac{\exp\left(F^t_{\hat{n}, y, x}\right)}{\sum_{n=1}^N \exp\left(F^t_{n, y, x}\right)}

    where:

    • T=3T = 3 is the total number of cascade branches (low, medium, and high resolution).
    • NN is the number of semantic target classes.
    • YtY_t and XtX_t are the spatial height and width of the predicted feature logit map FtF^t in branch tt.
    • Fn,y,xt∈RF^t_{n, y, x} \in \mathbb{R} is the unnormalized logit score for class nn at spatial coordinates (y,x)(y, x) in branch tt.
    • n^∈{1,…,N}\hat{n} \in \{1, \dots, N\} is the ground-truth class label at spatial coordinates (y,x)(y, x) for branch tt.
    • λt>0\lambda_t > 0 is the auxiliary loss weight for branch tt, set empirically to λ1=0.4\lambda_1 = 0.4 for the low-resolution branch (1/161/16 scale), λ2=0.4\lambda_2 = 0.4 for the medium-resolution branch (1/81/8 scale), and λ3=1.0\lambda_3 = 1.0 for the final high-resolution branch (1/41/4 scale).

    During inference, the low- and medium-resolution guidance heads are discarded, and only the high-resolution output is evaluated.

  4. Knowl 4 — Cityscapes Benchmark Real-Time Segmentation Performance

    data/table

    The table below compares the mean Intersection-over-Union (mIoU), forward inference time, and frame rate (frames per second, fps) on the Cityscapes test dataset at 1024×20481024 \times 2048 resolution using a single Nvidia TitanX GPU:

    Method Downsampling Ratio (DR) mIoU (%) Time (ms) Frame (fps)
    SegNet 4 57.0 60 16.7
    ENet 2 58.3 13 76.9
    SQ No 59.8 60 16.7
    CRF-RNN 2 62.5 700 1.4
    DeepLab 2 63.1 4000 0.25
    FCN-8S No 65.3 500 2.0
    Dilation10 No 67.1 4000 0.25
    FRRN 2 71.8 469 2.1
    PSPNet No 81.2 1288 0.78
    ICNet No 69.5 33 30.3
    ICNet†^\dagger No 70.6 33 30.3

    Methods trained using both fine and coarse annotations are denoted with †\dagger. The Downsampling Ratio (DR) specifies image downsampling applied during testing (e.g., DR = 4 tests at 256×512256 \times 512). ICNet processes the full 1024×20481024 \times 2048 image without test-time downsampling, achieving real-time performance of 30.3 fps30.3\text{ fps} (33 ms33\text{ ms}) with 69.5% mIoU69.5\%\text{ mIoU} (70.6% mIoU70.6\%\text{ mIoU} with coarse data), outperforming prior real-time methods (ENet and SQ) by approximately 10% mIoU10\%\text{ mIoU}.

  5. Knowl 5 — Multi-Resolution Cascade Branch Ablation in ICNet

    data/table

    An ablation study on the Cityscapes validation set (1024×20481024 \times 2048) evaluates the contribution of each cascade branch in ICNet compared to a baseline model (half-compressed PSPNet50):

    Items Baseline sub4 sub24 sub124
    mIoU (%) 67.9 59.6 66.5 67.7
    Time (ms) 170 18 25 33
    Frame (fps) 5.9 55.6 40.0 30.3
    Speedup 1×1\times 9.4×9.4\times 6.8×6.8\times 5.2×5.2\times
    Memory (GB) 9.2 0.6 1.1 1.6
    Memory save 1×1\times 15.3×15.3\times 8.4×8.4\times 5.8×5.8\times

    Here:

    • Baseline is PSPNet50 with channels compressed by half.
    • sub4 uses only the low-resolution (1/41/4 scale) branch.
    • sub24 incorporates both low- and medium-resolution (1/21/2 scale) branches.
    • sub124 represents the complete three-branch ICNet architecture.

    While sub4 runs at 55.6 fps55.6\text{ fps}, its accuracy drops to 59.6% mIoU59.6\%\text{ mIoU}. Progressively incorporating medium- and high-resolution branches (sub24 and sub124) recovers accuracy to 66.5%66.5\% and 67.7% mIoU67.7\%\text{ mIoU} with marginal time increments (+7 ms+7\text{ ms} and +8 ms+8\text{ ms}), achieving 5.2×5.2\times speedup and 5.8×5.8\times memory reduction compared to the baseline.

  6. Knowl 6 — Ablation Analysis on Fusion Units and Auxiliary Label Guidance

    data/table

    An ablation study on the Cityscapes validation set evaluates the Cascade Feature Fusion (CFF) unit versus standard deconvolution operations of varying kernel sizes, as well as the impact of Cascade Label Guidance (CLG):

    DC3 DC5 DC7 CFF CLG mIoU (%) Time (ms)
    ✓ ✓ 66.7 31
    ✓ ✓ 66.7 34
    ✓ ✓ 68.0 38
    ✓ ✓ 67.7 33
    ✓ 66.8 33

    In the table:

    • DC3, DC5, and DC7 denote replacing the bilinear upsampling and dilated convolution in CFF with standard deconvolution using kernel sizes 3×33 \times 3, 5×55 \times 5, and 7×77 \times 7, respectively.
    • CFF denotes the cascade feature fusion unit utilizing bilinear upsampling followed by a 3×33 \times 3 dilated convolution with dilation rate 2.
    • CLG denotes training with auxiliary cascade label guidance.

    CFF matches the accuracy of a large 7×77 \times 7 deconvolution layer (67.7%67.7\% vs 68.0% mIoU68.0\%\text{ mIoU}) while reducing inference time from 38 ms38\text{ ms} to 33 ms33\text{ ms}. Training without cascade label guidance degrades performance by 0.9% mIoU0.9\%\text{ mIoU} (66.8%66.8\% vs 67.7%67.7\%) without changing inference speed.

  7. Knowl 7 — Computational Complexity of Convolutional Feature Maps in Segmentation

    equation

    For a 2D convolutional transformation Φ:V→U\Phi: V \to U applied to an input feature map V∈Rc×h×wV \in \mathbb{R}^{c \times h \times w} to produce an output feature map U∈Rc′×h′×w′U \in \mathbb{R}^{c' \times h' \times w'}, using c′c' kernels of size c×k×kc \times k \times k with spatial stride ss (such that h′=h/sh' = h/s and w′=w/sw' = w/s), the total computational operation count is:

    O(Φ)≈c′ck2hws2O(\Phi) \approx \frac{c' c k^2 h w}{s^2}

    where:

    • c∈N+c \in \mathbb{N}^+ is the number of input channels.
    • c′∈N+c' \in \mathbb{N}^+ is the number of output channels (filters).
    • k∈N+k \in \mathbb{N}^+ is the kernel spatial dimension (k×kk \times k).
    • h,w∈N+h, w \in \mathbb{N}^+ are the input height and width.
    • s∈N+s \in \mathbb{N}^+ is the convolutional stride.

    The computational cost scales quadratically with input spatial dimensions (h⋅wh \cdot w). In dilated FCN architectures (such as PSPNet50), stage 4 and stage 5 maintain an identical spatial resolution (1/81/8 of the original input), but stage 5 is approximately four times computationally heavier than stage 4 because both the input channel depth cc and output filter count c′c' are doubled.

  8. Knowl 8 — Benchmark Results on CamVid and COCO-Stuff

    data/table

    Evaluation of ICNet against baseline segmentation networks on CamVid (resolution 720×960720 \times 960, 11 classes) and COCO-Stuff (resolution 640×640640 \times 640, 182 classes):

    CamVid (720×960720 \times 960) COCO-Stuff (640×640640 \times 640)
    Method mIoU (%) Time (ms) Frame (fps) Method mIoU (%) Time (ms) Frame (fps)
    SegNet 46.4 217 4.6 FCN 22.7 169 5.9
    DPN 60.1 830 1.2 DeepLab 26.9 124 8.1
    DeepLab 61.6 203 4.9 PSPNet50 32.6 151 6.6
    Dilation8 65.3 227 4.4 ICNet 29.1 28 35.7
    PSPNet50 69.1 185 5.4
    ICNet 67.1 36 27.8

    On CamVid, ICNet runs at 27.8 fps27.8\text{ fps} (36 ms36\text{ ms}), providing a 5.1×5.1\times speedup over uncompressed PSPNet50 (69.1% mIoU69.1\%\text{ mIoU} at 185 ms185\text{ ms}) while retaining 67.1% mIoU67.1\%\text{ mIoU}. On COCO-Stuff, ICNet achieves 35.7 fps35.7\text{ fps} (28 ms28\text{ ms}) with 29.1% mIoU29.1\%\text{ mIoU}, yielding a 5.4×5.4\times speedup over PSPNet50 (32.6% mIoU32.6\%\text{ mIoU} at 151 ms151\text{ ms}) and surpassing FCN (22.7%22.7\%) and DeepLab (26.9%26.9\%) in both speed and accuracy.

  9. Knowl 9 — Connected Component Scale Analysis for Cascaded Refinement

    empirical result

    To quantify the role of multi-resolution cascade branches in ICNet, segmentation accuracy is evaluated across ground-truth connected components grouped by pixel area Si∈[1,90000]S_i \in [1, 90000]. Components are divided into 30 bins with an interval width of K=3000K = 3000 pixels. For each region RiR_i, accuracy is computed as pi=si/Sip_i = s_i / S_i, where sis_i is the count of correctly predicted pixels in RiR_i, and the mean accuracy per bin is tracked across branch outputs:

    • The difference between sub24 (low + medium branches) and sub4 (low branch only) shows large positive accuracy gains concentrated in the lowest area bins (Si≤9000S_i \le 9000 pixels).
    • The difference between sub124 (all three branches) and sub24 provides further positive gains primarily in the smallest bins.

    This indicates that while the low-resolution branch successfully identifies large semantic regions (e.g., roads, buildings, sky), the medium- and high-resolution cascade branches specifically recover small and thin structural objects (such as traffic lights, poles, traffic signs, and distant pedestrians) whose features are lost during initial aggressive downsampling.

  10. Knowl 10 — Empirical Limits of Naive Acceleration Strategies for Semantic Segmentation

    data/table

    Direct strategies to accelerate segmentation models on high-resolution images (1024×20481024 \times 2048 Cityscapes validation set with PSPNet50) include input image downsampling, feature map downsampling, and filter pruning (ℓ1\ell_1-norm model compression):

    Downsampling Feature Map Model Compression (Kernel Pruning)
    Downsample Factor 8 16 32 Kernel Keeping Rate 1.0 0.5 0.25
    mIoU (%) 71.7 70.2 67.1 mIoU (%) 71.7 67.9 59.4
    Time (ms) 446 177 131 Time (ms) 446 170 72

    Additionally, evaluating input image downsampling on PSPNet50 yields:

    • Full scale (1.0×1.0\times): 71.7% mIoU71.7\%\text{ mIoU} at 446 ms446\text{ ms} (2.2 fps2.2\text{ fps}).
    • Half scale (0.5×0.5\times): 68.4% mIoU68.4\%\text{ mIoU} at 123 ms123\text{ ms} (8.1 fps8.1\text{ fps}).
    • Quarter scale (0.25×0.25\times): 60.7% mIoU60.7\%\text{ mIoU} at 42 ms42\text{ ms} (23.8 fps23.8\text{ fps}).

    None of these naive modifications reach real-time inference (≥30 fps\ge 30\text{ fps} or ≤33 ms\le 33\text{ ms}) without severe accuracy degradation: reducing input scale to 0.25×0.25\times or pruning 75%75\% of kernels fails to achieve 30 fps while dropping mIoU below 61%61\%.

Coverage note — None was omitted; all key contributions of ICNet—including network architecture, the CFF fusion unit, the cascade label guidance objective, convolution cost analysis, and experimental evaluations across Cityscapes, CamVid, and COCO-Stuff—are represented.

References

  1. 1.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: CVPR (2015)
  2. 2.Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Semantic image segmentation with deep convolutional nets and fully connected CRFs. In: ICLR (2015)
  3. 3.Badrinarayanan, V., Kendall, A., Cipolla, R.: SegNet: a deep convolutional encoder-decoder architecture for image segmentation. arXiv:1511.00561 (2015)
  4. 4.Noh, H., Hong, S., Han, B.: Learning deconvolution network for semantic segmentation. In: ICCV (2015)
  5. 5.Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: CVPR (2017)
  6. 6.Wu, Z., Shen, C., van den Hengel, A.: Wider or deeper: revisiting the ResNet model for visual recognition. arXiv:1611.10080 (2016)
  7. 7.Cordts, M., et al.: The cityscapes dataset for semantic urban scene understanding. In: CVPR (2016)
  8. 8.Paszke, A., Chaurasia, A., Kim, S., Culurciello, E.: ENet: a deep neural network architecture for real-time semantic segmentation. arXiv:1606.02147 (2016)
  9. 9.Treml, M., et al.: Speeding up semantic segmentation for autonomous driving. In: NIPS Workshop (2016)
  10. 10.Wang, P., et al.: Understanding convolution for semantic segmentation. arXiv:1702.08502 (2017)
  11. 11.Lin, G., Milan, A., Shen, C., Reid, I.D.: RefineNet: multi-path refinement networks for high-resolution semantic segmentation. In: CVPR (2017)
  12. 12.Pohlen, T., Hermans, A., Mathias, M., Leibe, B.: Full-resolution residual networks for semantic segmentation in street scenes. In: CVPR (2017)
  13. 13.Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. arXiv:1606.00915 (2016)
  14. 14.Yu, F., Koltun, V.: Multi-scale context aggregation by dilated convolutions. In: ICLR (2016)
  15. 15.Liu, Z., Li, X., Luo, P., Loy, C.C., Tang, X.: Semantic image segmentation via deep parsing network. In: ICCV (2015)
  16. 16.Zheng, S., et al.: Conditional random fields as recurrent neural networks. In: ICCV (2015)
  17. 17.Brostow, G.J., Fauqueur, J., Cipolla, R.: Semantic object classes in video: a high-definition ground truth database. Pattern Recognit. Lett. 30, 88–97 (2009)
  18. 18.Caesar, H., Uijlings, J., Ferrari, V.: Coco-stuff: thing and stuff classes in context. arXiv:1612.03716 (2016)
  19. 19.Liu, C., Yuen, J., Torralba, A.: Nonparametric scene parsing via label transfer. TPAMI 33, 2368–2382 (2011)
  20. 20.Chen, L., Yang, Y., Wang, J., Xu, W., Yuille, A.L.: Attention to scale: scale-aware semantic image segmentation. In: CVPR (2016)
  21. 21.Hariharan, B., Arbel´aez, P.A., Girshick, R.B., Malik, J.: Hypercolumns for object segmentation and fine-grained localization. In: CVPR (2015)
  22. 22.Xia, F., Wang, P., Chen, L.-C., Yuille, A.L.: Zoom better to see clearer: human and object parsing with hierarchical auto-zoom net. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9909, pp. 648–663. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46454-1_39
  23. 23.Girshick, R.: Fast R-CNN. In: ICCV (2015)
  24. 24.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: towards real-time object detection with region proposal networks. In: NIPS. (2015)
  25. 25.Redmon, J., Divvala, S.K., Girshick, R.B., Farhadi, A.: You only look once: unified, real-time object detection. In: CVPR (2016)
  26. 26.Redmon, J., Farhadi, A.: YOLO9000: better, faster, stronger. In: CVPR (2017)
  27. 27.Liu, W., et al.: SSD: single shot multibox detector. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9905, pp. 21–37. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46448-0_2
  28. 28.Romera, E., Alvarez, J.M., Bergasa, L.M., Arroyo, R.: Efficient ConvNet for real-time semantic segmentation. In: Intelligent Vehicles Symposium (IV) (2017)
  29. 29.Shelhamer, E., Rakelly, K., Hoffman, J., Darrell, T.: Clockwork convnets for video semantic segmentation. In: Hua, G., J´egou, H. (eds.) ECCV 2016. LNCS, vol. 9915, pp. 852–868. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-49409-8_69
  30. 30.Zhu, X., Xiong, Y., Dai, J., Yuan, L., Wei, Y.: Deep feature flow for video recognition. In: CVPR (2017)
  31. 31.Kundu, A., Vineet, V., Koltun, V.: Feature space optimization for semantic video segmentation. In: CVPR (2016)
  32. 32.Gadde, R., Jampani, V., Gehler, P.V.: Semantic video CNNs through representation warping. In: ICCV (2017)
  33. 33.Ronneberger, O., Fischer, P., Brox, T.: U-net: convolutional networks for biomedical image segmentation. In: Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F. (eds.) MICCAI 2015. LNCS, vol. 9351, pp. 234–241. Springer, Cham (2015). https://doi.org/10.1007/978-3-319-24574-4_28
  34. 34.Ghiasi, G., Fowlkes, C.C.: Laplacian pyramid reconstruction and refinement for semantic segmentation. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9907, pp. 519–534. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46487-9_32
  35. 35.Everingham, M., Gool, L.J.V., Williams, C.K.I., Winn, J.M., Zisserman, A.: The pascal visual object classes VOC challenge. IJCV 88, 303–338 (2010)
  36. 36.Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Semantic understanding of scenes through the ADE20K dataset. arXiv:1608.05442 (2016)
  37. 37.Jia, Y., et al.: Caffe: convolutional architecture for fast feature embedding. In: ACM MM (2014)
  38. 38.Iandola, F.N., Moskewicz, M.W., Ashraf, K., Han, S., Dally, W.J., Keutzer, K.: SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <1mb model size. arXiv:1602.07360 (2016)
  39. 39.Han, S., Mao, H., Dally, W.J.: Deep compression: compressing deep neural network with pruning, trained quantization and Huffman coding. In: ICLR (2016)
  40. 40.Han, S., et al.: DSD: regularizing deep neural networks with dense-sparse-dense training flow. In: ICLR (2017)
  41. 41.Li, H., Kadav, A., Durdanovic, I., Samet, H., Graf, H.P.: Pruning filters for efficient convnets. In: ICLR (2017)
  42. 42.Sturgess, P., Alahari, K., Ladicky, L., Torr, P.H.: Combining appearance and structure from motion features for road scene understanding. In: BMVC (2009)
  43. 43.Lin, T.-Y.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48

Citation

MLA
Zhao, H., et al. “ICNet for Real-Time Semantic Segmentation on High-Resolution Images”. arXiv, 2017, http://arxiv.org/abs/1704.08545v2.
APA
Zhao, H., Qi, X., Shen, X., Shi, J., & Jia, J. (2017). ICNet for Real-Time Semantic Segmentation on High-Resolution Images. arXiv. http://arxiv.org/abs/1704.08545v2
Chicago
Zhao, H., X. Qi, X. Shen, J. Shi, and J. Jia. 2017. “ICNet for Real-Time Semantic Segmentation on High-Resolution Images”. arXiv. http://arxiv.org/abs/1704.08545v2.
Harvard
Zhao, H. et al. (2017) “ICNet for Real-Time Semantic Segmentation on High-Resolution Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1704.08545v2.
Vancouver
1. Zhao H, Qi X, Shen X, Shi J, Jia J (2017) ICNet for Real-Time Semantic Segmentation on High-Resolution Images. arXiv

BibTeX

@article{zhao2017icnet,
  title = {ICNet for Real-Time Semantic Segmentation on High-Resolution Images},
  author = {Zhao, Hengshuang and Qi, Xiaojuan and Shen, Xiaoyong and Shi, Jianping and Jia, Jiaya},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1704.08545v2},
  eprint = {1704.08545}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF