Context Encoding for Semantic Segmentation

Hang ZhangKristin DanaJianping ShiZhongyue ZhangXiaogang WangAmbrish TyagiAmit Agrawal

article2018CVPR1,375 citations

Introduces a Context Encoding Module that captures global scene context to selectively emphasize relevant class feature maps, setting state-of-the-art semantic segmentation performance on standard benchmarks while adding minimal computational overhead.

Listen

Semantic segmentation—the computer vision task of labeling every pixel in an image with its corresponding object category—is essential for autonomous systems, scene understanding, and automated image analysis. While modern deep neural networks achieve dense spatial resolution, standard methods evaluate pixels in isolation. This isolation often ignores overall scene context, leading to obvious classification errors, such as misidentifying an indoor windowpane as an exterior door.

The article evaluates whether integrating global scene context into deep neural networks improves segmentation accuracy without adding substantial computational overhead. It demonstrates a new framework, Context Encoding Network (EncNet), which captures scene-level semantic context to emphasize relevant object categories and suppress irrelevant ones.

To achieve this, the authors designed a lightweight Context Encoding Module and introduced a complementary Semantic Encoding Loss. The module captures global feature statistics to predict scaling factors that highlight class-relevant feature maps. In parallel, the new loss function regularizes model training by requiring the network to predict the presence or absence of object categories across the entire scene, giving equal weight to large and small objects. The architecture was tested across major visual benchmarks, including PASCAL-Context, PASCAL VOC 2012, ADE20K, and the CIFAR-10 image classification dataset, alongside an efficient synchronized batch normalization implementation across multiple graphics processors.

The empirical findings demonstrate that explicit contextual modeling yields substantial performance gains. On PASCAL-Context, adding the module increased mean Intersection over Union (mIoU) from 41.0% to 47.6% over a standard fully convolutional baseline, with the full model achieving 51.7% mIoU. On PASCAL VOC 2012, the model achieved 85.9% mIoU with MS-COCO pre-training, outperforming competitive contemporary models. On the complex ADE20K dataset with 150 categories, a single EncNet model reached a test score of 0.5567, surpassing prior competition-winning entries. Furthermore, adding the module to a compact 14-layer network on CIFAR-10 achieved a low 3.45% error rate, matching the accuracy of networks requiring up to ten times more layers.

These results show that explicit global context modeling significantly improves segmentation accuracy—especially for small or easily confused objects—while adding only 3% to 5% extra computational cost. This provides engineering and product teams with a pathway to deploy more accurate computer vision models on existing hardware budgets without having to scale up model depth or computational footprints.

Organizations developing computer vision systems should adopt context-encoding mechanisms and multi-task scene presence loss functions to upgrade standard segmentation pipelines. Teams can directly integrate these lightweight modules into existing architectures and utilize the publicly released implementation, including the synchronized cross-processor batch normalization, to enhance training stability on high-resolution imagery.

The evidence supporting these findings is strong across standardized benchmarks and ablation tests. However, the evaluation remains focused on curated benchmark datasets. Real-world applications characterized by heavy visual occlusion, domain shift, or specialized embedded hardware constraints may require pilot testing to confirm that the observed accuracy and efficiency advantages translate directly into operational settings.

arXiv: 1803.08904
Cover for Context Encoding for Semantic Segmentation

Abstract

Recent work has made significant progress in improving spatial resolution for pixelwise labeling with Fully Convolutional Network (FCN) framework by employing Dilated/Atrous convolution, utilizing multi-scale features and refining boundaries. In this paper, we explore the impact of global contextual information in semantic segmentation by introducing the Context Encoding Module, which captures the semantic context of scenes and selectively highlights class-dependent featuremaps. The proposed Context Encoding Module significantly improves semantic segmentation results with only marginal extra computation cost over FCN. Our approach has achieved new state-of-the-art results 51.7% mIoU on PASCAL-Context, 85.9% mIoU on PASCAL VOC 2012. Our single model achieves a final score of 0.5567 on ADE20K test set, which surpass the winning entry of COCO-Place Challenge in 2017. In addition, we also explore how the Context Encoding Module can improve the feature representation of relatively shallow networks for the image classification on CIFAR-10 dataset. Our 14 layer network has achieved an error rate of 3.45%, which is comparable with state-of-the-art approaches with over 10 times more layers. The source code for the complete system are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Context Encoding Module
  • 2.1 Context Encoding Network (EncNet)
  • 2.2 Relation to Other Approaches
  • 3 Experimental Results
  • 3.1 Implementation Details
  • 3.2 Results on PASCAL-Context
  • 3.3 Results on PASCAL VOC 2012
  • 3.4 Results on ADE20K
  • 3.5 Image Classification Results on CIFAR-10
  • 4 Conclusion
  • A Implementation Details on Synchronized Cross-GPU Batch Normalization
  • References

Knowls

  1. Knowl 1 — Context Encoding Module and Featuremap Attention

    model/method

    The Context Encoding Module captures global scene semantics and selectively modulates class-dependent feature responses in semantic segmentation networks.

    Given an intermediate convolutional featuremap X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, viewed as a spatial collection of CC-dimensional feature vectors X={x1,x2,…,xN}X = \{x_1, x_2, \dots, x_N\} where N=H×WN = H \times W, the module processes XX through an Encoding Layer to extract an aggregated semantic context vector e∈RCe \in \mathbb{R}^C.

    To leverage this global contextual representation as a feedback mechanism, a channel-wise scaling vector γ∈RC\gamma \in \mathbb{R}^C is predicted via a fully connected layer followed by a sigmoid activation function:

    γ=δ(We)\gamma = \delta(W e)

    where W∈RC×CW \in \mathbb{R}^{C \times C} represents the learnable weight matrix and δ(⋅)\delta(\cdot) denotes the element-wise sigmoid activation δ(z)=11+e−z\delta(z) = \frac{1}{1 + e^{-z}}.

    The modulated output featuremap Y∈RC×H×WY \in \mathbb{R}^{C \times H \times W} is then computed via channel-wise multiplication between the input featuremap XX and the scaling factor γ\gamma:

    Y=X⊗γY = X \otimes \gamma

    where ⊗\otimes denotes channel-wise multiplication, scaling the entire spatial slice of channel cc by γc\gamma_c. This gating mechanism selectively emphasizes or de-emphasizes individual feature channels based on the global semantic scene context.

  2. Knowl 2 — Encoding Layer Formulation and Residual Aggregation

    equation

    The Encoding Layer maps an unordered set of CC-dimensional features X={x1,x2,…,xN}⊂RCX = \{x_1, x_2, \dots, x_N\} \subset \mathbb{R}^C (where N=H×WN = H \times W) into a compact global context vector e∈RCe \in \mathbb{R}^C using a learnable dictionary of KK visual codewords D={d1,d2,…,dK}⊂RCD = \{d_1, d_2, \dots, d_K\} \subset \mathbb{R}^C and corresponding positive smoothing factors S={s1,s2,…,sK}⊂R+S = \{s_1, s_2, \dots, s_K\} \subset \mathbb{R}^+.

    For each input feature xix_i and codeword dkd_k, the residual vector is defined as rik=xi−dkr_{ik} = x_i - d_k. The soft-assignment weight and residual for codeword dkd_k across all NN features is given by:

    ek=∑i=1Neik=∑i=1Nexp⁡(−sk∥rik∥2)∑j=1Kexp⁡(−sj∥rij∥2)rike_k = \sum_{i=1}^{N} e_{ik} = \sum_{i=1}^{N} \frac{\exp\left(-s_k \|r_{ik}\|^2\right)}{\sum_{j=1}^{K} \exp\left(-s_j \|r_{ij}\|^2\right)} r_{ik}

    To avoid imposing an artificial ordering on the KK independent codeword residuals and to reduce dimensionality, the codewords are aggregated via summation rather than concatenation:

    e=∑k=1Kϕ(ek)e = \sum_{k=1}^{K} \phi(e_k)

    where ϕ(⋅)\phi(\cdot) denotes Batch Normalization followed by a Rectified Linear Unit (ReLU) activation function, ϕ(z)=ReLU(BN(z))\phi(z) = \text{ReLU}(\text{BN}(z)). The resulting vector e∈RCe \in \mathbb{R}^C represents the orderless encoded global semantics of the scene.

  3. Knowl 3 — Semantic Encoding Loss

    model/method

    Semantic Encoding Loss (SE-loss) is an auxiliary global supervision objective designed to enforce the learning of scene context without requiring manual scene-level annotations.

    SE-loss regularizes the training by predicting the overall presence or absence of object categories in the image. An auxiliary branch consisting of a fully connected layer and a sigmoid activation is applied directly on top of the encoded semantic representation e∈RCe \in \mathbb{R}^C output by the Encoding Layer, predicting category existence probabilities y^∈[0,1]M\hat{y} \in [0, 1]^M, where MM is the number of semantic classes.

    The ground-truth multi-label presence vector y∈{0,1}My \in \{0, 1\}^M is automatically generated from the ground-truth per-pixel segmentation mask by applying a unique-element operation (ym=1y_m = 1 if class mm is present anywhere in the image mask, and 00 otherwise).

    The loss is optimized using multi-label binary cross-entropy:

    LSE=−1M∑m=1M[ymlog⁡y^m+(1−ym)log⁡(1−y^m)]\mathcal{L}_{\text{SE}} = -\frac{1}{M} \sum_{m=1}^{M} \left[ y_m \log \hat{y}_m + (1 - y_m) \log (1 - \hat{y}_m) \right]

    Unlike per-pixel cross-entropy segmentation loss, which is dominated by large background regions and large objects, SE-loss assigns equal weight to big and small object categories present in the image, significantly improving the segmentation of small objects.

  4. Knowl 4 — Context Encoding Network Architecture

    model/method

    Context Encoding Network (EncNet) is a semantic segmentation framework built upon a pre-trained Deep Residual Network (ResNet) backbone using dilated convolutions.

    The architectural pipeline operates as follows:

    1. Dilated convolution strategy is applied to stage 3 (stride 16) and stage 4 (stride 32) of the ResNet backbone, producing high-resolution feature representations with an overall output stride of 8 (featuremap spatial resolution of 1/81/8 the input image).
    2. A Context Encoding Module is placed on top of stage 4. It extracts encoded semantics ee using an Encoding Layer with K=32K=32 codewords and applies channel-wise featuremap attention scaling γ=δ(We)\gamma = \delta(We) to modulate the stage 4 featuremaps.
    3. The modulated featuremaps are passed into a final convolutional prediction layer to produce dense per-pixel logits, which are bilinearly upsampled by a factor of 8 to calculate the primary pixel-wise cross-entropy segmentation loss Lseg\mathcal{L}_{\text{seg}}.
    4. For deep regularization, SE-loss branches are attached to both stage 3 and stage 4 via dedicated Context Encoding Modules.
    5. The complete training objective is a weighted sum:

    L=Lseg+αLSE-stage4+αLSE-stage3\mathcal{L} = \mathcal{L}_{\text{seg}} + \alpha \mathcal{L}_{\text{SE-stage4}} + \alpha \mathcal{L}_{\text{SE-stage3}}

    where α\alpha is the loss weighting coefficient (empirically optimal at α=0.2\alpha = 0.2).

  5. Knowl 5 — Single-Synchronization Cross-GPU Batch Normalization

    algorithm

    Standard Cross-GPU Batch Normalization (SyncBN) calculates global mean across all GPUs and then performs a second communication round to calculate global variance. The single-synchronization SyncBN algorithm reduces inter-GPU communication to a single AllReduce-SUM operation per forward pass and per backward pass.

    Input: Per-device mini-batch tensors Xd={x1,…,xNd}X_d = \{x_1, \dots, x_{N_d}\} across all distributed devices d∈{1,…,G}d \in \{1, \dots, G\}, total samples N=∑d=1GNdN = \sum_{d=1}^G N_d, scale parameter γ\gamma, shift parameter β\beta, small constant ϵ>0\epsilon > 0
    Output: Normalized and scaled tensors Yd={y1,…,yNd}Y_d = \{y_1, \dots, y_{N_d}\} for each device dd
    // Forward Pass (Single Synchronization Step)
    for each device dd in parallel:
        Compute local sum: S1,d=∑i=1NdxiS_{1, d} = \sum_{i=1}^{N_d} x_i
        Compute local squared sum: S2,d=∑i=1Ndxi2S_{2, d} = \sum_{i=1}^{N_d} x_i^2
    Synchronize across all GPUs via a single AllReduce-SUM:
        S1=∑d=1GS1,d=∑i=1NxiS_1 = \sum_{d=1}^G S_{1, d} = \sum_{i=1}^N x_i
        S2=∑d=1GS2,d=∑i=1Nxi2S_2 = \sum_{d=1}^G S_{2, d} = \sum_{i=1}^N x_i^2
    for each device dd in parallel:
        Compute global mean: μ=S1N\mu = \frac{S_1}{N}
        Compute global variance: σ2=S2N−μ2\sigma^2 = \frac{S_2}{N} - \mu^2
        for each sample xix_i in XdX_d:
            yi=γxi−μσ2+ϵ+βy_i = \gamma \frac{x_i - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta
        return YdY_d

    During backpropagation, an identical single-synchronization strategy is used to aggregate gradients with respect to S1S_1 and S2S_2 across devices simultaneously.

  6. Knowl 6 — Ablation Study of EncNet on PASCAL-Context

    data/table

    Ablation experiments on the PASCAL-Context dataset evaluate the incremental contribution of the Context Encoding Module, the Semantic Encoding Loss (SE-loss), backbone depth (ResNet-50 vs. ResNet-101), and multi-scale (MS) testing. Evaluations are conducted across 59 object classes excluding background (with 60-class results reported in parentheses for comparison).

    Method BaseNet Encoding SE-loss MS pixAcc (%) mIoU (%)
    FCN Res50 73.4 41.0
    EncNet Res50 ✓ 78.1 47.6
    EncNet Res50 ✓ ✓ 79.4 49.2
    EncNet Res101 ✓ ✓ 80.4 51.7
    EncNet Res101 ✓ ✓ ✓ 81.2 52.6

    Key empirical findings include:

    • Adding the Context Encoding Module to a ResNet-50 FCN baseline increases mIoU from 41.0% to 47.6% (+6.6%) while introducing only 3%–5% extra computational cost.
    • Incorporating SE-loss provides an additional +1.6% mIoU improvement (reaching 49.2%).
    • Scaling the backbone to ResNet-101 improves mIoU to 51.7%.
    • Multi-scale evaluation achieves 52.6% mIoU on 59 classes (and 51.7% mIoU evaluated across all 60 classes with background, outperforming RefineNet Res152 at 47.3% and DeepLab-v2 Res101 at 45.7%).
    • Sensitivity analyses show that performance peaks at an SE-loss weight of α=0.2\alpha = 0.2 and saturates at K=32K = 32 dictionary codewords.
  7. Knowl 7 — Semantic Segmentation Benchmarks on PASCAL VOC 2012

    empirical result

    On the PASCAL VOC 2012 test benchmark (20 object classes plus background), EncNet achieves state-of-the-art performance under two distinct training regimes:

    1. Without MS-COCO pre-training: Trained on the augmented PASCAL VOC training set (10,582 images) and fine-tuned on the original PASCAL training set, EncNet achieves 82.9% mIoU, outperforming PSPNet (82.6%), ResNet38 (82.5%), and Piecewise (75.3%).

    2. With MS-COCO pre-training: Pre-trained on a 6.5K subset of MS-COCO (images containing at least 1,000 pixels belonging to the 20 PASCAL classes) and fine-tuned on PASCAL VOC, EncNet (ResNet-101 backbone) achieves 85.9% mIoU, outperforming DeepLabv3 (85.7%), PSPNet (85.4%), ResNet38 (84.9%), and RefineNet (84.2%) while maintaining lower computational complexity than PSPNet and DeepLabv3.

  8. Knowl 8 — Scene Parsing Benchmarks on ADE20K

    data/table

    EncNet was evaluated on the ADE20K scene parsing benchmark (150 stuff and object categories) on both the validation set (2,000 images) and test set (3,000 images).

    Method BaseNet Pixel Accuracy (%) mIoU (%)
    FCN Res50 74.57 34.38
    EncNet (ours) Res50 79.73 41.11
    EncNet (ours) Res101 81.69 44.65
    PSPNet Res101 81.39 43.29
    PSPNet Res269 81.69 44.94

    On the validation set, EncNet-50 outperforms the ResNet-50 FCN baseline by +6.73% mIoU and +5.16% pixel accuracy. EncNet-101 achieves 81.69% pixel accuracy and 44.65% mIoU, matching the performance of a much deeper PSPNet-269 backbone.

    On the ADE20K test set, a single EncNet-101 model achieves a final competition score of 0.5567, surpassing the single-model PSPNet-269 score (0.5538; 1st place in Places Challenge 2016) as well as the winning entry of the COCO-Place Challenge 2017 (CASIA_IVA_JD at 0.5547).

  9. Knowl 9 — Context Encoding for CIFAR-10 Classification and Stochastic Smoothing Regularization

    empirical result

    The Context Encoding Module can be adapted to shallow image classification networks by placing it on top of each basic residual block to predict scaling factors for the residual branches, thereby preserving identity mappings throughout the network. The module first reduces feature channels by a factor of 4 via a 1×11\times 1 convolution, applies the Encoding Layer with concatenated codeword vectors, and normalizes with an L2L_2 norm.

    During training, a stochastic regularization scheme is applied to the smoothing factors sks_k of the Encoding Layer: instead of remaining fixed, sks_k is randomly sampled from a uniform distribution sk∼U(0,1)s_k \sim U(0, 1) during forward and backward passes, and set to sk=0.5s_k = 0.5 during evaluation.

    On the CIFAR-10 test set, this allows shallow 14-layer networks to match the accuracy of models with over 10×10\times more layers:

    Method Depth Parameters Error (%)
    ResNet (pre-act) 1001 10.2M 4.62
    Wide ResNet 28×1028\times 10 28 36.5M 3.89
    ResNeXt-29 16×6416\times 64d 29 68.1M 3.58
    DenseNet-BC (k=40k=40) 190 25.6M 3.46
    ResNet 64d baseline 14 2.7M 4.93
    SE-ResNet 64d baseline 14 2.8M 4.65
    EncNet 16k64d (ours) 14 3.5M 3.96
    EncNet 32k128d (ours) 14 16.8M 3.45

    A 14-layer EncNet (32k128d) achieves a 3.45% error rate with 16.8M parameters, outperforming deeper architectures such as DenseNet-BC (190 layers, 3.46%) and ResNeXt-29 (29 layers, 3.58%).

Coverage note — None was omitted; all primary contributions, modules (Context Encoding, SE-loss, EncNet, SyncBN), equations, experimental settings, and results on PASCAL-Context, PASCAL VOC 2012, ADE20K, and CIFAR-10 are represented.

References

  1. 1.R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic. Netvlad: Cnn architecture for weakly supervised place recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5297–5307, 2016. 2
  2. 2.A. Arnab, S. Jayasumana, S. Zheng, and P. H. Torr. Higher order conditional random fields in deep neural networks. In European Conference on Computer Vision, pages 524–540. Springer, 2016. 6
  3. 3.V. Badrinarayanan, A. Kendall, and R. Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. arXiv preprint arXiv:1511.00561, 2015. 4, 7
  4. 4.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015. 1, 2, 4, 5, 7
  5. 5.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv:1606.00915, 2016. 2, 4, 5, 6, 7
  6. 6.L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017. 4, 5, 6, 7
  7. 7.M. Cimpoi, S. Maji, and A. Vedaldi. Deep filter banks for texture recognition and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3828–3836, 2015. 4
  8. 8.G. Csurka, C. Dance, L. Fan, J. Willamowski, and C. Bray. Visual categorization with bags of keypoints. In Workshop on statistical learning in computer vision, ECCV, volume 1, pages 1–2. Prague, 2004. 2
  9. 9.J. Dai, K. He, and J. Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1635–1643, 2015. 6
  10. 10.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009. 1, 2
  11. 11.V. Dumoulin, J. Shlens, M. Kudlur, A. Behboodi, F. Lemic, A. Wolisz, M. Molinaro, C. Hirche, M. Hayashi, E. Bagan, et al. A learned representation for artistic style. arXiv preprint arXiv:1610.07629, 2016. 4
  12. 12.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 4, 5, 6
  13. 13.L. Fei-Fei and P. Perona. A bayesian hierarchical model for learning natural scene categories. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), volume 2, pages 524–531. IEEE, 2005. 2
  14. 14.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014. 4
  15. 15.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Simultaneous detection and segmentation. In European Conference on Computer Vision, pages 297–312. Springer, 2014. 4
  16. 16.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Hypercolumns for object segmentation and fine-grained localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 447–456, 2015. 6
  17. 17.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015. 2, 4, 8
  18. 18.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015. 8
  19. 19.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. arXiv preprint arXiv:1603.05027, 2016. 8
  20. 20.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. arXiv preprint arXiv:1709.01507, 2017. 3, 4, 8
  21. 21.G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993, 2016. 8
  22. 22.X. Huang and S. Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017. 3, 4
  23. 23.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning, pages 448–456, 2015. 2, 4, 5, 8, 9
  24. 24.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In Advances in Neural Information Processing Systems, pages 2017–2025, 2015. 4
  25. 25.H. Jegou, M. Douze, C. Schmid, and P. P ´ erez. Aggregating local descriptors into a compact image representation. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 3304–3311. IEEE, 2010. 2
  26. 26.T. Joachims. Text categorization with support vector machines: Learning with many relevant features. In European conference on machine learning, pages 137–142. Springer, 1998. 2
  27. 27.G. Kang, J. Li, and D. Tao. Shakeout: A new regularized deep neural network training scheme. In AAAI, pages 1751–1757, 2016. 8
  28. 28.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. University of Toronto, Technical Report, 2009. 2, 8
  29. 29.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 1
  30. 30.T. Leung and J. Malik. Representing and recognizing the visual appearance of materials using three-dimensional textons. International journal of computer vision, 43(1):29–44, 2001. 2
  31. 31.G. Lin, A. Milan, C. Shen, and I. Reid. RefineNet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR, July 2017. 6, 7
  32. 32.G. Lin, C. Shen, A. van den Hengel, and I. Reid. Efficient piecewise training of deep structured models for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3194–3203, 2016. 6, 7
  33. 33.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. 7
  34. 34.S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia. Path aggregation network for instance segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 9
  35. 35.W. Liu, A. Rabinovich, and A. C. Berg. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015. 4, 5, 6
  36. 36.Z. Liu, X. Li, P. Luo, C.-C. Loy, and X. Tang. Semantic image segmentation via deep parsing network. In Proceedings of the IEEE International Conference on Computer Vision, pages 1377–1385, 2015. 7
  37. 37.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015. 1, 4, 6, 7
  38. 38.D. G. Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2):91–110, 2004. 2
  39. 39.P. Luo, G. Wang, L. Lin, and X. Wang. Deep dual learning for semantic image segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2718–2726, 2017. 7
  40. 40.R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille. The role of context for object detection and semantic segmentation in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 891–898, 2014. 4, 5, 6
  41. 41.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1520–1528, 2015. 4, 7
  42. 42.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017. 5, 9
  43. 43.C. Peng, T. Xiao, Z. Li, Y. Jiang, X. Zhang, K. Jia, G. Yu, and J. Sun. Megdet: A large mini-batch object detector. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 9
  44. 44.F. Perronnin, J. Sanchez, and T. Mensink. Improving the fisher kernel for large-scale image classification. In European conference on computer vision, pages 143–156. Springer, 2010. 2
  45. 45.G. Schwartz and K. Nishino. Material recognition from local appearance in global context. arXiv preprint arXiv:1611.09394, 2016. 5
  46. 46.J. Sivic, B. C. Russell, A. A. Efros, A. Zisserman, and W. T. Freeman. Discovering objects and their location in images. In Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, volume 1, pages 370–377. IEEE, 2005. 2
  47. 47.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of machine learning research, 15(1):1929–1958, 2014. 8
  48. 48.M. Varma and A. Zisserman. Classifying images of materials: Achieving viewpoint and illumination independence. In European Conference on Computer Vision, pages 255–271. Springer, 2002. 2
  49. 49.R. Vemulapalli, O. Tuzel, M.-Y. Liu, and R. Chellapa. Gaussian conditional random field network for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3224–3233, 2016. 7
  50. 50.G. Wang, P. Luo, L. Lin, and X. Wang. Learning object interactions and descriptions for semantic image segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5859–5867, 2017. 7
  51. 51.Z. Wu, C. Shen, and A. v. d. Hengel. Bridging category-level and instance-level semantic image segmentation. arXiv preprint arXiv:1605.06885, 2016. 6
  52. 52.Z. Wu, C. Shen, and A. v. d. Hengel. Wider or deeper: Revisiting the resnet model for visual recognition. arXiv preprint arXiv:1611.10080, 2016. 7
  53. 53.S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. arXiv preprint arXiv:1611.05431, 2016. 8
  54. 54.F. Yu and V. Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015. 1, 2, 4, 5, 7
  55. 55.F. Yu, V. Koltun, and T. Funkhouser. Dilated residual networks. arXiv preprint arXiv:1705.09914, 2017. 4
  56. 56.S. Zagoruyko and N. Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016. 8
  57. 57.H. Zhang and K. Dana. Multi-style generative network for real-time transfer. arXiv preprint arXiv:1703.06953, 2017. 3, 4
  58. 58.H. Zhang, J. Xue, and K. Dana. Deep ten: Texture encoding network. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. 2, 3
  59. 59.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia. Pyramid scene parsing network. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2, 4, 5, 7
  60. 60.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr. Conditional random fields as recurrent neural networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 1529–1537, 2015. 4, 6, 7
  61. 61.B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba. Scene parsing through ade20k dataset. In Proc. CVPR, 2017. 1, 2, 4, 5, 6, 7

Citation

MLA
Zhang, H., et al. “Context Encoding for Semantic Segmentation”. arXiv, 2018, http://arxiv.org/abs/1803.08904v1.
APA
Zhang, H., Dana, K., Shi, J., Zhang, Z., Wang, X., Tyagi, A., & Agrawal, A. (2018). Context Encoding for Semantic Segmentation. arXiv. http://arxiv.org/abs/1803.08904v1
Chicago
Zhang, H., K. Dana, J. Shi, et al. 2018. “Context Encoding for Semantic Segmentation”. arXiv. http://arxiv.org/abs/1803.08904v1.
Harvard
Zhang, H. et al. (2018) “Context Encoding for Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.08904v1.
Vancouver
1. Zhang H, Dana K, Shi J, Zhang Z, Wang X, Tyagi A, Agrawal A (2018) Context Encoding for Semantic Segmentation. arXiv

BibTeX

@article{zhang2018context,
  title = {Context Encoding for Semantic Segmentation},
  author = {Zhang, Hang and Dana, Kristin and Shi, Jianping and Zhang, Zhongyue and Wang, Xiaogang and Tyagi, Ambrish and Agrawal, Amit},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.08904v1},
  eprint = {1803.08904}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE