Large Kernel Matters — Improve Semantic Segmentation by Global Convolutional Network

Chao PengXiangyu ZhangGang YuGuiming LuoJian Sun

article2017CVPR1,618 citations

Proposes the Global Convolutional Network architecture with large separable kernels and residual boundary refinement to simultaneously address classification and localization challenges in semantic segmentation, achieving state-of-the-art accuracy on PASCAL VOC 2012 and Cityscapes.

Listen

Semantic segmentation—the computer vision task of assigning an accurate category label to every pixel in an image—is essential for autonomous driving, robotics, and image analysis. However, it faces a fundamental trade-off between category recognition and precise spatial localization. Category recognition requires broad context and invariance to transformations like rotation and shifting, whereas spatial localization demands precise sensitivity to pixel locations. Prior systems predominantly favored localization by using narrow, stacked computational filters, which restricted context and impaired recognition for large objects.

The article demonstrates an architecture called the Global Convolutional Network (GCN) coupled with a Boundary Refinement (BR) block to resolve this conflict. The core objective is to evaluate whether expanding the effective context area using large, computationally efficient convolutional filters improves pixel-level semantic labeling without sacrificing spatial accuracy.

The researchers designed an end-to-end framework based on high-capacity residual networks (ResNet-152) and benchmarked it on two standard public datasets: PASCAL VOC 2012, which contains diverse object classes, and Cityscapes, which comprises complex urban street scenes. To avoid the computational penalty of traditional large filters, the design decomposed broad two-dimensional filters into combinations of one-dimensional horizontal and vertical operations. The authors conducted extensive ablation experiments to isolate the effects of filter size, model parameters, and boundary alignment modules.

The experiments produced four critical findings. First, larger filter sizes consistently improved segmentation accuracy; expanding the filter size to span the entire feature map boosted accuracy on the PASCAL VOC validation set by 5.5 percentage points over small-filter baselines. Second, the decomposed large-filter structure outperformed both standard large filters and stacks of small filters while using significantly fewer parameters and avoiding training convergence issues. Third, error analysis showed that the large-filter network primarily improved the internal classification of objects (increasing internal accuracy to 95.0%), while the boundary refinement module improved alignment along object edges (raising boundary accuracy from 71.5% to 73.4%). Fourth, the full system established new state-of-the-art benchmarks, achieving 82.2% mean intersection-over-union on PASCAL VOC 2012 and 76.9% on Cityscapes, outperforming previous methods by 2.0 and 5.1 percentage points, respectively.

These findings demonstrate that semantic segmentation models do not need to choose between wide context recognition and spatial localization. The decomposed filter design delivers superior visual understanding with lower computational cost and fewer parameters than standard approaches. This provides direct benefits for real-world deployments by reducing processing overhead and memory usage while enhancing scene understanding in safety-critical applications such as autonomous navigation.

For engineering and development teams building visual perception systems, the article supports replacing traditional stacked small-kernel blocks with decomposed large-kernel modules and integrating residual boundary refinement. When targeting high accuracy on urban or multi-object scenes, teams should adopt multi-stage pre-training on broader datasets before final fine-tuning.

The evaluation is highly credible due to strong benchmark results, but readers should note certain limitations. The primary evaluations relied on deep, high-capacity base networks, and testing very large input images required cropping, multi-scale evaluation, and post-processing steps to reach peak performance. Further work is needed to validate real-time inference latency and efficiency on resource-constrained embedded hardware.

arXiv: 1703.02719
Cover for Large Kernel Matters — Improve Semantic Segmentation by Global Convolutional Network

Abstract

One of recent trends [30, 31, 14] in network architec- ture design is stacking small filters (e.g., 1x1 or 3x3) in the entire network because the stacked small filters is more ef- ficient than a large kernel, given the same computational complexity. However, in the field of semantic segmenta- tion, where we need to perform dense per-pixel prediction, we find that the large kernel (and effective receptive field) plays an important role when we have to perform the clas- sification and localization tasks simultaneously. Following our design principle, we propose a Global Convolutional Network to address both the classification and localization issues for the semantic segmentation. We also suggest a residual-based boundary refinement to further refine the ob- ject boundaries. Our approach achieves state-of-art perfor- mance on two public benchmarks and significantly outper- forms previous results, 82.2% (vs 80.2%) on PASCAL VOC 2012 dataset and 76.9% (vs 71.8%) on Cityscapes dataset.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Approach
  • 3.1 Global Convolutional Network
  • 3.2 Overall Framework
  • 4 Experiment
  • 4.1 Ablation Experiments
  • 4.1.1 Global Convolutional Network — Large Kernel Matters
  • 4.1.2 Global Convolutional Network for Pretrained Model
  • 4.2 PASCAL VOC 2012
  • 4.3 Cityscapes
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Global Convolutional Network Module Architecture

    model/method

    The Global Convolutional Network (GCN) module is designed to enable dense spatial connectivity over a large receptive field while retaining spatial localization and parameter efficiency. Given an input feature map of dimensions w×h×cw \times h \times c (where ww is width, hh is height, and cc is channel depth) and a target of CC output channels (e.g., semantic categories), the GCN module applies two parallel branches of symmetric, separable 1D convolutions:

    1. Branch 1: A 1×k1 \times k convolution transforming cc channels to CC channels, followed immediately by a k×1k \times 1 convolution from CC to CC channels.
    2. Branch 2: A k×1k \times 1 convolution transforming cc channels to CC channels, followed immediately by a 1×k1 \times k convolution from CC to CC channels.

    The outputs of both branches are merged via element-wise addition to produce an output tensor of shape w×h×Cw \times h \times C.

    Unlike separable convolutions used in Inception architectures, no non-linear activation functions (such as ReLU) are applied between the 1D convolutions within each branch. Compared to a standard full 2D k×kk \times k convolution, this separable formulation requires O(2k)O\left(\frac{2}{k}\right) of the computational complexity and parameter count, enabling the use of large kernel sizes (such as k=15k = 15 or k=25k = 25) that match or approach the spatial dimensions of the feature map.

  2. Knowl 2 — Residual Boundary Refinement Block

    model/method

    The Boundary Refinement (BR) block is a residual module designed to improve semantic segmentation localization along object boundaries. Given an input score map S∈Rw×h×CS \in \mathbb{R}^{w \times h \times C} (where ww and hh are spatial dimensions and CC is the number of semantic categories), the refined score map S~∈Rw×h×C\tilde{S} \in \mathbb{R}^{w \times h \times C} is defined by:

    S~=S+R(S)\tilde{S} = S + R(S)

    where R(⋅)R(\cdot) represents the residual mapping consisting of a sequence of two 3×33 \times 3 convolutional layers with CC channels and an intermediate Rectified Linear Unit (ReLU) activation:

    R(S)=Conv3×3(ReLU(Conv3×3(S)))R(S) = \text{Conv}_{3\times 3}\left(\text{ReLU}\left(\text{Conv}_{3\times 3}(S)\right)\right)

    Because it is fully convolutional and operates directly on semantic score maps, the Boundary Refinement block is trained end-to-end with the rest of the network via standard backpropagation, serving as an integrated alternative to post-processing methods such as Conditional Random Fields (CRFs).

  3. Knowl 3 — Multi-Scale Semantic Segmentation Framework with GCN and BR

    model/method

    The overall semantic segmentation framework combines a pretrained deep convolutional backbone (such as ResNet-152) with an FCN-4 multi-scale pyramid architecture:

    1. Multi-Scale Feature Extraction: Intermediate feature maps are extracted from four residual stages of the backbone: res-2, res-3, res-4, and res-5. For an input image of size 512×512512 \times 512, these feature maps have spatial resolutions of 128×128128 \times 128, 64×6464 \times 64, 32×3232 \times 32, and 16×1616 \times 16, with channel depths of 256, 512, 1024, and 2048, respectively.
    2. Score Map Generation: Each stage's feature map is fed into a separate Global Convolutional Network (GCN) module to generate a CC-channel class score map at that stage's spatial resolution.
    3. Hierarchical Boundary Refinement and Fusion: Each GCN score map is processed by a Boundary Refinement (BR) block. Deeper, lower-resolution score maps are upsampled by a factor of 2 via deconvolution (transposed convolution), passed through an additional BR block, and element-wise added to the BR-processed score map of the preceding stage.
    4. Final Upsampling: The fused score map is progressively upsampled and combined across all stages down to res-2 (128×128128 \times 128), and finally upsampled by deconvolution with a BR block to yield the full-resolution prediction map (512×512×C512 \times 512 \times C).
  4. Knowl 4 — Disentangled Effects of GCN and BR on Object Interior versus Boundary Accuracy

    empirical result

    To evaluate how Global Convolutional Networks (GCN) and Boundary Refinement (BR) address classification and localization separately, pixel predictions on the PASCAL VOC 2012 validation set were evaluated on two disjoint subsets of object pixels:

    • Boundary region: Pixels within a distance ≤7\le 7 pixels from ground-truth object boundaries.
    • Internal region: Pixels located at a distance >7> 7 pixels from boundaries.
    Model Boundary Accuracy (%) Internal Accuracy (%) Overall Mean IoU (%)
    Baseline (1×11 \times 1 conv) 71.3 93.9 69.0
    GCN (k=15k=15) 71.5 95.0 74.5
    GCN (k=15k=15) + BR 73.4 95.1 74.7

    The results demonstrate complementary mechanisms:

    1. GCN primarily increases the classification accuracy in the internal object regions (from 93.9% to 95.0%, a 1.1% gain) with negligible change in boundary accuracy (71.3% to 71.5%), confirming that large kernels act as transformation-invariant dense classifiers across object interiors.
    2. The Boundary Refinement block primarily improves localization accuracy in boundary regions (from 71.5% to 73.4%, a 1.9% gain), while leaving interior accuracy essentially unchanged (95.0% to 95.1%).
  5. Knowl 5 — Effect of Kernel Size $k$ in Global Convolutional Networks

    empirical result

    On the PASCAL VOC 2012 validation set (without Boundary Refinement, using a ResNet-152 backbone with 512×512512 \times 512 inputs where the top-most feature map is 16×1616 \times 16), the kernel size kk of the GCN module was evaluated across odd values from 3 to 15:

    Kernel size kk Baseline (1×11\times 1) 3 5 7 9 11 13 15
    Mean IoU (%) 69.0 70.1 71.1 72.8 73.4 73.7 74.0 74.5

    Segmentation accuracy increases monotonically with kernel size kk. Setting k=15k = 15 to match the spatial size of the top feature map (global convolution) achieves 74.5% mean IoU, surpassing the 1×11 \times 1 baseline by 5.5% mean IoU.

  6. Knowl 6 — Performance and Parameter Comparison of GCN versus Full Convolutions and Stacked Convolutions

    empirical result

    GCN was compared against standard full 2D k×kk \times k convolutions and stacked 3×33 \times 3 convolutions without intermediate non-linearities on the PASCAL VOC 2012 validation set.

    1. GCN vs. Full 2D k×kk \times k Convolution:

    Kernel size kk 3 5 7 9
    Mean IoU (GCN) (%) 70.1 71.1 72.8 73.4
    Mean IoU (Trivial Conv) (%) 69.8 70.4 69.6 68.8
    # Params after res-5 (GCN) 260K 434K 608K 782K
    # Params after res-5 (Trivial Conv) 387K 1075K 2107K 3484K

    Trivial full convolutions suffer performance degradation and training convergence issues when k≥7k \ge 7 due to parameter explosion, whereas GCN improves monotonically while maintaining a small parameter footprint.

    2. GCN vs. Stacked 3×33 \times 3 Convolutions:

    Equivalent kernel size kk 3 5 7 9 11
    Mean IoU (GCN) (%) 70.1 71.1 72.8 73.4 73.7
    Mean IoU (3×33\times 3 Stacks) (%) 69.8 71.8 71.3 69.5 67.5

    Stacked 3×33 \times 3 convolutions degrade rapidly past k=5k=5 (dropping to 67.5% at k=11k=11). Reducing the intermediate channel dimension mm of the stack at k=7k=7 to match parameter limits further degrades accuracy (m=2048→71.3%m=2048 \to 71.3\% with 75.9M params; m=1024→70.4%m=1024 \to 70.4\% with 28.5M params; m=210→68.8%m=210 \to 68.8\% with 4.3M params), trailing GCN (72.8% with 608K params).

  7. Knowl 7 — ResNet-GCN Architecture for Backbone Pretraining

    model/method

    The Global Convolutional Network structure can be incorporated directly into the backbone network building blocks. In ResNet50-GCN, the standard bottleneck structure (1×1→3×3→1×11 \times 1 \to 3 \times 3 \to 1 \times 1) in stages res-4 and res-5 is replaced by:

    1. Parallel separable GCN branches: (1×k→k×1)(1 \times k \to k \times 1) and (k×1→1×k)(k \times 1 \to 1 \times k), with Batch Normalization and ReLU applied after each 1D convolution layer.
    2. A final 1×11 \times 1 convolution projecting back to the bottleneck output channel depth.

    To match the FLOPs (3700 MFlops) and parameter count of standard ResNet-50:

    • res-4 (6 blocks): k=5k=5 with 85 intermediate channels for the separable layers and 1024 output channels.
    • res-5 (3 blocks): k=7k=7 with 128 intermediate channels for the separable layers and 2048 output channels.

    On ImageNet classification (top-5 error on 224×224224 \times 224 center crop), ResNet50-GCN achieves 7.9% error compared to 7.7% for standard ResNet-50. However, when fine-tuned for semantic segmentation on PASCAL VOC 2012 with a simple baseline head, ResNet50-GCN achieves 71.2% mean IoU versus 65.7% for standard ResNet-50 (a 5.5% improvement).

  8. Knowl 8 — PASCAL VOC 2012 Benchmark Results and Training Pipeline

    empirical result

    On the PASCAL VOC 2012 semantic segmentation benchmark (20 object categories plus background), the GCN framework with ResNet-152 was trained using a three-stage fine-tuning schedule with Stochastic Gradient Descent (SGD, batch size 1, momentum 0.99, weight decay 0.0005):

    1. Stage 1: Pre-training on 109,892 images combining MS COCO (retaining only the 20 VOC categories), the Semantic Boundaries Dataset (SBD), and VOC 2012 (input padded to 640×640640 \times 640).
    2. Stage 2: Fine-tuning on 10,582 images from SBD and VOC 2012 (input padded to 512×512512 \times 512).
    3. Stage 3: Fine-tuning on the 1,464 standard VOC 2012 training images (input padded to 512×512512 \times 512).
    Training Stage Baseline (1×11\times 1) Mean IoU (%) GCN Mean IoU (%) GCN + BR Mean IoU (%)
    Stage-1 69.6 74.1 75.0
    Stage-2 72.4 77.6 78.6
    Stage-3 74.0 78.7 80.3
    Stage-3 + Multi-Scale – – 80.4
    Stage-3 + Multi-Scale + CRF – – 81.0

    On the PASCAL VOC 2012 test server, the model achieved 82.2% mean IoU, exceeding previous published methods including CentraleSupelec Deep G-CRF (80.2%), Deeplabv2-CRF (79.7%), and LRR-4x ResNet (79.3%).

  9. Knowl 9 — Cityscapes Benchmark Results and Experimental Protocol

    empirical result

    The GCN framework was evaluated on the Cityscapes urban scene understanding dataset (19 evaluated classes, 5,000 fine-annotated images: 2,975 train, 500 val, 1,525 test; and 19,998 coarse-annotated images):

    • Setup: Input images (1024×20481024 \times 2048) were randomly cropped to 800×800800 \times 800, producing top-most feature maps of size 25×2525 \times 25. The GCN kernel size was correspondingly set to k=25k=25.
    • Training: Stage 1 trained on 22,973 images (coarse plus fine training sets); Stage 2 fine-tuned exclusively on the 2,975 fine training images.
    • Inference: Test images were split into four 1024×10241024 \times 1024 overlapping crops and their score maps fused.
    Validation Evaluation Phase GCN + BR Mean IoU (%)
    Stage-1 73.0
    Stage-2 76.9
    Stage-2 + Multi-Scale 77.2
    Stage-2 + Multi-Scale + CRF 77.4

    On the Cityscapes test set leaderboard, the model achieved 76.9% mean IoU, outperforming previous methods including LRR-4x (71.8%), Adelaide context (71.6%), DeepLabv2-CRF (70.4%), and Dilation10 (67.1%).

Coverage note — All primary contributed models (GCN, BR, ResNet-GCN), structural ablations, internal vs. boundary analyses, and benchmark evaluation protocols on PASCAL VOC 2012 and Cityscapes were included; qualitative visual segmentation figures (Figures 3, 6, 7) were omitted as their quantitative conclusions are fully represented in the empirical knowls.

References

  1. 1.A. Adams, J. Baek, and M. A. Davis. Fast high-dimensional filtering using the permutohedral lattice. In Computer Graphics Forum, volume 29, pages 753–762. Wiley Online Library, 2010. 2
  2. 2.A. Arnab, S. Jayasumana, S. Zheng, and P. H. Torr. Higher order conditional random fields in deep neural networks. In European Conference on Computer Vision, pages 524–540. Springer, 2016. 7
  3. 3.V. Badrinarayanan, A. Handa, and R. Cipolla. Segnet: A deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling. arXiv preprint arXiv:1505.07293, 2015. 2
  4. 4.J. T. Barron and B. Poole. The fast bilateral solver. ECCV, 2016. 2
  5. 5.S. Chandra and I. Kokkinos. Fast, exact and multi-scale inference for semantic image segmentation with deep gaussian crfs. arXiv preprint arXiv:1603.08358, 2016. 7
  6. 6.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015. 1, 2, 3, 6, 7
  7. 7.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv preprint arXiv:1606.00915, 2016. 2, 6, 7, 8
  8. 8.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. arXiv preprint arXiv:1604.01685, 2016. 4, 7
  9. 9.J. Dai, K. He, and J. Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1635–1643, 2015. 7
  10. 10.M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, 2015. 4
  11. 11.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 4
  12. 12.G. Ghiasi and C. C. Fowlkes. Laplacian pyramid reconstruction and refinement for semantic segmentation. In European Conference on Computer Vision, pages 519–534. Springer, 2016. 2, 7, 8
  13. 13.B. Hariharan, P. Arbelaez, L. Bourdev, S. Maji, and J. Malik. Semantic contours from inverse detectors. In 2011 International Conference on Computer Vision, pages 991–998. IEEE, 2011. 4
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. 1, 2, 3, 4
  15. 15.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of The 32nd International Conference on Machine Learning, pages 448–456, 2015. 6
  16. 16.V. Jampani, M. Kiefel, and P. V. Gehler. Learning sparse high dimensional filters: Image filtering, dense crfs and bilateral neural networks. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 2016. 2
  17. 17.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014. 4
  18. 18.V. Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. Adv. Neural Inf. Process. Syst, 2011. 2, 6
  19. 19.I. Krešo, D. Čaušević, J. Krapac, and S. Šegvić. Convolutional scale invariance for semantic segmentation. In German Conference on Pattern Recognition, pages 64–75. Springer, 2016. 8
  20. 20.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2, 4
  21. 21.G. Lin, C. Shen, A. van den Hengel, and I. Reid. Efficient piecewise training of deep structured models for semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. 2, 8
  22. 22.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. 6
  23. 23.W. Liu, A. Rabinovich, and A. C. Berg. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015. 2
  24. 24.Z. Liu, X. Li, P. Luo, C.-C. Loy, and X. Tang. Semantic image segmentation via deep parsing network. In Proceedings of the IEEE International Conference on Computer Vision, pages 1377–1385, 2015. 2, 6, 7, 8
  25. 25.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015. 1, 2, 3, 4
  26. 26.M. Mostajabi, P. Yadollahpour, and G. Shakhnarovich. Feedforward semantic segmentation with zoom-out features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3376–3385, 2015. 2, 7
  27. 27.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1520–1528, 2015. 2, 3
  28. 28.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. 4, 6
  29. 29.E. Shelhamer, J. Long, and T. Darrell. Fully convolutional networks for semantic segmentation. 2016. 2, 7, 8
  30. 30.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 1, 2, 5
  31. 31.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–9, 2015. 1, 2
  32. 32.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. arXiv preprint arXiv:1512.00567, 2015. 2, 3
  33. 33.Y. Wang, J. Liu, Y. Li, J. Yan, and H. Lu. Objectness-aware semantic segmentation. In Proceedings of the 2016 ACM on Multimedia Conference, pages 307–311. ACM, 2016. 7
  34. 34.Z. Wu, C. Shen, and A. v. d. Hengel. High-performance semantic segmentation using very deep fully convolutional networks. arXiv preprint arXiv:1604.04339, 2016. 7
  35. 35.S. Xie and Z. Tu. Holistically-nested edge detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 1395–1403, 2015. 3, 4
  36. 36.F. Yu and V. Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015. 2, 8
  37. 37.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr. Conditional random fields as recurrent neural networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 1529–1537, 2015. 2, 6, 7, 8
  38. 38.B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. Object detectors emerge in deep scene cnns. arXiv preprint arXiv:1412.6856, 2014. 3, 4

Citation

MLA
Peng, C., et al. “Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network”. arXiv, 2017, http://arxiv.org/abs/1703.02719v1.
APA
Peng, C., Zhang, X., Yu, G., Luo, G., & Sun, J. (2017). Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network. arXiv. http://arxiv.org/abs/1703.02719v1
Chicago
Peng, C., X. Zhang, G. Yu, G. Luo, and J. Sun. 2017. “Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network”. arXiv. http://arxiv.org/abs/1703.02719v1.
Harvard
Peng, C. et al. (2017) “Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1703.02719v1.
Vancouver
1. Peng C, Zhang X, Yu G, Luo G, Sun J (2017) Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network. arXiv

BibTeX

@article{peng2017large,
  title = {Large Kernel Matters -- Improve Semantic Segmentation by Global Convolutional Network},
  author = {Peng, Chao and Zhang, Xiangyu and Yu, Gang and Luo, Guiming and Sun, Jian},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1703.02719v1},
  eprint = {1703.02719}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE