Large Kernel Matters — Improve Semantic Segmentation by Global Convolutional Network
Chao PengXiangyu ZhangGang YuGuiming LuoJian Sun
Proposes the Global Convolutional Network architecture with large separable kernels and residual boundary refinement to simultaneously address classification and localization challenges in semantic segmentation, achieving state-of-the-art accuracy on PASCAL VOC 2012 and Cityscapes.
Semantic segmentation—the computer vision task of assigning an accurate category label to every pixel in an image—is essential for autonomous driving, robotics, and image analysis. However, it faces a fundamental trade-off between category recognition and precise spatial localization. Category recognition requires broad context and invariance to transformations like rotation and shifting, whereas spatial localization demands precise sensitivity to pixel locations. Prior systems predominantly favored localization by using narrow, stacked computational filters, which restricted context and impaired recognition for large objects.
The article demonstrates an architecture called the Global Convolutional Network (GCN) coupled with a Boundary Refinement (BR) block to resolve this conflict. The core objective is to evaluate whether expanding the effective context area using large, computationally efficient convolutional filters improves pixel-level semantic labeling without sacrificing spatial accuracy.
The researchers designed an end-to-end framework based on high-capacity residual networks (ResNet-152) and benchmarked it on two standard public datasets: PASCAL VOC 2012, which contains diverse object classes, and Cityscapes, which comprises complex urban street scenes. To avoid the computational penalty of traditional large filters, the design decomposed broad two-dimensional filters into combinations of one-dimensional horizontal and vertical operations. The authors conducted extensive ablation experiments to isolate the effects of filter size, model parameters, and boundary alignment modules.
The experiments produced four critical findings. First, larger filter sizes consistently improved segmentation accuracy; expanding the filter size to span the entire feature map boosted accuracy on the PASCAL VOC validation set by 5.5 percentage points over small-filter baselines. Second, the decomposed large-filter structure outperformed both standard large filters and stacks of small filters while using significantly fewer parameters and avoiding training convergence issues. Third, error analysis showed that the large-filter network primarily improved the internal classification of objects (increasing internal accuracy to 95.0%), while the boundary refinement module improved alignment along object edges (raising boundary accuracy from 71.5% to 73.4%). Fourth, the full system established new state-of-the-art benchmarks, achieving 82.2% mean intersection-over-union on PASCAL VOC 2012 and 76.9% on Cityscapes, outperforming previous methods by 2.0 and 5.1 percentage points, respectively.
These findings demonstrate that semantic segmentation models do not need to choose between wide context recognition and spatial localization. The decomposed filter design delivers superior visual understanding with lower computational cost and fewer parameters than standard approaches. This provides direct benefits for real-world deployments by reducing processing overhead and memory usage while enhancing scene understanding in safety-critical applications such as autonomous navigation.
For engineering and development teams building visual perception systems, the article supports replacing traditional stacked small-kernel blocks with decomposed large-kernel modules and integrating residual boundary refinement. When targeting high accuracy on urban or multi-object scenes, teams should adopt multi-stage pre-training on broader datasets before final fine-tuning.
The evaluation is highly credible due to strong benchmark results, but readers should note certain limitations. The primary evaluations relied on deep, high-capacity base networks, and testing very large input images required cropping, multi-scale evaluation, and post-processing steps to reach peak performance. Further work is needed to validate real-time inference latency and efficiency on resource-constrained embedded hardware.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It establishes the foundational fully convolutional network paradigm for dense per-pixel prediction that Global Convolutional Network builds upon.
- Paper: Understanding the Effective Receptive Field in Deep Convolutional Neural Networks, Wenjie Luo et al. (2016). It provides the theoretical and empirical foundation for effective receptive fields in deep neural networks, which directly motivates the source paper's large-kernel architecture design.
- Paper: DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, Liang-Chieh Chen et al. (2016). It establishes key techniques for dense prediction and boundary handling, including atrous convolution and multi-scale context aggregation, serving as a primary benchmark and motivation.
- Paper: Multi-Scale Context Aggregation by Dilated Convolutions, Fisher Yu et al. (2016). It introduces dilated convolutions to expand receptive fields without downsampling, offering an alternative mechanism that Global Convolutional Network contrasts with large kernel filters.
- Paper: Pyramid Scene Parsing Network, Hengshuang Zhao et al. (2016). It demonstrates the necessity of incorporating global scene context for dense semantic segmentation via pyramid pooling.
- Paper: RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation, Guosheng Lin et al. (2016). It develops multi-path residual refinement blocks across feature hierarchies to capture fine spatial details and precise object boundaries.
- Paper: SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation, Vijay Badrinarayanan et al. (2015). It introduces an encoder-decoder architecture with efficient index-based upsampling for recovering high-resolution segmentation details.
- Paper: Learning Deep Features for Discriminative Localization, Bolei Zhou et al. (2016). It analyzes the trade-off and coordination between deep feature classification and spatial localization in convolutional architectures.
- Paper: Rethinking Atrous Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2017). It advances global context capture in semantic segmentation by integrating image-level features and parallel multi-rate convolutions.
- Paper: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2018). It combines atrous spatial pyramid pooling with an explicit encoder-decoder structure to refine segmentation boundaries.
- Paper: BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation, Changqian Yu et al. (2018). It decouples spatial detail preservation from contextual semantic learning to achieve efficient real-time segmentation.
- Paper: Dual Attention Network for Scene Segmentation, Jun Fu et al. (2019). It advances beyond large convolutional kernels by employing spatial and channel self-attention to capture long-range contextual dependencies.
- Paper: CCNet: Criss-Cross Attention for Semantic Segmentation, Zilong Huang et al. (2019). It presents a lightweight criss-cross attention module to aggregate full-image contextual information more efficiently than conventional dense attention.
- Paper: Deep High-Resolution Representation Learning for Visual Recognition, Jingdong Wang et al. (2019). It maintains high-resolution representations throughout the entire network architecture rather than downsampling and upsampling features.
- Paper: Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers, Sixiao Zheng et al. (2020). It replaces convolutional receptive-field expansion mechanisms entirely by reframing semantic segmentation as a global sequence-to-sequence transformer task.
- Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). It provides a comprehensive survey synthesizing deep learning segmentation paradigms, including large receptive-field and context-aggregation strategies.
