CCNet: Criss-Cross Attention for Semantic Segmentation

Zilong HuangXinggang WangLichao HuangChang HuangYunchao WeiHumphrey ShiWenyu Liu

article2019ICCV3,087 citations

Introduces CCNet, a criss-cross attention network that captures global visual context for semantic segmentation while reducing GPU memory usage by eleven times and computation by 85 percent compared to standard non-local blocks.

Listen

Semantic segmentation assigns class labels to every pixel in an image and supports key applications such as autonomous driving and remote sensing. Conventional fully convolutional networks capture only local context, which reduces accuracy, while prior attention methods that gather full-image context incur high computational and memory costs that limit practical use.

The work set out to deliver full-image contextual information for semantic segmentation at far lower cost than existing attention mechanisms. The authors introduced the Criss-Cross Network (CCNet), which replaces dense attention with a recurrent criss-cross attention module. Each pixel aggregates information only along its horizontal and vertical paths; two successive modules allow every pixel to reach all other pixels. A category-consistent loss further encourages discriminative features, and the module extends naturally to three-dimensional video data.

Extensive experiments on Cityscapes, ADE20K, LIP, CamVid, and COCO show that CCNet reaches new state-of-the-art mean intersection-over-union scores of 81.9 percent, 45.76 percent, and 55.47 percent on the respective test or validation sets. The recurrent criss-cross module uses roughly one-eleventh the GPU memory and 85 percent fewer floating-point operations than a non-local block while still delivering higher accuracy. Adding the category-consistent loss yields an additional 0.7-point gain, and the same module improves instance segmentation when inserted into Mask R-CNN.

These results indicate that dense prediction tasks can obtain global context without prohibitive resource demands, lowering barriers to deployment on edge devices and in real-time systems. The approach also generalizes across image and video segmentation benchmarks.

The source code is publicly released, enabling immediate integration into existing fully convolutional pipelines. Further work could explore larger-scale video datasets and hardware-specific optimizations. The main limitations are that performance gains were measured on standard academic benchmarks rather than production-scale data, and the method still requires careful hyper-parameter tuning of the loss weights. Overall, the evidence supports confident adoption for accuracy-critical segmentation workloads where memory or compute budgets are constrained.

  • Paper: Non-local Neural Networks, Xiaolong Wang et al. (2018). Reading non-local neural networks first provides the essential background on non-local attention operations that CCNet adapts and makes more efficient via criss-cross paths.
Cover for CCNet: Criss-Cross Attention for Semantic Segmentation

Abstract

Contextual information is vital in visual understanding problems, such as semantic segmentation and object detection. We propose a Criss-Cross Network (CCNet) for obtaining full-image contextual information in a very effective and efficient way. Concretely, for each pixel, a novel criss-cross attention module harvests the contextual information of all the pixels on its criss-cross path. By taking a further recurrent operation, each pixel can finally capture the full-image dependencies. Besides, a category consistent loss is proposed to enforce the criss-cross attention module to produce more discriminative features. Overall, CCNet is with the following merits: 1) GPU memory friendly. Compared with the non-local block, the proposed recurrent criss-cross attention module requires 11x less GPU memory usage. 2) High computational efficiency. The recurrent criss-cross attention significantly reduces FLOPs by about 85% of the non-local block. 3) The state-of-the-art performance. We conduct extensive experiments on semantic segmentation benchmarks including Cityscapes, ADE20K, human parsing benchmark LIP, instance segmentation benchmark COCO, video segmentation benchmark CamVid. In particular, our CCNet achieves the mIoU scores of 81.9%, 45.76% and 55.47% on the Cityscapes test set, the ADE20K validation set and the LIP validation set respectively, which are the new state-of-the-art results. The source codes are available at \url{this https URL}.

Table of Contents

  • 2 RELATED WORK
  • 2.1 Semantic segmentation
  • 2.2 Contextual information aggregation
  • 2.3 Graph neural networks
  • 3 APPROACH
  • 3.1 Network Architecture
  • 3.2 Criss-Cross Attention
  • 3.3 Recurrent Criss-Cross Attention (RCCA)
  • 3.4 Learning Category Consistent Features
  • 3.5 3D Criss-Cross Attention
  • 4 EXPERIMENTS
  • 4.1 Datasets and Evaluation Metrics
  • 4.2 Implementation Details
  • 4.3 Experiments on Cityscapes
  • 4.3.1 Comparisons with state-of-the-arts
  • 4.4 Experiments on ADE20K
  • 4.5 Experiments on LIP
  • 4.6 Experiments on COCO
  • 4.7 Experiments on CamVid
  • 5 CONCLUSION AND FUTURE WORK
  • ACKNOWLEDGEMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Criss-Cross Attention Module for 2D Dense Context Aggregation

    model/method

    The Criss-Cross Attention (CCA) module captures contextual dependencies along the horizontal and vertical paths of each pixel at a reduced computational cost compared to dense self-attention mechanisms.

    Given an input feature map H∈RC×H×WH \in \mathbb{R}^{C \times H \times W}, the module applies two parallel 1×11 \times 1 convolutional layers to generate projected query and key representations Q,K∈RC′×H×WQ, K \in \mathbb{R}^{C' \times H \times W}, where C′<CC' < C reduces channel dimensionality. For each spatial location u=(ux,uy)u = (u_x, u_y), the feature vector is Qu∈RC′Q_u \in \mathbb{R}^{C'}. A set of key vectors Ωu∈R(H+W−1)×C′\Omega_u \in \mathbb{R}^{(H+W-1) \times C'} is extracted from KK corresponding to all positions located in the same row or column as uu. The affinity correlation di,ud_{i,u} between QuQ_u and the ii-th element Ωi,u∈Ωu\Omega_{i,u} \in \Omega_u is given by:

    di,u=QuΩi,uTd_{i,u} = Q_u \Omega_{i,u}^T

    where i∈{1,…,H+W−1}i \in \{1, \dots, H+W-1\}. Applying a softmax normalization along the channel dimension of the correlation tensor D∈R(H+W−1)×(H×W)D \in \mathbb{R}^{(H+W-1) \times (H \times W)} yields the sparse attention map A∈R(H+W−1)×(H×W)A \in \mathbb{R}^{(H+W-1) \times (H \times W)}.

    A third 1×11 \times 1 convolution transforms HH into an adapted feature representation V∈RC×H×WV \in \mathbb{R}^{C \times H \times W}. For each location uu, the set Φu∈R(H+W−1)×C\Phi_u \in \mathbb{R}^{(H+W-1) \times C} contains the feature vectors of VV lying on the same horizontal and vertical lines as uu. The aggregated feature Hu′H'_u is computed as:

    Hu′=∑i=0H+W−1Ai,uΦi,u+HuH'_u = \sum_{i=0}^{H+W-1} A_{i,u} \Phi_{i,u} + H_u

    where Hu′∈RCH'_u \in \mathbb{R}^C is the updated feature vector at location uu, and Ai,uA_{i,u} is the attention weight corresponding to index ii and position uu. This operation reduces the spatial complexity from O(N2)O(N^2) (where N=H×WN = H \times W) to O(NN)O(N\sqrt{N}) per operation.

  2. Knowl 2 — Recurrent Criss-Cross Attention for Full-Image Context Modeling

    model/method

    While a single pass through a Criss-Cross Attention (CCA) module collects context exclusively along a pixel's horizontal and vertical axes, a recurrent operation (RCCA) enables full-image, dense context propagation across all spatial locations without altering the parameter count.

    The RCCA module unrolls the CCA operation into RR successive loops where the network parameters of the CCA module are shared across all loops. In loop R=1R=1, the module takes the input feature map HH and produces intermediate features H′H' containing direct row and column dependencies. In loop R=2R=2, the module takes H′H' as input and generates H′′H''.

    For any target coordinate u=(ux,uy)u = (u_x, u_y) and an arbitrary coordinate θ=(θx,θy)\theta = (\theta_x, \theta_y) not initially in the criss-cross path of uu (i.e., θx≠ux\theta_x \neq u_x and θy≠uy\theta_y \neq u_y):

    1. In the first loop, the context from θ\theta propagates directly to the intermediate positions (ux,θy)(u_x, \theta_y) and (θx,uy)(\theta_x, u_y), because both lie on the horizontal and vertical axes of θ\theta.
    2. In the second loop, because (ux,θy)(u_x, \theta_y) and (θx,uy)(\theta_x, u_y) lie on the criss-cross path of uu, their updated representations (which now encode information from θ\theta) propagate into uu.

    Consequently, setting R=2R=2 allows every spatial location to harvest full-image dependencies across all H×WH \times W positions while retaining an overall time and memory complexity of O(NN)O(N\sqrt{N}) instead of O(N2)O(N^2).

  3. Knowl 3 — Category Consistent Loss with Robust Piece-Wise Distance

    equation

    To prevent feature over-smoothing in graph-like recurrent attention mechanisms and to encourage high intra-class compactness and inter-class separability, a category consistent loss (CCL) is applied to the dimension-reduced output features of the RCCA module:

    ℓ=ℓseg+αℓvar+βℓdis+γℓreg\ell = \ell_{\text{seg}} + \alpha \ell_{\text{var}} + \beta \ell_{\text{dis}} + \gamma \ell_{\text{reg}}

    where ℓseg\ell_{\text{seg}} is the standard cross-entropy segmentation loss, and α,β,γ\alpha, \beta, \gamma are weighting hyper-parameters set to α=1.0\alpha = 1.0, β=1.0\beta = 1.0, and γ=0.001\gamma = 0.001.

    The three constituent loss components are defined over the set of active classes CC present in the mini-batch:

    ℓvar=1∣C∣∑c∈C1Nc∑i=1Ncϕvar(hi,μc)\ell_{\text{var}} = \frac{1}{|C|} \sum_{c \in C} \frac{1}{N_c} \sum_{i=1}^{N_c} \phi_{\text{var}}(h_i, \mu_c)

    ℓdis=1∣C∣(∣C∣−1)∑ca∈C∑cb∈Cca≠cbϕdis(μca,μcb)\ell_{\text{dis}} = \frac{1}{|C|(|C|-1)} \sum_{c_a \in C} \sum_{\substack{c_b \in C \\ c_a \neq c_b}} \phi_{\text{dis}}(\mu_{c_a}, \mu_{c_b})

    ℓreg=1∣C∣∑c∈C∥μc∥\ell_{\text{reg}} = \frac{1}{|C|} \sum_{c \in C} \|\mu_c\|

    where NcN_c is the number of pixels belonging to category cc, hih_i is the feature vector at spatial position ii, and μc=1Nc∑i=1Nchi\mu_c = \frac{1}{N_c} \sum_{i=1}^{N_c} h_i is the cluster center for category cc.

    The distance functions ϕvar\phi_{\text{var}} and ϕdis\phi_{\text{dis}} employ margins δv=0.5\delta_v = 0.5 and δd=1.5\delta_d = 1.5. ϕvar\phi_{\text{var}} uses a piece-wise formulation (zero within δv\delta_v, quadratic in (δv,δd](\delta_v, \delta_d], and linear beyond δd\delta_d) to stabilize optimization against large gradient spikes:

    ϕvar(hi,μc)={∥μc−hi∥−δd+(δd−δv)2,if ∥μc−hi∥>δd(∥μc−hi∥−δv)2,if δv<∥μc−hi∥≤δd0,if ∥μc−hi∥≤δv\phi_{\text{var}}(h_i, \mu_c) = \begin{cases} \|\mu_c - h_i\| - \delta_d + (\delta_d - \delta_v)^2, & \text{if } \|\mu_c - h_i\| > \delta_d \\ (\|\mu_c - h_i\| - \delta_v)^2, & \text{if } \delta_v < \|\mu_c - h_i\| \le \delta_d \\ 0, & \text{if } \|\mu_c - h_i\| \le \delta_v \end{cases}

    ϕdis(μca,μcb)={(2δd−∥μca−μcb∥)2,if ∥μca−μcb∥≤2δd0,if ∥μca−μcb∥>2δd\phi_{\text{dis}}(\mu_{c_a}, \mu_{c_b}) = \begin{cases} (2\delta_d - \|\mu_{c_a} - \mu_{c_b}\|)^2, & \text{if } \|\mu_{c_a} - \mu_{c_b}\| \le 2\delta_d \\ 0, & \text{if } \|\mu_{c_a} - \mu_{c_b}\| > 2\delta_d \end{cases}

  4. Knowl 4 — 3D Criss-Cross Attention for Spatio-Temporal Video Context

    model/method

    The 3D Criss-Cross Attention module extends 2D spatial context aggregation to spatio-temporal video volumes H∈RC×T×H×WH \in \mathbb{R}^{C \times T \times H \times W}, where TT denotes the temporal/frame dimension.

    Two 1×1×11 \times 1 \times 1 convolutions map HH to feature representations Q,K∈RC′×T×H×WQ, K \in \mathbb{R}^{C' \times T \times H \times W}. For each voxel position u=(t,x,y)u = (t, x, y), the vector Qu∈RC′Q_u \in \mathbb{R}^{C'} interacts with a set Ωu∈R(T+H+W−2)×C′\Omega_u \in \mathbb{R}^{(T+H+W-2) \times C'} extracted from KK. Ωu\Omega_u contains all vectors sharing at least two coordinates with uu (i.e., varying along only the temporal dimension, horizontal spatial dimension, or vertical spatial dimension). The affinity correlation is:

    di,u=QuΩi,uTd_{i,u} = Q_u \Omega_{i,u}^T

    where i∈{1,…,T+H+W−2}i \in \{1, \dots, T+H+W-2\}. Applying a softmax along the first dimension yields the attention map A∈R(T+H+W−2)×T×H×WA \in \mathbb{R}^{(T+H+W-2) \times T \times H \times W}.

    A third 1×1×11 \times 1 \times 1 convolution produces V∈RC×T×H×WV \in \mathbb{R}^{C \times T \times H \times W}, from which a local set Φu∈R(T+H+W−2)×C\Phi_u \in \mathbb{R}^{(T+H+W-2) \times C} is formed along the 3D criss-cross structure centered at uu. The aggregated feature vector Hu′H'_u is computed as:

    Hu′=∑i=0T+H+W−2Ai,uΦi,u+HuH'_u = \sum_{i=0}^{T+H+W-2} A_{i,u} \Phi_{i,u} + H_u

    Repeating this module for R=3R = 3 iterations allows every voxel to aggregate dense contextual information across the entire spatial and temporal dimensions of the input sequence.

  5. Knowl 5 — CCNet Architecture for Semantic Segmentation

    model/method

    CCNet incorporates the Recurrent Criss-Cross Attention (RCCA) module into a deep fully convolutional architecture for dense semantic prediction:

    1. Feature Extraction: An input image is passed through a convolutional backbone (e.g., ResNet-101) modified by removing the last two downsampling operations and using dilated convolutions with dilation rates of 2 and 4 in stages 4 and 5, producing an output feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W} with an output stride of 8 (spatial dimensions H=Hin/8,W=Win/8H = H_{\text{in}}/8, W = W_{\text{in}}/8).
    2. Context Embedding: A 1×11 \times 1 convolution reduces the channel dimension of XX to form feature map HH. HH is processed through the RCCA module for R=2R=2 loops to produce the dense full-image contextual feature map H′′H''.
    3. Feature Fusion and Prediction: The contextual feature H′′H'' is concatenated along channels with the local feature map XX. One or more convolutional layers with batch normalization and non-linear activation fuse the concatenated features. A final segmentation layer predicts class probability maps.
  6. Knowl 6 — Efficiency and Resource Comparison: RCCA vs. Non-Local Self-Attention

    data/table

    On the Cityscapes validation set with an input resolution of 1×3×769×7691 \times 3 \times 769 \times 769 and a ResNet-50 backbone, Recurrent Criss-Cross Attention (R=2R=2) achieves higher segmentation accuracy than standard Non-Local (NL) attention while reducing computational cost and memory footprint by substantial margins.

    Method GFLOPs (Δ\Delta) Memory (MB Δ\Delta) mIoU (%)
    Baseline 0 0 73.3
    +NL 108 1411 77.3
    +NL (R=2R=2) 216 2820 78.7
    +RCCA (R=2R=2) 16.5 127 78.5

    Compared to the Non-Local block (+NL), RCCA (R=2R=2) uses 11×11\times less GPU memory overhead (127 MB127\text{ MB} vs. 1411 MB1411\text{ MB}) and reduces FLOPs by approximately 84.7%84.7\% (16.5 GFLOPs16.5\text{ GFLOPs} vs. 108 GFLOPs108\text{ GFLOPs}), while improving mIoU by 1.2%1.2\% over single-pass Non-Local attention (78.5%78.5\% vs. 77.3%77.3\%).

  7. Knowl 7 — Ablation of Recurrence Loops and Loss Functions on Cityscapes

    data/table

    Ablation experiments on the Cityscapes validation set evaluate the effect of varying RCCA loop counts (RR) and context modeling configurations on segmentation performance (mIoU).

    Setting / Method Backbone Memory / Details mIoU (%)
    Baseline ResNet-101 0 MB Δ\Delta, 0 GFLOPs Δ\Delta 75.1
    R=1R = 1 ResNet-101 53 MB Δ\Delta, 8.3 GFLOPs Δ\Delta 78.0
    R=2R = 2 ResNet-101 127 MB Δ\Delta, 16.5 GFLOPs Δ\Delta 79.8
    R=3R = 3 ResNet-101 208 MB Δ\Delta, 24.7 GFLOPs Δ\Delta 80.2
    ResNet50 + GCN ResNet-50 Global conv filters 76.2
    ResNet50 + PSP ResNet-50 Pyramid spatial pooling 76.4
    ResNet50 + ASPP ResNet-50 Multi-dilation conv 77.1
    ResNet50 + NL ResNet-50 Full self-attention 77.3
    ResNet50 + HV ResNet-50 Sequential horiz. + vert. 77.3
    ResNet50 + HVVH ResNet-50 Parallel horiz./vert. sum 77.8
    ResNet50 + RCCA (R=2R=2) ResNet-50 Criss-cross recurrence 78.5
    ResNet50 + RCCA + CCL ResNet-50 With Category Consistent Loss 79.3
    ResNet101 + RCCA (R=2R=2) + CCL ResNet-101 With Category Consistent Loss 80.5

    Adding a single criss-cross loop (R=1R=1) improves baseline ResNet-101 performance by 2.9%2.9\% mIoU. Increasing to R=2R=2 provides an additional 1.8%1.8\% mIoU gain, with diminishing returns at R=3R=3 (+0.4%+0.4\% mIoU). CCL provides a consistent +0.7%+0.7\% to +0.8%+0.8\% mIoU boost across both ResNet-50 and ResNet-101 backbones. Training stability tests (10 runs on ResNet-50) show that the piece-wise distance function achieves a 9/109/10 success rate (mean mIoU 79.3%79.3\%) versus 6/106/10 (mean mIoU 79.2%79.2\%) for a single quadratic distance function.

  8. Knowl 8 — Semantic Segmentation Benchmarks on Cityscapes Test Set

    data/table

    When trained on both the train-fine and val-fine sets of Cityscapes using a ResNet-101 backbone, CCNet achieves state-of-the-art performance on the official Cityscapes evaluation server.

    Method Backbone mIoU (%)
    DeepLab-v2 ResNet-101 70.4
    RefineNet ResNet-101 73.6
    GCN ResNet-101 76.9
    DUC ResNet-101 77.6
    SAC ResNet-101 78.1
    ResNet-38 WiderResNet-38 78.4
    PSPNet ResNet-101 78.4
    BiSeNet ResNet-101 78.9
    AAF ResNet-101 79.1
    DFN ResNet-101 79.3
    PSANet ResNet-101 80.1
    DenseASPP DenseNet-161 80.6
    CCNet ResNet-101 81.9

    CCNet outperforms all existing methods on the Cityscapes test benchmark, including PSANet (80.1%80.1\%) and DenseASPP with DenseNet-161 (80.6%80.6\%), achieving 81.9%81.9\% mIoU under single-scale evaluation.

  9. Knowl 9 — Evaluation on Scene Parsing (ADE20K) and Human Parsing (LIP)

    data/table

    CCNet achieves state-of-the-art accuracy on the ADE20K validation set (150 categories) and the Look Into Person (LIP) human parsing validation set (20 categories).

    Dataset Method Backbone Mean Acc (%) mIoU (%)
    ADE20K (val) RefineNet ResNet-152 – 40.70
    ADE20K (val) PSPNet ResNet-101 – 43.29
    ADE20K (val) PSANet ResNet-101 – 43.77
    ADE20K (val) EncNet ResNet-101 – 44.65
    ADE20K (val) CCNet ResNet-101 – 45.76
    LIP (val) DeepLab (ResNet-101) ResNet-101 55.63 44.80
    LIP (val) SAN VGG-16 55.09 44.81
    LIP (val) JPPNet ResNet-101 62.32 51.37
    LIP (val) CE2P ResNet-101 63.20 53.10
    LIP (val) CCNet ResNet-101 63.91 55.47

    On ADE20K validation, CCNet with CCL obtains 45.76%45.76\% mIoU, outperforming EncNet (44.65%44.65\%) by 1.11%1.11\%. On LIP validation, CCNet embedded in the CE2P framework achieves 55.47%55.47\% mIoU and 88.01%88.01\% pixel accuracy, outperforming the CE2P baseline (53.10%53.10\%) by 2.37%2.37\% mIoU.

  10. Knowl 10 — Generalization to Instance and Video Semantic Segmentation (COCO and CamVid)

    data/table

    Incorporating RCCA into Mask R-CNN on the COCO dataset and utilizing 3D-RCCA on the CamVid video dataset demonstrates generalization beyond 2D static semantic segmentation.

    Task / Dataset Model / Backbone Metric 1 Metric 2
    COCO (val) Instance Seg. ResNet-50 Baseline 38.2 APbox\text{AP}^{\text{box}} 34.8 APmask\text{AP}^{\text{mask}}
    COCO (val) Instance Seg. ResNet-50 + NL 39.0 APbox\text{AP}^{\text{box}} 35.5 APmask\text{AP}^{\text{mask}}
    COCO (val) Instance Seg. ResNet-50 + RCCA 39.3 APbox\text{AP}^{\text{box}} 36.1 APmask\text{AP}^{\text{mask}}
    COCO (val) Instance Seg. ResNet-101 Baseline 40.1 APbox\text{AP}^{\text{box}} 36.2 APmask\text{AP}^{\text{mask}}
    COCO (val) Instance Seg. ResNet-101 + NL 40.8 APbox\text{AP}^{\text{box}} 37.1 APmask\text{AP}^{\text{mask}}
    COCO (val) Instance Seg. ResNet-101 + RCCA 41.0 APbox\text{AP}^{\text{box}} 37.3 APmask\text{AP}^{\text{mask}}
    CamVid (test) Video Seg. PSPNet (ResNet-50) – 69.1 mIoU (%)
    CamVid (test) Video Seg. VideoGCRF (ResNet-101) – 75.2 mIoU (%)
    CamVid (test) Video Seg. CCNet3D (T=1T=1, ResNet-101) – 77.9 mIoU (%)
    CamVid (test) Video Seg. CCNet3D (T=5T=5, ResNet-101) – 79.1 mIoU (%)

    On COCO instance segmentation, adding RCCA before the final residual block of res4 outperforms both standard Mask R-CNN and Non-Local block (+NL) augmentations. On CamVid video segmentation, CCNet3D with T=5T=5 frames achieves 79.1%79.1\% mIoU, improving by 1.2%1.2\% over the single-frame CCNet3D (T=1T=1) and outperforming VideoGCRF (75.2%75.2\%).

Coverage note — None was omitted; all key architectural components (2D CCA, RCCA, CCL, 3D CCA, CCNet full pipeline) and all experimental benchmarks across Cityscapes, ADE20K, LIP, COCO, and CamVid have been fully captured.

References

  1. 1.J. Fritsch, T. Kuehnl, and A. Geiger, “A new performance measure and evaluation benchmark for road detection algorithms,” in ITSC, 2013, pp. 1693–1700.
  2. 2.R. T. Azuma, “A survey of augmented reality,” Presence: Teleoperators & Virtual Environments, vol. 6, no. 4, pp. 355–385, 1997.
  3. 3.M. Evening, Adobe Photoshop CS3 for photographers: a professional image editor’s guide to the creative use of Photoshop for the Macintosh and PC. Focal press, 2012.
  4. 4.Y. Song, Z. Huang, C. Shen, H. Shi, and D. A. Lange, “Deep learning-based automated image segmentation for concrete petrographic analysis,” Cement and Concrete Research, vol. 135, p. 106118, 2020.
  5. 5.Z. Zheng, Y. Zhong, J. Wang, and A. Ma, “Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery,” in CVPR, 2020.
  6. 6.M. T. Chiu, X. Xu, Y. Wei, Z. Huang, A. G. Schwing, R. Brunner, H. Khachatrian, H. Karapetyan, I. Dozier, G. Rose et al., “Agriculture-vision: A large aerial image database for agricultural pattern analysis,” in CVPR, 2020, pp. 2828–2838.
  7. 7.M. Tik Chiu, X. Xu, K. Wang, J. Hobbs, N. Hovakimyan, S. H. Huang, Thomas S et al., “The 1st agriculture-vision challenge: Methods and results,” in CVPR Workshops, 2020, pp. 48–49.
  8. 8.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR, 2015, pp. 3431–3440.
  9. 9.X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR, 2018, pp. 7794–7803.
  10. 10.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE TPAMI, vol. 40, no. 4, pp. 834–848, 2018.
  11. 11.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR, 2017, pp. 2881–2890.
  12. 12.L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587, 2017.
  13. 13.H. Ding, X. Jiang, B. Shuai, A. Qun Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmentation,” in CVPR, June 2018.
  14. 14.H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in CVPR, 2018.
  15. 15.F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE TNN, vol. 20, no. 1, pp. 61–80, 2008.
  16. 16.H. Zhao, Y. Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia, “Psanet: Point-wise spatial attention network for scene parsing,” in ECCV, 2018, pp. 270–286.
  17. 17.J. Cheng, L. Dong, and M. Lapata, “Long short-term memory-networks for machine reading,” 2016.
  18. 18.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS, 2017, pp. 5998–6008.
  19. 19.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR, 2016, pp. 3213–3223.
  20. 20.B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in CVPR, 2017.
  21. 21.X. Liang, K. Gong, X. Shen, and L. Lin, “Look into person: Joint body parsing & pose estimation network and a new benchmark,” IEEE TPAMI, vol. 41, no. 4, pp. 871–885, 2018.
  22. 22.G. J. Brostow, J. Fauqueur, and R. Cipolla, “Semantic object classes in video: A high-definition ground truth database,” Pattern Recognition Letters, vol. 30, no. 2, pp. 88–97, 2009.
  23. 23.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778.
  24. 24.Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,” in ICCV, 2019.
  25. 25.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Semantic image segmentation with deep convolutional nets and fully connected crfs,” ICLR, 2015.
  26. 26.F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” ICLR, 2016.
  27. 27.O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention, 2015, pp. 234–241.
  28. 28.L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” ECCV, 2018.
  29. 29.D. Lin, Y. Ji, D. Lischinski, D. Cohen-Or, and H. Huang, “Multi-scale context intertwining for semantic segmentation,” in ECCV, 2018, pp. 603–619.
  30. 30.B. Cheng, L.-C. Chen, Y. Wei, Y. Zhu, Z. Huang, J. Xiong, T. Huang, W.-M. Hwu, and H. Shi, “Spgnet: Semantic prediction guidance for scene parsing,” in ICCV, 2019.
  31. 31.G. Lin, A. Milan, C. Shen, and I. D. Reid, “Refinenet: Multi-path refinement networks for high-resolution semantic segmentation.” in CVPR, 2017.
  32. 32.C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning a discriminative feature network for semantic segmentation,” CVPR, 2018.
  33. 33.R. Zhang, S. Tang, Y. Zhang, J. Li, and S. Yan, “Scale-adaptive convolutions for scene parsing,” in ICCV, 2017, pp. 2031–2039.
  34. 34.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” ICCV, 2017.
  35. 35.Z. Liu, X. Li, P. Luo, C.-C. Loy, and X. Tang, “Semantic image segmentation via deep parsing network,” in ICCV, 2015, pp. 1377–1385.
  36. 36.T.-W. Ke, J.-J. Hwang, Z. Liu, and S. X. Yu, “Adaptive affinity field for semantic segmentation,” ECCV, 2018.
  37. 37.C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Bisenet: Bilateral segmentation network for real-time semantic segmentation,” ECCV, 2018.
  38. 38.P. Bilinski and V. Prisacariu, “Dense decoder shortcut connections for single-pass semantic segmentation,” in CVPR, 2018, pp. 6596–6605.
  39. 39.S. Chandra, C. Couprie, and I. Kokkinos, “Deep spatio-temporal random fields for efficient video segmentation,” in CVPR, 2018, pp. 8915–8924.
  40. 40.P.-Y. Huang, W.-T. Hsu, C.-Y. Chiu, T.-F. Wu, and M. Sun, “Efficient uncertainty estimation for semantic segmentation in videos,” in ECCV, 2018, pp. 520–535.
  41. 41.T. Ruan, T. Liu, Z. Huang, Y. Wei, S. Wei, and Y. Zhao, “Devil in the details: Towards accurate single and multiple human parsing,” in AAAI, vol. 33, 2019, pp. 4814–4821.
  42. 42.Z. Huang, C. Wang, X. Wang, W. Liu, and J. Wang, “Semantic image segmentation by scale-adaptive networks,” IEEE TIP, 2019.
  43. 43.Z. Wang, M. Yu, Y. Wei, R. Feris, J. Xiong, W.-m. Hwu, T. S. Huang, and H. Shi, “Differential treatment for stuff and things: A simple unsupervised domain adaptation method for semantic segmentation,” in CVPR, 2020, pp. 12 635–12 644.
  44. 44.Z. Wang, Y. Wei, R. Feris, J. Xiong, W.-M. Hwu, T. S. Huang, and H. Shi, “Alleviating semantic-level shift: A semi-supervised domain adaptation method for semantic segmentation,” in CVPR Workshops, 2020, pp. 936–937.
  45. 45.J. Jiao, Y. Wei, Z. Jie, H. Shi, R. W. Lau, and T. S. Huang, “Geometry-aware distillation for indoor semantic segmentation,” in CVPR, 2019, pp. 2869–2878.
  46. 46.Z. Huang, X. Wang, J. Wang, W. Liu, and J. Wang, “Weakly-supervised semantic segmentation network with deep seeded region growing,” in CVPR, 2018, pp. 7014–7023.
  47. 47.Y. Wei, H. Xiao, H. Shi, Z. Jie, J. Feng, and T. S. Huang, “Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation,” in CVPR, 2018, pp. 7268–7277.
  48. 48.R. Qian, Y. Wei, H. Shi, J. Li, J. Liu, and T. Huang, “Weakly supervised scene parsing with point-based distance metric learning,” in AAAI, vol. 33, 2019, pp. 8843–8850.
  49. 49.M. Yang, K. Yu, C. Zhang, Z. Li, and K. Yang, “Denseaspp for semantic segmentation in street scenes,” in CVPR, 2018, pp. 3684–3692.
  50. 50.L.-C. Chen, M. D. Collins, Y. Zhu, G. Papandreou, B. Zoph, F. Schroff, H. Adam, and J. Shlens, “Searching for efficient multi-scale architectures for dense image prediction,” NeurIPS, 2018.
  51. 51.L.-C. Chen, Y. Yang, J. Wang, W. Xu, and A. L. Yuille, “Attention to scale: Scale-aware semantic image segmentation,” in CVPR, 2016, pp. 3640–3649.
  52. 52.C. Liu, L.-C. Chen, F. Schroff, H. Adam, W. Hua, A. L. Yuille, and L. Fei-Fei, “Auto-deeplab: Hierarchical neural architecture search for semantic image segmentation,” in CVPR, 2019, pp. 82–92.
  53. 53.J. He, Z. Deng, L. Zhou, Y. Wang, and Y. Qiao, “Adaptive pyramid context network for semantic segmentation,” in CVPR, 2019, pp. 7519–7528.
  54. 54.S. Liu, S. De Mello, J. Gu, G. Zhong, M.-H. Yang, and J. Kautz, “Learning affinity via spatial propagation networks,” in NeurIPS, 2017, pp. 1520–1530.
  55. 55.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr, “Conditional random fields as recurrent neural networks,” in ICCV, 2015, pp. 1529–1537.
  56. 56.Y. Yuan and J. Wang, “Ocnet: Object context network for scene parsing,” arXiv preprint arXiv:1809.00916, 2018.
  57. 57.J. Fu, J. Liu, H. Tian, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” CVPR, 2019.
  58. 58.Y. Chen, M. Rohrbach, Z. Yan, Y. Shuicheng, J. Feng, and Y. Kalantidis, “Graph-based global reasoning networks,” in CVPR, 2019, pp. 433–442.
  59. 59.C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel mattersimprove semantic segmentation by global convolutional network,” in CVPR, 2017, pp. 1743–1751.
  60. 60.A. Sperduti and A. Starita, “Supervised neural networks for the classification of structures,” IEEE TNN, vol. 8, no. 3, pp. 714–735, 1997.
  61. 61.M. Gori, G. Monfardini, and F. Scarselli, “A new model for learning in graph domains,” in IJCNN, vol. 2, 2005, pp. 729–734.
  62. 62.M. Henaff, J. Bruna, and Y. LeCun, “Deep convolutional networks on graph-structured data,” arXiv preprint arXiv:1506.05163, 2015.
  63. 63.M. Defferrard, X. Bresson, and P. Vandergheynst, “Convolutional neural networks on graphs with fast localized spectral filtering,” in NeurIPS, 2016, pp. 3844–3852.
  64. 64.T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” arXiv preprint arXiv:1609.02907, 2016.
  65. 65.R. Levie, F. Monti, X. Bresson, and M. M. Bronstein, “Cayleynets: Graph convolutional neural networks with complex rational spectral filters,” IEEE TSP, vol. 67, no. 1, pp. 97–109, 2018.
  66. 66.J. Atwood and D. Towsley, “Diffusion-convolutional neural networks,” in NeurIPS, 2016, pp. 1993–2001.
  67. 67.M. Niepert, M. Ahmed, and K. Kutzkov, “Learning convolutional neural networks for graphs,” in ICML, 2016, pp. 2014–2023.
  68. 68.J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural message passing for quantum chemistry,” in ICML, 2017, pp. 1263–1272.
  69. 69.B. De Brabandere, D. Neven, and L. Van Gool, “Semantic instance segmentation with a discriminative loss function,” arXiv preprint arXiv:1708.02551, 2017.
  70. 70.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV, 2014, pp. 740–755.
  71. 71.G. J. Brostow, J. Shotton, J. Fauqueur, and R. Cipolla, “Segmentation and recognition using structure from motion point clouds,” in ECCV, 2008, pp. 44–57.
  72. 72.K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV, 2017, pp. 2980–2988.
  73. 73.P. Wang, P. Chen, Y. Yuan, D. Liu, Z. Huang, X. Hou, and G. Cottrell, “Understanding convolution for semantic segmentation,” in WACV, 2018, pp. 1451–1460.
  74. 74.Z. Wu, C. Shen, and A. v. d. Hengel, “Wider or deeper: Revisiting the resnet model for visual recognition,” arXiv preprint arXiv:1611.10080, 2016.
  75. 75.X. Liang, H. Zhou, and E. Xing, “Dynamic-structured semantic propagation network,” in CVPR, 2018, pp. 752–761.
  76. 76.T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” ECCV, 2018.
  77. 77.V. Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE TPAMI, vol. 39, no. 12, pp. 2481–2495, 2017.

Citation

MLA
Huang, Z., et al. “CCNet: Criss-Cross Attention for Semantic Segmentation”. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 603–12, https://doi.org/10.1109/ICCV.2019.00069.
APA
Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., & Liu, W. (2019). CCNet: Criss-Cross Attention for Semantic Segmentation. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 603–612. https://doi.org/10.1109/ICCV.2019.00069
Chicago
Huang, Z., X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu. 2019. “CCNet: Criss-Cross Attention for Semantic Segmentation”. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 603–12. https://doi.org/10.1109/ICCV.2019.00069.
Harvard
Huang, Z. et al. (2019) “CCNet: Criss-Cross Attention for Semantic Segmentation”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp. 603–612. Available at: https://doi.org/10.1109/ICCV.2019.00069.
Vancouver
1. Huang Z, Wang X, Huang L, Huang C, Wei Y, Liu W (2019) CCNet: Criss-Cross Attention for Semantic Segmentation. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp 603–612

BibTeX

@inproceedings{Huang_2019, title={CCNet: Criss-Cross Attention for Semantic Segmentation}, url={http://dx.doi.org/10.1109/ICCV.2019.00069}, DOI={10.1109/iccv.2019.00069}, booktitle={2019 IEEE/CVF International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Huang, Zilong and Wang, Xinggang and Huang, Lichao and Huang, Chang and Wei, Yunchao and Liu, Wenyu}, year={2019}, month=Oct, pages={603–612} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE