GCNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond

Yue CaoJiarui XuStephen LinFangyun WeiHan Hu

article2019ICCV1,975 citations

Develops a lightweight Global Context network that unifies Non-Local Networks and Squeeze-Excitation Networks, capturing long-range dependencies with significantly reduced computation across visual recognition tasks.

Listen

Modern computer vision models rely heavily on understanding full-scene context, known as long-range dependency, to accurately classify images, detect objects, and recognize actions in video. Standard convolutional networks struggle to capture this global context efficiently, while existing self-attention techniques—such as Non-Local Networks—demand high computational power and memory. This computational overhead limits these techniques to only one or two layers within a network architecture.

The article evaluates how Non-Local Networks model global context, identifies structural redundancies within them, and demonstrates a lightweight alternative architecture called the Global Context Network (GCNet).

To conduct the study, the authors performed empirical and statistical analyses of attention behaviors in Non-Local Networks across standard computer vision benchmarks, including COCO for object detection and segmentation, ImageNet for image classification, and Kinetics for action recognition. Based on findings from these analyses, they unified Non-Local Networks and Squeeze-Excitation Networks into a three-step framework covering context modeling, feature transformation, and feature fusion, and tested various architectural combinations against established baseline networks.

The analysis produced several key findings. First, Non-Local Networks learn attention maps that are almost identical across different positions within an image, proving that calculating expensive position-specific attention maps is largely redundant. Second, replacing these calculations with a shared, position-independent attention map maintains model accuracy while significantly reducing computational operations. Third, combining this simplified attention pooling with a compact bottleneck transformation and addition-based feature fusion creates a lightweight Global Context block that can be integrated across multiple network layers. Across benchmarks, applying this block across all stages improved object detection by 2.7% on COCO, image classification top-1 accuracy by 0.8% on ImageNet, and video action recognition by 1.1% on Kinetics, all while adding less than 0.26% computational overhead.

These findings indicate that artificial intelligence systems can achieve superior global context modeling without suffering the computational and financial burdens typically associated with attention mechanisms. Integrating lightweight global context across all network stages delivers higher recognition accuracy and lower latency compared to heavily restricted, single-layer deployments.

Teams developing and deploying computer vision architectures should adopt Global Context blocks across network backbones to boost visual accuracy with minimal computational cost. Engineering teams can leverage the authors' publicly released implementation to evaluate bottleneck ratios—such as a factor of 16 for standard efficiency or 4 for higher accuracy—based on specific hardware and latency budgets.

The findings are supported with high confidence across multiple standardized image and video benchmarks. However, the study evaluates these mechanisms within standard ResNet-style backbones on benchmark datasets, meaning performance should be re-evaluated when applying the architecture to custom edge-computing hardware or novel backbone designs.

  • Paper: Non-local Neural Networks, Xiaolong Wang et al. (2018). Non-Local Neural Networks establish the foundational query-specific global context aggregation module that GCNet directly analyzes and simplifies.
  • Paper: Squeeze-and-Excitation Networks, Jie Hu et al. (2018). Squeeze-and-Excitation Networks introduce the lightweight channel attention mechanism that GCNet explicitly unifies with simplified non-local blocks.
  • Paper: Group Normalization, Yuxin Wu et al. (2018). Group Normalization provides the batch-independent normalization technique used inside GCNet's bottleneck transform to ease optimization.
  • Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). Path Aggregation Network establishes the instance segmentation and feature pyramid framework utilized in GCNet's downstream visual recognition evaluations.
Cover for GCNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond

Abstract

The Non-Local Network (NLNet) presents a pioneering approach for capturing long-range dependencies, via aggregating query-specific global context to each query position. However, through a rigorous empirical analysis, we have found that the global contexts modeled by non-local network are almost the same for different query positions within an image. In this paper, we take advantage of this finding to create a simplified network based on a query-independent formulation, which maintains the accuracy of NLNet but with significantly less computation. We further observe that this simplified design shares similar structure with Squeeze-Excitation Network (SENet). Hence we unify them into a three-step general framework for global context modeling. Within the general framework, we design a better instantiation, called the global context (GC) block, which is lightweight and can effectively model the global context. The lightweight property allows us to apply it for multiple layers in a backbone network to construct a global context network (GCNet), which generally outperforms both simplified NLNet and SENet on major benchmarks for various recognition tasks. The code and configurations are released at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Analysis on Non-local Networks
  • 3.1 Revisiting the Non-local Block
  • 3.2 Analysis
  • 4 Method
  • 4.1 Simplifying the Non-local Block
  • 4.2 Global Context Modeling Framework
  • 4.3 Global Context Block
  • 5 Experiments
  • 5.1 Object Detection/Segmentation on COCO
  • 5.1.1 Ablation Study
  • 5.1.2 Experiments on Stronger Backbones
  • 5.2 Image Classification on ImageNet
  • 5.3 Action Recognition on Kinetics
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Global Context (GC) Block Architecture

    model/method

    The Global Context (GC) block is an architectural unit designed to capture long-range visual dependencies by combining query-independent global context modeling with lightweight channel transform and additive fusion.

    Let x={xi}i=1Np\mathbf{x} = \{\mathbf{x}_i\}_{i=1}^{N_p} denote the input feature map with channel dimension CC and spatial/spatiotemporal positions NpN_p (where Np=H⋅WN_p = H \cdot W for 2D images of height HH and width WW, or Np=H⋅W⋅TN_p = H \cdot W \cdot T for 3D video features). The output feature zi∈RC\mathbf{z}_i \in \mathbb{R}^C at position ii is defined as:

    zi=xi+Wv2 ReLU(LN(Wv1∑j=1Npexp⁡(Wkxj)∑m=1Npexp⁡(Wkxm)xj))\mathbf{z}_i = \mathbf{x}_i + W_{v2} \, \text{ReLU}\left(\text{LN}\left(W_{v1} \sum_{j=1}^{N_p} \frac{\exp(W_k \mathbf{x}_j)}{\sum_{m=1}^{N_p} \exp(W_k \mathbf{x}_m)} \mathbf{x}_j \right)\right)

    The block executes three sequential steps:

    1. Context Modeling (Global Attention Pooling): A 1×11\times 1 convolution with weight matrix Wk∈R1×CW_k \in \mathbb{R}^{1 \times C} maps each position to a scalar score, followed by spatial softmax normalization across all NpN_p positions to compute attention weights αj=exp⁡(Wkxj)∑m=1Npexp⁡(Wkxm)\alpha_j = \frac{\exp(W_k \mathbf{x}_j)}{\sum_{m=1}^{N_p} \exp(W_k \mathbf{x}_m)}. The global context vector is computed as the attention-weighted sum ∑j=1Npαjxj∈RC\sum_{j=1}^{N_p} \alpha_j \mathbf{x}_j \in \mathbb{R}^C.
    2. Feature Transform (Bottleneck with Layer Normalization): To reduce parameter complexity, the CC-dimensional context vector is projected down to a hidden bottleneck dimension C/rC/r via Wv1∈RCr×CW_{v1} \in \mathbb{R}^{\frac{C}{r} \times C} (where rr is the reduction ratio, default r=16r=16), passed through Layer Normalization (LN) and a ReLU non-linearity, and then projected back to CC channels via Wv2∈RC×CrW_{v2} \in \mathbb{R}^{C \times \frac{C}{r}}. Layer Normalization is placed before ReLU to ease optimization of the two-layer transform.
    3. Feature Fusion (Addition): The resulting channel-transformed global context vector is added element-wise (via broadcast addition) to the input feature vector xi\mathbf{x}_i at every position ii.
  2. Knowl 2 — Three-Step General Framework for Global Context Modeling

    model/method

    Global context modeling methods can be unified into an abstracted three-step formulation defined by:

    zi=F(xi, δ(∑j=1Npαjxj))\mathbf{z}_i = F\left(\mathbf{x}_i, \, \delta\left(\sum_{j=1}^{N_p} \alpha_j \mathbf{x}_j\right)\right)

    where xi∈RC\mathbf{x}_i \in \mathbb{R}^C is the feature vector at position ii, NpN_p is the total number of spatial/spatiotemporal positions, and zi∈RC\mathbf{z}_i \in \mathbb{R}^C is the output feature.

    The framework consists of:

    1. Context Modeling Module: Computes a pooled global context vector ∑j=1Npαjxj\sum_{j=1}^{N_p} \alpha_j \mathbf{x}_j as a weighted average over all positions jj with weighting coefficients αj\alpha_j.
    2. Feature Transform Module: A function δ(⋅)\delta(\cdot) that maps the aggregated context vector to capture channel-wise interdependencies.
    3. Fusion Module: A function F(xi,⋅)F(\mathbf{x}_i, \cdot) that merges the transformed global context feature into the local feature xi\mathbf{x}_i at each position ii.

    Different network blocks instantiate this general framework through specific choices:

    • Squeeze-and-Excitation (SE) Block: Sets αj=1Np\alpha_j = \frac{1}{N_p} (uniform global average pooling), instantiates δ(⋅)\delta(\cdot) as a bottleneck transform with a sigmoid activation (W2ReLU(W1(⋅))W_2 \text{ReLU}(W_1 (\cdot)) followed by sigmoid), and defines F(xi,y)=xi⊙yF(\mathbf{x}_i, \mathbf{y}) = \mathbf{x}_i \odot \mathbf{y} as channel-wise rescaling (multiplication).
    • Simplified Non-Local (SNL) Block: Sets αj=exp⁡(Wkxj)∑mexp⁡(Wkxm)\alpha_j = \frac{\exp(W_k \mathbf{x}_j)}{\sum_m \exp(W_k \mathbf{x}_m)} (global attention pooling), instantiates δ(⋅)\delta(\cdot) as a single linear 1×11\times 1 convolution WvW_v, and defines F(xi,y)=xi+yF(\mathbf{x}_i, \mathbf{y}) = \mathbf{x}_i + \mathbf{y} as broadcast element-wise addition.
    • Global Context (GC) Block: Combines global attention pooling for context modeling, a bottleneck transform with Layer Normalization and ReLU for feature transform, and broadcast element-wise addition for feature fusion.
  3. Knowl 3 — Simplified Non-Local (SNL) Block

    model/method

    The standard non-local block with embedded Gaussian computes query-specific pairwise attention maps between all query positions ii and key positions jj:

    zi=xi+Wz∑j=1Npexp⁡(⟨Wqxi,Wkxj⟩)∑m=1Npexp⁡(⟨Wqxi,Wkxm⟩)(Wvxj)\mathbf{z}_i = \mathbf{x}_i + W_z \sum_{j=1}^{N_p} \frac{\exp(\langle W_q \mathbf{x}_i, W_k \mathbf{x}_j \rangle)}{\sum_{m=1}^{N_p} \exp(\langle W_q \mathbf{x}_i, W_k \mathbf{x}_m \rangle)} (W_v \mathbf{x}_j)

    Because empirical analysis shows that trained non-local attention maps are virtually identical across all query positions ii, the query-dependent term WqxiW_q \mathbf{x}_i and the projection WzW_z can be removed without sacrificing accuracy, yielding a query-independent formulation:

    zi=xi+∑j=1Npexp⁡(Wkxj)∑m=1Npexp⁡(Wkxm)(Wvxj)\mathbf{z}_i = \mathbf{x}_i + \sum_{j=1}^{N_p} \frac{\exp(W_k \mathbf{x}_j)}{\sum_{m=1}^{N_p} \exp(W_k \mathbf{x}_m)} (W_v \mathbf{x}_j)

    By applying the distributive law of matrix multiplication over summation, the linear transformation WvW_v can be factored outside the spatial pooling operation:

    zi=xi+Wv∑j=1Npexp⁡(Wkxj)∑m=1Npexp⁡(Wkxm)xj\mathbf{z}_i = \mathbf{x}_i + W_v \sum_{j=1}^{N_p} \frac{\exp(W_k \mathbf{x}_j)}{\sum_{m=1}^{N_p} \exp(W_k \mathbf{x}_m)} \mathbf{x}_j

    This factorization shares the single aggregated global context feature across all query positions and reduces the computational complexity of the 1×11\times 1 convolution WvW_v from O(HWC2)\mathcal{O}(HWC^2) to O(C2)\mathcal{O}(C^2), where HH and WW denote spatial dimensions and CC is the channel dimension.

  4. Knowl 4 — Statistical Evidence of Query Independence in Non-Local Blocks

    data/table

    Empirical analysis across multiple instantiations of the non-local block (Gaussian, Embedded Gaussian, Dot Product, Concat) demonstrates that non-local blocks learn query-independent context maps despite their query-dependent formulation.

    The average distance between position vectors is measured as avg_dist=1Np2∑i=1Np∑j=1Npdist(vi,vj)\text{avg\_dist} = \frac{1}{N_p^2} \sum_{i=1}^{N_p} \sum_{j=1}^{N_p} \text{dist}(\mathbf{v}_i, \mathbf{v}_j), using cosine distance dist(vi,vj)=1−cos⁡(vi,vj)2\text{dist}(\mathbf{v}_i, \mathbf{v}_j) = \frac{1 - \cos(\mathbf{v}_i, \mathbf{v}_j)}{2} and Jensen-Shannon Divergence (JSD) for attention distributions:

    dist(ωi,ωj)=12∑k=1Np(ωiklog⁡2ωikωik+ωjk+ωjklog⁡2ωjkωik+ωjk)\text{dist}(\boldsymbol{\omega}_i, \boldsymbol{\omega}_j) = \frac{1}{2} \sum_{k=1}^{N_p} \left( \omega_{ik} \log \frac{2\omega_{ik}}{\omega_{ik} + \omega_{jk}} + \omega_{jk} \log \frac{2\omega_{jk}}{\omega_{ik} + \omega_{jk}} \right)

    Dataset Method APbbox\text{AP}^{\text{bbox}} APmask\text{AP}^{\text{mask}} Cosine Distance JSD-att
    input output att
    COCO Gaussian 38.0 34.8 0.397 0.062 0.177 0.065
    E-Gaussian 38.0 34.7 0.402 0.012 0.020 0.011
    Dot product 38.1 34.8 0.405 0.020 0.015 -
    Concat 38.0 34.9 0.393 0.003 0.004 -
    Dataset Method Top-1 Top-5 input output att JSD-att
    Kinetics Gaussian 76.0 92.3 0.345 0.056 0.056 0.021
    E-Gaussian 75.9 92.2 0.358 0.003 0.004 0.015
    Dot product 76.0 92.3 0.353 0.095 0.099 -
    Concat 75.4 92.2 0.354 0.048 0.049 -

    While the input features xi\mathbf{x}_i ('input') exhibit high cosine distance (~0.35 to 0.40), indicating clear spatial discriminability, the non-local block outputs before residual addition zi−xi\mathbf{z}_i - \mathbf{x}_i ('output') and the attention maps ωi\boldsymbol{\omega}_i ('att') have near-zero cosine distance and tiny JSD values. This confirms that the global context features generated by non-local layers are essentially invariant to query position.

  5. Knowl 5 — Comparison of Pooling and Fusion Strategies in Global Context Modeling

    data/table

    Ablation experiments evaluate the interaction between context pooling strategies (global average pooling avg vs. global attention pooling att) and feature fusion strategies (channel-wise multiplication/rescaling scale vs. broadcast addition add).

    Results on COCO 2017 validation (Mask R-CNN with ResNet-50 FPN) and ImageNet validation (ResNet-50 backbone):

    Method COCO 2017 Validation ImageNet Validation
    APbbox\text{AP}^{\text{bbox}} AP50bbox\text{AP}^{\text{bbox}}_{50} AP75bbox\text{AP}^{\text{bbox}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} AP75mask\text{AP}^{\text{mask}}_{75} Top-1 (%) Top-5 (%)
    Baseline 37.2 59.0 40.1 33.8 55.4 35.9 76.88 93.16
    avg+scale (SENet) 38.2 60.2 41.2 34.7 56.7 37.1 77.26 93.55
    avg+add 39.1 61.4 42.3 35.6 57.9 37.9 77.40 93.60
    att+scale 38.3 60.4 41.5 34.8 57.0 36.8 77.34 93.48
    att+add (GCNet) 39.4 61.6 42.4 35.7 58.4 37.6 77.70 93.66

    These results establish that:

    1. Broadcast addition (add) is substantially more effective than channel-wise scaling (scale) for incorporating global context into local representations across all tasks (e.g., avg+add outperforms avg+scale by 0.9% on APbbox\text{AP}^{\text{bbox}}).
    2. Attention pooling (att) consistently improves over uniform average pooling (avg), with the advantage being particularly pronounced on ImageNet image classification (+0.30% top-1 accuracy for att+add over avg+add).
  6. Knowl 6 — Bottleneck Transform and Layer Normalization Design

    data/table

    Ablations on COCO 2017 validation with Mask R-CNN ResNet-50 investigate the transform module δ(⋅)\delta(\cdot) structure and the reduction ratio rr, where the bottleneck reduces channel dimension from CC to C/rC/r.

    Transform Design APbbox\text{AP}^{\text{bbox}} AP50bbox\text{AP}^{\text{bbox}}_{50} AP75bbox\text{AP}^{\text{bbox}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} #param FLOPs
    Baseline 37.2 59.0 40.1 33.8 55.4 44.4M 279.4G
    w/o ratio (single 1×11\times 1) 39.4 61.8 42.8 35.9 58.6 64.4M 279.6G
    r16 (two linear layers) 38.8 61.0 42.3 35.3 57.6 46.9M 279.6G
    r16+ReLU 38.8 61.0 42.0 35.4 57.5 46.9M 279.6G
    r16+LN+ReLU (GC block) 39.4 61.6 42.4 35.7 58.4 46.9M 279.6G
    Reduction Ratio rr APbbox\text{AP}^{\text{bbox}} AP50bbox\text{AP}^{\text{bbox}}_{50} AP75bbox\text{AP}^{\text{bbox}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} #param FLOPs
    ratio 4 39.9 62.2 42.9 36.2 58.7 54.4M 279.6G
    ratio 8 39.5 62.1 42.5 35.9 58.1 49.4M 279.6G
    ratio 16 39.4 61.6 42.4 35.7 58.4 46.9M 279.6G
    ratio 32 39.1 61.6 42.4 35.7 58.1 45.7M 279.5G

    Two unnormalized linear layers (r16 or r16+ReLU) degrade performance due to optimization difficulty. Adding Layer Normalization (r16+LN+ReLU) overcomes optimization hurdles and matches the uncompressed 1×11\times 1 conv transform (w/o ratio) while saving 17.5M parameters. Decreasing the bottleneck ratio from r=32r=32 to r=4r=4 yields monotonic performance improvements from 39.1 to 39.9 APbbox\text{AP}^{\text{bbox}}.

  7. Knowl 7 — Layer-Wise and Stage-Wise Insertion of GC Blocks in Deep Backbones

    empirical result

    Because GC blocks are lightweight, they can be integrated into all residual blocks across multiple stages rather than restricted to a single block as in standard Non-Local Networks.

    Evaluating integration strategies on COCO 2017 validation using Mask R-CNN with ResNet-50 FPN demonstrates:

    1. Integration Point within Residual Block: Inserting the GC block after residual addition (afterAdd) yields 39.4% APbbox\text{AP}^{\text{bbox}} and 35.8% APmask\text{AP}^{\text{mask}}, while inserting it immediately after the final 1×11\times 1 convolution inside the residual branch (after1x1) yields 39.4% APbbox\text{AP}^{\text{bbox}} and 35.7% APmask\text{AP}^{\text{mask}}. The after1x1 location is chosen as the standard default.
    2. Stage Insertion: Adding GC blocks to individual ResNet stages improves the baseline (37.2% APbbox\text{AP}^{\text{bbox}}): stage c3 achieves 37.9%, stage c4 achieves 38.9%, and stage c5 achieves 38.7% APbbox\text{AP}^{\text{bbox}}. Higher semantic stages (c4, c5) benefit more from global context than lower stages (c3).
    3. Full Backbone Integration: Adding GC blocks across all residual blocks in all three stages (c3+c4+c5) achieves 39.4% APbbox\text{AP}^{\text{bbox}} and 35.7% APmask\text{AP}^{\text{mask}}, significantly outperforming single-stage insertions with only a nominal FLOP increase (279.4G to 279.6G, a +0.07% relative increase).
  8. Knowl 8 — Object Detection and Instance Segmentation Performance on COCO

    data/table

    GCNet evaluated with reduction ratios r=16r=16 and r=4r=4 across standard and advanced backbones on the COCO 2017 dataset using Mask R-CNN and Cascade R-CNN architectures.

    Backbone Method APbbox\text{AP}^{\text{bbox}} AP50bbox\text{AP}^{\text{bbox}}_{50} AP75bbox\text{AP}^{\text{bbox}}_{75} APmask\text{AP}^{\text{mask}} AP50mask\text{AP}^{\text{mask}}_{50} AP75mask\text{AP}^{\text{mask}}_{75} FLOPs
    COCO 2017 Validation Set
    ResNet-50 baseline 37.2 59.0 40.1 33.8 55.4 35.9 279.4G
    +GC r16 39.4 61.6 42.4 35.7 58.4 37.6 279.6G
    +GC r4 39.9 62.2 42.9 36.2 58.7 38.3 279.6G
    ResNet-101 baseline 39.8 61.3 42.9 36.0 57.9 38.3 354.0G
    +GC r16 41.1 63.6 45.0 37.4 60.1 39.6 354.3G
    +GC r4 41.7 63.7 45.5 37.6 60.5 39.8 354.3G
    ResNeXt-101 baseline 41.2 63.0 45.1 37.3 59.7 39.9 357.9G
    +GC r16 42.4 64.6 46.5 38.0 60.9 40.5 358.2G
    +GC r4 42.9 65.2 47.0 38.5 61.8 40.9 358.2G
    ResNeXt-101+Cascade baseline 44.7 63.0 48.5 38.3 59.9 41.3 536.9G
    +GC r16 45.9 64.8 50.0 39.3 61.8 42.1 537.2G
    +GC r4 46.5 65.4 50.7 39.7 62.5 42.7 537.3G
    ResNeXt-101+DCN+Cascade baseline 47.1 66.1 51.3 40.4 63.1 43.7 547.5G
    +GC r16 47.9 66.9 52.2 40.9 63.7 44.1 547.8G
    +GC r4 47.9 66.9 51.9 40.8 64.0 44.0 547.8G
    COCO 2017 Test-Dev Set
    ResNeXt-101+Cascade baseline 45.0 63.7 49.1 38.7 60.8 41.8 536.9G
    +GC r16 46.5 65.7 50.7 40.0 62.9 43.1 537.2G
    +GC r4 46.6 65.9 50.7 40.1 62.9 43.3 537.3G
    ResNeXt-101+DCN+Cascade baseline 47.7 66.7 52.0 41.0 63.9 44.3 547.5G
    +GC r16 48.3 67.5 52.7 41.5 64.6 45.0 547.8G
    +GC r4 48.4 67.6 52.7 41.5 64.6 45.0 547.8G

    The additions demonstrate that GC blocks provide consistent improvements across all network backbones, adding up to +2.7% APbbox\text{AP}^{\text{bbox}} and +2.4% APmask\text{AP}^{\text{mask}} on ResNet-50 and boosting the strongest baseline (ResNeXt-101 + DCN + Cascade R-CNN) by +0.8% APbbox\text{AP}^{\text{bbox}} and +0.5% APmask\text{AP}^{\text{mask}} with negligible computational overhead.

  9. Knowl 9 — ImageNet Image Classification Performance of GCNet

    data/table

    ImageNet 1K validation set performance for ResNet-50 models augmented with different global context blocks. Models are trained with standard two-stage training (120 epochs baseline training followed by 40 epochs finetuning with cosine decay):

    Method Top-1 Acc (%) Top-5 Acc (%) #params FLOPs
    Baseline (ResNet-50) 76.88 93.16 25.56M 3.86G
    +1 NL 77.20 93.51 27.66M 4.11G
    +1 SNL 77.28 93.60 26.61M 3.86G
    +1 GC 77.34 93.52 25.69M 3.86G
    +all GC (GC-ResNet-50) 77.70 93.66 28.08M 3.87G

    A single GC block (+1 GC) outperforms a single Non-Local block (+1 NL) by +0.14% top-1 accuracy while requiring 1.97M fewer parameters and 0.25G fewer FLOPs. Adding GC blocks to all residual blocks (+all GC) achieves 77.70% top-1 accuracy (+0.82% over baseline) with a relative FLOP increase of only 0.26%.

  10. Knowl 10 — Action Recognition Performance on Kinetics Validation Set

    data/table

    Evaluation of human action recognition on the Kinetics-400 dataset using a Slow-only ResNet-50 backbone with inflated 3D weights, trained on 8-frame clips and evaluated with 30 clips per video (reduction ratio r=4r=4 for GC blocks):

    Method Top-1 Acc (%) Top-5 Acc (%) #params FLOPs
    Baseline (Slow-only R50) 74.94 91.90 32.45M 39.29G
    +5 NL 75.95 92.29 39.81M 59.60G
    +5 SNL 75.76 92.44 36.13M 39.32G
    +5 GC 75.85 92.25 34.30M 39.31G
    +all GC (GC-ResNet-50) 76.00 92.34 42.45M 39.35G

    Inserting 5 GC blocks achieves comparable accuracy to inserting 5 standard Non-Local blocks (75.85% vs. 75.95% top-1, 92.25% vs. 92.29% top-5) while reducing FLOPs from 59.60G down to 39.31G (a 34% reduction in total model computation). Integrating GC blocks into all residual blocks (+all GC) yields the highest accuracy (76.00% top-1, +1.06% over baseline) at only 39.35 GFLOPs.

Coverage note — None was omitted; all key contributions including the theoretical framework, mathematical derivations of block simplifications, empirical analyses of non-local behaviors, and experimental results across COCO, ImageNet, and Kinetics are fully captured.

References

  1. 1.Z. Cai and N. Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6154–6162, 2018. 7
  2. 2.J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017. 2, 8
  3. 3.F. Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017. 2
  4. 4.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 764–773, 2017. 2, 7
  5. 5.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 6, 8
  6. 6.C. Feichtenhofer, H. Fan, J. Malik, and K. He. Slowfast networks for video recognition. arXiv preprint arXiv:1812.03982, 2018. 2, 8
  7. 7.J. Gehring, M. Auli, D. Grangier, and Y. N. Dauphin. A convolutional encoder model for neural machine translation. arXiv preprint arXiv:1611.02344, 2016. 2
  8. 8.J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin. Convolutional sequence to sequence learning. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 1243–1252. JMLR. org, 2017. 2
  9. 9.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 2, 6
  10. 10.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015. 2, 8
  11. 11.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
  12. 12.H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei. Relation networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3588–3597, 2018. 1, 2, 4
  13. 13.J. Hu, L. Shen, S. Albanie, G. Sun, and A. Vedaldi. Gather-excite: Exploiting feature context in convolutional neural networks. In Advances in Neural Information Processing Systems, pages 9423–9433, 2018. 2
  14. 14.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2018. 1, 2, 5, 8
  15. 15.G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017. 2
  16. 16.Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu. Ccnet: Criss-cross attention for semantic segmentation. arXiv preprint arXiv:1811.11721, 2018. 2
  17. 17.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. 6, 8
  18. 18.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2
  19. 19.Z. Li, C. Peng, G. Yu, X. Zhang, Y. Deng, and J. Sun. Detnet: A backbone network for object detection. arXiv preprint arXiv:1804.06215, 2018. 2
  20. 20.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature pyramid networks for object detection. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. 6
  21. 21.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 6
  22. 22.F. Massa and R. Girshick. maskrcnn-benchmark: Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch. https://github.com/facebookresearch/maskrcnn-benchmark, 2018. Accessed: [2019.03.22]. 6
  23. 23.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017. 6
  24. 24.Z. Qiu, T. Yao, and T. Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, pages 5533–5541, 2017. 2
  25. 25.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, pages 91–99, 2015. 6
  26. 26.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 2
  27. 27.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015. 2
  28. 28.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017. 1, 2, 4
  29. 29.P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017. 2
  30. 30.F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang. Residual attention network for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3156–3164, 2017. 2
  31. 31.X. Wang, R. Girshick, A. Gupta, and K. He. Non-local neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, 2018. 1, 2, 3, 4, 7, 8
  32. 32.S. Woo, J. Park, J.-Y. Lee, and I. So Kweon. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), pages 3–19, 2018. 2
  33. 33.J. Xie, T. He, Z. Zhang, H. Zhang, Z. Zhang, and M. Li. Bag of tricks for image classification with convolutional neural networks. arXiv preprint arXiv:1812.01187, 2018. 8
  34. 34.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1492–1500, 2017. 2, 7
  35. 35.S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy. Rethinking spatiotemporal feature learning: Speedaccuracy trade-offs in video classification. In Proceedings of the European Conference on Computer Vision (ECCV), pages 305–321, 2018. 2
  36. 36.Y. Yuan and J. Wang. Ocnet: Object context network for scene parsing. arXiv preprint arXiv:1809.00916, 2018. 2
  37. 37.S. Zagoruyko and N. Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016. 2
  38. 38.H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal. Context encoding for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7151–7160, 2018. 1
  39. 39.H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena. Self-attention generative adversarial networks. arXiv preprint arXiv:1805.08318, 2018. 2
  40. 40.X. Zhang, X. Zhou, M. Lin, and J. Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6848–6856, 2018. 2
  41. 41.H. Zhao, Y. Zhang, S. Liu, J. Shi, C. Change Loy, D. Lin, and J. Jia. Psanet: Point-wise spatial attention network for scene parsing. In Proceedings of the European Conference on Computer Vision (ECCV), pages 267–283, 2018. 2
  42. 42.X. Zhu, H. Hu, S. Lin, and J. Dai. Deformable convnets v2: More deformable, better results. arXiv preprint arXiv:1811.11168, 2018. 2, 7
  43. 43.B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le. Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8697–8710, 2018. 2

Citation

MLA
Cao, Y., et al. “GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond”. arXiv, 2019, http://arxiv.org/abs/1904.11492v1.
APA
Cao, Y., Xu, J., Lin, S., Wei, F., & Hu, H. (2019). GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond. arXiv. http://arxiv.org/abs/1904.11492v1
Chicago
Cao, Y., J. Xu, S. Lin, F. Wei, and H. Hu. 2019. “GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond”. arXiv. http://arxiv.org/abs/1904.11492v1.
Harvard
Cao, Y. et al. (2019) “GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.11492v1.
Vancouver
1. Cao Y, Xu J, Lin S, Wei F, Hu H (2019) GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond. arXiv

BibTeX

@article{cao2019gcnet,
  title = {GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond},
  author = {Cao, Yue and Xu, Jiarui and Lin, Stephen and Wei, Fangyun and Hu, Han},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.11492v1},
  eprint = {1904.11492}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE