Selective Kernel Networks

Xiang LiWenhai WangXiaolin HuJian Yang

article2019CVPR2,795 citations

Proposes Selective Kernel Networks, an architecture that uses attention-guided branch fusion to dynamically adjust neuron receptive field sizes according to input scale, outperforming standard convolutional networks with lower model complexity.

Listen

The article addresses the limitation in standard convolutional neural networks where neurons in each layer use fixed receptive field sizes, even though neuroscience shows that visual cortical neurons dynamically adjust these sizes based on stimulus properties such as contrast and scale. This fixed design restricts the ability of models to handle objects at varying scales efficiently, which is increasingly important for accurate image recognition in real-world applications.

The article set out to develop and test a mechanism that lets neurons adaptively select receptive field sizes during processing by combining information from multiple kernel sizes in a nonlinear way.

Researchers introduced the Selective Kernel convolution, built from split, fuse, and select operations. Multiple branches process the input with different kernel sizes, global information is aggregated to guide selection, and softmax attention weights the branches to form the output. They stacked these units into SKNets based on a ResNeXt backbone and evaluated them on ImageNet classification with over one million images, as well as on smaller CIFAR datasets. Additional tests embedded the units into lightweight models, and controlled experiments scaled target objects in validation images to observe attention shifts.

SKNet-50 reached 20.79 percent top-1 error on ImageNet, an improvement of 1.44 points over the ResNeXt-50 baseline and 0.33 points over SENet-50 at similar parameter counts and computation. Larger models such as SKNet-101 also outperformed prior attention-based and multi-scale networks. Attention analysis revealed that neurons assigned higher weights to larger kernels as object scale increased, with this adaptive behavior clearest in lower and middle layers across all 1,000 ImageNet categories. The approach also delivered consistent gains when added to compact architectures.

These results indicate that adaptive kernel selection improves recognition accuracy without substantial added cost and produces behavior closer to biological vision. The gains matter for deployment where both precision and efficiency matter, such as mobile or resource-constrained settings, and suggest that similar dynamic mechanisms could reduce the need for ever-larger fixed models.

The work points to further exploration of adaptive architectures in other vision tasks and automated network design. Future studies would benefit from testing on detection, segmentation, and video data, as well as from direct comparisons on hardware efficiency.

The main limitations are the focus on image classification benchmarks and the reliance on a single family of backbone networks; results on broader tasks or entirely new architectures remain untested. Confidence is high for the reported classification improvements and the observed adaptation pattern, yet caution is warranted when extrapolating beyond the evaluated conditions.

Cover for Selective Kernel Networks

Abstract

In standard Convolutional Neural Networks (CNNs), the receptive fields of artificial neurons in each layer are designed to share the same size. It is well-known in the neuroscience community that the receptive field size of visual cortical neurons are modulated by the stimulus, which has been rarely considered in constructing CNNs. We propose a dynamic selection mechanism in CNNs that allows each neuron to adaptively adjust its receptive field size based on multiple scales of input information. A building block called Selective Kernel (SK) unit is designed, in which multiple branches with different kernel sizes are fused using softmax attention that is guided by the information in these branches. Different attentions on these branches yield different sizes of the effective receptive fields of neurons in the fusion layer. Multiple SK units are stacked to a deep network termed Selective Kernel Networks (SKNets). On the ImageNet and CIFAR benchmarks, we empirically show that SKNet outperforms the existing state-of-the-art architectures with lower model complexity. Detailed analyses show that the neurons in SKNet can capture target objects with different scales, which verifies the capability of neurons for adaptively adjusting their receptive field sizes according to the input. The code and models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Selective Kernel Convolution
  • 3.2 Network Architecture
  • 4 Experiments
  • 4.1 ImageNet Classification
  • 4.2 CIFAR Classification
  • 4.3 Ablation Studies
  • 4.4 Analysis and Interpretation
  • 5 Conclusion
  • References
  • A Details of the Compared Models in Table 3
  • B Details of the Models in Table 4
  • C Details of the Compared Models in Figure 2
  • D Implementation Details on CIFAR Datasets (Section 4.2)
  • E More Examples of Dynamic Selection

Knowls

  1. Knowl 1 — Selective Kernel Convolution Mechanism

    model/method

    Selective Kernel (SK) convolution is a dynamic multi-scale feature extraction operator that adaptively adjusts the effective receptive field size of artificial neurons using a soft-attention gating mechanism. For an input feature map X∈RH′×W′×C′X \in \mathbb{R}^{H' \times W' \times C'}, the operator is structured into three consecutive phases: Split, Fuse, and Select.

    1. Split: Multiple parallel branches perform transformations with different kernel sizes. In a two-branch configuration (M=2M = 2), branch transformations F~:X→U~∈RH×W×C\tilde{\mathcal{F}}: X \to \tilde{U} \in \mathbb{R}^{H \times W \times C} and F^:X→U^∈RH×W×C\hat{\mathcal{F}}: X \to \hat{U} \in \mathbb{R}^{H \times W \times C} are applied, typically with kernel sizes 3×33\times 3 and 5×55\times 5, respectively. Each branch consists of grouped or depthwise convolutions followed by Batch Normalization and ReLU. To reduce computational complexity, a 5×55\times 5 convolution is realized using a 3×33\times 3 convolution with dilation rate D=2D = 2.

    2. Fuse: Information from all branches is first aggregated via element-wise summation:

    U=U~+U^U = \tilde{U} + \hat{U}

    Global spatial context is extracted via channel-wise global average pooling to obtain channel statistics s∈RCs \in \mathbb{R}^C, where the cc-th element is computed as:

    sc=1H×W∑i=1H∑j=1WUc(i,j)s_c = \frac{1}{H \times W} \sum_{i=1}^H \sum_{j=1}^W U_c(i, j)

    A compact feature vector z∈Rd×1z \in \mathbb{R}^{d \times 1} is generated using a fully connected layer with dimensionality reduction:

    z=δ(B(Ws))z = \delta(\mathcal{B}(W s))

    where δ\delta is the ReLU activation, B\mathcal{B} denotes Batch Normalization, W∈Rd×CW \in \mathbb{R}^{d \times C}, and the reduced dimensionality dd is determined by a reduction ratio rr and a lower bound threshold LL:

    d=max⁡(C/r,L)d = \max(C/r, L)

    1. Select: Channel-wise soft attention weights are computed across the branches via a softmax function conditioned on zz:

    ac=eAczeAcz+eBcz,bc=eBczeAcz+eBcza_c = \frac{e^{A_c z}}{e^{A_c z} + e^{B_c z}}, \quad b_c = \frac{e^{B_c z}}{e^{A_c z} + e^{B_c z}}

    where A,B∈RC×dA, B \in \mathbb{R}^{C \times d}, and Ac,Bc∈R1×dA_c, B_c \in \mathbb{R}^{1 \times d} are the cc-th rows of AA and BB, satisfying ac+bc=1a_c + b_c = 1. The final output feature map V=[V1,V2,…,VC]∈RH×W×CV = [V_1, V_2, \dots, V_C] \in \mathbb{R}^{H \times W \times C} is obtained by weighted summation of branch representations per channel:

    Vc=ac⋅U~c+bc⋅U^cV_c = a_c \cdot \tilde{U}_c + b_c \cdot \hat{U}_c

  2. Knowl 2 — Selective Kernel Network Architecture and Parameterization

    model/method

    Selective Kernel Networks (SKNets) adapt the ResNeXt backbone by replacing the standard 3×33\times 3 grouped convolutions in bottleneck blocks with Selective Kernel (SK) convolution units. An SK unit follows a 1×1 conv→SK conv→1×1 conv1\times 1 \text{ conv} \to \text{SK conv} \to 1\times 1 \text{ conv} structure with residual connections.

    An SK convolution block is parameterized by SK[M,G,r]\text{SK}[M, G, r]:

    • MM: the number of parallel convolutional kernel branches (default M=2M = 2, corresponding to 3×33\times 3 and dilated 3×33\times 3 with D=2D=2).
    • GG: the grouping number/cardinality of the convolutions in each branch (default G=32G = 32).
    • rr: the channel reduction ratio for computing the compact gating descriptor zz (default r=16r = 16, with minimum dimension L=32L = 32).

    The standard network variants across four feature stages are:

    • SKNet-26: Stacks {2,2,2,2}\{2, 2, 2, 2\} SK units across stages.
    • SKNet-50: Stacks {3,4,6,3}\{3, 4, 6, 3\} SK units across stages, yielding 27.5M parameters and 4.47 GFLOPs (compared to ResNeXt-50 with 25.0M parameters and 4.24 GFLOPs).
    • SKNet-101: Stacks {3,4,23,3}\{3, 4, 23, 3\} SK units across stages, yielding 48.9M parameters and 8.46 GFLOPs.
  3. Knowl 3 — ImageNet Classification Performance of SKNets

    empirical result

    On the ImageNet 2012 classification validation set (trained on 1.28M images, tested with single 224×224224\times 224 and 320×320320\times 320 crops), SKNets achieve lower top-1 error rates compared to contemporary baseline architectures with similar or greater computational budgets.

    Model Top-1 Err (224x224, %) Top-1 Err (320x320, %) Parameters (M) GFLOPs
    ResNeXt-50 22.23 21.05 25.0 4.24
    AttentionNeXt-56 21.76 – 31.9 6.32
    InceptionV3 – 21.20 27.1 5.73
    ResNeXt-50 + BAM 21.70 20.15 25.4 4.31
    ResNeXt-50 + CBAM 21.40 20.38 27.7 4.25
    SENet-50 21.12 19.71 27.7 4.25
    SKNet-50 (ours) 20.79 19.32 27.5 4.47
    ResNeXt-101 21.11 19.86 44.3 7.99
    Attention-92 – 19.50 51.3 10.43
    DPN-92 20.70 19.30 37.7 6.50
    DPN-98 20.20 18.90 61.6 11.70
    InceptionV4 – 20.00 42.0 12.31
    Inception-ResNetV2 – 19.90 55.0 13.22
    ResNeXt-101 + BAM 20.67 19.15 44.6 8.05
    ResNeXt-101 + CBAM 20.60 19.42 49.2 8.00
    SENet-101 20.58 18.61 49.2 8.00
    SKNet-101 (ours) 20.19 18.40 48.9 8.46

    SKNet-50 outperforms ResNeXt-101 (20.79% vs 21.11% top-1 error at 224×224224\times 224), despite ResNeXt-101 having 60% more parameters and 80% higher FLOPs. SKNet-50 and SKNet-101 also outperform SENet counterparts by 0.33% and 0.39% absolute top-1 error at 224×224224\times 224 evaluation.

  4. Knowl 4 — Efficiency of Selective Kernel Convolutions Relative to Depth, Width, and Cardinality Scaling

    empirical result

    When baseline ResNeXt-50 (32×4d32\times 4\text{d}) is expanded in depth, width, or cardinality to match the parameter and FLOP complexity of SKNet-50, the performance gains are substantially smaller than those obtained from adding dynamic Selective Kernel selection.

    Model Configuration Top-1 Err (%) Parameters (M) GFLOPs
    ResNeXt-50 (32×4d32\times 4\text{d}) 22.23 25.0 4.24
    ResNeXt-50, wider (+1/16+1/16 channels) 22.13 28.1 4.74
    ResNeXt-56, deeper (+2+2 blocks in stage 4) 22.04 27.3 4.67
    ResNeXt-50 (36×4d36\times 4\text{d}, cardinality 36) 22.00 27.6 4.70
    SKNet-50 (M=2,G=32,r=16M=2, G=32, r=16) 20.79 27.5 4.47

    Increasing ResNeXt-50 complexity via width, depth, or cardinality yields absolute top-1 improvements of 0.10%, 0.19%, and 0.23%, respectively. In contrast, SKNet-50 achieves a 1.44% absolute improvement over ResNeXt-50 under equivalent computational constraints.

  5. Knowl 5 — Ablation of Branch Kernel Configurations, Dilation, and Attention Selection

    empirical result

    Ablation studies on ImageNet validation with SKNet-50 investigate the effects of branch kernel size, dilation rates, branch quantity (MM), and dynamic attention selection versus unweighted summation:

    1. Dilation vs. Explicit Kernel Size: Fixing Branch 1 to 3×33\times 3 (D=1,G=32D=1, G=32):

      • Branch 2 as dilated 3×33\times 3 (D=2D=2, effective receptive field 5×55\times 5, G=32G=32) achieves 20.79% top-1 error with 27.5M parameters and 4.47 GFLOPs.
      • Branch 2 as standard 5×55\times 5 (D=1,G=64D=1, G=64) achieves 20.80% top-1 error with 28.1M parameters and 4.56 GFLOPs.
      • Using dilated 3×33\times 3 kernels provides equivalent receptive field benefits with lower parameter and FLOP costs.
    2. Branch Count and Selection Mechanism:

    K3 (3×33\times 3) K5 (3×3,D=23\times 3, D=2) K7 (3×3,D=33\times 3, D=3) SK Attention Top-1 Err (%) Parameters (M) GFLOPs
    ✓ 22.23 25.0 4.24
    ✓ 25.14 25.0 4.24
    ✓ 25.51 25.0 4.24
    ✓ ✓ 21.76 26.5 4.46
    ✓ ✓ ✓ 20.79 27.5 4.47
    ✓ ✓ 21.82 26.5 4.46
    ✓ ✓ ✓ 20.97 27.5 4.47
    ✓ ✓ 23.64 26.5 4.46
    ✓ ✓ ✓ 23.09 27.5 4.47
    ✓ ✓ ✓ 21.47 28.0 4.69
    ✓ ✓ ✓ ✓ 20.76 29.3 4.70

    Key takeaways:

    • Multi-branch architectures consistently outperform single-branch configurations (M>1M > 1 vs M=1M = 1).
    • SK soft attention aggregation consistently outperforms naive element-wise addition across all kernel combinations (e.g., 20.79% vs 21.76% for K3+K5).
    • Increasing branches from M=2M=2 to M=3M=3 yields only a minor performance gain (20.79% to 20.76%), making M=2M=2 the preferred efficiency/performance operating point.
  6. Knowl 6 — Integration of SK Convolutions into Lightweight Architectures

    empirical result

    Embedding Selective Kernel convolutions into compact lightweight architectures such as ShuffleNetV2 demonstrates broad generalization. In ShuffleNetV2 (0.5×0.5\times and 1.0×1.0\times configurations), each 3×33\times 3 depthwise convolution is replaced by an SK unit with M=2M=2 (K3 and K5 dilated paths), reduction ratio r=4r=4, and grouping GG equal to the stage channel count.

    Model Top-1 Err (%) MFLOPs Parameters (M)
    ShuffleNetV2 0.5×0.5\times (re-impl.) 38.41 40.39 1.40
    ShuffleNetV2 0.5×0.5\times + SE 36.34 40.85 1.56
    ShuffleNetV2 0.5×0.5\times + SK 35.35 42.58 1.48
    ShuffleNetV2 1.0×1.0\times (re-impl.) 30.57 140.35 2.45
    ShuffleNetV2 1.0×1.0\times + SE 29.47 141.73 2.66
    ShuffleNetV2 1.0×1.0\times + SK 28.36 145.66 2.63

    For depthwise SK units in ShuffleNetV2 1.0×1.0\times, omitting ReLU activation functions after both depthwise convolution paths achieves optimal top-1 error (28.36%), compared to placing ReLU after K3 only (28.65%), after K5 only (28.40%), or after both paths (28.49%).

  7. Knowl 7 — Scale-Adaptive Receptive Field Dynamics Across Depth and Categories

    empirical result

    When ImageNet validation images are modified by scaling the central target object across sizes 1.0×1.0\times, 1.5×1.5\times, and 2.0×2.0\times (via center cropping and resizing back to the original image dimensions), the attention allocation in SKNet-50 exhibits specific receptive field adaptation behaviors:

    1. Channel and Unit Level Adaptation: In low- and middle-level stages (e.g., SK units in stages 2 and 3 such as SK_2_3 and SK_3_4), enlarging the target object monotonically increases the soft attention weight bcb_c assigned to the larger kernel branch (5×55\times 5) across the vast majority of channels.
    2. Consistency Across Categories: The positive correlation between target object scale and attention difference (mean attention of kernel 5×55\times 5 minus kernel 3×33\times 3) holds universally across all 1,000 ImageNet semantic classes in early and middle stages.
    3. Loss of Scale Differentiation in Deepest Layers: In the highest network stages (e.g., SK_5_3), the scale-dependent attention shift largely disappears. This indicates that scale information is integrated and encoded directly in high-level feature vectors, making physical convolutional receptive field size modulation less influential in top layers.
  8. Knowl 8 — CIFAR-10 and CIFAR-100 Classification with SKNet-29

    empirical result

    On CIFAR-10 and CIFAR-100 datasets (consisting of 32×3232\times 32 images), SKNet is evaluated using an SKNet-29 backbone derived from ResNeXt-29 (16×32d16\times 32\text{d}). To mitigate overfitting on small-resolution inputs, the second branch in the SK unit is set to a 1×11\times 1 convolution rather than 5×55\times 5, while keeping the first branch at 3×33\times 3, configured with SK[M=2,G=16,r=32]\text{SK}[M=2, G=16, r=32].

    Model Parameters (M) CIFAR-10 Top-1 Err (%) CIFAR-100 Top-1 Err (%)
    ResNeXt-29 (16×32d16\times 32\text{d}) 25.2 3.87 18.56
    ResNeXt-29 (8×64d8\times 64\text{d}) 34.4 3.65 17.77
    ResNeXt-29 (16×64d16\times 64\text{d}) 68.1 3.58 17.31
    SENet-29 35.0 3.68 17.78
    SKNet-29 (ours) 27.7 3.47 17.33

    SKNet-29 achieves 3.47% on CIFAR-10 and 17.33% on CIFAR-100, outperforming SENet-29 and matching the performance of ResNeXt-29 (16×64d16\times 64\text{d}) while using 60% fewer parameters.

Coverage note — None was omitted; all primary contributions, mathematical formulations, network structures, empirical benchmark results, ablation experiments, and analytical findings are covered.

References

  1. 1.M. Abdi and S. Nahavandi. Multi-residual networks. arxiv preprint. arXiv preprint arXiv:1609.05672, 2016.
  2. 2.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014.
  3. 3.J. Carreira, H. Madeira, and J. G. Silva. Xception: A technique for the experimental evaluation of dependability in modern computers. Transactions on Software Engineering, 1998.
  4. 4.D. Chen, S. Zhang, W. Ouyang, J. Yang, and Y. Tai. Person search via a mask-guided two-stream cnn model. arXiv preprint arXiv:1807.08107, 2018.
  5. 5.Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng. Dual path networks. In NIPS, 2017.
  6. 6.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. arXiv preprint arXiv:1703.06211, 2017.
  7. 7.X. Gastaldi. Shake-shake regularization. arXiv preprint arXiv:1705.07485, 2017.
  8. 8.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015.
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  10. 10.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016.
  11. 11.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  12. 12.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. arXiv preprint arXiv:1709.01507, 2017.
  13. 13.G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger. Densely connected convolutional networks. In CVPR, 2017.
  14. 14.D. H. Hubel and T. N. Wiesel. Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology, 1962.
  15. 15.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  16. 16.L. Itti and C. Koch. Computational modelling of visual attention. Nature Reviews Neuroscience, 2001.
  17. 17.L. Itti, C. Koch, and E. Niebur. A model of saliency-based visual attention for rapid scene analysis. TPAMI, 1998.
  18. 18.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In NIPS, 2015.
  19. 19.Y. Jeon and J. Kim. Active convolution: Learning the shape of convolution for image classification. In CVPR, 2017.
  20. 20.X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool. Dynamic filter networks. In NIPS, 2016.
  21. 21.P. Kontschieder, M. Fiterau, A. Criminisi, and S. Rota Bulo. Deep neural decision forests. In ICCV, 2015.
  22. 22.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. 2009.
  23. 23.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  24. 24.H. Larochelle and G. E. Hinton. Learning to combine foveal glimpses with a third-order boltzmann machine. In NIPS, 2010.
  25. 25.G. Larsson, M. Maire, and G. Shakhnarovich. Fractalnet: Ultra-deep neural networks without residuals. arXiv preprint arXiv:1605.07648, 2016.
  26. 26.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1989.
  27. 27.N. Ma, X. Zhang, H.-T. Zheng, and J. Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. arXiv preprint arXiv:1807.11164, 2018.
  28. 28.V. Mnih, N. Heess, A. Graves, et al. Recurrent models of visual attention. In NIPS, 2014.
  29. 29.V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  30. 30.J. Nelson and B. Frost. Orientation-selective inhibition from beyond the classic visual receptive field. Brain Research, 1978.
  31. 31.B. A. Olshausen, C. H. Anderson, and D. C. Van Essen. A neurobiological model of visual attention and invariant pattern recognition based on dynamic routing of information. Journal of Neuroscience, 1993.
  32. 32.J. Park, S. Woo, J.-Y. Lee, and I. S. Kweon. Bam: Bottleneck attention module. arXiv preprint arXiv:1807.06514, 2018.
  33. 33.M. W. Pettet and C. D. Gilbert. Dynamic changes in receptive-field size in cat primary visual cortex. Proceedings of the National Academy of Sciences, 1992.
  34. 34.A. M. Rush, S. Chopra, and J. Weston. A neural attention model for abstractive sentence summarization. arXiv preprint arXiv:1509.00685, 2015.
  35. 35.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015.
  36. 36.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018.
  37. 37.M. P. Sceniak, D. L. Ringach, M. J. Hawken, and R. Shapley. Contrast's effect on spatial summation by macaque v1 neurons. Nature Neuroscience, 1999.
  38. 38.L. Spillmann, B. Dresp-Langley, and C.-h. Tseng. Beyond the classical receptive field: the effect of contextual stimuli. Journal of Vision, 2015.
  39. 39.R. K. Srivastava, K. Greff, and J. Schmidhuber. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  40. 40.K. Sun, M. Li, D. Liu, and J. Wang. Igcv3: Interleaved low-rank group convolutions for efficient deep neural networks. arXiv preprint arXiv:1806.00178, 2018.
  41. 41.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, 2017.
  42. 42.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  43. 43.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In CVPR, 2016.
  44. 44.F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang. Residual attention network for image classification. arXiv preprint arXiv:1704.06904, 2017.
  45. 45.S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon. Cbam: Convolutional block attention module. arXiv preprint arXiv:1807.06521, 2018.
  46. 46.G. Xie, J. Wang, T. Zhang, J. Lai, R. Hong, and G.-J. Qi. Igcv 2: Interleaved structured sparse convolutional neural networks. arXiv preprint arXiv:1804.06202, 2018.
  47. 47.S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In CVPR, 2017.
  48. 48.K. Xu, D. Li, N. Cassimatis, and X. Wang. Lcanet: End-to-end lipreading with cascaded attention-ctc. In International Conference on Automatic Face & Gesture Recognition, 2018.
  49. 49.Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo. Image captioning with semantic attention. In CVPR, 2016.
  50. 50.F. Yu and V. Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  51. 51.F. Yu, V. Koltun, and T. A. Funkhouser. Dilated residual networks. In CVPR, 2017.
  52. 52.K. Zhang, M. Sun, X. Han, X. Yuan, L. Guo, and T. Liu. Residual networks of residual networks: Multilevel residual networks. Transactions on Circuits and Systems for Video Technology, 2017.
  53. 53.T. Zhang, G.-J. Qi, B. Xiao, and J. Wang. Interleaved group convolutions. In CVPR, 2017.
  54. 54.X. Zhang, X. Zhou, M. Lin, and J. Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. arxiv 2017. arXiv preprint arXiv:1707.01083, 2017.
  55. 55.Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, and Y. Fu. Image super-resolution using very deep residual channel attention networks. arXiv preprint arXiv:1807.02758, 2018.

Citation

MLA
Li, X., et al. “Selective Kernel Networks”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 510–19, https://doi.org/10.1109/CVPR.2019.00060.
APA
Li, X., Wang, W., Hu, X., & Yang, J. (2019). Selective Kernel Networks. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 510–519. https://doi.org/10.1109/CVPR.2019.00060
Chicago
Li, X., W. Wang, X. Hu, and J. Yang. 2019. “Selective Kernel Networks”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 510–19. https://doi.org/10.1109/CVPR.2019.00060.
Harvard
Li, X. et al. (2019) “Selective Kernel Networks”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 510–519. Available at: https://doi.org/10.1109/CVPR.2019.00060.
Vancouver
1. Li X, Wang W, Hu X, Yang J (2019) Selective Kernel Networks. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 510–519

BibTeX

@inproceedings{Li_2019, title={Selective Kernel Networks}, url={http://dx.doi.org/10.1109/CVPR.2019.00060}, DOI={10.1109/cvpr.2019.00060}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Li, Xiang and Wang, Wenhai and Hu, Xiaolin and Yang, Jian}, year={2019}, month=June, pages={510–519} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE