BAM: Bottleneck Attention Module

Jongchan ParkSanghyun WooJoon-Young LeeIn-So Kweon

article2018BMVC1,340 citationsSilver Prize, Best Paper Award

Introduces a lightweight dual-pathway attention module placed at convolutional bottlenecks to improve image classification and object detection performance across diverse architectures with minimal computational overhead.

Listen

Modern artificial intelligence systems rely heavily on deep neural networks for visual recognition tasks, but traditional methods of improving accuracy—such as stacking more layers or widening networks—dramatically increase computational cost and memory requirements. As visual applications expand to resource-constrained environments like mobile and embedded systems, organizations face the operational challenge of improving model accuracy without incurring high latency and hardware expenses.

The article introduces and evaluates the Bottleneck Attention Module, a lightweight component designed to improve the accuracy of convolutional neural networks. The objective is to demonstrate that placing this module at key transition points in a network enhances performance across multiple visual recognition tasks with minimal computational overhead.

The researchers evaluated the module using standard benchmark datasets, including CIFAR-100 and ImageNet-1K for image classification, as well as VOC 2007 and MS COCO for object detection. The method computes attention through two separate, complementary streams: a channel branch that determines which features are important, and a spatial branch using dilated convolutions to determine where in an image to focus. These two streams are combined through element-wise addition and integrated specifically at the bottleneck locations where networks downsample image data.

The evaluation produced several key findings. First, integrating the module consistently reduced classification error across various baseline architectures, cutting error rates on CIFAR-100 by approximately 0.5 to 1.5 percentage points while adding almost no parameters; for instance, a ResNet-50 model with the module matched the accuracy of a ResNet-101 model while using roughly half the parameters. Second, the module improved large-scale image classification on ImageNet-1K across deep, wide, and compact architectures, reducing top-1 error by up to 1.77 percentage points in efficient mobile models. Third, the module improved object detection accuracy on both MS COCO and VOC 2007 benchmarks. Finally, comparative tests demonstrated that combining spatial and channel attention at bottleneck locations delivered better accuracy and parameter efficiency than existing alternatives such as Squeeze-and-Excitation modules or placing attention inside every convolutional block.

These findings indicate that organizations deploying computer vision can achieve higher accuracy without the financial and operational burdens of larger server infrastructure or excessive latency on edge devices. By refining features at critical network bottlenecks, models learn hierarchical representations that filter background noise early and focus on target objects in deeper layers, mimicking efficient human visual perception.

Engineering and deployment teams should consider integrating the module into existing vision pipelines, particularly where resource efficiency is critical, by using the demonstrated reduction ratio and dilation parameters. Future work should explore combining this attention mechanism with specialized network compression techniques to further optimize edge device performance.

Confidence in these findings is high given the broad evaluation across multiple architectures and standardized benchmarks. However, stakeholders should note that the reported gains vary by specific backbone model and that real-world deployment on proprietary, domain-specific visual data should be validated before full system adoption.

Cover for BAM: Bottleneck Attention Module

Abstract

Recent advances in deep neural networks have been developed via architecture search for stronger representational power. In this work, we focus on the effect of attention in general deep neural networks. We propose a simple and effective attention module, named Bottleneck Attention Module (BAM), that can be integrated with any feed-forward convolutional neural networks. Our module infers an attention map along two separate pathways, channel and spatial. We place our module at each bottleneck of models where the downsampling of feature maps occurs. Our module constructs a hierarchical attention at bottlenecks with a number of parameters and it is trainable in an end-to-end manner jointly with any feed-forward models. We validate our BAM through extensive experiments on CIFAR-100, ImageNet-1K, VOC 2007 and MS COCO benchmarks. Our experiments show consistent improvement in classification and detection performances with various models, demonstrating the wide applicability of BAM. The code and models will be publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Bottleneck Attention Module
  • 4 Experiments
  • 4.1 Ablation studies on CIFAR-100
  • 4.2 Classification Results on CIFAR-100
  • 4.3 Classification Results on ImageNet-1K
  • 4.4 Effectiveness of BAM with Compact Networks
  • 4.5 MS COCO Object Detection
  • 4.6 VOC 2007 Object Detection
  • 4.7 Comparison with Squeeze-and-Excitation[Hu et al.(2017)Hu, Shen, and Sun]
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Bottleneck Attention Module (BAM) Architecture and Refinement

    model/method

    The Bottleneck Attention Module (BAM) is an attention mechanism designed to refine an intermediate 3D feature representation F∈RC×H×W\mathbf{F} \in \mathbb{R}^{C \times H \times W} (where CC is the number of channels, HH is height, and WW is width) within deep convolutional neural networks. BAM decomposes the computation of a 3D attention map M(F)∈RC×H×W\mathbf{M}(\mathbf{F}) \in \mathbb{R}^{C \times H \times W} into two parallel branches: a channel attention branch Mc(F)∈RC\mathbf{M}_c(\mathbf{F}) \in \mathbb{R}^{C} that determines what feature attributes to emphasize, and a spatial attention branch Ms(F)∈RH×W\mathbf{M}_s(\mathbf{F}) \in \mathbb{R}^{H \times W} that determines where to focus.

    The final 3D attention map is produced by expanding both branch outputs to RC×H×W\mathbb{R}^{C \times H \times W}, summing them element-wise, and passing the result through a sigmoid activation function σ\sigma:

    M(F)=σ(Mc(F)+Ms(F))\mathbf{M}(\mathbf{F}) = \sigma\left(\mathbf{M}_c(\mathbf{F}) + \mathbf{M}_s(\mathbf{F})\right)

    The refined feature map F′∈RC×H×W\mathbf{F}' \in \mathbb{R}^{C \times H \times W} is then computed using a residual learning formulation:

    F′=F+F⊗M(F)\mathbf{F}' = \mathbf{F} + \mathbf{F} \otimes \mathbf{M}(\mathbf{F})

    where ⊗\otimes denotes element-wise multiplication. The residual connection ensures smooth gradient flow during end-to-end backpropagation.

  2. Knowl 2 — BAM Channel Attention Pathway

    model/method

    The channel attention branch computes inter-channel dependencies from the input tensor F∈RC×H×W\mathbf{F} \in \mathbb{R}^{C \times H \times W}. Global contextual information for each channel is first aggregated using global average pooling:

    Fc=AvgPool(F)∈RC×1×1\mathbf{F}_c = \text{AvgPool}(\mathbf{F}) \in \mathbb{R}^{C \times 1 \times 1}

    To model non-linear interactions across channels while constraining parameter overhead, Fc\mathbf{F}_c is passed through a multi-layer perceptron (MLP) with a single hidden layer of reduced dimension RC/r×1×1\mathbb{R}^{C/r \times 1 \times 1}, where rr is the reduction ratio (set by default to r=16r=16). A Batch Normalization (BN\text{BN}) layer is applied at the output to align the scale of the channel attention with the spatial branch output:

    Mc(F)=BN(W1(W0AvgPool(F)+b0)+b1)\mathbf{M}_c(\mathbf{F}) = \text{BN}\left(\mathbf{W}_1\left(\mathbf{W}_0\text{AvgPool}(\mathbf{F}) + \mathbf{b}_0\right) + \mathbf{b}_1\right)

    where W0∈RC/r×C\mathbf{W}_0 \in \mathbb{R}^{C/r \times C}, b0∈RC/r\mathbf{b}_0 \in \mathbb{R}^{C/r}, W1∈RC×C/r\mathbf{W}_1 \in \mathbb{R}^{C \times C/r}, and b1∈RC\mathbf{b}_1 \in \mathbb{R}^{C}.

  3. Knowl 3 — BAM Spatial Attention Pathway with Dilated Convolutions

    model/method

    The spatial attention branch computes a spatial weight map Ms(F)∈RH×W\mathbf{M}_s(\mathbf{F}) \in \mathbb{R}^{H \times W} to emphasize informative regions and suppress background noise. To capture contextual dependencies over wide spatial regions efficiently, the branch employs dilated convolutions arranged in a bottleneck topology:

    Ms(F)=BN(f31×1(f23×3(f13×3(f01×1(F)))))\mathbf{M}_s(\mathbf{F}) = \text{BN}\left(f_3^{1\times 1}\left(f_2^{3\times 3}\left(f_1^{3\times 3}\left(f_0^{1\times 1}(\mathbf{F})\right)\right)\right)\right)

    where:

    • f01×1f_0^{1\times 1} is a 1×11\times 1 convolution that projects F∈RC×H×W\mathbf{F} \in \mathbb{R}^{C \times H \times W} down to RC/r×H×W\mathbb{R}^{C/r \times H \times W} using reduction ratio r=16r=16.
    • f13×3f_1^{3\times 3} and f23×3f_2^{3\times 3} are two successive 3×33\times 3 dilated convolutions with dilation rate d=4d=4, operating in the reduced channel dimension C/rC/r to enlarge the effective receptive field without parameter inflation.
    • f31×1f_3^{1\times 1} is a 1×11\times 1 convolution reducing the channel dimension from C/rC/r to 11, yielding a tensor in R1×H×W\mathbb{R}^{1 \times H \times W}.
    • BN\text{BN} is a batch normalization layer applied to adjust the output scale prior to combination with the channel branch.
  4. Knowl 4 — Network Bottleneck Placement Strategy for Attention Modules

    model/method

    BAM is integrated specifically at the bottlenecks of convolutional neural network architectures—the transition interfaces between distinct stages where spatial downsampling occurs (e.g., via pooling or strided operations) and channel dimensions increase. Rather than inserting attention mechanisms inside every convolutional block, placing BAM at stage bottlenecks concentrates feature refinement at critical points of information flow, denoising low-level features at early bottlenecks and focusing on high-level semantics at later bottlenecks while minimizing computational (FLOPs) and parametric overhead.

  5. Knowl 5 — Ablation on BAM Branch Composition and Combination Operations

    empirical result

    Ablation experiments on the CIFAR-100 dataset using a ResNet-50 backbone (baseline top-1 error: 21.49%21.49\%) evaluate the effectiveness of individual attention branches and different combining operators:

    Configuration Channel Branch Spatial Branch Combining Op Top-1 Error (%)
    Baseline (ResNet-50) - - - 21.49
    Channel-only ✓ - 21.29
    Spatial-only ✓ - 21.24
    BAM (Element-wise Max) ✓ ✓ MAX 20.28
    BAM (Element-wise Product) ✓ ✓ PROD 20.21
    BAM (Element-wise Sum) ✓ ✓ SUM 20.00

    Combining both channel and spatial branches outperforms using either branch in isolation. Among the combining methods, element-wise summation achieves the lowest error (20.00%20.00\%) because it allows bidirectional, equal gradient distribution to both branches during backward propagation and retains complementary information from both pathways. Element-wise product produces inferior results because large gradients are assigned to small inputs, impeding convergence, while element-wise maximum routes gradients only to the dominant branch, destabilizing training.

  6. Knowl 6 — Hyperparameter Optimization for BAM Dilation Rate and Reduction Ratio

    empirical result

    Ablation experiments on CIFAR-100 with ResNet-50 evaluate the sensitivity of BAM to the spatial dilation rate dd and channel reduction ratio rr:

    Hyperparameter Value Parameters Top-1 Error (%)
    Baseline (ResNet-50) - 23.68M 21.49
    Dilation rate (dd) 1 24.07M 20.47
    2 24.07M 20.28
    4 24.07M 20.00
    6 24.07M 20.08
    Reduction ratio (rr) 4 26.30M 20.46
    8 24.62M 20.56
    16 24.07M 20.00
    32 23.87M 21.24

    Standard convolutions (d=1d=1) yield the highest error (20.47%20.47\%), while dilated convolutions expand the receptive field to capture context, with performance saturating at d=4d=4 (20.00%20.00\%). For the reduction ratio, r=16r=16 provides the best trade-off between representational capacity and regularization; lower ratios (r=4,8r=4, 8) incur higher parameter counts and suffer from overfitting.

  7. Knowl 7 — Bottleneck Placement vs. Inside-Block Insertion (BAM-C)

    empirical result

    Comparing the placement of BAM at network bottlenecks against inserting attention inside every convolutional block (designated BAM-C) across standard architectures on CIFAR-100:

    Architecture Params GFLOPs Top-1 Error (%)
    ResNet-50 Baseline 23.71M 1.22 21.49
    ResNet-50 + BAM-C 28.98M 1.37 20.88
    ResNet-50 + BAM 24.07M 1.25 20.00
    PreResNet-110 Baseline 1.73M 0.245 22.22
    PreResNet-110 + BAM-C 2.17M 0.275 21.29
    PreResNet-110 + BAM 1.73M 0.246 21.96
    WideResNet-28 (w=8w=8) Baseline 23.40M 3.36 20.40
    WideResNet-28 (w=8w=8) + BAM-C 23.78M 3.39 20.06
    WideResNet-28 (w=8w=8) + BAM 23.42M 3.37 19.06
    ResNeXt-29 (8×64d8\times 64\text{d}) Baseline 34.52M 4.99 18.18
    ResNeXt-29 (8×64d8\times 64\text{d}) + BAM-C 35.60M 5.07 18.15
    ResNeXt-29 (8×64d8\times 64\text{d}) + BAM 34.61M 5.00 16.71

    Placing BAM strictly at the bottlenecks achieves superior accuracy-to-overhead trade-offs compared to inserting modules inside every convolutional block, reducing error by larger margins while adding substantially fewer parameters and GFLOPs.

  8. Knowl 8 — ImageNet-1K Classification Performance Across Deep and Compact Architectures

    empirical result

    Evaluation of BAM on the ImageNet-1K validation set (single-crop 224×224224 \times 224 evaluation) across standard, wide, multi-branch, and mobile architectures:

    Architecture Params GFLOPs Top-1 Error (%) Top-5 Error (%)
    ResNet-18 11.69M 1.81 29.60 10.55
    ResNet-18 + BAM 11.71M (+0.02) 1.82 (+0.01) 28.88 10.01
    ResNet-50 25.56M 3.86 24.56 7.50
    ResNet-50 + BAM 25.92M (+0.36) 3.94 (+0.08) 24.02 7.18
    ResNet-101 44.55M 7.57 23.38 6.88
    ResNet-101 + BAM 44.91M (+0.36) 7.65 (+0.08) 22.44 6.29
    WideResNet-18 (w=1.5w=1.5) 25.88M 3.87 26.85 8.88
    WideResNet-18 (w=1.5w=1.5) + BAM 25.93M (+0.05) 3.88 (+0.01) 26.67 8.69
    WideResNet-18 (w=2.0w=2.0) 45.62M 6.70 25.63 8.20
    WideResNet-18 (w=2.0w=2.0) + BAM 45.71M (+0.09) 6.72 (+0.02) 25.00 7.81
    ResNeXt-50 (32×4d32\times 4\text{d}) 25.03M 3.77 22.85 6.48
    ResNeXt-50 (32×4d32\times 4\text{d}) + BAM 25.39M (+0.36) 3.85 (+0.08) 22.56 6.40
    MobileNet 4.23M 0.569 31.39 11.51
    MobileNet + BAM 4.32M (+0.09) 0.589 (+0.02) 30.58 10.90
    MobileNet (α=0.7\alpha=0.7) 2.30M 0.283 34.86 13.69
    MobileNet (α=0.7\alpha=0.7) + BAM 2.34M (+0.04) 0.292 (+0.009) 33.09 12.69
    MobileNet (ρ=192/224\rho=192/224) 4.23M 0.439 32.89 12.33
    MobileNet (ρ=192/224\rho=192/224) + BAM 4.32M (+0.09) 0.456 (+0.017) 31.56 11.60
    SqueezeNet v1.1 1.24M 0.290 43.09 20.48
    SqueezeNet v1.1 + BAM 1.26M (+0.02) 0.304 (+0.014) 41.83 19.58

    Inserting only three BAM modules across the entire network reduces error rates consistently across all network scales with negligible overhead.

  9. Knowl 9 — BAM vs. Squeeze-and-Excitation (SE) on CIFAR-100

    empirical result

    Comparative performance of BAM versus Squeeze-and-Excitation (SE) modules on the CIFAR-100 benchmark:

    Architecture Params GFLOPs Top-1 Error (%)
    ResNet-50 Baseline 23.71M 1.22 21.49
    ResNet-50 + SE 26.24M 1.23 20.72
    ResNet-50 + BAM 24.07M 1.25 20.00
    PreResNet-110 Baseline 1.73M 0.245 22.22
    PreResNet-110 + SE 1.93M 0.245 21.85
    PreResNet-110 + BAM 1.73M 0.246 21.96
    WideResNet-28 (w=8w=8) Baseline 23.40M 3.36 20.40
    WideResNet-28 (w=8w=8) + SE 23.58M 3.36 19.85
    WideResNet-28 (w=8w=8) + BAM 23.42M 3.37 19.06
    ResNeXt-29 (16×64d16\times 64\text{d}) Baseline 68.25M 9.88 17.25
    ResNeXt-29 (16×64d16\times 64\text{d}) + SE 68.81M 9.88 16.52
    ResNeXt-29 (16×64d16\times 64\text{d}) + BAM 68.34M 9.90 16.39

    BAM achieves lower top-1 classification error than SE on ResNet-50 (20.00%20.00\% vs. 20.72%20.72\%), WideResNet-28 (19.06%19.06\% vs. 19.85%19.85\%), and ResNeXt-29 (16.39%16.39\% vs. 16.52%16.52\%) while using fewer parameters because BAM is placed only at stage bottlenecks rather than within every convolutional block.

  10. Knowl 10 — Object Detection Performance on MS COCO and PASCAL VOC 2007

    empirical result

    Performance of BAM integrated into object detection frameworks on MS COCO and PASCAL VOC 2007:

    (a) MS COCO Detection (Faster-RCNN with ResNet-101)
    Architecture [email protected] [email protected] mAP@[0.5, 0.95]
    ResNet-101 Baseline 48.4 30.7 29.1
    ResNet-101 + BAM 50.2 32.5 30.4
    (b) PASCAL VOC 2007 Detection
    Backbone Detector Params (M) [email protected] (%)
    VGG-16 SSD 26.5 77.8
    VGG-16 StairNet 32.0 78.9
    VGG-16 StairNet + BAM 32.1 79.3
    MobileNet SSD 5.81 68.1
    MobileNet StairNet 5.98 70.1
    MobileNet StairNet + BAM 6.00 70.6

    On MS COCO, adding BAM to the ResNet-101 backbone of Faster-RCNN improves [email protected] by +1.8%+1.8\% and overall mAP@[0.5, 0.95] by +1.3%+1.3\%. On VOC 2007, inserting BAM right before classifiers in StairNet increases [email protected] by +0.4%+0.4\% on VGG-16 and +0.5%+0.5\% on MobileNet with minimal parameter increase (+0.1M+0.1\text{M} and +0.02M+0.02\text{M} params).

Coverage note — Qualitative visual attention heatmaps across network stages were omitted as they serve illustrative analysis rather than standalone quantitative contributions.

References

  1. 1.Pytorch. http://pytorch.org/. Accessed: 2018-04-20.
  2. 2.Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu. Multiple object recognition with visual attention. 2014.
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. 2014.
  4. 4.Sean Bell, C Lawrence Zitnick, Kavita Bala, and Ross Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2016.
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv preprint arXiv:1606.00915, 2016.
  6. 6.Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, and Tat-Seng Chua. Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2017.
  7. 7.François Chollet. Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357, 2016.
  8. 8.Maurizio Corbetta and Gordon L Shulman. Control of goal-directed and stimulus-driven attention in the brain. In Nature reviews neuroscience 3.3, 2002.
  9. 9.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. CoRR, abs/1703.06211, 1(2):3, 2017.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2009.
  11. 11.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672–2680, 2014.
  12. 12.Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra. Draw: A recurrent neural network for image generation. 2015.
  13. 13.Dongyoon Han, Jiwhan Kim, and Junmo Kim. Deep pyramidal residual networks. In Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, pages 6307–6315. IEEE, 2017.
  14. 14.Bharath Hariharan, Pablo Arbeláez, Ross Girshick, and Jitendra Malik. Hypercolumns for object segmentation and fine-grained localization. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2015.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2016.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In Proc. of European Conf. on Computer Vision (ECCV), 2016.
  17. 17.Joy Hirsch and Christine A Curcio. The spatial resolution capacity of human foveal retina. Vision research, 29(9):1095–1101, 1989.
  18. 18.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  19. 19.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. arXiv preprint arXiv:1709.01507, 2017.
  20. 20.Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993, 2016.
  21. 21.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In Proc. of European Conf. on Computer Vision (ECCV), 2016.
  22. 22.Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <0.5mb model size. arXiv preprint arXiv:1602.07360, 2016.
  23. 23.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. 2015.
  24. 24.Laurent Itti, Christof Koch, and Ernst Niebur. A model of saliency-based visual attention for rapid scene analysis. In IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI), 1998.
  25. 25.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. In Proc. of Neural Information Processing Systems (NIPS), 2015.
  26. 26.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. In Advances in neural information processing systems, pages 2017–2025, 2015.
  27. 27.Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool. Dynamic filter networks. In Advances in Neural Information Processing Systems, pages 667–675, 2016.
  28. 28.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  29. 29.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images.
  30. 30.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Proc. of Neural Information Processing Systems (NIPS), 2012.
  31. 31.Hugo Larochelle and Geoffrey E Hinton. Learning to combine foveal glimpses with a third-order boltzmann machine. In Proc. of Neural Information Processing Systems (NIPS), 2010.
  32. 32.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proc. of European Conf. on Computer Vision (ECCV), 2014.
  33. 33.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In Proc. of European Conf. on Computer Vision (ECCV), 2016.
  34. 34.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2015.
  35. 35.Volodymyr Mnih, Nicolas Heess, Alex Graves, et al. Recurrent models of visual attention." advances in neural information processing systems. In Proc. of Neural Information Processing Systems (NIPS), 2014.
  36. 36.Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. Dual attention networks for multimodal reasoning and matching. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2156–2164, 2017.
  37. 37.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Proc. of Neural Information Processing Systems (NIPS), 2015.
  38. 38.Ronald A Rensink. The dynamic representation of scenes. In Visual cognition 7.1-3, 2000.
  39. 39.Woo Sanghyun, Hwang Soonmin, and Kweon In So. Stairnet: Top-down semantic aggregation for accurate one shot detection. In Proc. of Winter Conf. on Applications of Computer Vision (WACV), 2018.
  40. 40.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  41. 41.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2015.
  42. 42.Fei Wang, Mengqing Jiang, Chen Qian, Shuo Yang, Cheng Li, Honggang Zhang, Xiaogang Wang, and Xiaoou Tang. Residual attention network for image classification. arXiv preprint arXiv:1704.06904, 2017.
  43. 43.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. arXiv preprint arXiv:1611.05431, 2016.
  44. 44.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. 2015.
  45. 45.Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2016.
  46. 46.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. 2015.
  47. 47.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  48. 48.Matthew D Zeiler. Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012.
  49. 49.Yousong Zhu, Chaoyang Zhao, Jinqiao Wang, Xu Zhao, Yi Wu, and Hanqing Lu. Couplenet: Coupling global structure with local parts for object detection. In Proc. of IntâA˘Zl Conf. on Computer Vision (ICCV) ´ , 2017.

Citation

MLA
Park, J., et al. “BAM: Bottleneck Attention Module”. arXiv, 2018, http://arxiv.org/abs/1807.06514v2.
APA
Park, J., Woo, S., Lee, J.-Y., & Kweon, I. S. (2018). BAM: Bottleneck Attention Module. arXiv. http://arxiv.org/abs/1807.06514v2
Chicago
Park, J., S. Woo, J.-Y. Lee, and I. S. Kweon. 2018. “BAM: Bottleneck Attention Module”. arXiv. http://arxiv.org/abs/1807.06514v2.
Harvard
Park, J. et al. (2018) “BAM: Bottleneck Attention Module”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1807.06514v2.
Vancouver
1. Park J, Woo S, Lee J-Y, Kweon IS (2018) BAM: Bottleneck Attention Module. arXiv

BibTeX

@article{park2018bam,
  title = {BAM: Bottleneck Attention Module},
  author = {Park, Jongchan and Woo, Sanghyun and Lee, Joon-Young and Kweon, In So},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1807.06514v2},
  eprint = {1807.06514}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors