Attention to Scale: Scale-Aware Semantic Image Segmentation

Liang-Chieh ChenYi YangJiang WangWei XuAlan L. Yuille

article2015CVPR1,394 citations

Proposes an attention mechanism that dynamically weights multi-scale features at each pixel, improving semantic image segmentation accuracy over standard pooling baselines while providing interpretable diagnostics of scale selection.

Listen

Semantic image segmentation—the process of assigning a category label to every individual pixel in a digital image—is a foundational technology for high-stakes applications such as autonomous driving, medical imaging, image editing, and augmented reality. A persistent challenge in this domain is handling objects of drastically varying sizes within the same scene. Standard deep learning models often struggle to segment small details and broad contextual regions simultaneously because traditional methods for merging multi-scale image features rely on rigid, uniform operations like average-pooling or max-pooling that treat all image regions equally.

The main objective of the article is to design, evaluate, and demonstrate an attention-based deep learning mechanism that adaptively weights multi-scale image features at every individual pixel location, paired with scale-specific extra supervision to improve segmentation accuracy.

To accomplish this, the authors extended a leading convolutional network framework (DeepLab-LargeFOV) into a shared multi-scale architecture. Rather than relying on static feature-merging rules, they trained a compact neural attention model directly with the primary network in a single, end-to-end training pipeline. The approach was systematically evaluated across three standard benchmark datasets: PASCAL-Person-Part, PASCAL VOC 2012, and a 10,000-image subset of MS-COCO 2014, testing various input scale combinations (such as full, three-quarter, and half resolutions) with and without scale-level supervision.

The experimental findings show significant, consistent performance gains across all benchmarks. First, the proposed attention model consistently outperformed standard pooling strategies across all evaluated datasets, achieving a mean intersection-over-union score of 56.39% on PASCAL-Person-Part and 71.5% on the PASCAL VOC 2012 test set without conditional random field post-processing. Second, injecting extra supervision at each individual scale proved vital, delivering notable performance boosts across all merging configurations and preventing feature degradation when scaling to three input resolutions. Third, the attention model provides clear diagnostic transparency by generating interpretable weight maps showing that full-resolution processing focuses on fine, small-scale details while downscaled inputs automatically capture large objects and broad background context.

These findings indicate that dynamic, pixel-level scale weighting offers a practical and explainable upgrade for visual understanding systems. The unified training approach avoids cumbersome multi-stage workflows, maintaining manageable training runtimes of approximately 21 hours on a single graphics processing unit and fast per-image inference times of about 350 milliseconds. While it does not require manual annotations for scale selection, it provides engineering and safety teams with visual interpretability into why a network prioritizes specific visual features.

Organizations developing computer vision systems should integrate learned multi-scale attention and scale-specific loss supervision into their segmentation architectures rather than relying on static feature pooling. For maximum accuracy, practitioners should combine this attention mechanism with complementary refinement techniques, such as conditional random fields or domain transforms, and employ scale-jittering data augmentation. Further research and data collection are recommended to address challenging edge cases, specifically highly unusual human poses, extreme visual occlusions such as clothing, and very small or imbalanced object categories that still present recognition difficulties.

arXiv: 1511.03339
Cover for Attention to Scale: Scale-Aware Semantic Image Segmentation

Abstract

Incorporating multi-scale features in fully convolutional neural networks (FCNs) has been a key element to achieving state-of-the-art performance on semantic image segmentation. One common way to extract multi-scale features is to feed multiple resized input images to a shared deep network and then merge the resulting features for pixelwise classification. In this work, we propose an attention mechanism that learns to softly weight the multi-scale features at each pixel location. We adapt a state-of-the-art semantic image segmentation model, which we jointly train with multi-scale input images and the attention model. The proposed attention model not only outperforms average- and max-pooling, but allows us to diagnostically visualize the importance of features at different positions and scales. Moreover, we show that adding extra supervision to the output at each scale is essential to achieving excellent performance when merging multi-scale features. We demonstrate the effectiveness of our model with extensive experiments on three challenging datasets, including PASCAL-Person-Part, PASCAL VOC 2012 and a subset of MS-COCO 2014.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model
  • 3.1 Review of DeepLab
  • 3.2 Attention model for scales
  • 3.3 Extra supervision
  • 4 Experimental Evaluations
  • 4.1 PASCAL-Person-Part
  • 4.2 PASCAL VOC 2012
  • 4.3 Subset of MS-COCO
  • 5 Conclusion
  • A More qualitative results
  • References

Knowls

  1. Knowl 1 — Scale-attention fusion of shared multi-scale FCNs

    model/method

    The proposed share-net processes resized versions of the same image with a single DeepLab network whose weights are shared across all input scales. Let SS be the number of scales, let s∈{1,…,S}s\in\{1,\ldots,S\} index a scale, let ii index a position on the common finest-scale output grid, and let c∈{1,…,C}c\in\{1,\ldots,C\} index a semantic class. The FCN produces a real-valued score map fi,csf^s_{i,c} for each scale; all score maps are resized to the finest-scale resolution by bilinear interpolation.

    A second fully convolutional attention network produces a real-valued scale logit hish^s_i at every position. The scale weights are normalized with a softmax over scales and are shared across semantic channels:

    wis=exp⁡(his)∑t=1Sexp⁡(hit),wis≥0,∑s=1Swis=1.w^s_i=\frac{\exp(h^s_i)}{\sum_{t=1}^{S}\exp(h^t_i)},\qquad w^s_i\geq 0,\qquad \sum_{s=1}^{S}w^s_i=1.

    The fused class score at position ii is the weighted sum

    gi,c=∑s=1Swisfi,cs.g_{i,c}=\sum_{s=1}^{S}w^s_i f^s_{i,c}.

    The final semantic prediction is obtained by applying a softmax over classes to gi,cg_{i,c}. The attention weights provide a position-dependent, differentiable choice of scale: average pooling is recovered by fixing wis=1/Sw^s_i=1/S, while max pooling replaces the weighted sum by a maximum over scale-specific scores. Because the weighting is differentiable, the attention network and all shared FCN branches are trained jointly without pixel-level ground-truth scale annotations.

  2. Knowl 2 — Attention-network architecture and scale-specific interpretation

    model/method

    The scale-attention network is itself a fully convolutional network. It takes the convolutionalized VGG-16 fc7fc7 features as input, applies a first convolution with 512 filters of spatial size 3×33\times3, and applies a second convolution with SS filters of spatial size 1×11\times1. The SS output channels are the logits hish^s_i used to compute one attention map per input scale. Since the output is dense, every image position receives a separate distribution over scales.

    The learned maps are intended to diagnose which scale supplies useful evidence at each position. In the reported visualizations, the scale-11 map generally emphasizes small objects or parts, the scale-0.750.75 map emphasizes intermediate-scale structures, and the scale-0.50.5 map emphasizes large objects and background context. This scale specialization is spatially varying rather than a single global preference for one input resolution.

  3. Knowl 3 — Extra supervision for every scale branch

    model/method

    The training objective supervises both the fused output and the individual FCN output at every input scale. If pi(c)p_i(c) is the final softmax probability for class cc at position ii, pis(c)p^s_i(c) is the softmax probability from the FCN branch at scale ss, and yiy_i is the corresponding ground-truth class, the total loss is the sum of 1+S1+S cross-entropy losses, with unit weight for every term:

    L=−∑i∈Ω0log⁡pi(yi)−∑s=1S∑i∈Ωslog⁡pis(yis).\mathcal{L}=-\sum_{i\in\Omega_0}\log p_i(y_i)-\sum_{s=1}^{S}\sum_{i\in\Omega_s}\log p^s_i(y^s_i).

    Here Ω0\Omega_0 and Ωs\Omega_s are the output grids of the fused prediction and scale-ss branch, respectively, and yisy^s_i denotes the pixel annotation downsampled to the resolution of branch ss. The extra branch losses train the score maps being merged to be discriminative before pooling or attention. Across the experiments, this additional supervision is essential for obtaining strong results, especially when three scale-specific predictions are merged.

  4. Knowl 4 — DeepLab-LargeFOV share-net backbone

    model/method

    The base network is DeepLab-LargeFOV built from VGG-16. The original fully connected layers are converted to convolutional layers so that the network produces dense score maps. Five stride-2 pooling layers would normally give an output stride of 32; DeepLab replaces the relevant downsampling with the atrous algorithm to reduce the stride to 8, then bilinearly upsamples the final score maps by a factor of 8 to image resolution. The LargeFOV variant uses atrous filters in the convolutional form of VGG-16 fc6fc6 to enlarge the receptive field.

    For scale-aware segmentation, the same backbone parameters process every resized input image. The scale branches, attention network, fused classifier, and extra branch classifiers are optimized end-to-end from ImageNet-pretrained VGG-16 parameters. This changes multi-scale processing from a fixed average or maximum over independently useful features into a jointly learned, spatially adaptive fusion.

  5. Knowl 5 — Experimental protocol and computational cost

    experimental setup

    Performance is measured by mean pixel intersection-over-union, averaged across semantic classes. Experiments use two or three input scales, principally {1,0.5}\{1,0.5\} and {1,0.75,0.5}\{1,0.75,0.5\}. The PASCAL-Person-Part experiment uses 1,716 person-containing training images and 1,817 validation images, with six merged person-part classes—Head, Torso, Upper Arms, Lower Arms, Upper Legs, and Lower Legs—plus background. PASCAL VOC 2012 uses 20 foreground classes and background, with the standard augmented training annotations and validation/test evaluation. The MS-COCO experiment uses 10,000 randomly selected training images and 1,500 validation images from the 2014 dataset, which contains 80 foreground classes and background.

    Training uses SGD with mini-batches of 30 images, initial learning rate 0.0010.001 for the network and 0.010.01 for the final classifier, momentum 0.90.9, and weight decay 0.00050.0005. The learning rate is multiplied by 0.10.1 after 2,000 iterations. Fine-tuning takes approximately 21 hours on an NVIDIA Tesla K40 GPU; jointly processing the scaled inputs takes roughly twice the training time of a single-scale DeepLab-LargeFOV. Average inference time for one PASCAL image is 350 ms.

  6. Knowl 6 — PASCAL-Person-Part ablation results

    data/table

    The PASCAL-Person-Part validation experiment compares the single-scale DeepLab-LargeFOV baseline with two- and three-scale fusion, with and without extra supervision. The metric is mean pixel IOU in percent. Extra supervision consistently raises performance for every merging method, and attention is the best merger for both scale sets.

    Model or merger without extra supervision with extra supervision
    DeepLab-LargeFOV baseline 51.91 –
    Scales {1,0.5}\{1,0.5\}
    Max pooling 52.90 55.26
    Average pooling 52.71 55.17
    Attention 53.49 55.85
    Scales {1,0.75,0.5}\{1,0.75,0.5\}
    Max pooling 53.02 55.78
    Average pooling 52.56 55.72
    Attention 53.12 56.39

    The best result is 56.39%, obtained with three scales, attention fusion, and extra supervision. It exceeds the single-scale baseline by 4.48 percentage points and the reported DeepLab-MSc-LargeFOV skip-net result of 53.72% by 2.67 percentage points. With three scales and no extra supervision, attention reaches only 53.12%, demonstrating that merely adding more scales does not guarantee useful fusion.

  7. Knowl 7 — PASCAL VOC 2012 results

    data/table

    On the PASCAL VOC 2012 validation set with ImageNet-pretrained DeepLab-LargeFOV, the proposed attention merger improves over average and max pooling, while extra supervision is particularly important for three-scale fusion. Values are mean pixel IOU in percent.

    Model or merger without extra supervision with extra supervision
    DeepLab-LargeFOV baseline 62.28 –
    Scales {1,0.5}\{1,0.5\}
    Max pooling 64.81 67.43
    Average pooling 64.86 67.79
    Attention 65.27 68.24
    Scales {1,0.75,0.5}\{1,0.75,0.5\}
    Max pooling 65.15 67.79
    Average pooling 63.92 67.98
    Attention 64.37 69.08

    The best validation score is 69.08%, a 6.80-point improvement over the 62.28% baseline. On the test set, the single-scale DeepLab-LargeFOV, DeepLab-MSc-LargeFOV, and the paper's attention model obtain 65.1%, 67.0%, and 71.5% respectively when pretrained with ImageNet. The attention model therefore improves by 6.4 points over DeepLab-LargeFOV and 4.5 points over DeepLab-MSc-LargeFOV. With MS-COCO pretraining, the best validation result is 71.42% versus a 67.58% baseline, and the best test result with fully connected CRF post-processing is 75.1%.

  8. Knowl 8 — MS-COCO scale-fusion results

    data/table

    On the selected MS-COCO 2014 validation subset, multi-scale fusion remains beneficial despite the larger number of classes and greater scale variation. The table reports mean pixel IOU in percent for all classes; the final two columns compare training without and with extra branch supervision.

    Model or merger without extra supervision with extra supervision
    DeepLab-LargeFOV baseline 31.22 –
    Scales {1,0.5}\{1,0.5\}
    Max pooling 32.95 34.70
    Average pooling 33.69 35.14
    Attention 34.03 35.41
    Scales {1,0.75,0.5}\{1,0.75,0.5\}
    Max pooling 33.58 35.08
    Average pooling 33.74 35.72
    Attention 33.42 35.78

    The best all-class result is 35.78%, which improves over the 31.22% baseline by 4.56 points and over the reported DeepLab-MSc-LargeFOV score of 31.61% by 4.17 points. For the frequently occurring person class, the baseline is 68.76%; three-scale attention with extra supervision reaches 72.72%, while two-scale attention with extra supervision reaches 72.20%. The all-class gains are smaller and average pooling is close to attention, which the paper attributes to very low accuracy on small, imbalanced object classes such as fork, mouse, and toothbrush.

  9. Knowl 9 — Learned attention maps provide scale diagnostics

    empirical result

    The attention mechanism produces interpretable spatial maps rather than only a final segmentation. Across the qualitative experiments, the scale-11 map tends to highlight small objects or parts, the scale-0.750.75 map tends to highlight middle-sized structures, and the scale-0.50.5 map tends to highlight the largest objects and background context. Examples include small person parts at scale 11, middle-scale torsos and legs at scale 0.750.75, and large legs or background at scale 0.50.5.

    The corresponding max-pooling maps are less informative because max pooling provides no learned, continuously varying importance values. The learned maps therefore serve as a diagnostic of which resized input contributes to each pixel's decision, while the same attention model also improves quantitative segmentation over fixed average- and max-pooling in the principal experiments.

  10. Knowl 10 — Sensitivity to attention design and number of scales

    empirical result

    The selected attention design is not highly sensitive to moderate architectural changes, but the choice of feature level and scale range matters. Replacing the two-layer attention network with a one-layer network, changing the first kernel from 3×33\times3 to 1×11\times1, or varying the number of first-layer filters causes only a 0.1%0.1\%--0.4%0.4\% degradation. Using convolutionalized VGG-16 fc8fc8 features instead of the chosen fc7fc7 features reduces performance by approximately 0.5%0.5\%; experiments using fc6fc6 and fc7fc7 have similar behavior.

    Adding a fourth input scale, {1,0.75,0.5,0.25}\{1,0.75,0.5,0.25\}, lowers performance by approximately 0.5%0.5\%. The paper suggests that the score maps from scale 0.250.25 are too spatially small to contribute useful information. These results support using three input scales at most and using fc7fc7 features as the attention input.

  11. Knowl 11 — Reported limitations and failure cases

    limitation

    The proposed model does not surpass the strongest contemporaneous segmentation systems that jointly train fully connected CRF or other structured components with the FCN. The paper presents scale attention as complementary to those structured models rather than as a replacement for them.

    On MS-COCO, improvements are limited by severe class imbalance and poor recognition of very small object categories. In the person-part experiments, qualitative failures arise from extremely difficult human poses and from confusion between clothing and person parts when the parts are occluded. The paper identifies additional data as a possible remedy for pose failures, while clothing-versus-part ambiguity remains difficult because the visual evidence for the underlying part is often unavailable.

Coverage note — No substantial contributed component was omitted; proof-free background and supplementary qualitative examples were excluded because they do not add distinct load-bearing method or result content.

References

  1. 1.M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In CVPR, 2014. 5
  2. 2.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik. Contour detection and hierarchical image segmentation. PAMI, 33(5):898–916, 2011. 2
  3. 3.R. T. Azuma. A survey of augmented reality. Presence: Teleoperators and virtual environments, 6(4):355–385, 1997. 1
  4. 4.J. Ba, V. Mnih, and K. Kavukcuoglu. Multiple object recognition with visual attention. arXiv:1412.7755, 2014. 2
  5. 5.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. In ICLR, 2015. 2, 3, 4
  6. 6.Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle, et al. Greedy layer-wise training of deep networks. In NIPS, 2007. 2, 4
  7. 7.J. C. Caicedo and S. Lazebnik. Active object localization with deep reinforcement learning. In ICCV, 2015. 2
  8. 8.C. Cao, X. Liu, Y. Yang, Y. Yu, J. Wang, Z. Wang, Y. Huang, L. Wang, C. Huang, W. Xu, et al. Look and think twice: Capturing top-down visual attention with feedback convolutional neural networks. In ICCV, 2015. 2
  9. 9.K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia. ABC-CNN: An attention based convolutional neural network for visual question answering. arXiv:1511.05960, 2015. 2
  10. 10.L.-C. Chen, J. T. Barron, G. Papandreou, K. Murphy, and A. L. Yuille. Semantic image segmentation with task-specific edge detection using cnns and a discriminatively trained domain transform. arXiv:1511.03328, 2015. 7
  11. 11.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015. 1, 2, 3, 4, 5, 6, 7
  12. 12.L.-C. Chen, A. Schwing, A. Yuille, and R. Urtasun. Learning deep structured models. In ICML, 2015. 7
  13. 13.X. Chen, R. Mottaghi, X. Liu, S. Fidler, R. Urtasun, and A. Yuille. Detect what you can: Detecting and representing objects using holistic models and body parts. In CVPR, 2014. 2, 4, 9
  14. 14.D. Ciresan, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In CVPR, 2012. 1, 3
  15. 15.J. Dai, K. He, and J. Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In ICCV, 2015. 1, 2, 3, 5, 7
  16. 16.D. Eigen and R. Fergus. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In ICCV, 2015. 2
  17. 17.M. Evening. Adobe Photoshop CS2 for Photographers: A professional image editor’s guide to the creative use of Photoshop for the Macintosh and PC. Taylor & Francis, 2005. 1
  18. 18.M. Everingham, S. A. Eslami, L. V. Gool, C. K. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. IJCV, 111(1):98–136, 2014. 1, 2, 4, 5, 9
  19. 19.C. Farabet, C. Couprie, L. Najman, and Y. LeCun. Learning hierarchical features for scene labeling. PAMI, 35(8):1915–1929, 2013. 1, 2
  20. 20.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. PAMI, 32(9):1627–1645, 2010. 1, 3
  21. 21.L. Florack, B. T. H. Romeny, M. Viergever, and J. Koenderink. The gaussian scale-space paradigm and the multi-scale local jet. IJCV, 18(1):61–75, 1996. 2
  22. 22.J. Fritsch, T. Kuhnl, and A. Geiger. A new performance measure and evaluation benchmark for road detection algorithms. In Intelligent Transportation Systems-(ITSC), 2013 16th International IEEE Conference on, pages 1693–1700. IEEE, 2013. 1
  23. 23.E. S. L. Gastal and M. M. Oliveira. Domain transform for edge-aware image and video processing. In SIGGRAPH, 2011. 7
  24. 24.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 2
  25. 25.K. Gregor, I. Danihelka, A. Graves, and D. Wierstra. Draw: A recurrent neural network for image generation. In ICML, 2015. 2, 3
  26. 26.B. Hariharan, P. Arbelaez, L. Bourdev, S. Maji, and J. Malik. Semantic contours from inverse detectors. In ICCV, 2011. 5
  27. 27.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Hypercolumns for object segmentation and fine-grained localization. In CVPR, 2015. 1, 2
  28. 28.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In ECCV. 2014. 2
  29. 29.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv:1408.5093, 2014. 4
  30. 30.P. Kr¨ahenb¨uhl and V. Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In NIPS, 2011. 7
  31. 31.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 2
  32. 32.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 2
  33. 33.C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu. Deeply-supervised nets. In AISTATS, 2015. 2, 4
  34. 34.G. Lin, C. Shen, I. Reid, et al. Efficient piecewise training of deep structured models for semantic segmentation. arXiv:1504.01013, 2015. 1, 2, 7
  35. 35.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 2, 4, 7, 8, 9
  36. 36.W. Liu, A. Rabinovich, and A. C. Berg. Parsenet: Looking wider to see better. arXiv:1506.04579, 2015. 2, 7
  37. 37.Z. Liu, X. Li, P. Luo, C. C. Loy, and X. Tang. Semantic image segmentation via deep parsing network. In ICCV, 2015. 1, 3, 7
  38. 38.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1, 2, 3, 4
  39. 39.S. Mallat. A Wavelet Tour of Signal Processing. Acad. Press, 2 edition, 1999. 3
  40. 40.V. Mnih, N. Heess, A. Graves, et al. Recurrent models of visual attention. In NIPS, 2014. 2
  41. 41.M. Mostajabi, P. Yadollahpour, and G. Shakhnarovich. Feedforward semantic segmentation with zoom-out features. In CVPR, 2015. 1, 2, 7
  42. 42.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. arXiv:1505.04366, 2015. 1, 2
  43. 43.G. Papandreou, L.-C. Chen, K. Murphy, and A. L. Yuille. Weakly- and semi-supervised learning of a dcnn for semantic image segmentation. In ICCV, 2015. 7
  44. 44.G. Papandreou, I. Kokkinos, and P.-A. Savalle. Untangling local and global deformations in deep convolutional networks for image classification and sliding window detection. In CVPR, 2015. 1, 2, 3
  45. 45.P. H. Pinheiro and R. Collobert. Recurrent convolutional neural networks for scene parsing. arXiv:1306.2795, 2013. 1, 2
  46. 46.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. IJCV, pages 1–42, 2015. 5
  47. 47.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In ICLR, 2014. 2
  48. 48.S. Sharma, R. Kiros, and R. Salakhutdinov. Action recognition using visual attention. arXiv:1511.04119, 2015. 2
  49. 49.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 2, 3, 4, 5
  50. 50.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. arXiv:1409.4842, 2014. 2, 4
  51. 51.J. Wang and A. Yuille. Semantic part segmentation using compositional model combining shape and appearance. In CVPR, 2015. 4
  52. 52.P. Wang, X. Shen, Z. Lin, S. Cohen, B. Price, and A. Yuille. Joint object and part segmentation using deep learned potentials. In ICCV, 2015. 4
  53. 53.T. Xiao, Y. Xu, K. Yang, J. Zhang, Y. Peng, and Z. Zhang. The application of two-level attention models in deep convolutional neural network for fine-grained image classification. In CVPR, 2015. 2
  54. 54.S. Xie and Z. Tu. Holistically-nested edge detection. In ICCV, 2015. 1, 2, 4
  55. 55.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv:1502.03044, 2015. 2, 3
  56. 56.L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville. Describing videos by exploiting temporal structure. In ICCV, 2015. 2, 3
  57. 57.D. Yoo, S. Park, J.-Y. Lee, A. S. Paek, and I. So Kweon. Attentionnet: Aggregating weak directions for accurate object detection. In ICCV, 2015. 2
  58. 58.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. Torr. Conditional random fields as recurrent neural networks. In ICCV, 2015. 1, 2, 3, 5, 7

Citation

MLA
Chen, L.-C., et al. “Attention to Scale: Scale-aware Semantic Image Segmentation”. arXiv, 2015, http://arxiv.org/abs/1511.03339v2.
APA
Chen, L.-C., Yang, Y., Wang, J., Xu, W., & Yuille, A. L. (2015). Attention to Scale: Scale-aware Semantic Image Segmentation. arXiv. http://arxiv.org/abs/1511.03339v2
Chicago
Chen, L.-C., Y. Yang, J. Wang, W. Xu, and A. L. Yuille. 2015. “Attention to Scale: Scale-aware Semantic Image Segmentation”. arXiv. http://arxiv.org/abs/1511.03339v2.
Harvard
Chen, L.-C. et al. (2015) “Attention to Scale: Scale-aware Semantic Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1511.03339v2.
Vancouver
1. Chen L-C, Yang Y, Wang J, Xu W, Yuille AL (2015) Attention to Scale: Scale-aware Semantic Image Segmentation. arXiv

BibTeX

@article{chen2015attention,
  title = {Attention to Scale: Scale-aware Semantic Image Segmentation},
  author = {Chen, Liang-Chieh and Yang, Yi and Wang, Jiang and Xu, Wei and Yuille, Alan L.},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1511.03339v2},
  eprint = {1511.03339}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE