Visual saliency based on multiscale deep features

Guanbin LiYizhou Yu

article2015CVPR1,360 citations

Introduces a multiscale deep convolutional framework for visual saliency detection that integrates spatial coherence refinement across multiple segmentation levels and establishes state-of-the-art accuracy alongside the 4,447-image HKU-IS benchmark dataset.

Listen

Visual saliency estimation aims to identify the regions in an image that naturally draw human attention, serving as a critical foundational component for applications such as image cropping, categorization, and object recognition. Traditional methods primarily rely on handcrafted visual features and heuristic assumptions, such as presuming that salient subjects are centrally located or distinct in color from image boundaries. However, these conventional approaches frequently fail in complex real-world scenes containing multiple focal points, low visual contrast, or cluttered backgrounds.

The article demonstrates that high-performance visual saliency models can be constructed using multiscale features extracted via deep convolutional neural networks. It designs and evaluates an integrated saliency framework that pairs multiscale deep representations with spatial coherence refinement and hierarchical segmentation fusion, while introducing a large-scale, challenging benchmark dataset to advance model evaluation.

The approach extracts feature representations from three nested visual scales—the target image region, its immediate neighborhood, and the entire image context—using a pre-trained deep convolutional neural network. These multiscale representations feed into fully connected neural layers that compute region saliency scores. To enhance boundary accuracy, the framework incorporates a spatial coherence refinement step based on edge strength and fuses saliency predictions across 15 hierarchical segmentation levels via linear regression. The authors evaluated the system across multiple standard datasets (including MSRA-B, SOD, and iCoSeg) and created the HKU-IS benchmark, a curated dataset of 4,447 challenging images with pixelwise annotations from multiple human reviewers.

The findings confirm that the proposed model substantially outperforms existing state-of-the-art methods across all evaluated benchmarks. On the new HKU-IS dataset, the model improved the overall F-measure score by 13.2% and reduced the mean absolute error by 35.1% relative to the best-performing existing methods. On the standard MSRA-B dataset, it achieved an 86.4% precision and 87.0% recall, raising the F-measure by 5.0% and lowering mean absolute error by 5.7%. Component evaluations established that combining all three nested context scales is essential for optimal performance, while spatial coherence refinement and multi-level segmentation fusion provided significant additive gains in accuracy.

These results demonstrate that deep neural networks can successfully model contrast and relative visual importance rather than just isolated object categories. By removing the reliance on rigid geometric assumptions like center-biases, the approach enhances the reliability of downstream computer vision pipelines in unconstrained visual environments. Furthermore, the model processes standard testing images in approximately 8 seconds, demonstrating practical feasibility for automated image processing workflows.

Organizations developing automated visual inspection, media processing, or robotic vision tools should transition away from handcrafted saliency heuristics toward deep multiscale contrast architectures. Future development should focus on optimizing runtime efficiency to enable real-time video processing and exploring end-to-end network training across segmentation layers.

Confidence in these findings is high given the extensive cross-validation and consistent performance across diverse standard datasets. Potential limitations include computational dependencies on multi-level superpixel segmentations during preprocessing and the initial 20-hour model training requirement, which may necessitate further refinement for resource-constrained or real-time deployment settings.

arXiv: 1503.08663
Cover for Visual saliency based on multiscale deep features

Abstract

Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this CVPR 2015 paper, we discover that a high-quality visual saliency model can be trained with multiscale features extracted using a popular deep learning architecture, convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for extracting features at three different scales. We then propose a refinement method to enhance the spatial coherence of our saliency results. Finally, aggregating multiple saliency maps computed for different levels of image segmentation can further boost the performance, yielding saliency maps better than those generated from a single segmentation. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotation. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks, improving the F-Measure by 5.0% and 13.2% respectively on the MSRA-B dataset and our new dataset (HKU-IS), and lowering the mean absolute error by 5.7% and 35.1% respectively on these two datasets.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 Saliency Inference with Deep Features
  • 2.1 Multiscale Feature Extraction
  • 2.2 Neural Network Training
  • 3 The Complete Algorithm
  • 3.1 Multi-Level Region Decomposition
  • 3.2 Spatial Coherence
  • 3.3 Saliency Map Fusion
  • 4 A New Dataset
  • 5 Experimental Results
  • 5.1 Dataset
  • 5.2 Evaluation Criteria
  • 5.3 Comparison with the State of the Art
  • 5.4 Component-wise Efficacy
  • References

Knowls

  1. Knowl 1 — Multiscale Deep Features (S-3CNN) for Visual Saliency Representation

    model/method

    To capture visual contrast at multiple spatial contexts, an image region RR is represented by extracting deep convolutional neural network (CNN) features from three nested rectangular bounding boxes using an 8-layer CNN (comprising five convolutional and three fully connected layers) pre-trained on ImageNet:

    1. Feature A (Local Region Appearance): Extracted from the bounding box of region RR. Pixels outside RR but inside the bounding box are filled with the mean pixel values computed over all ImageNet training images (becoming zero after mean subtraction). The box is warped to 227×227227 \times 227 pixels and forward-propagated to extract a 4096-dimensional activation vector from the second fully connected layer (extfc7 ext{fc}_7).
    2. Feature B (Neighborhood Contrast): Extracted from the bounding box enclosing region RR along with all its immediately adjacent neighboring regions, leaving internal pixel values intact. After warping to 227×227227 \times 227 pixels, extfc7 ext{fc}_7 activations yield a second 4096-dimensional vector capturing contrast against local surroundings.
    3. Feature C (Global Context and Location): Extracted from the entire image, where region RR is masked out with ImageNet mean pixel values to indicate its position. The full masked image is warped to 227×227227 \times 227 pixels to extract a third 4096-dimensional extfc7 ext{fc}_7 vector encoding global uniqueness and spatial location.

    The concatenation of these three feature vectors forms a 1228812288-dimensional descriptor, denoted as S-3CNN\text{S-3CNN}.

  2. Knowl 2 — Deep Neural Network Architecture for Region Saliency Prediction

    model/method

    Region-level saliency is predicted from the 1228812288-dimensional multiscale deep feature vector (S-3CNN\text{S-3CNN}) via a multi-layer neural network regressor:

    • Architecture: The network consists of the concatenated S-3CNN\text{S-3CNN} input, followed by two fully connected hidden layers containing 300300 neurons each, and an output layer that applies a two-way softmax function to output a probability distribution over binary saliency states (salient vs. non-salient).
    • Training Sample Selection: Training images are decomposed into non-overlapping regions. To ensure clean training supervision, only regions in which at least 70%70\% of constituent pixels share the same binary ground-truth saliency label are used as training samples, assigning them a region target of 11 or 00 accordingly.
    • Optimization: Network weights are trained by minimizing the least-squares prediction error accumulated over all selected regions across all training images.
    • Inference: The trained regressor is evaluated on every segmented region of an image to produce a scalar saliency score, which is subsequently assigned to all pixels inside that region.
  3. Knowl 3 — Graph-Based Spatial Coherence Energy Minimization for Saliency Refinement

    equation

    To eliminate boundary artifacts and noise from region-based predictions, an initial saliency score aiIa_i^I for superpixel PiP_i (computed as the mean saliency value over all pixels in PiP_i) is refined into an optimized score aiRa_i^R by minimizing the quadratic cost function:

    ∑i(aiR−aiI)2+∑i,jwij(aiR−ajR)2\sum_i \left(a_i^R - a_i^I\right)^2 + \sum_{i,j} w_{ij} \left(a_i^R - a_j^R\right)^2

    where wijw_{ij} is the spatial coherence weight between superpixels PiP_i and PjP_j. The problem reduces to solving a sparse linear system.

    The pairwise weight wijw_{ij} is defined on an undirected superpixel graph as:

    wij=exp⁡(−d2(Pi,Pj)2σ2)w_{ij} = \exp\left( -\frac{d^2(P_i, P_j)}{2\sigma^2} \right)

    where σ\sigma is the standard deviation of all pairwise graph distances. For adjacent superpixels PiP_i and PjP_j, the distance d(Pi,Pj)d(P_i, P_j) is computed from the Ultrametric Contour Map (UCM) edge strength ES(p)∈[0,1]ES(p) \in [0, 1] across their shared boundary:

    d(Pi,Pj)=∑p∈(ΩPi∩Pj∪Pi∩ΩPj)ES(p)∣ΩPi∩Pj∪Pi∩ΩPj∣d(P_i, P_j) = \frac{\sum_{p \in (\Omega_{P_i} \cap P_j \cup P_i \cap \Omega_{P_j})} ES(p)}{|\Omega_{P_i} \cap P_j \cup P_i \cap \Omega_{P_j}|}

    where ΩP\Omega_P denotes the outer boundary pixels of superpixel PP. For non-adjacent superpixels, d(Pi,Pj)d(P_i, P_j) is the shortest path distance on the graph.

  4. Knowl 4 — Hierarchical Multi-Level Image Decomposition and Saliency Map Fusion

    model/method

    To capture visual saliency at multiple granularities, an image is processed across M=15M = 15 hierarchical segmentation levels:

    1. Multi-Level Decomposition: Initial superpixels are generated via graph-based segmentation, and an Ultrametric Contour Map (UCM) region-merge tree is constructed. By varying the edge strength threshold, M=15M = 15 non-overlapping segmentation levels S={S1,S2,…,SM}S = \{S_1, S_2, \dots, S_M\} are produced. The target number of regions is set to 300300 at the finest level (S1S_1) and 2020 at the coarsest level (S15S_{15}), with region counts at intermediate levels following a geometric series.
    2. Map Generation and Refinement: For each segmentation level k∈{1,…,M}k \in \{1, \dots, M\}, region saliency prediction and spatial coherence optimization are applied independently to yield a refined saliency map A(k)A^{(k)}.
    3. Least-Squares Fusion: The final saliency map AA is formed as a linear combination of all MM maps:

    A=∑k=1MαkA(k)A = \sum_{k=1}^M \alpha_k A^{(k)}

    where the weights {αk}k=1M\{\alpha_k\}_{k=1}^M are learned on a validation dataset IvI_v by solving:

    {αk}k=1M=arg⁡min⁡α1,…,αM∑i∈Iv∥Ai−∑k=1MαkAi(k)∥F2\{\alpha_k\}_{k=1}^M = \arg\min_{\alpha_1, \dots, \alpha_M} \sum_{i \in I_v} \left\| A_i - \sum_{k=1}^M \alpha_k A_i^{(k)} \right\|_F^2

    with AiA_i denoting the ground-truth binary saliency map of validation image ii and ∥⋅∥F\|\cdot\|_F denoting the Frobenius norm.

  5. Knowl 5 — HKU-IS Visual Saliency Benchmark Dataset Construction

    experimental setup

    The HKU-IS benchmark dataset comprises 44474447 challenging natural images with pixelwise ground-truth binary annotations. Images were selected according to three explicit difficulty criteria designed to avoid common center/boundary biases:

    1. The image contains multiple disconnected salient objects (50.34%50.34\% of HKU-IS vs. 6.24%6.24\% of MSRA-B).
    2. At least one salient object touches the image boundary (21%21\% of HKU-IS vs. 13%13\% of MSRA-B).
    3. Color contrast—defined as the minimum Chi-square distance between the color histograms of any salient object and surrounding background—is <0.7< 0.7 (mean color contrast is 0.690.69 on HKU-IS vs. 0.780.78 on MSRA-B).

    Annotation and Quality Filtering: Three independent human annotators segmented each image using an interactive segmentation tool. For binary masks {a(1),a(2),a(3)}\{a^{(1)}, a^{(2)}, a^{(3)}\}, label consistency CC is computed as:

    C=∑x(∏p=13ax(p))∑x1(∑p=13ax(p)≠0)C = \frac{\sum_x \left(\prod_{p=1}^3 a_x^{(p)}\right)}{\sum_x \mathbf{1}\left(\sum_{p=1}^3 a_x^{(p)} \ne 0\right)}

    Images with C<0.9C < 0.9 were discarded from the initial pool of 73207320 images. Ground truth gx∈{0,1}g_x \in \{0, 1\} for each pixel xx of the remaining 44474447 images was set by majority vote:

    gx=1(∑p=13ax(p)≥2)g_x = \mathbf{1}\left(\sum_{p=1}^3 a_x^{(p)} \ge 2\right)

    The standard benchmark split designates 25002500 images for training, 500500 for validation, and 14471447 for testing.

  6. Knowl 6 — Salient Object Detection Benchmark Results for Multiscale Deep Features (MDF)

    empirical result

    The Multiscale Deep Features (MDF) method was evaluated against nine existing saliency algorithms (DRFI, wCtr*, MR, RC, HS, GS, SF, FT, SR) across multiple benchmarks using adaptive-threshold F-measure (FβF_\beta with β2=0.3\beta^2 = 0.3, thresholded at twice the image mean saliency Ta=2W×H∑x,yS(x,y)T_a = \frac{2}{W \times H} \sum_{x,y} S(x, y)) and Mean Absolute Error (MAE):

    • MSRA-B (2000 test images): MDF achieves 86.4%86.4\% precision and 87.0%87.0\% recall (compared to second-best MR at 84.8%84.8\% precision and 76.3%76.3\% recall), improving overall F-measure by 5.0%5.0\% and reducing MAE by 5.7%5.7\% relative to second-best wCtr*.
    • HKU-IS (1447 test images): MDF improves F-measure by 13.2%13.2\% over second-best DRFI (0.800.80 vs. 0.710.71), achieving a 9.0%9.0\% gain in precision, a 5.7%5.7\% gain in recall, and a 35.1%35.1\% reduction in MAE relative to wCtr*.
    • SOD and iCoSeg: MDF reduces MAE by 17.1%17.1\% on SOD and 26.3%26.3\% on iCoSeg compared to the respective runner-up methods.
    • Precision-recall curves demonstrate that MDF achieves the highest precision across almost the entire recall spectrum on MSRA-B, HKU-IS, SOD, and iCoSeg.
  7. Knowl 7 — Ablation of S-3CNN Multiscale Feature Components

    empirical result

    To evaluate the contribution of each scale within the S-3CNN\text{S-3CNN} descriptor, five alternative regressors were trained on MSRA-B under identical conditions using subsets of the features:

    • Single-scale models (Feature A only, Feature B only, Feature C only): Each single-scale model yielded substantially lower precision, recall, and adaptive F-measure across testing images. Feature A alone lacks contrast information against the background; Feature C alone lacks fine local boundary localization.
    • Two-scale combinations (Features A+B, Features A+C): Models trained on pairs of components performed noticeably better than any single-component model, confirming the synergy between local shape and contextual contrast.
    • Full S-3CNN (Features A+B+C): The concatenated descriptor achieved the highest precision, recall, and F-measure throughout the entire precision-recall curve, showing that region-level appearance, local neighborhood contrast, and global context are complementary.
  8. Knowl 8 — Ablation of Spatial Coherence Optimization and Multi-Level Segmentation Fusion

    empirical result

    Evaluations on the MSRA-B testing set isolate the individual benefits of spatial coherence optimization and multi-level fusion:

    • Spatial Coherence Refinement: Turning on graph-based spatial coherence optimization yields a consistent upward shift in precision-recall curves across individual segmentation layers (tested on the top three best-performing individual layers) as well as on the multi-level fused model, effectively filtering out segmentation noise and enforcing consistent scores within homogeneous areas.
    • Multi-Level Segmentation Fusion: Aggregating saliency maps across 1515 levels of hierarchical segmentation outperforms any single segmentation level on its own, improving average precision by 2.15%2.15\% and average recall by 3.47%3.47\% relative to the best-performing individual segmentation layer.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.R. Achanta, S. Hemami, F. Estrada, and S. Susstrunk. Frequency-tuned salient region detection. In CVPR, 2009. 2, 6, 7
  2. 2.S. Alpert, M. Galun, R. Basri, and A. Brandt. Image segmentation by probabilistic bottom-up aggregation and cue integration. In CVPR, 2007. 5
  3. 3.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik. Contour detection and hierarchical image segmentation. TPAMI, 33(5):898–916, 2011. 4
  4. 4.S. Avidan and A. Shamir. Seam carving for content-aware image resizing. ACM Trans. Graphics, 26(3), 2007. 1
  5. 5.D. Batra, A. Kowdle, D. Parikh, J. Luo, and T. Chen. icoseg: Interactive co-segmentation with intelligent scribble guidance. In CVPR, 2010. 6
  6. 6.A. Borji and L. Itti. State-of-the-art in visual attention modeling. TPAMI, 35(1):185–207, 2013. 1
  7. 7.K.-Y. Chang, T.-L. Liu, H.-T. Chen, and S.-H. Lai. Fusing generic objectness and visual saliency for salient object detection. In ICCV, 2011. 2
  8. 8.M.-M. Cheng, N. J. Mitra, X. Huang, P. H. S. Torr, and S.-M. Hu. Global contrast based salient region detection. TPAMI, 2014. 2, 6, 7
  9. 9.M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V. Vineet, and N. Crook. Efficient salient region detection with soft image abstraction. In ICCV, 2013. 2
  10. 10.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 1, 2, 3
  11. 11.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. arXiv preprint arXiv:1310.1531, 2013. 2
  12. 12.C. Farabet, C. Couprie, L. Najman, and Y. LeCun. Learning hierarchical features for scene labeling. TPAMI, 35(8):1915 – 1929, 2013. 1, 2
  13. 13.P. F. Felzenszwalb and D. P. Huttenlocher. Efficient graph-based image segmentation. IJCV, 59(2):167–181, 2004. 4
  14. 14.K. Fukushima. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological cybernetics, 36(4):193–202, 1980. 1
  15. 15.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 1, 2, 3
  16. 16.S. Goferman, L. Zelnik-Manor, and A. Tal. Context-aware saliency detection. TPAMI, 34(10):1915–1926, 2012. 2
  17. 17.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Simultaneous detection and segmentation. In ECCV. 1
  18. 18.X. Hou and L. Zhang. Saliency detection: A spectral residual approach. In CVPR, 2007. 2, 6, 7
  19. 19.L. Itti, C. Koch, and E. Niebur. A model of saliency-based visual attention for rapid scene analysis. TPAMI, 20(11):1254–1259, 1998. 2
  20. 20.Y. Jia and M. Han. Category-independent object-level saliency detection. In ICCV, 2013. 2
  21. 21.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014. 3
  22. 22.H. Jiang, J. Wang, Z. Yuan, Y. Wu, N. Zheng, and S. Li. Salient object detection: A discriminative regional feature integration approach. In CVPR, 2013. 2, 4, 5, 6, 7
  23. 23.T. Judd, K. Ehinger, F. Durand, and A. Torralba. Learning to predict where humans look. In ICCV, 2009. 2
  24. 24.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, 2012. 1, 2
  25. 25.R. Liu, J. Cao, Z. Lin, and S. Shan. Adaptive partial differential equation learning for visual saliency detection. In CVPR, 2014. 2
  26. 26.T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum. Learning to detect a salient object. TPAMI, 33(2):353–367, 2011. 2, 5
  27. 27.L. Mai, Y. Niu, and F. Liu. Saliency aggregation: A data-driven approach. In CVPR, 2013. 5
  28. 28.D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001. 5
  29. 29.F. Perazzi, P. Krahenbuhl, Y. Pritch, and A. Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, 2012. 6, 7
  30. 30.A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. arXiv preprint arXiv:1403.6382, 2014. 2
  31. 31.C. Rother, L. Bordeaux, Y. Hamadi, and A. Blake. Autocollage. ACM Trans. Graphics, 25(3):847–852, 2006. 1
  32. 32.U. Rutishauser, D. Walther, C. Koch, and P. Perona. Is bottom-up attention useful for object recognition? In CVPR, 2004. 1
  33. 33.X. Shen and Y. Wu. A unified approach to salient object detection via low rank matrix recovery. In CVPR, 2012. 2
  34. 34.D. Simakov, Y. Caspi, E. Shechtman, and M. Irani. Summarizing visual data using bidirectional similarity. In CVPR, 2008. 1
  35. 35.Y. Wei, F. Wen, W. Zhu, and J. Sun. Geodesic saliency using background priors. In ECCV. 2012. 2, 6, 7
  36. 36.R. Wu, Y. Yu, and W. Wang. Scale: Supervised and cascaded laplacian eigenmaps for visual object recognition based on nearest neighbors. In CVPR, 2013. 1
  37. 37.Q. Yan, L. Xu, J. Shi, and J. Jia. Hierarchical saliency detection. In CVPR, 2013. 6, 7, 8
  38. 38.C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang. Saliency detection via graph-based manifold ranking. In CVPR, 2013. 6, 7
  39. 39.R. Zhao, W. Ouyang, and X. Wang. Unsupervised salience learning for person re-identification. In CVPR, 2013. 1
  40. 40.W. Zhu, S. Liang, Y. Wei, and J. Sun. Saliency optimization from robust background detection. In CVPR, 2014. 2, 5, 6, 7, 8

Citation

MLA
Li, G., and Y. Yu. “Visual Saliency Based on Multiscale Deep Features”. arXiv, 2015, http://arxiv.org/abs/1503.08663v3.
APA
Li, G., & Yu, Y. (2015). Visual Saliency Based on Multiscale Deep Features. arXiv. http://arxiv.org/abs/1503.08663v3
Chicago
Li, G., and Y. Yu. 2015. “Visual Saliency Based on Multiscale Deep Features”. arXiv. http://arxiv.org/abs/1503.08663v3.
Harvard
Li, G. and Yu, Y. (2015) “Visual Saliency Based on Multiscale Deep Features”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1503.08663v3.
Vancouver
1. Li G, Yu Y (2015) Visual Saliency Based on Multiscale Deep Features. arXiv

BibTeX

@article{li2015visual,
  title = {Visual Saliency Based on Multiscale Deep Features},
  author = {Li, Guanbin and Yu, Yizhou},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1503.08663v3},
  eprint = {1503.08663}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE