Salient Object Detection: A Benchmark

Ali BorjiMing-Ming ChengHuaizu JiangJia Li

article2015IEEE Transactions on Image Processing1,841 citations

Establishes a comprehensive benchmark by evaluating forty models across six datasets, analyzing factors like center bias and scene complexity to expose key failure modes and guide future salient object detection research.

Listen

Visual attention systems enable computer vision applications—such as autonomous robotics, image editing, compression, and retrieval—to automatically identify and isolate primary foreground objects in cluttered real-world images. However, the field has suffered from ambiguous problem definitions, conflicting performance evaluations, and untested model progression across varied image datasets.

The article establishes a large-scale standardized benchmark to rigorously assess the progress, accuracy, and operational efficiency of single-image salient object detection models against related visual attention approaches.

The authors conducted a comprehensive empirical evaluation comparing 41 computational models (29 dedicated salient object detection models, 10 human fixation prediction models, 1 generic object proposal method, and 1 location baseline) across 7 diverse public image datasets totaling tens of thousands of images. Performance was benchmarked across multiple standard evaluation criteria, including precision-recall curves, weighted harmonic mean accuracy scores, receiver operating characteristics, mean absolute error, and image processing runtimes.

The benchmark demonstrates five key findings. First, dedicated salient object detection algorithms significantly outperformed traditional fixation prediction and object proposal models, demonstrating that segmenting full foreground objects requires distinct computational formulations. Second, data-driven regional methods (most notably DRFI, followed by models like QCUT and RBD) established the highest overall accuracy across the benchmark datasets. Third, regional superpixel groupings and boundary background assumptions proved to be the most effective design elements, outperforming pixel-level operations in both boundary delineation and processing speed. Fourth, performance dropped sharply across all algorithms when handling cluttered scenes, small objects, or off-center subjects, revealing a widespread over-reliance on center biases. Fifth, processing speeds varied by several orders of magnitude, ranging from roughly 0.017 seconds to over 100 seconds per standard image.

These findings indicate that real-world deployment risks can be substantially reduced by prioritizing region-based, learning-driven models that combine high segmentation precision with sub-second execution speeds. Practitioners must note that relying on algorithms tuned solely for centered, single-object images poses significant operational failure risks in complex, real-world environments with low-contrast boundaries or multiple focal points.

To move the field forward, development efforts should transition toward robust deep learning architectures and train models on complex, multi-object images without positional biases. Future benchmarks must also incorporate images lacking salient targets entirely, while expanding into multi-image domains such as video streams and multi-view datasets.

While the empirical findings are highly reliable for single-image static processing, caution is advised when generalizing these performance rankings to real-time video, multi-camera setups, or active robotic vision systems that fall outside this study's single-frame benchmark scope.

  • Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). Cheng et al. introduced foundational global contrast-based salient region detection algorithms and benchmark datasets that this survey directly evaluates and builds upon.
  • Paper: State-of-the-Art in Visual Attention Modeling, Ali Borji et al. (2013). This comprehensive taxonomy of visual attention models establishes the theoretical baseline and metrics for distinguishing eye-fixation prediction from salient object detection.
  • Paper: Graph-Based Visual Saliency, Jonathan Harel et al. (2006). The Graph-Based Visual Saliency framework is a classical fixation prediction baseline that the benchmark contrasts against object-level saliency detectors.
  • Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, L. Itti et al. (1998). Itti, Koch, and Niebur formulate the seminal bottom-up computational architecture of visual saliency that underpins the entire field evaluated in the benchmark.
  • Paper: Unbiased look at dataset bias, A. Torralba et al. (2011). Torralba and Efros provide the critical methodology for analyzing dataset bias and cross-dataset generalization that the benchmark adapts to salient object detection datasets.
Cover for Salient Object Detection: A Benchmark

Abstract

We extensively compare, qualitatively and quantitatively, 40 state-of-the-art models (28 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over 6 challenging datasets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted just two years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for state-of-the-art models, provide useful hints towards constructing more challenging large scale datasets and better saliency models. Finally, we propose probable solutions for tackling several open problems such as evaluation scores and dataset bias, which also suggest future research directions in the rapidly-growing field of salient object detection.

Table of Contents

  • I Introduction
  • II Salient Object Detection Benchmark
  • II-A Compared Models
  • II-B Datasets
  • II-C Evaluation Measures
  • II-D Quantitative Comparison of Models
  • II-E Qualitative Comparison of Models
  • III Performance Analysis
  • III-A Analysis of Segmentation Methods
  • III-B Analysis of Center Bias
  • III-C Analysis of Salient Object Existence
  • III-D Analysis of Worst and Best Cases for Top Models
  • III-E Runtime Analysis
  • IV Discussions and Conclusions
  • References

Knowls

  1. Knowl 1 — Comprehensive Benchmark Performance and Ranking of Salient Object Detection Models

    data/table

    Across an extensive benchmark of 41 models (29 salient object detection models, 10 human fixation prediction models, 1 object proposal model, and 1 baseline) evaluated over multiple datasets (MSRA10K, THUR15K, ECSSD, JuddDB, DUT-OMRON, and PASCAL-S), models designed specifically for salient object detection consistently and significantly outperform fixation prediction models and generic object proposal methods.

    Model FβF_\beta Rank FβwF_\beta^w Rank AUC Rank MAE Rank AdpT Rank SCut Rank Center-Bias Rank Overall Rank
    DRFI 1 3 1 4 2 2 1 1
    QCUT 2 6 4 1 1 8 5 2
    RBD 5 1 5 3 6 3 4 3
    ST 3 2 3 7 7 1 7 4
    DSR 6 4 2 2 3 21 10 5
    MC 4 10 6 12 4 15 11 6
    GMR 7 5 11 10 5 17 12 7
    CHM 10 12 8 5 9 10 15 8
    HDCT 9 11 9 6 8 6 2 9
    HS 8 7 12 19 16 7 3 10
    BMS (Fixation) 15 14 16 13 10 23 17 11
    RC 11 9 17 14 13 14 16 13
    OBJ (Objectness) 27 21 24 36 29 20 31 26
    AAM (Baseline) 30 23 25 33 25 22 41 24
    IT (Fixation) 38 41 41 21 31 41 39 40
    AC 37 39 38 22 40 39 34 41

    The overall rank is calculated based on the mean of AUC, 1−MAE1 - \text{MAE}, maximal FβF_\beta, adaptive thresholding score (extAdpT ext{AdpT}), SaliencyCut score (extSCut ext{SCut}), and weighted F-measure (FβwF_\beta^w). DRFI consistently achieves the highest rank across area under the ROC curve (AUC), maximum FβF_\beta, and robustness to center-bias variations. The top 6 models are DRFI, QCUT, RBD, ST, DSR, and MC.

  2. Knowl 2 — Definition and Characterization of Salient Object Detection

    definition

    Salient Object Detection (SOD) is formulated as a figure/ground segmentation task that aims to selectively process visual scenes through two primary sequential objectives:

    1. Salient Object Identification: Automatically detect attention-grabbing visual objects or regions in a single scene.
    2. Whole Object Segmentation: Uniformly highlight and segment the entire foreground object region from the background clutter.

    The output of a salient object detection model is a continuous 2D saliency map S∈[0,1]W×HS \in [0, 1]^{W \times H} (or scaled to [0,255][0, 255]), where the intensity of each pixel (x,y)(x, y) represents its probability of belonging to a salient foreground object.

    Salient object detection differs distinctly from related visual attention tasks:

    • Human Fixation Prediction: Aims to predict human eye gaze locations, yielding sparse, blob-like saliency peaks rather than solid, uniformly highlighted objects.
    • Object Proposal Generation: Generates category-independent candidate bounding boxes or candidate regions without evaluating relative visual saliency.
    • Traditional Image Segmentation: Partitions an image into perceptually coherent and homogeneous regions without separating semantic salient foreground from non-salient background.
  3. Knowl 3 — Quantitative Evaluation Metrics for Salient Object Detection

    equation

    Salient object detection models are evaluated by comparing a predicted saliency map S(x,y)∈[0,255]S(x, y) \in [0, 255] with a ground-truth binary foreground mask G(x,y)∈{0,1}G(x, y) \in \{0, 1\} of dimensions W×HW \times H. Given a binarized prediction mask M∈{0,1}W×HM \in \{0, 1\}^{W \times H} and letting ∣⋅∣|\cdot| denote the count of non-zero entries:

    Precision=∣M∩G∣∣M∣,Recall=∣M∩G∣∣G∣\text{Precision} = \frac{|M \cap G|}{|M|}, \qquad \text{Recall} = \frac{|M \cap G|}{|G|}

    TPR=∣M∩G∣∣G∣,FPR=∣M∩Gˉ∣∣Gˉ∣\text{TPR} = \frac{|M \cap G|}{|G|}, \qquad \text{FPR} = \frac{|M \cap \bar{G}|}{|\bar{G}|}

    where Gˉ\bar{G} is the spatial complement of the ground truth (background pixels).

    Adaptive Binarization Threshold (TaT_a): Ta=2W×H∑x=1W∑y=1HS(x,y)T_a = \frac{2}{W \times H} \sum_{x=1}^W \sum_{y=1}^H S(x, y)

    F-Measure (FβF_\beta): Weighted harmonic mean of Precision and Recall parameterized with β2=0.3\beta^2 = 0.3 to weight precision more heavily than recall: Fβ=(1+β2)Precision×Recallβ2Precision+RecallF_\beta = \frac{(1 + \beta^2) \text{Precision} \times \text{Recall}}{\beta^2 \text{Precision} + \text{Recall}}

    Mean Absolute Error (MAE): Measures the continuous deviation between the normalized saliency map Sˉ(x,y)∈[0,1]\bar{S}(x, y) \in [0, 1] and normalized ground truth Gˉ(x,y)∈{0,1}\bar{G}(x, y) \in \{0, 1\}: MAE=1W×H∑x=1W∑y=1H∣Sˉ(x,y)−Gˉ(x,y)∣\text{MAE} = \frac{1}{W \times H} \sum_{x=1}^W \sum_{y=1}^H |\bar{S}(x, y) - \bar{G}(x, y)|

    While the Area Under the ROC Curve (AUC) measures the area under the True Positive Rate (TPR) versus False Positive Rate (FPR) curve across fixed thresholds Tf∈[0,255]T_f \in [0, 255], Precision-Recall curves and FβF_\beta are more informative because background pixels heavily outnumber foreground pixels in natural scenes.

  4. Knowl 4 — Algorithmic and Structural Design Factors of Top-Performing Saliency Models

    model/method

    Analysis of top-performing salient object detection methods (including DRFI, DSR, RBD, MC, and ST) demonstrates three critical architectural design choices that drive superior detection accuracy:

    1. Superpixel/Region-Based Primitives: High-ranking models operate on superpixel segmentations rather than individual pixels or rectangular image patches. Regions enable the extraction of expressive, multi-scale statistical features (such as regional color histograms and texture descriptors), preserve sharp object boundaries, and substantially lower computational time by reducing the primitive count.
    2. Explicit Background Border Prior: Top models incorporate the prior that image boundary margins belong to the background. Unlike an object center prior, the border background prior is robust across variations in object location and scale.
    3. Data-Driven Discriminative Feature Integration: DRFI, the leading method, trains a regression model (Random Forest) over a 93-dimensional feature vector combining regional appearance, contrast, and boundary properties using annotated training data. Learning integration rules automatically from human annotations outperforms heuristic, hand-crafted cue combination schemes.
  5. Knowl 5 — Impact of and Model Robustness to Center Bias

    empirical result

    The performance of salient object detection models is strongly affected by center bias (the tendency of photographer-captured objects to lie near the image center). The Average Annotation Map (AAM) baseline achieves high AUC and FβF_\beta on strongly center-biased datasets (e.g., MSRA10K, ECSSD), outperforming multiple published models solely via positional bias.

    When evaluated on an off-center test subset of 1,000 images from MSRA10K (where object centroids have a normalized distance to image center >0.247> 0.247) and the two-object SED2 dataset:

    • The baseline AAM model drops dramatically (maximal Fβ=0.328F_\beta = 0.328 on off-center MSRA10K vs. 0.5800.580 on full MSRA10K; AAM is the lowest-performing model on SED2).
    • Models that explicitly rely heavily on center priors (e.g., Context-Based Saliency CB) suffer severe drops; CB drops by 0.1220.122 in maximal FβF_\beta and 0.1220.122 in AUC.
    • Robust models leveraging background border priors and regional contrast (DRFI and DSR) maintain top performance, with DRFI's FβF_\beta, AUC, and MAE shifting by only 0.050.05, 0.050.05, and 0.0090.009 respectively.
  6. Knowl 6 — Comparative Evaluation of Post-Processing Segmentation Strategies for Saliency Maps

    empirical result

    Converting continuous saliency maps into discrete binary object masks was evaluated across three strategies:

    1. Fixed Thresholding: Sweeping fixed thresholds Tf∈[0,255]T_f \in [0, 255] to extract the maximal FβF_\beta along the Precision-Recall curve.
    2. Image-Dependent Adaptive Thresholding: Binarizing using Ta=2×mean(S)T_a = 2 \times \text{mean}(S).
    3. SaliencyCut: Initializing an iterative GrabCut energy minimization with a loose saliency threshold to enforce local graph-based label consistency and appearance modeling.

    Empirical Findings:

    • On datasets with single dominant objects (MSRA10K, ECSSD, THUR15K, DUT-OMRON), SaliencyCut combined with top saliency maps (e.g., DRFI, RBD) achieves the highest segmentation performance (Fβ=0.905F_\beta = 0.905 for DRFI on MSRA10K via SaliencyCut vs. 0.8380.838 via adaptive thresholding).
    • On datasets with multiple salient objects, cluttered scenes, or off-center objects (SED2, JuddDB, PASCAL-S), SaliencyCut performs worse than thresholding methods because its formulation enforces a single dominant salient object assumption.
  7. Knowl 7 — Dataset Complexity Taxonomy and Characterization for Saliency Benchmarks

    experimental setup

    Seven benchmark datasets were characterized along three quantitative properties:

    1. Positional / Center Bias: Quantified by the distribution of normalized Euclidean distances from salient object centroids to the image center. Objects in ECSSD are closest to the image center (highest center bias), whereas objects in SED2 and JuddDB exhibit the largest dispersion from the center.
    2. Structural Clutter and Scene Complexity: Quantified by applying the Felzenszwalb-Huttenlocher graph segmentation algorithm and counting the resulting superpixels. High superpixel counts in foreground and background indicate high structural clutter. JuddDB (average of 493 superpixels in background regions) and PASCAL-S are the most complex datasets, whereas SED2 has the lowest superpixel count.
    3. Normalized Object Size: Saliency mask area normalized by total image area. MSRA10K and ECSSD contain larger foreground objects, while SED2 and JuddDB contain smaller objects.
  8. Knowl 8 — Failure of Salient Object Detection Models on Background-Only Images

    limitation

    Virtually all salient object detection models operate under the implicit assumption that every input image contains at least one dominant salient object.

    When evaluated on background-only images (scenes consisting purely of natural textures, repetitive patterns, or scene clutter without distinct semantic objects):

    • Top models (including DRFI, DSR, and MC) fail to produce null (all-zero) saliency maps and instead generate high-intensity false positive activations on high-contrast background texture regions.
    • Standard quantitative metrics (Precision-Recall curves, FβF_\beta, and ROC AUC) cannot be computed on background-only images because the true positive ground-truth pixel set is empty (∣G∣=0|G| = 0).
    • Mean Absolute Error (MAE) is uninformative for this failure mode due to the standard practice of min-max normalizing output saliency maps to [0,255][0, 255] during post-processing.
  9. Knowl 9 — Metric Disparity Between Precision-Recall and ROC Curves in Saliency Evaluation

    theoretical result

    In salient object detection, the number of non-salient background pixels (negative class Gˉ\bar{G}) is substantially larger than the number of salient foreground pixels (positive class GG).

    Because the False Positive Rate denominator is the total negative count: FPR=∣M∩Gˉ∣∣Gˉ∣\text{FPR} = \frac{|M \cap \bar{G}|}{|\bar{G}|} large increases in false positive detections ∣M∩Gˉ∣|M \cap \bar{G}| yield only small numerical increases in FPR\text{FPR}, causing ROC curves to appear compressed toward the upper-left axis and producing unrealistically high AUC scores.

    In contrast, Precision directly penalizes false positive assignments against total predicted positive pixels: Precision=∣M∩G∣∣M∣=∣M∩G∣∣M∩G∣+∣M∩Gˉ∣\text{Precision} = \frac{|M \cap G|}{|M|} = \frac{|M \cap G|}{|M \cap G| + |M \cap \bar{G}|} Consequently, Precision-Recall curves and FβF_\beta-scores provide a significantly more sensitive and reliable evaluation metric for distinguishing algorithm performance.

  10. Knowl 10 — Inference Runtime Efficiency vs. Accuracy Trade-Offs

    empirical result

    Evaluation of model execution time across all 10,000 images of MSRA10K (typical resolution 400×300400 \times 300) on an Intel Xeon E5645 2.40 GHz CPU with 8 GB RAM categorizes models into three operational tiers:

    1. High-Speed Heuristic Models: Histogram Contrast (HC) is the fastest at 0.017 seconds/image, followed by Global Contrast (GC) at 0.037s, Spectral Residual (SR) at 0.040s, Image Signature (SS) at 0.053s, and Frequency-Tuned (FT) at 0.072s.
    2. High-Accuracy Balanced Models: Region-based top performers maintain practical runtimes, including Global Manifold Ranking (GMR) at 0.149s, Markov Chain (MC) at 0.195s, Robust Background Detection (RBD) at 0.269s, and Discriminative Regional Feature Integration (DRFI) at 0.697s.
    3. Computationally Intensive Models: Several models require tens to hundreds of seconds per image, including Context-Aware Saliency (CA) at 40.9s, Spatial-Visual Objectness (SVO) at 56.5s, Saliency Tree (ST) at 79.1s, and Low/Mid-Level Cues (LMLC) at 140.0s, and Layered Boundary Information (LBI) at 251.0s.

    RBD, MC, and DRFI offer the most effective trade-offs between weighted F-measure (FβwF_\beta^w) accuracy and computational cost.

Coverage note — Omitted peripheral discussions of related multi-image saliency applications (video saliency, RGB-D saliency, and co-saliency benchmarks) and general vision applications (such as image retargeting, visual tracking, and image captioning), which are cited in the review sections rather than evaluated within the paper's benchmark experiments.

References

  1. 1.A. Borji, D. N. Sihite, and L. Itti, "Salient object detection: A benchmark," in ECCV, 2012, pp. 414–429.
  2. 2.A. Borji and L. Itti, "State-of-the-art in visual attention modeling," IEEE TPAMI, vol. 35, no. 1, pp. 185–207, 2013.
  3. 3.A. Borji, D. Sihite, and L. Itti, "Quantitative analysis of human-model agreement in visual saliency modeling: A comparative study," IEEE TIP, vol. 22, no. 1, pp. 55–69, 2013.
  4. 4.M. Hayhoe and D. Ballard, "Eye movements in natural behavior," Trends in cognitive sciences, pp. 188–194, 2005.
  5. 5.L. Itti and C. Koch, "Computational modelling of visual attention," Nature reviews neuroscience, vol. 2, no. 3, pp. 194–203, 2001.
  6. 6.A. M. Treisman and G. Gelade, "A feature-integration theory of attention," Cognitive Psychology, pp. 97–136, 1980.
  7. 7.J. M. Wolfe, K. R. Cave, and S. L. Franzel, "Guided search: an alternative to the feature integration model for visual search." J. Exp. Psychol. Human., vol. 15, no. 3, p. 419, 1989.
  8. 8.J. M. Wolfe, "Guidance of visual search by preattentive information," in Neurobiology of Attention, 2005, pp. 101–104.
  9. 9.C. Koch and S. Ullman, "Shifts in selective visual attention: towards the underlying neural circuitry," in Matters of Intelligence, 1987, pp. 115–141.
  10. 10.L. Itti, C. Koch, and E. Niebur, "A model of saliency-based visual attention for rapid scene analysis," IEEE TPAMI, 1998.
  11. 11.D. Parkhurst, K. Law, and E. Niebur, "Modeling the role of salience in the allocation of overt visual attention," Vision Research, vol. 42, no. 1, pp. 107–123, 2002.
  12. 12.J. Li, Y. Tian, T. Huang, and W. Gao, "Probabilistic multi-task learning for visual saliency estimation in video," IJCV, vol. 90, no. 2, pp. 150–165, Nov. 2010.
  13. 13.A. Borji and L. Itti, "Exploiting local and global patch rarities for saliency detection," in IEEE CVPR, 2012, pp. 478–485.
  14. 14.A. Borji, "Boosting bottom-up and top-down visual features for saliency estimation," in IEEE CVPR, 2012, pp. 438–445.
  15. 15.K. Koehler, F. Guo, S. Zhang, and M. P. Eckstein, "What do saliency models predict?" J. Vision, 2014.
  16. 16.J. Li, Y. Tian, and T. Huang, "Visual saliency with statistical priors," IJCV, vol. 107, no. 3, pp. 239–253, 2014.
  17. 17.T. Liu, J. Sun, N. Zheng, X. Tang, and H.-Y. Shum, "Learning to detect a salient object," in IEEE CVPR, 2007, pp. 1–8.
  18. 18.R. Achanta, S. Hemami, F. Estrada, and S. Süsstrunk, "Frequency-tuned salient region detection," in IEEE CVPR, 2009.
  19. 19.Y. Tian, J. Li, S. Yu, and T. Huang, "Learning complementary saliency priors for foreground object segmentation in complex scenes," IJCV, 2014.
  20. 20.J. Wang, L. Quan, J. Sun, X. Tang, and H.-Y. Shum, "Picture collage," in IEEE CVPR, vol. 1, 2006, pp. 347–354.
  21. 21.A. Borji, D. N. Sihite, and L. Itti, "What stands out in a scene? a study of human explicit saliency judgment," Vision Research, vol. 91, no. 0, pp. 62–77, 2013.
  22. 22.U. Rutishauser, D. Walther, C. Koch, and P. Perona, "Is bottom-up attention useful for object recognition?" in IEEE CVPR, 2004.
  23. 23.C. Kanan and G. Cottrell, "Robust classification of objects, faces, and flowers using natural image statistics," in IEEE CVPR, 2010, pp. 2472–2479.
  24. 24.F. Moosmann, D. Larlus, and F. Jurie, "Learning saliency maps for object categorization," in ECCV Workshop, 2006.
  25. 25.A. Borji, M. N. Ahmadabadi, and B. N. Araabi, "Cost-sensitive learning of top-down modulation for attentional control," Machine Vision and Applications, 2011.
  26. 26.A. Borji and L. Itti, "Scene classification with a sparse set of salient regions," in IEEE ICRA, 2011, pp. 1902–1908.
  27. 27.H. Shen, S. Li, C. Zhu, H. Chang, and J. Zhang, "Moving object detection in aerial video based on spatiotemporal saliency," Chinese Journal of Aeronautics, 2013.
  28. 28.A. Borji, M.-M. Cheng, H. Jiang, and J. Li, "Salient object detection: A survey," arXiv preprint arXiv:1411.5878, 2014.
  29. 29.Z. Ren, S. Gao, L.-T. Chia, and I. Tsang, "Region-based saliency detection and its application in object recognition," IEEE TCSVT, vol. PP, no. 99, pp. 1–1, 2013.
  30. 30.M. Guo, Y. Zhao, C. Zhang, and Z. Chen, "Fast object detection based on selective visual attention," Neurocomputing, vol. 144, pp. 184–197, 2014.
  31. 31.C. Guo and L. Zhang, "A novel multiresolution spatiotemporal saliency detection model and its applications in image and video compression," IEEE TIP, 2010.
  32. 32.L. Itti, "Automatic foveation for video compression using a neurobiological model of visual attention," IEEE TIP, 2004.
  33. 33.Y.-F. Ma, X.-S. Hua, L. Lu, and H.-J. Zhang, "A generic framework of user attention model and its application in video summarization," IEEE TMM, 2005.
  34. 34.Y. J. Lee, J. Ghosh, and K. Grauman, "Discovering important people and objects for egocentric video summarization," in IEEE CVPR, 2012, pp. 1346–1353.
  35. 35.Q.-G. Ji, Z.-D. Fang, Z.-H. Xie, and Z.-M. Lu, "Video abstraction based on the visual attention model and online clustering," Signal Processing: Image Communication, 2012.
  36. 36.S. Goferman, A. Tal, and L. Zelnik-Manor, "Puzzle-like collage," Computer Graphics Forum, 2010.
  37. 37.H. Huang, L. Zhang, and H.-C. Zhang, "Arcimboldo-like collage using internet images," ACM TOG, vol. 30, no. 6, p. 155, 2011.
  38. 38.A. Ninassi, O. Le Meur, P. Le Callet, and D. Barbba, "Does where you gaze on an image affect your perception of quality? applying visual attention to image quality metric," in IEEE ICIP, 2007.
  39. 39.H. Liu and I. Heynderickx, "Studying the added value of visual attention in objective image quality metrics based on eye movement data," in IEEE ICIP, 2009, pp. 3097–3100.
  40. 40.A. Li, X. She, and Q. Sun, "Color image quality assessment combining saliency and fsim," in ICDIP, vol. 8878, 2013.
  41. 41.W. Zhang, A. Borji, Z. Wang, P. Le Callet, and H. Liu, "The application of visual saliency models in objective image quality assessment: A statistical evaluation," IEEE Trans. on Neural Networks and Learning Systems, 2015.
  42. 42.M. Donoser, M. Urschler, M. Hirzer, and H. Bischof, "Saliency driven total variation segmentation," in IEEE ICCV, 2009.
  43. 43.Q. Li, Y. Zhou, and J. Yang, "Saliency based image segmentation," in ICMT, 2011, pp. 5068–5071.
  44. 44.C. Qin, G. Zhang, Y. Zhou, W. Tao, and Z. Cao, "Integration of the saliency-based seed extraction and random walks for image segmentation," Neurocomputing, vol. 129, 2013.
  45. 45.M. Johnson-Roberson, J. Bohg, M. Bjorkman, and D. Kragic, "Attention-based active 3d point cloud segmentation," in IEEE IROS, 2010, pp. 1165–1170.
  46. 46.T. Chen, M.-M. Cheng, P. Tan, A. Shamir, and S.-M. Hu, "Sketch2photo: internet image montage," ACM TOG, 2009.
  47. 47.S. Feng, D. Xu, and X. Yang, "Attention-driven salient edge (s) and region (s) extraction with application to CBIR," Signal Processing, vol. 90, no. 1, pp. 1–15, 2010.
  48. 48.J. Sun, J. Xie, J. Liu, and T. Sikora, "Image adaptation and dynamic browsing based on two-layer saliency combination," IEEE Trans. Broadcasting, vol. 59, no. 4, pp. 602–613, 2013.
  49. 49.L. Li, S. Jiang, Z. Zha, Z. Wu, and Q. Huang, "Partial-duplicate image retrieval via saliency-guided visually matching," IEEE MultiMedia, vol. 20, no. 3, pp. 13–23, 2013.
  50. 50.A. Y.-S. Chia, S. Zhuo, R. K. Gupta, Y.-W. Tai, S.-Y. Cho, P. Tan, and S. Lin, "Semantic colorization with internet images," ACM TOG, vol. 30, no. 6, p. 156, 2011.
  51. 51.H. Liu, L. Zhang, and H. Huang, "Web-image driven best views of 3d shapes," The Visual Computer, 2012.
  52. 52.R. Margolin, L. Zelnik-Manor, and A. Tal, "Saliency for image manipulation," The Visual Computer, pp. 1–12, 2013.
  53. 53.C. Goldberg, T. Chen, F.-L. Zhang, A. Shamir, and S.-M. Hu, "Data-driven object manipulation in images," Computer Graphics Forum, vol. 31, pp. 265–274, 2012.
  54. 54.S. Stalder, H. Grabner, and L. Van Gool, "Dynamic objectness for adaptive tracking," in ACCV, 2012.
  55. 55.J. Li, M. Levine, X. An, X. Xu, and H. He, "Visual saliency based on scale-space analysis in the frequency domain," IEEE TPAMI, vol. 35, no. 4, pp. 996–1010, 2013.
  56. 56.G. M. García, D. A. Klein, J. Stuckler, S. Frintrop, and A. B. Cremers, "Adaptive multi-cue 3d tracking of arbitrary objects," in Pattern Recognition, 2012, pp. 357–366.
  57. 57.A. Borji, S. Frintrop, D. N. Sihite, and L. Itti, "Adaptive object tracking by learning background context," in IEEE CVPR, 2012.
  58. 58.D. A. Klein, D. Schulz, S. Frintrop, and A. B. Cremers, "Adaptive real-time video-tracking for arbitrary objects," in IEEE IROS, 2010, pp. 772–777.
  59. 59.S. Frintrop and M. Kessel, "Most salient region tracking," in IEEE ICRA, 2009, pp. 1869–1874.
  60. 60.G. Zhang, Z. Yuan, N. Zheng, X. Sheng, and T. Liu, "Visual saliency based object tracking," in ACCV, 2010.
  61. 61.A. Karpathy, S. Miller, and L. Fei-Fei, "Object discovery in 3d scenes via shape analysis," in ICRA, 2013, pp. 2088–2095.
  62. 62.S. Frintrop, G. M. García, and A. B. Cremers, "A cognitive approach for object discovery," in IEEE ICPR, 2014.
  63. 63.D. Meger, P.-E. Forssen, K. Lai, S. Helmer, S. McCann, T. Southey, M. Baumann, J. J. Little, and D. G. Lowe, "Curious george: An attentive semantic robot," Robotics and Autonomous Systems, vol. 56, no. 6, pp. 503–511, 2008.
  64. 64.Y. Sugano, Y. Matsushita, and Y. Sato, "Calibration-free gaze sensing using saliency maps," in IEEE CVPR, 2010.
  65. 65.A. Borji and L. Itti, "Defending yarbus: Eye movements reveal observers' task," J. Vision, vol. 14, no. 3, p. 29, 2014.
  66. 66.R. Achanta, F. Estrada, P. Wils, and S. Süsstrunk, "Salient region detection and segmentation," in Comp. Vis. Sys., 2008.
  67. 67.S. Goferman, L. Zelnik-Manor, and A. Tal, "Context-aware saliency detection," IEEE TPAMI, vol. 34, no. 10, pp. 1915–1926, 2012.
  68. 68.R. Achanta and S. Süsstrunk, "Saliency detection using maximum symmetric surround," in IEEE ICIP, 2010, pp. 2653–2656.
  69. 69.E. Rahtu, J. Kannala, M. Salo, and J. Heikkilä, "Segmenting salient objects from images and videos," in ECCV, 2010.
  70. 70.M.-M. Cheng, N. J. Mitra, X. Huang, P. H. S. Torr, and S.-M. Hu, "Global contrast based salient region detection," IEEE TPAMI, 2015.
  71. 71.L. Duan, C. Wu, J. Miao, L. Qing, and Y. Fu, "Visual saliency detection by spatially weighted dissimilarity," in IEEE CVPR, 2011, pp. 473–480.
  72. 72.K.-Y. Chang, T.-L. Liu, H.-T. Chen, and S.-H. Lai, "Fusing generic objectness and visual saliency for salient object detection," in IEEE ICCV, 2011, pp. 914–921.
  73. 73.H. Jiang, J. Wang, Z. Yuan, T. Liu, and N. Zheng, "Automatic salient object segmentation based on context and shape prior," in BMVC, 2011.
  74. 74.H. R. Tavakoli, E. Rahtu, and J. Heikkilä, "Fast and efficient saliency detection using sparse sampling and kernel density estimation," in Scandinavian Conference on Image Analysis, 2011, pp. 666–675.
  75. 75.F. Perazzi, P. Krähenbühl, Y. Pritch, and A. Hornung, "Saliency filters: Contrast based filtering for salient region detection," in IEEE CVPR, 2012, pp. 733–740.
  76. 76.Y. Xie, H. Lu, and M.-H. Yang, "Bayesian saliency via low and mid level cues," IEEE TIP, vol. 22, no. 5, 2013.
  77. 77.Q. Yan, L. Xu, J. Shi, and J. Jia, "Hierarchical saliency detection," in IEEE CVPR, 2013, pp. 1155–1162.
  78. 78.C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang, "Saliency detection via graph-based manifold ranking," in IEEE CVPR, 2013.
  79. 79.J. Wang, H. Jiang, Z. Yuan, M.-M. Cheng, X. Hu, and N. Zheng, "Salient object detection: A discriminative regional feature integration approach," International Journal of Computer Vision, vol. 123, no. 2, pp. 251–268, 2017.
  80. 80.R. Margolin, A. Tal, and L. Zelnik-Manor, "What makes a patch distinct?" in IEEE CVPR, 2013, pp. 1139–1146.
  81. 81.P. Siva, C. Russell, T. Xiang, and L. Agapito, "Looking beyond the image: Unsupervised learning for object saliency and detection," in IEEE CVPR, 2013, pp. 3238–3245.
  82. 82.M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V. Vineet, and N. Crook, "Efficient salient region detection with soft image abstraction," in IEEE ICCV, 2013, pp. 1529–1536.
  83. 83.X. Li, Y. Li, C. Shen, A. R. Dick, and A. van den Hengel, "Contextual hypergraph modeling for salient object detection," in IEEE ICCV, 2013, pp. 3328–3335.
  84. 84.X. Li, H. Lu, L. Zhang, X. Ruan, and M.-H. Yang, "Saliency detection via dense and sparse reconstruction," in IEEE ICCV, 2013.
  85. 85.B. Jiang, L. Zhang, H. Lu, C. Yang, and M.-H. Yang, "Saliency detection via absorbing markov chain," in IEEE ICCV, 2013.
  86. 86.P. Jiang, H. Ling, J. Yu, and J. Peng, "Salient region detection by ufo: Uniqueness, focusness and objectness," in IEEE ICCV, 2013.
  87. 87.C. Yang, L. Zhang, and H. Lu, "Graph-regularized saliency detection with convex-hull-based center prior," IEEE Signal Processing Letters, vol. 20, no. 7, pp. 637–640, 2013.
  88. 88.W. Zhu, S. Liang, Y. Wei, and J. Sun, "Saliency optimization from robust background detection," in IEEE CVPR, 2014.
  89. 89.J. Kim, D. Han, Y.-W. Tai, and J. Kim, "Salient region detection via high-dimensional color transform," in IEEE CVPR, 2014.
  90. 90.Z. Liu, W. Zou, and O. Le Meur, "Saliency tree: A novel saliency detection framework," IEEE TIP, 2013.
  91. 91.C. Aytekin, S. Kiranyaz, and M. Gabbouj, "Automatic object segmentation by quantum cuts," in IEEE ICPR, 2014, pp. 112–117.
  92. 92.N. D. Bruce and J. K. Tsotsos, "Saliency based on information maximization," in NIPS, 2005, pp. 155–162.
  93. 93.J. Harel, C. Koch, and P. Perona, "Graph-based visual saliency," in NIPS, 2007, pp. 545–552.
  94. 94.X. Hou and L. Zhang, "Saliency detection: A spectral residual approach," in IEEE CVPR, 2007, pp. 1–8.
  95. 95.L. Zhang, M. H. Tong, T. K. Marks, H. Shan, and G. W. Cottrell, "Sun: A bayesian framework for saliency using natural statistics," J. Vision, vol. 8, no. 7, pp. 32, 1–20, 2008.
  96. 96.H. J. Seo and P. Milanfar, "Static and space-time visual saliency detection by self-resemblance," J. Vision, 2009.
  97. 97.N. Murray, M. Vanrell, X. Otazu, and C. A. Parraga, "Saliency estimation using a non-parametric low-level vision model," in IEEE CVPR, 2011, pp. 433–440.
  98. 98.X. Hou, J. Harel, and C. Koch, "Image signature: Highlighting sparse salient regions," IEEE TPAMI, vol. 34, no. 1, 2012.
  99. 99.E. Erdem and A. Erdem, "Visual saliency estimation by nonlinearly integrating features using region covariances," J. Vision, vol. 13, no. 4, pp. 11, 1–20, 2013.
  100. 100.J. Zhang and S. Sclaroff, "Saliency detection: A boolean map approach," in IEEE ICCV, 2013, pp. 153–160.
  101. 101.B. Alexe, T. Deselaers, and V. Ferrari, "What is an object?" in IEEE CVPR, 2010, pp. 73–80.
  102. 102.THUR15000, "http://mmcheng.net/gsal/."
  103. 103.A. Borji, "What is a salient object? a dataset and a baseline model for salient object detection," in IEEE TIP, 2014.
  104. 104.S. Alpert, M. Galun, R. Basri, and A. Brandt, "Image segmentation by probabilistic bottom-up aggregation and cue integration," in IEEE CVPR, 2007, pp. 1–8.
  105. 105.Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, "The secrets of salient object segmentation," in IEEE CVPR, 2014.
  106. 106.P. F. Felzenszwalb and D. P. Huttenlocher, "Efficient graph-based image segmentation," IJCV, pp. 167–181, 2004.
  107. 107.V. Movahedi and J. H. Elder, "Design and perceptual validation of performance measures for salient object segmentation," in IEEE CVPRW, 2010.
  108. 108.R. Margolin, L. Zelnik-Manor, and A. Tal, "How to evaluate foreground maps?" in IEEE CVPR, 2014.
  109. 109.C. Rother, V. Kolmogorov, and A. Blake, ""GrabCut": interactive foreground extraction using iterated graph cuts," ACM TOG, vol. 23, no. 3, pp. 309–314, 2004.
  110. 110.T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum, "Learning to detect a salient object," IEEE TPAMI, vol. 33, no. 2, pp. 353–367, 2011.
  111. 111.D. R. Martin, C. C. Fowlkes, and J. Malik, "Learning to detect natural image boundaries using local brightness, color, and texture cues," IEEE TPAMI, vol. 26, no. 5, pp. 530–549, 2004.
  112. 112.S. Avidan and A. Shamir, "Seam carving for content-aware image resizing," in ACM TOG, vol. 26, no. 3, 2007, p. 10.
  113. 113.J. Davis and M. Goadrich, "The relationship between precision-recall and roc curves," in ICML, 2006, pp. 233–240.
  114. 114.J.-Y. Zhu, J. Wu, Y. Wei, E. Chang, and Z. Tu, "Unsupervised object class discovery via saliency-guided multiple class learning," in IEEE CVPR, 2012, pp. 3218–3225.
  115. 115.J. He, J. Feng, X. Liu, T. Cheng, T.-H. Lin, H. Chung, and S.-F. Chang, "Mobile product search with bag of hash bits and boundary reranking," in IEEE CVPR. IEEE, 2012, pp. 3005–3012.
  116. 116.P. Wang, J. Wang, G. Zeng, J. Feng, H. Zha, and S. Li, "Salient object detection for searched web images via global saliency," in IEEE CVPR, 2012, pp. 3194–3201.
  117. 117.J. Zhang, M. Sameki, S. Ma, B. Price, R. Mech, X. Shen, M. Betke, S. Sclaroff, and Z. Lin, "Salient object subitizing," in IEEE CVPR, 2015.
  118. 118.H. Peng, B. Li, W. Xiong, W. Hu, and R. Ji, "Rgbd salient object detection: a benchmark and algorithms," in ECCV, 2014, pp. 92–109.
  119. 119.A. K. Mishra, Y. Aloimonos, L. F. Cheong, and A. Kassim, "Active visual segmentation," IEEE TPAMI, vol. 34, 2012.
  120. 120.L. Mai, Y. Niu, and F. Liu, "Saliency aggregation: A data-driven approach," in IEEE CVPR, 2013, pp. 1131–1138.
  121. 121.O. Le Meur and Z. Liu, "Saliency aggregation: Does unity make strength?" in ACCV. Springer, 2014, pp. 18–32.
  122. 122.A. Borji, D. N. Sihite, and L. Itti, "Objects do not predict fixations better than early saliency: A re-analysis of einhauser et al.'s data," J. Vision, vol. 13, no. 10, p. 18, 2013.
  123. 123.A. Krizhevsky, I. Sutskever, and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," in Advances in neural information processing systems, 2012, pp. 1097–1105.
  124. 124.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions," in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015.
  125. 125.R. B. Girshick, J. Donahue, T. Darrell, and J. Malik, "Rich feature hierarchies for accurate object detection and semantic segmentation," in 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, June 23-28, 2014, 2014, pp. 580–587.
  126. 126.R. Zhao, W. Ouyang, H. Li, and X. Wang, "Saliency detection by multi-context deep learning," in IEEE CVPR, 2015, pp. 1265–1274.
  127. 127.S. He, R. W. Lau, W. Liu, Z. Huang, and Q. Yang, "SuperCNN: A superpixelwise convolutional neural network for salient object detection," IJCV, 2015.
  128. 128.Y. Lin, S. Kong, D. Wang, and Y. Zhuang, "Saliency detection within a deep convolutional architecture," in Workshops at AAAI Conference on Artificial Intelligence, 2014.
  129. 129.G. Li and Y. Yu, "Visual saliency based on multiscale deep features," CoRR, vol. abs/1503.08663, 2015.
  130. 130.A. Borji and J. Tanner, "Reconciling saliency and object center-bias hypotheses in explaining free-viewing fixations," arXiv preprint arXiv:1503.08853, 2015.
  131. 131.G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, "Baby talk: Understanding and generating simple image descriptions," in IEEE CVPR, 2011, pp. 1601–1608.
  132. 132.A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth, "Describing objects by their attributes," in IEEE CVPR, 2009, pp. 1778–1785.
  133. 133.L. Itti and M. A. Arbib, "Attention and the minimal subscene," Action to language via the mirror neuron system, 2006.
  134. 134.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, "Vqa: Visual question answering," arXiv preprint arXiv:1505.00468, 2015.
  135. 135.H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt et al., "From captions to visual concepts and back," arXiv preprint arXiv:1411.4952, 2014.
  136. 136.C. L. Zitnick, D. Parikh, and L. Vanderwende, "Learning the visual interpretation of sentences," in IEEE ICCV, 2013, pp. 1681–1688.
  137. 137.K. Yun, Y. Peng, D. Samaras, G. J. Zelinsky, and T. Berg, "Studying relationships between human gaze, description, and computer vision," in IEEE CVPR, 2013, pp. 739–746.
  138. 138.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick, "Microsoft coco captions: Data collection and evaluation server," arXiv preprint arXiv:1504.00325, 2015.
  139. 139.D. Geman, S. Geman, N. Hallonquist, and L. Younes, "Visual turing test for computer vision systems," Proceedings of the National Academy of Sciences, vol. 112, no. 12, pp. 3618–3623, 2015.
  140. 140.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., "Imagenet large scale visual recognition challenge," IJCV, pp. 1–42, 2014.

Citation

MLA
Borji, A., et al. “Salient Object Detection: A Benchmark”. IEEE Transactions on Image Processing, vol. 24, no. 12, 2015, pp. 5706–22, https://doi.org/10.1109/TIP.2015.2487833.
APA
Borji, A., Cheng, M.-M., Jiang, H., & Li, J. (2015). Salient Object Detection: A Benchmark. IEEE Transactions on Image Processing, 24(12), 5706–5722. https://doi.org/10.1109/TIP.2015.2487833
Chicago
Borji, A., M.-M. Cheng, H. Jiang, and J. Li. 2015. “Salient Object Detection: A Benchmark”. IEEE Transactions on Image Processing 24 (12): 5706–22. https://doi.org/10.1109/TIP.2015.2487833.
Harvard
Borji, A. et al. (2015) “Salient Object Detection: A Benchmark”, IEEE Transactions on Image Processing, 24(12), pp. 5706–5722. Available at: https://doi.org/10.1109/TIP.2015.2487833.
Vancouver
1. Borji A, Cheng M-M, Jiang H, Li J (2015) Salient Object Detection: A Benchmark. IEEE Transactions on Image Processing 24:5706–5722

BibTeX

@article{Borji_2015, title={Salient Object Detection: A Benchmark}, volume={24}, ISSN={1941-0042}, url={http://dx.doi.org/10.1109/TIP.2015.2487833}, DOI={10.1109/tip.2015.2487833}, number={12}, journal={IEEE Transactions on Image Processing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Borji, Ali and Cheng, Ming-Ming and Jiang, Huaizu and Li, Jia}, year={2015}, month=Dec, pages={5706–5722} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF