The Secrets of Salient Object Segmentation

Yin LiXiaodi HouChristof KochJames M. RehgAlan L. Yuille

article2014CVPR1,331 citations

Exposes critical design biases in salient object benchmarks and introduces a unified dataset with joint fixation and segmentation ground truth alongside a method that directly connects human visual attention to object segmentation.

Listen

Visual saliency—the ability of computer vision systems to identify and isolate important visual information in an image—is divided into two isolated research areas: predicting human eye fixations and segmenting complete salient objects. Existing benchmarks for salient object segmentation have relied heavily on simplified, "textbook" datasets where objects exhibit unnaturally high color contrast and distinct boundaries. Consequently, algorithms tuned on these datasets perform poorly when applied to realistic, complex natural scenes, obscuring real technical progress.

The article aims to evaluate the consistency of human saliency annotations, diagnose the dataset design bias affecting existing benchmarks, demonstrate the direct connection between human eye fixations and salient object boundaries, and present an improved segmentation model that bridges the gap between fixation prediction and object segmentation.

The authors conducted psychophysical experiments to create PASCAL-S, a new benchmark dataset comprising 850 natural images augmented with both eye-tracking fixation data from 8 human subjects and unrestricted salient object segmentations from 12 annotators. They benchmarked seven fixation prediction algorithms and four leading salient object segmentation models across multiple standard datasets, analyzing core image statistics such as local and global color contrast, boundary strength, and object size. Furthermore, they trained a machine learning regression model (random forest) using shape metrics and spatial fixation distributions to rank generic object proposals generated by Constrained Parametric Min-Cuts.

The analysis yielded several critical findings. First, salient object segmentation is a well-grounded task with high human agreement, achieving an inter-subject consistency score of 0.972 on PASCAL-S. Second, popular salient object benchmarks suffer from severe dataset design bias; when top segmentation algorithms were evaluated on realistic datasets rather than the standard benchmark, their accuracy dropped by approximately 31% to 34%. Third, fixation algorithms proved highly competitive at identifying salient objects once the influence of center bias and dataset design flaws was removed. Finally, the proposed model combining generic object proposals with fixation-based scoring outperformed existing state-of-the-art salient object segmentation algorithms across all evaluated benchmarks, improving accuracy by up to 11.82% while requiring only a small pool of candidate segments.

These results demonstrate that the long-standing divide between fixation prediction and object segmentation is unnecessary and driven primarily by biased benchmark design. Deploying computer vision models that were validated only on artificial, high-contrast datasets introduces substantial operational performance risks in real-world environments. By decoupling the task into generic region proposal generation followed by fixation-guided ranking, vision systems can achieve superior segmentation performance without relying on rigid, category-specific training or brittle contrast heuristics.

Organizations developing vision systems should decouple object proposal generation from saliency scoring and re-evaluate their models on unbiased benchmarks that separate image collection from annotation. Future research should prioritize expanding realistic multi-annotator datasets and refining proposal generators to better capture small salient objects, which remain a primary limitation of the current candidate generation pipeline.

  • Paper: Unbiased look at dataset bias, Antonio Torralba et al. (2011). This paper establishes the foundational methodology for identifying and measuring dataset bias in computer vision benchmarks, directly inspiring the source's investigation of design bias in salient object datasets.
  • Paper: State-of-the-Art in Visual Attention Modeling, Ali Borji et al. (2013). This comprehensive review categorizes computational visual attention and fixation models, providing the conceptual background required to understand the disconnection between fixation prediction and salient object segmentation.
  • Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). This work introduces regional contrast-based saliency methods that represent the prevailing paradigm and standard benchmarked techniques evaluated and critiqued in the source.
  • Paper: Graph-Based Visual Saliency, Jonathan Harel et al. (2006). This seminal paper formulates graph-based fixation prediction, serving as a primary baseline for human eye-gaze modeling analyzed in the source.
  • Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, Laurent Itti et al. (1998). This foundational work defines biologically motivated bottom-up visual saliency maps that underpin the visual attention field.
Cover for The Secrets of Salient Object Segmentation

Abstract

In this paper we provide an extensive evaluation of fixation prediction and salient object segmentation algorithms as well as statistics of major datasets. Our analysis identifies serious design flaws of existing salient object benchmarks, called the dataset design bias, by over emphasizing the stereotypical concepts of saliency. The dataset design bias does not only create the discomforting disconnection between fixations and salient object segmentation, but also misleads the algorithm designing. Based on our analysis, we propose a new high quality dataset that offers both fixation and salient object segmentation ground-truth. With fixations and salient object being presented simultaneously, we are able to bridge the gap between fixations and salient objects, and propose a novel method for salient object segmentation. Finally, we report significant benchmark progress on three existing datasets of segmenting salient objects

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Fixation prediction
  • 2.2 Salient object segmentation
  • 2.3 Objectness, object proposal, and foreground segments
  • 2.4 Datasets and dataset bias
  • 3 Dataset Analysis
  • 3.1 Psychophysical experiments on the PASCAL-S dataset
  • 3.2 Evaluating dataset consistency
  • 3.3 Benchmarking
  • 3.4 Dataset design bias
  • 3.5 Fixations and F-measure
  • 4 From Fixations to Salient Object Detection
  • 4.1 Salient object, object proposal and fixations
  • 4.2 The model
  • 4.3 Limits of the model
  • 4.4 Results
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — CPMC-Fixation Salient Object Segmentation Framework

    model/method

    The framework decouples salient object segmentation into two independent stages: category-independent object proposal generation and fixation-guided saliency scoring.

    1. Proposal Generation: Constrained Parametric Min-Cuts (CPMC) generates an over-complete pool of candidate figure-ground segmentation masks by initializing foreground seeds on a uniform grid and solving multiple parameterized min-cut optimization problems.

    2. Feature Extraction: For each candidate segment mask, a 33-dimensional feature vector is extracted comprising 16 geometric shape descriptors and 17 fixation energy distribution descriptors computed against a continuous or discrete fixation saliency map.

    3. Saliency Scoring: A Random Regression Forest composed of 30 decision trees predicts the saliency score of each candidate mask. The forest is trained to predict the Intersection-over-Union (IoU) overlap between the candidate proposal and the ground-truth salient object masks by minimizing the Mean Squared Error (MSE) at each decision tree split.

    4. Mask Formation: In the test phase, candidate segments are scored independently. The top-KK scoring segments (typically K=20K = 20) are averaged at the pixel level to form a continuous saliency distribution, which is then binarized via simple thresholding. Because CPMC proposals already conform to natural image boundaries, explicit post-processing boundary cut algorithms are not required.

  2. Knowl 2 — 33-Dimensional Feature Representation for Salient Segment Scoring

    model/method

    To score the saliency of an object proposal mask without using class-specific appearance features, each candidate segment is encoded into a 33-dimensional feature vector combining geometric shape properties and spatial fixation distributions:

    • Shape Features (16 dimensions):

      • Area (11 dim)
      • Centroid coordinates (x,y)(x, y) (22 dims)
      • Convex area (11 dim)
      • Euler number (11 dim)
      • Perimeter (11 dim)
      • Major and minor axis lengths (22 dims)
      • Eccentricity (11 dim)
      • Orientation (11 dim)
      • Equivalent diameter (11 dim)
      • Solidity (11 dim)
      • Extent (11 dim)
      • Width and height (22 dims)
    • Fixation Distribution Features (17 dimensions):

      • Minimum and maximum fixation energy within the segment (22 dims)
      • Mean fixation energy within the segment (11 dim)
      • Fixation-weighted centroid coordinates within the segment (22 dims)
      • Fixation Energy Ratio (11 dim), defined as:

    Fixation Energy Ratio=∑p∈SE(p)∑p∈IE(p)\text{Fixation Energy Ratio} = \frac{\sum_{p \in \mathcal{S}} E(p)}{\sum_{p \in \mathcal{I}} E(p)}

    where S\mathcal{S} is the set of pixels in the segment mask, I\mathcal{I} is the full image, and E(p)E(p) represents the fixation energy (discrete hit count for human fixations, continuous probability in [0,1][0, 1] for algorithmic predictions) at pixel pp.

    • 4×34 \times 3 Spatial Histogram of Fixations (1212 dims), extracted across the object candidate mask after rotating and aligning its major axis to capture internal fixation distribution patterns.
  3. Knowl 3 — Comparative Evaluation of Salient Object Segmentation Benchmarks

    data/table

    The benchmark compares standalone salient object detectors (FT, GC, PCAS, SF), standalone fixation prediction algorithms (AIM, AWS, DVA, GBVS, ITTI, SIG, SUN with an added Gaussian center bias of σ=0.4×image width\sigma = 0.4 \times \text{image width}), baseline methods, and the CPMC + Fixation model (K=20K=20) across the FT, IS, and PASCAL-S datasets using FF-measure (FβF_\beta with β2=0.3\beta^2 = 0.3).

    Algorithm FT IS PASCAL-S
    Salient Object Algorithms
    FT 0.7427 0.4736 0.4325
    GC 0.8383 0.6261 0.6072
    PCAS 0.8646 0.6558 0.6275
    SF 0.8850 0.5555 0.5570
    Original Fixation (+ Center Bias)
    AIM 0.7148 0.4522 0.6216
    AWS 0.7240 0.6121 0.5906
    DVA 0.6592 0.3764 0.5223
    GBVS 0.7093 0.5308 0.6186
    ITTI 0.6816 0.4452 0.6079
    SIG 0.6959 0.5131 0.5850
    SUN 0.6708 0.3314 0.5281
    Baseline Models
    CPMC Ranking 0.2661 0.3918 0.5799
    CPMC + Human Fixations N/A 0.7863 0.7756
    CPMC Best (Oracle top 200) 0.9496 0.8416 0.8699
    GT Segments + Human Fixations N/A N/A 0.9201
    CPMC + Fixation (Proposed, K=20K=20)
    CPMC + AIM 0.8920 0.6728 0.7204
    CPMC + AWS 0.8998 0.7241 0.7224
    CPMC + DVA 0.8700 0.6377 0.7112
    CPMC + GBVS 0.9097 0.7264 0.7454
    CPMC + ITTI 0.8950 0.6827 0.7288
    CPMC + SIG 0.8908 0.7255 0.7214
    CPMC + SUN 0.8635 0.6249 0.7058

    The CPMC + GBVS combination achieves the highest FF-measure across all three benchmarks. Compared to the best existing salient object models (PCAS on PASCAL-S: 0.6275; PCAS on IS: 0.6558; SF on FT: 0.8850), CPMC + GBVS achieves improvements of +11.82%+11.82\%, +7.06%+7.06\%, and +2.47%+2.47\%, respectively.

  4. Knowl 4 — Dataset Design Bias in Salient Object Benchmarks

    definition

    Dataset design bias occurs when the image selection process is coupled with the annotation concept rather than conducted independently. In salient object datasets, this bias arises when curators intentionally select images containing clear, isolated, high-contrast foreground objects against simple backgrounds, thereby disproportionately sampling positive saliency traits and suppressing ambiguous negative examples.

    Quantitative analysis of dataset statistics reveals this bias:

    • Local Color Contrast: Measured as the χ2\chi^2 distance between RGB color histograms of 5×55 \times 5 foreground and background patches along object boundaries. The FT dataset exhibits unnaturally high boundary contrast distributions compared to IS and PASCAL-S.
    • Global Color Contrast: Measured as the χ2\chi^2 distance between the full foreground object RGB histogram and the full background RGB histogram. FT is skewed toward extreme contrast values.
    • Local Boundary Strength: Evaluated via the mean globalized Probability of Boundary (gPB) response in 3×33 \times 3 patches along object edges. FT displays significantly higher edge responses than natural scenes.

    Due to dataset design bias in FT, top-performing salient object models overfit textbook saliency cues, suffering an average performance degradation of 30.88%30.88\% on IS and 33.70%33.70\% on PASCAL-S.

  5. Knowl 5 — PASCAL-S Dataset Protocol

    experimental setup

    The PASCAL-S benchmark dataset is built on the 850 natural validation images of the PASCAL VOC 2010 segmentation challenge to ensure independent image selection and annotation, pairing eye fixation tracking with salient object masks.

    • Fixation Acquisition: 8 human subjects engaged in free-viewing exploration for 2 seconds per image. Eye gaze was recorded at 125 Hz using an EyeLink 1000 eye-tracker, with re-calibration performed every 25 images.
    • Full Ground-Truth Segmentation: Full manual segmentation was performed to isolate all candidate objects in each scene prior to saliency labeling, adhering to three constraints: sub-parts (e.g., human faces) were not segmented separately, disconnected regions of a single object were isolated into separate segments, and hollow objects (such as bike wheels) were filled as solid regions.
    • Salient Object Annotation: 12 subjects selected salient objects by clicking on segments within the full segmentation masks. No constraints were imposed on viewing time or on the number of selectable objects. Segment saliency was computed as the fraction of subjects who clicked on that segment.
  6. Knowl 6 — Upper-Bound Performance of Saliency Selector and CPMC Proposals

    empirical result

    Decoupled oracle evaluations establish the independent upper-bound potentials of the segment scoring mechanism and the proposal generation stage:

    • Selector Upper Bound: Training the random regression forest selector on ground-truth object segments paired with human fixation maps on PASCAL-S yields an FF-measure of 0.92010.9201 (Precision P=0.9328P = 0.9328, Recall R=0.7989R = 0.7989). This confirms that spatial fixation distributions paired with geometric shape descriptors provide sufficient information to accurately isolate salient objects given accurate candidate boundaries.

    • Proposal Upper Bound (CPMC Best): Greedily matching CPMC candidate segments from the first 200 proposals to ground-truth salient objects achieves an oracle FF-measure of 0.86990.8699 (P=0.8687,R=0.8830P = 0.8687, R = 0.8830) on PASCAL-S, 0.94960.9496 (P=0.9494,R=0.9517P = 0.9494, R = 0.9517) on FT, and 0.84160.8416 (P=0.8572,R=0.6982P = 0.8572, R = 0.6982) on IS.

  7. Knowl 7 — Convergence Efficiency of Fixation-Based Proposal Ranking over Proposal Count $K$

    empirical result

    The CPMC + Fixation model achieves stable, near-optimal FF-measure performance with a small number of top proposals (K=20K = 20). In contrast, the original CPMC category-independent ranking function fails to converge even when utilizing K=200K = 200 proposals.

    On the PASCAL-S test split, the CPMC ranking function achieves an FF-measure below 0.550.55 at K=10K = 10 and reaches approximately 0.680.68 at K=200K = 200. By contrast, CPMC + Fixation models (using AWS, GBVS, AIM, or SIG) reach FF-measures between 0.700.70 and 0.750.75 at K=20K = 20, matching or exceeding their performance at K=200K = 200. Fixation features enable rapid selection of salient foreground objects without requiring evaluation of hundreds of candidate segments.

  8. Knowl 8 — Inter-Subject Consistency in Eye Fixation and Salient Object Annotation

    empirical result

    Human consistency across subjects was measured using a split-half cross-validation protocol (randomly partitioning subjects into two equal 50%50\% splits, generating a test map from one split, and evaluating it against the remaining ground-truth split):

    • Fixation Prediction Consistency (AUC): Evaluated using saliency maps generated from test fixations blurred with a 2D Gaussian kernel (σ=0.05×image width\sigma = 0.05 \times \text{image width}):

      • PASCAL-S: 0.8350.835
      • Bruce: 0.8300.830
      • Cerf: 0.9030.903
      • IS: 0.8360.836
      • Judd: 0.8670.867
    • Salient Object Segmentation Consistency (FF-measure): Evaluated on binary consensus masks thresholded at Th=0.5Th = 0.5 (agreement by at least half of the annotators in the split):

      • PASCAL-S: 0.9720.972
      • IS: 0.9000.900

    These results establish that salient object segmentation displays high human agreement across annotators, demonstrating that the task is statistically well-posed despite open-ended instructions.

  9. Knowl 9 — Small Object Omission Limitation of CPMC Proposal Initialization

    limitation

    The CPMC + Fixation framework struggles to detect and segment small salient objects. Because CPMC initializes foreground seed points uniformly over the image grid, small objects frequently fall between seed locations and fail to generate corresponding candidate masks in the proposal pool. Since the framework relies entirely on re-ranking candidate proposals without generating new segment boundaries, any object missed during initial proposal generation cannot be recovered in the final saliency mask.

Coverage note — No substantial contributed material was omitted. All key contributions—including the PASCAL-S dataset, dataset design bias analysis, the 33D feature proposal scoring method, benchmark comparisons, oracle limits, convergence analysis, and model limitations—are fully represented.

References

  1. 1.R. Achanta, S. Hemami, F. Estrada, and S. Susstrunk. Frequency-tuned salient region detection. In CVPR, pages 1597–1604. IEEE, 2009.
  2. 2.B. Alexe, T. Deselaers, and V. Ferrari. Measuring the objectness of image windows. TPAMI, IEEE, pages 2189–2202, 2012.
  3. 3.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik. Contour detection and hierarchical image segmentation. TPAMI, IEEE, 33(5):898–916, 2011.
  4. 4.A. Borji, D. N. Sihite, and L. Itti. Salient object detection: A benchmark. In ECCV, pages 414–429. Springer, 2012.
  5. 5.A. Borji, D. N. Sihite, and L. Itti. What stands out in a scene? a study of human explicit saliency judgment. Vision research, 91:62–77, 2013.
  6. 6.N. Bruce and J. Tsotsos. Saliency based on information maximization. In NIPS, pages 155–162, 2005.
  7. 7.J. Carreira and C. Sminchisescu. Constrained parametric min-cuts for automatic object segmentation. In CVPR, pages 3241–3248. IEEE, 2010.
  8. 8.M. Cerf, J. Harel, W. Einh¨auser, and C. Koch. Predicting human gaze using low-level saliency combined with face detection. NIPS, 20, 2008.
  9. 9.M.-M. Cheng, G.-X. Zhang, N. J. Mitra, X. Huang, and S.-M. Hu. Global contrast based salient region detection. In CVPR, pages 409–416. IEEE, 2011.
  10. 10.M.-M. Cheng, Z. Zhang, W.-Y. Lin, and P. H. S. Torr. BING: Binarized normed gradients for objectness estimation at 300fps. In IEEE CVPR, 2014.
  11. 11.I. Endres and D. Hoiem. Category independent object proposals. In ECCV, pages 575–588. Springer, 2010.
  12. 12.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2010 (VOC2010) Results.
  13. 13.A. Garcia-Diaz, V. Lebor´an, X. R. Fdez-Vidal, and X. M. Pardo. On the relationship between optical variability, visual saliency, and eye fixations: A computational approach. Journal of Vision, 12(6), 2012.
  14. 14.J. Harel, C. Koch, and P. Perona. Graph-based visual saliency. In NIPS, pages 545–552, 2006.
  15. 15.X. Hou, J. Harel, and C. Koch. Image signature: Highlighting sparse salient regions. TPAMI, IEEE, 34(1):194–201, 2012.
  16. 16.X. Hou and L. Zhang. Dynamic visual attention: Searching for coding length increments. In NIPS, pages 681–688, 2008.
  17. 17.L. Itti, C. Koch, and E. Niebur. A model of saliency-based visual attention for rapid scene analysis. TPAMI, IEEE, 20(11):1254–1259, 1998.
  18. 18.T. Judd, K. Ehinger, F. Durand, and A. Torralba. Learning to predict where humans look. In ICCV, pages 2106–2113. IEEE, 2009.
  19. 19.F. Li, J. Carreira, and C. Sminchisescu. Object recognition as ranking holistic figure-ground hypotheses. In CVPR, pages 1712–1719. IEEE, 2010.
  20. 20.J. Li, M. D. Levine, X. An, X. Xu, and H. He. Visual saliency based on scale-space analysis in the frequency domain. TPAMI, IEEE, 2013.
  21. 21.T. Liu, J. Sun, N.-N. Zheng, X. Tang, and H.-Y. Shum. Learning to detect a salient object. In CVPR, 2007.
  22. 22.R. Margolin, A. Tal, and L. Zelnik-Manor. What makes a patch distinct? In CVPR. IEEE, 2013.
  23. 23.F. Perazzi, P. Krahenbuhl, Y. Pritch, and A. Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, pages 733–740. IEEE, 2012.
  24. 24.B. W. Tatler, R. J. Baddeley, and I. D. Gilchrist. Visual correlates of fixation selection: Effects of scale and time. Vision research, 45(5):643–659, 2005.
  25. 25.A. Torralba and A. A. Efros. Unbiased look at dataset bias. In CVPR, pages 1521–1528. IEEE, 2011.
  26. 26.L. Zhang, M. H. Tong, T. K. Marks, H. Shan, and G. W. Cottrell. Sun: A bayesian framework for saliency using natural statistics. Journal of Vision, 8(7), 2008.

Citation

MLA
Li, Y., et al. “The Secrets of Salient Object Segmentation”. arXiv, 2014, http://arxiv.org/abs/1406.2807v2.
APA
Li, Y., Hou, X., Koch, C., Rehg, J. M., & Yuille, A. L. (2014). The Secrets of Salient Object Segmentation. arXiv. http://arxiv.org/abs/1406.2807v2
Chicago
Li, Y., X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille. 2014. “The Secrets of Salient Object Segmentation”. arXiv. http://arxiv.org/abs/1406.2807v2.
Harvard
Li, Y. et al. (2014) “The Secrets of Salient Object Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1406.2807v2.
Vancouver
1. Li Y, Hou X, Koch C, Rehg JM, Yuille AL (2014) The Secrets of Salient Object Segmentation. arXiv

BibTeX

@article{li2014the,
  title = {The Secrets of Salient Object Segmentation},
  author = {Li, Yin and Hou, Xiaodi and Koch, Christof and Rehg, James M. and Yuille, Alan L.},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1406.2807v2},
  eprint = {1406.2807}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE