Salient Object Detection: A Benchmark
Ali BorjiMing-Ming ChengHuaizu JiangJia Li
Establishes a comprehensive benchmark by evaluating forty models across six datasets, analyzing factors like center bias and scene complexity to expose key failure modes and guide future salient object detection research.
Visual attention systems enable computer vision applications—such as autonomous robotics, image editing, compression, and retrieval—to automatically identify and isolate primary foreground objects in cluttered real-world images. However, the field has suffered from ambiguous problem definitions, conflicting performance evaluations, and untested model progression across varied image datasets.
The article establishes a large-scale standardized benchmark to rigorously assess the progress, accuracy, and operational efficiency of single-image salient object detection models against related visual attention approaches.
The authors conducted a comprehensive empirical evaluation comparing 41 computational models (29 dedicated salient object detection models, 10 human fixation prediction models, 1 generic object proposal method, and 1 location baseline) across 7 diverse public image datasets totaling tens of thousands of images. Performance was benchmarked across multiple standard evaluation criteria, including precision-recall curves, weighted harmonic mean accuracy scores, receiver operating characteristics, mean absolute error, and image processing runtimes.
The benchmark demonstrates five key findings. First, dedicated salient object detection algorithms significantly outperformed traditional fixation prediction and object proposal models, demonstrating that segmenting full foreground objects requires distinct computational formulations. Second, data-driven regional methods (most notably DRFI, followed by models like QCUT and RBD) established the highest overall accuracy across the benchmark datasets. Third, regional superpixel groupings and boundary background assumptions proved to be the most effective design elements, outperforming pixel-level operations in both boundary delineation and processing speed. Fourth, performance dropped sharply across all algorithms when handling cluttered scenes, small objects, or off-center subjects, revealing a widespread over-reliance on center biases. Fifth, processing speeds varied by several orders of magnitude, ranging from roughly 0.017 seconds to over 100 seconds per standard image.
These findings indicate that real-world deployment risks can be substantially reduced by prioritizing region-based, learning-driven models that combine high segmentation precision with sub-second execution speeds. Practitioners must note that relying on algorithms tuned solely for centered, single-object images poses significant operational failure risks in complex, real-world environments with low-contrast boundaries or multiple focal points.
To move the field forward, development efforts should transition toward robust deep learning architectures and train models on complex, multi-object images without positional biases. Future benchmarks must also incorporate images lacking salient targets entirely, while expanding into multi-image domains such as video streams and multi-view datasets.
While the empirical findings are highly reliable for single-image static processing, caution is advised when generalizing these performance rankings to real-time video, multi-camera setups, or active robotic vision systems that fall outside this study's single-frame benchmark scope.
- Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). Cheng et al. introduced foundational global contrast-based salient region detection algorithms and benchmark datasets that this survey directly evaluates and builds upon.
- Paper: State-of-the-Art in Visual Attention Modeling, Ali Borji et al. (2013). This comprehensive taxonomy of visual attention models establishes the theoretical baseline and metrics for distinguishing eye-fixation prediction from salient object detection.
- Paper: Graph-Based Visual Saliency, Jonathan Harel et al. (2006). The Graph-Based Visual Saliency framework is a classical fixation prediction baseline that the benchmark contrasts against object-level saliency detectors.
- Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, L. Itti et al. (1998). Itti, Koch, and Niebur formulate the seminal bottom-up computational architecture of visual saliency that underpins the entire field evaluated in the benchmark.
- Paper: Unbiased look at dataset bias, A. Torralba et al. (2011). Torralba and Efros provide the critical methodology for analyzing dataset bias and cross-dataset generalization that the benchmark adapts to salient object detection datasets.
- Paper: Structure-Measure: A New Way to Evaluate Foreground Maps, Deng-Ping Fan et al. (2017). This work directly addresses the benchmark's call for better evaluation scores by proposing the Structure-measure metric to overcome traditional pixel-level evaluation limitations.
- Paper: U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection, Xuebin Qin et al. (2020). U2-Net advances salient object detection beyond traditional handcrafted models into deep nested architectures evaluated on the standard benchmark datasets established by the survey.
- Paper: Object Detection With Deep Learning: A Review, Zhong-Qiu Zhao et al. (2018). This review surveys the subsequent deep learning era of object detection and incorporates salient object detection into a broader deep visual localization framework.
