Deeply Supervised Salient Object Detection with Short Connections
Qibin HouMing-Ming ChengXiao-Wei HuAli BorjiZhuowen TuPhilip Torr
Introduces short connections to deeply supervised skip-layer architectures to capture multi-scale features, achieving state-of-the-art salient object detection accuracy and fast inference across five standard benchmarks.
Salient object detection aims to identify and segment the most visually distinctive objects in an image from the background. This capability serves as an essential preliminary step for various practical applications, such as image and video compression, content-aware editing, object recognition, and visual tracking. While deep learning methods have substantially advanced computer vision, standard fully convolutional networks struggle to simultaneously resolve broad scene context and fine boundary details. Previous architectures either rely on coarse semantic predictions or produce noisy edge maps, failing to cleanly extract complete salient objects.
The main objective of the article is to demonstrate a fully convolutional neural network that effectively combines high-level semantic localization with rich low-level spatial details by introducing top-down short connections into a deeply supervised architecture. The article also comprehensively evaluates how training dataset composition impacts detection performance across standardized benchmarks.
To achieve this, the authors designed a network based on standard backbone models (VGGNet and ResNet-101) augmented with side-output layers and top-down skip connections that transfer high-level features from deeper layers directly to shallower layers. The model was trained and benchmarked across five standard datasets (MSRA-B, ECSSD, HKU-IS, PASCAL-S, and SOD) using universally agreed metrics: precision-recall curves, F-measure (evaluating overall accuracy), and mean absolute error (MAE, evaluating pixel-level prediction errors). In addition, an auxiliary classification branch was introduced to predict whether an image actually contains a salient object, and an exhaustive cross-dataset evaluation of eleven training set combinations was performed.
The findings show that the proposed architecture achieves state-of-the-art accuracy across all five test benchmarks. The model improved the best existing F-measure scores by approximately 1 percentage point on challenging datasets like ECSSD and SOD, while reducing MAE significantly (by over 1 percentage point on MSRA-B and PASCAL-S). Processing speed is high: the network computes a prediction map for a 300 × 400 image in roughly 0.08 seconds (under 0.5 seconds when using a conditional random field refinement step), which is more than ten times faster than competing deep methods. Furthermore, the multi-dataset analysis revealed that dataset quality and scene diversity matter more than raw data volume; simply expanding training image counts does not guarantee better performance, and training models on individual datasets introduces significant performance bias across different evaluation benchmarks.
These results demonstrate that fast, highly accurate saliency detection can be deployed into real-time operational pipelines without requiring complex, computationally expensive post-processing routines or manual feature engineering. By resolving both object localization and boundary detail in a unified end-to-end framework, systems can achieve higher operational reliability and throughput in real-world visual applications.
For future development and fair benchmarking, the article recommends adopting a combined, multi-source training set (specifically the 9,103-image composite set designated as Scheme 11) to eliminate dataset bias. Developers facing scenes where salient objects might be entirely absent should implement the auxiliary existence-prediction branch. To address remaining failure modes—such as complex backgrounds, low foreground-background contrast, and transparent objects—future work should explore segment-level prior knowledge and more challenging datasets containing complex, cluttered environments.
- Paper: Holistically-Nested Edge Detection, Saining Xie et al. (2015). This paper establishes the deeply supervised holistically-nested edge detection (HED) architecture with side-outputs, which the source directly adapts and enhances with short connections for salient object detection.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). This work introduces fully convolutional networks and skip connections for dense pixel-level prediction, providing the foundational paradigm underlying the source paper's saliency segmentation framework.
- Paper: Salient Object Detection: A Benchmark, Ali Borji et al. (2015). This comprehensive benchmark formalizes salient object detection evaluation protocols and datasets, providing the experimental foundation used to assess the source method.
- Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). This paper defines fundamental global contrast-based principles and benchmark datasets for salient region detection that modern deep learning models build upon.
- Paper: BASNet: Boundary-Aware Salient Object Detection, Xuebin Qin et al. (2019). This paper extends deeply supervised salient object detection by incorporating a boundary-aware residual refinement network and hybrid multi-level losses.
- Paper: U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection, Xuebin Qin et al. (2020). This work advances multi-scale deep supervision in salient object detection by designing a nested U-structure that captures fine-grained boundaries without relying on pre-trained backbones.
- Paper: Structure-Measure: A New Way to Evaluate Foreground Maps, Deng-Ping Fan et al. (2017). This work introduces the Structure-measure metric to overcome pixel-level evaluation limitations in salient object detection models like the one proposed in the source.
- Paper: Enhanced-alignment Measure for Binary Foreground Map Evaluation, Deng-Ping Fan et al. (2018). This article develops the Enhanced-alignment measure to better assess structural and global alignment in binary saliency maps produced by salient object detection frameworks.
