Deeply Supervised Salient Object Detection with Short Connections

Qibin HouMing-Ming ChengXiao-Wei HuAli BorjiZhuowen TuPhilip Torr

article2016TPAMI1,453 citations

Introduces short connections to deeply supervised skip-layer architectures to capture multi-scale features, achieving state-of-the-art salient object detection accuracy and fast inference across five standard benchmarks.

Listen

Salient object detection aims to identify and segment the most visually distinctive objects in an image from the background. This capability serves as an essential preliminary step for various practical applications, such as image and video compression, content-aware editing, object recognition, and visual tracking. While deep learning methods have substantially advanced computer vision, standard fully convolutional networks struggle to simultaneously resolve broad scene context and fine boundary details. Previous architectures either rely on coarse semantic predictions or produce noisy edge maps, failing to cleanly extract complete salient objects.

The main objective of the article is to demonstrate a fully convolutional neural network that effectively combines high-level semantic localization with rich low-level spatial details by introducing top-down short connections into a deeply supervised architecture. The article also comprehensively evaluates how training dataset composition impacts detection performance across standardized benchmarks.

To achieve this, the authors designed a network based on standard backbone models (VGGNet and ResNet-101) augmented with side-output layers and top-down skip connections that transfer high-level features from deeper layers directly to shallower layers. The model was trained and benchmarked across five standard datasets (MSRA-B, ECSSD, HKU-IS, PASCAL-S, and SOD) using universally agreed metrics: precision-recall curves, F-measure (evaluating overall accuracy), and mean absolute error (MAE, evaluating pixel-level prediction errors). In addition, an auxiliary classification branch was introduced to predict whether an image actually contains a salient object, and an exhaustive cross-dataset evaluation of eleven training set combinations was performed.

The findings show that the proposed architecture achieves state-of-the-art accuracy across all five test benchmarks. The model improved the best existing F-measure scores by approximately 1 percentage point on challenging datasets like ECSSD and SOD, while reducing MAE significantly (by over 1 percentage point on MSRA-B and PASCAL-S). Processing speed is high: the network computes a prediction map for a 300 × 400 image in roughly 0.08 seconds (under 0.5 seconds when using a conditional random field refinement step), which is more than ten times faster than competing deep methods. Furthermore, the multi-dataset analysis revealed that dataset quality and scene diversity matter more than raw data volume; simply expanding training image counts does not guarantee better performance, and training models on individual datasets introduces significant performance bias across different evaluation benchmarks.

These results demonstrate that fast, highly accurate saliency detection can be deployed into real-time operational pipelines without requiring complex, computationally expensive post-processing routines or manual feature engineering. By resolving both object localization and boundary detail in a unified end-to-end framework, systems can achieve higher operational reliability and throughput in real-world visual applications.

For future development and fair benchmarking, the article recommends adopting a combined, multi-source training set (specifically the 9,103-image composite set designated as Scheme 11) to eliminate dataset bias. Developers facing scenes where salient objects might be entirely absent should implement the auxiliary existence-prediction branch. To address remaining failure modes—such as complex backgrounds, low foreground-background contrast, and transparent objects—future work should explore segment-level prior knowledge and more challenging datasets containing complex, cluttered environments.

  • Paper: Holistically-Nested Edge Detection, Saining Xie et al. (2015). This paper establishes the deeply supervised holistically-nested edge detection (HED) architecture with side-outputs, which the source directly adapts and enhances with short connections for salient object detection.
  • Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). This work introduces fully convolutional networks and skip connections for dense pixel-level prediction, providing the foundational paradigm underlying the source paper's saliency segmentation framework.
  • Paper: Salient Object Detection: A Benchmark, Ali Borji et al. (2015). This comprehensive benchmark formalizes salient object detection evaluation protocols and datasets, providing the experimental foundation used to assess the source method.
  • Paper: Global contrast based salient region detection, Ming-Ming Cheng et al. (2011). This paper defines fundamental global contrast-based principles and benchmark datasets for salient region detection that modern deep learning models build upon.
Cover for Deeply Supervised Salient Object Detection with Short Connections

Abstract

Recent progress on saliency detection is substantial, benefiting mostly from the explosive development of Convolutional Neural Networks (CNNs). Semantic segmentation and saliency detection algorithms developed lately have been mostly based on Fully Convolutional Neural Networks (FCNs). There is still a large room for improvement over the generic FCN models that do not explicitly deal with the scale-space problem. Holistically-Nested Edge Detector (HED) provides a skip-layer structure with deep supervision for edge and boundary detection, but the performance gain of HED on salience detection is not obvious. In this paper, we propose a new method for saliency detection by introducing short connections to the skip-layer structures within the HED architecture. Our framework provides rich multi-scale feature maps at each layer, a property that is critically needed to perform segment detection. Our method produces state-of-the-art results on 5 widely tested salient object detection benchmarks, with advantages in terms of efficiency (0.15 seconds per image), effectiveness, and simplicity over the existing algorithms.

Table of Contents

  • I Introduction
  • II Related Works
  • II-A CNN-Based Saliency Models
  • II-B Skip-Layer Structures
  • III Deep Supervision with Short Connections
  • III-A Observations
  • III-B HED-based saliency detection
  • III-B1 HED architecture
  • III-B2 Enhanced HED architecture
  • III-C Short connections
  • III-C1 Formulation
  • III-C2 Construction
  • III-D Implementation Details
  • III-D1 Inference
  • III-D2 Smoothing Method
  • III-D3 Parameters
  • IV Experiments and Analyses
  • IV-A Datasets
  • IV-B Evaluation Metrics
  • IV-C Ablation Analysis
  • IV-C1 Various Short Connection Patterns
  • IV-C2 Details of Side-Output Layers
  • IV-C3 Upsampling Operation
  • IV-C4 Data Augmentation
  • IV-C5 Different Backbones
  • IV-C6 The Proposed CRF Model
  • IV-D Comparison with the State-of-the-art
  • IV-D1 Visual Comparison
  • IV-D2 PR Curve
  • IV-D3 F-measure and MAE
  • IV-E The Existence of Saliency
  • IV-F Timing
  • V Discussion
  • V-A Failure Case Analysis
  • V-B Benchmarking Training Set
  • V-B1 Dataset Quality Measuring
  • V-B2 Beyond Training on Individual Datasets
  • VI Conclusion
  • References

Knowls

  1. Knowl 1 — Architecture of Deeply Supervised Salient Object Detector with Short Connections

    model/method

    The Deeply Supervised Salient Object Detection (DSS) framework extends fully convolutional skip-layer architectures by incorporating top-down short connections between side outputs to integrate high-level semantic localization with low-level spatial detail.

    When built on a VGGNet backbone, the network defines 6 side-output paths connected to conv1_2, conv2_2, conv3_3, conv4_3, conv5_3, and the final pooling layer pool5 (representing 6 spatial scales). Unlike standard edge detectors (such as HED) that use single 1×11 \times 1 convolutions per side output, each DSS side output path comprises three convolutional layers:

    1. For conv1_2 and conv2_2: two convolutional layers with 128 channels and 3×33 \times 3 filters, each followed by a ReLU activation, terminated by a 1×11 \times 1 convolution producing a 1-channel score map.
    2. For conv3_3 and conv4_3: two convolutional layers with 256 channels and 5×55 \times 5 filters (with ReLUs), followed by a 1×11 \times 1 convolution.
    3. For conv5_3: two convolutional layers with 512 channels and 5×55 \times 5 filters (with ReLUs), followed by a 1×11 \times 1 convolution.
    4. For pool5: two convolutional layers with 512 channels and 7×77 \times 7 filters (with ReLUs), followed by a 1×11 \times 1 convolution.

    When implemented with a ResNet-101 backbone, the network connects 5 side outputs to layers conv1, res2c, res3b3, res4b22, and res5c using corresponding convolutional side branches. Top-down short connections route feature maps from deeper stages to shallower stages to simultaneously localize salient objects and sharpen boundary definitions.

  2. Knowl 2 — Side-Output Feature Formulation with Top-Down Short Connections

    equation

    Let A^side(m)\hat{A}_{\text{side}}^{(m)} denote the primary score map activations from the mm-th side output branch (for m∈{1,2,…,6}m \in \{1, 2, \dots, 6\}, ordered from shallowest m=1m=1 to deepest m=6m=6). Let R~side(m)\tilde{R}_{\text{side}}^{(m)} denote the refined side activation map after short connection aggregation, and let rimr_i^m denote the learnable scalar/convolutional weight connecting deeper side output ii (i>mi > m) to shallower side output mm.

    In the DSS architecture (Pattern 3), the short-connection side activations are defined hierarchically:

    R~side(m)={∑i=36rimR~side(i)+A^side(m),for m∈{1,2}r5mR~side(5)+r6mR~side(6)+A^side(m),for m∈{3,4}A^side(m),for m∈{5,6}\tilde{R}_{\text{side}}^{(m)} = \begin{cases} \sum_{i=3}^6 r_i^m \tilde{R}_{\text{side}}^{(i)} + \hat{A}_{\text{side}}^{(m)}, & \text{for } m \in \{1, 2\} \\ r_5^m \tilde{R}_{\text{side}}^{(5)} + r_6^m \tilde{R}_{\text{side}}^{(6)} + \hat{A}_{\text{side}}^{(m)}, & \text{for } m \in \{3, 4\} \\ \hat{A}_{\text{side}}^{(m)}, & \text{for } m \in \{5, 6\} \end{cases}

    To match spatial resolutions across scales, score maps from deeper stages R~side(i)\tilde{R}_{\text{side}}^{(i)} are upsampled to the spatial dimensions of stage mm using in-network bilinear interpolation (with scaling factors such as 2×2\times or 4×4\times). The upsampled maps and the native score map A^side(m)\hat{A}_{\text{side}}^{(m)} are concatenated channel-wise and fused using a 1×11 \times 1 convolutional layer to generate the combined side activation R~side(m)\tilde{R}_{\text{side}}^{(m)}.

  3. Knowl 3 — Deep Supervision and Weighted Fusion Loss Formulation

    equation

    For an input image X={xj∣j=1,…,∣X∣}X = \{x_j \mid j=1, \dots, |X|\} and a ground truth binary saliency map Z={zj∣j=1,…,∣Z∣}Z = \{z_j \mid j=1, \dots, |Z|\} where zj∈{0,1}z_j \in \{0, 1\}, the overall network loss L~final\tilde{\mathcal{L}}_{\text{final}} is the sum of side-output supervision losses L~side\tilde{\mathcal{L}}_{\text{side}} and a weighted fusion loss L~fuse\tilde{\mathcal{L}}_{\text{fuse}}:

    L~final(W,w~,f,r)=L~fuse(W,w~,f,r)+L~side(W,w~,r)\tilde{\mathcal{L}}_{\text{final}}(W, \tilde{\mathbf{w}}, \mathbf{f}, \mathbf{r}) = \tilde{\mathcal{L}}_{\text{fuse}}(W, \tilde{\mathbf{w}}, \mathbf{f}, \mathbf{r}) + \tilde{\mathcal{L}}_{\text{side}}(W, \tilde{\mathbf{w}}, \mathbf{r})

    where WW is the backbone network parameter set, w~=(w~(1),…,w~(M^))\tilde{\mathbf{w}} = (\tilde{w}^{(1)}, \dots, \tilde{w}^{(\hat{M})}) are side classifier weights, f=(f1,…,fM^)\mathbf{f} = (f_1, \dots, f_{\hat{M}}) are fusion layer weights, and r={rim∣i>m}\mathbf{r} = \{r_i^m \mid i > m\} are short-connection parameters across M^=6\hat{M}=6 side outputs.

    The deep supervision side loss is given by:

    L~side(W,w~,r)=∑m=1M^αml~side(m)(W,w~(m),r)\tilde{\mathcal{L}}_{\text{side}}(W, \tilde{\mathbf{w}}, \mathbf{r}) = \sum_{m=1}^{\hat{M}} \alpha_m \tilde{l}_{\text{side}}^{(m)}(W, \tilde{w}^{(m)}, \mathbf{r})

    where αm=1\alpha_m = 1 is the side loss weight, and l~side(m)\tilde{l}_{\text{side}}^{(m)} is the standard pixel-wise binary cross-entropy loss:

    l~side(m)(W,w~(m),r)=−∑j=1∣Z∣[zjlog⁡Pr⁡(zj=1∣X;W,w~(m),r)+(1−zj)log⁡Pr⁡(zj=0∣X;W,w~(m),r)]\tilde{l}_{\text{side}}^{(m)}(W, \tilde{w}^{(m)}, \mathbf{r}) = -\sum_{j=1}^{|Z|} \left[ z_j \log \operatorname{Pr}(z_j = 1 \mid X; W, \tilde{w}^{(m)}, \mathbf{r}) + (1 - z_j) \log \operatorname{Pr}(z_j = 0 \mid X; W, \tilde{w}^{(m)}, \mathbf{r}) \right]

    with Pr⁡(zj=1∣X;W,w~(m),r)=h(R~side(m)(j))\operatorname{Pr}(z_j = 1 \mid X; W, \tilde{w}^{(m)}, \mathbf{r}) = h(\tilde{R}_{\text{side}}^{(m)}(j)), where h(a)=11+e−ah(a) = \frac{1}{1 + e^{-a}} is the sigmoid function.

    The fusion loss is computed as:

    L~fuse(W,w~,f,r)=−∑j=1∣Z∣[zjlog⁡h(∑m=1M^fmR~side(m)(j))+(1−zj)log⁡(1−h(∑m=1M^fmR~side(m)(j)))]\tilde{\mathcal{L}}_{\text{fuse}}(W, \tilde{\mathbf{w}}, \mathbf{f}, \mathbf{r}) = -\sum_{j=1}^{|Z|} \left[ z_j \log h\left(\sum_{m=1}^{\hat{M}} f_m \tilde{R}_{\text{side}}^{(m)}(j)\right) + (1 - z_j) \log \left(1 - h\left(\sum_{m=1}^{\hat{M}} f_m \tilde{R}_{\text{side}}^{(m)}(j)\right)\right) \right]
  4. Knowl 4 — Multi-Side Output Inference Averaging Strategy

    model/method

    During inference, individual side saliency prediction maps are defined as Z~m=h(R~side(m))\tilde{Z}_m = h(\tilde{R}_{\text{side}}^{(m)}), where h(⋅)h(\cdot) is the element-wise sigmoid function. Because the shallowest side output (m=1m=1) contains high-frequency background noise and the deepest side output (m=6m=6) lacks spatial regularity, both m=1m=1 and m=6m=6 are excluded from the primary fusion.

    The intermediate fusion map Z~fuse\tilde{Z}_{\text{fuse}} aggregates side outputs 2, 3, and 4:

    Z~fuse=h(∑m=24fmR~side(m))\tilde{Z}_{\text{fuse}} = h\left( \sum_{m=2}^{4} f_m \tilde{R}_{\text{side}}^{(m)} \right)

    To preserve fine boundary details that may be smoothed during weighted combination, the final inferred continuous saliency map Z~final\tilde{Z}_{\text{final}} is obtained by taking the pixel-wise mean of the fusion output and individual intermediate side predictions:

    Z~final=Mean⁡(Z~fuse,Z~2,Z~3,Z~4)\tilde{Z}_{\text{final}} = \operatorname{Mean}\left(\tilde{Z}_{\text{fuse}}, \tilde{Z}_2, \tilde{Z}_3, \tilde{Z}_4\right)

    This ensembling strategy recovers missing details and consistently improves test accuracy compared to using Z~fuse\tilde{Z}_{\text{fuse}} alone.

  5. Knowl 5 — Modulated Unary Conditional Random Field for Boundary Refinement

    model/method

    To enhance spatial coherence and align saliency boundaries with image edges, a fully connected Conditional Random Field (CRF) is applied during inference. The CRF energy function over binary pixel labels x={xi}\mathbf{x} = \{x_i\} is:

    E(x)=∑iθi(xi)+∑i<jθij(xi,xj)E(\mathbf{x}) = \sum_i \theta_i(x_i) + \sum_{i < j} \theta_{ij}(x_i, x_j)

    Unlike standard CRF formulations that directly use negative log likelihoods, the unary potential incorporates a modulating scaling factor h(S^i)h(\hat{S}_i) in the denominator to increase the confidence of foreground predictions:

    θi(xi)=−log⁡S^iτh(S^i)\theta_i(x_i) = -\frac{\log \hat{S}_i}{\tau h(\hat{S}_i)}

    where S^i∈[0,1]\hat{S}_i \in [0, 1] is the normalized predicted saliency value for pixel xix_i, h(⋅)h(\cdot) is the sigmoid function, and τ=1.05\tau = 1.05 is a scale hyperparameter. This modulation decreases the Mean Absolute Error (MAE) by approximately 0.3 points by suppressing false positive noise.

    The pairwise potential consists of bilateral appearance and spatial Gaussian filters:

    θij(xi,xj)=μ(xi,xj)[w1exp⁡(−∥pi−pj∥22σα2−∥Ii−Ij∥22σβ2)+w2exp⁡(−∥pi−pj∥22σγ2)]\theta_{ij}(x_i, x_j) = \mu(x_i, x_j) \left[ w_1 \exp\left( -\frac{\|p_i - p_j\|^2}{2\sigma_\alpha^2} - \frac{\|I_i - I_j\|^2}{2\sigma_\beta^2} \right) + w_2 \exp\left( -\frac{\|p_i - p_j\|^2}{2\sigma_\gamma^2} \right) \right]

    where μ(xi,xj)=1\mu(x_i, x_j) = 1 if xi≠xjx_i \neq x_j and 00 otherwise; pip_i and IiI_i are pixel coordinates and RGB color vectors; and parameters are set to w1=3.0w_1 = 3.0, w2=3.0w_2 = 3.0, σα=60.0\sigma_\alpha = 60.0, σβ=8.0\sigma_\beta = 8.0, and σγ=5.0\sigma_\gamma = 5.0 via validation cross-validation.

  6. Knowl 6 — Saliency Existence Prediction Subnetwork

    model/method

    To prevent false detections on non-salient images (e.g., pure textures or uniform backgrounds), a classification subnetwork branch is integrated into the model to predict the existence of salient objects at the image level.

    The branch connects to the backbone convolutional features and consists of:

    1. A Global Average Pooling (GAP) layer that pools variable-sized feature maps into a fixed-length vector.
    2. A Multi-Layer Perceptron (MLP) consisting of three fully connected layers: two hidden layers with 1,024 neurons each, followed by an output layer with 2 neurons parameterized with a softmax loss.

    During joint training on balanced datasets containing background images (without salient objects) and salient images, gradients originating from non-salient background images are blocked from backpropagating into the pixel-level salient object detection branch. This gradient stopping prevents background noise from degrading pixel-level segment feature representations while achieving high existence classification accuracy (98.84%98.84\% on JSOD, 99.05%99.05\% on MSRA-B, and 96.80%96.80\% on ECSSD).

  7. Knowl 7 — Benchmark Performance of DSS on Salient Object Detection Datasets

    data/table

    The DSS model was evaluated against 11 salient object detection methods (including CNN-based methods DCL, DHS, RFCN, ELD, MC, MDF, DS and classical methods DRFI, DSR, CHM, RC) across 5 standard benchmark datasets: MSRA-B, ECSSD, HKU-IS, PASCALS, and SOD. All deep models were trained on 2,500 images from MSRA-B for fair comparison.

    Evaluation metrics are maximum F-measure (FβF_\beta with β2=0.3\beta^2 = 0.3) and Mean Absolute Error (MAE):

    Method MSRA-B ECSSD HKU-IS PASCALS SOD
    FβF_\beta MAE FβF_\beta MAE FβF_\beta MAE FβF_\beta MAE FβF_\beta MAE
    RC 0.817 0.138 0.741 0.187 0.726 0.165 0.640 0.225 0.657 0.242
    CHM 0.809 0.138 0.722 0.195 0.728 0.158 0.631 0.222 0.655 0.249
    DSR 0.812 0.119 0.737 0.173 0.735 0.140 0.646 0.204 0.655 0.234
    DRFI 0.855 0.119 0.787 0.166 0.783 0.143 0.679 0.221 0.712 0.215
    MC 0.872 0.062 0.822 0.107 0.781 0.098 0.721 0.147 0.708 0.184
    ELD 0.914 0.042 0.865 0.081 0.844 0.071 0.767 0.121 0.760 0.154
    MDF 0.885 0.104 0.833 0.108 0.860 0.129 0.764 0.145 0.785 0.155
    DS - - 0.810 0.160 - - 0.818 0.170 0.781 0.150
    RFCN 0.926 0.062 0.898 0.097 0.895 0.079 0.827 0.118 0.805 0.161
    DHS - - 0.905 0.061 0.892 0.052 0.820 0.091 0.823 0.127
    DCL+^+ 0.916 0.047 0.898 0.071 0.907 0.048 0.822 0.108 0.832 0.126
    Ours (VGG) 0.927 0.028 0.915 0.052 0.913 0.039 0.830 0.080 0.842 0.118
    Ours†^{\dagger} (ResNet-101) 0.936 0.030 0.928 0.048 0.920 0.035 0.838 0.092 0.850 0.119

    DSS with VGGNet achieves the top score across all five benchmarks among VGG-based methods. Replacing the backbone with ResNet-101 yields an additional ∼1.0\sim 1.0 point increase in FβF_\beta on average. Processing speed is 0.08 seconds per 300×400300 \times 400 image without CRF, and under 0.5 seconds per image with CRF.

  8. Knowl 8 — Ablation Study of Network Topologies and Side-Output Structures

    data/table

    Ablation experiments on the PASCALS dataset evaluate the impact of side-output depth, convolutional complexity, and top-down connection patterns.

    Scheme Architecture Configuration FβF_\beta
    1 Hypercolumns 0.818
    2 Original HED (1 conv per side output, 5 side outputs) 0.791
    3 Enhanced HED (3 convs per side output, 6 side outputs) 0.816
    4 Pattern 1 (Direct adjacent short connections: R~side(m)=rm+1mR~side(m+1)+A^side(m)\tilde{R}_{\text{side}}^{(m)} = r_{m+1}^m \tilde{R}_{\text{side}}^{(m+1)} + \hat{A}_{\text{side}}^{(m)}) 0.816
    5 Pattern 2 (Two-hop connections: R~side(m)=∑i=m+1m+2rimR~side(i)+A^side(m)\tilde{R}_{\text{side}}^{(m)} = \sum_{i=m+1}^{m+2} r_i^m \tilde{R}_{\text{side}}^{(i)} + \hat{A}_{\text{side}}^{(m)}) 0.824
    6 Pattern 3 (Proposed full top-down short connections) 0.830

    Key takeaways from side-output parameter ablations:

    1. Enhancing HED with an extra pool5 side output and 2 extra conv layers per branch improves FβF_\beta by 2.5 points over original HED (0.791 to 0.816).
    2. Adding multi-stage top-down short connections (Pattern 3) yields an additional 1.4 point gain over Enhanced HED (0.816 to 0.830).
    3. Reducing side outputs from 2 convolutional layers to 1 drops FβF_\beta by 1.5 points (0.830 to 0.815).
    4. Increasing filter channel capacities beyond baseline does not improve FβF_\beta (remains 0.830), while reducing spatial kernel sizes in deep side outputs decreases performance (0.830 to 0.820).
  9. Knowl 9 — Analysis of Training Dataset Selection and Composite Training Protocol

    empirical result

    Empirical evaluations on training dataset composition reveal that cross-dataset saliency detection performance is sensitive to the training distribution and not purely a function of dataset size:

    1. Domain Bias: When trained on a single dataset, a model almost always attains its peak performance on the matching test set (e.g., DUT-OMRON achieves Fβ=0.828F_\beta = 0.828 when trained on DUT-OMRON, but drops to 0.7200.720–0.7640.764 when trained on ECSSD, MSRA-B, or HKU-IS). Comparing models trained on different datasets therefore introduces severe benchmarking bias.
    2. Dataset Size vs. Quality: Training on a larger image collection does not guarantee superior generalizability. For example, training on ECSSD (1,000 images) achieves higher SOD test performance (Fβ=0.840F_\beta = 0.840) than training on MSRA-B (2,500 images, Fβ=0.836F_\beta = 0.836) or DUT-OMRON (3,103 images, Fβ=0.814F_\beta = 0.814).
    3. Composite Benchmark Recommendation: To establish an unbiased and uniform benchmarking protocol, the authors recommend Scheme 11, combining MSRA-B (MB), ECSSD (E), DUT-OMRON (D), and HKU-IS (H) into a 9,103-image training set. Without post-processing, this combination achieves balanced top performance across all evaluation sets: MSRA-B (Fβ=0.923F_\beta = 0.923, MAE =0.046= 0.046), HKU-IS (Fβ=0.927F_\beta = 0.927, MAE =0.042= 0.042), PASCALS (Fβ=0.844F_\beta = 0.844, MAE =0.091= 0.091), SOD (Fβ=0.864F_\beta = 0.864, MAE =0.113= 0.113), and DUT-OMRON (Fβ=0.843F_\beta = 0.843, MAE =0.051= 0.051).
  10. Knowl 10 — Structural Failure Modes in Deep Salient Object Detection

    limitation

    Analysis of failure cases in the DSS framework reveals three recurring failure modes common to deep convolutional salient object detection:

    1. Incomplete Object Coverage: Salient objects are occasionally only partially segmented, leaving peripheral components or appendages unextracted due to spatial resolution loss in convolutional downsampling.
    2. Low Contrast and Background Clutter: In scenes with complex background textures or minimal foreground-background color contrast, non-salient background regions are mistakenly predicted as salient, or the main body of the salient object is missed.
    3. Transparent and Semi-Transparent Objects: Fully detecting and segmenting transparent objects (such as glass containers) remains difficult because internal features blend directly with background content, causing fragmented saliency activations.

    Proposed remedies include integrating segment-level region priors/superpixel voting to enforce texture/color consistency, training on more complex cluttered scenes, and developing more expressive feature representations.

Coverage note — None was omitted; all key architectural components, short connection formulations, loss functions, CRF and existence subnetwork designs, ablation studies, benchmark comparisons, dataset analyses, and failure analyses are covered.

References

  1. 1.Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P. Torr, “Deeply supervised salient object detection with short connections,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  2. 2.C. Guo and L. Zhang, “A novel multiresolution spatiotemporal saliency detection model and its applications in image and video compression,” IEEE Trans. Image Process., vol. 19, no. 1, pp. 185–198, 2010.
  3. 3.J. Guo, T. Ren, L. Huang, X. Liu, M.-M. Cheng, and G. Wu, “Video salient object detection via cross-frame cellular automata,” in Int. Conf. Multimedia and Expo, 2017, pp. 325–330.
  4. 4.M. Donoser, M. Urschler, M. Hirzer, and H. Bischof, “Saliency driven total variation segmentation,” in Int. Conf. Comput. Vis., 2009, pp. 817–824.
  5. 5.M.-M. Cheng, F.-L. Zhang, N. J. Mitra, X. Huang, and S.-M. Hu, “Repfinder: finding approximately repeated scene elements for image editing,” in ACM Trans. Graph., vol. 29, no. 4, 2010, p. 83.
  6. 6.G.-X. Zhang, M.-M. Cheng, S.-M. Hu, and R. R. Martin, “A shape-preserving approach to image resizing,” Comput. Graph. Forum, vol. 28, no. 7, pp. 1897–1906, 2009.
  7. 7.U. Rutishauser, D. Walther, C. Koch, and P. Perona, “Is bottom-up attention useful for object recognition?” in IEEE Conf. Comput. Vis. Pattern Recog., 2004.
  8. 8.Y. Wei, X. Liang, Y. Chen, X. Shen, M.-M. Cheng, Y. Zhao, and S. Yan, “Stc: A simple to complex framework for weakly-supervised semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 11, pp. 2314–2320, 2017.
  9. 9.Q. Hou, P. K. Dokania, D. Massiceti, Y. Wei, M.-M. Cheng, and P. H. S. Torr, “Mining pixels: Weakly supervised semantic segmentation using image labels,” in EMMCVPR, 2017.
  10. 10.Y. Wei, J. Feng, X. Liang, M.-M. Cheng, Y. Zhao, and S. Yan, “Object region mining with adversarial erasing: A simple classification to semantic segmentation approach,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  11. 11.Y. Wei, X. Liang, Y. Chen, Z. Jie, Y. Xiao, Y. Zhao, and S. Yan, “Learning to segment with image-level annotations,” Pattern Recognition, vol. 59, pp. 234–244, 2016.
  12. 12.A. Borji, S. Frintrop, D. N. Sihite, and L. Itti, “Adaptive object tracking by learning background context,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2012.
  13. 13.P. L. Rosin and Y.-K. Lai, “Artistic minimal rendering with lines and blocks,” Graphical Models, vol. 75, no. 4, pp. 208–229, 2013.
  14. 14.J. Han, E. J. Pauwels, and P. De Zeeuw, “Fast saliency-aware multi-modality image fusion,” Neurocomputing, pp. 70–80, 2013.
  15. 15.T. Chen, M.-M. Cheng, P. Tan, A. Shamir, and S.-M. Hu, “Sketch2photo: Internet image montage,” ACM Trans. Graph., vol. 28, no. 5, pp. 124:1–10, 2009.
  16. 16.S.-M. Hu, T. Chen, K. Xu, M.-M. Cheng, and R. R. Martin, “Internet visual media processing: a survey with graphics and vision applications,” The Vis. Comput., vol. 29, no. 5, pp. 393–405, 2013.
  17. 17.H. Liu, L. Zhang, and H. Huang, “Web-image driven best views of 3d shapes,” The Vis. Comput., pp. 1–9, 2012.
  18. 18.J.-Y. Zhu, J. Wu, Y. Wei, E. Chang, and Z. Tu, “Unsupervised object class discovery via saliency-guided multiple class learning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012, pp. 3218–3225.
  19. 19.Y. Gao, M. Wang, D. Tao, R. Ji, and Q. Dai, “3-D object retrieval and recognition with hypergraph analysis,” IEEE Trans. Image Process., vol. 21, no. 9, pp. 4290–4303, 2012.
  20. 20.M.-M. Cheng, Q.-B. Hou, S.-H. Zhang, and P. L. Rosin, “Intelligent visual media processing: When graphics meets vision,” J. Comput. Sci. Tech., 2017.
  21. 21.A. Abdulmunem, Y.-K. Lai, and X. Sun, “Saliency guided local and global descriptors for effective action recognition,” Computational Visual Media, vol. 2, no. 1, pp. 97–106, 2016.
  22. 22.L. Itti and C. Koch, “Computational modeling of visual attention,” Nature reviews neuroscience, vol. 2, no. 3, pp. 194–203, 2001.
  23. 23.A. Borji, M.-M. Cheng, H. Jiang, and J. Li, “Salient object detection: A benchmark,” IEEE Trans. Image Process., vol. 24, no. 12, pp. 5706–5722, 2015.
  24. 24.A. Borji and L. Itti, “State-of-the-art in visual attention modeling,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 1, pp. 185–207, 2013.
  25. 25.J. Wang, H. Jiang, Z. Yuan, M.-M. Cheng, X. Hu, and N. Zheng, “Salient object detection: A discriminative regional feature integration approach,” Int. J. Comput. Vis., vol. 123, no. 2, pp. 251–268, 2017, http://people.cs.umass.edu/~hzjiang/.
  26. 26.S. Xie and Z. Tu, “Holistically-nested edge detection,” Int. J. Comput. Vis., vol. 125, no. 1, pp. 3–18, 2017.
  27. 27.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Adv. Neural Inform. Process. Syst., 2012, pp. 1097–1105.
  28. 28.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Int. Conf. Learn. Represent., 2015.
  29. 29.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 3431–3440.
  30. 30.Y. Liu, M.-M. Cheng, X. Hu, K. Wang, and X. Bai, “Richer convolutional features for edge detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  31. 31.J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan, “Perceptual generative adversarial networks for small object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  32. 32.R. Girshick, “Fast r-cnn,” in Int. Conf. Comput. Vis., 2015, pp. 1440–1448.
  33. 33.L. Zhang, L. Lin, X. Liang, and K. He, “Is faster r-cnn doing well for pedestrian detection?” in Eur. Conf. Comput. Vis., 2016, pp. 443–457.
  34. 34.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” P. IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  35. 35.G. Li and Y. Yu, “Deep contrast learning for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  36. 36.N. Liu and J. Han, “Dhsnet: Deep hierarchical saliency network for salient object detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 678–686.
  37. 37.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik, “Hypercolumns for object segmentation and fine-grained localization,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 447–456.
  38. 38.L. Itti, C. Koch, and E. Niebur, “A model of saliency-based visual attention for rapid scene analysis,” IEEE Trans. Pattern Anal. Mach. Intell., no. 11, pp. 1254–1259, 1998.
  39. 39.Y. Xie, H. Lu, and M.-H. Yang, “Bayesian saliency via low and mid level cues,” IEEE Trans. Image Process., vol. 22, no. 5, pp. 1689–1698, 2013.
  40. 40.W. Qi, M.-M. Cheng, A. Borji, H. Lu, and L.-F. Bai, “Saliencyrank: Two-stage manifold ranking for salient object detection,” Computational Visual Media, vol. 1, no. 4, pp. 309–320, 2015.
  41. 41.M.-M. Cheng, N. J. Mitra, X. Huang, P. H. S. Torr, and S.-M. Hu, “Global contrast based salient region detection,” vol. 37, no. 3, pp. 569–582, 2015.
  42. 42.M.-M. Cheng, J. Warrell, W.-Y. Lin, S. Zheng, V. Vineet, and N. Crook, “Efficient salient region detection with soft image abstraction,” in Int. Conf. Comput. Vis., 2013, pp. 1529–1536.
  43. 43.T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum, “Learning to detect a salient object,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 2, pp. 353–367, 2011.
  44. 44.A. Borji and L. Itti, “Exploiting local and global patch rarities for saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012, pp. 478–485.
  45. 45.A. Borji, M.-M. Cheng, Q. Hou, H. Jiang, and J. Li, “Salient object detection: A survey,” arXiv preprint arXiv:1411.5878, 2014.
  46. 46.S. He, R. Lau, W. Liu, Z. Huang, and Q. Yang, “Supercnn: A superpixelwise convolutional neural network for salient object detection,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 330–344, 2015, http://www.shengfenghe.com/.
  47. 47.G. Li and Y. Yu, “Visual saliency based on multiscale deep features,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 5455–5463, http://i.cs.hku.hk/~yzyu/vision.html.
  48. 48.L. Wang, H. Lu, X. Ruan, and M.-H. Yang, “Deep networks for saliency detection via local estimation and global search,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 3183–3192.
  49. 49.R. Zhao, W. Ouyang, H. Li, and X. Wang, “Saliency detection by multi-context deep learning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2015, pp. 1265–1274, https://github.com/Robert0812/deepsaldet.
  50. 50.L. Gayoung, T. Yu-Wing, and K. Junmo, “Deep saliency with encoded low level distance map and high level features,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, https://github.com/gylee1103/SaliencyELD.
  51. 51.L. Wang, L. Wang, H. Lu, P. Zhang, and X. Ruan, “Saliency detection with recurrent fully convolutional networks,” in Eur. Conf. Comput. Vis., 2016.
  52. 52.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in ACM Int. Conf. Multimedia, 2014, pp. 675–678.
  53. 53.P. Krähenbühl and V. Koltun, “Efficient inference in fully connected crfs with gaussian edge potentials,” in Adv. Neural Inform. Process. Syst., 2011.
  54. 54.Q. Yan, L. Xu, J. Shi, and J. Jia, “Hierarchical saliency detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2013, pp. 1155–1162.
  55. 55.Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, “The secrets of salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2014, pp. 280–287.
  56. 56.D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Int. Conf. Comput. Vis., 2001, pp. 416–423.
  57. 57.V. Movahedi and J. H. Elder, “Design and perceptual validation of performance measures for salient object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2010, pp. 49–56.
  58. 58.F. Perazzi, P. Krähenbühl, Y. Pritch, and A. Hornung, “Saliency filters: Contrast based filtering for salient region detection,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012, pp. 733–740.
  59. 59.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  60. 60.X. Li, L. Zhao, L. Wei, M.-H. Yang, F. Wu, Y. Zhuang, H. Ling, and J. Wang, “Deepsaliency: Multi-task deep neural network model for salient object detection,” IEEE Trans. Image Process., vol. 25, no. 8, pp. 3919 – 3930, 2016, https://github.com/zlmzju/DeepSaliency.
  61. 61.X. Li, Y. Li, C. Shen, A. Dick, and A. Van Den Hengel, “Contextual hypergraph modeling for salient object detection,” in Int. Conf. Comput. Vis., 2013, pp. 3328–3335.
  62. 62.X. Li, H. Lu, L. Zhang, X. Ruan, and M.-H. Yang, “Saliency detection via dense and sparse reconstruction,” in Int. Conf. Comput. Vis., 2013, pp. 2976–2983.
  63. 63.H. Jiang, M.-M. Cheng, S.-J. Li, A. Borji, and J. Wang, “Joint salient object detection and existence prediction,” Front. Comput. Sci., 2017.
  64. 64.P. Wang, J. Wang, G. Zeng, J. Feng, H. Zha, and S. Li, “Salient object detection for searched web images via global saliency,” in IEEE Conf. Comput. Vis. Pattern Recog., 2012.
  65. 65.D.-P. Fan, M.-M. Cheng, Y. Liu, T. Li, and A. Borji, “Structure-measure: A New Way to Evaluate Foreground Maps,” in Int. Conf. Comput. Vis., 2017.

Citation

MLA
Hou, Q., et al. “Deeply Supervised Salient Object Detection with Short Connections”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 4, 2019, pp. 815–28, https://doi.org/10.1109/TPAMI.2018.2815688.
APA
Hou, Q., Cheng, M.-M., Hu, X., Borji, A., Tu, Z., & Torr, P. H. S. (2019). Deeply Supervised Salient Object Detection with Short Connections. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(4), 815–828. https://doi.org/10.1109/TPAMI.2018.2815688
Chicago
Hou, Q., M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P. H. S. Torr. 2019. “Deeply Supervised Salient Object Detection with Short Connections”. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (4): 815–28. https://doi.org/10.1109/TPAMI.2018.2815688.
Harvard
Hou, Q. et al. (2019) “Deeply Supervised Salient Object Detection with Short Connections”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(4), pp. 815–828. Available at: https://doi.org/10.1109/TPAMI.2018.2815688.
Vancouver
1. Hou Q, Cheng M-M, Hu X, Borji A, Tu Z, Torr PHS (2019) Deeply Supervised Salient Object Detection with Short Connections. IEEE Transactions on Pattern Analysis and Machine Intelligence 41:815–828

BibTeX

@article{Hou_2019, title={Deeply Supervised Salient Object Detection with Short Connections}, volume={41}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2018.2815688}, DOI={10.1109/tpami.2018.2815688}, number={4}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Hou, Qibin and Cheng, Ming-Ming and Hu, Xiaowei and Borji, Ali and Tu, Zhuowen and Torr, Philip H. S.}, year={2019}, month=Apr, pages={815–828} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF