Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

Yi WangRuili WangXin FanTianzhu WangXiangjian He

article2023CVPR125 citations

Proposes Multiple Enhancement Network (MENet), which integrates human visual system mechanisms through a dual-branch decoder, multiscale feature enhancement modules, and a multi-level hybrid loss across pixel, region, and object scales to achieve state-of-the-art salient object detection in complex scenes.

Listen

Salient object detection aims to automatically identify and segment the most visually prominent objects in digital images, mimicking human vision. While this capability is critical for downstream computer vision tasks such as automated surveillance, image editing, video summarization, and autonomous navigation, existing models often struggle in cluttered scenes, low-contrast environments, and situations involving complex object boundaries. The article addresses these shortcomings by introducing the Multiple Enhancement Network (MENet), a deep learning architecture designed to significantly improve object detection accuracy and boundary precision.

The main objective of the article is to design, implement, and validate a novel network architecture and training scheme inspired by human visual cognition. Specifically, the article demonstrates how combining multi-level loss supervision with dual-stream iterative refinement produces cleaner object contours and superior overall segmentation accuracy compared to existing state-of-the-art approaches.

To accomplish this, the authors constructed an encoder-decoder architecture that decouples image processing into two parallel, non-interacting streams: one dedicated to high-frequency edge details and the other focused on low-frequency interior body regions. A flexible multiscale feature enhancement module alters input sequences across four iterative training rounds to alternately refine global context and fine details. Additionally, the network is trained using a composite loss function that simultaneously evaluates pixel-level accuracy, regional consistency across image quadrants, and object-level foreground contrast. The model was trained on standard benchmark datasets and thoroughly evaluated against 16 leading algorithms across six public image datasets representing varying levels of visual complexity.

The evaluation yielded several key findings. First, MENet consistently outperformed competing models across all six standard benchmark datasets, delivering the lowest prediction error and highest boundary fidelity on major benchmarks such as DUTS-TE, DUT-OMRON, and HKU-IS. Second, an ablation study confirmed that four iterative enhancement rounds optimize performance, striking the best balance between broad contextual discovery and fine-grained boundary refinement. Third, introducing region-level and object-level similarity constraints substantially improved segmentation completeness; partitioning the image into four regional quadrants reduced mean absolute error by approximately 4.7% to 8.2% across the datasets compared to single-region evaluation. Finally, the model maintained practical efficiency, processing test images at roughly 45 frames per second.

These findings indicate that incorporating human visual principles—such as separating boundary detection from interior region analysis and evaluating spatial consistency at multiple scales—overcomes major limitations in computer vision systems. By generating sharper boundaries without sacrificing processing speed, this approach provides a viable, high-performance visual processing component that can enhance real-time vision applications while reducing downstream errors caused by noisy segmentations.

Stakeholders and engineering teams developing automated vision pipelines should consider adopting dual-stream architectures and multi-level loss formulations to enhance segmentation quality. Future technical efforts should focus on validating and adapting this framework for specialized, high-stakes environments, such as medical imagery or autonomous driving under challenging lighting. While the article establishes high confidence in MENet's performance across standard benchmarks, the authors note that heavily blurred and extremely low-contrast real-world scenes remain an ongoing challenge that warrants continued research.

No sufficiently relevant recommendations were found.

Cover for Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

Abstract

Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Framework Overview
  • 3.2. Multiscale Feature Enhancement
  • 3.3. Iterative Enhancement
  • 4. Supervision Strategy
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Comparison with the state-of-the-arts
  • 5.3. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Architecture of the Multiple Enhancement Network

    model/method

    The Multiple Enhancement Network (MENet) is an encoder-decoder architecture designed for salient object detection (SOD) in complex scenes. It integrates mechanisms of the human visual system (HVS), including frequency decomposition, dual-branch boundary/body refinement, and iterative feature aggregation.

    Given an input image I∈R3×W×HI \in \mathbb{R}^{3 \times W \times H}, a convolutional backbone (e.g., ResNet-50) extracts N=5N=5 multi-stage feature maps B={bi}i=1NB = \{b_i\}_{i=1}^N. These multi-stage features are compressed into 64-channel feature blocks gB={gbi}i=1NgB = \{gb_i\}_{i=1}^N via a Gradient Feature Encoder (GF-Encoder) and aB={abi}i=1NaB = \{ab_i\}_{i=1}^N via an Adaptive Feature Encoder (AF-Encoder).

    MENet decomposes saliency feature learning into two non-interacting streams:

    1. Gradient Branch (GME): Processes gBgB under boundary gradient supervision Lg(k)\mathcal{L}_g^{(k)} to learn high-frequency contour features EgB(k)={Egbi(k)}i=1NEgB^{(k)} = \{Egb_i^{(k)}\}_{i=1}^N.
    2. Adaptive Branch (AME): Processes aBaB to capture low-frequency internal body features EaB(k)={Eabi(k)}i=1NEaB^{(k)} = \{Eab_i^{(k)}\}_{i=1}^N without direct branch-specific supervision.

    There is deliberately no intermediate cross-talk between the two branches during propagation to prevent noisy boundary estimates from corrupting global consistency. At iteration kk, the enhanced features are concatenated to generate the saliency map: S(k)=Linear(Concat(EgB(k),EaB(k)))S^{(k)} = \text{Linear}(\text{Concat}(EgB^{(k)}, EaB^{(k)}))

    For subsequent iterations, EgB(k)EgB^{(k)} and EaB(k)EaB^{(k)} are combined into EagB=Concat(EgB(k),EaB(k))EagB = \text{Concat}(EgB^{(k)}, EaB^{(k)}), re-squeezed to 64 channels by an Enhanced GF-Encoder and Enhanced AF-Encoder, and re-fed into GME and AME alongside the original backbone features.

  2. Knowl 2 — Multiscale Feature Enhancement Module

    model/method

    The Multiscale Feature Enhancement Module (ME-Module) is a five-stage aggregation block that acts as either a global context enhancer or a detail enhancer depending on the spatial resolution ordering of its inputs.

    The module receives input features from three ports:

    • Port #1: Squeezed backbone features from the current stream (gBgB or aBaB).
    • Port #2: Enhanced features from iteration k−1k-1 (EgB(k−1)EgB^{(k-1)} or EaB(k−1)EaB^{(k-1)}).
    • Port #3: Enhanced features from iteration k−2k-2 (EgB(k−2)EgB^{(k-2)} or EaB(k−2)EaB^{(k-2)}).

    At stage j∈{1,2,3,4,5}j \in \{1, 2, 3, 4, 5\}, the sub-features FbijFbij from each of the three ports are fused with the output of stage j−1j-1 via pixel-wise addition. The combined representation then undergoes:

    1. Atrous Spatial Pyramid Pooling (ASPP): Expands the receptive field and captures multi-scale contextual dependencies.
    2. ME-Attention Module: Applies global and local attention to selectively emphasize salient object positions across spatial and channel dimensions.
    3. Interpolator: Performs scale-matched upsampling or downsampling to align spatial dimensions with the requirement of stage j+1j+1.

    By reversing the order of the multi-scale input feature list (sorting from largest-to-smallest versus smallest-to-largest), the ME-Module flexibly alternates between top-down global feature propagation and bottom-up detail refinement.

  3. Knowl 3 — Iterative Enhancement Algorithm for MENet

    algorithm

    MENet applies a four-round alternating iterative enhancement process to mimic the human visual system's alternating focus between overall scene context and fine boundaries.

    Let V(⋅)V(\cdot) denote an operator that reverses the list ordering of multi-scale feature maps. The initial 64-channel backbone feature blocks gBgB and aBaB have their sub-features ordered from smallest to largest spatial resolution.

    Input: Backbone gradient features gB={gb1,…,gb5}gB = \{gb_1, \dots, gb_5\}, adaptive features aB={ab1,…,ab5}aB = \{ab_1, \dots, ab_5\}, iteration limit K=4K = 4
    Output: Final saliency map S(K)S^{(K)}
    Initialize EgB(0)←NullEgB^{(0)} \leftarrow \text{Null}, EgB(−1)←NullEgB^{(-1)} \leftarrow \text{Null}
    Initialize EaB(0)←NullEaB^{(0)} \leftarrow \text{Null}, EaB(−1)←NullEaB^{(-1)} \leftarrow \text{Null}
    for k=1k = 1 to KK do
        EgB(k)←GME(V(gB),V(EgB(k−1)),V(EgB(k−2)))EgB^{(k)} \leftarrow \text{GME}(V(gB), V(EgB^{(k-1)}), V(EgB^{(k-2)}))
        EaB(k)←AME(V(aB),V(EaB(k−1)),V(EaB(k−2)))EaB^{(k)} \leftarrow \text{AME}(V(aB), V(EaB^{(k-1)}), V(EaB^{(k-2)}))
        S(k)←Linear(Concat(EgB(k),EaB(k)))S^{(k)} \leftarrow \text{Linear}(\text{Concat}(EgB^{(k)}, EaB^{(k)}))
        if k<Kk < K then
            EagB←Concat(EgB(k),EaB(k))EagB \leftarrow \text{Concat}(EgB^{(k)}, EaB^{(k)})
            gB←EnhancedGFEncoder(EagB)gB \leftarrow \text{EnhancedGFEncoder}(EagB)
            aB←EnhancedAFEncoder(EagB)aB \leftarrow \text{EnhancedAFEncoder}(EagB)
        end if
    end for
    return S(K)S^{(K)}

    Odd iterations (k=1,3k=1, 3) reverse the feature sequence to process from largest to smallest, extracting comprehensive global semantic context. Even iterations (k=2,4k=2, 4) reverse the sequence again to process from smallest to largest, refining sharp local boundary details.

  4. Knowl 4 — Overall Supervision Objective and Multi-Level Hybrid Loss

    equation

    The overall training objective L\mathcal{L} of MENet is the sum of gradient supervision Lg(k)\mathcal{L}_g^{(k)} and multi-level hybrid saliency loss Ls(k)\mathcal{L}_s^{(k)} across all K=4K=4 enhancement rounds: L=∑k=14(α1(k)Lg(k)+α2(k)Ls(k))\mathcal{L} = \sum_{k=1}^{4} \left( \alpha_1^{(k)} \mathcal{L}_g^{(k)} + \alpha_2^{(k)} \mathcal{L}_s^{(k)} \right) where balancing weights are set to α1(k)=α2(k)=0.5\alpha_1^{(k)} = \alpha_2^{(k)} = 0.5 for all k∈{1,2,3,4}k \in \{1, 2, 3, 4\}.

    The multi-level hybrid loss Ls\mathcal{L}_s evaluates pixel-level, region-level, and object-level similarities between the predicted saliency map S∈[0,1]W×HS \in [0, 1]^{W \times H} and the ground-truth mask G∈{0,1}W×HG \in \{0, 1\}^{W \times H}: Ls=β1Lsbce+β2Lsreg+β3Lsobj\mathcal{L}_s = \beta_1 \mathcal{L}_{s_{bce}} + \beta_2 \mathcal{L}_{s_{reg}} + \beta_3 \mathcal{L}_{s_{obj}} where the hyperparameters are set to β1=0.4\beta_1 = 0.4, β2=0.4\beta_2 = 0.4, and β3=0.2\beta_3 = 0.2.

    The pixel-level loss Lsbce\mathcal{L}_{s_{bce}} is standard Binary Cross-Entropy: Lsbce=−∑x=1W∑y=1H(G(x,y)log⁡S(x,y)+(1−G(x,y))log⁡(1−S(x,y)))\mathcal{L}_{s_{bce}} = -\sum_{x=1}^W \sum_{y=1}^H \left( G(x,y) \log S(x,y) + (1 - G(x,y)) \log (1 - S(x,y)) \right)

  5. Knowl 5 — Region-Level Saliency Loss Formulation

    equation

    The region-level loss Lsreg\mathcal{L}_{s_{reg}} enforces structural and overlap consistency by partitioning both the predicted saliency map SS and ground-truth mask GG into four equal, non-overlapping quadrants SiS_i and GiG_i (i∈{1,2,3,4}i \in \{1, 2, 3, 4\}): Lsreg=1−∑i=14ωi(θ1SSIMi+θ2IoUi)\mathcal{L}_{s_{reg}} = 1 - \sum_{i=1}^{4} \omega_i (\theta_1 \text{SSIM}_i + \theta_2 \text{IoU}_i) where θ1=0.5\theta_1 = 0.5, θ2=0.5\theta_2 = 0.5, and ωi\omega_i is the weight assigned to quadrant ii, defined as the ratio of predicted foreground area to ground-truth foreground area in that sub-region.

    Structural Similarity for region ii is formulated as the product of luminance, contrast, and structure comparisons: SSIMi=(2μSiμGi+c1μSi2+μGi2+c1)⋅(2σSiσGi+c2σSi2+σGi2+c2)⋅(σSiGi+c3σSiσGi+c3)\text{SSIM}_i = \left( \frac{2 \mu_{S_i} \mu_{G_i} + c_1}{\mu_{S_i}^2 + \mu_{G_i}^2 + c_1} \right) \cdot \left( \frac{2 \sigma_{S_i} \sigma_{G_i} + c_2}{\sigma_{S_i}^2 + \sigma_{G_i}^2 + c_2} \right) \cdot \left( \frac{\sigma_{S_i G_i} + c_3}{\sigma_{S_i} \sigma_{G_i} + c_3} \right) where μSi\mu_{S_i} and μGi\mu_{G_i} are region means, σSi\sigma_{S_i} and σGi\sigma_{G_i} are standard deviations, σSiGi\sigma_{S_i G_i} is the covariance between SiS_i and GiG_i, and c1,c2,c3c_1, c_2, c_3 are small stabilization constants.

    The regional Intersection over Union is defined as: IoUi=∑p∈region iSi(p)Gi(p)∑p∈region i(Gi(p)+Si(p)−Gi(p)Si(p))\text{IoU}_i = \frac{\sum_{p \in \text{region } i} S_i(p) G_i(p)}{\sum_{p \in \text{region } i} (G_i(p) + S_i(p) - G_i(p) S_i(p))}

  6. Knowl 6 — Object-Level Saliency Loss Formulation

    equation

    The object-level loss Lsobj\mathcal{L}_{s_{obj}} measures the holistic contrast and distribution uniformity of salient object foregrounds. It models foreground-background contrast via the luminance component of SSIM and foreground distribution uniformity via the coefficient of variation (ratio of standard deviation to mean).

    Let SoS_o and GoG_o represent the predicted and ground-truth saliency values restricted strictly to the foreground pixels. The loss is formulated as: Lsobj=1−1μSo2+μGo22μSoμGo+λσSoμSo\mathcal{L}_{s_{obj}} = 1 - \frac{1}{\frac{\mu_{S_o}^2 + \mu_{G_o}^2}{2\mu_{S_o}\mu_{G_o}} + \lambda \frac{\sigma_{S_o}}{\mu_{S_o}}} where μSo\mu_{S_o} and μGo\mu_{G_o} are the means of SoS_o and GoG_o, σSo\sigma_{S_o} is the standard deviation of SoS_o, and λ\lambda is a weighting coefficient.

    Because the ground-truth foreground mean is binary normalized such that μGo=1\mu_{G_o} = 1, the expression simplifies in implementation to: Lsobj=1−2μSoμSo2+1+2λσSo\mathcal{L}_{s_{obj}} = 1 - \frac{2\mu_{S_o}}{\mu_{S_o}^2 + 1 + 2\lambda \sigma_{S_o}}

  7. Knowl 7 — Gradient Boundary Supervision Loss

    equation

    Gradient supervision guides the Gradient ME-Module (GME) to learn sharp boundary representations corresponding to high spatial frequencies.

    Let Gg∈[0,1]W×HG_g \in [0, 1]^{W \times H} denote the gradient map extracted from the ground-truth mask GG, and let Sg∈[0,1]W×HS_g \in [0, 1]^{W \times H} denote the predicted gradient map produced by the GME branch at iteration kk. The gradient loss Lg\mathcal{L}_g is defined as a pixel-wise Binary Cross-Entropy loss: Lg=−∑x=1W∑y=1H(Gg(x,y)log⁡Sg(x,y)+(1−Gg(x,y))log⁡(1−Sg(x,y)))\mathcal{L}_g = -\sum_{x=1}^W \sum_{y=1}^H \left( G_g(x,y) \log S_g(x,y) + (1 - G_g(x,y)) \log (1 - S_g(x,y)) \right)

  8. Knowl 8 — Experimental Settings and Training Details for MENet

    experimental setup

    MENet is implemented in PyTorch 1.12 and trained on an NVIDIA A100 GPU (40 GB memory) with an AMD EPYC 7742 CPU.

    • Backbone & Pre-training: ResNet-50 initialized with ImageNet pre-trained weights; other parameters initialized randomly from a normal distribution.
    • Training Dataset: Fine-tuned on the DUTS-TR dataset (10,553 images).
    • Data Augmentation: Multi-scale input scaling to dimensions [352×352][352 \times 352], [320×320][320 \times 320], [288×288][288 \times 288], [256×256][256 \times 256], and [224×224][224 \times 224].
    • Optimizer: Stochastic Gradient Descent (SGD) with momentum 0.9, weight decay 0.0005, batch size 24, trained for 99 epochs.
    • Learning Rate Schedule: 'Poly' learning rate decay strategy with maximum learning rate 0.00025 for the backbone and 0.0025 for all other modules.
    • Inference Speed: 0.022 seconds per 352×352352 \times 352 image (45 frames per second).
    • Benchmark Test Datasets: DUT-OMRON (5,168 images), DUTS-TE (5,019 images), HKU-IS (4,447 images), PASCAL-S (850 images), ECSSD (1,000 images), and SOD (300 images).
    • Evaluation Metrics: Mean Absolute Error (MAE ↓\downarrow), maximum F-measure (MaxF↑\text{MaxF} \uparrow), mean F-measure (mF↑\text{mF} \uparrow), mean Enhanced-alignment measure (mEm↑\text{mEm} \uparrow), and S-measure (Sm↑S_m \uparrow).
  9. Knowl 9 — Quantitative SOD Benchmark Results of MENet

    data/table

    MENet was benchmarked against 16 state-of-the-art salient object detection methods across six standard datasets. MENet with a ResNet-50 backbone achieves the overall highest scores across metrics, consistently obtaining the lowest MAE values.

    Dataset Metric BASNet MINet (Res50) LDF ICON (Res50) MENet (Res50)
    DUT-OMRON MAE ↓\downarrow 0.0565 0.0559 0.0517 0.0569 0.0450
    MaxF ↑\uparrow 0.8053 0.8098 0.8199 0.8254 0.8337
    mEm ↑\uparrow 0.8691 0.8734 0.8814 0.8791 0.8911
    Sm ↑\uparrow 0.8362 0.8329 0.8392 0.8445 0.8496
    DUTS-TE MAE ↓\downarrow 0.0472 0.0373 0.0336 0.0370 0.0281
    MaxF ↑\uparrow 0.8589 0.8833 0.8968 0.8917 0.9123
    mEm ↑\uparrow 0.8790 0.9132 0.9232 0.9142 0.9368
    Sm ↑\uparrow 0.8660 0.8842 0.8924 0.8889 0.9049
    HKU-IS MAE ↓\downarrow 0.0322 0.0292 0.0275 0.0289 0.0234
    MaxF ↑\uparrow 0.9284 0.9349 0.9394 0.9395 0.9483
    mEm ↑\uparrow 0.9458 0.9600 0.9597 0.9585 0.9657
    Sm ↑\uparrow 0.9090 0.9189 0.9196 0.9202 0.9274
    PASCAL-S MAE ↓\downarrow 0.0758 0.0643 0.0596 0.0644 0.0535
    MaxF ↑\uparrow 0.8539 0.8665 0.8741 0.8757 0.8896
    mEm ↑\uparrow 0.8527 0.8981 0.9048 0.8931 0.9132
    Sm ↑\uparrow 0.8380 0.8563 0.8630 0.8611 0.8721
    ECSSD MAE ↓\downarrow 0.0370 0.0342 0.0335 0.0318 0.0307
    MaxF ↑\uparrow 0.9425 0.9475 0.9501 0.9503 0.9549
    mEm ↑\uparrow 0.9210 0.9532 0.9509 0.9543 0.9544
    Sm ↑\uparrow 0.9163 0.9250 0.9245 0.9290 0.9279
    SOD MAE ↓\downarrow 0.1124 - - 0.0841 0.0874
    MaxF ↑\uparrow 0.8487 - - 0.8790 0.8780
    mEm ↑\uparrow 0.7793 - - 0.8516 0.8381
    Sm ↑\uparrow 0.7721 - - 0.8238 0.8089
  10. Knowl 10 — Ablation Analysis of Iteration Count and Multi-Level Hybrid Loss Components

    empirical result

    Ablation experiments evaluate the impact of the number of iterative enhancements and the loss function components on MENet performance:

    1. Number of Iterative Enhancements: Comparing 1, 2, 3, and 4 iterations across all datasets confirms that 4 iterations achieve optimal results. On DUTS-TE, MAE drops from 0.0496 (1 iteration) to 0.0310 (2 iterations), temporarily rises to 0.0389 (3 iterations) as global features are re-extracted, and achieves its minimum at 0.0281 (4 iterations). MaxF similarly increases from 0.8560 (1 iteration) to 0.9123 (4 iterations).

    2. Contribution of Loss Components:

      • Starting from baseline pixel BCE loss Lsbce\mathcal{L}_{s_{bce}} alone on DUTS-TE yields MAE=0.0308,MaxF=0.9030\text{MAE} = 0.0308, \text{MaxF} = 0.9030.
      • Adding region-level loss Lsreg\mathcal{L}_{s_{reg}} with four sub-regions improves MAE to 0.0305.
      • Incorporating object-level loss Lsobj\mathcal{L}_{s_{obj}} alongside Lsreg\mathcal{L}_{s_{reg}} improves MAE to 0.0295 and MaxF to 0.9097.
      • Including gradient supervision Lg\mathcal{L}_g with all components achieves the best performance: MAE=0.0281,MaxF=0.9123,mEm=0.9368,Sm=0.9049\text{MAE} = 0.0281, \text{MaxF} = 0.9123, \text{mEm} = 0.9368, S_m = 0.9049.
    3. Sub-Region Partitioning in Region Loss: Partitioning the region-level loss Lsreg\mathcal{L}_{s_{reg}} into 4 sub-regions versus treating the image as 1 single region reduces MAE across all six datasets: by 4.66% on DUT-OMRON (0.0450 vs 0.0472), 4.75% on DUTS-TE (0.0281 vs 0.0295), 8.24% on HKU-IS (0.0234 vs 0.0255), 7.12% on PASCAL-S (0.0535 vs 0.0576), 5.83% on ECSSD (0.0307 vs 0.0326), and 7.42% on SOD (0.0874 vs 0.0944).

Coverage note — None was omitted; all contributed aspects including model architecture, enhancement modules, loss functions, algorithms, experimental setup, benchmarking, and ablation analyses are completely covered.

References

  1. 1.Illya Bakurov, Marco Buzzelli, Raimondo Schettini, Mauro Castelli, and Leonardo Vanneschi. Structural similarity index (ssim) revisited: A data-driven approach. Expert Syst. Appl., 189:116087, 2022. 5
  2. 2.Léon Bottou. Stochastic gradient descent tricks. In Neural networks: Tricks of the trade, pages 421–436. Springer, 2012. 5
  3. 3.Tao Chen, Yazhou Yao, Lei Zhang, Qiong Wang, Guosen Xie, and Fumin Shen. Saliency guided inter-and intra-class relation constraints for weakly supervised semantic segmentation. IEEE TMM, 2022. 1
  4. 4.Zhe Chen, Ruili Wang, Zhen Zhang, Huibin Wang, and Lizhong Xu. Background–foreground interaction for moving object detection in dynamic scenes. Inf. Sci., 483:65–81, 2019. 1
  5. 5.Ming-Ming Cheng and Deng-Ping Fan. Structure-measure: A new way to evaluate foreground maps. IJCV, 129(9):2622–2638, 2021. 2, 5, 6
  6. 6.Yimian Dai, Fabian Gieseke, Stefan Oehmcke, Yiquan Wu, and Kobus Barnard. Attentional feature fusion. In WACV, pages 3560–3569, 2021. 2
  7. 7.Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. A tutorial on the cross-entropy method. Ann. Oper. Res., 134(1):19–67, 2005. 2
  8. 8.D.P. Fan, C. Gong, Y. Cao, B. Ren, M.M. Cheng, and A. Borji. Enhanced-alignment measure for binary foreground map evaluation. In IJCAI, page 698–704, 2018. 6
  9. 9.Deng-Ping Fan, Ming-Ming Cheng, Jiang-Jiang Liu, Shang-Hua Gao, Qibin Hou, and Ali Borji. Salient objects in clutter: Bringing salient object detection to the foreground. In ECCV, pages 186–202, 2018. 1
  10. 10.Deng-Ping Fan, Jing Zhang, Gang Xu, Ming-Ming Cheng, and Ling Shao. Salient objects in clutter. IEEE TPAMI, 45(2):2344–2366, 2022. 1
  11. 11.Mengyang Feng, Huchuan Lu, and Errui Ding. Attentive feedback network for boundary-aware salient object detection. In CVPR, pages 1623–1632, 2019. 1, 2, 3, 6
  12. 12.Lucas Fidon, Wenqi Li, Luis C Garcia-Peraza-Herrera, Jinendra Ekanayake, Neil Kitchen, Sébastien Ourselin, and Tom Vercauteren. Generalised wasserstein dice score for imbalanced multi-class segmentation using holistic convolutional networks. In International MICCAI brainlesion workshop, pages 64–76. Springer, 2017. 2
  13. 13.Ashish Kumar Gupta, Ayan Seal, Mukesh Prasad, and Pritee Khanna. Salient object detection techniques in computer vision—a survey. Entropy, 22(1174):1–49, 2020. 1, 3
  14. 14.Filiz Gurkan, Llukman Cerkezi, Ozgun Cirakman, and Bilge Gunsel. Tdiot: Target-driven inference for deep video object tracking. IEEE TIP, 30:7938–7951, 2021. 1
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 3, 6
  16. 16.Xiaowei Hu, Chi Wing Fu, Lei Zhu, Tianyu Wang, and Pheng Ann Heng. Sac-net: Spatial attenuation context for salient object detection. IEEE TCSVT, 31(3):1079–1090, 2021. 1, 2, 3, 6, 7
  17. 17.Qi Jia, Shuilian Yao, Yu Liu, Xin Fan, Risheng Liu, and Zhongxuan Luo. Segment, magnify and reiterate: Detecting camouflaged objects the hard way. In CVPR, pages 4713–4722, 2022. 1
  18. 18.Kurt Koffka. Principles of Gestalt psychology. Routledge, 2013. 2
  19. 19.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Proc. Adv. neural inf. proces. syst., volume 25, pages 1097–1105, Red Hook, NY, USA, 2012. 5
  20. 20.Guanbin Li and Yizhou Yu. Visual saliency based on multiscale deep features. In CVPR, pages 5455–5463, 2015. 5, 6, 8
  21. 21.Linghui Li, Sheng Tang, Yongdong Zhang, Lixi Deng, and Qi Tian. Gla: Global–local attention for image description. IEEE TMM, 20(3):726–737, 2017. 4
  22. 22.Xin Li, Fan Yang, Hong Cheng, Wei Liu, and Dinggang Shen. Contour knowledge transfer for salient object detection. In ECCV, pages 355–370, 2018. 1, 2
  23. 23.Yin Li, Xiaodi Hou, Christof Koch, James M. Rehg, and Alan L. Yuille. The secrets of salient object segmentation. In CVPR, pages 280–287, 2014. 5, 6, 8
  24. 24.J. J. Liu, Q. Hou, M. M. Cheng, J. Feng, and J. Jiang. A simple pooling-based design for real-time salient object detection. In CVPR, page 3917–3926, 2019. 1
  25. 25.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015. 1
  26. 26.Yuxin Mao, Jing Zhang, Zhexiong Wan, Yuchao Dai, Aixuan Li, Yunqiu Lv, Xinyu Tian, Deng-Ping Fan, and Nick Barnes. Generative transformer for accurate and reliable salient object detection. arXiv e-prints, pages arXiv–2104, 2021. 3
  27. 27.Ran Margolin, Lihi Zelnik Manor, and Ayellet Tal. How to evaluate foreground maps. In CVPR, pages 248–255, Columbus, OH, USA, 2014. 6
  28. 28.Vida Movahedi and James H. Elder. Design and perceptual validation of performance measures for salient object segmentation. In CVPRW, pages 49–56, 2010. 6, 8
  29. 29.Federico Perazzi, Philipp Krähenbóuhl, Yael Pritch, and Alexander Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, pages 733–740, 2012. 6
  30. 30.S Plainis and IJ Murray. Neurophysiological interpretation of human visual reaction times: effect of contrast, spatial frequency and luminance. Neuropsychologia, 38(12):1555–1564, 2000. 2
  31. 31.Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R. Zaiane, and Martin Jagersand. U2-net: Going deeper with nested u-structure for salient object detection. PR, 106:107404, 2020. 3, 6
  32. 32.Xuebin Qin, Zichen Zhang, Chenyang Huang, Chao Gao, Masood Dehghan, and Martin Jagersand. Basnet: Boundary-aware salient object detection. In CVPR, pages 7471–7481, 2019. 2, 3, 6
  33. 33.Shuang Qiu, Yao Zhao, Jianbo Jiao, Yunchao Wei, and Shikui Wei. Referring image segmentation by generative adversarial learning. IEEE TMM, 22(5):1333–1344, 2020. 1
  34. 34.Yu Qiu, Yun Liu, Yanan Chen, Jianwen Zhang, Jinchao Zhu, and Jing Xu. A2sppnet: Attentive atrous spatial pyramid pooling network for salient object detection. IEEE TMM, pages 1–1, 2022. 2, 4
  35. 35.Md Atiqur Rahman and Yang Wang. Optimizing intersection-over-union in deep neural networks for image segmentation. In International symposium on Vis. Comput., pages 234–244. Springer, 2016. 2, 5
  36. 36.Qinghua Ren, Shijian Lu, Jinxia Zhang, and Renjie Hu. Salient object detection by fusing local and global contexts. IEEE TMM, 23:1442–1453, 2021. 1, 3, 6
  37. 37.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. Unet: Convolutional networks for biomedical image segmentation. In MICCAI, page 234–241, 2015. 2, 3
  38. 38.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 6
  39. 39.Lijun Wang, Huchuan Lu, Yifan Wang, Mengyang Feng, Dong Wang, Baocai Yin, and Xiang Ruan. Learning to detect salient objects with image-level supervision. In CVPR, pages 3796–3805, 2017. 5, 6, 8
  40. 40.Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, and Ruigang Yang. Salient object detection in the deep learning era: An in-depth survey. IEEE TPAMI, 44(6):3239–3259, 2021. 1, 3
  41. 41.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 13(4):600–612, 2004. 2, 5
  42. 42.Jun Wei, Shuhui Wang, and Qingming Huang. F^3 net: fusion, feedback and focus for salient object detection. In AAAI, volume 34, pages 12321–12328, 2020. 2
  43. 43.Jun Wei, Shuhui Wang, Zhe Wu, Chi Su, Qingming Huang, and Qi Tian. Label decoupling framework for salient object detection. In CVPR, pages 13022–13031, 2020. 1, 2, 3, 6, 7
  44. 44.Runmin Wu, Mengyang Feng, Wenlong Guan, Dong Wang, Huchuan Lu, and Errui Ding. A mutual learning method for salient object detection with intertwined multi-supervision. In CVPR, pages 8142–8151, 2019. 3, 6
  45. 45.Yu-Huan Wu, Yun Liu, Le Zhang, Ming-Ming Cheng, and Bo Ren. Edn: Salient object detection via extremely-downsampled network. IEEE TIP, 31:3125–3136, 2022. 2, 3, 6, 7
  46. 46.Zhe Wu, Li Su, and Qingming Huang. Cascaded partial decoder for fast and accurate salient object detection. In CVPR, pages 3902–3911, 2019. 3, 6
  47. 47.Binwei Xu, Haoran Liang, Ronghua Liang, and Peng Chen. Locate globally, segment locally: A progressive architecture with knowledge review network for salient object detection. In AAAI, volume 35, pages 3004–3012, 2021. 1, 2, 3, 6, 7
  48. 48.Qiong Yan, Li Xu, Jianping Shi, and Jiaya Jia. Hierarchical saliency detection. In CVPR, pages 1155–1162, 2013. 6, 8
  49. 49.Chuan Yang, Lihe Zhang, Huchuan Lu, Xiang Ruan, and Ming-Hsuan Yang. Saliency detection via graph-based manifold ranking. In CVPR, pages 3166–3173, 2013. 5, 6, 8
  50. 50.Lihe Zhang, Jie Wu, Tiantian Wang, Ali Borji, Guohua Wei, and Huchuan Lu. A multistage refinement network for salient object detection. IEEE TIP, 29:3534–3545, 2020. 3, 6, 7
  51. 51.Yuan-fang Zhang, Jiangbin Zheng, Wenjing Jia, Wenfeng Huang, Long Li, Nian Liu, Fei Li, and Xiangjian He. Deep rgb-d saliency detection without depth. IEEE TMM, 24:755–767, 2021. 2
  52. 52.Zongjian Zhang, Qiang Wu, Yang Wang, and Fang Chen. Exploring pairwise relationships adaptively from linguistic context in image captioning. IEEE TMM, 24:3101–3113, 2022. 1
  53. 53.Jia Xing Zhao, Jiang Jiang Liu, Deng Ping Fan, Yang Cao, Jufeng Yang, and Ming Ming Cheng. Egnet: Edge guidance network for salient object detection. In ICCV, pages 8779–8788, 2019. 1, 2, 3, 6
  54. 54.Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu, and Lei Zhang. Suppress and balance: A simple gated network for salient object detection. In ECCV, pages 23–28, 2020. 3, 6
  55. 55.Li Zhaoping. Understanding vision: theory, models, and data. OUP Oxford, 2014. 1
  56. 56.Huajun Zhou, Xiaohua Xie, Jian-Huang Lai, Zixuan Chen, and Lingxiao Yang. Interactive two-stream decoder for accurate and fast saliency detection. In CVPR, pages 9141–9150, 2020. 2
  57. 57.Lei Zhu, Jiaxing Chen, Xiaowei Hu, Chi-Wing Fu, Xuemiao Xu, Jing Qin, and Pheng-Ann Heng. Aggregating attentional dilated features for salient object detection. IEEE TCSVT, 30(10):3358–3371, 2020. 1, 2, 3, 6, 7
  58. 58.Wencheng Zhu, Jiwen Lu, Jiahao Li, and Jie Zhou. Dsnet: A flexible detect-to-summarize network for video summarization. IEEE TIP, 30:948–962, 2021. 1
  59. 59.Mingchen Zhuge, Deng-Ping Fan, Nian Liu, Dingwen Zhang, Dong Xu, and Ling Shao. Salient object detection via integrity learning. IEEE TPAMI, 2022. 2, 3, 6, 7
  60. 60.Ming Zong, Ruili Wang, Xiubo Chen, Zhe Chen, and Yuanhao Gong. Motion saliency based multi-stream multiplier resnets for action recognition. Image Vis Comput., 107:104108, 2021. 1

Citation

MLA
Wang, Y., et al. “Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 10031–40, https://openaccess.thecvf.com/content/CVPR2023/html/Wang_Pixels_Regions_and_Objects_Multiple_Enhancement_for_Salient_Object_Detection_CVPR_2023_paper.html.
APA
Wang, Y., Wang, R., Fan, X., Wang, T., & He, X. (2023). Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10031–10040. https://openaccess.thecvf.com/content/CVPR2023/html/Wang_Pixels_Regions_and_Objects_Multiple_Enhancement_for_Salient_Object_Detection_CVPR_2023_paper.html
Chicago
Wang, Y., R. Wang, X. Fan, T. Wang, and X. He. 2023. “Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10031–40. https://openaccess.thecvf.com/content/CVPR2023/html/Wang_Pixels_Regions_and_Objects_Multiple_Enhancement_for_Salient_Object_Detection_CVPR_2023_paper.html.
Harvard
Wang, Y. et al. (2023) “Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10031–10040. Available at: https://openaccess.thecvf.com/content/CVPR2023/html/Wang_Pixels_Regions_and_Objects_Multiple_Enhancement_for_Salient_Object_Detection_CVPR_2023_paper.html.
Vancouver
1. Wang Y, Wang R, Fan X, Wang T, He X (2023) Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp 10031–10040

BibTeX

@InProceedings{Wang_2023_CVPR,
    author    = {Wang, Yi and Wang, Ruili and Fan, Xin and Wang, Tianzhu and He, Xiangjian},
    title     = {Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2023},
    pages     = {10031-10040}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE