BASNet: Boundary-Aware Salient Object Detection

Xuebin QinZichen ZhangChenyang HuangChao GaoMasood DehghanMartin Jägersand

article2019CVPR1,464 citations

Proposes a predict-refine neural network coupled with a multi-level hybrid loss that integrates BCE, SSIM, and IoU to generate sharp, boundary-accurate salient object segmentations at real-time speeds.

Listen

Automated salient object detection identifies and segments the most visually prominent objects in an image, serving as a critical foundation for downstream computer vision tasks such as visual tracking, image manipulation, and user interface optimization. While deep learning methods have significantly advanced regional segmentation accuracy, existing approaches struggle to delineate sharp, well-defined boundaries and capture fine structural details. Conventional training loss functions fail to provide high confidence along object edges, frequently leaving output boundaries blurry or degraded.

The article introduces and evaluates BASNet, a boundary-aware salient object detection framework designed to produce highly accurate region segmentations with crisp, well-defined edges. The architecture adopts a two-stage predict-and-refine structure: a primary encoder-decoder network predicts coarse saliency maps, and a residual refinement module enhances the output in a single pass. To train the system, the article introduces a hybrid loss combining three complementary objectives across hierarchical levels: pixel-level cross-entropy for steady gradient flow, patch-level structural similarity for edge awareness, and map-level intersection-over-union for overall foreground quality. The approach was trained on 10,553 images and benchmarked across six public datasets comprising thousands of challenging visual scenes.

The findings confirm that the proposed method outperforms 15 leading baseline methods across both regional and boundary metrics. Specifically, BASNet achieved substantial gains in boundary accuracy, improving the relaxed boundary measure by approximately 3.4% to 6.2% across all evaluated benchmark datasets. The framework maintained a high overall accuracy, producing the lowest average absolute error on five out of six benchmark tests. In addition to high accuracy, the system achieved a processing speed of over 25 frames per second on standard hardware, demonstrating real-time operational efficiency without requiring heavy post-processing steps like conditional random fields.

These results establish that incorporating patch-level structural information alongside region losses eliminates the trade-off between boundary clarity and computational efficiency. For technology leaders and visual application developers, the framework offers a practical path toward higher-quality visual segmentation without incurring additional post-processing delays or heavy computational costs. Practitioners seeking to deploy automated image editing, object tracking, or interface optimization tools can leverage this modular framework to upgrade existing pipelines. Future efforts should focus on extending this predict-refine modular framework and multi-level loss formulation to adjacent computer vision tasks, such as road extraction or medical image segmentation.

Confidence in these findings is supported by consistent quantitative improvements across six diverse benchmark datasets and extensive ablation testing. However, the evaluation relies on resizing images to a standardized resolution during testing, meaning fine details in ultra-high-resolution imagery may require further scaling validation before deployment in specialized industrial domains.

Cover for BASNet: Boundary-Aware Salient Object Detection

Abstract

Deep Convolutional Neural Networks have been adopted for salient object detection and achieved the state-of-the-art performance. Most of the previous works however focus on region accuracy but not on the boundary quality. In this paper, we propose a predict-refine architecture, BASNet, and a new hybrid loss for Boundary-Aware Salient object detection. Specifically, the architecture is composed of a densely supervised Encoder-Decoder network and a residual refinement module, which are respectively in charge of saliency prediction and saliency map refinement. The hybrid loss guides the network to learn the transformation between the input image and the ground truth in a three-level hierarchy - pixel-, patch- and map- level - by fusing Binary Cross Entropy (BCE), Structural SIMilarity (SSIM) and Intersection-over-Union (IoU) losses. Equipped with the hybrid loss, the proposed predict-refine architecture is able to effectively segment the salient object regions and accurately predict the fine structures with clear boundaries. Experimental results on six public datasets show that our method outperforms the state-of-the-art methods both in terms of regional and boundary evaluation measures. Our method runs at over 25 fps on a single GPU. The code is available at: https://github.com/NathanUA/BASNet.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. BASNet
  • 3.1. Overview of Network Architecture
  • 3.2. Predict Module
  • 3.3. Refine Module
  • 3.4. Hybrid Loss
  • 4. Experimental Results
  • 4.1. Datasets
  • 4.2. Implementation and Experimental Setup
  • 4.3. Evaluation Metrics
  • 4.4. Ablation Study
  • 4.5. Comparison with State-of-the-arts
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — BASNet Predict-Refine Architecture Overview

    model/method

    BASNet is an end-to-end convolutional neural network for boundary-aware salient object detection designed using a two-stage predict-refine framework. The network comprises two primary sub-networks:

    1. Predict Module (Encoder-Decoder): A deeply supervised, U-Net-like encoder-decoder network that receives an input RGB image and produces a primary coarse saliency map ScoarseS_{coarse}, alongside multi-scale intermediate side-output saliency predictions used for deep supervision.

    2. Residual Refinement Module (RRM): A lightweight residual encoder-decoder network operating at the input image scale. It takes the coarse saliency map ScoarseS_{coarse} from the predict module and learns a residual map SresidualS_{residual} to correct blurry or inaccurate boundary transitions and regional prediction uncertainties, producing the final refined saliency map:

    Srefined=Scoarse+SresidualS_{refined} = S_{coarse} + S_{residual}

    Deep supervision is applied to all intermediate side outputs of the predict module as well as the output of the refinement module, training the complete system end-to-end.

  2. Knowl 2 — Predict Module Architecture in BASNet

    model/method

    The Predict Module of BASNet is an encoder-decoder network designed to capture high-level global context and low-level spatial details simultaneously:

    • Encoder: Consists of an input convolutional layer and six stages of basic residual blocks adapted from ResNet-34. The initial convolutional layer has 64 filters of size 3×33 \times 3 with stride 1 and no subsequent max-pooling (in contrast to standard ResNet-34 7×77 \times 7 filters with stride 2 and max-pooling), allowing the initial feature representations to maintain full input resolution. To maintain an equivalent receptive field to standard ResNet-34 despite this higher initial resolution, two extra stages are appended after the fourth stage of ResNet-34; both added stages consist of three basic residual blocks with 512 channels, each preceded by a non-overlapping 2×22 \times 2 max-pooling operation.
    • Bridge Stage: Placed between the encoder and decoder, consisting of three convolutional layers, each with 512 filters of size 3×33 \times 3 and dilation rate 2, Batch Normalization, and ReLU activations.
    • Decoder: Symmetrical to the encoder with six stages. Each decoder stage consists of three convolutional layers followed by Batch Normalization and ReLU. The input to each decoder stage is the concatenation of the bilinearly upsampled feature map from the preceding decoder stage and the corresponding encoder stage's feature map via skip connection.
    • Side Outputs: The multi-channel output of the bridge stage and each of the six decoder stages is fed to a 3×33 \times 3 convolutional layer, bilinearly upsampled to the input spatial resolution, and mapped through a sigmoid activation to yield 7 side saliency maps during training. The final decoder stage output is taken as ScoarseS_{coarse} and fed into the refinement module.
  3. Knowl 3 — Residual Refinement Module Architecture in BASNet

    model/method

    The Residual Refinement Module (RRM) in BASNet refines the coarse saliency prediction ScoarseS_{coarse} at the original single scale by learning the residual error map SresidualS_{residual} relative to the ground truth. It is structured as a compact encoder-decoder network:

    • Encoder: Composed of four stages. Each stage consists of a single convolutional layer with 64 filters of size 3×33 \times 3, Batch Normalization, and ReLU activation, followed by a non-overlapping 2×22 \times 2 max-pooling layer for downsampling.
    • Bridge Stage: A single convolutional layer with 64 filters of size 3×33 \times 3, Batch Normalization, and ReLU activation.
    • Decoder: Composed of four stages symmetrical to the encoder. Each stage consists of a single convolutional layer with 64 filters of size 3×33 \times 3, Batch Normalization, and ReLU activation. Features from the previous decoder stage are upsampled via bilinear interpolation and concatenated with the corresponding encoder stage features before passing through the stage's convolutional layer.
    • Output Layer: A 3×33 \times 3 convolution with 1 output channel followed by a sigmoid activation function to generate the residual map SresidualS_{residual}. The final output is formed by summing ScoarseS_{coarse} and SresidualS_{residual}.
  4. Knowl 4 — Multi-Level Hybrid Loss Function

    equation

    BASNet trains all supervised outputs using a multi-level hybrid loss that combines pixel-level, patch-level, and map-level loss terms. The total training loss is defined over K=8K = 8 outputs (7 side outputs from the predict module and 1 from the refinement module):

    L=∑k=1Kαkℓ(k)\mathcal{L} = \sum_{k=1}^{K} \alpha_k \ell^{(k)}

    where αk\alpha_k is the loss weight for the kk-th output, and the hybrid loss ℓ(k)\ell^{(k)} is defined as:

    ℓ(k)=ℓbce(k)+ℓssim(k)+ℓiou(k)\ell^{(k)} = \ell_{bce}^{(k)} + \ell_{ssim}^{(k)} + \ell_{iou}^{(k)}

    The individual loss components are:

    1. Binary Cross Entropy Loss (Pixel-level): ℓbce=−∑(r,c)[G(r,c)log⁡(S(r,c))+(1−G(r,c))log⁡(1−S(r,c))]\ell_{bce} = -\sum_{(r,c)} \left[ G(r,c) \log(S(r,c)) + (1 - G(r,c)) \log(1 - S(r,c)) \right] where G(r,c)∈{0,1}G(r,c) \in \{0, 1\} is the ground-truth binary label at pixel coordinate (r,c)(r,c) and S(r,c)∈[0,1]S(r,c) \in [0, 1] is the predicted saliency probability.

    2. Structural Similarity Loss (Patch-level): For corresponding N×NN \times N patches xx and yy cropped from the predicted probability map SS and ground-truth mask GG: ℓssim=1−(2μxμy+C1)(2σxy+C2)(μx2+μy2+C1)(σx2+σy2+C2)\ell_{ssim} = 1 - \frac{(2\mu_x \mu_y + C_1)(2\sigma_{xy} + C_2)}{(\mu_x^2 + \mu_y^2 + C_1)(\sigma_x^2 + \sigma_y^2 + C_2)} where μx,μy\mu_x, \mu_y and σx,σy\sigma_x, \sigma_y are the means and standard deviations of xx and yy, σxy\sigma_{xy} is their covariance, and C1=0.012,C2=0.032C_1 = 0.01^2, C_2 = 0.03^2 prevent division by zero.

    3. Intersection-over-Union Loss (Map-level): ℓiou=1−∑r=1H∑c=1WS(r,c)G(r,c)∑r=1H∑c=1W[S(r,c)+G(r,c)−S(r,c)G(r,c)]\ell_{iou} = 1 - \frac{\sum_{r=1}^{H} \sum_{c=1}^{W} S(r,c) G(r,c)}{\sum_{r=1}^{H} \sum_{c=1}^{W} \left[ S(r,c) + G(r,c) - S(r,c)G(r,c) \right]} where HH and WW denote the height and width of the map.

  5. Knowl 5 — Relaxed Boundary F-Measure Evaluation Metric

    definition

    The relaxed boundary F-measure (relaxFβb\text{relax}F_\beta^b) evaluates the geometric boundary quality of a predicted saliency map against a ground-truth binary mask:

    1. The continuous saliency map SS is binarized at a threshold of 0.5 to produce a binary mask SbwS_{bw}.
    2. A 1-pixel wide boundary mask is extracted by computing the XOR difference between SbwS_{bw} and its morphological erosion SerdS_{erd} using standard erosion: Boundary(Sbw)=XOR(Sbw,Serd)\text{Boundary}(S_{bw}) = \text{XOR}(S_{bw}, S_{erd}) The same procedure extracts the boundary of the ground-truth mask GG.
    3. Relaxed Boundary Precision (relaxPrecisionb\text{relaxPrecision}^b): The fraction of predicted boundary pixels that lie within a Euclidean distance of ρ\rho pixels from any ground-truth boundary pixel.
    4. Relaxed Boundary Recall (relaxRecallb\text{relaxRecall}^b): The fraction of ground-truth boundary pixels that lie within a Euclidean distance of ρ\rho pixels from any predicted boundary pixel.
    5. Using the slack parameter ρ=3\rho = 3 and β2=0.3\beta^2 = 0.3 (weighting precision more than recall), the relaxed boundary F-measure is calculated as:

    relaxFβb=(1+β2)×relaxPrecisionb×relaxRecallbβ2×relaxPrecisionb+relaxRecallb\text{relax}F_\beta^b = \frac{(1 + \beta^2) \times \text{relaxPrecision}^b \times \text{relaxRecall}^b}{\beta^2 \times \text{relaxPrecision}^b + \text{relaxRecall}^b}

  6. Knowl 6 — Training Configuration and Implementation Parameters for BASNet

    experimental setup

    BASNet is trained with the following configuration:

    • Training Dataset: DUTS-TR containing 10,553 images, augmented via horizontal flipping to 21,106 images.
    • Data Preprocessing: Each image is resized to 256×256256 \times 256 pixels and randomly cropped to 224×224224 \times 224 pixels during training.
    • Network Initialization: The corresponding encoder weights are initialized from ImageNet-pretrained ResNet-34; remaining convolutional layers are initialized via Xavier uniform initialization.
    • Optimization: Trained using the Adam optimizer with initial learning rate lr=10−3\text{lr} = 10^{-3}, β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, ϵ=10−8\epsilon = 10^{-8}, and weight decay =0= 0. Batch size is set to 8.
    • Convergence: Trained for 400k iterations (approx. 125 hours on an NVIDIA GTX 1080ti GPU with 11GB RAM and an AMD Ryzen 1800x CPU with 32GB RAM).
    • Inference: Test images are resized to 256×256256 \times 256, processed in a single forward pass, and the resulting 256×256256 \times 256 output map is bilinearly upsampled to the original image dimensions. Inference speed is 0.040 seconds per image (>25 fps) on a single GTX 1080ti GPU without CRF post-processing.
  7. Knowl 7 — Ablation Study of BASNet Architecture and Loss Components

    data/table

    The table below presents the ablation study on the ECSSD dataset, comparing different network architectures and loss formulations against the baseline U-Net. Evaluated metrics are maximum F-measure (max⁡Fβ\max F_\beta), relaxed boundary F-measure (relaxFβb\text{relax}F_\beta^b), and Mean Absolute Error (MAE).

    Configurations max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE
    Architecture
    Baseline U-Net + ℓbce\ell_{bce} 0.896 0.669 0.066
    En-De + ℓbce\ell_{bce} 0.929 0.767 0.047
    En-De+Sup + ℓbce\ell_{bce} 0.934 0.805 0.040
    En-De+Sup+RRM_LC + ℓbce\ell_{bce} 0.936 0.803 0.040
    En-De+Sup+RRM_MS + ℓbce\ell_{bce} 0.935 0.804 0.042
    En-De+Sup+RRM_Ours + ℓbce\ell_{bce} 0.937 0.806 0.042
    Loss
    En-De+Sup+RRM_Ours + ℓssim\ell_{ssim} 0.924 0.808 0.042
    En-De+Sup+RRM_Ours + ℓiou\ell_{iou} 0.933 0.795 0.039
    En-De+Sup+RRM_Ours + ℓbs\ell_{bs} 0.940 0.815 0.040
    En-De+Sup+RRM_Ours + ℓbi\ell_{bi} 0.940 0.813 0.038
    En-De+Sup+RRM_Ours + ℓbsi\ell_{bsi} 0.942 0.826 0.037

    Notation:

    • En-De: Modified Encoder-Decoder predict module.
    • Sup: Densely supervised side outputs.
    • RRM_LC: Local context residual refinement module.
    • RRM_MS: Multi-scale residual refinement module.
    • RRM_Ours: Proposed encoder-decoder residual refinement module.
    • Loss terms: ℓbi=ℓbce+ℓiou\ell_{bi} = \ell_{bce} + \ell_{iou}, ℓbs=ℓbce+ℓssim\ell_{bs} = \ell_{bce} + \ell_{ssim}, ℓbsi=ℓbce+ℓssim+ℓiou\ell_{bsi} = \ell_{bce} + \ell_{ssim} + \ell_{iou}.

    The full model (En-De+Sup+RRM_Ours trained with ℓbsi\ell_{bsi}) demonstrates superior performance across all metrics, with relaxFβb\text{relax}F_\beta^b improving from 0.669 (baseline) to 0.826 and MAE dropping from 0.066 to 0.037.

  8. Knowl 8 — Benchmark Comparison of BASNet Against State-of-the-Art Methods

    data/table

    Performance comparison of BASNet against 15 state-of-the-art salient object detection methods across six standard benchmark datasets (SOD, ECSSD, DUT-OMRON, PASCAL-S, HKU-IS, and DUTS-TE) using maximum F-measure (max⁡Fβ\max F_\beta, higher is better), relaxed boundary F-measure (relaxFβb\text{relax}F_\beta^b, higher is better), and MAE (lower is better):

    Method SOD ECSSD DUT-OMRON PASCAL-S HKU-IS DUTS-TE
    max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE max⁡Fβ\max F_\beta relaxFβb\text{relax}F_\beta^b MAE
    Ours 0.851 0.603 0.114 0.942 0.826 0.037 0.805 0.694 0.056 0.854 0.660 0.076 0.928 0.807 0.032 0.860 0.758 0.047
    PiCANetR 0.856 0.528 0.104 0.935 0.775 0.046 0.803 0.632 0.065 0.857 0.598 0.076 0.918 0.765 0.043 0.860 0.696 0.050
    BMPM 0.856 0.562 0.108 0.928 0.770 0.045 0.774 0.612 0.064 0.850 0.617 0.074 0.921 0.773 0.039 0.852 0.699 0.048
    R^3Net+ 0.850 0.431 0.125 0.934 0.759 0.040 0.795 0.599 0.063 0.834 0.538 0.092 0.915 0.740 0.036 0.828 0.601 0.058
    PAGRN - - - 0.927 0.747 0.061 0.771 0.582 0.071 0.847 0.594 0.0895 0.918 0.762 0.048 0.854 0.692 0.055
    RADF+ 0.838 0.476 0.126 0.923 0.720 0.049 0.791 0.579 0.061 0.830 0.515 0.097 0.914 0.725 0.039 0.821 0.608 0.061
    DGRL 0.848 0.502 0.106 0.925 0.753 0.042 0.779 0.584 0.063 0.848 0.569 0.074 0.913 0.744 0.037 0.834 0.656 0.051
    RAS 0.851 0.544 0.124 0.921 0.741 0.056 0.786 0.615 0.062 0.829 0.560 0.101 0.913 0.748 0.045 0.831 0.656 0.059
    C2S 0.823 0.457 0.124 0.910 0.708 0.055 0.758 0.565 0.072 0.840 0.543 0.082 0.896 0.717 0.048 0.807 0.607 0.062
    LFR 0.828 0.479 0.123 0.911 0.694 0.052 0.740 0.508 0.103 0.801 0.499 0.107 0.911 0.731 0.040 0.778 0.556 0.083
    DSS+ 0.846 0.444 0.124 0.921 0.696 0.052 0.781 0.559 0.063 0.831 0.499 0.093 0.916 0.706 0.040 0.825 0.606 0.056
    NLDF+ 0.841 0.475 0.125 0.905 0.666 0.063 0.753 0.514 0.080 0.822 0.495 0.098 0.902 0.694 0.048 0.813 0.591 0.065
    SRM 0.843 0.392 0.128 0.917 0.672 0.054 0.769 0.523 0.069 0.838 0.509 0.084 0.906 0.680 0.046 0.826 0.592 0.058
    Amulet 0.798 0.454 0.144 0.915 0.711 0.059 0.743 0.528 0.098 0.828 0.541 0.100 0.897 0.716 0.051 0.778 0.568 0.084
    UCF 0.808 0.471 0.148 0.903 0.669 0.069 0.730 0.480 0.120 0.814 0.493 0.115 0.888 0.679 0.062 0.773 0.518 0.112
    MDF 0.746 0.311 0.192 0.832 0.472 0.105 0.694 0.406 0.092 0.759 0.343 0.142 0.860 0.594 0.129 0.729 0.447 0.099

    BASNet achieves state-of-the-art results across datasets, delivering gains in boundary quality with relaxFβb\text{relax}F_\beta^b improvements of 4.1% on SOD, 5.1% on ECSSD, 6.2% on DUT-OMRON, 6.2% on PASCAL-S, 3.4% on HKU-IS, and 5.9% on DUTS-TE over previous best-performing methods without requiring any CRF post-processing.

Coverage note — None was omitted; all key contributions including network architecture components, hybrid loss formulation, evaluation metrics, experimental settings, ablation studies, and state-of-the-art benchmark comparisons have been captured.

References

  1. 1.Radhakrishna Achanta, Sheila Hemami, Francisco Estrada, and Sabine Susstrunk. Frequency-tuned salient region detection. In Computer vision and pattern recognition, 2009. cvpr 2009. ieee conference on, pages 1597–1604. IEEE, 2009.
  2. 2.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis & Machine Intelligence, (12):2481–2495, 2017.
  3. 3.Ali Borji, Ming-Ming Cheng, Huaizu Jiang, and Jia Li. Salient object detection: A benchmark. IEEE Trans. Image Processing, 24(12):5706–5722, 2015.
  4. 4.Shuhan Chen, Xiuli Tan, Ben Wang, and Xuelong Hu. Reverse attention for salient object detection. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part IX, pages 236–252, 2018.
  5. 5.Pieter-Tjerk de Boer, Dirk P. Kroese, Shie Mannor, and Reuven Y. Rubinstein. A tutorial on the cross-entropy method. Annals OR, 134(1):19–67, 2005.
  6. 6.Zijun Deng, Xiaowei Hu, Lei Zhu, Xuemiao Xu, Jing Qin, Guoqiang Han, and Pheng-Ann Heng. R3net: Recurrent residual refinement network for saliency detection. IJCAI, 2018.
  7. 7.Marc Ehrig and Jer´ ome Euzenat. Relaxed precision and recall for ontology matching. In Proc. K-Cap 2005 workshop on Integrating ontology, pages 25–32. No commercial editor., 2005.
  8. 8.Lucas Fidon, Wenqi Li, Luis C. Herrera, Jinendra Ekanayake, Neil Kitchen, Sebastien Ourselin, and Tom Vercauteren. Generalised wasserstein dice score for imbalanced multi-class segmentation using holistic convolutional networks. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries - Third International Workshop, BrainLes 2017, Held in Conjunction with MICCAI 2017, Quebec City, QC, Canada, pages 64–76, 2017.
  9. 9.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
  10. 10.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2010, Chia Laguna Resort, Sardinia, Italy, May 13-15, 2010, pages 249–256, 2010.
  11. 11.Stas Goferman, Lihi Zelnik-Manor, and Ayellet Tal. Context-aware saliency detection. IEEE transactions on pattern analysis and machine intelligence, 34(10):1915–1926, 2012.
  12. 12.Prakhar Gupta, Shubh Gupta, Ajaykrishnan Jayagopal, Sourav Pal, and Ritwik Sinha. Saliency prediction for mobile user interfaces. In 2018 IEEE Winter Conference on Applications of Computer Vision, WACV 2018, Lake Tahoe, NV, USA, March 12-15, 2018, pages 1529–1538, 2018.
  13. 13.Richard HR Hahnloser and H Sebastian Seung. Permitted and forbidden sets in symmetric threshold-linear networks. In Advances in Neural Information Processing Systems, pages 217–223, 2001.
  14. 14.Robert M Haralick, Stanley R Sternberg, and Xinhua Zhuang. Image analysis using mathematical morphology. IEEE transactions on pattern analysis and machine intelligence, (4):532–550, 1987.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In European conference on computer vision, pages 346–361. Springer, 2014.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  17. 17.Qibin Hou, Ming-Ming Cheng, Xiaowei Hu, Ali Borji, Zhuowen Tu, and Philip Torr. Deeply supervised salient object detection with short connections. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5300–5309. IEEE, 2017.
  18. 18.Ping Hu, Bing Shuai, Jun Liu, and Gang Wang. Deep level sets for salient object detection. In CVPR, volume 1, page 2, 2017.
  19. 19.Xiaowei Hu, Lei Zhu, Jing Qin, Chi-Wing Fu, and Pheng-Ann Heng. Recurrently aggregating deep features for salient object detection. In Proceedings of AAAI-18, New Orleans, Louisiana, USA, pages 6943–6950, 2018.
  20. 20.Xun Huang, Chengyao Shen, Xavier Boix, and Qi Zhao. Salicon: Reducing the semantic gap in saliency prediction by adapting deep neural networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 262–270, 2015.
  21. 21.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  22. 22.Md Amirul Islam, Mahmoud Kalash, Mrigank Rochan, Neil DB Bruce, and Yang Wang. Salient object detection using a context-aware refinement network.
  23. 23.Paul Jaccard. The distribution of the flora in the alpine zone. 1. New phytologist, 11(2):37–50, 1912.
  24. 24.Martin J¨agersand. Saliency maps and attention selection in scale and spatial coordinates: An information theoretic approach. In ICCV, pages 195–202, 1995.
  25. 25.Timor Kadir and Michael Brady. Saliency, scale and image description. International Journal of Computer Vision, 45(2):83–105, 2001.
  26. 26.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.Philipp Kr¨ahenb¨uhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In Advances in neural information processing systems, pages 109–117, 2011.
  28. 28.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States., pages 1106–1114, 2012.
  29. 29.Srinivas SS Kruthiventi, Vennela Gudisa, Jaley H Dholakiya, and R Venkatesh Babu. Saliency unified: A deep architecture for simultaneous eye fixation prediction and salient object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5781–5790, 2016.
  30. 30.Jason Kuen, Zhenhua Wang, and Gang Wang. Recurrent attentional networks for saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3668–3677, 2016.
  31. 31.Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu. Deeply-supervised nets. In Artificial Intelligence and Statistics, pages 562–570, 2015.
  32. 32.Hyemin Lee and Daijin Kim. Salient region-based online object tracking. In 2018 IEEE Winter Conference on Applications of Computer Vision, WACV 2018, Lake Tahoe, NV, USA, March 12-15, 2018, pages 1170–1177, 2018.
  33. 33.Guanbin Li and Yizhou Yu. Visual saliency based on multiscale deep features. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5455–5463, 2015.
  34. 34.Guanbin Li and Yizhou Yu. Deep contrast learning for salient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 478–487, 2016.
  35. 35.Guanbin Li and Yizhou Yu. Visual saliency detection based on multiscale deep cnn features. IEEE Transactions on Image Processing, 25(11):5012–5024, 2016.
  36. 36.Xin Li, Fan Yang, Hong Cheng, Wei Liu, and Dinggang Shen. Contour knowledge transfer for salient object detection. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part XV, pages 370–385, 2018.
  37. 37.Yin Li, Xiaodi Hou, Christof Koch, James M Rehg, and Alan L Yuille. The secrets of salient object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 280–287, 2014.
  38. 38.Nian Liu and Junwei Han. Dhsnet: Deep hierarchical saliency network for salient object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 678–686, 2016.
  39. 39.Nian Liu, Junwei Han, and Ming-Hsuan Yang. Picanet: Learning pixel-wise contextual attention for saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3089–3098, 2018.
  40. 40.Nian Liu, Junwei Han, Dingwen Zhang, Shifeng Wen, and Tianming Liu. Predicting eye fixations using convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 362–370, 2015.
  41. 41.Zhiming Luo, Akshaya Mishra, Andrew Achkar, Justin Eichel, Shaozi Li, and Pierre-Marc Jodoin. Non-local deep features for salient object detection. In Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, pages 6593–6601. IEEE, 2017.
  42. 42.Gell´ert M´attyus, Wenjie Luo, and Raquel Urtasun. Deeproadmapper: Extracting road topology from aerial images.
  43. 43.Roey Mechrez, Eli Shechtman, and Lihi Zelnik-Manor. Saliency driven image manipulation. In 2018 IEEE Winter Conference on Applications of Computer Vision, WACV 2018, Lake Tahoe, NV, USA, March 12-15, 2018, pages 1368–1376, 2018.
  44. 44.Volodymyr Mnih and Geoffrey E Hinton. Learning to detect roads in high-resolution aerial images. In European Conference on Computer Vision, pages 210–223. Springer, 2010.
  45. 45.Vida Movahedi and James H Elder. Design and perceptual validation of performance measures for salient object segmentation. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on, pages 49–56. IEEE, 2010.
  46. 46.David Mumford and Jayant Shah. Optimal approximations by piecewise smooth functions and associated variational problems. Communications on pure and applied mathematics, 42(5):577–685, 1989.
  47. 47.Gattigorla Nagendar, Digvijay Singh, Vineeth N. Balasubramanian, and C. V. Jawahar. Neuro-iou: Learning a surrogate loss for semantic segmentation. In British Machine Vision Conference 2018, BMVC 2018, Northumbria University, Newcastle, UK, September 3-6, 2018, page 278, 2018.
  48. 48.Stanley Osher and James A Sethian. Fronts propagating with curvature-dependent speed: algorithms based on hamiltonjacobi formulations. Journal of computational physics, 79(1):12–49, 1988.
  49. 49.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS-W, 2017.
  50. 50.Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun. Large kernel mattersimprove semantic segmentation by global convolutional network. In Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, pages 1743–1751. IEEE, 2017.
  51. 51.Federico Perazzi, Philipp Kr¨ahenb¨uhl, Yael Pritch, and Alexander Hornung. Saliency filters: Contrast based filtering for salient region detection. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 733–740. IEEE, 2012.
  52. 52.Xuebin Qin, Shida He, Camilo Perez Quintero, Abhineet Singh, Masood Dehghan, and Martin J¨agersand. Real-time salient closed boundary tracking via line segments perceptual grouping. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2017, Vancouver, BC, Canada, September 24-28, 2017, pages 4284–4289, 2017.
  53. 53.Xuebin Qin, Shida He, Xiucheng Yang, Masood Dehghan, Qiming Qin, and Martin Jagersand. Accurate outline extraction of individual building from very high-resolution optical images. IEEE Geoscience and Remote Sensing Letters, (99):1–5, 2018.
  54. 54.Xuebin Qin, Shida He, Zichen Zhang, Masood Dehghan, and Martin Jagersand. Bylabel: A boundary based semiautomatic image annotation tool. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1804–1813. IEEE, 2018.
  55. 55.Xuebin Qin, Shida He, Zichen Vincent Zhang, Masood Dehghan, and Martin J¨agersand. Real-time salient closed boundary tracking using perceptual grouping and shape priors. In British Machine Vision Conference 2017, BMVC 2017, London, UK, September 4-7, 2017, 2017.
  56. 56.Md Atiqur Rahman and Yang Wang. Optimizing intersection-over-union in deep neural networks for image segmentation. In International Symposium on Visual Computing, pages 234–244. Springer, 2016.
  57. 57.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. Unet: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  58. 58.Shunta Saito, Takayoshi Yamashita, and Yoshimitsu Aoki. Multiple object extraction from aerial imagery with convolutional neural networks. Electronic Imaging, 2016(10):1–9, 2016.
  59. 59.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  60. 60.R Sai Srivatsa and R Venkatesh Babu. Salient object detection via objectness measure. In Image Processing (ICIP), 2015 IEEE International Conference on, pages 4481–4485. IEEE, 2015.
  61. 61.Lijun Wang, Huchuan Lu, Xiang Ruan, and Ming-Hsuan Yang. Deep networks for saliency detection via local estimation and global search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3183–3192, 2015.
  62. 62.Lijun Wang, Huchuan Lu, Yifan Wang, Mengyang Feng, Dong Wang, Baocai Yin, and Xiang Ruan. Learning to detect salient objects with image-level supervision. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(CVPR), pages 136–145, 2017.
  63. 63.Linzhao Wang, Lijun Wang, Huchuan Lu, Pingping Zhang, and Xiang Ruan. Salient object detection with recurrent fully convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  64. 64.Tiantian Wang, Ali Borji, Lihe Zhang, Pingping Zhang, and Huchuan Lu. A stagewise refinement model for detecting salient objects in images. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 4039–4048, 2017.
  65. 65.Tiantian Wang, Lihe Zhang, Shuo Wang, Huchuan Lu, Gang Yang, Xiang Ruan, and Ali Borji. Detect globally, refine locally: A novel approach to saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3127–3135, 2018.
  66. 66.Zhou Wang, Eero P Simoncelli, and Alan C Bovik. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, volume 2, pages 1398–1402. Ieee, 2003.
  67. 67.Saining Xie and Zhuowen Tu. Holistically-nested edge detection. In Proceedings of the IEEE international conference on computer vision, pages 1395–1403, 2015.
  68. 68.Qiong Yan, Li Xu, Jianping Shi, and Jiaya Jia. Hierarchical saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1155–1162, 2013.
  69. 69.Chuan Yang, Lihe Zhang, Huchuan Lu, Xiang Ruan, and Ming-Hsuan Yang. Saliency detection via graph-based manifold ranking. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3166–3173, 2013.
  70. 70.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  71. 71.Jianming Zhang, Stan Sclaroff, Zhe Lin, Xiaohui Shen, Brian Price, and Radomir Mech. Minimum barrier salient object detection at 80 fps. In Proceedings of the IEEE international conference on computer vision, pages 1404–1412, 2015.
  72. 72.Lu Zhang, Ju Dai, Huchuan Lu, You He, and Gang Wang. A bi-directional message passing model for salient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1741–1750, 2018.
  73. 73.Pingping Zhang, Wei Liu, Huchuan Lu, and Chunhua Shen. Salient object detection by lossless feature reflection. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden., pages 1149–1155, 2018.
  74. 74.Pingping Zhang, Dong Wang, Huchuan Lu, Hongyu Wang, and Xiang Ruan. Amulet: Aggregating multi-level convolutional features for salient object detection. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 202–211, 2017.
  75. 75.Pingping Zhang, Dong Wang, Huchuan Lu, Hongyu Wang, and Baocai Yin. Learning uncertain convolutional features for accurate saliency detection. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 212–221, 2017.
  76. 76.Xiaoning Zhang, Tiantian Wang, Jinqing Qi, Huchuan Lu, and Gang Wang. Progressive attention guided recurrent network for salient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 714–722, 2018.
  77. 77.Zhengxin Zhang, Qingjie Liu, and Yunhong Wang. Road extraction by deep residual u-net. IEEE Geoscience and Remote Sensing Letters, 2018.
  78. 78.Kai Zhao, Shanghua Gao, Qibin Hou, Dandan Li, and MingMing Cheng. Optimizing the f-measure for threshold-free salient object detection. CoRR, abs/1805.07567, 2018.
  79. 79.Rui Zhao, Wanli Ouyang, Hongsheng Li, and Xiaogang Wang. Saliency detection by multi-context deep learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1265–1274, 2015.
  80. 80.Wangjiang Zhu, Shuang Liang, Yichen Wei, and Jian Sun. Saliency optimization from robust background detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2814–2821, 2014.

Citation

MLA
Qin, X., et al. “BASNet: Boundary-Aware Salient Object Detection”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 7471–81, https://doi.org/10.1109/CVPR.2019.00766.
APA
Qin, X., Zhang, Z., Huang, C., Gao, C., Dehghan, M., & Jagersand, M. (2019). BASNet: Boundary-Aware Salient Object Detection. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7471–7481. https://doi.org/10.1109/CVPR.2019.00766
Chicago
Qin, X., Z. Zhang, C. Huang, C. Gao, M. Dehghan, and M. Jagersand. 2019. “BASNet: Boundary-Aware Salient Object Detection”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7471–81. https://doi.org/10.1109/CVPR.2019.00766.
Harvard
Qin, X. et al. (2019) “BASNet: Boundary-Aware Salient Object Detection”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7471–7481. Available at: https://doi.org/10.1109/CVPR.2019.00766.
Vancouver
1. Qin X, Zhang Z, Huang C, Gao C, Dehghan M, Jagersand M (2019) BASNet: Boundary-Aware Salient Object Detection. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7471–7481

BibTeX

@inproceedings{Qin_2019, title={BASNet: Boundary-Aware Salient Object Detection}, url={http://dx.doi.org/10.1109/CVPR.2019.00766}, DOI={10.1109/cvpr.2019.00766}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Qin, Xuebin and Zhang, Zichen and Huang, Chenyang and Gao, Chao and Dehghan, Masood and Jagersand, Martin}, year={2019}, month=June, pages={7471–7481} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE