I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

Hongwei ZhuPeng LiHaoran XieXuefeng YanDong LiangDapeng ChenMingqiang WeiJing Qin

article2022AAAI238 citations

Proposes a boundary-guided separated attention network that mirrors human perception by decoupling foreground and background streams to locate camouflaged objects with highly ambiguous boundaries, outperforming sixteen state-of-the-art methods across standard benchmarks.

Listen

Detecting objects that visually blend into their environments is critical for high-stakes applications such as medical image segmentation, search-and-rescue operations in harsh conditions, and surveillance. Traditional computer vision methods and standard prominent-object detectors struggle in this domain because camouflaged targets actively conceal themselves, sharing colors, textures, and ambiguous boundaries with their surrounding backgrounds.

The article evaluates a human-inspired artificial intelligence architecture called the Boundary-Guided Separated Attention Network (BSA-Net). Its primary objective is to demonstrate that coordinating dual attention mechanisms—one focusing on foreground details while the other analyzes background cues—alongside targeted boundary guidance significantly enhances the detection accuracy of camouflaged objects.

The evaluated framework employs a coarse-to-fine learning strategy built on a multi-scale feature backbone. It utilizes two dedicated processing streams: a normal attention stream to isolate foreground information and a reverse attention stream that erases target interiors to evaluate the background. A specialized boundary module then integrates edge representations directly into these features. The system was validated across three benchmark datasets—CAMO (1,250 images), CHAMELEON (76 images), and the large-scale COD10K dataset (10,000 images)—and benchmarked against 16 leading baseline and state-of-the-art methods using four standard segmentation accuracy and error metrics.

The evaluation yielded several key findings. First, BSA-Net outperformed all 16 competing methods across all evaluated benchmarks, achieving the highest overall structural similarity and lowest pixel-level error. Second, on the large-scale COD10K dataset, BSA-Net demonstrated substantial gains over leading models such as SINet, improving weighted precision-recall metrics by approximately 0.148 while reducing error rates by 0.017. Third, ablation experiments confirmed that both the separated attention mechanism and the boundary guider module were essential to performance, with each contributing measurable improvements over baseline feature extraction. Finally, qualitative visual analysis verified that the network produces crisper, more complete contours and effectively identifies subtle and small-scale hidden targets that other models miss.

These findings indicate that treating foreground and background features separately provides a more robust foundation for separating targets from confusing clutter than conventional single-stream approaches. By substantially improving boundary delineation and reducing omissions, the model reduces the operational risk of missed detections in mission-critical tasks such as polyp or lung infection segmentation in healthcare, as well as rapid hazard identification during rescue missions.

Organizations developing automated image segmentation pipelines should consider adopting dual-stream foreground-background separation and explicit boundary conditioning modules to enhance edge-detection accuracy. For future development, the article recommends investigating synthetic camouflage generation to augment training pipelines and incorporating depth data alongside standard imagery to further boost detection robustness.

Confidence in these findings is supported by consistent performance across three independent benchmark datasets. However, stakeholders should note that real-world deployment will depend on operational image resolutions and environmental conditions that may differ from standard experimental datasets, warranting targeted validation before deployment in safety-critical workflows.

Cover for I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

Abstract

Can you find me? By simulating how humans to discover the so-called ‘perfectly’-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object’s boundary) between an image’s background and foreground: the reverse attention stream helps erase the camouflaged object’s interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of the boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that our BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs.

Table of Contents

  • Abstract
  • Introduction
  • Related Work
  • Salient Object Detection (SOD)
  • Camouflaged Object Detection (COD)
  • Methodology
  • Network Architecture
  • Residual Multi-scale Feature Extractor
  • Separated Attention
  • Boundary Guider
  • Loss Function
  • Settings
  • Comparison with SOTA Methods
  • Experiment
  • Ablation Study
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Boundary-Guided Separated Attention Network (BSA-Net) Architecture

    model/method

    BSA-Net is a coarse-to-fine deep neural network designed for camouflaged object detection (COD) that explicitly decouples foreground and background attention while incorporating boundary guidance.

    Given an input image I∈RW×H×3I \in \mathbb{R}^{W \times H \times 3}, where WW and HH denote the width and height, feature extraction and prediction proceed as follows:

    1. Backbone Feature Extraction: A Res2Net backbone extracts multi-level feature representations FiF_i for stage i∈{1,2,3,4,5}i \in \{1, 2, 3, 4, 5\}.
    2. Context Enrichment: Multi-level features F2,F3,F4,F5F_2, F_3, F_4, F_5 are processed in parallel by Residual Multi-scale Feature Extractor (RMFE) blocks to generate context-enhanced feature maps RMFEi\text{RMFE}_i.
    3. Separated Attention and Boundary Guidance: The features are fed into Separated Attention (SEA) modules containing two streams: a reverse attention stream that erases the object interior to attend to background context, and a normal attention stream that restores the interior to focus on the foreground. Both streams are modulated by a Boundary Guider (BG) module using boundary predictions from a dedicated Boundary Detector.
    4. Coarse Predictions and Refinement: Coarse prediction maps CiC_i (i∈{1,2,3,4}i \in \{1, 2, 3, 4\}) are obtained from the reverse attention streams under deep supervision. Shuffle Attention (SHA) blocks and feature fusion units then refine the outputs across layers, yielding four refined saliency maps RiR_i (i∈{1,2,3,4}i \in \{1, 2, 3, 4\}). The finest prediction R1R_1 serves as the final binary segmentation output during inference.
  2. Knowl 2 — Separated Attention (SEA) Module

    model/method

    The Separated Attention (SEA) module coordinates foreground and background representations to isolate the camouflaged object boundary.

    Given the ii-th layer multi-scale feature RMFEi∈RC×Hi×Wi\text{RMFE}_i \in \mathbb{R}^{C \times H_i \times W_i} and the upsampled coarse prediction from the (i+1)(i+1)-th deeper layer Ci+1C_{i+1}, the foreground and background attention maps are defined as:

    Wfai=σ(Ci+1)W_{\text{fai}} = \sigma(C_{i+1}) Wbai=1−σ(Ci+1)W_{\text{bai}} = 1 - \sigma(C_{i+1})

    where σ(⋅)\sigma(\cdot) is the sigmoid function, and both attention maps are channel-expanded to 64 channels. The background-focused and foreground-focused stream features are computed as:

    Bai=Outi=Convs(RMFEi⊗expand(Wbai))B_{ai} = \text{Out}_i = \text{Conv}_s(\text{RMFE}_i \otimes \text{expand}(W_{\text{bai}})) Fai=Convs(RMFEi⊗expand(Wfai))F_{ai} = \text{Conv}_s(\text{RMFE}_i \otimes \text{expand}(W_{\text{fai}}))

    where Convs\text{Conv}_s denotes a 1×11 \times 1 convolution, ⊗\otimes is element-wise multiplication, and Outi\text{Out}_i represents the coarse output map supervised by ground truth. To balance the contributions of the two streams across spatial and channel dimensions, a Multi-scale Channel Attention Module (MS-CAM) computes a scale weight matrix:

    W(X)=G(σ(G(X)))+G(σ(L(X)))W(X) = G(\sigma(G(X))) + G(\sigma(L(X)))

    where G(X)G(X) applies global average pooling followed by convolution, and L(X)L(X) applies pointwise convolution for local feature extraction. Using X=Bai+FaiX = B_{ai} + F_{ai}, the two streams are gated, modulated by the Boundary Guider module (BGi\text{BG}_i), and integrated:

    SEAFi=BGi(W(Bai+Fai)⊗Bai,BM)\text{SEA}_{Fi} = \text{BG}_i(W(B_{ai} + F_{ai}) \otimes B_{ai}, BM) SEABi=BGi((1−W(Bai+Fai))⊗Fai,BM)\text{SEA}_{Bi} = \text{BG}_i((1 - W(B_{ai} + F_{ai})) \otimes F_{ai}, BM) SEAi=SEAFi⊕SEABi,i∈{2,3,4}\text{SEA}_i = \text{SEA}_{Fi} \oplus \text{SEA}_{Bi}, \quad i \in \{2, 3, 4\}

    where BMBM is the predicted boundary map and ⊕\oplus denotes element-wise addition.

  3. Knowl 3 — Boundary Guider (BG) Module and Boundary Detector

    model/method

    The Boundary Guider (BG) module injects boundary priors directly into the attention stream feature maps using conditional batch normalization / adaptive space normalization.

    Given an Attention Stream Map ASMi\text{ASM}_i from the SEA module and the predicted boundary map BMBM, the boundary-guided feature map BGMi\text{BGM}_i is computed as:

    BGMi=CB(ASMi)⊗γ(BM)⊕β(BM)\text{BGM}_i = \text{CB}(\text{ASM}_i) \otimes \gamma(BM) \oplus \beta(BM)

    where CB\text{CB} represents a 3×33 \times 3 convolution followed by batch normalization; γ(BM)\gamma(BM) and β(BM)\beta(BM) are affine transformation parameters generated by applying two separate 3×33 \times 3 convolutional layers to BMBM, projecting the single-channel boundary map to 64 channels to match the feature channels of CB(ASMi)\text{CB}(\text{ASM}_i); ⊗\otimes and ⊕\oplus denote element-wise multiplication and addition, respectively.

    The boundary prediction map BMBM is produced by a dedicated boundary detector network that concatenates four multi-level feature maps from the backbone network and passes them through a convolutional layer supervised by the ground-truth boundary map BGBG (computed as the spatial gradient map of the binary ground-truth mask).

  4. Knowl 4 — Residual Multi-scale Feature Extractor (RMFE)

    model/method

    The Residual Multi-scale Feature Extractor (RMFE) expands the receptive field of backbone features FiF_i (i∈{2,3,4,5}i \in \{2, 3, 4, 5\}) using four cascaded residual branches.

    For an input feature map FiF_i, branch outputs Boutki\text{Bout}_k^i for k∈{1,2,3,4}k \in \{1, 2, 3, 4\} are computed as:

    Boutki={Convr(Fi)k=1Convr(Fi⊕Boutk−1i)k=2,3,4\text{Bout}_k^i = \begin{cases} \text{Conv}_r(F_i) & k = 1 \\ \text{Conv}_r(F_i \oplus \text{Bout}_{k-1}^i) & k = 2, 3, 4 \end{cases}

    where ⊕\oplus denotes element-wise addition, and Convr(⋅)\text{Conv}_r(\cdot) denotes a composite block consisting of a 1×11 \times 1 convolution (to reduce channel dimensionality) followed by asymmetric 1×31 \times 3 and 3×13 \times 1 convolutions to enlarge receptive fields while reducing computational cost.

    The four branch outputs are concatenated, projected to 64 channels via a 1×11 \times 1 convolution, and added to the projected input feature:

    RMFEi=Conv(Fi)⊕Conv(Catk=14(Boutki))\text{RMFE}_i = \text{Conv}(F_i) \oplus \text{Conv}\left(\text{Cat}_{k=1}^4(\text{Bout}_k^i)\right)

    where Conv(⋅)\text{Conv}(\cdot) denotes a 1×11 \times 1 convolution and Catk=14\text{Cat}_{k=1}^4 denotes concatenation along the channel dimension.

  5. Knowl 5 — Multi-Task Supervision Loss Function for BSA-Net

    equation

    BSA-Net is trained end-to-end using a multi-task loss that supervises 9 target maps: 4 coarse segmentation maps (C1,C2,C3,C4C_1, C_2, C_3, C_4), 4 refined segmentation maps (R1,R2,R3,R4R_1, R_2, R_3, R_4), and 1 boundary map (BB):

    L=∑i=14[Lt(Ci,G)+Lt(Ri,G)]+Lbce(B,BG)\mathcal{L} = \sum_{i=1}^4 \left[ \mathcal{L}_t(C_i, G) + \mathcal{L}_t(R_i, G) \right] + \mathcal{L}_{\text{bce}}(B, BG)

    where GG is the binary ground-truth segmentation mask, BGBG is the ground-truth boundary mask, and Lbce\mathcal{L}_{\text{bce}} is standard binary cross-entropy loss.

    For each segmentation map prediction PP, the combined segmentation loss Lt(P,G)\mathcal{L}_t(P, G) is:

    Lt=LWbce+LIOU\mathcal{L}_t = \mathcal{L}_{\text{Wbce}} + \mathcal{L}_{\text{IOU}}

    where LIOU\mathcal{L}_{\text{IOU}} is the intersection-over-union loss, and LWbce\mathcal{L}_{\text{Wbce}} is a weighted binary cross-entropy loss that assigns higher weights to hard-to-predict pixels:

    LWbce=−∑n=1Nwn[Gnln⁡(Pn)+(1−Gn)ln⁡(1−Pn)]\mathcal{L}_{\text{Wbce}} = -\sum_{n=1}^N w_n \left[ G_n \ln(P_n) + (1 - G_n) \ln(1 - P_n) \right]

    with pixel weight wn=σ(∣Pn−Gn∣)w_n = \sigma(|P_n - G_n|), where NN is the total number of image pixels, and Pn,GnP_n, G_n denote predicted and ground-truth values at pixel nn.

  6. Knowl 6 — Comparative Performance on COD Benchmark Datasets

    data/table

    BSA-Net was evaluated against 16 state-of-the-art salient object detection (SOD) and camouflaged object detection (COD) methods across three benchmark datasets: CAMO, CHAMELEON, and COD10K. The evaluation metrics are Structure-measure (Sα↑S_\alpha \uparrow), Enhanced-alignment measure (Eϕ↑E_\phi \uparrow), Weighted F-measure (Fβω↑F_\beta^\omega \uparrow), and Mean Absolute Error (MAE↓\text{MAE} \downarrow).

    Methods CAMO CHAMELEON COD10K
    Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow
    UNet++ (2018) 0.599 0.653 0.392 0.149 0.695 0.762 0.501 0.094 0.623 0.672 0.350 0.086
    PiCANet (2018) 0.609 0.584 0.356 0.156 0.769 0.749 0.536 0.085 0.649 0.643 0.322 0.090
    MSRCNN (2019) 0.617 0.669 0.454 0.133 0.637 0.686 0.443 0.091 0.641 0.706 0.419 0.073
    BASNet (2019) 0.618 0.661 0.413 0.159 0.687 0.721 0.474 0.118 0.634 0.678 0.365 0.105
    PFANet (2019) 0.659 0.622 0.391 0.172 0.679 0.648 0.378 0.144 0.636 0.618 0.286 0.128
    CPD (2019) 0.726 0.729 0.550 0.115 0.853 0.866 0.706 0.052 0.747 0.770 0.508 0.059
    HTC (2019) 0.476 0.442 0.174 0.172 0.517 0.489 0.204 0.129 0.548 0.520 0.221 0.088
    EGNet (2019) 0.732 0.768 0.583 0.104 0.848 0.870 0.702 0.050 0.737 0.779 0.509 0.056
    ANet-SRM (2019) 0.682 0.685 0.484 0.126 - - - - - - - -
    SINet (2020) 0.751 0.771 0.606 0.100 0.869 0.891 0.740 0.044 0.771 0.806 0.551 0.051
    PraNet (2020) 0.769 0.824 0.663 0.094 0.860 0.907 0.763 0.044 0.789 0.861 0.629 0.045
    MCIF-Net (2021) 0.784 0.845 0.677 0.084 - - - - 0.787 0.872 0.636 0.042
    TANet (2021) 0.793 0.834 0.690 0.083 0.888 0.911 0.786 0.036 0.803 0.848 0.629 0.041
    R-MGL (2021) 0.775 0.847 0.673 0.088 0.893 0.923 0.813 0.030 0.814 0.865 0.666 0.035
    PFNet (2021) 0.782 0.841 0.695 0.085 0.882 0.931 0.810 0.033 0.800 0.877 0.660 0.040
    Rank-Net (2021) 0.787 0.838 0.696 0.080 0.890 0.935 0.822 0.030 0.804 0.880 0.673 0.037
    BSA-Net (Ours) 0.796 0.851 0.717 0.079 0.895 0.946 0.841 0.027 0.818 0.891 0.699 0.034

    BSA-Net outperforms all sixteen comparison models on all four metrics across CAMO, CHAMELEON, and COD10K datasets.

  7. Knowl 7 — Ablation Study on BSA-Net Modules

    data/table

    An ablation study evaluated the progressive contributions of the Separated Attention (SEA) and Boundary Guider (BG) modules starting from a baseline composed of the Res2Net backbone and Residual Multi-scale Feature Extractor (RMFE).

    Model CAMO-Test CHAMELEON COD10K-Test
    Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow Sα↑S_\alpha \uparrow Eϕ↑E_\phi \uparrow Fβω↑F_\beta^\omega \uparrow MAE↓\text{MAE} \downarrow
    a. Baseline 0.783 0.840 0.691 0.083 0.879 0.933 0.807 0.032 0.805 0.880 0.672 0.037
    b. Baseline+SEA 0.796 0.854 0.716 0.080 0.882 0.934 0.815 0.030 0.811 0.881 0.685 0.037
    c. Baseline+BG 0.786 0.848 0.704 0.081 0.875 0.921 0.808 0.031 0.812 0.883 0.686 0.036
    d. Ours (Full BSA-Net) 0.796 0.851 0.717 0.079 0.895 0.946 0.841 0.027 0.818 0.891 0.699 0.034
    • Adding SEA (Model b vs. Model a) consistently boosts FβωF_\beta^\omega from 0.691 to 0.716 on CAMO-Test, 0.807 to 0.815 on CHAMELEON, and 0.672 to 0.685 on COD10K-Test, confirming that separating foreground and background attention streams helps resolve ambiguity at camouflage boundaries.
    • Adding BG alone (Model c vs. Model a) also improves boundary sensitivity.
    • The full model combining both SEA and BG (Model d) achieves the best overall performance across all three datasets.
  8. Knowl 8 — Experimental Setup and Training Configuration for BSA-Net

    experimental setup

    BSA-Net is implemented in PyTorch and evaluated under the following settings:

    • Benchmark Datasets:
      • CAMO: 1250 camouflaged images (1000 train, 250 test) and 1250 non-camouflaged images from MS-COCO (1000 train, 250 test).
      • CHAMELEON: 76 images used exclusively for testing.
      • COD10K: 6000 training images and 4000 test images.
    • Network Initialization: Res2Net (blocks 1 to 5) pretrained on ImageNet serves as the backbone. Convolutional and linear layers are initialized with Kaiming normal initialization.
    • Optimization: Adam optimizer with an initial learning rate of 8×10−58 \times 10^{-5} and weight decay of 0.10.1.
    • Training Parameters: Batch size of 36, trained for up to 35 epochs. Input images are resized to 384×384384 \times 384 using bilinear interpolation.
    • Hardware: Single Nvidia GeForce RTX 3090 GPU (24GB memory).

Coverage note — None was omitted; all key architectural components (RMFE, SEA, BG, Boundary Detector), multi-task loss formulation, benchmark evaluations, ablation studies, and implementation details are fully captured.

References

  1. 1.Bhajantri, N. U.; and Nagabhushan, P. 2006. Camouflage defect identification: a novel approach. In 9th International Conference on Information Technology (ICIT’06), 145–148. IEEE.
  2. 2.Chen, K.; Pang, J.; Wang, J.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Shi, J.; Ouyang, W.; et al. 2019. Hybrid task cascade for instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4974–4983.
  3. 3.Chen, S.; Tan, X.; Wang, B.; and Hu, X. 2018. Reverse attention for salient object detection. In Proceedings of the European Conference on Computer Vision (ECCV), 234–250.
  4. 4.Chu, H.; Hsu, W.; Mitra, N. J.; Cohen-Or, D.; Wong, T.; and Lee, T. 2010. Camouflage images. ACM Trans. Graph., 29(4): 51:1–51:8.
  5. 5.Dai, Y.; Gieseke, F.; Oehmcke, S.; Wu, Y.; and Barnard, K. 2021. Attentional Feature Fusion. In IEEE Winter Conference on Applications of Computer Vision, WACV 2021, Waikoloa, HI, USA, January 3-8, 2021, 3559–3568. IEEE.
  6. 6.Dong, B.; Zhuge, M.; Wang, Y.; Bi, H.; and Chen, G. 2021. Towards Accurate Camouflaged Object Detection with Mixture Convolution and Interactive Fusion. CoRR, abs/2101.05687.
  7. 7.Fan, D.; Gong, C.; Cao, Y.; Ren, B.; Cheng, M.; and Borji, A. 2018. Enhanced-alignment Measure for Binary Foreground Map Evaluation. In Lang, J., ed., Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, 698–704. ijcai.org.
  8. 8.Fan, D.-P.; Cheng, M.-M.; Liu, Y.; Li, T.; and Borji, A. 2017. Structure-measure: A new way to evaluate foreground maps. In Proceedings of the IEEE international conference on computer vision, 4548–4557.
  9. 9.Fan, D.-P.; Ji, G.-P.; Sun, G.; Cheng, M.-M.; Shen, J.; and Shao, L. 2020a. Camouflaged object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2777–2787.
  10. 10.Fan, D.-P.; Ji, G.-P.; Zhou, T.; Chen, G.; Fu, H.; Shen, J.; and Shao, L. 2020b. Pranet: Parallel reverse attention network for polyp segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 263–273. Springer.
  11. 11.Feng, M.; Lu, H.; and Ding, E. 2019. Attentive feedback network for boundary-aware salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1623–1632.
  12. 12.Gao, S.; Cheng, M.; Zhao, K.; Zhang, X.; Yang, M.; and Torr, P. H. S. 2021. Res2Net: A New Multi-Scale Backbone Architecture. IEEE Trans. Pattern Anal. Mach. Intell., 43(2): 652–662.
  13. 13.Huang, Z.; Huang, L.; Gong, Y.; Huang, C.; and Wang, X. 2019. Mask scoring r-cnn. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6409–6418.
  14. 14.Huerta, I.; Rowe, D.; Mozerov, M.; and Gonzalez, J. 2007. Improving background subtraction based on a casuistry of colour-motion segmentation problems. In Iberian Conference on Pattern Recognition and Image Analysis, 475–482. Springer.
  15. 15.Kavitha, C.; Rao, B. P.; and Govardhan, A. 2011. An efficient content based image retrieval using color and texture of image sub blocks. International Journal of Engineering Science and Technology (IJEST), 3(2): 1060–1068.
  16. 16.Le, T.; Nguyen, T. V.; Nie, Z.; Tran, M.; and Sugimoto, A. 2019. Anabranch network for camouflaged object segmentation. Comput. Vis. Image Underst., 184: 45–56.
  17. 17.Li, A.; Zhang, J.; Lv, Y.; Liu, B.; Zhang, T.; and Dai, Y. 2021. Uncertainty-Aware Joint Salient Object and Camouflaged Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, 10071–10081. Computer Vision Foundation / IEEE.
  18. 18.Liu, J.; Zhang, J.; and Barnes, N. 2021. Confidence-Aware Learning for Camouflaged Object Detection. CoRR, abs/2106.11641.
  19. 19.Liu, N.; Han, J.; and Yang, M.-H. 2018. Picanet: Learning pixel-wise contextual attention for saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3089–3098.
  20. 20.Lv, Y.; Zhang, J.; Dai, Y.; Li, A.; Liu, B.; Barnes, N.; and Fan, D. 2021. Simultaneously Localize, Segment and Rank the Camouflaged Objects. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, 11591–11601. Computer Vision Foundation / IEEE.
  21. 21.Mao, Y.; Zhang, J.; Wan, Z.; Dai, Y.; Li, A.; Lv, Y.; Tian, X.; Fan, D.; and Barnes, N. 2021. Transformer Transforms Salient Object Detection and Camouflaged Object Detection. CoRR, abs/2104.10127.
  22. 22.Margolin, R.; Zelnik-Manor, L.; and Tal, A. 2014. How to evaluate foreground maps? In Proceedings of the IEEE conference on computer vision and pattern recognition, 248–255.
  23. 23.Mei, H.; Ji, G.; Wei, Z.; Yang, X.; Wei, X.; and Fan, D. 2021. Camouflaged Object Segmentation With Distraction Mining. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, 8772–8781. Computer Vision Foundation / IEEE.
  24. 24.Pang, Y.; Zhao, X.; Zhang, L.; and Lu, H. 2020. Multi-scale interactive network for salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9413–9422.
  25. 25.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Wallach, H. M.; Larochelle, H.; Beygelzimer, A.; d’Alche-Buc, F.; Fox, E. B.; and Garnett, R., eds., Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, 8024–8035.
  26. 26.Perazzi, F.; Krähenbühl, P.; Pritch, Y.; and Hornung, A. 2012. Saliency filters: Contrast based filtering for salient region detection. In 2012 IEEE conference on computer vision and pattern recognition, 733–740. IEEE.
  27. 27.Qin, X.; Zhang, Z.; Huang, C.; Gao, C.; Dehghan, M.; and Jagersand, M. 2019. Basnet: Boundary-aware salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7479–7489.
  28. 28.Ren, J.; Hu, X.; Zhu, L.; Xu, X.; Xu, Y.; Wang, W.; Deng, Z.; and Heng, P. 2021. Deep Texture-Aware Features for Camouflaged Object Detection. CoRR, abs/2102.02996.
  29. 29.Skurowski, P.; Abdulameer, H.; Błaszczyk, J.; Depta, T.; Kornacki, A.; and Kozieł, P. 2018. Animal camouflage analysis: Chameleon database. Unpublished Manuscript.
  30. 30.Sun, Y.; Chen, G.; Zhou, T.; Zhang, Y.; and Liu, N. 2021. Context-aware Cross-level Fusion Network for Camouflaged Object Detection. In Zhou, Z., ed., Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, 1025–1031. ijcai.org.
  31. 31.Tankus, A.; and Yeshurun, Y. 2001. Convexity-Based Visual Camouflage Breaking. Comput. Vis. Image Underst., 82(3): 208–237.
  32. 32.Wei, J.; Wang, S.; and Huang, Q. 2020. F3Net: Fusion, Feedback and Focus for Salient Object Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 12321–12328.
  33. 33.Wei, J.; Wang, S.; Wu, Z.; Su, C.; Huang, Q.; and Tian, Q. 2020. Label Decoupling Framework for Salient Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13022–13031. IEEE.
  34. 34.Wu, Y.; Gao, S.; Mei, J.; Xu, J.; Fan, D.; Zhang, R.; and Cheng, M. 2021. JCS: An Explainable COVID-19 Diagnosis System by Joint Classification and Segmentation. IEEE Trans. Image Process., 30: 3113–3126.
  35. 35.Wu, Z.; Su, L.; and Huang, Q. 2019a. Cascaded partial decoder for fast and accurate salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3907–3916.
  36. 36.Wu, Z.; Su, L.; and Huang, Q. 2019b. Stacked cross refinement network for edge-aware salient object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 7264–7273.
  37. 37.Xue, F.; Guoying, C.; Hong, R.; and Gu, J. 2015. Camouflage texture evaluation using a saliency map. Multim. Syst., 21(2): 169–175.
  38. 38.Yan, J.; Le, T.; Nguyen, K.; Tran, M.; Do, T.; and Nguyen, T. V. 2021. MirrorNet: Bio-Inspired Camouflaged Object Segmentation. IEEE Access, 9: 43290–43300.
  39. 39.Yang, Q.-L. Z. Y.-B. 2021. SA-Net: Shuffle Attention for Deep Convolutional Neural Networks. arXiv preprint arXiv:2102.00240.
  40. 40.Zhai, Q.; Li, X.; Yang, F.; Chen, C.; Cheng, H.; and Fan, D. 2021. Mutual Graph Learning for Camouflaged Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, 12997–13007. Computer Vision Foundation / IEEE.
  41. 41.Zhang, J.; Lv, Y.; Xiang, M.; Li, A.; Dai, Y.; and Zhong, Y. 2021. Depth-Guided Camouflaged Object Detection. CoRR, abs/2106.13217.
  42. 42.Zhao, J.-X.; Liu, J.-J.; Fan, D.-P.; Cao, Y.; Yang, J.; and Cheng, M.-M. 2019. EGNet: Edge guidance network for salient object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8779–8788.
  43. 43.Zhao, T.; and Wu, X. 2019. Pyramid feature attention network for saliency detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3085–3094.
  44. 44.Zhou, Z.; Siddiquee, M. M. R.; Tajbakhsh, N.; and Liang, J. 2018. Unet++: A nested u-net architecture for medical image segmentation. In Deep learning in medical image analysis and multimodal learning for clinical decision support, 3–11. Springer.
  45. 45.Zhu, L.; Chen, J.; Hu, X.; Fu, C.-W.; Xu, X.; Qin, J.; and Heng, P.-A. 2019. Aggregating attentional dilated features for salient object detection. IEEE Transactions on Circuits and Systems for Video Technology, 30(10): 3358–3371.

Citation

MLA
Zhu, H., et al. “I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 3608–16, https://doi.org/10.1609/AAAI.V36I3.20273.
APA
Zhu, H., Li, P., Xie, H., Yan, X., Liang, D., Chen, D., Wei, M., & Qin, J. (2022). I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), 3608–3616. https://doi.org/10.1609/AAAI.V36I3.20273
Chicago
Zhu, H., P. Li, H. Xie, et al. 2022. “I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (3): 3608–16. https://doi.org/10.1609/AAAI.V36I3.20273.
Harvard
Zhu, H. et al. (2022) “I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), pp. 3608–3616. Available at: https://doi.org/10.1609/AAAI.V36I3.20273.
Vancouver
1. Zhu H, Li P, Xie H, Yan X, Liang D, Chen D, Wei M, Qin J (2022) I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence 36:3608–3616

BibTeX

@article{Zhu_2022, title={I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I3.20273}, DOI={10.1609/aaai.v36i3.20273}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhu, Hongwei and Li, Peng and Xie, Haoran and Yan, Xuefeng and Liang, Dong and Chen, Dapeng and Wei, Mingqiang and Qin, Jing}, year={2022}, month=June, pages={3608–3616} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF