Built independently by an author, for readers. Read the story and support ChapterPal

keyword

salient object detection

Salient object detection is a computer vision task that aims to automatically identify and segment the most visually prominent or attention-grabbing objects within an image or video, simulating the human visual attention mechanism. Instead of predicting human gaze fixations as isolated focus points, this process extracts the complete spatial extent of dominant foreground objects, generating accurate pixel-level masks that clearly separate them from the surrounding background. Unlike conventional semantic segmentation or standard object detection, which classify every object into predetermined category labels, salient object detection determines visual significance based on contrast, feature distinction, and contextual cues regardless of the specific semantic class.

10 items

GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Ding Jia, Jianyuan Guo, Kai Han, Han Wu, Chao Zhang, Chang Xu, Xinghao Chen

Why you should read this

Proposes GeminiFusion, a multimodal vision transformer framework that achieves linear computational complexity by combining intra-modal and inter-modal attention at aligned spatial positions, matching the efficiency of unimodal models while outperforming token exchange and full cross-attention across diverse segmentation, detection, and translation benchmarks.

Cross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal features, and demonstrate exchange based methods underperform cross-attention mechanisms, while the computational demand of the latter inevitably restricts its use with longer sequences. To surmount the computational challenges, we propose GeminiFusion, a pixel-wise fusion approach that capitalizes on aligned cross-modal representations. GeminiFusion elegantly combines intra-modal and inter-modal attentions, dynamically integrating complementary information across modalities. We employ a layer-adaptive noise to adaptively control their interplay on a per-layer basis, thereby achieving a harmonized fusion process. Notably, GeminiFusion maintains linear complexity with respect to the number of input tokens, ensuring this multimodal framework operates with efficiency comparable to unimodal networks. Comprehensive evaluations across multimodal image-to-image translation, 3D object detection and arbitrary-modal semantic segmentation tasks, including RGB, depth, LiDAR, event data, etc. demonstrate the superior performance of our GeminiFusion against leading-edge techniques. The PyTorch code is available here.

Added

2026-10-03

Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing

Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing

Wujie Zhou, Shaohua Dong, Caie Xu, Yaguan Qian

OrganizationsZhejiang University of Science and Technology

Why you should read this

Proposes an edge-aware guidance fusion network that incorporates prior edge maps, specialized cross-modal fusion modules, and multitask deep supervision to significantly improve object boundary localization in RGB-thermal scene parsing.

RGB–thermal scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing methods fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, we propose an edge-aware guidance fusion network (EGFNet) for RGB–thermal scene parsing. First, we introduce a prior edge map generated using the RGB and thermal images to capture detailed information in the prediction map and then embed the prior edge information in the feature maps. To effectively fuse the RGB and thermal information, we propose a multimodal fusion module that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, we propose a global information module and a semantic information module to extract rich semantic information from the high-level features. For decoding, we use simple elementwise addition for cascaded feature fusion. Finally, to improve the parsing accuracy, we apply multitask deep supervision to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with state-of-the-art methods. The code and results can be found at https://github.com/ShaohuaDong2021/EGFNet.

Added

2026-09-26

Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection

Yi Wang, Ruili Wang, Xin Fan, Tianzhu Wang, Xiangjian He

OrganizationsDalian University of TechnologyMassey UniversityUniversity of Nottingham

Why you should read this

Proposes Multiple Enhancement Network (MENet), which integrates human visual system mechanisms through a dual-branch decoder, multiscale feature enhancement modules, and a multi-level hybrid loss across pixel, region, and object scales to achieve state-of-the-art salient object detection in complex scenes.

Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet.

Added

2026-09-26

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, Wenqi Shao

OrganizationsShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversityShenzhen Institute of Advanced Technology, Chinese Academy of SciencesUniversity of AdelaideUniversity of Hong KongZhejiang University

Why you should read this

Presents MMT-Bench, an extensive evaluation benchmark spanning over 31,000 questions across 162 vision-language tasks, to expose key performance limits in advanced multimodal models and map their progress toward general visual intelligence.

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, and reasoning. MMT-Bench comprises 31,325 meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering 32 core meta-tasks and 162 subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving 32 LVLMs such as the proprietary GPT-4o, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.

Added

2026-09-26

I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

Hongwei Zhu, Peng Li, Haoran Xie, Xuefeng Yan, Dong Liang, Dapeng Chen, Mingqiang Wei, Jing Qin

OrganizationsHong Kong Polytechnic UniversityHuaweiLingnan UniversityMIIT Key Laboratory of Pattern Analysis and Machine IntelligenceNanjing University of Aeronautics and Astronautics

Why you should read this

Proposes a boundary-guided separated attention network that mirrors human perception by decoupling foreground and background streams to locate camouflaged objects with highly ambiguous boundaries, outperforming sixteen state-of-the-art methods across standard benchmarks.

Can you find me? By simulating how humans to discover the so-called ‘perfectly’-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object’s boundary) between an image’s background and foreground: the reverse attention stream helps erase the camouflaged object’s interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of the boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that our BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs.

Added

2026-09-26

Visual saliency based on multiscale deep features

Visual saliency based on multiscale deep features

Guanbin Li, Yizhou Yu

OrganizationsUniversity of Hong Kong

Why you should read this

Introduces a multiscale deep convolutional framework for visual saliency detection that integrates spatial coherence refinement across multiple segmentation levels and establishes state-of-the-art accuracy alongside the 4,447-image HKU-IS benchmark dataset.

Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this CVPR 2015 paper, we discover that a high-quality visual saliency model can be trained with multiscale features extracted using a popular deep learning architecture, convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for extracting features at three different scales. We then propose a refinement method to enhance the spatial coherence of our saliency results. Finally, aggregating multiple saliency maps computed for different levels of image segmentation can further boost the performance, yielding saliency maps better than those generated from a single segmentation. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotation. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks, improving the F-Measure by 5.0% and 13.2% respectively on the MSRA-B dataset and our new dataset (HKU-IS), and lowering the mean absolute error by 5.7% and 35.1% respectively on these two datasets.

Added

2026-09-25

Deeply Supervised Salient Object Detection with Short Connections

Deeply Supervised Salient Object Detection with Short Connections

Qibin Hou, Ming-Ming Cheng, Xiao-Wei Hu, Ali Borji, Zhuowen Tu, Philip Torr

OrganizationsNankai UniversityUniversity of California, San DiegoUniversity of Central FloridaUniversity of Oxford

Why you should read this

Introduces short connections to deeply supervised skip-layer architectures to capture multi-scale features, achieving state-of-the-art salient object detection accuracy and fast inference across five standard benchmarks.

Recent progress on saliency detection is substantial, benefiting mostly from the explosive development of Convolutional Neural Networks (CNNs). Semantic segmentation and saliency detection algorithms developed lately have been mostly based on Fully Convolutional Neural Networks (FCNs). There is still a large room for improvement over the generic FCN models that do not explicitly deal with the scale-space problem. Holistically-Nested Edge Detector (HED) provides a skip-layer structure with deep supervision for edge and boundary detection, but the performance gain of HED on salience detection is not obvious. In this paper, we propose a new method for saliency detection by introducing short connections to the skip-layer structures within the HED architecture. Our framework provides rich multi-scale feature maps at each layer, a property that is critically needed to perform segment detection. Our method produces state-of-the-art results on 5 widely tested salient object detection benchmarks, with advantages in terms of efficiency (0.15 seconds per image), effectiveness, and simplicity over the existing algorithms.

Added

2026-09-25

Salient Object Detection: A Benchmark

Salient Object Detection: A Benchmark

Ali Borji, Ming-Ming Cheng, Huaizu Jiang, Jia Li

OrganizationsBeihang UniversityNankai UniversityUniversity of Wisconsin-MilwaukeeXi'an Jiaotong University

Why you should read this

Establishes a comprehensive benchmark by evaluating forty models across six datasets, analyzing factors like center bias and scene complexity to expose key failure modes and guide future salient object detection research.

We extensively compare, qualitatively and quantitatively, 40 state-of-the-art models (28 salient object detection, 10 fixation prediction, 1 objectness, and 1 baseline) over 6 challenging datasets for the purpose of benchmarking salient object detection and segmentation methods. From the results obtained so far, our evaluation shows a consistent rapid progress over the last few years in terms of both accuracy and running time. The top contenders in this benchmark significantly outperform the models identified as the best in the previous benchmark conducted just two years ago. We find that the models designed specifically for salient object detection generally work better than models in closely related areas, which in turn provides a precise definition and suggests an appropriate treatment of this problem that distinguishes it from other problems. In particular, we analyze the influences of center bias and scene complexity in model performance, which, along with the hard cases for state-of-the-art models, provide useful hints towards constructing more challenging large scale datasets and better saliency models. Finally, we propose probable solutions for tackling several open problems such as evaluation scores and dataset bias, which also suggest future research directions in the rapidly-growing field of salient object detection.

Added

2026-09-18

PraNet: Parallel Reverse Attention Network for Polyp Segmentation

PraNet: Parallel Reverse Attention Network for Polyp Segmentation

Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, Ling Shao

OrganizationsInception Institute of AIMohamed bin Zayed University of Artificial IntelligenceWuhan University

Why you should read this

Introduces PraNet, an efficient deep network that resolves vague lesion boundaries in colonoscopy images by coupling high-level context aggregation with reverse attention modules to achieve accurate, real-time colorectal polyp segmentation.

Colonoscopy is an effective technique for detecting colorectal polyps, which are highly related to colorectal cancer. In clinical practice, segmenting polyps from colonoscopy images is of great importance since it provides valuable information for diagnosis and surgery. However, accurate polyp segmentation is a challenging task, for two major reasons: (i) the same type of polyps has a diversity of size, color and texture; and (ii) the boundary between a polyp and its surrounding mucosa is not sharp. To address these challenges, we propose a parallel reverse attention network (PraNet) for accurate polyp segmentation in colonoscopy images. Specifically, we first aggregate the features in high-level layers using a parallel partial decoder (PPD). Based on the combined feature, we then generate a global map as the initial guidance area for the following components. In addition, we mine the boundary cues using a reverse attention (RA) module, which is able to establish the relationship between areas and boundary cues. Thanks to the recurrent cooperation mechanism between areas and boundaries, our PraNet is capable of calibrating any misaligned predictions, improving the segmentation accuracy. Quantitative and qualitative evaluations on five challenging datasets across six metrics show that our PraNet improves the segmentation accuracy significantly, and presents a number of advantages in terms of generalizability, and real-time segmentation efficiency.

Added

2026-09-18

U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection

U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection

Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R Zaiane, Martin Jägersand

OrganizationsUniversity of Alberta

Why you should read this

Introduces U2-Net, a two-level nested U-structure network that captures multi-scale contextual information through residual U-blocks, enabling accurate salient object detection trained completely from scratch without pre-trained classification backbones.

In this paper, we design a simple yet powerful deep network architecture, U2^2-Net, for salient object detection (SOD). The architecture of our U2^2-Net is a two-level nested U-structure. The design has the following advantages: (1) it is able to capture more contextual information from different scales thanks to the mixture of receptive fields of different sizes in our proposed ReSidual U-blocks (RSU), (2) it increases the depth of the whole architecture without significantly increasing the computational cost because of the pooling operations used in these RSU blocks. This architecture enables us to train a deep network from scratch without using backbones from image classification tasks. We instantiate two models of the proposed architecture, U2^2-Net (176.3 MB, 30 FPS on GTX 1080Ti GPU) and U2^2-Net†^{\dagger} (4.7 MB, 40 FPS), to facilitate the usage in different environments. Both models achieve competitive performance on six SOD datasets. The code is available: this https URL.

Added

2026-09-16