keyword
atrous spatial pyramid
An atrous spatial pyramid is a neural network architectural module designed to capture multiscale visual information and contextual relationships in digital images without reducing spatial resolution. Commonly employed in computer vision tasks like semantic segmentation and salient object detection, it applies multiple parallel dilated convolutions with distinct sampling rates to an input feature map alongside standard operations such as pointwise convolutions and global pooling. By expanding filter receptive fields without increasing the total parameter count or relying on repetitive downsampling, the module enables deep learning models to robustly recognize, delineate, and segment objects of varying sizes while preserving fine boundary and spatial details.
2 items

Pixels, Regions, and Objects: Multiple Enhancement for Salient Object Detection
Yi Wang, Ruili Wang, Xin Fan, Tianzhu Wang, Xiangjian He
Why you should read this
Proposes Multiple Enhancement Network (MENet), which integrates human visual system mechanisms through a dual-branch decoder, multiscale feature enhancement modules, and a multi-level hybrid loss across pixel, region, and object scales to achieve state-of-the-art salient object detection in complex scenes.
Salient object detection (SOD) aims to mimic the human visual system (HVS) and cognition mechanisms to identify and segment salient objects. However, due to the complexity of these mechanisms, current methods are not perfect. Accuracy and robustness need to be further improved, particularly in complex scenes with multiple objects and background clutter. To address this issue, we propose a novel approach called Multiple Enhancement Network (MENet) that adopts the boundary sensibility, content integrity, iterative refinement, and frequency decomposition mechanisms of HVS. A multi-level hybrid loss is firstly designed to guide the network to learn pixel-level, region-level, and object-level features. A flexible multiscale feature enhancement module (ME-Module) is then designed to gradually aggregate and refine global or detailed features by changing the size order of the input feature sequence. An iterative training strategy is used to enhance boundary features and adaptive features in the dual-branch decoder of MENet. Comprehensive evaluations on six challenging benchmark datasets show that MENet achieves state-of-the-art results. Both the codes and results are publicly available at https://github.com/yiwangtz/MENet.
Added
2026-09-26

BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, Nong Sang
Why you should read this
Proposes BiSeNet, a bilateral architecture that decouples spatial detail retention from wide-context feature extraction to resolve the trade-off between speed and accuracy in real-time semantic segmentation, achieving over 100 frames per second on standard benchmarks.
Semantic segmentation requires both rich spatial information and sizeable receptive field. However, modern approaches usually compromise spatial resolution to achieve real-time inference speed, which leads to poor performance. In this paper, we address this dilemma with a novel Bilateral Segmentation Network (BiSeNet). We first design a Spatial Path with a small stride to preserve the spatial information and generate high-resolution features. Meanwhile, a Context Path with a fast downsampling strategy is employed to obtain sufficient receptive field. On top of the two paths, we introduce a new Feature Fusion Module to combine features efficiently. The proposed architecture makes a right balance between the speed and segmentation performance on Cityscapes, CamVid, and COCO-Stuff datasets. Specifically, for a 2048x1024 input, we achieve 68.4% Mean IOU on the Cityscapes test dataset with speed of 105 FPS on one NVIDIA Titan XP card, which is significantly faster than the existing methods with comparable performance.
Added
2026-09-14
