Built independently by an author, for readers. Read the story and support ChapterPal

keyword

real-time semantic segmentation

Real-time semantic segmentation is a computer vision task that assigns a predefined semantic category to every pixel in an image while operating at high inference speeds and low latency, typically matching or exceeding standard video frame rates. Unlike offline semantic segmentation models that prioritize dense pixel-level accuracy at the expense of heavy computational demands, real-time approaches are designed to strike an optimal trade-off between segmentation accuracy and processing efficiency. These models frequently utilize lightweight neural network architectures, multi-resolution or multi-branch processing pipelines that separate fine spatial details from broad contextual information, and streamlined feature fusion mechanisms. This capability enables efficient, low-power visual scene understanding on embedded devices and graphics processing units, making it essential for latency-critical applications such as autonomous driving, robotic navigation, real-time video surveillance, and mobile augmented reality.

9 items

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer

Jian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang

OrganizationsAustralian National UniversityBaidu

Why you should read this

Presents a dual-resolution transformer architecture that combines GPU-friendly linear attention with cross-resolution context sharing to outperform pure CNN models in real-time semantic segmentation across benchmarks like Cityscapes and CamVid.

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmentation, which achieves better trade-off between performance and efficiency than CNN-based models. To achieve high inference efficiency on GPU-like devices, our RTFormer leverages GPU-Friendly Attention with linear complexity and discards the multi-head mechanism. Besides, we find that cross-resolution attention is more efficient to gather global context information for high-resolution branch by spreading the high level knowledge learned from low-resolution branch. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our proposed RTFormer, it achieves state-of-the-art on Cityscapes, CamVid and COCostuff, and shows promising results on ADE20K. Code is available at PaddleSeg[24]: https://github.com/PaddlePaddle/PaddleSeg.

Added

2026-10-06

Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing

Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing

Wujie Zhou, Shaohua Dong, Caie Xu, Yaguan Qian

OrganizationsZhejiang University of Science and Technology

Why you should read this

Proposes an edge-aware guidance fusion network that incorporates prior edge maps, specialized cross-modal fusion modules, and multitask deep supervision to significantly improve object boundary localization in RGB-thermal scene parsing.

RGB–thermal scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing methods fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, we propose an edge-aware guidance fusion network (EGFNet) for RGB–thermal scene parsing. First, we introduce a prior edge map generated using the RGB and thermal images to capture detailed information in the prediction map and then embed the prior edge information in the feature maps. To effectively fuse the RGB and thermal information, we propose a multimodal fusion module that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, we propose a global information module and a semantic information module to extract rich semantic information from the high-level features. For decoding, we use simple elementwise addition for cascaded feature fusion. Finally, to improve the parsing accuracy, we apply multitask deep supervision to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with state-of-the-art methods. The code and results can be found at https://github.com/ShaohuaDong2021/EGFNet.

Added

2026-09-26

PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

Jiacong Xu, Zixiang Xiong, Shankar P. Bhattacharyya

OrganizationsJohns Hopkins UniversityTexas A&M University

Why you should read this

Proposes a three-branch real-time semantic segmentation architecture inspired by PID controllers that introduces a dedicated boundary branch to prevent contextual feature overshoot, achieving state-of-the-art speed-accuracy trade-offs on Cityscapes and CamVid.

Two-branch network architecture has shown its efficiency and effectiveness in real-time semantic segmentation tasks. However, direct fusion of high-resolution details and low-frequency context has the drawback of detailed features being easily overwhelmed by surrounding contextual information. This overshoot phenomenon limits the improvement of the segmentation accuracy of existing two-branch models. In this paper, we make a connection between Convolutional Neural Networks (CNN) and Proportional-Integral-Derivative (PID) controllers and reveal that a two-branch network is equivalent to a Proportional-Integral (PI) controller, which inherently suffers from similar overshoot issues. To alleviate this problem, we propose a novel three-branch network architecture: PIDNet, which contains three branches to parse detailed, context and boundary information, respectively, and employs boundary attention to guide the fusion of detailed and context branches. Our family of PIDNets achieve the best trade-off between inference speed and accuracy and their accuracy surpasses all the existing models with similar inference speed on the Cityscapes and CamVid datasets. Specifically, PIDNet-S achieves 78.6% mIOU with inference speed of 93.2 FPS on Cityscapes and 80.1% mIOU with speed of 153.7 FPS on CamVid.

Added

2026-09-26

UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery

UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery

Libo Wang, Rui Li, Ce Zhang, Shenghui Fang, Chenxi Duan, Xiaoliang Meng, Peter M. Atkinson

Why you should read this

Proposes UNetFormer, an efficient hybrid segmentation network that combines a lightweight CNN encoder with a dual global-local attention decoder to achieve real-time inference speeds over 300 FPS alongside state-of-the-art accuracy on remote sensing benchmarks.

Semantic segmentation of remotely sensed urban scene images is required in a wide range of practical applications, such as land cover mapping, urban change detection, environmental protection, and economic this http URL by rapid developments in deep learning technologies, the convolutional neural network (CNN) has dominated semantic segmentation for many years. CNN adopts hierarchical feature representation, demonstrating strong capabilities for local information extraction. However, the local property of the convolution layer limits the network from capturing the global context. Recently, as a hot topic in the domain of computer vision, Transformer has demonstrated its great potential in global information modelling, boosting many vision-related tasks such as image classification, object detection, and particularly semantic segmentation. In this paper, we propose a Transformer-based decoder and construct a UNet-like Transformer (UNetFormer) for real-time urban scene segmentation. For efficient segmentation, the UNetFormer selects the lightweight ResNet18 as the encoder and develops an efficient global-local attention mechanism to model both global and local information in the decoder. Extensive experiments reveal that our method not only runs faster but also produces higher accuracy compared with state-of-the-art lightweight models. Specifically, the proposed UNetFormer achieved 67.8% and 52.4% mIoU on the UAVid and LoveDA datasets, respectively, while the inference speed can achieve up to 322.4 FPS with a 512x512 input on a single NVIDIA GTX 3090 GPU. In further exploration, the proposed Transformer-based decoder combined with a Swin Transformer encoder also achieves the state-of-the-art result (91.3% F1 and 84.1% mIoU) on the Vaihingen dataset. The source code will be freely available at this https URL.

Added

2026-09-25

Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges

Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges

Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, Klaus Dietmayer

OrganizationsRobert Bosch GmbHUlm UniversityUniversity of Karlsruhe

Why you should read this

Systematizes multi-sensor fusion methodologies across camera, LiDAR, and radar modalities to provide clear architectural guidelines on what, when, and how to integrate data for autonomous object detection and semantic segmentation.

Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs, Radars), and multiple sensing modalities can be fused to exploit their complementary properties. In this context, many methods have been proposed for deep multi-modal perception problems. However, there is no general guideline for network architecture design, and questions of "what to fuse", "when to fuse", and "how to fuse" remain open. This review paper attempts to systematically summarize methodologies and discuss challenges for deep multi-modal object detection and semantic segmentation in autonomous driving. To this end, we first provide an overview of on-board sensors on test vehicles, open datasets, and background information for object detection and semantic segmentation in autonomous driving research. We then summarize the fusion methodologies and discuss challenges and open questions. In the appendix, we provide tables that summarize topics and methods. We also provide an interactive online platform to navigate each reference: this https URL.

Added

2026-09-25

BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation

BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation

Changqian Yu, Changxin Gao, Jingbo Wang, Gang Yu, Chunhua Shen, Nong Sang

OrganizationsHuazhong University of Science and TechnologyTencentThe Chinese University of Hong KongUniversity of Adelaide

Why you should read this

Proposes a dual-branch bilateral network that separates high-resolution spatial details from deep semantic context, achieving an optimal trade-off between speed and accuracy with 72.6% mIoU at 156 frames per second on Cityscapes.

The low-level details and high-level semantics are both essential to the semantic segmentation task. However, to speed up the model inference, current approaches almost always sacrifice the low-level details, which leads to a considerable accuracy decrease. We propose to treat these spatial details and categorical semantics separately to achieve high accuracy and high efficiency for realtime semantic segmentation. To this end, we propose an efficient and effective architecture with a good trade-off between speed and accuracy, termed Bilateral Segmentation Network (BiSeNet V2). This architecture involves: (i) a Detail Branch, with wide channels and shallow layers to capture low-level details and generate high-resolution feature representation; (ii) a Semantic Branch, with narrow channels and deep layers to obtain high-level semantic context. The Semantic Branch is lightweight due to reducing the channel capacity and a fast-downsampling strategy. Furthermore, we design a Guided Aggregation Layer to enhance mutual connections and fuse both types of feature representation. Besides, a booster training strategy is designed to improve the segmentation performance without any extra inference cost. Extensive quantitative and qualitative evaluations demonstrate that the proposed architecture performs favourably against a few state-of-the-art real-time semantic segmentation approaches. Specifically, for a 2,048x1,024 input, we achieve 72.6% Mean IoU on the Cityscapes test set with a speed of 156 FPS on one NVIDIA GeForce GTX 1080 Ti card, which is significantly faster than existing methods, yet we achieve better segmentation accuracy.

Added

2026-09-24

ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation

ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation

Adam Paszke, Abhishek Chaurasia, Sangpil Kim, Eugenio Culurciello

OrganizationsPurdue UniversityUniversity of Warsaw

Why you should read this

Proposes ENet, a lightweight deep neural network architecture that enables real-time semantic segmentation on resource-constrained embedded devices by reducing computational cost and parameter size by over an order of magnitude without sacrificing accuracy.

The ability to perform pixel-wise semantic segmentation in real-time is of paramount importance in mobile applications. Recent deep neural networks aimed at this task have the disadvantage of requiring a large number of floating point operations and have long run-times that hinder their usability. In this paper, we propose a novel deep neural network architecture named ENet (efficient neural network), created specifically for tasks requiring low latency operation. ENet is up to 18×\times faster, requires 75×\times less FLOPs, has 79×\times less parameters, and provides similar or better accuracy to existing models. We have tested it on CamVid, Cityscapes and SUN datasets and report on comparisons with existing state-of-the-art methods, and the trade-offs between accuracy and processing time of a network. We present performance measurements of the proposed architecture on embedded systems and suggest possible software improvements that could make ENet even faster.

Added

2026-09-14

BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation

BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation

Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, Nong Sang

OrganizationsMegvii TechnologyPeking University

Why you should read this

Proposes BiSeNet, a bilateral architecture that decouples spatial detail retention from wide-context feature extraction to resolve the trade-off between speed and accuracy in real-time semantic segmentation, achieving over 100 frames per second on standard benchmarks.

Semantic segmentation requires both rich spatial information and sizeable receptive field. However, modern approaches usually compromise spatial resolution to achieve real-time inference speed, which leads to poor performance. In this paper, we address this dilemma with a novel Bilateral Segmentation Network (BiSeNet). We first design a Spatial Path with a small stride to preserve the spatial information and generate high-resolution features. Meanwhile, a Context Path with a fast downsampling strategy is employed to obtain sufficient receptive field. On top of the two paths, we introduce a new Feature Fusion Module to combine features efficiently. The proposed architecture makes a right balance between the speed and segmentation performance on Cityscapes, CamVid, and COCO-Stuff datasets. Specifically, for a 2048x1024 input, we achieve 68.4% Mean IOU on the Cityscapes test dataset with speed of 105 FPS on one NVIDIA Titan XP card, which is significantly faster than the existing methods with comparable performance.

Added

2026-09-14