keyword
scene understanding
Scene understanding is a subfield of computer vision and artificial intelligence focused on enabling computational systems to comprehensively perceive, analyze, and interpret the contents, spatial structure, and contextual relationships within an environment from visual data. Rather than merely recognizing isolated objects or classifying global image types, scene understanding integrates low-level perception with high-level reasoning across two- and three-dimensional representations. This encompasses tasks such as semantic and instance segmentation of discrete objects and amorphous regions, estimating geometric layouts and depth, resolving physical and support relationships, and inferring functional affordances. By constructing a holistic, semantically rich model of how entities interact within a physical space, scene understanding serves as a foundational capability for autonomous driving, robotic manipulation, augmented reality, and embodied agents navigating complex real-world environments.
31 items

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Dorsa Sadigh, Leonidas J. Guibas, Fei Xia
Why you should read this
Introduces SpatialVLM, a model trained on a two-billion-example 3D spatial reasoning dataset that equips vision-language architectures with direct metric distance and physical size estimation for robotics.
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: this https URL
Added
2026-10-05

OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
Zilong Deng, Federico Tombari, Marc Pollefeys, Johanna Wald, Dániel Baráth
Why you should read this
Presents a real-time open-vocabulary 3D instance mapping system that decouples geometry reconstruction from semantic inference using selective-view vision-language querying, eliminating dense feature fusion while maintaining high temporal consistency during online exploration.
Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely on the closed-set assumption or dense per-pixel language fusion, which limits scalability and temporal consistency. We introduce OVI-MAP that decouples instance reconstruction from semantic inference. We propose to build a class-agnostic 3D instance map that is incrementally constructed from RGB-D input, while semantic features are extracted only from a small set of automatically selected views using vision-language models. This design enables stable instance tracking and zero-shot semantic labeling throughout online exploration. Our system operates in real time and outperforms state-of-the-art open-vocabulary mapping baselines on standard benchmarks.
Added
2026-09-29

FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
Alexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario, Boyang Sun, Rishabh Dabral, Leonidas Guibas, Christian Theobalt, Marc Pollefeys, Francis Engelmann, Daniel Barath
Why you should read this
Develops a framework for reconstructing simulation-ready articulated 3D digital twins directly from in-the-wild egocentric interaction videos, recovering part kinematics and dynamic geometry without requiring CAD priors or controlled multi-state captures.
We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunRec operates directly on in-the-wild human interaction sequences to recover interactable 3D scenes. It automatically discovers articulated parts, estimates their kinematic parameters, tracks their 3D motion, and reconstructs static and moving geometry in canonical space, yielding simulation-compatible meshes. Across new real and simulated benchmarks, FunRec surpasses prior work by a large margin, achieving up to +50 mIoU improvement in part segmentation, 5-10 times lower articulation and pose errors, and significantly higher reconstruction accuracy. We further demonstrate applications on URDF/USD export for simulation, hand-guided affordance mapping and robot-scene interaction.
Added
2026-09-29

An Embodied Generalist Agent in 3D World
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, Siyuan Huang
Why you should read this
Presents LEO, a multimodal generalist agent trained on unified vision-language-action sequences to perform complex 3D perception, spatial reasoning, robotic manipulation, and embodied navigation.
Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics. However, several significant challenges remain: (i) most of these models rely on 2D images yet exhibit a limited capacity for 3D input; (ii) these models rarely explore the tasks inherently defined in 3D world, e.g., 3D grounding, embodied reasoning and acting. We argue these limitations significantly hinder current models from performing real-world tasks and approaching general intelligence. To this end, we introduce LEO, an embodied multi-modal generalist agent that excels in perceiving, grounding, reasoning, planning, and acting in the 3D world. LEO is trained with a unified task interface, model architecture, and objective in two stages: (i) 3D vision-language (VL) alignment and (ii) 3D vision-language-action (VLA) instruction tuning. We collect large-scale datasets comprising diverse object-level and scene-level tasks, which require considerable understanding of and interaction with the 3D world. Moreover, we meticulously design an LLM-assisted pipeline to produce high-quality 3D VL data. Through extensive experiments, we demonstrate LEO’s remarkable proficiency across a wide spectrum of tasks, including 3D captioning, question answering, embodied reasoning, navigation and manipulation. Our ablative studies and scaling analyses further provide valuable insights for developing future embodied generalist agents. Code and data are available on project page.
Added
2026-09-28

ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, Silvio Savarese
Why you should read this
Introduces a model-agnostic pre-training framework that aligns 3D point cloud encoders with frozen vision-language models using synthesized multimodal triplets, substantially boosting zero-shot and standard 3D recognition performance across various architectures.
The recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from other modalities, such as language. Inspired by this, leveraging multimodal information for 3D modality could be promising to improve 3D understanding under the restricted data regime, but this line of research is not well studied. Therefore, we introduce ULIP to learn a unified representation of image, text, and 3D point cloud by pre-training with object triplets from the three modalities. To overcome the shortage of training triplets, ULIP leverages a pre-trained vision-language model that has already learned a common visual and textual space by training with massive image-text pairs. Then, ULIP learns a 3D representation space aligned with the common image-text space, using a small number of automatically synthesized triplets. ULIP is agnostic to 3D backbone networks and can easily be integrated into any 3D architecture. Experiments show that ULIP effectively improves the performance of multiple recent 3D backbones by simply pre-training them on ShapeNet55 using our framework, achieving state-of-the-art performance in both standard 3D classification and zero-shot 3D classification on ModelNet40 and ScanObjectNN. ULIP also improves the performance of PointMLP by around 3% in 3D classification on ScanObjectNN, and outperforms PointCLIP by 28.8% on top-1 accuracy for zero-shot 3D classification on ModelNet40. Our code and pre-trained models will be released.
Added
2026-09-26

Generative Semantic Segmentation
Jiaqi Chen, Jiachen Lu, Xiatian Zhu, Li Zhang
Why you should read this
Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.
We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.
Added
2026-09-26

A Review on Deep Learning Techniques Applied to Semantic Segmentation
Alberto Garcia-Garcia, Sergio Orts-Escolano, Sergiu Oprea, Victor Villena-Martinez, Jose Garcia-Rodriguez
Why you should read this
Surveys deep learning methods for semantic segmentation by systematically categorizing leading architectures, comparing quantitative benchmark performance across standard datasets, and identifying key future research directions for computer vision applications.
Image semantic segmentation is more and more being of interest for computer vision and machine learning researchers. Many applications on the rise need accurate and efficient segmentation mechanisms: autonomous driving, indoor navigation, and even virtual or augmented reality systems to name a few. This demand coincides with the rise of deep learning approaches in almost every field or application target related to computer vision, including semantic segmentation or scene understanding. This paper provides a review on deep learning methods for semantic segmentation applied to various application areas. Firstly, we describe the terminology of this field as well as mandatory background concepts. Next, the main datasets and challenges are exposed to help researchers decide which are the ones that best suit their needs and their targets. Then, existing methods are reviewed, highlighting their contributions and their significance in the field. Finally, quantitative results are given for the described methods and the datasets in which they were evaluated, following up with a discussion of the results. At last, we point out a set of promising future works and draw our own conclusions about the state of the art of semantic segmentation using deep learning techniques.
Added
2026-09-25

Semantic Scene Completion from a Single Depth Image
Shuran Song, Fisher Yu, Andy Zeng, Angel X. Chang, Manolis Savva, Thomas Funkhouser
Why you should read this
Introduces an end-to-end 3D convolutional network that jointly predicts complete volumetric occupancy and semantic labels from a single depth image, alongside the large-scale SUNCG dataset for learning 3D contextual scene completion.
This paper focuses on semantic scene completion, a task for producing a complete 3D voxel representation of volumetric occupancy and semantic labels for a scene from a single-view depth map observation. Previous work has considered scene completion and semantic labeling of depth maps separately. However, we observe that these two problems are tightly intertwined. To leverage the coupled nature of these two tasks, we introduce the semantic scene completion network (SSCNet), an end-to-end 3D convolutional network that takes a single depth image as input and simultaneously outputs occupancy and semantic labels for all voxels in the camera view frustum. Our network uses a dilation-based 3D context module to efficiently expand the receptive field and enable 3D context learning. To train our network, we construct SUNCG - a manually created large-scale dataset of synthetic 3D scenes with dense volumetric annotations. Our experiments demonstrate that the joint model outperforms methods addressing each task in isolation and outperforms alternative approaches on the semantic scene completion task.
Added
2026-09-25

ICNet for Real-Time Semantic Segmentation on High-Resolution Images
Hengshuang Zhao, Xiaojuan Qi, Xiaoyong Shen, Jianping Shi, Jiaya Jia
Why you should read this
Introduces an image cascade network that combines multi-resolution branches with cascade feature fusion to deliver real-time, high-accuracy semantic segmentation on high-resolution images.
We focus on the challenging task of real-time semantic segmentation in this paper. It finds many practical applications and yet is with fundamental difficulty of reducing a large portion of computation for pixel-wise label inference. We propose an image cascade network (ICNet) that incorporates multi-resolution branches under proper label guidance to address this challenge. We provide in-depth analysis of our framework and introduce the cascade feature fusion unit to quickly achieve high-quality segmentation. Our system yields real-time inference on a single GPU card with decent quality results evaluated on challenging datasets like Cityscapes, CamVid and COCO-Stuff.
Added
2026-09-24

OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, Roozbeh Mottaghi
Why you should read this
Introduces OK-VQA, a benchmark of over 14,000 questions designed to evaluate visual question answering systems on queries that require external, world knowledge beyond the visual content of the image.
Visual Question Answering (VQA) in its ideal form lets us study reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most VQA benchmarks to date are focused on questions such as simple counting, visual attributes, and object detection that do not require reasoning or knowledge beyond what is in the image. In this paper, we address the task of knowledge-based visual question answering and provide a benchmark, called OK-VQA, where the image content is not sufficient to answer the questions, encouraging methods that rely on external knowledge resources. Our new dataset includes more than 14,000 questions that require external knowledge to answer. We show that the performance of the state-of-the-art VQA models degrades drastically in this new setting. Our analysis shows that our knowledge-based VQA task is diverse, difficult, and large compared to previous knowledge-based VQA datasets. We hope that this dataset enables researchers to open up new avenues for research in this domain. See this http URL to download and browse the dataset.
Added
2026-09-24

COCO-Stuff: Thing and Stuff Classes in Context
Holger Caesar, Jasper Uijlings, Vittorio Ferrari
Why you should read this
Presents COCO-Stuff, which enriches the 164,000-image COCO benchmark with dense pixel-level annotations for 91 background categories to enable joint contextual modeling and full-image semantic segmentation of amorphous regions and discrete objects.
Semantic classes can be either things (objects with a well-defined shape, e.g. car, person) or stuff (amorphous background regions, e.g. grass, sky). While lots of classification and detection works focus on thing classes, less attention has been given to stuff classes. Nonetheless, stuff classes are important as they allow to explain important aspects of an image, including (1) scene type; (2) which thing classes are likely to be present and their location (through contextual reasoning); (3) physical attributes, material types and geometric properties of the scene. To understand stuff and things in context we introduce COCO-Stuff, which augments all 164K images of the COCO 2017 dataset with pixel-wise annotations for 91 stuff classes. We introduce an efficient stuff annotation protocol based on superpixels, which leverages the original thing annotations. We quantify the speed versus quality trade-off of our protocol and explore the relation between annotation time and boundary complexity. Furthermore, we use COCO-Stuff to analyze: (a) the importance of stuff and thing classes in terms of their surface cover and how frequently they are mentioned in image captions; (b) the spatial relations between stuff and things, highlighting the rich contextual relations that make our dataset unique; (c) the performance of a modern semantic segmentation method on stuff and thing classes, and whether stuff is easier to segment than things.
Added
2026-09-24

DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving
Chenyi Chen, Ari Seff, Alain Kornhauser, Jianxiong Xiao
Why you should read this
Proposes a direct perception framework for autonomous driving that bridges the gap between full scene parsing and end-to-end control by learning compact driving affordances that enable simple controllers to drive across diverse simulated and real-world environments.
Today, there are two major paradigms for vision-based autonomous driving systems: mediated perception approaches that parse an entire scene to make a driving decision, and behavior reflex approaches that directly map an input image to a driving action by a regressor. In this paper, we propose a third paradigm: a direct perception approach to estimate the affordance for driving. We propose to map an input image to a small number of key perception indicators that directly relate to the affordance of a road/traffic state for driving. Our representation provides a set of compact yet complete descriptions of the scene to enable a simple controller to drive autonomously. Falling in between the two extremes of mediated perception and behavior reflex, we argue that our direct perception representation provides the right level of abstraction. To demonstrate this, we train a deep Convolutional Neural Network using recording from 12 hours of human driving in a video game and show that our model can work well to drive a car in a very diverse set of virtual environments. We also train a model for car distance estimation on the KITTI dataset. Results show that our direct perception approach can generalize well to real driving images. Source code and data are available on our project website.
Added
2026-09-18

Segmenter: Transformer for Semantic Segmentation
Robin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia Schmid
Why you should read this
Introduces Segmenter, an end-to-end Vision Transformer architecture for semantic segmentation that captures global context throughout the network using a mask transformer decoder, surpassing convolutional baselines on the ADE20K and Pascal Context benchmarks.
Image segmentation is often ambiguous at the level of individual image patches and requires contextual information to reach label consensus. In this paper we introduce Segmenter, a transformer model for semantic segmentation. In contrast to convolution-based methods, our approach allows to model global context already at the first layer and throughout the network. We build on the recent Vision Transformer (ViT) and extend it to semantic segmentation. To do so, we rely on the output embeddings corresponding to image patches and obtain class labels from these embeddings with a point-wise linear decoder or a mask transformer decoder. We leverage models pre-trained for image classification and show that we can fine-tune them on moderate sized datasets available for semantic segmentation. The linear decoder allows to obtain excellent results already, but the performance can be further improved by a mask transformer generating class masks. We conduct an extensive ablation study to show the impact of the different parameters, in particular the performance is better for large models and small patch sizes. Segmenter attains excellent results for semantic segmentation. It outperforms the state of the art on both ADE20K and Pascal Context datasets and is competitive on Cityscapes.
Added
2026-09-17

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
Qingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang, Niki Trigoni, Andrew Markham
Why you should read this
Presents RandLA-Net, a lightweight neural network that couples random point sampling with local feature aggregation to perform semantic segmentation on million-point 3D point clouds up to 200 times faster than existing methods.
We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Extensive experiments show that our RandLA-Net can process 1 million points in a single pass with up to 200X faster than existing approaches. Moreover, our RandLA-Net clearly surpasses state-of-the-art approaches for semantic segmentation on two large-scale benchmarks Semantic3D and SemanticKITTI.
Added
2026-09-16

SUN RGB-D: A RGB-D scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, Jianxiong Xiao
Why you should read this
Introduces a large-scale RGB-D dataset of over 10,000 images paired with dense 2D and 3D bounding annotations to benchmark core scene understanding tasks, including 3D object detection, semantic segmentation, and room layout estimation.
Scene understanding is one of the most important and challenging tasks in computer vision. Although remarkable progress has been made in the past few decades, the performance of general-purpose scene understanding is still far from satisfactory. The recent arrival of affordable depth sensors in consumer markets enables us to obtain reliable depth maps at a very low cost, greatly simplifying some common challenges in computer vision and enabling breakthroughs for several tasks, such as body pose estimation [5], intrinsic image [1], and 3D modeling [2]. RGB-D sensors have also enabled rapid progress for scene understanding. However, although we can download color images from the Internet easily, it is hard to obtain RGB-D data online. Therefore the existing RGB-D recognition benchmarks, such as NYU Depth v2 [4], are an order-of-magnitude smaller than modern recognition datasets (e.g. PASCAL VOC) for color images. Although these small datasets successfully bootstrapped initial research and enabled early progress in the past few years, the size limit is now the common bottleneck in advancing research to the next level. It causes easy overfitting of the algorithm during evaluation, and it cannot support learning for data-hungry algorithms that achieve state-of-the-art performance in color-based recognition tasks. Furthermore, although the RGB-D images these datasets provide contain depth maps, the annotation and evaluation metrics are mostly in the 2D image domain but not directly in 3D, as shown in Figure 1. However, scene understanding will be much more useful in the real 3D space. Hence we advocate that the community should use the depth map to reasoning about scenes and evaluating algorithms in 3D. We introduce SUN RGB-D, a dataset containing 10,335 RGB-D images with dense annotations in both 2D and 3D, for both objects and rooms. The goal of our dataset construction is to obtain an image dataset captured by various RGB-D sensors ( Intel RealSense, Asus Xtion LIVE PRO, Microsoft Kinect versions 1 and 2 ) at a similar scale as the PASCAL VOC object detection benchmark. We capture 3,784 images using Kinect v2 and 1,159 images using Intel RealSense. We included the 1,449 images from the NYU Depth V2 captured by Kinect v1 and We also choosed 554 realistic scene images from the Berkeley B3DO Dataset captured by Kinect v1, manually selected 3,389 distinguished frames without significant motion blur from the SUN3D videos captured by Asus Xtion. In total, we obtain 10,335 RGB-D images. To improve the depth map quality, we take short videos and use our SIFT+ICP algorithm to do the refinement. For each image, we annotate the objects with both 2D polygons and 3D bounding boxes and the room layout with 3D polygons. For the 10,335 RGB-D images, we have 146,617 2D polygons and 64,595 3D bounding boxes (with accurate orientations for objects) annotated. Therefore, there are 14.2 objects in each image on average. In total, there are 47 scene categories and 800 object categories. To evaluate all major tasks that work towards total scene understanding, we focus on six important recognition tasks towards total scene understanding that integrates objects, room layout and scene class, includes scene categorization, semantic segmentation, object detection, object orientation, room layout estimation, as well as a final total scene understanding task that integrates everything. The final task for our scene understanding benchmark is to estimate the whole scene including objects and room layout in 3D [3]. This task is also referred to “Basic Level Scene Understanding” [6]. We propose this benchmark task as the final goal to integrate both object detection and room layout estimation to obtain a total scene understanding, recognizing and localizing all the objects and the room structure. We choose state-of-the-art algorithms to evaluate each task. For the tasks without existing algorithm or implementation, we adapt popular algorithms from other tasks. For each task, whenever possible, we try to evaluate algorithms using color, depth, as well as RGB-D images to study the relative importance of color and depth, and gauge to what extent the information from both is complementary. Various evaluation results show that we can apply standard techniques that invented design for color (e.g. hand craft features, deep learning features, detector, sift flow label transfer) to depth domain and it can achieve comparable performance for various tasks.In most of cases, when we combining these two source of information the performance get improved. By constructing a PASCAL-scale dataset using various sensors and defining a benchmark for all major scene understanding tasks with 3D evaluation metrics, we hope to lay the foundation for advancing RGB-D scene understanding in the coming years.
Added
2026-09-16

Playing for Data: Ground Truth from Computer Games
Stephan R. Richter, Vibhav Vineet, Stefan Roth, Vladlen Koltun
Why you should read this
Demonstrates how intercepting communication between commercial video games and graphics hardware enables automatic generation of pixel-accurate semantic annotations, cutting the real-world manual training data needed for semantic segmentation by two-thirds.
Recent progress in computer vision has been driven by high-capacity models trained on large datasets. Unfortunately, creating large datasets with pixel-level labels has been extremely costly due to the amount of human effort required. In this paper, we present an approach to rapidly creating pixel-accurate semantic label maps for images extracted from modern computer games. Although the source code and the internal operation of commercial games are inaccessible, we show that associations between image patches can be reconstructed from the communication between the game and the graphics hardware. This enables rapid propagation of semantic labels within and across images synthesized by the game, with no access to the source code or the content. We validate the presented approach by producing dense pixel-level semantic annotations for 25 thousand images synthesized by a photorealistic open-world computer game. Experiments on semantic segmentation datasets show that using the acquired data to supplement real-world images significantly increases accuracy and that the acquired data enables reducing the amount of hand-labeled real-world data: models trained with game data and just 1/3 of the CamVid training set outperform models trained on the complete CamVid training set.
Added
2026-09-16

ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation
Adam Paszke, Abhishek Chaurasia, Sangpil Kim, Eugenio Culurciello
Why you should read this
Proposes ENet, a lightweight deep neural network architecture that enables real-time semantic segmentation on resource-constrained embedded devices by reducing computational cost and parameter size by over an order of magnitude without sacrificing accuracy.
The ability to perform pixel-wise semantic segmentation in real-time is of paramount importance in mobile applications. Recent deep neural networks aimed at this task have the disadvantage of requiring a large number of floating point operations and have long run-times that hinder their usability. In this paper, we propose a novel deep neural network architecture named ENet (efficient neural network), created specifically for tasks requiring low latency operation. ENet is up to 18 faster, requires 75 less FLOPs, has 79 less parameters, and provides similar or better accuracy to existing models. We have tested it on CamVid, Cityscapes and SUN datasets and report on comparisons with existing state-of-the-art methods, and the trade-offs between accuracy and processing time of a network. We present performance measurements of the proposed architecture on embedded systems and suggest possible software improvements that could make ENet even faster.
Added
2026-09-14

Semantic Understanding of Scenes Through the ADE20K Dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, Antonio Torralba
Why you should read this
Presents ADE20K, an open-vocabulary dataset of 25,000 densely annotated images spanning over 3,000 object and part classes, establishing standardized benchmarks and baselines for fine-grained scene parsing and instance segmentation.
Semantic understanding of visual scenes is one of the holy grails of computer vision. Despite efforts of the community in data collection, there are still few image datasets covering a wide range of scenes and object categories with pixel-wise annotations for scene understanding. In this work, we present a densely annotated dataset ADE20K, which spans diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. Totally there are 25k images of the complex everyday scenes containing a variety of objects in their natural spatial context. On average there are 19.5 instances and 10.5 object classes per image. Based on ADE20K, we construct benchmarks for scene parsing and instance segmentation. We provide baseline performances on both of the benchmarks and re-implement the state-of-the-art models for open source. We further evaluate the effect of synchronized batch normalization and find that a reasonably large batch size is crucial for the semantic segmentation performance. We show that the networks trained on ADE20K are able to segment a wide variety of scenes and objects1.
Added
2026-09-14

Unified Perceptual Parsing for Scene Understanding
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, Jian Sun
Why you should read this
Develops UPerNet, a multi-task framework that learns from heterogeneous annotations to simultaneously parse visual concepts across scenes, objects, parts, and textures for comprehensive image understanding.
Humans recognize the visual world at multiple levels: we effortlessly categorize scenes and detect objects inside, while also identifying the textures and surfaces of the objects along with their different compositional parts. In this paper, we study a new task called Unified Perceptual Parsing, which requires the machine vision systems to recognize as many visual concepts as possible from a given image. A multi-task framework called UPerNet and a training strategy are developed to learn from heterogeneous image annotations. We benchmark our framework on Unified Perceptual Parsing and show that it is able to effectively segment a wide range of concepts from images. The trained networks are further applied to discover visual knowledge in natural scenes. Models are available at \url{this https URL}.
Added
2026-09-14

CCNet: Criss-Cross Attention for Semantic Segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, Humphrey Shi, Wenyu Liu
Why you should read this
Introduces CCNet, a criss-cross attention network that captures global visual context for semantic segmentation while reducing GPU memory usage by eleven times and computation by 85 percent compared to standard non-local blocks.
Contextual information is vital in visual understanding problems, such as semantic segmentation and object detection. We propose a Criss-Cross Network (CCNet) for obtaining full-image contextual information in a very effective and efficient way. Concretely, for each pixel, a novel criss-cross attention module harvests the contextual information of all the pixels on its criss-cross path. By taking a further recurrent operation, each pixel can finally capture the full-image dependencies. Besides, a category consistent loss is proposed to enforce the criss-cross attention module to produce more discriminative features. Overall, CCNet is with the following merits: 1) GPU memory friendly. Compared with the non-local block, the proposed recurrent criss-cross attention module requires 11x less GPU memory usage. 2) High computational efficiency. The recurrent criss-cross attention significantly reduces FLOPs by about 85% of the non-local block. 3) The state-of-the-art performance. We conduct extensive experiments on semantic segmentation benchmarks including Cityscapes, ADE20K, human parsing benchmark LIP, instance segmentation benchmark COCO, video segmentation benchmark CamVid. In particular, our CCNet achieves the mIoU scores of 81.9%, 45.76% and 55.47% on the Cityscapes test set, the ADE20K validation set and the LIP validation set respectively, which are the new state-of-the-art results. The source codes are available at \url{this https URL}.
Added
2026-09-12
