Built independently by an author, for readers. Read the story and support ChapterPal

keyword

feature pyramid network

A feature pyramid network is a deep learning architecture designed to extract and combine multi-scale visual features from an input image to improve computer vision tasks such as object detection and instance segmentation. In standard convolutional neural networks, deeper layers capture rich semantic context at lower spatial resolutions, while earlier layers retain fine spatial details at higher resolutions. A feature pyramid network bridges this gap by augmenting a standard bottom-up convolutional pathway with a top-down pathway linked by lateral skip connections, effectively merging high-level semantic information with high-resolution spatial details. This structure produces a multi-level pyramid of feature maps where every level is both spatially detailed and semantically strong, enabling detection algorithms to accurately identify objects across widely varying scales with minimal additional computational overhead.

8 items

QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

Chenhongyi Yang, Zehao Huang, Naiyan Wang

OrganizationsTuSimpleUniversity of Edinburgh

Why you should read this

Proposes a cascaded sparse query mechanism that dramatically accelerates high-resolution feature pyramid detectors by predicting coarse object locations on low-resolution maps to guide sparse computation only where small objects exist.

While general object detection with deep learning has achieved great success in the past few years, the performance and efficiency of detecting small objects are far from satisfactory. The most common and effective way to promote small object detection is to use high-resolution images or feature maps. However, both approaches induce costly computation since the computational cost grows squarely as the size of images and features increases. To get the best of two worlds, we propose QueryDet that uses a novel query mechanism to accelerate the inference speed of feature-pyramid based object detectors. The pipeline composes two steps: it first predicts the coarse locations of small objects on low-resolution features and then computes the accurate detection results using high-resolution features sparsely guided by those coarse positions. In this way, we can not only harvest the benefit of high-resolution feature maps but also avoid useless computation for the background area. On the popular COCO dataset, the proposed method improves the detection mAP by 1.0 and mAP-small by 2.0, and the high-resolution inference speed is improved to 3.0× on average. On VisDrone dataset, which contains more small objects, we create a new state-of-the-art while gaining a 2.3× high-resolution acceleration on average. Code is available at https://github.com/ChenhongyiYang/QueryDet-PyTorch.

Added

2026-09-26

MMDetection: Open MMLab Detection Toolbox and Benchmark

MMDetection: Open MMLab Detection Toolbox and Benchmark

Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, KC Jimmy, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, Dahua Lin

OrganizationsBeijing Institute of TechnologyHuazhong University of Science and TechnologyMicrosoftNanjing UniversityNanyang Technological UniversityNortheastern UniversityPeking UniversitySenseTimeSun Yat-sen UniversityThe Chinese University of Hong KongThe Hong Kong University of Science and TechnologyUniversity of SydneyZhejiang University

Why you should read this

Introduces MMDetection, an open-source framework and benchmark that unifies numerous object detection and instance segmentation algorithms into a modular codebase with over 200 pre-trained models to facilitate fair comparison and rapid detector development.

We present MMDetection, an object detection toolbox that contains a rich set of object detection and instance segmentation methods as well as related components and modules. The toolbox started from a codebase of MMDet team who won the detection track of COCO Challenge 2018. It gradually evolves into a unified platform that covers many popular detection methods and contemporary modules. It not only includes training and inference codes, but also provides weights for more than 200 network models. We believe this toolbox is by far the most complete detection toolbox. In this paper, we introduce the various features of this toolbox. In addition, we also conduct a benchmarking study on different methods, components, and their hyper-parameters. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to reimplement existing methods and develop their own new detectors. Code and models are available at this https URL. The project is under active development and we will keep this document updated.

Added

2026-09-13

CSPNet: A New Backbone that can Enhance Learning Capability of CNN

CSPNet: A New Backbone that can Enhance Learning Capability of CNN

Chien-Yao Wang, Hong-Yuan Mark Liao, I-Hau Yeh, Yueh-Hua Wu, Ping-Yang Chen, Jun-Wei Hsieh

OrganizationsAcademia SinicaElan Microelectronics CorporationNational Chiao Tung University

Why you should read this

Introduces Cross Stage Partial Network (CSPNet) to eliminate duplicate gradient information across network stages, reducing inference computation by 20% while maintaining or improving accuracy across popular architectures on ImageNet and MS COCO.

Neural networks have enabled state-of-the-art approaches to achieve incredible results on computer vision tasks such as object detection. However, such success greatly relies on costly computation resources, which hinders people with cheap devices from appreciating the advanced technology. In this paper, we propose Cross Stage Partial Network (CSPNet) to mitigate the problem that previous works require heavy inference computations from the network architecture perspective. We attribute the problem to the duplicate gradient information within network optimization. The proposed networks respect the variability of the gradients by integrating feature maps from the beginning and the end of a network stage, which, in our experiments, reduces computations by 20% with equivalent or even superior accuracy on the ImageNet dataset, and significantly outperforms state-of-the-art approaches in terms of AP50 on the MS COCO object detection dataset. The CSPNet is easy to implement and general enough to cope with architectures based on ResNet, ResNeXt, and DenseNet. Source code is at this https URL.

Added

2026-09-13

Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers

Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers

Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H. S. Torr, Li Zhang

OrganizationsFudan UniversityMetaTencentUniversity of OxfordUniversity of Surrey

Why you should read this

Introduces the SEgmentation TRansformer (SETR), replacing traditional convolutional encoder-decoder models with a pure sequence-to-sequence transformer that extracts global visual context across every layer to achieve state-of-the-art accuracy on major benchmarks like ADE20K.

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for segmentation, the latest efforts have been focused on increasing the receptive field, through either dilated/atrous convolutions or inserting attention modules. However, the encoder-decoder based FCN architecture remains unchanged. In this paper, we aim to provide an alternative perspective by treating semantic segmentation as a sequence-to-sequence prediction task. Specifically, we deploy a pure transformer (ie, without convolution and resolution reduction) to encode an image as a sequence of patches. With the global context modeled in every layer of the transformer, this encoder can be combined with a simple decoder to provide a powerful segmentation model, termed SEgmentation TRansformer (SETR). Extensive experiments show that SETR achieves new state of the art on ADE20K (50.28% mIoU), Pascal Context (55.83% mIoU) and competitive results on Cityscapes. Particularly, we achieve the first position in the highly competitive ADE20K test server leaderboard on the day of submission.

Added

2026-09-11

Creative Commons License
Deep High-Resolution Representation Learning for Visual Recognition

Deep High-Resolution Representation Learning for Visual Recognition

Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, Bin Xiao

OrganizationsMicrosoft

Why you should read this

Proposes HRNet, a vision backbone that maintains high-resolution feature maps across parallel multiresolution streams to deliver spatially precise representations for pose estimation, semantic segmentation, and object detection.

High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions \emph{in series} (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams \emph{in parallel}; (ii) Repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at~{\url{this https URL}}.

Added

2026-09-10

EfficientDet: Scalable and Efficient Object Detection

EfficientDet: Scalable and Efficient Object Detection

Mingxing Tan, Ruoming Pang, Quoc V. Le

OrganizationsGoogle

Why you should read this

Introduces EfficientDet, a scalable object detection architecture combining a weighted bi-directional feature pyramid network and compound scaling to achieve state-of-the-art COCO accuracy with up to nine times fewer parameters and drastically lower computational cost than previous models.

Model efficiency has become increasingly important in computer vision. In this paper, we systematically study neural network architecture design choices for object detection and propose several key optimizations to improve efficiency. First, we propose a weighted bi-directional feature pyramid network (BiFPN), which allows easy and fast multiscale feature fusion; Second, we propose a compound scaling method that uniformly scales the resolution, depth, and width for all backbone, feature network, and box/class prediction networks at the same time. Based on these optimizations and better backbones, we have developed a new family of object detectors, called EfficientDet, which consistently achieve much better efficiency than prior art across a wide spectrum of resource constraints. In particular, with single model and single-scale, our EfficientDet-D7 achieves state-of-the-art 55.1 AP on COCO test-dev with 77M parameters and 410B FLOPs, being 4x - 9x smaller and using 13x - 42x fewer FLOPs than previous detectors. Code is available at this https URL.

Added

2026-09-08