Built independently by an author, for readers. Read the story and support ChapterPal

keyword

anchor-free detection

Anchor-free detection is an approach in computer vision object detection that identifies and locates objects directly without relying on predefined bounding box templates known as anchor boxes. Traditional anchor-based detectors place multiple candidate boxes of fixed scales and aspect ratios across an image and predict offsets from them, whereas anchor-free detectors predict object presence and geometric boundaries directly from spatial locations, center points, or keypoints within feature representations. By regressing dimensions such as the distances from a point to the object edges or by grouping keypoint pairs, anchor-free methods eliminate the need for manually tuned anchor hyperparameters, reduce computational complexity, and mitigate sample imbalance issues while effectively generalizing across varied object shapes and sizes.

4 items

TOOD: Task-aligned One-stage Object Detection

TOOD: Task-aligned One-stage Object Detection

Chengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott, Weilin Huang

OrganizationsAlibaba GroupByteDanceIntellifusion Inc.Malong LLCMeituan

Why you should read this

Proposes a task-aligned one-stage object detection framework that resolves spatial misalignment between classification and localization features via an interactive prediction head and alignment-based sample assignment, achieving 51.1 AP on MS-COCO with fewer parameters and FLOPs than existing detectors.

One-stage object detection is commonly implemented by optimizing two sub-tasks: object classification and localization, using heads with two parallel branches, which might lead to a certain level of spatial misalignment in predictions between the two tasks. In this work, we propose a Task-aligned One-stage Object Detection (TOOD) that explicitly aligns the two tasks in a learning-based manner. First, we design a novel Task-aligned Head (T-Head) which offers a better balance between learning task-interactive and task-specific features, as well as a greater flexibility to learn the alignment via a task-aligned predictor. Second, we propose Task Alignment Learning (TAL) to explicitly pull closer (or even unify) the optimal anchors for the two tasks during training via a designed sample assignment scheme and a task-aligned loss. Extensive experiments are conducted on MS-COCO, where TOOD achieves a 51.1 AP at single-model single-scale testing. This surpasses the recent one-stage detectors by a large margin, such as ATSS (47.7 AP), GFL (48.2 AP), and PAA (49.0 AP), with fewer parameters and FLOPs. Qualitative results also demonstrate the effectiveness of TOOD for better aligning the tasks of object classification and localization. Code is available at this https URL.

Added

2026-09-25

YOLOv11: An Overview of the Key Architectural Enhancements

YOLOv11: An Overview of the Key Architectural Enhancements

Rahima Khanam, Muhammad Hussain

OrganizationsUniversity of Huddersfield

Why you should read this

Presents an architectural breakdown of YOLOv11, detailing how components like C3k2 blocks and C2PSA attention mechanisms optimize feature extraction and speed-accuracy trade-offs across detection, segmentation, and pose estimation tasks.

This study presents an architectural analysis of YOLOv11, the latest iteration in the YOLO (You Only Look Once) series of object detection models. We examine the models architectural innovations, including the introduction of the C3k2 (Cross Stage Partial with kernel size 2) block, SPPF (Spatial Pyramid Pooling - Fast), and C2PSA (Convolutional block with Parallel Spatial Attention) components, which contribute in improving the models performance in several ways such as enhanced feature extraction. The paper explores YOLOv11's expanded capabilities across various computer vision tasks, including object detection, instance segmentation, pose estimation, and oriented object detection (OBB). We review the model's performance improvements in terms of mean Average Precision (mAP) and computational efficiency compared to its predecessors, with a focus on the trade-off between parameter count and accuracy. Additionally, the study discusses YOLOv11's versatility across different model sizes, from nano to extra-large, catering to diverse application needs from edge devices to high-performance computing environments. Our research provides insights into YOLOv11's position within the broader landscape of object detection and its potential impact on real-time computer vision applications.

Added

2026-09-24

Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection

Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection

Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, Stan Z. Li

OrganizationsAerospace Information Research Institute, Chinese Academy of SciencesBeijing University of Posts and TelecommunicationsInstitute of Automation, Chinese Academy of SciencesUniversity of Chinese Academy of SciencesWestlake University

Why you should read this

Reveals that positive and negative sample selection is the essential difference between anchor-based and anchor-free object detectors, introducing an Adaptive Training Sample Selection (ATSS) strategy that unifies both paradigms and boosts detection accuracy without added overhead.

Object detection has been dominated by anchor-based detectors for several years. Recently, anchor-free detectors have become popular due to the proposal of FPN and Focal Loss. In this paper, we first point out that the essential difference between anchor-based and anchor-free detection is actually how to define positive and negative training samples, which leads to the performance gap between them. If they adopt the same definition of positive and negative samples during training, there is no obvious difference in the final performance, no matter regressing from a box or a point. This shows that how to select positive and negative training samples is important for current object detectors. Then, we propose an Adaptive Training Sample Selection (ATSS) to automatically select positive and negative samples according to statistical characteristics of object. It significantly improves the performance of anchor-based and anchor-free detectors and bridges the gap between them. Finally, we discuss the necessity of tiling multiple anchors per location on the image to detect objects. Extensive experiments conducted on MS COCO support our aforementioned analysis and conclusions. With the newly introduced ATSS, we improve state-of-the-art detectors by a large margin to 50.7%50.7\% AP without introducing any overhead. The code is available at this https URL

Added

2026-09-16