Built independently by an author, for readers. Read the story and support ChapterPal

keyword

NMS-free training

NMS-free training is an object detection optimization methodology designed to eliminate the need for non-maximum suppression post-processing during inference by teaching the neural network to output exactly one distinct bounding box prediction per object. In traditional object detectors, training relies on one-to-many label assignments where multiple candidate anchors or grid cells are assigned to a single ground-truth object, necessitating non-maximum suppression at test time to filter out redundant overlapping predictions and causing additional inference latency. NMS-free training overcomes this limitation by incorporating one-to-one label matching or dual label assignment strategies, where a one-to-one prediction head is trained alongside a traditional auxiliary branch to maintain rich gradient flow during optimization. As a result, the deployed model directly outputs deduplicated predictions, enabling true end-to-end execution, lower latency, and deterministic inference on edge and real-time computing platforms without sacrificing detection accuracy.

1 item

YOLOv11: An Overview of the Key Architectural Enhancements

YOLOv11: An Overview of the Key Architectural Enhancements

Rahima Khanam, Muhammad Hussain

OrganizationsUniversity of Huddersfield

Why you should read this

Presents an architectural breakdown of YOLOv11, detailing how components like C3k2 blocks and C2PSA attention mechanisms optimize feature extraction and speed-accuracy trade-offs across detection, segmentation, and pose estimation tasks.

This study presents an architectural analysis of YOLOv11, the latest iteration in the YOLO (You Only Look Once) series of object detection models. We examine the models architectural innovations, including the introduction of the C3k2 (Cross Stage Partial with kernel size 2) block, SPPF (Spatial Pyramid Pooling - Fast), and C2PSA (Convolutional block with Parallel Spatial Attention) components, which contribute in improving the models performance in several ways such as enhanced feature extraction. The paper explores YOLOv11's expanded capabilities across various computer vision tasks, including object detection, instance segmentation, pose estimation, and oriented object detection (OBB). We review the model's performance improvements in terms of mean Average Precision (mAP) and computational efficiency compared to its predecessors, with a focus on the trade-off between parameter count and accuracy. Additionally, the study discusses YOLOv11's versatility across different model sizes, from nano to extra-large, catering to diverse application needs from edge devices to high-performance computing environments. Our research provides insights into YOLOv11's position within the broader landscape of object detection and its potential impact on real-time computer vision applications.

Added

2026-09-24