Built independently by an author, for readers. Read the story and support ChapterPal

keyword

YOLOv11

YOLOv11, also known as YOLO11, is a real-time deep learning model architecture developed by Ultralytics for computer vision tasks within the You Only Look Once family. It is designed to perform various visual perception tasks, including object detection, instance segmentation, image classification, pose estimation, and oriented bounding box detection. Built upon earlier iterations of the YOLO framework, YOLOv11 incorporates architectural refinements such as enhanced convolutional blocks and spatial attention mechanisms to improve feature extraction and achieve higher accuracy with reduced parameter counts and computational overhead. Available in multiple scalable model sizes ranging from lightweight configurations for edge devices to larger models for enterprise systems, it provides a versatile solution for real-time image and video processing.

1 item

YOLOv11: An Overview of the Key Architectural Enhancements

YOLOv11: An Overview of the Key Architectural Enhancements

Rahima Khanam, Muhammad Hussain

OrganizationsUniversity of Huddersfield

Why you should read this

Presents an architectural breakdown of YOLOv11, detailing how components like C3k2 blocks and C2PSA attention mechanisms optimize feature extraction and speed-accuracy trade-offs across detection, segmentation, and pose estimation tasks.

This study presents an architectural analysis of YOLOv11, the latest iteration in the YOLO (You Only Look Once) series of object detection models. We examine the models architectural innovations, including the introduction of the C3k2 (Cross Stage Partial with kernel size 2) block, SPPF (Spatial Pyramid Pooling - Fast), and C2PSA (Convolutional block with Parallel Spatial Attention) components, which contribute in improving the models performance in several ways such as enhanced feature extraction. The paper explores YOLOv11's expanded capabilities across various computer vision tasks, including object detection, instance segmentation, pose estimation, and oriented object detection (OBB). We review the model's performance improvements in terms of mean Average Precision (mAP) and computational efficiency compared to its predecessors, with a focus on the trade-off between parameter count and accuracy. Additionally, the study discusses YOLOv11's versatility across different model sizes, from nano to extra-large, catering to diverse application needs from edge devices to high-performance computing environments. Our research provides insights into YOLOv11's position within the broader landscape of object detection and its potential impact on real-time computer vision applications.

Added

2026-09-24