YOLOv11: An Overview of the Key Architectural Enhancements

Rahima KhanamMuhammad Hussain

article2024arXiv3,605 citations

Presents an architectural breakdown of YOLOv11, detailing how components like C3k2 blocks and C2PSA attention mechanisms optimize feature extraction and speed-accuracy trade-offs across detection, segmentation, and pose estimation tasks.

Listen

Organizations increasingly rely on real-time computer vision to power automated systems across sectors such as manufacturing, autonomous transportation, retail, and healthcare. Achieving high detection accuracy while maintaining low latency and minimal computational overhead remains a primary operational challenge, particularly when deploying models to edge devices and resource-constrained hardware.

The article evaluates the architectural innovations, operational efficiency, and multi-task capabilities of YOLOv11, the latest iteration in the You Only Look Once real-time vision series. It analyzes how these structural design changes enhance performance and expand practical use cases compared to prior models.

The authors conducted an architectural analysis and performance benchmark review of YOLOv11 against predecessor models ranging from YOLOv5 to YOLOv10. The evaluation examined model variants spanning from nano to extra-large, assessing core metrics including mean Average Precision on the standard COCO dataset, inference speed in milliseconds, and total parameter counts across multiple vision tasks.

The key findings demonstrate clear performance gains across the model family. First, YOLOv11 establishes a superior performance frontier on standard benchmarks, with the extra-large model reaching approximately 54.5% mean Average Precision at 13 milliseconds of latency. Second, the architecture delivers notable parameter efficiency; the medium variant matches or exceeds previous accuracy levels while utilizing 22% fewer parameters than its YOLOv8 equivalent. Third, the small variant delivers high accuracy of approximately 47% within a ultra-low latency window of 2 to 6 milliseconds, operating at speeds previously limited to less capable models. Finally, the framework broadens multi-task versatility by supporting native object detection, instance segmentation, pose estimation, image classification, and oriented bounding box detection within a unified framework.

These findings indicate that organizations can deploy higher-accuracy vision models at reduced infrastructure and hardware costs. The lower computational footprint allows complex visual inspection and tracking tasks to run directly on edge devices without requiring cloud processing pipelines. The inclusion of spatial attention and oriented detection also reduces failure rates when identifying occluded or rotated items in demanding operational settings.

Decision-makers should evaluate YOLOv11 for upcoming computer vision implementations or system upgrades. Teams should select model sizes tailored to their compute constraints, leveraging smaller variants for fast edge inference and larger models for centralized high-precision tasks. Pilot testing in production environments is recommended to benchmark task-specific workflows such as defect identification or tracking before full-scale migration.

Confidence in the reported benchmark metrics is high, though readers should note that reported gains reflect standardized public datasets. Real-world performance may vary based on environmental lighting, domain-specific camera resolutions, and target hardware configurations.

arXiv: 2410.17725
  • Paper: YOLOv12: Attention-Centric Real-Time Object Detectors, Yunjie Tian et al. (2025). This paper introduces YOLOv12, extending the real-time detector lineage by adopting attention-centric architectures that build upon and compare directly against YOLOv11 baselines.
Cover for YOLOv11: An Overview of the Key Architectural Enhancements

Abstract

This study presents an architectural analysis of YOLOv11, the latest iteration in the YOLO (You Only Look Once) series of object detection models. We examine the models architectural innovations, including the introduction of the C3k2 (Cross Stage Partial with kernel size 2) block, SPPF (Spatial Pyramid Pooling - Fast), and C2PSA (Convolutional block with Parallel Spatial Attention) components, which contribute in improving the models performance in several ways such as enhanced feature extraction. The paper explores YOLOv11's expanded capabilities across various computer vision tasks, including object detection, instance segmentation, pose estimation, and oriented object detection (OBB). We review the model's performance improvements in terms of mean Average Precision (mAP) and computational efficiency compared to its predecessors, with a focus on the trade-off between parameter count and accuracy. Additionally, the study discusses YOLOv11's versatility across different model sizes, from nano to extra-large, catering to diverse application needs from edge devices to high-performance computing environments. Our research provides insights into YOLOv11's position within the broader landscape of object detection and its potential impact on real-time computer vision applications.

Table of Contents

  • 1 Introduction
  • 2 Evolution of YOLO models
  • 3 What is YOLOv11?
  • 4 Architectural footprint of Yolov11
  • 4.1 Backbone
  • 4.1.1 Convolutional Layers
  • 4.1.2 SPPF and C2PSA
  • 4.2 Neck
  • 4.2.1 C3k2 Block
  • 4.2.2 Attention Mechanism
  • 4.3 Head
  • 4.3.1 C3k2 Block
  • 4.3.2 CBS Blocks
  • 4.3.3 Final Convolutional Layers and Detect Layer
  • 5 Key Computer Vision Tasks Supported by YOLO11
  • 6 Advancements and Key Features of YOLOv11
  • 7 Discussion
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — YOLOv11 Architectural Overview: Backbone, Neck, and Head

    model/method

    YOLOv11 is an end-to-end real-time object detector that builds upon and refines the YOLOv8 unified neural network design. Its architecture is divided into three primary sub-networks:

    1. Backbone: Serves as the primary multi-scale feature extractor. It receives raw images and passes them through initial convolutional layers to progressively reduce spatial resolution while increasing channel depth. It replaces the C2f blocks used in YOLOv8 with C3k2 (Cross Stage Partial with kernel size 2) blocks. The backbone concludes with a Spatial Pyramid Pooling - Fast (SPPF) block followed immediately by a Cross Stage Partial with Spatial Attention (C2PSA) block.

    2. Neck: Functions as the intermediate multi-scale feature aggregator. It combines feature maps from different stages of the backbone using upsampling and concatenation pathways, integrating C3k2 modules and C2PSA attention to enhance cross-scale spatial representations.

    3. Head: Performs dense bounding box regression and object classification across multiple scale paths. It refines incoming aggregated features via C3k2 blocks and Convolution-BatchNorm-SiLU (CBS) layers, terminating in 2D convolutional layers and a unified Detect layer.

  2. Knowl 2 — C3k2 Block Architecture in YOLOv11

    model/method

    The C3k2 (Cross Stage Partial with kernel size 2) block is a core architectural enhancement in YOLOv11, substituting the C2f blocks of YOLOv8 across the backbone, neck, and prediction head.

    Key design mechanisms of the C3k2 block include:

    • Dual Smaller Convolutions: Instead of executing a single large convolution within the bottleneck, C3k2 divides the computation across two smaller convolutional layers with kernel size k=2k=2. This reduces computational floating-point operations and parameter count while preserving feature extraction capacity.
    • Configurable Bottleneck via c3k Parameter:
      • When c3k=False\text{c3k} = \text{False}, the C3k2 block adopts a standard bottleneck structure analogous to the C2f block.
      • When c3k=True\text{c3k} = \text{True}, the bottleneck is replaced by a C3 module, enabling deeper and more intricate feature extraction.
    • Customizable Kernel Dimensions: The underlying C3k component allows adaptive kernel sizing to adjust the receptive field for specific scale and depth requirements.
  3. Knowl 3 — C2PSA (Cross Stage Partial with Spatial Attention) Module

    model/method

    The C2PSA (Cross Stage Partial with Spatial Attention) block is an attention mechanism integrated into YOLOv11. It is positioned directly after the Spatial Pyramid Pooling - Fast (SPPF) module in the backbone and utilized within the neck pathways.

    The C2PSA block enhances feature maps by:

    • Performing spatial pooling across feature channels to identify and emphasize critical spatial regions within an image.
    • Dynamically modulating multi-scale feature maps so the network concentrates on relevant objects while attenuating background regions.
    • Improving localization and classification performance for difficult cases, such as small objects or partially occluded targets, distinguishing YOLOv11 from YOLOv8 (which lacks this spatial attention module).
  4. Knowl 4 — Empirical Benchmark: Accuracy and Latency of YOLOv11 vs. Prior YOLO Generations

    empirical result

    Benchmarked on the MS COCO dataset using TensorRT 10 FP16 on an NVIDIA T4 GPU, YOLOv11 establishes an improved accuracy-versus-latency Pareto frontier across all model scales compared to YOLOv5, YOLOv6-3.0, YOLOv7, YOLOv8, YOLOv9, YOLOv10, and PP-YOLOE+:

    • Efficiency at Medium Scale: YOLOv11m exceeds the mAP50−95\text{mAP}_{50-95} accuracy of YOLOv8m while requiring 22%22\% fewer parameters.
    • High-Accuracy Regime: YOLOv11x achieves approximately 54.5%54.5\% mAP50−95\text{mAP}_{50-95} at 13 ms13\text{ ms} latency, attaining higher accuracy than all previous YOLO variants at that latency.
    • Low-Latency Regime (2–6 ms2\text{--}6\text{ ms}): YOLOv11s achieves approximately 47%47\% mAP50−95\text{mAP}_{50-95} at an inference latency between 2 ms2\text{ ms} and 3 ms3\text{ ms}, matching the accuracy of larger models from earlier generations at speeds previously associated only with lower-accuracy compact models.
    • Nano Scale: The YOLOv11n variant achieves faster inference speeds and higher frames per second (FPS) alongside improved mAP compared to YOLOv8n.
  5. Knowl 5 — YOLOv11 Vision Tasks and Model Variants

    model/method

    YOLOv11 supports six computer vision tasks across five model scales: Nano (nano / n), Small (small / s), Medium (medium / m), Large (large / l), and Extra-Large (xlarge / x). All variants support inference, validation, training, and export:

    • Object Detection (YOLOv11): Standard 2D bounding box regression and classification.
    • Instance Segmentation (YOLOv11-seg): Pixel-level instance mask delineation.
    • Pose Estimation (YOLOv11-pose): Keypoint detection for human body posture and joint tracking.
    • Oriented Object Detection (YOLOv11-obb): Rotated bounding box prediction with angular orientation for aerial and angled imagery.
    • Image Classification (YOLOv11-cls): Whole-image category classification.
    • Object Tracking: Temporal multi-frame identity tracking built on top of object detection.
    Model Variants Inference Validation Training Export
    YOLOv11 nano, small, medium, large, xlarge ✓ ✓ ✓ ✓
    YOLOv11-seg nano-seg, small-seg, medium-seg, large-seg, xlarge-seg ✓ ✓ ✓ ✓
    YOLOv11-pose nano-pose, small-pose, medium-pose, large-pose, xlarge-pose ✓ ✓ ✓ ✓
    YOLOv11-obb nano-obb, small-obb, medium-obb, large-obb, xlarge-obb ✓ ✓ ✓ ✓
    YOLOv11-cls nano-cls, small-cls, medium-cls, large-cls, xlarge-cls ✓ ✓ ✓ ✓
  6. Knowl 6 — YOLOv11 Prediction Head and CBS Refinement Layers

    model/method

    In YOLOv11, the prediction head processes multi-scale feature maps passed from the neck to produce localized bounding boxes and class predictions through a structured layer sequence:

    1. C3k2 Processing: Multiple C3k2 modules refine features across multi-scale pathways at distinct depths.
    2. CBS Layers: Convolution-BatchNorm-SiLU layers follow the C3k2 blocks to stabilize and normalize intermediate representations:
      • 2D Convolution (Conv2D): Extracts scale-specific localized features.
      • Batch Normalization (BatchNorm): Normalizes feature distributions across mini-batches to ensure stable gradient flow.
      • SiLU Activation: Applies the Sigmoid Linear Unit non-linearity, SiLU(x)=x⋅σ(x)=x1+e−x\text{SiLU}(x) = x \cdot \sigma(x) = \frac{x}{1 + e^{-x}}.
    3. Detect Layer: The final branch applies Conv2D layers to map features to the required output tensor dimensions, yielding bounding box coordinates, objectness confidence scores, and categorical class probabilities.
  7. Knowl 7 — Historical Progression of YOLO Architectures from YOLOv1 to YOLOv10

    data/table

    The architectural progression of the YOLO detector family prior to YOLOv11 encompasses key innovations across model tasks and backends from 2015 to 2024:

    Release Year Tasks Key Contributions Framework
    YOLO 2015 Object Detection, Basic Classification Single-stage object detector Darknet
    YOLOv2 2016 Object Detection, Improved Classification Multi-scale training, dimension clustering Darknet
    YOLOv3 2018 Object Detection, Multi-scale Detection SPP block, Darknet-53 backbone Darknet
    YOLOv4 2020 Object Detection, Basic Object Tracking Mish activation, CSPDarknet-53 backbone Darknet
    YOLOv5 2020 Object Detection, Basic Instance Segmentation Anchor-free detection, SWISH activation, PANet PyTorch
    YOLOv6 2022 Object Detection, Instance Segmentation Self-attention, anchor-free object detection PyTorch
    YOLOv7 2022 Object Detection, Object Tracking, Instance Segmentation Transformers, E-ELAN reparameterisation PyTorch
    YOLOv8 2023 Object Detection, Instance/Panoptic Segmentation, Keypoints GANs, anchor-free detection PyTorch
    YOLOv9 2024 Object Detection, Instance Segmentation PGI and GELAN PyTorch
    YOLOv10 2024 Object Detection Consistent dual assignments for NMS-free training PyTorch

Coverage note — General introductory descriptions of computer vision applications, standard deep learning background, and generic industrial use-case listings were omitted as non-contributed background.

References

  1. 1.Milan Sonka, Vaclav Hlavac, and Roger Boyle. Image processing, analysis and machine vision. Springer, 2013.
  2. 2.Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye. Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3):257–276, 2023.
  3. 3.Zhong-Qiu Zhao, Peng Zheng, Shou-tao Xu, and Xindong Wu. Object detection with deep learning: A review. IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019.
  4. 4.Muhammad Hussain and Rahima Khanam. In-depth review of yolov1 to yolov10 variants for enhanced photovoltaic defect detection. In Solar, volume 4, pages 351–386. MDPI, 2024.
  5. 5.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016.
  6. 6.Juan Du. Understanding of object detection based on cnn family and yolo. In Journal of Physics: Conference Series, volume 1004, page 012029. IOP Publishing, 2018.
  7. 7.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7263–7271, 2017.
  8. 8.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
  9. 9.Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
  10. 10.Roboflow Blog Jacob Solawetz. What is yolov5? a guide for beginners., 2020. Accessed: 21 October 2024.
  11. 11.Chuyi Li, Lulu Li, Hongliang Jiang, Kaiheng Weng, Yifei Geng, Liang Li, Zaidan Ke, Qingyuan Li, Meng Cheng, Weiqiang Nie, et al. Yolov6: A single-stage object detection framework for industrial applications. arXiv preprint arXiv:2209.02976, 2022.
  12. 12.Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7464–7475, 2023.
  13. 13.Francesco Jacob Solawetz. What is yolov8? the ultimate guide, 2023. Accessed: 21 October 2024.
  14. 14.Chien-Yao Wang, I-Hau Yeh, and Hong-Yuan Mark Liao. Yolov9: Learning what you want to learn using programmable gradient information. arXiv preprint arXiv:2402.13616, 2024.
  15. 15.Ao Wang, Hui Chen, Lihao Liu, Kai Chen, Zijia Lin, Jungong Han, and Guiguang Ding. Yolov10: Real-time end-to-end object detection. arXiv preprint arXiv:2405.14458, 2024.
  16. 16.Glenn Jocher and Jing Qiu. Ultralytics yolo11, 2024.
  17. 17.Rahima Khanam, Muhammad Hussain, Richard Hill, and Paul Allen. A comprehensive review of convolutional neural networks for defect detection in industrial applications. IEEE Access, 2024.
  18. 18.Satya Mallick. Yolo - learnopencv. https://learnopencv.com/yolo11/, 2024. Accessed: 2024-10-21.
  19. 19.Jingwen Feng, Qiaofeng An, Jiahao Zhang, Shuxun Zhou, Guangwei Du, and Kai Yang. Application of yolov7-tiny in the detection of steel surface defects. In 2024 5th International Seminar on Artificial Intelligence, Networking and Information Technology (AINIT), pages 2241–2245. IEEE, 2024.
  20. 20.Ultralytics. Instance segmentation and tracking, 2024. Accessed: 2024-10-21.
  21. 21.Ultralytics Abirami Vina. Ultralytics yolo11 has arrived: Redefine what’s possible in ai, 2024. Accessed: 2024-10-21.
  22. 22.Viso.AI Gaudenz Boesch. Yolov11: A new iteration of “you only look once. https://viso.ai/computer-vision/yolov11/, 2024. Accessed: 2024-10-21.
  23. 23.Ultralytics. Ultralytics yolov11. https://docs.ultralytics.com/models/yolo11/s, 2024. Accessed: 21-Oct-2024.
  24. 24.Rahima Khanam and Muhammad Hussain. What is yolov5: A deep look into the internal features of the popular object detector. arXiv preprint arXiv:2407.20892, 2024.
  25. 25.DigitalOcean. What’s new in yolov11 transforming object detection once again part 1, 2024. Accessed: 2024-10-21.
  26. 26.Muhammad Hussain and Hussain Al-Aqrabi. Child emotion recognition via custom lightweight cnn architecture. In Kids Cybersecurity Using Computational Intelligence Techniques, pages 165–174. Springer, 2023.
  27. 27.Burcu Ataer Aydin, Muhammad Hussain, Richard Hill, and Hussain Al-Aqrabi. Domain modelling for a lightweight convolutional network focused on automated exudate detection in retinal fundus images. In 2023 9th International Conference on Information Technology Trends (ITT), pages 145–150. IEEE, 2023.

Citation

MLA
Khanam, R., and M. Hussain. “YOLOv11: An Overview of the Key Architectural Enhancements”. arXiv, 2024, http://arxiv.org/abs/2410.17725v1.
APA
Khanam, R., & Hussain, M. (2024). YOLOv11: An Overview of the Key Architectural Enhancements. arXiv. http://arxiv.org/abs/2410.17725v1
Chicago
Khanam, R., and M. Hussain. 2024. “YOLOv11: An Overview of the Key Architectural Enhancements”. arXiv. http://arxiv.org/abs/2410.17725v1.
Harvard
Khanam, R. and Hussain, M. (2024) “YOLOv11: An Overview of the Key Architectural Enhancements”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.17725v1.
Vancouver
1. Khanam R, Hussain M (2024) YOLOv11: An Overview of the Key Architectural Enhancements. arXiv

BibTeX

@article{khanam2024yolov11,
  title = {YOLOv11: An Overview of the Key Architectural Enhancements},
  author = {Khanam, Rahima and Hussain, Muhammad},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.17725v1},
  eprint = {2410.17725}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/