keyword
spatial pyramid
A spatial pyramid is a multi-scale representation technique in computer vision that divides an image or feature map into hierarchical grids of varying spatial resolutions to capture information at multiple levels of detail. By segmenting visual data into progressively finer sub-regions or applying pooling operations across different receptive field sizes, a spatial pyramid preserves spatial geometry while summarizing local feature distributions. This structure enables neural networks and pattern recognition systems to aggregate contextual cues across diverse scales, improve scale invariance, and produce fixed-length representations from visual inputs of arbitrary dimensions.
2 items

YOLOv11: An Overview of the Key Architectural Enhancements
Rahima Khanam, Muhammad Hussain
Why you should read this
Presents an architectural breakdown of YOLOv11, detailing how components like C3k2 blocks and C2PSA attention mechanisms optimize feature extraction and speed-accuracy trade-offs across detection, segmentation, and pose estimation tasks.
This study presents an architectural analysis of YOLOv11, the latest iteration in the YOLO (You Only Look Once) series of object detection models. We examine the models architectural innovations, including the introduction of the C3k2 (Cross Stage Partial with kernel size 2) block, SPPF (Spatial Pyramid Pooling - Fast), and C2PSA (Convolutional block with Parallel Spatial Attention) components, which contribute in improving the models performance in several ways such as enhanced feature extraction. The paper explores YOLOv11's expanded capabilities across various computer vision tasks, including object detection, instance segmentation, pose estimation, and oriented object detection (OBB). We review the model's performance improvements in terms of mean Average Precision (mAP) and computational efficiency compared to its predecessors, with a focus on the trade-off between parameter count and accuracy. Additionally, the study discusses YOLOv11's versatility across different model sizes, from nano to extra-large, catering to diverse application needs from edge devices to high-performance computing environments. Our research provides insights into YOLOv11's position within the broader landscape of object detection and its potential impact on real-time computer vision applications.
Added
2026-09-24

SSD: Single Shot MultiBox Detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg
Why you should read this
Introduces a fast single-stage object detector that predicts bounding boxes and class scores across multi-scale feature maps, matching two-stage detection accuracy at real-time speeds without generating region proposals.
We present a method for detecting objects in images using a single deep neural network. Our approach, named SSD, discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per feature map location. At prediction time, the network generates scores for the presence of each object category in each default box and produces adjustments to the box to better match the object shape. Additionally, the network combines predictions from multiple feature maps with different resolutions to naturally handle objects of various sizes. Our SSD model is simple relative to methods that require object proposals because it completely eliminates proposal generation and subsequent pixel or feature resampling stage and encapsulates all computation in a single network. This makes SSD easy to train and straightforward to integrate into systems that require a detection component. Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified framework for both training and inference. Compared to other single stage methods, SSD has much better accuracy, even with a smaller input image size. For input, SSD achieves 72.1% mAP on VOC2007 test at 58 FPS on a Nvidia Titan X and for input, SSD achieves 75.1% mAP, outperforming a comparable state of the art Faster R-CNN model. Code is available at this https URL .
Added
2026-09-06
