Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RetinaNet architecture

RetinaNet architecture is a single-stage deep learning framework designed for dense object detection in computer vision. It typically combines a convolutional backbone network, such as a residual network, with a Feature Pyramid Network to extract multi-scale feature representations across different object sizes. Connected to this pyramid are two dedicated subnetworks: a classification subnetwork that predicts the category of candidate bounding boxes and a box regression subnetwork that refines their spatial coordinates. To overcome the severe class imbalance between background regions and foreground objects during training, the architecture employs focal loss, a specialized loss function that dynamically down-weights easily classified background examples. This structural and training design allows RetinaNet to maintain the processing speed of single-stage detectors while achieving accuracy comparable to slower two-stage detection systems.

1 item

Stand-Alone Self-Attention in Vision Models

Stand-Alone Self-Attention in Vision Models

Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, Jonathon Shlens

OrganizationsGoogle

Why you should read this

Demonstrates that replacing spatial convolutions entirely with stand-alone self-attention in ResNet models achieves superior ImageNet accuracy and competitive COCO detection performance with significantly fewer parameters and FLOPs.

Convolutions are a fundamental building block of modern computer vision systems. Recent approaches have argued for going beyond convolutions in order to capture long-range dependencies. These efforts focus on augmenting convolutional models with content-based interactions, such as self-attention and non-local means, to achieve gains on a number of vision tasks. The natural question that arises is whether attention can be a stand-alone primitive for vision models instead of serving as just an augmentation on top of convolutions. In developing and testing a pure self-attention vision model, we verify that self-attention can indeed be an effective stand-alone layer. A simple procedure of replacing all instances of spatial convolutions with a form of self-attention applied to ResNet model produces a fully self-attentional model that outperforms the baseline on ImageNet classification with 12% fewer FLOPS and 29% fewer parameters. On COCO object detection, a pure self-attention model matches the mAP of a baseline RetinaNet while having 39% fewer FLOPS and 34% fewer parameters. Detailed ablation studies demonstrate that self-attention is especially impactful when used in later layers. These results establish that stand-alone self-attention is an important addition to the vision practitioner's toolbox.

Added

2026-09-25