Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Retentive Networks

A Retentive Network is a deep learning architecture designed for sequential and structural data processing as an efficient alternative to the standard Transformer. It replaces the traditional self-attention mechanism with a retention mechanism that applies an explicit decay factor to encode positional distance and contextual dependencies. The architecture uniquely supports three equivalent computational representations: a parallel formulation for fast model training, a recurrent formulation that enables constant-time inference and linear memory complexity per step, and a chunkwise recurrent formulation for efficient long-sequence processing. By combining the high training parallelism and modeling capacity of attention-based models with the low memory footprint and inference efficiency of recurrent models, Retentive Networks achieve linear computational complexity while maintaining strong performance across diverse tasks.

1 item

RMT: Retentive Networks Meet Vision Transformers

RMT: Retentive Networks Meet Vision Transformers

Qihang Fan, Huaibo Huang, Mingrui Chen, Hongmin Liu, Ran He

Why you should read this

Proposes a vision backbone that adapts RetNet's decay mechanism into a 2D Manhattan distance-based spatial prior and decomposes self-attention to achieve linear complexity while outperforming standard Vision Transformers across image classification, object detection, and semantic segmentation.

Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and bears a quadratic computational complexity, thereby constraining the applicability of ViT. To alleviate these issues, we draw inspiration from the recent Retentive Network (RetNet) in the field of NLP, and propose RMT, a strong vision backbone with explicit spatial prior for general purposes. Specifically, we extend the RetNet’s temporal decay mechanism to the spatial domain, and propose a spatial decay matrix based on the Manhattan distance to introduce the explicit spatial prior to Self-Attention. Additionally, an attention decomposition form that adeptly adapts to explicit spatial prior is proposed, aiming to reduce the computational burden of modeling global information without disrupting the spatial decay matrix. Based on the spatial decay matrix and the attention decomposition form, we can flexibly integrate explicit spatial prior into the vision backbone with linear complexity. Extensive experiments demonstrate that RMT exhibits exceptional performance across various vision tasks. Specifically, without extra training data, RMT achieves 84.8% and 86.1% top-1 acc on ImageNet-1K with 27M/4.5GFLOPs and 96M/18.2GFLOPs. For downstream tasks, RMT achieves 54.5 box AP and 47.2 mask AP on the COCO detection task, and 52.8 mIoU on the ADE20K semantic segmentation task.

Added

2026-10-05