keyword
Retentive Networks
A Retentive Network is a deep learning architecture designed for sequential and structural data processing as an efficient alternative to the standard Transformer. It replaces the traditional self-attention mechanism with a retention mechanism that applies an explicit decay factor to encode positional distance and contextual dependencies. The architecture uniquely supports three equivalent computational representations: a parallel formulation for fast model training, a recurrent formulation that enables constant-time inference and linear memory complexity per step, and a chunkwise recurrent formulation for efficient long-sequence processing. By combining the high training parallelism and modeling capacity of attention-based models with the low memory footprint and inference efficiency of recurrent models, Retentive Networks achieve linear computational complexity while maintaining strong performance across diverse tasks.
1 item

