Vision Transformers Need Registers
Timothée DarcetMaxime OquabJulien MairalPiotr Bojanowski
Introduces dedicated register tokens to Vision Transformers to eliminate high-norm feature map artifacts, producing cleaner attention maps and setting a new state of the art on dense visual prediction tasks.
Modern vision transformers serve as foundational models across computer vision tasks, but they frequently develop internal processing artifacts that degrade local feature quality and model interpretability. While earlier architectures like DINO produced clean, interpretable attention maps that enabled automated object discovery, newer and more capable models such as DINOv2 exhibit severe local anomalies despite achieving high overall accuracy. This issue prevents high-performing models from being reliably used in dense spatial prediction tasks and unsupervised object detection workflows.
The article aims to identify the root cause of these artifacts across supervised, text-supervised, and self-supervised vision transformers and to demonstrate a lightweight architectural solution that completely removes them.
The researchers analyzed the internal activations of large vision models—specifically evaluating models like DINOv2, DeiT-III, and OpenCLIP—across different layer depths, model sizes, and training durations. They applied linear diagnostic models to probe the information stored inside individual tokens, measuring spatial position accuracy, pixel reconstruction error, and global classification ability. To eliminate the identified issue, the authors introduced dedicated placeholder tokens, termed "registers," into the input sequence during pretraining. These registers act as temporary computational storage and are discarded after processing, requiring no changes to the downstream model interface.
The analysis revealed that artifacts consist of high-norm outlier tokens that have output norms roughly 10 times higher than regular tokens, accounting for approximately 2% of the sequence. These outliers emerge midway through deep networks (around layer 15 in 40-layer models) and only appear when training large models (ViT-Large and above) for extended periods. The models selectively repurpose tokens from low-information, uniform background patches to hold global context, which inadvertently discards critical local spatial details and corrupts attention maps. Adding learnable register tokens entirely absorbs this behavior, eliminating outlier patches from image tokens across all tested training regimes. Implementing 4 registers eliminates the artifacts while maintaining or slightly improving downstream classification, semantic segmentation, and depth estimation accuracy. Furthermore, in unsupervised object discovery using the LOST algorithm on VOC 2007, adding registers restored DINOv2's localization performance from 35.3 to 55.4 CorLoc.
These findings demonstrate that vision transformers naturally require dedicated internal scratchpad memory to aggregate global image context during inference. Without dedicated tokens, the network hijacks arbitrary background patches, compromising downstream tasks that rely on smooth, reliable spatial features. The register token mechanism provides this scratchpad explicitly, resolving performance bottlenecks in dense prediction and restoring visual interpretability with negligible computational overhead (adding less than 2% to total training FLOPs when using 4 registers).
Engineering and research teams training large vision transformer backbones should adopt register tokens as a standard architectural default during pretraining. Setting 4 register tokens provides the best trade-off between spatial smoothness, global classification accuracy, and compute overhead. Teams utilizing models for dense spatial tasks or automated object discovery should prioritize models trained with registers to prevent background artifacts from corrupting downstream linear decoders and attention-based algorithms.
Confidence in these findings is high across standard supervised, text-supervised, and self-distillation transformer frameworks. However, the study notes that the exact training dynamics driving why certain architectures trigger artifacts faster than others remain partially unexplained. Additionally, while the register modification significantly improves object discovery performance in models like DINOv2, it does not fully close the gap to original DINO baselines, indicating that further research is required into regularizing how registers specialize during training.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the foundational Vision Transformer (ViT) architecture whose feature map artifacts and sequence representation mechanisms are analyzed and remedied by the source paper.
- Paper: Emerging Properties in Self-Supervised Vision Transformers, Mathilde Caron et al. (2021). Demonstrates the emergence of interpretable spatial attention maps and dense visual representations in self-supervised ViTs (DINO), establishing the baseline properties and downstream dense prediction tasks directly evaluated in the source.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Introduces dedicated distillation tokens in ViT architectures, establishing early precedents for adding specialized auxiliary tokens to the input sequence.
- Paper: Going deeper with Image Transformers, Hugo Touvron et al. (2021). Explores the internal token interactions and stabilization techniques in deep Vision Transformers, providing key context for handling token-level bottlenecks.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). Analyzes the internal representations and global information propagation across ViT layers, offering crucial analytical background on how ViTs differ from CNNs.
- Paper: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, Sihyun Yu et al. (2025). Leverages rich representations from modern self-supervised Vision Transformers to regularize and dramatically accelerate the training of diffusion transformers.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). Proposes alternative architectural designs to eliminate visual artifacts and depth degradation in vision transformers through foveal perception mechanisms.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). Applies large-scale transformer representations directly to downstream dense geometric and 3D visual prediction tasks.
