Vision Transformers Need Registers

Timothée DarcetMaxime OquabJulien MairalPiotr Bojanowski

article2024ICLR1,002 citationsOutstanding Paper Award

Introduces dedicated register tokens to Vision Transformers to eliminate high-norm feature map artifacts, producing cleaner attention maps and setting a new state of the art on dense visual prediction tasks.

Listen

Modern vision transformers serve as foundational models across computer vision tasks, but they frequently develop internal processing artifacts that degrade local feature quality and model interpretability. While earlier architectures like DINO produced clean, interpretable attention maps that enabled automated object discovery, newer and more capable models such as DINOv2 exhibit severe local anomalies despite achieving high overall accuracy. This issue prevents high-performing models from being reliably used in dense spatial prediction tasks and unsupervised object detection workflows.

The article aims to identify the root cause of these artifacts across supervised, text-supervised, and self-supervised vision transformers and to demonstrate a lightweight architectural solution that completely removes them.

The researchers analyzed the internal activations of large vision models—specifically evaluating models like DINOv2, DeiT-III, and OpenCLIP—across different layer depths, model sizes, and training durations. They applied linear diagnostic models to probe the information stored inside individual tokens, measuring spatial position accuracy, pixel reconstruction error, and global classification ability. To eliminate the identified issue, the authors introduced dedicated placeholder tokens, termed "registers," into the input sequence during pretraining. These registers act as temporary computational storage and are discarded after processing, requiring no changes to the downstream model interface.

The analysis revealed that artifacts consist of high-norm outlier tokens that have output norms roughly 10 times higher than regular tokens, accounting for approximately 2% of the sequence. These outliers emerge midway through deep networks (around layer 15 in 40-layer models) and only appear when training large models (ViT-Large and above) for extended periods. The models selectively repurpose tokens from low-information, uniform background patches to hold global context, which inadvertently discards critical local spatial details and corrupts attention maps. Adding learnable register tokens entirely absorbs this behavior, eliminating outlier patches from image tokens across all tested training regimes. Implementing 4 registers eliminates the artifacts while maintaining or slightly improving downstream classification, semantic segmentation, and depth estimation accuracy. Furthermore, in unsupervised object discovery using the LOST algorithm on VOC 2007, adding registers restored DINOv2's localization performance from 35.3 to 55.4 CorLoc.

These findings demonstrate that vision transformers naturally require dedicated internal scratchpad memory to aggregate global image context during inference. Without dedicated tokens, the network hijacks arbitrary background patches, compromising downstream tasks that rely on smooth, reliable spatial features. The register token mechanism provides this scratchpad explicitly, resolving performance bottlenecks in dense prediction and restoring visual interpretability with negligible computational overhead (adding less than 2% to total training FLOPs when using 4 registers).

Engineering and research teams training large vision transformer backbones should adopt register tokens as a standard architectural default during pretraining. Setting 4 register tokens provides the best trade-off between spatial smoothness, global classification accuracy, and compute overhead. Teams utilizing models for dense spatial tasks or automated object discovery should prioritize models trained with registers to prevent background artifacts from corrupting downstream linear decoders and attention-based algorithms.

Confidence in these findings is high across standard supervised, text-supervised, and self-distillation transformer frameworks. However, the study notes that the exact training dynamics driving why certain architectures trigger artifacts faster than others remain partially unexplained. Additionally, while the register modification significantly improves object discovery performance in models like DINOv2, it does not fully close the gap to original DINO baselines, indicating that further research is required into regularizing how registers specialize during training.

Cover for Vision Transformers Need Registers

Abstract

Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 2.1 Artifacts in the local features of DINOv2
  • 2.2 Hypothesis and remediation
  • 3 Experiments
  • 3.1 Training algorithms and data
  • 3.2 Evaluation of the proposed solution
  • 3.3 Object discovery
  • 3.4 Qualitative evaluation of registers
  • 4 Related Work
  • 5 Conclusion
  • References
  • A Interpolation artifacts and outlier position distribution
  • B Complexity analysis
  • C Analysis of LOST performance
  • D Behavior of models trained with registers
  • D.1 Norms
  • D.2 Information held by tokens
  • D.3 Positional focus
  • E Masked autoencoders
  • F Behavior per attention head
  • G Variance on token information probing
  • H Qualitative Results

Knowls

  1. Knowl 1 — Vision Transformer Register Tokens

    model/method

    To prevent Vision Transformers (ViTs) from corrupting visual patch tokens into internal computation buffers, register tokens are introduced into the sequence.

    Given an image divided into patch embeddings along with a standard class token [CLS][\text{CLS}], a set of NN learnable embedding vectors [REG1],[REG2],…,[REGN][\text{REG}_1], [\text{REG}_2], \dots, [\text{REG}_N] is initialized. These register tokens are concatenated to the sequence directly following the patch embedding layer. The augmented sequence is processed through all self-attention and feed-forward transformer layers. At the output of the final layer, the register tokens are entirely discarded during both pretraining and downstream inference, leaving only the [CLS][\text{CLS}] token and visual patch tokens as output representations.

    By providing explicit non-patch memory slots, the architecture allows the transformer to store, route, and aggregate global context without discarding local patch information. In standard setups, N=4N = 4 registers are used.

  2. Knowl 2 — Identification and Emergence of High-Norm Token Artifacts in Vision Transformers

    empirical result

    Modern vision transformers trained with supervision (DeiT-III), text supervision (OpenCLIP), or self-supervision (DINOv2) exhibit sharp localized artifacts in their feature maps and attention maps, characterized by a small fraction (around 2% to 2.4%) of patch tokens having an L2L_2 norm roughly 10 times higher than regular tokens (e.g., norms exceeding 150 versus typical norms between 0 and 100 in DINOv2 ViT-g/14).

    Empirical analysis of this artifact phenomenon demonstrates that:

    1. Outlier tokens emerge only in sufficiently large vision transformers (observed in ViT-Large, ViT-Huge, and ViT-giant, but absent in ViT-Tiny, ViT-Small, and ViT-Base for DINOv2).
    2. They appear during training only after sufficient duration (emerging after approximately one-third of the training iterations).
    3. Within the network depth, outliers differentiate from normal patch tokens around intermediate layers (e.g., layer 15 in a 40-layer ViT-giant).
    4. Outlier tokens preferentially occur in visually redundant, uniform background image areas where input patch embeddings have high cosine similarity with their four spatial neighbors.
  3. Knowl 3 — Information Divergence in High-Norm Outlier Tokens

    empirical result

    Probing high-norm outlier patch tokens (L2L_2 norm >150> 150) versus normal patch tokens in a 40-layer DINOv2 ViT-g model reveals that outlier tokens discard local patch information to aggregate global image context:

    1. Position Prediction: A linear probe trained on patch embeddings to predict original image position achieves 41.7%41.7\% top-1 accuracy (average distance error 0.790.79) on normal tokens, but only 22.8%22.8\% top-1 accuracy (average distance error 5.095.09) on outlier tokens.
    2. Pixel Reconstruction: A linear probe trained to reconstruct original input patch pixel values achieves an L2L_2 reconstruction error of 18.3818.38 on normal tokens versus 25.2325.23 on outlier tokens.
    3. Global Semantic Classification: A logistic regression classifier trained on single extracted patch tokens across 14 vision benchmarks achieves vastly higher accuracy when trained on outlier tokens than on normal tokens (e.g., on FGVC-Aircraft, outlier tokens yield 79.1%79.1\% top-1 accuracy vs. 17.1%17.1\% on normal tokens; on Stanford Cars, 85.2%85.2\% vs. 10.8%10.8\%).
  4. Knowl 4 — Downstream Task Evaluation of Vision Transformers with Register Tokens

    data/table

    Vision transformer models trained with register tokens (N=4N=4) remove high-norm feature artifacts while preserving or improving performance on linear probing and dense prediction benchmarks.

    Model ImageNet Top-1 (%) ADE20k mIoU NYUd depth rmse ↓\downarrow
    DeiT-III 84.7 38.9 0.511
    DeiT-III + reg 84.7 39.1 0.512
    OpenCLIP 78.2 26.6 0.702
    OpenCLIP + reg 78.1 26.7 0.661
    DINOv2 84.3 46.6 0.378
    DINOv2 + reg 84.8 47.9 0.366

    For OpenCLIP, zero-shot ImageNet top-1 classification accuracy is 59.9%59.9\% without registers and 60.1%60.1\% with registers. Adding registers provides gains in dense downstream tasks such as ADE20k semantic segmentation (+1.3+1.3 mIoU for DINOv2) and NYUd depth estimation (reducing RMSE from 0.3780.378 to 0.3660.366 on DINOv2 and from 0.7020.702 to 0.6610.661 on OpenCLIP) while maintaining image classification capability.

  5. Knowl 5 — Single-Token Linear Probing Classification Across Benchmarks

    data/table

    Linear probing classification performance using a single extracted token representation from a frozen DINOv2-g backbone across 14 image classification datasets highlights that outlier tokens function as global image summaries.

    Token Type IN1k P205 Airc. CF10 CF100 CUB Cal101 Cars DTD Flow. Food Pets SUN VOC
    86.0 66.4 87.3 99.4 94.5 91.3 96.9 91.5 85.2 99.7 94.7 96.9 78.6 89.1
    normal patch 65.8 53.1 17.1 97.1 81.3 18.6 73.2 10.8 63.1 59.5 74.2 47.8 37.7 70.8
    outlier patch 69.0 55.1 79.1 99.3 93.7 84.9 97.6 85.2 84.9 99.6 93.5 94.1 78.5 89.7

    Outlier patch tokens carry significantly higher classification accuracy than normal patch tokens across all evaluated datasets, closely matching the representation performance of the designated [CLS][\text{CLS}] token.

  6. Knowl 6 — Unsupervised Object Discovery Performance with Register Tokens

    data/table

    Evaluating unsupervised object discovery using the Localizing Objects with Self-Supervised Transformers (LOST) algorithm demonstrates that register tokens restore object discovery capability in modern ViT backbones. Performance is reported using the Correct Localization (CorLoc %) metric across PASCAL VOC 2007, PASCAL VOC 2012, and COCO 20k.

    Model VOC 2007 VOC 2012 COCO 20k
    DeiT-III 11.7 13.1 10.7
    DeiT-III + reg 27.1 32.7 25.1
    OpenCLIP 38.8 44.3 31.0
    OpenCLIP + reg 37.1 42.0 27.9
    DINOv2 35.3 40.2 26.9
    DINOv2 + reg 55.4 60.0 42.0

    Adding register tokens to DINOv2 improves CorLoc by +20.1%+20.1\% on VOC 2007, +19.8%+19.8\% on VOC 2012, and +15.1%+15.1\% on COCO 20k. For DeiT-III, registers more than double CorLoc across all datasets.

  7. Knowl 7 — Sensitivity to Register Count and Computational Overhead

    empirical result

    Varying the number of register tokens N∈{0,1,2,4,8,16}N \in \{0, 1, 2, 4, 8, 16\} in DINOv2 ViT-L/14 indicates:

    1. Adding a single register token (N=1N=1) is sufficient to eliminate high-norm outlier artifacts from the attention and feature maps.
    2. Dense downstream prediction benchmarks (ADE20k semantic segmentation and NYUd depth estimation) achieve optimal performance around N=4N = 4 registers.
    3. ImageNet classification accuracy steadily increases with more registers up to N=16N = 16.
    4. Parameter count increase is negligible across all values of NN. The FLOP count increase remains below 2%2\% for N=4N = 4 and reaches at most 6%6\% for N=16N = 16.
  8. Knowl 8 — Absorption of Outlier Behavior into Registers

    empirical result

    When Vision Transformers are trained with register tokens, the distribution of token norms and information content shifts:

    1. Norm Distribution: Visual patch token norms become strictly unimodal without high-norm tails. The high-norm distribution is completely absorbed into the register tokens, with individual registers displaying quantized norm levels.
    2. Information Transfer: On the FGVC-Aircraft linear probing benchmark, a model without registers achieves 15.5%15.5\% accuracy on normal patches and 73.3%73.3\% on outlier patches. With 1 register, normal patches achieve 14.5%14.5\%, outlier patches cease to exist, and the register token achieves 71.1%71.1\% accuracy ([CLS][\text{CLS}] token scores 84.6%84.6\% and 85.2%85.2\% respectively).
    3. Preservation of Patch Information: Linear probing for position prediction on non-outlier patch tokens scores 66.3%66.3\% top-1 accuracy without registers and 65.8%65.8\% with 4 registers; L2L_2 pixel reconstruction error remains essentially identical (15.915.9 vs. 16.016.0).
  9. Knowl 9 — Absence of High-Norm Outliers in Masked Autoencoders

    empirical result

    Masked Autoencoders (MAE) do not produce high-norm outlier tokens or attention map artifacts. This absence is attributed to the MAE pretraining objective, which relies exclusively on a local pixel-reconstruction loss applied directly on patch tokens without requiring sequence-level or global representation aggregation during pretraining.

Coverage note — Qualitative attention visualization figures and the analysis of bicubic interpolation antialiasing during position embedding resizing were summarized in the respective empirical result knowls rather than extracted as standalone entries.

References

  1. 1.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. In ICLR, 2021.
  2. 2.Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. Recurrent memory transformer. In NeurIPS, 2022.
  3. 3.Mikhail S Burtsev, Yuri Kuratov, Anton Peganov, and Grigory V Sapunov. Memory transformer. arXiv preprint arXiv:2006.11527, 2020.
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020.
  5. 5.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  6. 6.Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong. When vision transformers outperform resnets without pre-training or strong data augmentations. In ICLR, 2022.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  8. 8.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  10. 10.Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. You only look at one sequence: Rethinking transformer in vision through object detection. In NeurIPS, 2021.
  11. 11.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  12. 12.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.
  13. 13.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip. 2021.
  14. 14.Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In ICML, 2021.
  15. 15.Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Andrew Brock, Evan Shelhamer, Olivier J. H'enaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, and João Carreira. Perceiver io: A general architecture for structured inputs & outputs. In ICLR, 2022.
  16. 16.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  17. 17.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012.
  18. 18.Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. In NeurIPS, 2020.
  19. 19.David G Lowe. Distinctive image features from scale-invariant keypoints. IJCV, 2004.
  20. 20.Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  21. 21.Bill Psomas, Ioannis Kakogeorgiou, Konstantinos Karantzalos, and Yannis Avrithis. Keep it simpool: Who said supervised transformers suffer from attention deficit? In ICCV, 2023.
  22. 22.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  23. 23.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. IJCV, 2015.
  24. 24.Mark Sandler, Andrey Zhmoginov, Max Vladymyrov, and Andrew Jackson. Fine-tuning image transformers using learnable memory. In CVPR, 2022.
  25. 25.Baifeng Shi, Siyu Gai, Trevor Darrell, and Xin Wang. Toast: Transfer learning via attention steering, 2023.
  26. 26.Oriane Siméoni, Gilles Puy, Huy V Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, and Jean Ponce. Localizing objects with self-supervised transformers and no labels. In BMVC, 2021.
  27. 27.Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, and Ming-Hsuan Yang. Vidt: An efficient and effective fully transformer-based object detector. In ICLR, 2021.
  28. 28.Antonio Torralba and Alexei A. Efros. Unbiased look at dataset bias. In CVPR, 2011.
  29. 29.Hugo Touvron, Matthieu Cord, and Hervé Jégou. Deit iii: Revenge of the vit. In ECCV, 2022.
  30. 30.Xudong Wang, Rohit Girdhar, Stella X Yu, and Ishan Misra. Cut and learn for unsupervised object detection and instance segmentation. In CVPR, 2023.
  31. 31.Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You. Adaptive computation with elastic input sequence. In ICML, 2023.
  32. 32.Yaodong Yu, Tianzhe Chu, Shengbang Tong, Ziyang Wu, Druv Pai, Sam Buchanan, and Yi Ma. Emergence of segmentation with minimalistic white-box transformers. In CPAL, 2024.
  33. 33.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021.
  34. 34.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. In ICLR, 2022.

Citation

MLA
Darcet, T., et al. “Vision Transformers Need Registers”. arXiv, 2023, http://arxiv.org/abs/2309.16588v2.
APA
Darcet, T., Oquab, M., Mairal, J., & Bojanowski, P. (2023). Vision Transformers Need Registers. arXiv. http://arxiv.org/abs/2309.16588v2
Chicago
Darcet, T., M. Oquab, J. Mairal, and P. Bojanowski. 2023. “Vision Transformers Need Registers”. arXiv. http://arxiv.org/abs/2309.16588v2.
Harvard
Darcet, T. et al. (2023) “Vision Transformers Need Registers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.16588v2.
Vancouver
1. Darcet T, Oquab M, Mairal J, Bojanowski P (2023) Vision Transformers Need Registers. arXiv

BibTeX

@article{darcet2023vision,
  title = {Vision Transformers Need Registers},
  author = {Darcet, Timothée and Oquab, Maxime and Mairal, Julien and Bojanowski, Piotr},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.16588v2},
  eprint = {2309.16588}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors