DINOv2: Learning Robust Visual Features without Supervision

Maxime OquabTimothée DarcetThéo MoutakanniHuy VoMarc SzafraniecVasil KhalidovPierre FernandezDaniel HazizaFrancisco MassaAlaaeldin El-Nouby

article2023Trans. Mach. Learn. Res.10,484 citationsOutstanding paper finalist, TMLR 2024

Demonstrates that scaling self-supervised vision transformers to one billion parameters with automated data filtering produces universal visual representations that surpass leading text-supervised models across both image- and pixel-level tasks without fine-tuning.

Listen

Foundation models pretrained without human annotation have revolutionized natural language processing, but computer vision has largely depended either on text-guided supervision—which struggles with fine, pixel-level details—or self-supervised training on small, uncurated datasets that fail to generalize. The article investigates whether self-supervised learning alone, when scaled across both model architecture and vast quantities of curated data, can produce robust, general-purpose visual representations that work out of the box across diverse image and pixel-level tasks.

To demonstrate this, the authors built an automated, text-free data curation pipeline that filtered and rebalanced a massive web crawl into a 142-million-image pretraining dataset named LVD-142M. They trained a family of Vision Transformers (DINOv2) scaling up to a 1.1-billion-parameter architecture using a discriminative self-supervised objective that combines image- and patch-level losses, enhanced with computational optimizations such as memory-efficient attention, sequence packing, and distributed model sharding. Smaller model variants were created through knowledge distillation directly from the largest model, and the entire suite was evaluated across diverse tasks including image classification, instance retrieval, semantic segmentation, and depth estimation without fine-tuning the underlying feature extractors.

Key findings show that DINOv2 substantially surpasses prior self-supervised models and matches or outperforms leading weakly supervised, text-guided models such as OpenCLIP. On ImageNet-1k, frozen DINOv2 features achieved 86.5% top-1 accuracy with a simple linear classifier, marking a 4.2% improvement over previous self-supervised state-of-the-art models. In instance retrieval benchmarks such as Oxford-Hard, the model outperformed previous self-supervised approaches by 41% mean average precision and weakly supervised models by 34%. For dense pixel-level tasks such as semantic segmentation and monocular depth estimation, linear probes applied to frozen features delivered results competitive with fully fine-tuned specialized architectures. Furthermore, knowledge distillation proved superior to training smaller architectures from scratch across all evaluated benchmarks, and the training optimizations reduced computational memory requirements threefold while doubling execution speed.

These findings indicate that task-specific fine-tuning is no longer essential to achieve state-of-the-art visual performance, significantly lowering deployment complexity, engineering timelines, and inference overhead. The emergence of granular spatial understanding and object-part correspondences without supervision suggests that pretraining solely on visual data provides a stronger geometric foundation than text-aligned alternatives. Consequently, organizations can leverage frozen DINOv2 backbones to power a broad array of downstream visual applications through lightweight task heads.

Decision-makers adopting these models should consider integrating them into downstream pipelines while accounting for identified operational boundaries. While fairness evaluations showed minimal harmful associations across demographic attributes, the model retains geographical and socioeconomic performance disparities, exhibiting lower accuracy on imagery from low-income households and non-Western regions such as Africa. Stakeholders should conduct localized validation and bias testing before deploying these models in critical socio-technical applications.

arXiv: 2304.07193
Cover for DINOv2: Learning Robust Visual Features without Supervision

Abstract

The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any system by producing all-purpose visual features, i.e., features that work across image distributions and tasks without finetuning. This work shows that existing pretraining methods, especially self-supervised methods, can produce such features if trained on enough curated data from diverse sources. We revisit existing approaches and combine different techniques to scale our pretraining in terms of data and model size. Most of the technical contributions aim at accelerating and stabilizing the training at scale. In terms of data, we propose an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature. In terms of models, we train a ViT model (Dosovitskiy et al., 2020) with 1B parameters and distill it into a series of smaller models that surpass the best available all-purpose features, OpenCLIP (Ilharco et al., 2021) on most of the benchmarks at image and pixel levels.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Data Processing
  • 4 Discriminative Self-supervised Pre-training
  • 5 Efficient implementation
  • 6 Ablation Studies
  • 6.1 Improved Training Recipe
  • 6.2 Pretraining Data Source
  • 6.3 Model Size and Data
  • 6.4 Loss Components
  • 6.5 Impact of Knowledge Distillation
  • 6.6 Impact of Resolution
  • 7 Results
  • 7.1 ImageNet Classification
  • 7.2 Additional Image and Video classification Benchmarks
  • 7.3 Instance Recognition
  • 7.4 Dense Recognition Tasks
  • 7.5 Qualitative Results
  • 8 Fairness and Bias Analysis
  • 8.1 Geographical Fairness
  • 8.2 Gender, Skintones and Age
  • 9 Estimating the Environmental Impact of Training our Models
  • 10 Future work and Discussion
  • References
  • A Data Processing
  • A.1 Data selection
  • A.2 Image similarity
  • A.3 Deduplication
  • A.4 Retrieval
  • B Implementation Details
  • B.1 Unsupervised pre-training
  • B.2 High-Resolution adaptation
  • B.3 Linear probing evaluation
  • C List of Datasets used

Knowls

  1. Knowl 1 — DINOv2 Discriminative Self-Supervised Pre-training Framework

    model/method

    DINOv2 trains Vision Transformer (ViT) backbones using a discriminative self-supervised learning framework that combines image-level and patch-level objectives in a student-teacher setup:

    1. Image-Level Objective (LDINO\mathcal{L}_{\text{DINO}}): Computes cross-entropy between the output probability distributions of the student and teacher class tokens ([CLS][\text{CLS}]) extracted from different augmented views (crops) of the same input image: LDINO=−∑k=1Kpt(k)log⁡ps(k)\mathcal{L}_{\text{DINO}} = - \sum_{k=1}^K p_t(k) \log p_s(k) where ps=softmax(gs([CLS]s)/τs)p_s = \text{softmax}(g_s([\text{CLS}]_s)/\tau_s) is the student distribution produced by an MLP projection head gsg_s with temperature parameter aus au_s, and ptp_t is the teacher distribution over K=131,072K = 131,072 prototype vectors. The teacher distribution is normalized using 3 iterations of the Sinkhorn-Knopp algorithm.

    2. Patch-Level Objective (LiBOT\mathcal{L}_{\text{iBOT}}): A masked image modeling objective applied to masked patch tokens: LiBOT=−∑i∈M∑k=1Kpt,i(k)log⁡ps,i(k)\mathcal{L}_{\text{iBOT}} = - \sum_{i \in \mathcal{M}} \sum_{k=1}^K p_{t,i}(k) \log p_{s,i}(k) where M\mathcal{M} is the subset of patch indices masked in the student input (the teacher processes the entire unmasked image), and ps,i,pt,ip_{s,i}, p_{t,i} are normalized probability distributions over prototype scores produced by dedicated iBOT MLP projection heads.

    3. Untied Projection Heads: Unlike earlier iBOT implementations that share MLP parameters across image and patch objectives, DINOv2 uses separate, unshared projection heads for the class token (DINO) and patch tokens (iBOT), which is essential for training stability and performance when scaling to billion-parameter models.

    4. Exponential Moving Average Teacher: The teacher weights θt\theta_t are updated as an exponential moving average of student weights θs\theta_s: θt←λθt+(1−λ)θs\theta_t \leftarrow \lambda \theta_t + (1 - \lambda) \theta_s where the momentum coefficient λ\lambda follows a cosine schedule from 0.9940.994 to 1.01.0 across training.

    5. Total Training Loss: Ltotal=LDINO+LiBOT+λkoleoLKoLeo\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{DINO}} + \mathcal{L}_{\text{iBOT}} + \lambda_{\text{koleo}} \mathcal{L}_{\text{KoLeo}} where λkoleo=0.1\lambda_{\text{koleo}} = 0.1 weights the KoLeo regularization term.

  2. Knowl 2 — Automatic Self-Supervised Data Curation Pipeline for LVD-142M

    algorithm

    DINOv2 constructs the curated pre-training dataset LVD-142M (142,109,386 images) from an unfiltered pool of 1.2 billion crawled web images without requiring text annotations or metadata.

    Input: Raw crawled web image pool UrawU_{\text{raw}}, set of curated source datasets DcuratedD_{\text{curated}}
    Output: Curated pre-training dataset LVDLVD
    U←FilterUnsafeAndBlurFaces(Uraw)U \leftarrow \text{FilterUnsafeAndBlurFaces}(U_{\text{raw}})
    Ecopy←ExtractCopyDetectionEmbeddings(U)E_{\text{copy}} \leftarrow \text{ExtractCopyDetectionEmbeddings}(U)
    // Self-Deduplication
    G←BuildKNNGraph(Ecopy,k=64)G \leftarrow \text{BuildKNNGraph}(E_{\text{copy}}, k=64)
    C←FindConnectedComponents(G,similarity threshold=0.60)C \leftarrow \text{FindConnectedComponents}(G, \text{similarity threshold}=0.60)
    Uself_dedup←RetainOnePerComponent(U,C)U_{\text{self\_dedup}} \leftarrow \text{RetainOnePerComponent}(U, C)
    // Relative Deduplication against downstream test/validation sets
    for each image s∈Uself_dedups \in U_{\text{self\_dedup}} do
        if max cosine similarity of ss to any evaluation split >0.45> 0.45 then
            Discard ss
    Uclean←Remaining pool of 744M imagesU_{\text{clean}} \leftarrow \text{Remaining pool of } 744\text{M images}
    // Feature Extraction and Retrieval
    Eretrieval←ExtractViTH16Embeddings(Uclean)E_{\text{retrieval}} \leftarrow \text{ExtractViTH16Embeddings}(U_{\text{clean}})
    LVD←∅LVD \leftarrow \emptyset
    for each source dataset Dq∈DcuratedD_q \in D_{\text{curated}} do
        if ∣Dq∣≥106|D_q| \ge 10^6 (e.g., ImageNet-22k, ImageNet-1k, Google Landmarks v2) then
            for each image q∈Dqq \in D_q do
                Nq←RetrieveNearestNeighbors(q,Eretrieval,N=4 or k=32)N_q \leftarrow \text{RetrieveNearestNeighbors}(q, E_{\text{retrieval}}, N=4 \text{ or } k=32)
                LVD←LVD∪{q}∪NqLVD \leftarrow LVD \cup \{q\} \cup N_q
        else
            // Cluster-based retrieval for smaller datasets
            \mathcal{C} \leftarrow \text{DistributedKMeans}(E_{\text{retrieval}}, K=100000)
            for each cluster c∈Cc \in \mathcal{C} do
                if CountQueryMatches(c,Dqc, D_q) >3> 3 then
                    Sc←SampleImages(c,limit=10000)S_c \leftarrow \text{SampleImages}(c, \text{limit}=10000)
                    LVD←LVD∪ScLVD \leftarrow LVD \cup S_c (capped at 10610^6 retrieved images per dataset)
    return LVDLVD
  3. Knowl 3 — Kozachenko-Leonenko (KoLeo) Regularization for Feature Spreading

    equation

    The KoLeo regularizer is derived from the Kozachenko-Leonenko differential entropy estimator. It encourages the visual representations produced within a batch to span uniformly across the unit hypersphere, preventing representation collapse and improving nearest-neighbor instance retrieval.

    For a batch of nn feature vectors (x1,…,xn)∈Rd(\mathbf{x}_1, \dots, \mathbf{x}_n) \in \mathbb{R}^d that are ℓ2\ell_2-normalized such that ∥xi∥2=1\|\mathbf{x}_i\|_2 = 1 for all i∈{1,…,n}i \in \{1, \dots, n\}, the loss LKoLeo\mathcal{L}_{\text{KoLeo}} is defined as:

    LKoLeo=−1n∑i=1nlog⁡(dn,i)\mathcal{L}_{\text{KoLeo}} = - \frac{1}{n} \sum_{i=1}^n \log(d_{n,i})

    where dn,id_{n,i} is the minimum Euclidean distance between vector xi\mathbf{x}_i and every other feature vector in the current batch:

    dn,i=min⁡j≠i∥xi−xj∥2d_{n,i} = \min_{j \neq i} \|\mathbf{x}_i - \mathbf{x}_j\|_2

    In DINOv2, LKoLeo\mathcal{L}_{\text{KoLeo}} is computed on the student class tokens ([CLS][\text{CLS}]) of the first global crops with a loss weight of λkoleo=0.1\lambda_{\text{koleo}} = 0.1 independently across the samples assigned to each GPU worker without inter-GPU communication.

  4. Knowl 4 — High-Throughput and Memory-Efficient Training Techniques in DINOv2

    model/method

    DINOv2 introduces four structural and systems-level optimizations enabling stable, high-throughput self-supervised pre-training of billion-parameter models at scale, running ≈2×\approx 2\times faster with 1/31/3 the GPU memory of earlier iBOT implementations:

    1. Sequence Packing for Multi-Crop Inputs: DINO pre-training forwards global crops (resolution 224×224224 \times 224) and local crops (resolution 98×9898 \times 98) with unequal token sequence lengths. Rather than executing multiple separate forward and backward passes per crop size, all sequences are concatenated into a single long sequence and forwarded in one operation. A block-diagonal attention mask is applied to prevent cross-attention between tokens originating from different image crops.

    2. Efficient Stochastic Depth: Standard stochastic depth computes all residual branch operations and zeroes out dropped paths. DINOv2 uses a fused kernel that randomly permutes the batch indices and slices the first (1−d)×B(1 - d) \times B elements along the batch dimension (with drop rate d=0.4d = 0.4), bypassing both compute and memory overhead for the dropped 40%40\% of tokens.

    3. Hardware-Aligned Vision Transformer Architectures: GPU Tensor Cores attain peak compute throughput when head embedding dimensions are multiples of 64 and overall model embedding dimensions are multiples of 256. DINOv2 modifies the ViT-g backbone from 1408 embedding dimension with 16 heads (88 dim/head) to an embedding dimension of 1536 with 24 heads (64 dim/head), 40 transformer blocks, and SwiGLU feed-forward networks (1.1B total parameters).

    4. PyTorch Fully-Sharded Data Parallel (FSDP): Shards the 16 GB model state (student, teacher, and AdamW first and second moments) across GPU nodes. Backbone weights and gradient all-reductions are broadcast in float16 precision (cutting cross-GPU communication by ≈50%\approx 50\% relative to DDP), while MLP head gradients are reduced in float32 to ensure numerical stability.

  5. Knowl 5 — Self-Supervised Knowledge Distillation from ViT-g Teacher to Smaller ViT Backbones

    model/method

    Rather than training small Vision Transformer architectures (ViT-S, ViT-B, ViT-L) from scratch on the full curated dataset, DINOv2 distills knowledge directly from the largest pre-trained checkpoint (a frozen ViT-g/14 with 1.1B parameters).

    The self-supervised distillation procedure adapts the pre-training loop with the following modifications:

    • The teacher network is replaced with the fixed, frozen ViT-g/14 teacher.
    • Masking and stochastic depth are disabled (d=0d = 0) in the student network.
    • The patch-level LiBOT\mathcal{L}_{\text{iBOT}} loss is computed directly on the two unmasked global crops between student and frozen ViT-g patch tokens.
    • An exponential moving average (EMA) of the student weights is maintained throughout training and serves as the final evaluated model.
    • Standard MLP feed-forward blocks are retained in distilled models rather than SwiGLU layers.

    Distilling from ViT-g outperforms training identical architectures from scratch across every evaluated vision benchmark. For example, a distilled ViT-L/14 attains 86.3% ImageNet-1k top-1 accuracy (versus 84.5% trained from scratch) and 76.3% on instance retrieval (versus 71.3% from scratch).

  6. Knowl 6 — High-Resolution Pre-training Adaptation for Vision Transformers

    model/method

    Pre-training Vision Transformers exclusively at high spatial resolution is computationally expensive (e.g., training at 416×416416 \times 416 or 518×518518 \times 518 requires ≈3×\approx 3\times more compute than at 224×224224 \times 224). DINOv2 adopts a two-stage pre-training protocol:

    1. Base Pre-training: The network is pre-trained for 625,000 iterations on standard 224×224224 \times 224 images.
    2. High-Resolution Adaptation: The pre-trained weights are loaded, and pre-training is resumed for an additional short phase of 10,000 iterations at 518×518518 \times 518 resolution (or 416×416416 \times 416 in ablation experiments). All optimization and learning rate schedules from base training are compressed into this 10k-iteration window, with the base learning rate scaled down.

    This short adaptation phase matches the downstream semantic segmentation (ADE-20k) and classification performance of full-length high-resolution pre-training while consuming only a small fraction of the total computational budget.

  7. Knowl 7 — ImageNet Classification and Robustness Benchmarks for Frozen DINOv2 Features

    empirical result

    Evaluating frozen visual representations on ImageNet-1k and out-of-distribution robustness benchmarks via a single linear classification probe at 224×224224 \times 224 resolution demonstrates that DINOv2 substantially surpasses prior self-supervised methods and performs competitively with weakly-supervised models trained on billions of text-image pairs.

    Method Backbone Pre-train Data Text Sup. INet-1k (val) ImageNet-ReaL ImageNet-V2
    Weakly-supervised
    CLIP ViT-L/14 WIT-400M ✓ 84.3 88.1 75.3
    SWAG ViT-H/14 IG3.6B ✓ 85.7 88.7 77.6
    OpenCLIP ViT-H/14 LAION-2B ✓ 84.4 88.4 75.5
    OpenCLIP ViT-G/14 LAION-2B ✓ 86.2 89.4 77.2
    EVA-CLIP ViT-g/14 custom mixture ✓ 86.4 89.3 77.4
    Self-supervised
    MAE ViT-H/14 ImageNet-1k ×\times 76.6 83.3 64.8
    DINO ViT-B/8 ImageNet-1k ×\times 79.2 85.5 68.2
    MSN ViT-L/7 ImageNet-1k ×\times 80.7 86.0 69.7
    EsViT Swin-B/14 ImageNet-1k ×\times 81.3 87.0 70.4
    Mugs ViT-L/16 ImageNet-1k ×\times 82.1 86.9 70.8
    iBOT ViT-L/16 ImageNet-22k ×\times 82.3 87.5 72.4
    DINOv2 (Self-supervised)
    DINOv2-S ViT-S/14 LVD-142M ×\times 81.1 86.6 70.9
    DINOv2-B ViT-B/14 LVD-142M ×\times 84.5 88.3 75.1
    DINOv2-L ViT-L/14 LVD-142M ×\times 86.3 89.5 78.0
    DINOv2-g ViT-g/14 LVD-142M ×\times 86.5 89.6 78.4

    On out-of-distribution robustness benchmarks evaluated with linear probing on frozen features at resolution 224:

    • ImageNet-A: DINOv2-g achieves 75.9% accuracy, exceeding OpenCLIP-G (63.8%) and iBOT ViT-L/16 (41.5%).
    • ImageNet-R: DINOv2-g achieves 78.8% accuracy (iBOT ViT-L/16: 51.0%, OpenCLIP-G: 87.8%).
    • ImageNet-C (mean corruption error, lower is better): DINOv2-g achieves 28.2 (iBOT ViT-L/16: 43.9, OpenCLIP-G: 45.3).
    • ImageNet-Sketch: DINOv2-g achieves 62.5% accuracy (iBOT ViT-L/16: 38.5%, OpenCLIP-G: 66.4%).

    When supervised fine-tuning is applied end-to-end to DINOv2-g, performance increases modestly from 86.5% to 88.5% at resolution 224 (+2.0%+2.0\%) and from 86.7% to 88.9% at resolution 448 (+2.2%+2.2\%), indicating that DINOv2 features already operate near peak capacity out of the box without task-specific fine-tuning.

  8. Knowl 8 — Dense Visual Recognition Performance of DINOv2 on Semantic Segmentation and Monocular Depth Estimation

    empirical result

    Frozen patch tokens from DINOv2 yield high accuracy on pixel-level dense vision tasks without fine-tuning backbone weights:

    ADE20k (mIoU) Cityscapes (mIoU) Pascal VOC (mIoU)
    Method Arch. lin. +ms lin. +ms lin. +ms
    OpenCLIP ViT-G/14 39.3 46.0 60.3 70.3 71.4 79.2
    MAE ViT-H/14 33.3 30.7 58.4 61.0 67.6 63.3
    DINO ViT-B/8 31.8 35.2 56.9 66.2 66.4 75.6
    iBOT ViT-L/16 44.6 47.5 64.8 74.5 82.3 84.3
    DINOv2-S ViT-S/14 44.3 47.2 66.6 77.1 81.1 82.6
    DINOv2-B ViT-B/14 47.3 51.3 69.4 80.0 82.5 84.9
    DINOv2-L ViT-L/14 47.7 53.1 70.3 80.9 82.1 86.0
    DINOv2-g ViT-g/14 49.0 53.0 71.3 81.0 83.0 86.2

    (Evaluation setups: lin. trains a single linear classifier on patch tokens; +ms concatenates patch tokens from the final 4 layers at resolution 640 with multiscale augmentations. Integrating frozen DINOv2-g into Mask2Former with ViT-Adapter achieves 60.2 mIoU on ADE20k).

    On monocular depth estimation evaluated using Root Mean Squared Error (RMSE, lower is better):

    • NYU-Depth V2: DINOv2-g achieves an RMSE of 0.344 with a 1-layer linear probe, 0.298 with a 4-layer linear probe, and 0.279 using a DPT decoder head (compared to OpenCLIP-G at 0.541 / 0.510 / 0.414, and iBOT ViT-L/16 at 0.417 / 0.387 / 0.358).
    • KITTI (Eigen split): DINOv2-g achieves an RMSE of 2.62 (1-layer), 2.35 (4-layer), and 2.11 (DPT), outperforming OpenCLIP-G (3.57 / 3.21 / 2.56).
    • SUN-RGBD (Zero-shot transfer from NYU-Depth V2): DINOv2-g achieves an RMSE of 0.402 (1-layer), 0.362 (4-layer), and 0.338 (DPT), outperforming OpenCLIP-G (0.537 / 0.476 / 0.408).
  9. Knowl 9 — Instance-Level Retrieval, Video Action Recognition, and Fine-Grained Classification

    empirical result

    DINOv2 features transfer across category-level classification, video action recognition, and non-parametric instance retrieval benchmarks:

    Revisiting Oxford Revisiting Paris Met AmsterTime
    Method Arch. Medium Hard Medium Hard GAP ACC mAP
    OpenCLIP ViT-G/14 50.7 19.7 79.2 60.2 6.5 34.4 24.6
    MAE ViT-H/14 11.7 2.2 19.9 4.7 7.5 30.5 4.2
    DINO ViT-B/8 40.1 13.7 65.3 35.3 17.1 43.9 24.6
    iBOT ViT-L/16 39.0 12.7 70.7 47.0 25.1 54.8 26.7
    DINOv2-S ViT-S/14 68.8 43.2 84.6 68.5 29.4 57.7 43.5
    DINOv2-B ViT-B/14 72.9 49.5 90.3 78.6 36.7 66.1 45.6
    DINOv2-L ViT-L/14 75.1 54.0 92.7 83.5 40.0 71.6 50.0
    DINOv2-g ViT-g/14 73.6 52.3 92.1 82.6 36.8 76.5 46.7
    • Fine-Grained Classification: On 12 transfer benchmarks (Food-101, CIFAR-10, CIFAR-100, SUN397, Stanford Cars, FGVC-Aircraft, VOC 2007, DTD, Oxford Pets, Caltech 101, Flowers-102, CUB-200), DINOv2-g achieves an average linear probe accuracy of 92.1%, outperforming OpenCLIP-G (91.9%) and iBOT ViT-L (86.6%). On iNaturalist 2018 and iNaturalist 2021, DINOv2-g reaches 81.6% and 85.7% accuracy, exceeding OpenCLIP-G (73.0% and 76.0%).
    • Video Action Recognition: Linear probing on frozen frame features yields 78.4% top-1 accuracy on Kinetics-400, 91.2% on UCF-101, and 38.3% on Something-Something v2 (SSv2) for DINOv2-g, outperforming OpenCLIP-G (78.3% K400, 90.7% UCF-101, 35.8% SSv2) and iBOT ViT-L (72.6% K400, 88.6% UCF-101, 38.7% SSv2).
  10. Knowl 10 — Pretraining Data Scaling and Loss Component Ablations

    empirical result

    Ablation experiments evaluate the contribution of individual components, training data sources, and loss terms relative to the baseline iBOT method:

    1. Step-by-Step Training Recipe Ablation (ViT-L on ImageNet-22k):

      • Baseline iBOT: 72.9% k-NN / 82.3% linear probe on ImageNet-1k.
      • +LayerScale and Stochastic Depth (d=0.4d = 0.4): 75.4% k-NN / 82.0% linear (necessary for preventing loss divergence and NaN values at scale).
      • +128k Prototypes: 76.6% k-NN / 81.9% linear.
      • +KoLeo Regularizer: 78.9% k-NN / 82.5% linear (+2.3%+2.3\% k-NN improvement).
      • +SwiGLU FFN + Patch size 14 + Teacher momentum 0.994 + Warmup tuning: 80.5% k-NN / 83.8% linear.
      • +Batch size 3,072 + Sinkhorn-Knopp + Untying projection heads (= DINOv2): 82.0% k-NN / 84.5% linear.
    2. Curated vs. Uncurated vs. ImageNet-22k Data (ViT-g/14 trained for identical iterations):

      • Curated LVD-142M achieves 85.8% on INet-1k, 73.9% on ImageNet-A, 47.7 mIoU on ADE-20k, 64.6 mAP on Oxford-M, and 82.3% on iNaturalist 2018.
      • Uncurated 142M crawl samples achieve 83.3% on INet-1k, 59.4% on ImageNet-A, 48.5 mIoU on ADE-20k, 54.3 mAP on Oxford-M, and 68.0% on iNaturalist 2018.
      • ImageNet-22k (14M images) achieves 85.9% on INet-1k, 73.5% on ImageNet-A, 46.6 mIoU on ADE-20k, 62.5 mAP on Oxford-M, and 81.1% on iNaturalist 2018.
      • Curation provides large gains over uncurated web data across tasks, and dataset scale beyond ImageNet-22k improves performance in out-of-domain evaluation tasks.
    3. Loss Term Ablations (ViT-g/14):

      • Ablating the LKoLeo\mathcal{L}_{\text{KoLeo}} loss drops Oxford-M retrieval from 63.9% to 55.6% mAP (−8.3%-8.3\%) and ImageNet-A accuracy from 72.8% to 70.6%.
      • Ablating the patch-level masked image modeling loss (LiBOT\mathcal{L}_{\text{iBOT}}) drops ADE-20k segmentation from 47.1 to 44.2 mIoU (−2.9%-2.9\%).
  11. Knowl 11 — Fairness Assessment, Demographic Biases, and Carbon Footprint of DINOv2

    empirical result

    Systematic evaluation of geographic disparities, demographic label associations, and environmental costs for DINOv2-g reveals the following characteristics:

    1. Geographic and Income Fairness (Dollar Street): On recognizing 94 household concepts across 54 countries and income levels using frozen features:

      • DINOv2-g scores higher across income buckets than SEERv2 (Low: 67.4% vs. 59.7%; Medium: 83.3% vs. 78.5%; High: 90.5% vs. 86.6%).
      • Across continents, DINOv2-g scores 74.0% in Africa, 81.6% in Asia, 86.2% in the Americas, and 89.7% in Europe (SEERv2: 65.9%, 76.3%, 81.1%, 85.6%).
      • A substantial disparity remains: performance in Africa is 25.7% lower than in Europe, and accuracy on low-income households is 31.7% lower than on high-income households.
    2. Demographic Label Association (Casual Conversations dataset): Evaluating a linear classifier trained on 619 ImageNet-22k classes across gender, skin tone, and age groups:

      • 0.0% associations with harmful "Non-Human" labels across all demographic groups.
      • Near-zero "Crime" label associations (0.2% on darker-skinned males, 0.1% on the 30–45 age group, 0.0% elsewhere).
      • Frequent predictions of "Possibly-Human" classes (e.g., Beard, Glasses, Scarf) for male subjects due to the high frequency of facial hair attributes in ImageNet-22k.
    3. Carbon Footprint and Energy Consumption:

      • Pre-training a single DINOv2 ViT-g model consumed 22,016 GPU-hours on NVIDIA A100-40GB GPUs (400W TDP).
      • Assuming a Power Usage Effectiveness (PUE) of 1.1 and US average carbon intensity of 0.385 kg CO2e/kWh0.385\text{ kg CO}_2\text{e/kWh}, pre-training emitted 3.7 tCO2eq3.7\text{ tCO}_2\text{eq} and consumed 9.7 MWh (compared to an estimated 22.4 MWh / 118.9 MWh for retraining OpenCLIP ViT-L / ViT-G).
      • Cumulative emissions across the entire project were estimated between 0.5k0.5\text{k} and 1.0k tCO2eq1.0\text{k tCO}_2\text{eq} (approx200,000\\approx 200,000 GPU-days).

Coverage note — Qualitative visual PCA component figures, patch-matching visualizations, and the full multi-dataset breakdown tables from the appendix were omitted in favor of self-contained quantitative benchmark summaries and exact methodological specifications.

References

  1. 1.Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel. Deep vit features as dense visual descriptors. In ECCV workshop on "What is Motion For?", 2022.
  2. 2.Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. In ICLR, 2020.
  3. 3.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. In ECCV, 2022.
  4. 4.Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In CVPR, 2023.
  5. 5.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. In ICML, 2022.
  6. 6.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. In ICLR, 2021.
  7. 7.Jan Beirlant, Edward J Dudewicz, László Györfi, Edward C Van der Meulen, et al. Nonparametric entropy estimation: An overview. International Journal of Mathematical and Statistical Sciences, 6(1):17–39, 1997.
  8. 8.Maxim Berman, Hervé Jégou, Vedaldi Andrea, Iasonas Kokkinos, and Matthijs Douze. MultiGrain: a unified image embedding for classes and instances. arXiv preprint arXiv:1902.05509, 2019.
  9. 9.Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. Are we done with imagenet? arXiv preprint arXiv:2006.07159, 2020.
  10. 10.Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. In CVPR, 2023.
  11. 11.Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. AdaBins: Depth estimation using adaptive bins. In CVPR, 2021.
  12. 12.Piotr Bojanowski and Armand Joulin. Unsupervised learning by predicting noise. In ICML, 2017.
  13. 13.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  14. 14.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In ECCV, 2014.
  15. 15.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020.
  16. 16.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018.
  17. 17.Mathilde Caron, Piotr Bojanowski, Julien Mairal, and Armand Joulin. Unsupervised pre-training of image features on non-curated data. In ICCV, 2019.
  18. 18.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In NeurIPS, 2020.
  19. 19.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  20. 20.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  21. 21.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020.
  22. 22.Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al. Symbolic discovery of optimization algorithms. arXiv preprint arXiv:2302.06675, 2023a.
  23. 23.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In CVPR, 2021.
  24. 24.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV, 2021.
  25. 25.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. In ICLR, 2023b.
  26. 26.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022.
  27. 27.Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In CVPR, 2023.
  28. 28.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  29. 29.M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi. Describing textures in the wild. In CVPR, 2014.
  30. 30.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  31. 31.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. In NeurIPS, 2022.
  32. 32.Terrance De Vries, Ishan Misra, Changhan Wang, and Laurens Van der Maaten. Does object recognition work for everyone? In CVPR workshops, 2019.
  33. 33.Sylvain Delattre and Nicolas Fournier. On the kozachenko–leonenko entropy estimator. Journal of Statistical Planning and Inference, 185:69–93, 2017.
  34. 34.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  35. 35.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  36. 36.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015.
  37. 37.Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. Discriminative unsupervised feature learning with exemplar convolutional neural networks. TPAMI, 2016.
  38. 38.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  39. 39.Matthijs Douze, Hervé Jégou, Harsimrat Sandhawalia, Laurent Amsaleg, and Cordelia Schmid. Evaluation of gist descriptors for web-scale image search. In CIVR, 2009.
  40. 40.Quentin Duval, Ishan Misra, and Nicolas Ballas. A simple recipe for competitive low-compute self supervised vision models. arXiv preprint arXiv:2301.09451, 2023.
  41. 41.Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Hervé Jegou, and Edouard Grave. Are large-scale datasets necessary for self-supervised pre-training? arXiv preprint arXiv:2112.10740, 2021.
  42. 42.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010.
  43. 43.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. In CVPR, 2023.
  44. 44.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In CVPR, 2004.
  45. 45.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. IJRR, 2013.
  46. 46.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. In ICLR, 2018.
  47. 47.Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Omnimae: Single model masked pretraining on images and videos. In CVPR, 2023.
  48. 48.Priya Goyal, Dhruv Mahajan, Abhinav Gupta, and Ishan Misra. Scaling and benchmarking self-supervised visual representation learning. In ICCV, 2019.
  49. 49.Priya Goyal, Mathilde Caron, Benjamin Lefaudeux, Min Xu, Pengchao Wang, Vivek Pai, Mannat Singh, Vitaliy Liptchinsky, Ishan Misra, Armand Joulin, et al. Self-supervised pretraining of visual features in the wild. preprint arXiv:2103.01988, 2021.
  50. 50.Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Mannat Singh, Ishan Misra, Levent Sagun, Armand Joulin, and Piotr Bojanowski. Vision models are more robust and fair when pretrained on uncurated images without supervision. arXiv preprint arXiv:2202.08360, 2022a.
  51. 51.Priya Goyal, Adriana Romero Soriano, Caner Hazirbas, Levent Sagun, and Nicolas Usunier. Fairness indicators for systematic assessments of visual feature extractors. In FAcct, 2022b.
  52. 52.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The "something something" video database for learning and evaluating visual common sense. In ICCV, 2017.
  53. 53.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. Bootstrap your own latent: A new approach to self-supervised learning. In NeurIPS, 2020.
  54. 54.Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In CVPR, 2006.
  55. 55.Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T Freeman. Unsupervised semantic segmentation by distilling feature correspondences. In ICLR, 2022.
  56. 56.Caner Hazirbas, Joanna Bitton, Brian Dolhansky, Jacqueline Pan, Albert Gordo, and Cristian Canton Ferrer. Towards measuring fairness in ai: the casual conversations dataset. IEEE Transactions on Biometrics, Behavior, and Identity Science, 2021.
  57. 57.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  58. 58.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.
  59. 59.Olivier J Hénaff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, SM Eslami, and Aaron van den Oord. Data-efficient image recognition with contrastive predictive coding. PMLR, 2019.
  60. 60.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In ICLR, 2019.
  61. 61.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In ICCV, 2021a.
  62. 62.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In CVPR, 2021b.
  63. 63.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In NeurIPS Deep Learning Workshop, 2014.
  64. 64.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  65. 65.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In ECCV, 2016.
  66. 66.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip. 2021.
  67. 67.Herve Jegou, Matthijs Douze, and Cordelia Schmid. Product quantization for nearest neighbor search. TPAMI, 2010.
  68. 68.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 2019.
  69. 69.Armand Joulin, Laurens Van Der Maaten, Allan Jabri, and Nicolas Vasilache. Learning visual features from large weakly supervised data. In ECCV, 2016.
  70. 70.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  71. 71.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 3DRR, 2013.
  72. 72.Mario Michael Krell, Matej Kosec, Sergio P. Perez, and Andrew Fitzgibbon. Efficient sequence packing without cross-contamination: Accelerating large language models without impacting performance, 2022.
  73. 73.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  74. 74.Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, and Daniel Haziza. xformers: A modular and hackable transformer modelling library. https://github.com/facebookresearch/xformers, 2022.
  75. 75.Chunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao, Bin Xiao, Xiyang Dai, Lu Yuan, and Jianfeng Gao. Efficient self-supervised vision transformers for representation learning. In ICLR, 2022a.
  76. 76.Zhenyu Li, Xuyang Wang, Xianming Liu, and Junjun Jiang. Binsformer: Revisiting adaptive bins for monocular depth estimation. arXiv preprint arXiv:2204.00987, 2022b.
  77. 77.Tatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert, and Alex Rogozhnikov. Cape: Encoding relative positions with continuous augmented positional embeddings. In NeurIPS, 2021.
  78. 78.Huajun Liu, Fuqiang Liu, Xinyi Fan, and Dong Huang. Polarized self-attention: towards high-quality pixel-wise regression. arXiv preprint arXiv:2107.00782, 2021.
  79. 79.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, 2018.
  80. 80.S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi. Fine-grained visual classification of aircraft. Technical report, 2013.
  81. 81.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In CVPR, 2020.
  82. 82.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In ICVGIP, 2008.
  83. 83.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV, 2016.
  84. 84.Dolev Ofri-Amar, Michal Geyer, Yoni Kasten, and Tali Dekel. Neural congealing: Aligning images to a joint semantic atlas. In CVPR, 2023.
  85. 85.Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. Cats and dogs. In CVPR, 2012.
  86. 86.Deepak Pathak, Philipp Krähenbñhl, Jeff Donahue, Trevor Darrell, and Alexei Efros. Context encoders: Feature learning by inpainting. In CVPR, 2016.
  87. 87.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  88. 88.Ed Pizzi, Sreya Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze. A self-supervised descriptor for image copy detection. In CVPR, 2022.
  89. 89.Filip Radenović, Ahmet Iscen, Giorgos Tolias, Yannis Avrithis, and Ondřej Chum. Revisiting oxford and paris: Large-scale image retrieval benchmarking. In CVPR, 2018a.
  90. 90.Filip Radenović, Giorgos Tolias, and Ondřej Chum. Fine-tuning cnn image retrieval with no human annotation. TPAMI, 2018b.
  91. 91.Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444, 2017.
  92. 92.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  93. 93.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  94. 94.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020.
  95. 95.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In ICCV, 2021.
  96. 96.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In ICML, 2019.
  97. 97.Jerome Revaud, Jon Almazán, Rafael S Rezende, and Cesar Roberto de Souza. Learning with average precision: Training image retrieval with a listwise loss. In ICCV, 2019.
  98. 98.Yangjun Ruan, Saurabh Singh, Warren Morningstar, Alexander A Alemi, Sergey Ioffe, Ian Fischer, and Joshua V Dillon. Weighted ensemble self-supervised learning. In ICLR, 2023.
  99. 99.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. IJCV, 2015.
  100. 100.Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, and Hervé Jégou. Spreading vectors for similarity search. In ICLR, 2019.
  101. 101.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. In NeurIPS Data Centric AI Workshop, 2021.
  102. 102.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. In NeurIPS, 2022.
  103. 103.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  104. 104.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012.
  105. 105.Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan, Ross Girshick, Piotr Dollár, and Laurens van der Maaten. Revisiting Weakly Supervised Pre-Training of Visual Perception Models. In CVPR, 2022.
  106. 106.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015.
  107. 107.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  108. 108.Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your vit? data, augmentation, and regularization in vision transformers. TMLR, 2021.
  109. 109.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. ACL, 2019.
  110. 110.Yonglong Tian, Olivier J Henaff, and Aäron van den Oord. Divide and contrast: Self-supervised learning from uncurated data. In ICCV, 2021.
  111. 111.Giorgos Tolias, Ronan Sicre, and Hervé Jégou. Particular object retrieval with integral max-pooling of cnn activations. In ICLR, 2016.
  112. 112.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In NeurIPS, 2022.
  113. 113.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou. Fixing the train-test resolution discrepancy. In NeurIPS, 2019.
  114. 114.Hugo Touvron, Matthieu Cord, and Hervé Jégou. Deit iii: Revenge of the vit. In ECCV, 2022.
  115. 115.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  116. 116.Narek Tumanyan, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Splicing vit features for semantic appearance transfer. In CVPR, 2022.
  117. 117.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018.
  118. 118.Grant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber, Serge Belongie, and Oisin Mac Aodha. Benchmarking representation learning for natural world image collections. In CVPR, 2021.
  119. 119.Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al. Internimage: Exploring large-scale vision foundation models with deformable convolutions. In CVPR, 2022.
  120. 120.Xiaolong Wang, Allan Jabri, and Alexei A Efros. Learning correspondence from the cycle-consistency of time. In CVPR, 2019.
  121. 121.Frederik Warburg, Soren Hauberg, Manuel Lopez-Antequera, Pau Gargallo, Yubin Kuang, and Javier Civera. Mapillary street-level sequences: A dataset for lifelong place recognition. In CVPR, 2020.
  122. 122.Philippe Weinzaepfel, Thomas Lucas, Diane Larlus, and Yannis Kalantidis. Learning super-features for image retrieval. In ICLR, 2021.
  123. 123.P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona. Caltech-UCSD Birds 200. Technical Report CNS-TR-2010-001, 2010.
  124. 124.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. In LREC, 2020.
  125. 125.Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. Google landmarks dataset v2 – a large-scale benchmark for instance-level recognition and retrieval. In CVPR, 2020.
  126. 126.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018.
  127. 127.J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In CVPR, 2010.
  128. 128.Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, Christoph Feichtenhofer, et al. Masked autoencoders that listen. arXiv preprint arXiv:2207.06405, 2022.
  129. 129.I Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri, and Dhruv Mahajan. Billion-scale semi-supervised learning for image classification. arXiv preprint arXiv:1905.00546, 2019.
  130. 130.Burak Yildiz, Seyran Khademi, Ronald Maria Siebes, and Jan van Gemert. Amstertime: A visual place recognition benchmark dataset for severe domain shift. In ICPR, 2022.
  131. 131.Nikolaos-Antonios Ypsilantis, Noa Garcia, Guangxing Han, Sarah Ibrahimi, Nanne Van Noord, and Giorgos Tolias. The met dataset: Instance-level recognition for artworks. In NeurIPS Datasets and Benchmarks Track, 2021.
  132. 132.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In CVPR, 2022.
  133. 133.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In ECCV, 2016.
  134. 134.Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. Learning deep features for scene recognition using places database. In NeurIPS, 2014.
  135. 135.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017.
  136. 136.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. In ICLR, 2022a.
  137. 137.Pan Zhou, Yichen Zhou, Chenyang Si, Weihao Yu, Teck Khim Ng, and Shuicheng Yan. Mugs: A multi-granular self-supervised learning framework. arXiv preprint arXiv:2203.14415, 2022b.

Citation

MLA
Oquab, M., et al. “DINOv2: Learning Robust Visual Features Without Supervision”. arXiv, 2023, http://arxiv.org/abs/2304.07193v2.
APA
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., … Bojanowski, P. (2023). DINOv2: Learning Robust Visual Features without Supervision. arXiv. http://arxiv.org/abs/2304.07193v2
Chicago
Oquab, M., T. Darcet, T. Moutakanni, et al. 2023. “DINOv2: Learning Robust Visual Features Without Supervision”. arXiv. http://arxiv.org/abs/2304.07193v2.
Harvard
Oquab, M. et al. (2023) “DINOv2: Learning Robust Visual Features without Supervision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.07193v2.
Vancouver
1. Oquab M, Darcet T, Moutakanni T, et al (2023) DINOv2: Learning Robust Visual Features without Supervision. arXiv

BibTeX

@article{oquab2023dinov2,
  title = {DINOv2: Learning Robust Visual Features without Supervision},
  author = {Oquab, Maxime and Darcet, Timothée and Moutakanni, Théo and Vo, Huy and Szafraniec, Marc and Khalidov, Vasil and Fernandez, Pierre and Haziza, Daniel and Massa, Francisco and El-Nouby, Alaaeldin and Assran, Mahmoud and Ballas, Nicolas and Galuba, Wojciech and Howes, Russell and Huang, Po-Yao and Li, Shang-Wen and Misra, Ishan and Rabbat, Michael and Sharma, Vasu and Synnaeve, Gabriel and Xu, Hu and Jegou, Hervé and Mairal, Julien and Labatut, Patrick and Joulin, Armand and Bojanowski, Piotr},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.07193v2},
  eprint = {2304.07193}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF