VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

Adrien BardesJean PonceYann LeCun

article2021ICLR1,429 citations

Introduces VICReg, a self-supervised representation learning approach that prevents informational collapse through explicit variance and covariance regularization, matching state-of-the-art visual recognition performance without requiring negative pairs, momentum encoders, or asymmetric network architectures.

Listen

Self-supervised visual learning enables models to learn rich data representations without human-labeled annotations, typically by ensuring different views of an image map to similar representations. A major operational challenge in these joint-embedding systems is representation collapse, where the system minimizes error trivially by mapping all inputs to a single constant output. Existing remedies rely heavily on complex workarounds—such as large batches of negative samples, memory banks, vector quantization, or asymmetric engineering tricks like stop-gradients and momentum encoders. These constraints complicate training and restrict systems to symmetric, single-modality setups.

The article evaluates a new framework called VICReg (Variance-Invariance-Covariance Regularization). Its main objective is to eliminate representation collapse using an explicit, modular loss function that preserves information across embedding dimensions without requiring architectural tricks, normalization schemes, or weight sharing between network branches.

To evaluate this framework, the authors conducted extensive self-supervised pretraining experiments using standard vision backbones (principally ResNet-50) on the 1,000-class ImageNet dataset across up to 1,000 epochs. The method applies a three-part objective: learning invariance by minimizing distances between transformed views, preserving variance along each dimension using a hinge loss, and enforcing covariance penalties to decorrelate individual embedding variables. The quality of the learned representations was tested across standard benchmarks, including ImageNet linear and semi-supervised evaluations, multi-task transfers (scene classification, multi-label recognition, object detection, and instance segmentation), cross-modal image-text retrieval on MS-COCO, and audio classification on ESC-50.

The results establish several key findings. First, on standard ImageNet linear evaluation, the proposed method achieves a top-1 accuracy of 73.2%, which matches the performance of leading self-supervised methods like Barlow Twins (73.2%) and approaches complex asymmetric frameworks like BYOL (74.3%). Second, the framework demonstrates robust multi-modal and asymmetric capability: on MS-COCO image-to-text retrieval, it achieves a 33.6% Recall@1, outperforming both Barlow Twins (31.4%) and contrastive baselines (30.3%), while also improving raw-versus-spectral audio classification on ESC-50 by 5.7% over supervised baselines. Third, the method maintains stable performance when branches use entirely different network architectures (such as pairing a ResNet with a Vision Transformer), showing only minor performance decreases where other frameworks experience severe degradation or complete incompatibility. Finally, incorporating the variance regularization term into existing asymmetric frameworks such as BYOL and SimSiam accelerates convergence and improves classification accuracy.

These findings indicate that explicit variance and covariance constraints provide a simpler, more interpretable, and mathematically transparent mechanism to prevent model collapse. By removing dependencies on shared weights, batch-wide normalizations, and predictor sub-networks, this approach substantially reduces architectural constraints. Consequently, engineering teams can readily apply joint-embedding self-supervised learning to disparate multi-modal signals—such as combining raw audio with spectrograms or pairing text with visual data—without requiring mirrored network architectures.

Based on these results, machine learning practitioners developing self-supervised workflows should adopt explicit variance and covariance regularization, particularly when handling multi-modal pipelines or architectures with asymmetric branches. When applying this framework, practitioners should ensure the expander network is configured with high output dimensionality (such as 8,192 units), as representation quality scales significantly with expander width.

The reported results carry high confidence across standard computer vision and multimodal retrieval benchmarks, demonstrating run-to-run variation below 0.1% accuracy. However, certain trade-offs remain: the approach slightly underperforms specialized clustering techniques on dense object detection tasks, depends heavily on expander width, and requires initial tuning of loss balancing coefficients to avoid training instability.

  • Paper: Barlow Twins: Self-Supervised Learning via Redundancy Reduction, Jure Zbontar et al. (2021). Barlow Twins introduced redundancy reduction via cross-correlation matrix decorrelation to prevent representation collapse in self-supervised learning, directly establishing the foundation that VICReg builds upon and extends with variance regularization.
  • Paper: Exploring Simple Siamese Representation Learning, Xinlei Chen et al. (2021). SimSiam analyzes the collapse problem in non-contrastive Siamese architectures using architectural dynamics and stop-gradients, motivating VICReg's search for an explicit, principled regularization objective.
  • Paper: Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere, Tongzhou Wang et al. (2020). This paper formalizes the core self-supervised representation principles of alignment and uniformity on the hypersphere, which conceptually parallel VICReg's invariance and variance-covariance criteria.
  • Paper: Momentum Contrast for Unsupervised Visual Representation Learning, Kaiming He et al. (2020). MoCo establishes how contrastive learning prevents representation collapse through negative pair queues and momentum updates, providing essential context for why VICReg aims to avoid these asymmetric heuristics.
  • Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). Contrastive Multiview Coding frames multi-view representation learning as maximizing mutual information across views, outlining the multiview objective that non-contrastive methods like VICReg regularize.
  • Paper: Deep CORAL: Correlation Alignment for Deep Domain Adaptation, Baochen Sun et al. (2016). Deep CORAL details the optimization of second-order covariance alignment penalties in deep networks, underpinning the mechanics of covariance regularization used in VICReg.
  • Paper: Deep Canonical Correlation Analysis, Galen Andrew et al. (2013). Deep Canonical Correlation Analysis establishes the theoretical groundwork for maximizing correlation and decorrelating multi-view latent dimensions in deep neural networks.
Cover for VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

Abstract

Recent self-supervised methods for image representation learning are based on maximizing the agreement between embedding vectors from different views of the same image. A trivial solution is obtained when the encoder outputs constant vectors. This collapse problem is often avoided through implicit biases in the learning architecture, that often lack a clear justification or interpretation. In this paper, we introduce VICReg (Variance-Invariance-Covariance Regularization), a method that explicitly avoids the collapse problem with a simple regularization term on the variance of the embeddings along each dimension individually. VICReg combines the variance term with a decorrelation mechanism based on redundancy reduction and covariance regularization, and achieves results on par with the state of the art on several downstream tasks. In addition, we show that incorporating our new variance term into other methods helps stabilize the training and leads to performance improvements.

Table of Contents

  • 1 Introduction
  • 2 VICReg: intuition
  • 3 Related work
  • 4 VICReg: detailed description
  • 4.1 Method
  • 4.2 Implementation details
  • 5 Results
  • 5.1 Evaluation on ImageNet
  • 5.2 Transfer to other downstream tasks
  • 5.3 Multi-modal pretraining on MS-COCO
  • 6 Analysis
  • 7 Conclusion
  • References
  • A Algorithm
  • B Relation to other self-supervised methods
  • C Additional implementation details
  • C.1 Data augmentation
  • C.2 ImageNet evaluation
  • C.3 Transfer learning
  • C.4 Analysis
  • D Additional results
  • D.1 Other ResNet architectures
  • D.2 Pretraining and evaluation on ESC-50 audio classification
  • D.3 K-nearest-neighbors
  • D.4 Loss function coefficients.
  • D.5 Normalizations
  • D.6 Expander network architecture
  • D.7 Batch size
  • D.8 Combination with BYOL and SimSiam
  • E Running time

Knowls

  1. Knowl 1 — VICReg Architecture and Self-Supervised Learning Framework

    model/method

    VICReg (Variance-Invariance-Covariance Regularization) is a self-supervised representation learning method for joint-embedding architectures designed to prevent representation collapse without requiring negative pair mining, memory banks, asymmetric stop-gradient operations, momentum teacher encoders, output quantization, or feature normalization.

    Given an input image ii sampled from dataset D\mathcal{D}, two distinct views x=t(i)x = t(i) and x′=t′(i)x' = t'(i) are generated by applying stochastic augmentations t,t′∼Tt, t' \sim \mathcal{T}. Each view is mapped into a representation vector by an encoder network:

    y=fθ(x),y′=fθ′′(x′)y = f_\theta(x), \quad y' = f'_{\theta'}(x')

    The representations y,y′∈Rdyy, y' \in \mathbb{R}^{d_y} (typically output by a backbone such as ResNet-50) are mapped by a non-linear expander network into higher-dimensional embeddings:

    z=hϕ(y),z′=hϕ′′(y′)z = h_\phi(y), \quad z' = h'_{\phi'}(y')

    where z,z′∈Rdz, z' \in \mathbb{R}^d with d>dyd > d_y (e.g., d=8192d = 8192 vs. dy=2048d_y = 2048). The expander serves two functions: discarding nuisance variance between representations while mapping them into a higher-dimensional space where linear decorrelation effectively reduces higher-order statistical dependencies. After pretraining, the expander is discarded and representations yy are used for downstream tasks.

    Unlike traditional Siamese frameworks, VICReg does not enforce weight sharing (fθ=fθ′′f_\theta = f'_{\theta'} is optional), identical branch architectures, or identical input modalities, because collapse prevention is enforced by separate regularization terms applied to each branch independently.

  2. Knowl 2 — VICReg Objective Function and Loss Formulation

    equation

    Let Z=[z1,…,zn]∈Rn×dZ = [z_1, \dots, z_n] \in \mathbb{R}^{n \times d} and Z′=[z1′,…,zn′]∈Rn×dZ' = [z'_1, \dots, z'_n] \in \mathbb{R}^{n \times d} denote batches of nn embedding vectors of dimension dd produced by the two branches of a joint-embedding architecture. The VICReg objective function consists of three complementary loss terms:

    1. Invariance Criterion (ss): The mean-squared Euclidean distance between embeddings corresponding to views of the same sample without normalization:

    s(Z,Z′)=1n∑i=1n∥zi−zi′∥22s(Z, Z') = \frac{1}{n} \sum_{i=1}^n \|z_i - z'_i\|_2^2

    1. Variance Regularization (vv): A hinge function enforcing the standard deviation of each embedding dimension along the batch dimension to remain above a target threshold γ\gamma (set to γ=1\gamma = 1):

    v(Z)=1d∑j=1dmax⁡(0,γ−S(zj,ϵ))v(Z) = \frac{1}{d} \sum_{j=1}^d \max\left(0, \gamma - S(z^j, \epsilon)\right)

    where zj∈Rnz^j \in \mathbb{R}^n is the vector of values at dimension jj across the batch, ϵ>0\epsilon > 0 is a small constant (e.g., 10−410^{-4}) to guarantee numerical stability, and S(x,ϵ)=Var(x)+ϵS(x, \epsilon) = \sqrt{\text{Var}(x) + \epsilon} is the regularized standard deviation:

    Var(x)=1n−1∑i=1n(xi−xˉ)2,xˉ=1n∑i=1nxi\text{Var}(x) = \frac{1}{n-1} \sum_{i=1}^n (x_i - \bar{x})^2, \quad \bar{x} = \frac{1}{n} \sum_{i=1}^n x_i

    1. Covariance Regularization (cc): A penalty driving the off-diagonal elements of the batch covariance matrix towards zero to decorrelate embedding dimensions and eliminate informational redundancy:

    c(Z)=1d∑i≠j[C(Z)]i,j2c(Z) = \frac{1}{d} \sum_{i \neq j} [C(Z)]_{i,j}^2

    where the batch sample covariance matrix C(Z)∈Rd×dC(Z) \in \mathbb{R}^{d \times d} is given by:

    C(Z)=1n−1∑i=1n(zi−zˉ)(zi−zˉ)T,zˉ=1n∑i=1nziC(Z) = \frac{1}{n-1} \sum_{i=1}^n (z_i - \bar{z})(z_i - \bar{z})^T, \quad \bar{z} = \frac{1}{n} \sum_{i=1}^n z_i

    1. Total VICReg Loss (ℓ\ell): The complete loss is a weighted sum parameterized by scaling coefficients λ,μ,ν>0\lambda, \mu, \nu > 0:

    ℓ(Z,Z′)=λs(Z,Z′)+μ[v(Z)+v(Z′)]+ν[c(Z)+c(Z′)]\ell(Z, Z') = \lambda s(Z, Z') + \mu \left[v(Z) + v(Z')\right] + \nu \left[c(Z) + c(Z')\right]

    The total objective L=∑I∈D∑t,t′∼Tℓ(ZI,Z′I)\mathcal{L} = \sum_{I \in \mathcal{D}} \sum_{t,t' \sim \mathcal{T}} \ell(Z^I, Z'^I) is minimized over encoder parameters θ\theta and expander parameters ϕ\phi across batches II.

  3. Knowl 3 — Gradient Dynamics of Standard Deviation vs Variance in Hinge Loss Regularization

    theoretical result

    In VICReg, defining the variance regularization term v(Z)v(Z) using the regularized standard deviation S(x,ϵ)=Var(x)+ϵS(x, \epsilon) = \sqrt{\text{Var}(x) + \epsilon} rather than the sample variance Var(x)\text{Var}(x) directly is strictly necessary to prevent complete representation collapse.

    If the sample variance were used directly in the hinge criterion vvar(x)=max⁡(0,γ−Var(x))v_{\text{var}}(x) = \max(0, \gamma - \text{Var}(x)), the derivative with respect to sample component xix_i would be:

    ∂Var(x)∂xi=2n−1(xi−xˉ)\frac{\partial \text{Var}(x)}{\partial x_i} = \frac{2}{n - 1}(x_i - \bar{x})

    When representations begin to collapse and shrink towards a single point, xi→xˉx_i \to \bar{x}, causing ∂Var(x)∂xi→0\frac{\partial \text{Var}(x)}{\partial x_i} \to 0. Consequently, the repulsive gradient vanishes precisely when collapse occurs, preventing the network from recovering from collapsed states.

    In contrast, taking the derivative of the regularized standard deviation S(x,ϵ)=Var(x)+ϵS(x, \epsilon) = \sqrt{\text{Var}(x) + \epsilon} yields:

    ∂S(x,ϵ)∂xi=xi−xˉ(n−1)Var(x)+ϵ\frac{\partial S(x, \epsilon)}{\partial x_i} = \frac{x_i - \bar{x}}{(n - 1)\sqrt{\text{Var}(x) + \epsilon}}

    For small ϵ\epsilon, as Var(x)→0\text{Var}(x) \to 0, the ratio xi−xˉVar(x)\frac{x_i - \bar{x}}{\sqrt{\text{Var}(x)}} maintains unit scale in direction, providing a stable, non-vanishing repulsive force that drives embeddings apart and maintains batch variance at γ\gamma.

  4. Knowl 4 — VICReg PyTorch Implementation Algorithm

    algorithm

    The core optimization step of VICReg computes representations, applies the regularized standard deviation hinge loss along batch dimensions, computes centered sample covariance matrices, and calculates mean squared error invariance.

    import torch
    import torch.nn.functional as F
    
    def off_diagonal(x):
        # Return a flattened view of the off-diagonal elements of a 2D square matrix
        n, m = x.shape
        assert n == m
        return x.flatten()[:-1].view(n - 1, n + 1)[:, 1:].flatten()
    
    def vicreg_loss(z_a, z_b, sim_coeff=25.0, std_coeff=25.0, cov_coeff=1.0, epsilon=1e-4, gamma=1.0):
        N, D = z_a.shape
    
        # 1. Invariance loss: Mean squared Euclidean distance
        sim_loss = F.mse_loss(z_a, z_b)
    
        # 2. Variance loss: Hinge loss on regularized standard deviation
        std_z_a = torch.sqrt(z_a.var(dim=0) + epsilon)
        std_z_b = torch.sqrt(z_b.var(dim=0) + epsilon)
        std_loss = torch.mean(F.relu(gamma - std_z_a)) + torch.mean(F.relu(gamma - std_z_b))
    
        # 3. Covariance loss: Off-diagonal terms of sample covariance
        z_a_centered = z_a - z_a.mean(dim=0)
        z_b_centered = z_b - z_b.mean(dim=0)
        
        cov_z_a = (z_a_centered.T @ z_a_centered) / (N - 1)
        cov_z_b = (z_b_centered.T @ z_b_centered) / (N - 1)
        
        cov_loss = (off_diagonal(cov_z_a).pow(2).sum() / D) + (off_diagonal(cov_z_b).pow(2).sum() / D)
    
        # Total weighted loss
        loss = sim_coeff * sim_loss + std_coeff * std_loss + cov_coeff * cov_loss
        return loss
    

    During training on ImageNet-1K, this loss is minimized using the LARS optimizer with weight decay 10−610^{-6} and cosine learning rate decay.

  5. Knowl 5 — ImageNet Pretraining and Evaluation Experimental Protocol

    experimental setup

    The standard VICReg pretraining and evaluation benchmark setup on ImageNet-1K comprises:

    1. Architecture: Backbone encoder fθf_\theta is a standard ResNet-50 producing representation dimension dy=2048d_y = 2048. Expander network hϕh_\phi is a 3-layer multilayer perceptron where each layer has 8192 units. The first two layers are followed by Batch Normalization and ReLU activations; the third layer is linear without normalization.

    2. Pretraining Optimization: Pretrained for 1000 epochs using the LARS optimizer with weight decay 10−610^{-6}, batch size n=2048n = 2048, and learning rate schedule:

    lr=batch_size256×base_lr\text{lr} = \frac{\text{batch\_size}}{256} \times \text{base\_lr}

    with base_lr=0.2\text{base\_lr} = 0.2. Learning rate follows a cosine decay schedule over epochs starting from 0, with 10 warmup epochs and final value 0.002. Loss coefficients in Eq. (6) are λ=25\lambda = 25, μ=25\mu = 25, ν=1\nu = 1, with ϵ=10−4\epsilon = 10^{-4} and target variance γ=1.0\gamma = 1.0.

    1. Data Augmentation: Two 224×224224 \times 224 crops sampled uniformly with scale (0.08,1.0)(0.08, 1.0), followed by random horizontal flip (p=0.5p=0.5), color jitter (p=0.8p=0.8, brightness/contrast/saturation 0.40.4, hue 0.10.1), grayscale conversion (p=0.2p=0.2), Gaussian blur (p=0.5p=0.5, kernel size 23), and solarization (p=0.1p=0.1).

    2. Downstream ImageNet Evaluations:

      • Linear probe: Frozen ResNet-50 backbone trained with SGD, batch size 256, learning rate 0.02, cosine decay for 100 epochs, weight decay 10−610^{-6}.
      • Semi-supervised: Fine-tuning the backbone and linear head on 1% or 10% labeled subsets of ImageNet for 20 epochs using SGD with no weight decay.
  6. Knowl 6 — ImageNet Linear and Semi-Supervised Classification Benchmarks

    data/table

    Performance of representations learned by ResNet-50 pretrained with VICReg for 1000 epochs on ImageNet-1K without labels, evaluated under linear classification on frozen features and semi-supervised classification with 1% and 10% labeled splits:

    Linear Semi-supervised
    Method Top-1 (%) Top-5 (%) Top-1 (%) Top-5 (%)
    1% 10% 1% 10%
    Supervised 76.5 - 25.4 56.4 48.4 80.4
    MoCo 60.6 - - - - -
    PIRL 63.6 - - - 57.2 83.8
    SimCLR 69.3 89.0 48.3 65.6 75.5 87.8
    MoCo v2 71.1 - - - - -
    SimSiam 71.3 - - - - -
    SwAV 71.8 - - - - -
    InfoMin Aug 73.0 91.1 - - - -
    OBoW 73.8 - - - 82.9 90.7
    BYOL 74.3 91.6 53.2 68.8 78.4 89.0
    SwAV (w/ multi-crop) 75.3 - 53.9 70.2 78.5 89.9
    Barlow Twins 73.2 91.0 55.0 69.7 79.2 89.3
    VICReg (ours) 73.2 91.1 54.8 69.5 79.4 89.5

    VICReg matches Barlow Twins (73.2% Top-1 linear, 54.8% vs 55.0% on 1% semi-supervised, and 69.5% vs 69.7% on 10% semi-supervised) without computing cross-correlations between branches or requiring batch standardization.

  7. Knowl 7 — Transfer Learning Performance on Downstream Vision Tasks

    data/table

    Evaluation of frozen ResNet-50 features pretrained with VICReg transferred to image classification, object detection, and instance segmentation:

    Linear Classification Object Detection Segmentation
    Method Places205 VOC07 iNat18 VOC07+12 det COCO det COCO seg
    Top-1 (%) mAP (%) Top-1 (%) AP50\text{AP}_{50} (%) AP (%) AP (%)
    Supervised 53.2 87.5 46.7 81.3 39.0 35.4
    MoCo v2 51.8 86.4 38.6 82.5 39.8 36.1
    BYOL 54.0 86.6 47.6 - 40.4 37.0
    SwAV (w/ multi-crop) 56.7 88.9 48.6 82.6 41.6 37.8
    Barlow Twins 54.1 86.2 46.5 82.6 40.0 36.7
    VICReg (ours) 54.3 86.6 47.0 82.4 39.4 36.4

    On linear classification benchmarks (Places205 scene recognition, Pascal VOC07 multi-label, and iNaturalist2018 fine-grained), VICReg achieves performance on par with or outperforming Barlow Twins and BYOL. On object detection (Faster R-CNN with C4 backbone on VOC07+12) and instance segmentation (Mask R-CNN with FPN backbone on COCO), VICReg performs comparably to concurrent self-supervised approaches.

  8. Knowl 8 — Branch Independence and Cross-Architecture Pretraining

    empirical result

    Because VICReg regularizes the variance and covariance of each branch independently, it does not require shared weights, shared network architectures, or identical input statistics between branches.

    1. Unshared Weights and Heterogeneous Vision Backbones: Table 5 evaluates linear classification accuracy (100 pretraining epochs on ImageNet-1K) across Shared Weights (SW), Different Weights (DW), and Different Architectures (DA):
    Method SW (R50/R50) DW (R50/R50) DA (R50/R101) DA (R50/ViT-S)
    BYOL 69.3 Failed Failed Failed
    SimCLR 64.4 63.1 63.9 63.5
    Barlow Twins 68.7 64.2 65.3 63.9
    VICReg 68.6 66.5 68.1 66.2

    When switching from SW to DW, VICReg drops by only 2.1% (68.6% →\to 66.5%) compared to a 4.5% drop for Barlow Twins (68.7% →\to 64.2%). With heterogeneous architectures, VICReg outperforms Barlow Twins by 2.8% on ResNet-50 / ResNet-101 and 2.3% on ResNet-50 / ViT-S.

    1. Multi-Modal Pretraining on MS-COCO: On 5K cross-modal retrieval pairing image (ResNet-152) and text (word embedding + GRU), VICReg achieves Image-to-Text R@1 of 33.6% (vs 31.4% Barlow Twins, 30.3% VSE++) and Text-to-Image R@1 of 45.2% (vs 42.9% Barlow Twins, 41.3% VSE++).

    2. Cross-Modal Audio Pretraining on ESC-50: Pretraining a 1D ResNet-18 (raw waveform) jointly with a 2D ResNet-18 (mel spectrogram) yields 78.4% validation Top-1 accuracy, outperforming Barlow Twins (75.4%) and supervised baseline (72.7%).

  9. Knowl 9 — Interaction of Variance and Covariance Regularization with Asymmetric SSL Components

    empirical result

    Ablation of variance regularization (VR), covariance regularization (CR), batch normalization (BN), predictor modules (PR), stop-gradient (SG), and momentum encoders (ME) on ImageNet-1K linear evaluation (100 pretraining epochs):

    Method ME SG PR BN No Reg (%) Var Reg (%) Var/Cov Reg (%)
    BYOL ✓ ✓ ✓ ✓ 69.3 70.2 69.5
    SimSiam ✓ ✓ ✓ 67.9 68.1 67.6
    SimSiam ✓ ✓ 35.1 67.3 67.1
    SimSiam ✓ Collapse 56.8 66.1
    VICReg ✓ Collapse 56.2 67.3
    VICReg ✓ ✓ Collapse 57.1 68.7
    VICReg ✓ Collapse 57.5 68.6
    VICReg Collapse 56.5 67.4

    Key mechanistic observations include:

    • Without variance regularization, joint-embedding architectures without stop-gradient/momentum collapse completely.
    • Adding variance regularization prevents collapse across all settings and removes the necessity for predictor networks or stop-gradient operations.
    • Adding variance regularization to BYOL increases performance from 69.3% to 70.2% at 100 epochs, accelerating convergence by maintaining representation standard deviations throughout training.
    • Variance regularization stabilizes SimSiam without batch normalization, jumping from 35.1% to 67.3% Top-1 accuracy.
  10. Knowl 10 — Ablation of Normalization, Loss Weights, Expander Width, and Batch Size

    data/table

    Empirical analyses of VICReg architectural and hyperparameter choices on ImageNet-1K linear probe accuracy (100 epochs):

    1. Normalization of Embeddings and Representations:
    Representation Normalization Embedding Normalization Top-1 (%)
    Standardization None 68.6
    Standardization Standardization 68.4
    None Standardization 67.4
    None ℓ2\ell_2-normalization 65.1
    None None 67.2

    Applying ℓ2\ell_2 projection onto the unit sphere forces the batch standard deviation to 1/d1/\sqrt{d}, reducing accuracy by 3.5% (65.1%). Unconstrained embeddings allow covariance values to spread across wider ranges, improving optimization.

    1. Loss Term Weighting (λ,μ,ν\lambda, \mu, \nu): Setting μ=0\mu = 0 (no variance regularization) leads to immediate collapse regardless of λ\lambda or ν\nu. Optimal performance is achieved when λ=μ=25\lambda = \mu = 25 and ν=1\nu = 1 (68.6%). Unbalanced settings where ν>μ\nu > \mu or λ≠μ\lambda \neq \mu cause instability.

    2. Expander Dimensionality: Performance increases monotonically with expander width (dd): 256 (55.9%), 512 (59.2%), 1024 (62.4%), 2048 (65.1%), 4096 (67.3%), 8192 (68.6%), and 16384 (68.8%), saturating near 8192.

    3. Batch Size Robustness: Varying pretraining batch size maintains strong accuracy: 128 (67.3%), 256 (67.9%), 512 (68.2%), 1024 (68.3%), 2048 (68.6%), and 4096 (67.8%).

Coverage note — No substantial contributed material was omitted; the knowls cover the VICReg formulation, theoretical gradient dynamics, optimization algorithms, multi-modal transfer experiments, architectural ablations, and benchmark evaluations.

References

  1. 1.Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. In ICLR, 2020.
  2. 2.Philip Bachman, R Devon Hjelm, and William Buchwalter. Learning representations by maximizing mutual information across views. In NeurIPS, 2019.
  3. 3.Miguel A. Bautista, Artsiom Sanakoyeu, Ekaterina Sutter, and Björn Ommer. Cliquecnn: Deep unsupervised exemplar learning. In NeurIPS, 2016.
  4. 4.Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Sackinger, and Roopak Shah. Signature verification using a “siamese” time delay neural network. In NeurIPS, 1994.
  5. 5.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning. In ECCV, 2018.
  6. 6.Mathilde Caron, Piotr Bojanowski, Julien Mairal, and Armand Joulin. Unsupervised pre-training of image features on non-curated data. In ICCV, 2019.
  7. 7.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In NeurIPS, 2020.
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020a.
  9. 9.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. In NeurIPS, 2020b.
  10. 10.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In CVPR, 2020.
  11. 11.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020c.
  12. 12.Sumit Chopra, Raia Hadsell, and Yann LeCun. Learning a similarity metric discriminatively, with application to face verification. In CVPR, 2005.
  13. 13.Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In NeurIPS, 2013.
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  16. 16.Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe. Whitening for self-supervised representation learning, 2021.
  17. 17.Mark Everingham, Luc Van Gool, John Winn Christopher K. I. Williams, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010.
  18. 18.Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. Vse++: Improving visual-semantic embeddings with hard negatives. In BMVC, 2018.
  19. 19.Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin. Liblinear: A library for large linear classification. JMLR, 2008.
  20. 20.Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord. Learning representations by predicting bags of visual words. In CVPR, 2020.
  21. 21.Spyros Gidaris, Andrei Bursuc, Gilles Puy, Nikos Komodakis, Matthieu Cord, and Patrick Pérez. Online bag-of-visual-words generation for unsupervised representation learning. In CVPR, 2021.
  22. 22.Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
  23. 23.Priya Goyal, Quentin Duval, Jeremy Reizenstein, Matthew Leavitt, Min Xu, Benjamin Lefaudeux, Mannat Singh, Vinicius Reis, Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Ishan Misra. Vissl. https://github.com/facebookresearch/vissl, 2021.
  24. 24.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. Bootstrap your own latent: A new approach to self-supervised learning. In NeurIPS, 2020.
  25. 25.Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In CVPR, 2006.
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  27. 27.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  28. 28.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  29. 29.Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2015.
  30. 30.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. In ICLR, 2019.
  31. 31.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018.
  32. 32.Jiabo Huang, Qi Dong andShaogang Gong, and Xiatian Zhu. Unsupervised deep learning by neighbourhood discovery. In ICML, 2019.
  33. 33.Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, S. M. Ali Eslami, and Aäron van den Oord. Data-efficient image recognition with contrastive predictive coding. In ICML, 2019.
  34. 34.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  35. 35.Junnan Li, Pan Zhou, Caiming Xiong, and Steven C.H. Hoi. Prototypical contrastive learning of unsupervised representations. In ICLR, 2021.
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. Microsoft coco: Common objects in context. In ECCV, 2014.
  37. 37.Ilya Loshchilov and Frank Hutter. Sgdr: stochastic gradient descent with warm restarts. In ICLR, 2017.
  38. 38.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In CVPR, 2020.
  39. 39.Karol J. Piczak. ESC: Dataset for Environmental Sound Classification. In Proceedings of the 23rd Annual ACM Conference on Multimedia, 2015.
  40. 40.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, 2015.
  41. 41.Pierre H. Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, and Michal Valko. Byol works even without batch statistics. arXiv preprint arXiv:2010.10241, 2020.
  42. 42.Yonglong Tian, Dilip Krishnan, , and Phillip Isola. Contrastive multiview coding. arXiv preprint arXiv:1906.05849v4, 2019.
  43. 43.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning. In NeurIPS, 2020.
  44. 44.Yuandong Tian, Xinlei Chen, and Surya Ganguli. Understanding self-supervised learning dynamics without contrastive pairs. arXiv preprint arXiv:2102.06810, 2021.
  45. 45.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  46. 46.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019.
  47. 47.Zhirong Wu, Yuanjun Xiong, Stella Yu, , and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018.
  48. 48.Junyuan Xie, Ross Girshick, and Ali Farhadi. Unsupervised deep embedding for clustering analysis. In ICML, 2016.
  49. 49.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017.
  50. 50.Xueting Yan, Ishan Misra, Abhinav Gupta, Deepti Ghadiyaram, and Dhruv Mahajan. Clusterfit: Improving generalization of visual representations. In CVPR, 2020.
  51. 51.Jianwei Yang, Devi Parikh, and Dhruv Batra. Joint unsupervised learning of deep representations and image clusters. In CVPR, 2016.
  52. 52.Mang Ye, Xu Zhang, Pong C Yuen, and Shih-Fu Chang. Unsupervised embedding learning via invariant and spreading instance feature. In CVPR, 2019.
  53. 53.Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888, 2017.
  54. 54.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  55. 55.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arxiv:2103.03230, 2021.
  56. 56.Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. Learning deep features for scene recognition using places database. In NeurIPS, 2014.
  57. 57.Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins. Local aggregation for unsupervised learning of visual embeddings. In ICCV, 2019.

Citation

MLA
Bardes, A., et al. “VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning”. arXiv, 2021, http://arxiv.org/abs/2105.04906v3.
APA
Bardes, A., Ponce, J., & LeCun, Y. (2021). VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. arXiv. http://arxiv.org/abs/2105.04906v3
Chicago
Bardes, A., J. Ponce, and Y. LeCun. 2021. “VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning”. arXiv. http://arxiv.org/abs/2105.04906v3.
Harvard
Bardes, A., Ponce, J. and LeCun, Y. (2021) “VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2105.04906v3.
Vancouver
1. Bardes A, Ponce J, LeCun Y (2021) VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. arXiv

BibTeX

@article{bardes2021vicreg,
  title = {VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning},
  author = {Bardes, Adrien and Ponce, Jean and LeCun, Yann},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2105.04906v3},
  eprint = {2105.04906}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors