Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Mahmoud AssranQuentin DuvalIshan MisraPiotr BojanowskiPascal VincentMichael G. RabbatYann LeCunNicolas Ballas

article2023CVPR734 citations

Introduces the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a scalable self-supervised method that learns semantic image representations by predicting latent target features without relying on pixel-level generation or hand-crafted data augmentations.

Abstract

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Method
  • 4 Related Work
  • 5 Image Classification
  • 6 Local Prediction Tasks
  • 7 Scalability
  • 8 Predictor Visualizations
  • 9 Ablations
  • 10 Conclusion
  • References
  • A Implementation Details
  • A.1 Pretraining
  • A.2 Downstream Tasks
  • B Broader Related Work
  • C Additional Ablations
  • D Finetuning on the full ImageNet
  • E RCDM Visualizations
  • E.1 Encoder Visualization

Knowls

  1. Knowl 1 — Image-Based Joint-Embedding Predictive Architecture (I-JEPA)

    model/method

    The Image-based Joint-Embedding Predictive Architecture (I-JEPA) is a non-generative self-supervised learning framework designed to learn semantic visual representations from images without requiring hand-crafted data augmentations.

    Given an image yy partitioned into NN non-overlapping patches, I-JEPA extracts abstract representation targets by passing the unmasked image through a target encoder fθˉf_{\bar{\theta}}, producing patch-level representations sy={sy1,…,syN}s_y = \{s_{y1}, \dots, s_{yN}\}. A single, spatially distributed context block xx is formed by sampling a large image region and removing any patches overlapping with the target regions. The context block xx is processed by a context encoder fθf_\theta to produce latent patch features sx={sxj}j∈Bxs_x = \{s_{xj}\}_{j \in B_x}, where BxB_x is the set of context patch indices.

    A narrow Vision Transformer predictor network gϕg_\phi takes the context representations sxs_x and a set of learnable mask tokens (each enriched with positional embeddings corresponding to target patch locations) as inputs to predict the latent patch representations of MM target blocks: s^y(i)=gϕ(sx,{mj}j∈Bi)\hat{s}_y(i) = g_\phi(s_x, \{m_j\}_{j \in B_i}) for i∈{1,…,M}i \in \{1, \dots, M\}. The context encoder fθf_\theta and predictor gϕg_\phi are optimized via gradient descent on the prediction error in representation space, while the target encoder parameters θˉ\bar{\theta} are updated at each iteration as an exponential moving average (EMA) of θ\theta to prevent representation collapse.

  2. Knowl 2 — Multi-Block Context and Target Masking Strategy

    model/method

    The multi-block masking strategy in I-JEPA creates a non-trivial self-supervised pretext task that guides the model toward high-level semantic abstractions rather than low-level pixel interpolation:

    1. Target Block Sampling: Given an input image divided into NN non-overlapping patches, M=4M = 4 target blocks B1,B2,B3,B4B_1, B_2, B_3, B_4 are independently sampled. Each target block is sampled with a scale in the range of [0.15,0.2][0.15, 0.2] relative to the full image and an aspect ratio in the range [0.75,1.5][0.75, 1.5]. Target masking is applied to the output feature representations of the target encoder rather than the input image pixels.

    2. Context Block Sampling: A single context block is sampled from the image with a large scale in the range [0.85,1.0][0.85, 1.0] and an aspect ratio of 1.01.0.

    3. Overlap Removal (Context Pruning): Any patch in the context block that overlaps with any of the MM target blocks is removed from the context. The resulting context BxB_x is spatially distributed across the image, providing an informative yet sparse input that requires the context encoder to process fewer patches efficiently.

    Because the targets are sufficiently large blocks and the context is spatially surrounding rather than directly adjacent at the local patch level, the network cannot solve the prediction task through simple local continuous interpolation.

  3. Knowl 3 — I-JEPA Prediction Loss and Optimization Dynamics

    equation

    The training objective of I-JEPA is the mean squared error (L2L_2 distance) between predicted patch embeddings and target patch embeddings across all sampled target blocks:

    LI-JEPA(θ,ϕ;θˉ)=1M∑i=1M∑j∈Bi∥s^yj−syj∥22\mathcal{L}_{\text{I-JEPA}}(\theta, \phi; \bar{\theta}) = \frac{1}{M} \sum_{i=1}^M \sum_{j \in B_i} \|\hat{s}_{yj} - s_{yj}\|_2^2

    where:

    • MM is the number of target blocks (typically M=4M=4).
    • BiB_i is the set of patch indices belonging to the ii-th target block.
    • s^yj∈Rd\hat{s}_{yj} \in \mathbb{R}^d is the predicted patch-level embedding at patch index jj generated by the predictor network gϕ(sx,{mk}k∈Bi)g_\phi(s_x, \{m_k\}_{k \in B_i}).
    • syj∈Rds_{yj} \in \mathbb{R}^d is the target patch-level embedding at patch index jj produced by the target encoder fθˉ(y)f_{\bar{\theta}}(y).

    The context encoder parameters θ\theta and predictor parameters ϕ\phi are updated by gradient descent minimizing LI-JEPA\mathcal{L}_{\text{I-JEPA}}. To avoid representation collapse (where encoders map all inputs to a constant vector), gradients are not backpropagated into the target encoder; instead, the target encoder parameters θˉ\bar{\theta} are updated at step tt via an exponential moving average (EMA):

    θˉt=τθˉt−1+(1−τ)θt\bar{\theta}_t = \tau \bar{\theta}_{t-1} + (1 - \tau) \theta_t

    where τ∈[0,1)\tau \in [0, 1) is the momentum parameter.

  4. Knowl 4 — Impact of Representation-Space vs. Pixel-Space Target Prediction

    data/table

    Computing the predictive loss in abstract representation space versus pixel reconstruction space is a critical factor for linear probe performance in self-supervised learning.

    Targets Architecture Pretraining Epochs ImageNet-1K (1% Labels) Top-1 (%)
    Target-Encoder Output ViT-L/16 500 66.9
    Pixels ViT-L/16 800 40.7

    Under identical linear evaluation protocols on ImageNet-1K using only 1% of the available labels, predicting latent targets produced by an EMA target encoder achieves 66.9% Top-1 accuracy after 500 epochs. In contrast, predicting raw pixel values under the same architecture degrades performance to 40.7% Top-1 accuracy even after 800 epochs. Predicting in latent representation space enables the encoder to discard irrelevant high-frequency pixel variations and capture abstract semantic features.

  5. Knowl 5 — Masking Strategy Ablation in Predictive Self-Supervised Learning

    data/table

    Comparing different masking strategies shows that multi-block masking is necessary for I-JEPA to produce high-level semantic representations.

    Mask Strategy Target Specification Frequency Context Specification ImageNet-1K (1% Labels) Top-1 (%)
    multi-block Block scale (0.15,0.2)(0.15, 0.2) 4 Block scale (0.85,1.0)∖(0.85, 1.0) \setminus Targets 54.2
    rasterized Quadrant 3 Complement quadrant 15.5
    block Block scale 0.60.6 1 Complement 20.2
    random Random patches (60%60\%) 1 Complement 17.6

    All configurations evaluate a ViT-B/16 pretrained for 300 epochs on ImageNet-1K and tested via linear probing with 1% labels:

    • Multi-block masking outperforms rasterized quadrant masking (54.2% vs 15.5%), single large block masking (20.2%), and random patch masking (17.6%).
    • Random patch masking and single large block masking allow the network to rely on low-level local continuity or fail to provide sufficient context, whereas multi-block masking provides a spatially distributed context and requires predicting large semantic blocks.
  6. Knowl 6 — ImageNet-1K Classification Benchmarks for I-JEPA

    data/table

    I-JEPA achieves state-of-the-art results among methods that do not use hand-crafted view augmentations, and approaches or matches methods with view augmentations on frozen linear evaluation and low-shot evaluation on ImageNet-1K.

    Method Architecture Epochs Linear Probing Top-1 (%) 1% Semi-Supervised Top-1 (%)
    Methods without view data augmentations
    data2vec ViT-L/16 1600 77.3 73.3
    MAE ViT-B/16 1600 68.0 –
    MAE ViT-L/16 1600 76.0 67.1
    MAE ViT-H/14 1600 77.2 71.5
    CAE ViT-B/16 1600 70.4 –
    CAE ViT-L/16 1600 78.1 –
    I-JEPA ViT-B/16 600 72.9 –
    I-JEPA ViT-L/16 600 77.5 69.4
    I-JEPA ViT-H/14 300 79.3 73.3
    I-JEPA ViT-H/16448_{448} 300 81.1 77.3
    Methods with extra view data augmentations
    SimCLR v2 ResNet-152 (2×2\times) 800 79.1 70.2
    BYOL ResNet-200 (2×2\times) 800 – 71.2
    DINO ViT-B/8 300 80.1 70.0
    MSN ViT-B/4 300 – 75.7
    iBOT ViT-L/16 250 81.0 –
    iBOT ViT-B/16 400 – 69.7

    When pretrained at resolution 448×448448 \times 448 (ViT-H/16448_{448}), I-JEPA reaches 81.1% linear probing accuracy and 77.3% on 1% low-shot ImageNet-1K, surpassing view-invariant approaches without needing hand-crafted view augmentations during pretraining.

  7. Knowl 7 — Transfer Learning Performance on Semantic and Low-Level Tasks

    data/table

    Linear probe evaluation across diverse semantic classification datasets (CIFAR100, Places205, iNaturalist 2018) and low-level diagnostic tasks (Clevr object counting and depth prediction) shows that I-JEPA learns representations suitable for both high-level semantic understanding and low-level geometry.

    Method Architecture CIFAR100 (%) Places205 (%) iNat18 (%) Clevr/Count (%) Clevr/Dist (%)
    Methods without view data augmentations
    data2vec ViT-L/16 81.6 54.6 28.1 85.3 71.3
    MAE ViT-H/14 77.3 55.0 32.9 90.5 72.4
    I-JEPA ViT-H/14 87.5 58.4 47.6 86.7 72.4
    Methods with extra view data augmentations
    DINO ViT-B/8 84.9 57.9 55.9 86.6 53.4
    iBOT ViT-L/16 88.3 60.4 57.3 85.7 62.8

    Key observations:

    • I-JEPA significantly outperforms generative/reconstruction baselines (MAE and data2vec) on semantic transfer tasks, scoring 87.5% on CIFAR100 (vs 77.3% for MAE) and 47.6% on iNat18 (vs 32.9% for MAE).
    • On low-level tasks, view-invariance methods (DINO and iBOT) suffer on depth prediction (53.4% and 62.8% on Clevr/Dist), whereas I-JEPA preserves local geometric properties, achieving 72.4%.
  8. Knowl 8 — Computational Efficiency and Scaling Properties of I-JEPA

    empirical result

    I-JEPA exhibits higher computational efficiency than both pixel-reconstruction methods and multi-view invariance methods:

    • Pretraining Compute: Pretraining a ViT-Huge/14 (ViT-H/14) on ImageNet-1K with I-JEPA requires under 1200 GPU hours (under 72 hours on 16 A100 GPUs for 300 epochs). This is over 2.5×2.5\times faster than pretraining a much smaller ViT-Small/16 with iBOT (which processes multiple augmented views per image) and over 10×10\times more compute-efficient than training a ViT-H/14 with MAE (1600 epochs).
    • Iteration Speed vs. Convergence Rate: Although computing targets in representation space introduces an overhead of approximately 7% compute time per iteration compared to pixel-space targets (like MAE), I-JEPA requires roughly 5×5\times fewer iterations to converge.
    • Single-View Processing: By relying strictly on a single image view per iteration and predicting in representation space, I-JEPA scales efficiently with model capacity without suffering from the quadratic cost explosion of generating and encoding multiple augmented views.
  9. Knowl 9 — Scaling Dataset Size and Model Size in I-JEPA

    data/table

    Scaling pretraining data from ImageNet-1K to ImageNet-22K and scaling model capacity from ViT-Huge to ViT-Giant improves downstream performance across transfer benchmarks.

    Pretrain Dataset Architecture Equivalent IN1k Epochs CIFAR100 (%) Places205 (%) iNat18 (%) Clevr/Count (%) Clevr/Dist (%)
    IN1k ViT-H/14 300 87.5 58.4 47.6 86.7 72.4
    IN22k ViT-H/14 900 89.5 57.8 50.5 88.6 75.0
    IN22k ViT-G/16 600 89.5 59.1 55.3 86.7 73.0
    • Increasing dataset scale (IN1k to IN22k) on ViT-H/14 improves performance on fine-grained semantic classification (iNat18: 47.6% →\rightarrow 50.5%) and low-level spatial tasks (Clevr/Dist: 72.4% →\rightarrow 75.0%).
    • Scaling model size to ViT-Giant/16 (ViT-G/16) on IN22k further boosts complex semantic classification (Places205 to 59.1%, iNat18 to 55.3%), though the larger patch size (16×1616 \times 16 vs 14×1414 \times 14) slightly reduces accuracy on fine-grained spatial counting and distance tasks.
  10. Knowl 10 — Predictor Latent Representation Visualization via Generative Probing

    empirical result

    Qualitative evaluation of the I-JEPA predictor representations using a frozen-encoder generative decoder (trained following the Representation-Conditioned Diffusion Model / RCDM framework) reveals what information the predictor captures:

    • The context encoder and predictor parameters are frozen, and a conditional generative decoder is trained to reconstruct pixel patches from the average-pooled predictor output embeddings conditioned on positional mask tokens.
    • Samples drawn across different random seeds produce consistent high-level object components, orientations, and correct poses corresponding to the target mask bounding box (e.g., accurately predicting the top of a vehicle or the back and head angle of a bird).
    • While semantic parts and spatial poses remain consistent, low-level details (such as precise surface textures) and background patterns vary across generative samples. This confirms that the predictor learns high-level semantic and positional structure while discarding irrelevant pixel-level noise.

Coverage note — None was omitted; all primary architectural components, mathematical formulations, core ablations, empirical benchmarks (ImageNet-1K linear/low-shot, semantic and low-level transfer, scaling laws), and visualization analyses were extracted.

References

  1. 1.Mahmoud Assran, Randall Balestriero, Quentin Duval, Florian Bordes, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, and Nicolas Ballas. The hidden uniform cluster prior in self-supervised learning. arXiv preprint arXiv:2210.07277, 2022. 1, 13
  2. 2.Mahmoud Assran, Nicolas Ballas, Lluis Castrejon, and Michael Rabbat. Supervision accelerates pre-training in contrastive semi-supervised learning of visual representations. arXiv preprint arXiv:2006.10803, 2020. 13
  3. 3.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. arXiv preprint arXiv:2204.07141, 2022. 1, 2, 3, 5, 6, 12, 13, 16, 17
  4. 4.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Armand Joulin, Nicolas Ballas, and Michael Rabbat. Semi-supervised learning of visual features by non-parametrically predicting view assignments with support samples. arXiv preprint arXiv:2104.13963, 2021. 3, 13
  5. 5.Philip Bachman, R Devon Hjelm, and William Buchwalter. Learning representations by maximizing mutual information across views. Advances in neural information processing systems, 32, 2019. 13
  6. 6.Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli. Efficient self-supervised learning with contextualized target representations for vision, speech and language. arXiv preprint arXiv:2212.07525, 2022. 5
  7. 7.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022. 1, 2, 3, 4, 5, 6, 13
  8. 8.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 1, 3, 4, 13
  9. 9.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906, 2021. 1, 3, 13
  10. 10.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicregl: Self-supervised learning of local visual features. arXiv preprint arXiv:2210.01571, 2022. 1, 13
  11. 11.Florian Bordes, Randall Balestriero, Quentin Garrido, Adrien Bardes, and Pascal Vincent. Guillotine regularization: Improving deep networks generalization by removing their head. arXiv preprint arXiv:2206.13378, 2022. 13
  12. 12.Florian Bordes, Randall Balestriero, and Pascal Vincent. High fidelity visualization of what your self-supervised representation knows about. Transactions on Machine Learning Research, 2022. 7, 16
  13. 13.John Bridle, Anthony Heading, and David MacKay. Unsupervised classifiers, mutual information and’phantom targets. Advances in neural information processing systems, 4, 1991. 13
  14. 14.Jane Bromley, James W Bentz, Leon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Sackinger, and Roopak Shah. Signature verification using a “siamese” time delay neural network. International Journal of Pattern Recognition and Artificial Intelligence, 7(04):669–688, 1993. 1, 3
  15. 15.Zhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang, Davide Modolo, Rahul Bhotika, Zhuowen Tu, and Stefano Soatto. Semi-supervised vision transformers at scale. arXiv preprint arXiv:2208.05688, 2022. 13
  16. 16.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882, 2020. 1, 6
  17. 17.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021. 1, 3, 4, 5, 6, 12, 13
  18. 18.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR, 2020. 13
  19. 19.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. preprint arXiv:2002.05709, 2020. 1, 2, 13
  20. 20.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020. 5
  21. 21.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context autoencoder for self-supervised representation learning. arXiv preprint arXiv:2202.03026, 2022. 5
  22. 22.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020. 12, 13
  23. 23.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020. 1, 3, 13
  24. 24.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021. 4, 5
  25. 25.Yubei Chen, Adrien Bardes, Zengyi Li, and Yann LeCun. Intra-instance vicreg: Bag of self-supervised image patch embedding. arXiv preprint arXiv:2206.08954, 2022. 13
  26. 26.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), volume 1, pages 886–893. Ieee, 2005. 4
  27. 27.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1
  28. 28.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3, 4, 12, 13
  29. 29.Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Herve Jegou, and Edouard Grave. Are large-scale datasets necessary for self-supervised pre-training? arXiv preprint arXiv:2112.10740, 2021. 13
  30. 30.Karl Friston. A theory of cortical responses. Philosophical transactions of the Royal Society B: Biological sciences, 360(1456):815–836, 2005. 1
  31. 31.Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Perez, and Matthieu Cord. Learning representations by predicting bags of visual words. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6928–6938, 2020. 13
  32. 32.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT press, 2016. 13
  33. 33.Priya Goyal, Quentin Duval, Jeremy Reizenstein, Matthew Leavitt, Min Xu, Benjamin Lefaudeux, Mannat Singh, Vinicius Reis, Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Ishan Misra. Vissl. https://github.com/facebookresearch/vissl, 2021. 12
  34. 34.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020. 1, 3, 5, 12, 13
  35. 35.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021. 1, 2, 3, 4, 5, 6, 12, 13, 15, 16
  36. 36.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722, 2019. 1, 3, 12, 13
  37. 37.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 3
  38. 38.Olivier Henaff. Data-efficient image recognition with contrastive predictive coding. In International conference on machine learning, pages 4182–4192. PMLR, 2020. 13
  39. 39.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018. 13
  40. 40.Weihua Hu, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi Sugiyama. Learning discrete representations via information maximizing self-augmented training. In International conference on machine learning, pages 1558–1567. PMLR, 2017. 13
  41. 41.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2901–2910, 2017. 12
  42. 42.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, 2017. 6
  43. 43.Andreas Krause, Pietro Perona, and Ryan Gomes. Discriminative clustering by regularized information maximization. Advances in neural information processing systems, 23, 2010. 13
  44. 44.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 12
  45. 45.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Learning representations for automatic colorization. 2016. 4
  46. 46.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Colorization as a proxy task for visual understanding. 2017. 4
  47. 47.Yann LeCun. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. 2022. 2, 3
  48. 48.Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and Fujie Huang. A tutorial on energy-based learning. Predicting structured data, 1(0), 2006. 2
  49. 49.Ralph Linsker. Self-organization in a perceptual network. Computer, 21(3):105–117, 1988. 13
  50. 50.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 12
  51. 51.Yi Ma, Doris Tsao, and Heung-Yeung Shum. On the principles of parsimony and self-consistency for the emergence of intelligence. Frontiers of Information Technology & Electronic Engineering, pages 1–26, 2022. 13
  52. 52.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6707–6717, 2020. 13
  53. 53.Jovana Mitrovic, Brian McWilliams, Jacob Walker, Lars Buesing, and Charles Blundell. Representation learning via invariant causal mechanisms. arXiv preprint arXiv:2010.07922, 2020. 13
  54. 54.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 13
  55. 55.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019. 12
  56. 56.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2536–2544, 2016. 1, 4
  57. 57.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 4
  58. 58.Rajesh PN Rao and Dana H Ballard. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects. Nature neuroscience, 2(1):79–87, 1999. 1
  59. 59.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. 5, 12
  60. 60.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780, 2017. 12
  61. 61.Yuandong Tian, Xinlei Chen, and Surya Ganguli. Understanding self-supervised learning dynamics without contrastive pairs. In International Conference on Machine Learning, pages 10268–10278. PMLR, 2021. 13
  62. 62.Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic. On mutual information maximization for representation learning. arXiv preprint arXiv:1907.13625, 2019. 13
  63. 63.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8769–8778, 2018. 12
  64. 64.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 3
  65. 65.Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, Pierre-Antoine Manzagol, and Leon Bottou. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of machine learning research, 11(12), 2010. 1, 4, 13
  66. 66.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. arXiv preprint arXiv:2112.09133, 2021. 1, 13
  67. 67.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3733–3742, 2018. 13
  68. 68.Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. Unsupervised data augmentation. arXiv preprint arXiv:1904.12848, 2019. 13
  69. 69.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. arXiv preprint arXiv:2111.09886, 2021. 1, 4
  70. 70.Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks, 2017. 12
  71. 71.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stephane Deny. Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arXiv:2103.03230, 2021. 1, 3, 13
  72. 72.Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby. A large-scale study of representation learning with the visual task adaptation benchmark, 2019. 12
  73. 73.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. 2016. 4
  74. 74.Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. Learning deep features for scene recognition using places database. Advances in neural information processing systems, 27, 2014. 12
  75. 75.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. Ibot: Image bert pre-training with online tokenizer. arXiv preprint arXiv:2111.07832, 2021. 2, 4, 5, 6, 12, 13

Citation

MLA
Assran, M., et al. “Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture”. arXiv, 2023, http://arxiv.org/abs/2301.08243v3.
APA
Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., & Ballas, N. (2023). Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. arXiv. http://arxiv.org/abs/2301.08243v3
Chicago
Assran, M., Q. Duval, I. Misra, et al. 2023. “Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture”. arXiv. http://arxiv.org/abs/2301.08243v3.
Harvard
Assran, M. et al. (2023) “Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.08243v3.
Vancouver
1. Assran M, Duval Q, Misra I, Bojanowski P, Vincent P, Rabbat M, LeCun Y, Ballas N (2023) Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. arXiv

BibTeX

@article{assran2023self,
  title = {Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture},
  author = {Assran, Mahmoud and Duval, Quentin and Misra, Ishan and Bojanowski, Piotr and Vincent, Pascal and Rabbat, Michael and LeCun, Yann and Ballas, Nicolas},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.08243v3},
  eprint = {2301.08243}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE