CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

Philippe WeinzaepfelVincent LeroyThomas LucasRomain BrégierYohann CabonVaibhav AroraLeonid AntsfeldBoris ChidlovskiiGabriela CsurkaJérôme Revaud

article2022NeurIPS188 citations

Proposes a self-supervised cross-view image completion framework that learns spatial and geometric relationships across viewpoint pairs, effectively transferring to both monocular and binocular 3D vision downstream tasks such as depth estimation, optical flow, and relative camera pose regression.

Listen

Modern computer vision models increasingly rely on self-supervised pre-training to learn useful representations from unlabeled images before adapting to specific downstream tasks. While masked image modeling techniques excel at capturing high-level semantic features for tasks like image classification, they do not inherently teach models the spatial and geometric relationships required for three-dimensional (3D) vision. Consequently, tasks such as depth estimation, optical flow, and camera tracking often require separate, task-specific architectures and costly supervised training data.

The article demonstrates a novel self-supervised pre-training framework called Cross-View Completion (CroCo), designed specifically to learn 3D scene geometry and spatial relationships from unlabeled image pairs without human supervision.

The approach uses a Vision Transformer architecture consisting of a shared encoder and an attention-based decoder. The model receives two images depicting the same scene from different viewpoints. An aggressive 90% of the visual patches from the first image are masked out, and the model must reconstruct these hidden patches by conditioning its predictions on the visible patches and the unmasked second reference image. To evaluate this method, the model was pre-trained on roughly 1.82 million synthetic image pairs of indoor environments generated via a 3D simulator, then tested across multiple single-image (monocular) and two-image (binocular) downstream tasks.

The findings establish that CroCo significantly improves performance on geometric and 3D vision applications. In monocular depth estimation on the standard NYUv2 benchmark, CroCo achieved an accuracy of 85.6%, outperforming competing self-supervised models such as Masked Autoencoders (79.6%) and Multi-Modal Masked Autoencoders (83.0%). Across eight dense 2D and 3D regression tasks on the Taskonomy benchmark, CroCo secured the top performance on six tasks and ranked best overall. For two-image tasks, the generic pre-trained architecture transferred directly without specialized engineering: it reduced optical flow endpoint errors on the MPI-Sintel benchmark by approximately 1.6 to 1.7 pixels compared to baseline models, and achieved a competitive median position error of 5.0 centimeters in relative camera pose estimation.

These results imply that forcing a neural network to reconcile visual discrepancies across distinct viewpoints naturally encodes 3D geometric awareness into the model representations. This reduces the need for expensive multi-modal ground-truth labels and eliminates the requirement for complex, task-specific model architectures in downstream 3D applications. The findings also reveal a trade-off: because the model was trained on indoor scenes rather than object-centric datasets, its accuracy on purely semantic tasks like ImageNet classification is lower than models pre-trained specifically for high-level object recognition.

Organizations developing spatial computing, robotics, autonomous navigation, or augmented reality systems should consider viewpoint-conditioned pre-training as a cost-effective strategy to build foundational geometric models. Future technical initiatives should expand pre-training data beyond synthetic indoor environments to real-world multi-view imagery and explore hybrid datasets that balance geometric learning with semantic classification.

Confidence in these findings is high across geometric and spatial transfer tasks, supported by consistent ablation studies on masking ratios, dataset co-visibility, and decoder variants. However, decision-makers should note that the current pre-training relies entirely on synthetic indoor datasets, and the fixed 224-by-224 input resolution limits test performance on large visual displacements and high-resolution inputs.

arXiv: 2210.10716naver/croco
  • Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). It introduces the foundational masked autoencoder (MAE) framework for self-supervised representation learning via masked visual patch reconstruction, which CroCo directly adapts and generalizes to multi-view image pairs.
  • Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). It demonstrates that simple masked image modeling using direct raw-pixel regression provides an effective pre-training signal, establishing core design principles underlying masked visual representation learning.
  • Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). It pioneers masked image modeling for Vision Transformers by recovering corrupted visual tokens, setting the stage for self-supervised reconstruction objectives in vision.
  • Paper: Context Encoders: Feature Learning by Inpainting, Deepak Pathak et al. (2016). It introduces the foundational concept of visual representation learning through inpainting and spatial context completion.
  • Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It establishes the principle of learning visual representations by contrasting shared information across multiple views of a scene, motivating cross-view self-supervision.
  • Paper: Unsupervised Monocular Depth Estimation with Left-Right Consistency, Clément Godard et al. (2016). It formalizes cross-view consistency and stereo image reconstruction as effective self-supervised mechanisms for learning geometric scene representations like depth.
  • Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It builds upon multi-view visual representation concepts by training large-scale transformers to directly predict unified 3D attributes like poses, depth, and dense point clouds across multiple views.
  • Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). It explores cross-view geometric synthesis by conditioning generative diffusion models on viewpoint transformations to reconstruct 3D objects from single views.
  • Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). It extends cross-view feature alignment and geometric conditioning to generate synchronized, multiview-consistent images for 3D reconstruction.
  • Paper: Depth Anything V2, Lihe Yang et al. (2024). It scales foundational monocular depth estimation representations learned from extensive synthetic and real image distributions, advancing downstream geometric perception.
Cover for CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

Abstract

Masked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-art performance when finetuned for high-level semantic tasks, e.g. image classification and object detection. In this paper we instead seek to learn representations that transfer well to a wide variety of 3D vision and lower-level geometric downstream tasks, such as depth prediction or optical flow estimation. Inspired by MIM, we propose an unsupervised representation learning task trained from pairs of images showing the same scene from different viewpoints. More precisely, we propose the pretext task of cross-view completion where the first input image is partially masked, and this masked content has to be reconstructed from the visible content and the second image. In single-view MIM, the masked content often cannot be inferred precisely from the visible portion only, so the model learns to act as a prior influenced by high-level semantics. In contrast, this ambiguity can be resolved with cross-view completion from the second unmasked image, on the condition that the model is able to understand the spatial relationship between the two images. Our experiments show that our pretext task leads to significantly improved performance for monocular 3D vision downstream tasks such as depth estimation. In addition, our model can be directly applied to binocular downstream tasks like optical flow or relative camera pose estimation, for which we obtain competitive results without bells and whistles, i.e., using a generic architecture without any task-specific design.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Cross-view Completion Pre-training
  • 4 Experimental results
  • 4.1 Monocular transfer tasks
  • 4.1.1 Ablations
  • 4.1.2 Comparison to the state of the art
  • 4.2 Applications to binocular tasks
  • 5 Discussion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Cross-View Completion (CroCo) Pre-training Architecture and Task

    model/method

    Cross-view Completion (CroCo) is a self-supervised representation learning framework designed to learn 3D scene geometry and spatial relationships from pairs of unlabeled images showing the same scene from different viewpoints.

    Given two views x1x_1 and x2x_2 of a 3D scene, each image is divided into NN non-overlapping spatial patches p1={p11,…,p1N}p_1 = \{p_1^1, \dots, p_1^N\} and p2={p21,…,p2N}p_2 = \{p_2^1, \dots, p_2^N\}. A fraction r∈[0,1]r \in [0, 1] (masking ratio, typically r=0.9r = 0.9) of patches from p1p_1 is randomly masked out, leaving visible patches p~1={p1i∣mi=0}\tilde{p}_1 = \{p_1^i \mid m_i = 0\}, where mi∈{0,1}m_i \in \{0, 1\} is the binary mask indicator for patch ii.

    A Siamese Vision Transformer (ViT) encoder EθE_\theta independently extracts token representations from the visible patches of the first image p~1\tilde{p}_1 and all patches of the reference image p2p_2:

    z1=Eθ(p~1),z2=Eθ(p2)z_1 = E_\theta(\tilde{p}_1), \quad z_2 = E_\theta(p_2)

    A transformer decoder DϕD_\phi takes the encoded visible tokens z1z_1 padded with a learned mask token emaske_{\text{mask}} at masked patch locations, and conditions on the reference tokens z2z_2 to reconstruct the original first image patches p1p_1:

    p^1=Dϕ(Eθ(p~1);Eθ(p2))\hat{p}_1 = D_\phi(E_\theta(\tilde{p}_1); E_\theta(p_2))

    To solve this reconstruction task under heavy masking, the network cannot rely solely on monocular semantic priors; it must infer 3D scene geometry and cross-view spatial correspondence between x1x_1 and x2x_2.

  2. Knowl 2 — CroCo Reconstruction Loss Function

    equation

    The CroCo self-supervised objective trains the Siamese encoder EθE_\theta and conditional decoder DϕD_\phi by minimizing the Mean Squared Error (MSE) pixel reconstruction loss computed exclusively over the masked patches p1∖p~1p_1 \setminus \tilde{p}_1 of the first image:

    L(x1,x2)=1∣p1∖p~1∣∑p1i∈p1∖p~1∥p^1i−p1i∥2\mathcal{L}(x_1, x_2) = \frac{1}{|p_1 \setminus \tilde{p}_1|} \sum_{p_1^i \in p_1 \setminus \tilde{p}_1} \| \hat{p}_1^i - p_1^i \|^2

    where p1i∈RP2×3p_1^i \in \mathbb{R}^{P^2 \times 3} is the target RGB pixel vector for the ii-th patch (with patch side length P=16P=16, yielding 16×16×3=76816 \times 16 \times 3 = 768 values), and p^1i\hat{p}_1^i is the decoder output for patch ii.

    In normalized-target pre-training, each target ground-truth patch p1ip_1^i is normalized by subtracting its mean pixel value and dividing by its standard deviation across all pixels within that patch before computing the MSE loss.

  3. Knowl 3 — CrossBlock and CatBlock Attention Mechanisms in the CroCo Decoder

    model/method

    The CroCo decoder combines tokens from the masked input view and the reference view using one of two attention-based block designs:

    1. CrossBlock: Uses alternating self-attention and cross-attention. Tokens from the first image (visible tokens plus repeated learned mask tokens emaske_{\text{mask}}) undergo multi-head self-attention, followed by multi-head cross-attention where the first image tokens serve as queries and the encoded reference image tokens Eθ(p2)E_\theta(p_2) serve as keys and values, followed by a Multi-Layer Perceptron (MLP). This avoids computing joint self-attention across all 2N2N tokens, lowering computational complexity (FLOPs) at the expense of additional parameters in the cross-attention layers.

    2. CatBlock: Concatenates the NN tokens of the first image (visible and masked) with the NN tokens of the reference image along the sequence dimension, adding a learned view embedding to differentiate between views. The combined sequence of 2N2N tokens passes through standard transformer blocks (multi-head self-attention over the joint sequence followed by an MLP). Only the NN tokens corresponding to the first image are projected to predict patch pixels.

  4. Knowl 4 — Decoder Architecture and Target Normalization Ablation

    data/table

    Ablations on target normalization and decoder design (using a ViT-Base/16 encoder, 8 decoder blocks, 90% masking ratio, and 400 epochs of pre-training on the Habitat dataset) show that per-patch pixel normalization consistently improves downstream performance across semantic segmentation (ADE20k mIoU), depth estimation (NYUv2 δ1\delta_1 accuracy), and the Taskonomy benchmark (average L1 error ×1000\times 1000 and average rank across 8 tasks).

    Normalized Decoder ADE20k ↑\uparrow NYUv2 ↑\uparrow Taskonomy ↓\downarrow FLOPs / Params
    Target Block mIoU depth (δ1\delta_1) avg. rank (Encoder + Decoder)
    No CrossBlock 39.0 83.4 34.74 2.50 50.2G / 120M
    Yes CrossBlock 40.6 85.6 33.00 1.63 50.2G / 120M
    Yes CatBlock 41.3 86.2 33.35 1.88 58.5G / 111M

    CatBlock achieves slightly higher performance on semantic segmentation and depth estimation at the cost of higher FLOPs (58.5G vs. 50.2G), while CrossBlock achieves better average ranking across Taskonomy dense regression tasks.

  5. Knowl 5 — Optimal Masking Ratio in Cross-View Completion

    empirical result

    Unlike single-view Masked Image Modeling (e.g., Masked Autoencoders), which typically uses masking ratios around 75%75\%, cross-view completion achieves peak downstream performance across semantic segmentation (ADE20k), monocular depth prediction (NYUv2), and Taskonomy dense prediction tasks at a high masking ratio of r=0.90r = 0.90 (90%90\% of patches masked).

    Because the reference view provides complementary visual information of the scene, retaining only 10%10\% of visible patches from the target view provides sufficient anchor points to establish correspondence and reconstruct the missing 90%90\% of the image without collapsing the difficulty of the pretext task.

  6. Knowl 6 — Comparison of Pre-training Methods on Monocular 3D and Semantic Downstream Tasks

    data/table

    Evaluation of ViT-Base/16 pre-trained representations across monocular high-level semantic tasks (ImageNet-1K linear probing, ADE20k semantic segmentation mIoU) and dense 3D/geometric tasks (NYUv2 depth δ1\delta_1 accuracy at threshold 1.25, and Taskonomy 8-task L1-loss ×1000\times 1000 and average rank):

    Method Pre-train Data IN1K Top-1 ↑\uparrow ADE20k ↑\uparrow NYUv2 ↑\uparrow Taskonomy Avg ↓\downarrow Taskonomy Rank ↓\downarrow
    DINO IN1K 78.2 44.7 66.8 39.07 5.00
    MAE IN1K 68.0 46.1 79.6 36.09 2.13
    MultiMAE IN1K 60.2 46.4 83.0 36.17 2.75
    MAE Habitat 32.5 40.3 79.0 35.65 2.88
    CroCo Habitat 37.0 40.6 85.6 33.00 1.25

    When pre-trained on identical indoor scene pairs (Habitat), CroCo outperforms single-view MAE by +6.6%+6.6\% on NYUv2 depth accuracy (85.6 vs 79.0) and achieves the best overall score and average rank (1.25/5) across Taskonomy tasks, outperforming models pre-trained on ImageNet-1K (including MultiMAE, which utilizes multi-modal supervision).

  7. Knowl 7 — Requirement of True Multi-View Pairs vs. 2D Synthetic Transformations

    empirical result

    Cross-view completion pre-training requires genuine 3D viewpoint variation rather than artificial 2D geometric augmentations.

    When CroCo is pre-trained using image pairs generated via 2D synthetic transformations (homographies, rotations, scaling, cropping) applied to single images rather than pairs sampled from two distinct camera viewpoints in 3D scenes, downstream performance degrades drastically after 400 epochs:

    • ADE20k semantic segmentation mIoU drops from 38.838.8 to 27.027.0.
    • NYUv2 depth δ1\delta_1 accuracy drops from 86.886.8 to 66.166.1.
    • Taskonomy average L1 error rises from 33.5633.56 to 47.8147.81 (average rank worsens from 1.251.25 to 1.751.75).

    When reference images are generated by synthetic 2D transforms, the network solves the completion task by directly fitting the global transformation shortcut without learning underlying 3D scene geometry.

  8. Knowl 8 — Binocular Transfer: Optical Flow Regression with CroCo

    empirical result

    CroCo transfers directly to binocular optical flow estimation without task-specific 4D cost volumes or complex architectures. The pre-trained encoder and decoder are fine-tuned on 40,000 synthetic image pairs from AutoFlow by changing the output linear projection to regress 2 flow channels (u,v)(u, v) per pixel under an MSE loss.

    Evaluation on the MPI-Sintel dataset training split (average endpoint error, AEPE, where lower is better):

    Encoder Init Decoder Init Sintel Clean (AEPE) ↓\downarrow Sintel Final (AEPE) ↓\downarrow
    Random Random 18.81 18.97
    MAE (IN1K) Random 4.68 5.16
    MAE (Habitat) Random 4.63 5.24
    CroCo (Habitat) CroCo (Habitat) 3.00 3.60

    Pre-training both the encoder and decoder with CroCo cross-view completion reduces the optical flow error by approximately 1.61.6 pixels compared to initializing only the encoder from an MAE auto-completion model.

  9. Knowl 9 — Binocular Transfer: Relative Camera Pose Regression

    model/method

    CroCo can be applied to Relative Pose Regression (RPR) between two camera views (x1,x2)(x_1, x_2). A prediction head coupled with a differentiable Procrustes layer is placed on top of the CroCo model to output a rigid transformation (R,t)∈SO(3)×R3(R, t) \in SO(3) \times \mathbb{R}^3, ensuring that RR is a mathematically valid rotation matrix.

    The network is fine-tuned using the loss:

    Lpose=∥R−R^∥F2+λ∥t−t^∥2\mathcal{L}_{\text{pose}} = \| R - \hat{R} \|_F^2 + \lambda \| t - \hat{t} \|^2

    where ∥⋅∥F\| \cdot \|_F is the Frobenius norm, (R,t)(R, t) is the predicted transformation, (R^,t^)(\hat{R}, \hat{t}) is the ground truth relative pose, and λ=100\lambda = 100.

    On the 7-scenes benchmark (evaluating median translation and rotation errors across test scenes relative to retrieved training images via AP-GeM-18 descriptors), CroCo achieves an average error of 5.0 cm/3.46∘5.0\text{ cm} / 3.46^\circ, significantly outperforming a model initialized with MAE (Habitat) (24.8 cm/13.09∘24.8\text{ cm} / 13.09^\circ), RelocNet (21 cm/6.74∘21\text{ cm} / 6.74^\circ), and NC-EssNet (21 cm/7.50∘21\text{ cm} / 7.50^\circ) without requiring multiple-pose fusion or temporal sequences.

  10. Knowl 10 — CroCo Pre-training Dataset and Implementation Specifications

    experimental setup

    The pre-training dataset (termed Habitat) comprises 1,821,391 synthetic image pairs generated from 3D indoor meshes of HM3D, ScanNet, Replica, and ReplicaCAD using the Habitat simulator. In each scene, up to 1,000 viewpoint pairs are randomly sampled subject to having a co-visibility ratio greater than 50%50\%.

    Pre-training specifications:

    • Resolution & Patching: 224×224224 \times 224 pixels with patch size 16×1616 \times 16 (N=196N=196 tokens per image).
    • Encoder Backbone: ViT-Base/16 (12 transformer blocks, 768 feature dimensions, 12 attention heads).
    • Decoder: 8 blocks, 512 feature dimensions, 16 attention heads (interfaced with the encoder via a linear projection layer).
    • Optimization: AdamW optimizer, base learning rate 1.5×10−41.5 \times 10^{-4}, effective batch size 256, 400 epochs, cosine learning rate schedule with 40 warmup epochs.
  11. Knowl 11 — Limitations of Domain Transfer and Semantic Pre-training in CroCo

    limitation

    Models pre-trained with CroCo exhibit specific limitations:

    1. Semantic Task Degradation: CroCo representations obtain lower classification performance on ImageNet-1K linear probing (37.0%37.0\%) compared to models pre-trained directly on ImageNet-1K using DINO (78.2%78.2\%) or MAE (68.0%68.0\%). This stems primarily from pre-training on indoor scene environments (Habitat) rather than object-centric datasets.
    2. Paired Data Requirement: CroCo requires multi-view image pairs observing the same underlying 3D scene with sufficient co-visibility (>50%>50\%), limiting its direct application on unorganized, single-image collections unless multi-view pairs are mined via Structure-from-Motion (SfM) or camera pose metadata.
    3. Test-Time Resolution for Binocular Displacement: Fine-tuning CroCo on fixed 224×224224 \times 224 resolutions degrades performance on optical flow displacements exceeding the receptive patch tile bounds.

Coverage note — None was omitted; all key contributions including pretext task design, decoder architectures, pre-training setup, monocular transfer comparisons, binocular downstream evaluations (optical flow and relative pose regression), ablations, and stated limitations are represented.

References

  1. 1.Yuki M. Asano, Christian Rupprecht, and Andrea Vedaldi. A Critical Analysis of Self-supervision, or what We Can Learn from a Single Image. In ICLR, 2020.
  2. 2.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. In ECCV, 2022.
  3. 3.Sara Atito, Muhammad Awais, and Josef Kittler. SiT: Self-supervised vIsion Transformer. arXiv preprint arXiv:2104.03602, 2021.
  4. 4.Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir. MultiMAE: Multi-modal Multi-task Masked Autoencoders. In ECCV, 2022.
  5. 5.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language. arXiv preprint arXiv:2202.03555, 2022.
  6. 6.Vassileios Balntas, Shuda Li, and Victor Prisacariu. RelocNet: Continuous Metric Learning Relocalisation Using Neural Nets. In ECCV, 2018.
  7. 7.Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT Pre-Training of Image Transformers. In ICLR, 2022.
  8. 8.Romain Brégier. Deep regression on manifolds: a 3D rotation case study. In 3DV, 2021.
  9. 9.Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for optical flow evaluation. In ECCV, 2012.
  10. 10.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018.
  11. 11.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised Learning of Visual Features by Contrasting Cluster Assignments. In NeurIPS, 2020.
  12. 12.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging Properties in Self-Supervised Vision Transformers. In ICCV, 2021.
  13. 13.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative Pretraining From Pixels. In ICML, 2020.
  14. 14.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A Simple Framework for Contrastive Learning of Visual Representations. In ICML, 2020.
  15. 15.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E. Hinton. Big self-supervised models are strong semi-supervised learners. In NeurIPS, 2020.
  16. 16.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context Autoencoder for Self-Supervised Representation Learning. arXiv preprint arXiv:2202.03026, 2022.
  17. 17.Antonio Criminisi, Patrick Pérez, and Kentaro Toyama. Region Filling and Object Removal by Exemplar-Based Image Inpainting. IEEE Trans. Image Processing, 2004.
  18. 18.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. In CVPR, 2017.
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL HLT, 2019.
  20. 20.Mingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi, and Ping Luo. CamNet: Coarse-to-Fine Retrieval for Camera Re-Localization. In ICCV, 2019.
  21. 21.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR, 2021.
  22. 22.Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Häusser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. FlowNet: Learning Optical Flow with Convolutional Networks. In ICCV, 2015.
  23. 23.Alexey Dosovitskiy, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. Discriminative Unsupervised Feature Learning with Convolutional Neural Networks. In NeurIPS, 2014.
  24. 24.Mohamed El Banani and Justin Johnson. Bootstrap Your Own Correspondences. In ICCV, 2021.
  25. 25.Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Hervé Jegou, and Edouard Grave. Are Large-scale Datasets Necessary for Self-Supervised Pre-training? arXiv preprint arXiv:2112.10740, 2021.
  26. 26.Linus Ericsson, Henry Gouk, and Timothy M. Hospedales. How Well Do Self-Supervised Models Transfer? In CVPR, 2021.
  27. 27.Yuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang, and Furu Wei. Corrupted image modeling for self-supervised visual pre-training. arXiv preprint arXiv:2202.03382, 2022.
  28. 28.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised Representation Learning by Predicting Image Rotations. In ICLR, 2018.
  29. 29.Ben Glocker, Shahram Izadi, Jamie Shotton, and Antonio Criminisi. Real-time RGB-D camera relocalization. In ISMAR, 2013.
  30. 30.Clément Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised Monocular Depth Estimation with Left-Right Consistency. In CVPR, 2017.
  31. 31.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. Bootstrap your own latent - A new approach to self-supervised learning. In NeurIPS, 2020.
  32. 32.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked Autoencoders are Scalable Vision Learners. In CVPR, 2022.
  33. 33.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum Contrast for Unsupervised Visual Representation Learning. In CVPR, 2020.
  34. 34.Ji Hou, Saining Xie, Benjamin Graham, Angela Dai, and Matthias Niessner. Pri3D: Can 3D Priors Help 2D Representation Learning? In ICCV, 2021.
  35. 35.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. A lightweight optical flow CNN - revisiting data fidelity and regularization. IEEE Trans. PAMI, 2021.
  36. 36.Longlong Jing and Yingli Tian. Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey. IEEE Trans. PAMI, 2021.
  37. 37.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring Plain Vision Transformer Backbones for Object Detection. In ECCV, 2022.
  38. 38.Zhaowen Li, Zhiyang Chen, Fan Yang, Wei Li, Yousong Zhu, Chaoyang Zhao, Rui Deng, Liwei Wu, Rui Zhao, Ming Tang, and Jinqiao Wang. MST: Masked Self-Supervised Transformer for Visual Representation. In NeurIPS, 2021.
  39. 39.Zhengqi Li and Noah Snavely. MegaDepth: Learning Single-View Depth Prediction from Internet Photos. In CVPR, 2018.
  40. 40.Songtao Liu, Zeming Li, and Jian Sun. Self-EMD: Self-Supervised Object Detection without ImageNet. arXiv preprint arXiv:2011.13677, 2020.
  41. 41.Yunze Liu, Li Yi, Shanghang Zhang, Qingnan Fan, Thomas Funkhouser, and Hao Dong. P4Contrast: Contrastive Learning with Pairs of Point-Pixel Pairs for RGB-D Scene Understanding. arXiv preprint arXiv:2012.13089, 2020.
  42. 42.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A ConvNet for the 2020s. In CVPR, 2022.
  43. 43.Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization. In ICLR, 2019.
  44. 44.Nikolaus Mayer, Eddy Ilg, Philip Häusser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In CVPR, 2016.
  45. 45.Simon Meister, Junhwa Hur, and Stefan Roth. UnFlow: Unsupervised Learning of Optical Flow with a Bidirectional Census Loss. In AAAI, 2018.
  46. 46.Mehdi Noroozi and Paolo Favaro. Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles. In ECCV, 2016.
  47. 47.Mehdi Noroozi, Ananth Vinjimoor, Paolo Favaro, and Hamed Pirsiavash. Boosting Self-Supervised Learning via Knowledge Transfer. In CVPR, 2018.
  48. 48.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS, 2019.
  49. 49.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. Context Encoders: Feature Learning by Inpainting. In CVPR, 2016.
  50. 50.Pedro O. Pinheiro, Amjad Almahairi, Ryan Y. Benmalek, Florian Golemo, and Aaron Courville. Unsupervised Learning of Dense Visual Representations. In NeurIPS, 2020.
  51. 51.Filip Radenović, Giorgos Tolias, and Ondřej Chum. Fine-tuning CNN image retrieval with no human annotation. IEEE trans. PAMI, 2018.
  52. 52.Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alexander Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI. In NeurIPS datasets and benchmarks, 2021.
  53. 53.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In ICCV, 2021.
  54. 54.Jerome Revaud, Jon Almazan, Rafael Sampaio de Rezende, and Cesar Roberto de Souza. Learning with Average Precision: Training Image Retrieval with a Listwise Loss. In ICCV, 2019.
  55. 55.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. IJCV, 2015.
  56. 56.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied AI Research. In ICCV, 2019.
  57. 57.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from RGBD images. In ECCV, 2012.
  58. 58.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. The Replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797, 2019.
  59. 59.Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu. AutoFlow: Learning a better training set for optical flow. In CVPR, 2021.
  60. 60.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. In CVPR, 2018.
  61. 61.Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Training Home Assistants to Rearrange their Habitat. In NeurIPS, 2021.
  62. 62.Zachary Teed and Jia Deng. RAFT: recurrent all-pairs field transforms for optical flow. In ECCV, 2020.
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  64. 64.Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol. Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion. JMLR, 2010.
  65. 65.Luya Wang, Feng Liang, Yangguang Li, Honggang Zhang, Wanli Ouyang, and Jing Shao. RePre: Improving Self-Supervised Vision Transformer with Reconstructive Pre-training. arXiv preprint arXiv:2201.06857, 2022.
  66. 66.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Dense Contrastive Learning for Self-Supervised Visual Pre-Training. In CVPR, 2021.
  67. 67.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked Feature Prediction for Self-Supervised Visual Pre-Training. In CVPR, 2022.
  68. 68.Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance-level discrimination. In CVPR, 2018.
  69. 69.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning. In CVPR, 2021.
  70. 70.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. SimMIM: A Simple Framework for Masked Image Modeling. In CVPR, 2022.
  71. 71.Yuwen Xiong, Mengye Ren, Wenyuan Zeng, and Raquel Urtasun. Self-Supervised Representation Learning from Flow Equivariance. In ICCV, 2021.
  72. 72.Ceyuan Yang, Zhirong Wu, Bolei Zhou, and Stephen Lin. Instance Localization for Self-supervised Detection Pretraining. In CVPR, 2021.
  73. 73.Zhenheng Yang, Peng Wang, Wei Xu, Liang Zhao, and Ramakant Nevatia. Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency. In AAAI, 2018.
  74. 74.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. HRFormer: High-Resolution Transformer for Dense Prediction. In NeurIPS, 2021.
  75. 75.Amir Zamir, Alexander Sax, William Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In CVPR, 2018.
  76. 76.Sameh Zarif, Ibrahima Faye, and Dayang Rohaya. Image Completion: Survey and Comparative Study. IJPRAI, 2015.
  77. 77.Yucheng Zhao, Guangting Wang, Chong Luo, Wenjun Zeng, and Zheng-Jun Zha. Self-supervised visual representations learning by contrastive mask prediction. In ICCV, 2021.
  78. 78.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K Dataset. In CVPR, 2017.
  79. 79.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. iBOT: Image BERT Pre-training with Online Tokenizer. In ICLR, 2022.
  80. 80.Qunjie Zhou, Torsten Sattler, Marc Pollefeys, and Laura Leal-Taixe. To Learn or Not to Learn: Visual Localization from Essential Matrices. In ICRA, 2020.
  81. 81.Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. Unsupervised Learning of Depth and Ego-Motion from Video. In CVPR, 2017.

Citation

MLA
Weinzaepfel, P., et al. “CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 3502–16, https://proceedings.neurips.cc/paper_files/paper/2022/file/16e71d1a24b98a02c17b1be1f634f979-Paper-Conference.pdf.
APA
Weinzaepfel, P., Leroy, V., Lucas, T., BRÉGIER, R., Cabon, Y., ARORA, V., Antsfeld, L., Chidlovskii, B., Csurka, G., & Revaud, J. (2022). CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion. Advances in Neural Information Processing Systems, 35, 3502–3516. https://proceedings.neurips.cc/paper_files/paper/2022/file/16e71d1a24b98a02c17b1be1f634f979-Paper-Conference.pdf
Chicago
Weinzaepfel, P., V. Leroy, T. Lucas, et al. 2022. “CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion”. Advances in Neural Information Processing Systems 35: 3502–16. https://proceedings.neurips.cc/paper_files/paper/2022/file/16e71d1a24b98a02c17b1be1f634f979-Paper-Conference.pdf.
Harvard
Weinzaepfel, P. et al. (2022) “CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 3502–3516. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/16e71d1a24b98a02c17b1be1f634f979-Paper-Conference.pdf.
Vancouver
1. Weinzaepfel P, Leroy V, Lucas T, BRÉGIER R, Cabon Y, ARORA V, Antsfeld L, Chidlovskii B, Csurka G, Revaud J (2022) CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 3502–3516

BibTeX

@inproceedings{weinzaepfel2022croco,
  title = {CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion},
  author = {Weinzaepfel, Philippe and Leroy, Vincent and Lucas, Thomas and BRÉGIER, Romain and Cabon, Yohann and ARORA, Vaibhav and Antsfeld, Leonid and Chidlovskii, Boris and Csurka, Gabriela and Revaud, Jerome},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {3502-3516},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/16e71d1a24b98a02c17b1be1f634f979-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors