CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion
Philippe WeinzaepfelVincent LeroyThomas LucasRomain BrégierYohann CabonVaibhav AroraLeonid AntsfeldBoris ChidlovskiiGabriela CsurkaJérôme Revaud
Proposes a self-supervised cross-view image completion framework that learns spatial and geometric relationships across viewpoint pairs, effectively transferring to both monocular and binocular 3D vision downstream tasks such as depth estimation, optical flow, and relative camera pose regression.
Modern computer vision models increasingly rely on self-supervised pre-training to learn useful representations from unlabeled images before adapting to specific downstream tasks. While masked image modeling techniques excel at capturing high-level semantic features for tasks like image classification, they do not inherently teach models the spatial and geometric relationships required for three-dimensional (3D) vision. Consequently, tasks such as depth estimation, optical flow, and camera tracking often require separate, task-specific architectures and costly supervised training data.
The article demonstrates a novel self-supervised pre-training framework called Cross-View Completion (CroCo), designed specifically to learn 3D scene geometry and spatial relationships from unlabeled image pairs without human supervision.
The approach uses a Vision Transformer architecture consisting of a shared encoder and an attention-based decoder. The model receives two images depicting the same scene from different viewpoints. An aggressive 90% of the visual patches from the first image are masked out, and the model must reconstruct these hidden patches by conditioning its predictions on the visible patches and the unmasked second reference image. To evaluate this method, the model was pre-trained on roughly 1.82 million synthetic image pairs of indoor environments generated via a 3D simulator, then tested across multiple single-image (monocular) and two-image (binocular) downstream tasks.
The findings establish that CroCo significantly improves performance on geometric and 3D vision applications. In monocular depth estimation on the standard NYUv2 benchmark, CroCo achieved an accuracy of 85.6%, outperforming competing self-supervised models such as Masked Autoencoders (79.6%) and Multi-Modal Masked Autoencoders (83.0%). Across eight dense 2D and 3D regression tasks on the Taskonomy benchmark, CroCo secured the top performance on six tasks and ranked best overall. For two-image tasks, the generic pre-trained architecture transferred directly without specialized engineering: it reduced optical flow endpoint errors on the MPI-Sintel benchmark by approximately 1.6 to 1.7 pixels compared to baseline models, and achieved a competitive median position error of 5.0 centimeters in relative camera pose estimation.
These results imply that forcing a neural network to reconcile visual discrepancies across distinct viewpoints naturally encodes 3D geometric awareness into the model representations. This reduces the need for expensive multi-modal ground-truth labels and eliminates the requirement for complex, task-specific model architectures in downstream 3D applications. The findings also reveal a trade-off: because the model was trained on indoor scenes rather than object-centric datasets, its accuracy on purely semantic tasks like ImageNet classification is lower than models pre-trained specifically for high-level object recognition.
Organizations developing spatial computing, robotics, autonomous navigation, or augmented reality systems should consider viewpoint-conditioned pre-training as a cost-effective strategy to build foundational geometric models. Future technical initiatives should expand pre-training data beyond synthetic indoor environments to real-world multi-view imagery and explore hybrid datasets that balance geometric learning with semantic classification.
Confidence in these findings is high across geometric and spatial transfer tasks, supported by consistent ablation studies on masking ratios, dataset co-visibility, and decoder variants. However, decision-makers should note that the current pre-training relies entirely on synthetic indoor datasets, and the fixed 224-by-224 input resolution limits test performance on large visual displacements and high-resolution inputs.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). It introduces the foundational masked autoencoder (MAE) framework for self-supervised representation learning via masked visual patch reconstruction, which CroCo directly adapts and generalizes to multi-view image pairs.
- Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). It demonstrates that simple masked image modeling using direct raw-pixel regression provides an effective pre-training signal, establishing core design principles underlying masked visual representation learning.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). It pioneers masked image modeling for Vision Transformers by recovering corrupted visual tokens, setting the stage for self-supervised reconstruction objectives in vision.
- Paper: Context Encoders: Feature Learning by Inpainting, Deepak Pathak et al. (2016). It introduces the foundational concept of visual representation learning through inpainting and spatial context completion.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It establishes the principle of learning visual representations by contrasting shared information across multiple views of a scene, motivating cross-view self-supervision.
- Paper: Unsupervised Monocular Depth Estimation with Left-Right Consistency, Clément Godard et al. (2016). It formalizes cross-view consistency and stereo image reconstruction as effective self-supervised mechanisms for learning geometric scene representations like depth.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It builds upon multi-view visual representation concepts by training large-scale transformers to directly predict unified 3D attributes like poses, depth, and dense point clouds across multiple views.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). It explores cross-view geometric synthesis by conditioning generative diffusion models on viewpoint transformations to reconstruct 3D objects from single views.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). It extends cross-view feature alignment and geometric conditioning to generate synchronized, multiview-consistent images for 3D reconstruction.
- Paper: Depth Anything V2, Lihe Yang et al. (2024). It scales foundational monocular depth estimation representations learned from extensive synthetic and real image distributions, advancing downstream geometric perception.
