Siamese Image Modeling for Self-Supervised Vision Representation Learning

Chenxin TaoXizhou ZhuWeijie SuGao HuangBin LiJie ZhouYu QiaoXiaogang WangJifeng Dai

article2023CVPR112 citations

Proposes Siamese Image Modeling, a self-supervised framework that bridges the gap between instance discrimination and masked image modeling by predicting dense representations across differently augmented views to achieve both semantic alignment and spatial sensitivity.

Listen

Modern computer vision relies heavily on self-supervised pre-training, which trains artificial intelligence models on vast amounts of visual data without requiring costly human annotations. Currently, two competing pre-training approaches dominate the field: Instance Discrimination, which compares different views of an entire image, and Masked Image Modeling, which learns by reconstructing missing parts of a single masked image. However, practitioners face a fundamental trade-off. Instance Discrimination creates high-level semantic representations that separate categories well but fails at detailed, location-sensitive tasks like object detection. Conversely, Masked Image Modeling excels at capturing fine spatial details but struggles with broad category alignment and sample-efficient classification.

The article evaluates a unified framework called Siamese Image Modeling, designed to resolve this dilemma. The researchers developed a system that uses a two-branch neural network to predict the detailed, local representations of an augmented image view from a different, partially masked view of the same image. By strictly aligning the relative spatial positions between these two views, the model simultaneously learns high-level semantic concepts and precise spatial geometry using a single, unified training loss.

The experimental evaluation demonstrated consistent performance gains across multiple standard computer vision benchmarks. In image classification, the proposed method achieved superior accuracy while outperforming standard Masked Image Modeling by 10.0 percentage points in linear evaluation and by 14.0 percentage points in data-limited scenarios using only 1% of labeled training data. In dense prediction tasks, it outperformed leading Instance Discrimination models by 4.2 points in object detection and exceeded pure baseline models by 3.0 points in complex scene segmentation. Furthermore, the model showed significant advantages in real-world challenges, gaining 1.6 points on rare categories in long-tailed object detection and improving robustness against image corruptions and perturbations by 4.5 to 6.1 points over existing baselines.

These findings show that organizations do not need to choose between models optimized for high-level classification and those tailored for detailed localization. A single pre-training framework can deliver state-of-the-art results across both domains, reducing the engineering overhead and infrastructure cost of maintaining specialized pre-trained models. This capability is especially impactful for applications facing severe label scarcity, rare objects, or harsh operational environments where visual inputs may be distorted.

Decision-makers should consider adopting dense cross-view reconstruction for foundational vision models, particularly when deploying systems for medical imaging, autonomous navigation, or quality inspection where fine spatial accuracy and data efficiency are critical. However, pre-training this architecture requires significant computational power and memory. Future initiatives should focus on developing more lightweight, compute-efficient training variants and verifying that source training datasets do not introduce unintended operational biases before large-scale commercial deployment.

arXiv: 2206.01204
Cover for Siamese Image Modeling for Self-Supervised Vision Representation Learning

Abstract

Self-supervised learning (SSL) has delivered superior performance on a variety of downstream vision tasks. Two main-stream SSL frameworks have been proposed, i.e., Instance Discrimination (ID) and Masked Image Modeling (MIM). ID pulls together representations from different views of the same image, while avoiding feature collapse. It lacks spatial sensitivity, which requires modeling the local structure within each image. On the other hand, MIM reconstructs the original content given a masked image. It instead does not have good semantic alignment, which requires projecting semantically similar views into nearby representations. To address this dilemma, we observe that (1) semantic alignment can be achieved by matching different image views with strong augmentations; (2) spatial sensitivity can benefit from predicting dense representations with masked images. Driven by these analysis, we propose Siamese Image Modeling (SiameseIM), which predicts the dense representations of an augmented view, based on another masked view from the same image but with different augmentations. SiameseIM uses a Siamese network with two branches. The online branch encodes the first view, and predicts the second view's representation according to the relative positions between these two views. The target branch produces the target by encoding the second view. SiameseIM can surpass both ID and MIM on a wide range of downstream tasks, including ImageNet finetuning and linear probing, COCO and LVIS detection, and ADE20k semantic segmentation. The improvement is more significant in few-shot, long-tail and robustness-concerned scenarios. Code shall be released.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Method
  • 3.1. Augmented Inputs
  • 3.2. Prediction Targets
  • 3.3. Loss Function
  • 3.4. Discussion
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Siamese Image Modeling Framework

    model/method

    Siamese Image Modeling (SiameseIM) is a self-supervised vision representation learning framework designed to simultaneously achieve semantic alignment (linear separability) and spatial sensitivity (dense localization) using a single dense objective. SiameseIM operates on two differently augmented views, xax_a and xbx_b, cropped from the same image:

    1. Online Branch: Consists of an encoder (a Vision Transformer backbone followed by a Transformer projector) and a Transformer decoder. A masking strategy (such as blockwise masking) is applied to xax_a, and the online encoder processes only the visible patch tokens into latent representations ya∈RNv×Dy_a \in \mathbb{R}^{N_v \times D}, where NvN_v is the number of visible patches and DD is the embedding dimension. The online decoder takes yay_a, learnable mask tokens mm corresponding to the locations of patches in xbx_b, and positional/scale embeddings derived from the relative spatial coordinates between xax_a and xbx_b, outputting predicted dense representations yb∈RN×Dy_b \in \mathbb{R}^{N \times D} for all NN tokens of view xbx_b.

    2. Target Branch: Consists of a momentum encoder whose weights are updated via an exponential moving average (EMA) of the online encoder. The target encoder receives the entire unmasked view xbx_b and outputs the target token representations zb∈RN×Dz_b \in \mathbb{R}^{N \times D}.

    During pre-training, a dense loss aligns yby_b with zbz_b. For downstream transfer tasks (e.g., classification, detection, segmentation), only the pre-trained online backbone is retained.

  2. Knowl 2 — Relative Positional and Scale Embeddings for Inter-View Token Alignment

    equation

    To enable the online decoder to predict the dense representations of target view xbx_b from visible patches of input view xax_a, positional embeddings are defined relative to the top-left origin of xax_a. Let the bounding box coordinates of the two cropped views in the original image coordinate frame be (i1,j1,h1,w1)(i_1, j_1, h_1, w_1) for xax_a and (i2,j2,h2,w2)(i_2, j_2, h_2, w_2) for xbx_b, representing top, left, height, and width, respectively. Let NhN_h and NwN_w denote the token grid dimensions along height and width (where N=Nh×NwN = N_h \times N_w), and (u,v)(u, v) denote token grid indices (1≤u≤Nh1 \le u \le N_h, 1≤v≤Nw1 \le v \le N_w).

    The unnormalized 2D coordinate positions for patches in xax_a and mask tokens predicting xbx_b are given by:

    p~a(u,v)=(u−1,v−1)\tilde{p}_a^{(u,v)} = (u-1, v-1)

    p~b(u,v)=(h2h1(u−1)+i2−i1h1Nh,  w2w1(v−1)+j2−j1w1Nw)\tilde{p}_b^{(u,v)} = \left(\frac{h_2}{h_1}(u-1) + \frac{i_2-i_1}{h_1}N_h, \; \frac{w_2}{w_1}(v-1) + \frac{j_2-j_1}{w_1}N_w\right)

    The relative scale change between the views is parameterized as:

    s=(10log⁡h2h1,  10log⁡w2w1)s = \left(10\log\frac{h_2}{h_1}, \; 10\log\frac{w_2}{w_1}\right)

    The final positional embeddings are computed using 2D sine-cosine positional encodings PE(⋅)\mathrm{PE}(\cdot):

    pa(u,v)=PE(p~a(u,v))p_a^{(u,v)} = \mathrm{PE}(\tilde{p}_a^{(u,v)})

    pb(u,v)=Linear(Concat(PE(p~b(u,v)),PE(s)))p_b^{(u,v)} = \mathrm{Linear}\left(\mathrm{Concat}\left(\mathrm{PE}(\tilde{p}_b^{(u,v)}), \mathrm{PE}(s)\right)\right)

    The online decoder g(⋅)g(\cdot) computes the predictions yb∈RN×Dy_b \in \mathbb{R}^{N \times D} by combining the visible tokens yay_a, mask token mm, and positional embeddings:

    yb=g(Concat(ya+pa,  {m+pb(u,v)}u=1,v=1Nh,Nw))y_b = g\left(\mathrm{Concat}\left(y_a + p_a, \; \{m + p_b^{(u,v)}\}_{u=1,v=1}^{N_h, N_w}\right)\right)

  3. Knowl 3 — Token-Level Dense Contrastive Loss Formulation

    equation

    SiameseIM trains the network using a dense, token-level UniGrad contrastive loss over all tokens. Let ybi∈RDy_b^i \in \mathbb{R}^D denote the ii-th predicted token representation from the online decoder, and let zbi∈RDz_b^i \in \mathbb{R}^D denote the corresponding target representation produced by the target branch for i=1,…,Ni = 1, \dots, N. The set N={zbj}j=1N\mathcal{N} = \{z_b^j\}_{j=1}^N represents the negative sample set consisting of all target branch token representations.

    The loss is defined as:

    L=E{ybi,zbi}[−∥ybi−zbi∥2+λ∑u∈N(uTybi)2]L = \mathbb{E}_{\{y_b^i, z_b^i\}} \left[ -\|y_b^i - z_b^i\|^2 + \lambda \sum_{u \in \mathcal{N}} (u^T y_b^i)^2 \right]

    where λ\lambda is a balancing hyperparameter set to λ=0.02\lambda = 0.02.

    Target tokens zbiz_b^i are normalized with LayerNorm without learnable affine parameters, while online predictions ybiy_b^i receive no normalization. By computing the covariance matrix of negative samples first, this formulation consumes O(D2)\mathcal{O}(D^2) memory rather than the O(∣N∣)\mathcal{O}(|\mathcal{N}|) memory required by standard InfoNCE loss, enabling tractable token-level contrastive optimization across full token sets.

  4. Knowl 4 — Benchmark Evaluation of SiameseIM Across Downstream Vision Tasks

    data/table

    Using a standard ViT-B/16 backbone, SiameseIM outperforms both instance discrimination baselines (MoCo-v3, DINO) and masked image modeling baselines (MAE, BEiT) across image classification, dense object detection, instance segmentation, semantic segmentation, long-tail detection, and out-of-distribution robustness.

    Method Pretrain ImageNet COCO ADE20k LVIS (rare)
    Epochs FT LIN APb\text{AP}^b APm\text{AP}^m mIoU APrareb\text{AP}^b_{rare} APrarem\text{AP}^m_{rare}
    MoCo-v3 600 83.0 76.7 47.9 42.7 47.3 25.5 25.8
    BEiT 800 83.2 - 49.8 44.4 47.1 - -
    MAE 1600 83.6 68.0 51.6 45.9 48.1 29.3 29.1
    SiameseIM 400 83.7 76.8 50.7 44.9 49.6 28.9 27.7
    SiameseIM 1600 84.1 78.0 52.1 46.2 51.1 30.9 30.1

    On ImageNet-1k, SiameseIM achieves 84.1%84.1\% fine-tuning top-1 accuracy (FT), 78.0%78.0\% linear probing accuracy (LIN, +10.0%+10.0\% over MAE), and 65.1%65.1\% top-1 accuracy under 1%1\% few-shot fine-tuning (+1.7%+1.7\% over MoCo-v3, +14.0%+14.0\% over MAE). On ADE20k semantic segmentation, SiameseIM attains 51.151.1 mIoU (+3.0+3.0 over MAE, +3.8+3.8 over MoCo-v3). On long-tail LVIS detection, SiameseIM achieves 30.9 APrareb30.9\ \text{AP}^b_{rare} (+1.6+1.6 over MAE) and 30.1 APrarem30.1\ \text{AP}^m_{rare} (+1.0+1.0 over MAE). On robustness benchmarks (average of IN-A, IN-R, IN-Sketch top-1 accuracies and IN-C 1−mCE1-\text{mCE}), SiameseIM achieves an average score of 47.947.9, compared to 43.443.4 for MoCo-v3 and 41.841.8 for MAE.

  5. Knowl 5 — Cross-View Feature Prediction versus Intra-View Pixel Reconstruction

    empirical result

    The choice of prediction target interacts critically with whether pre-training uses a single view or multiple cross-views:

    1. Single-View Setting: When the input and target are generated from the identical cropped view (standard Masked Image Modeling), predicting raw pixels is superior to predicting latent features, yielding 82.8%82.8\% vs. 81.0%81.0\% ImageNet fine-tuning accuracy, 62.3%62.3\% vs. 48.7%48.7\% linear probing accuracy, and 47.347.3 vs. 43.543.5 COCO detection APb\text{AP}^b.

    2. Cross-View Setting: When predicting content across two differently augmented views xax_a and xbx_b, predicting latent features substantially outperforms predicting raw pixels. Latent feature prediction achieves 82.9%82.9\% ImageNet fine-tuning, 69.6%69.6\% linear probing, and 48.548.5 COCO APb\text{AP}^b, whereas raw pixel prediction degrades to 78.7%78.7\% fine-tuning, 46.2%46.2\% linear probing, and 38.138.1 COCO APb\text{AP}^b.

    Predicting raw pixels across distinct views is an overly difficult pretext task due to severe pixel-level discrepancies, while predicting target encoder features abstracts away irrelevant low-level pixel variance and forces the extraction of shared semantic structures.

  6. Knowl 6 — Effect of Cross-View Inputs and Color Augmentations on Semantic Alignment

    empirical result

    In self-supervised pre-training with dense loss, introducing cross-view matching and color augmentations is crucial for developing linear separability:

    • Cross-View Input: Changing the input and target from the same view to two different spatial crops increases ImageNet linear probing performance by approximately 11%11\% (from 62.3%62.3\% to 73.1%73.1\% under equivalent 400-epoch ablation settings), demonstrating that enforcing consistency across distinct crops produces strong semantic alignment.
    • Color Augmentation: Adding strong color jittering in cross-view pre-training provides a +3.5%+3.5\% gain in linear probing (from 69.6%69.6\% to 73.1%73.1\%) and improves ImageNet fine-tuning from 82.9%82.9\% to 83.0%83.0\%. Conversely, applying color augmentations within a single-view reconstruction framework degrades linear probing (from 62.3%62.3\% to 59.9%59.9\%) and detection APb\text{AP}^b (from 47.347.3 to 46.346.3) because identical-view color distortion leaks target information and interferes with reconstruction.
  7. Knowl 7 — Sufficiency of Dense Supervision over Global Contrastive Loss

    empirical result

    Applying contrastive supervision purely at the dense token level is sufficient to learn both high semantic linear separability and spatial sensitivity, outperforming global image-level contrastive losses:

    • Training SiameseIM with a token-level dense UniGrad loss achieves 73.6%73.6\% linear probing accuracy and 48.7 APb48.7\ \text{AP}^b on COCO object detection (400 epochs pre-training).
    • Replacing the dense loss with a global pooling UniGrad loss over the whole image representation drops linear probing performance to 72.0%72.0\% and reduces COCO object detection performance to 45.9 APb45.9\ \text{AP}^b.
    • Adding masking to a global contrastive baseline (MoCo-v3 with mask) yields only 72.2%72.2\% linear probing and 45.0 APb45.0\ \text{AP}^b.

    Dense token supervision aligned via relative spatial coordinates provides +2.8 APb+2.8\ \text{AP}^b on detection and +1.6%+1.6\% on linear probing over global supervision without requiring any auxiliary global objective.

  8. Knowl 8 — Ablation of Masking Strategies, Normalization Schemes, and Loss Functions in SiameseIM

    empirical result

    Ablation studies on SiameseIM with a 400-epoch pre-trained ViT-B/16 identify the optimal architectural and optimization configurations:

    • Masking Strategy: Blockwise masking outperforms random patch masking on both linear probing (74.7%74.7\% vs. 73.6%73.6\%) and COCO object detection (50.050.0 vs. 48.7 APb48.7\ \text{AP}^b). Masking contiguous regions prevents simple local patch interpolation and forces the model to learn long-range contextual dependencies.
    • Normalization: Replacing LayerNorm (LN) with BatchNorm (BN) in the projector and decoder improves linear probing from 73.1%73.1\% to 73.6%73.6\% and detection from 47.947.9 to 48.7 APb48.7\ \text{AP}^b. In the loss formulation, MAE-like normalization (applying LN without affine parameters only to target tokens) achieves 76.8%76.8\% linear probing, outperforming MoCo-like normalization (BN without affine parameters followed by ℓ2\ell_2 normalization on both target and prediction) at 74.7%74.7\%.
    • Loss Formulation: The dense UniGrad loss achieves 83.7%83.7\% fine-tuning and 76.8%76.8\% linear probing, slightly outperforming a dense ℓ2\ell_2 regression loss (83.3%83.3\% fine-tuning and 76.5%76.5\% linear probing).
  9. Knowl 9 — Pre-Training Computational Overhead from Unmasked Target Branch Processing

    limitation

    SiameseIM has lower pre-training computational efficiency compared to pure Masked Image Modeling methods such as Masked Autoencoders (MAE). In MAE, masked tokens are completely omitted from the encoder, and no separate target network is evaluated. In SiameseIM, the target branch requires passing the full target image crop xbx_b (all NN unmasked tokens) through a full Vision Transformer momentum encoder to generate dense feature targets zbz_b, increasing the memory footprint and per-epoch training compute.

Coverage note — None. All core contributed methodological designs, mathematical formulations, key downstream empirical benchmarks, ablation studies, and stated limitations have been fully extracted into knowls.

References

  1. 1.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. arXiv preprint arXiv:2204.07141, 2022. 1, 3, 11
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 7
  3. 3.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022. 3
  4. 4.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 1, 3, 4, 6, 11, 12
  5. 5.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906, 2021. 1, 3
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, volume 33, pages 1877–1901, 2020. 3
  7. 7.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, pages 9650–9660, 2021. 3, 4
  8. 8.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In ICML, pages 1691–1703. PMLR, 2020. 3
  9. 9.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, pages 1597–1607. PMLR, 2020. 1, 3, 4, 5, 7
  10. 10.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context autoencoder for self-supervised representation learning. arXiv preprint arXiv:2202.03026, 2022. 3
  11. 11.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020. 3
  12. 12.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In CVPR, pages 15750–15758, 2021. 1, 3, 7
  13. 13.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV, pages 9640–9649, 2021. 3, 4, 6, 7, 11
  14. 14.ImageNet contributors. Imagenet terms of access. https://image-net.org/download, 2020. 13
  15. 15.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020. 12
  16. 16.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009. 2, 6
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 3
  18. 18.Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Peco: Perceptual codebook for bert pre-training of vision transformers. arXiv preprint arXiv:2111.12710, 2021. 3
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 3, 6, 7, 11
  20. 20.Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Hervé Jegou, and Edouard Grave. Are large-scale datasets necessary for self-supervised pre-training? arXiv preprint arXiv:2112.10740, 2021. 3
  21. 21.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. In NeurIPS, volume 33, pages 21271–21284, 2020. 1, 3
  22. 22.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, pages 5356–5364, 2019. 2, 6, 12
  23. 23.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021. 1, 3, 4, 5, 6, 7, 11
  24. 24.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, pages 9729–9738, 2020. 1, 3, 4, 7
  25. 25.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 3
  26. 26.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In ICCV, pages 8340–8349, 2021. 2, 6, 7, 13
  27. 27.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261, 2019. 2, 6, 7, 13
  28. 28.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In CVPR, pages 15262–15271, 2021. 2, 6, 7, 13
  29. 29.Tianyu Hua, Wenxiao Wang, Zihui Xue, Sucheng Ren, Yue Wang, and Hang Zhao. On feature decorrelation in self-supervised learning. In ICCV, pages 9598–9608, 2021. 3
  30. 30.Zhicheng Huang, Xiaojie Jin, Chengze Lu, Qibin Hou, Ming-Ming Cheng, Dongmei Fu, Xiaohui Shen, and Jiashi Feng. Contrastive masked autoencoders are stronger vision learners. arXiv preprint arXiv:2207.13532, 2022. 4, 6
  31. 31.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, pages 448–456. PMLR, 2015. 6, 7, 11
  32. 32.Longlong Jing and Yingli Tian. Self-supervised visual feature learning with deep neural networks: A survey. TPAMI, 43(11):4037–4058, 2020. 1
  33. 33.Chunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao, Bin Xiao, Xiyang Dai, Lu Yuan, and Jianfeng Gao. Efficient self-supervised vision transformers for representation learning. arXiv preprint arXiv:2106.09785, 2021. 3
  34. 34.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. arXiv preprint arXiv:2203.16527, 2022. 11, 12
  35. 35.Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. Benchmarking detection transfer learning with vision transformers. arXiv preprint arXiv:2111.11429, 2021. 1
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 2, 6, 11
  37. 37.Songtao Liu, Zeming Li, and Jian Sun. Self-emd: Self-supervised object detection without imagenet. arXiv preprint arXiv:2011.13677, 2020. 3
  38. 38.Julien Mairal. Cyanure: An open-source toolbox for empirical risk minimization for python, c++, and soon more. arXiv preprint arXiv:1912.08165, 2019. 11
  39. 39.Shlok Mishra, Joshua Robinson, Huiwen Chang, David Jacobs, Aaron Sarna, Aaron Maschinot, and Dilip Krishnan. A simple, efficient and scalable contrastive masked autoencoder for learning visual representations. arXiv preprint arXiv:2210.16870, 2022. 4, 6
  40. 40.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 3
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 3
  42. 42.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, pages 8821–8831. PMLR, 2021. 3
  43. 43.Chenxin Tao, Honghui Wang, Xizhou Zhu, Jiahua Dong, Shiji Song, Gao Huang, and Jifeng Dai. Exploring the equivalence of siamese self-supervised learning via a unified gradient framework. arXiv preprint arXiv:2112.05141, 2021. 1, 3, 5, 6
  44. 44.Nenad Tomasev, Ioana Bica, Brian McWilliams, Lars Buesing, Razvan Pascanu, Charles Blundell, and Jovana Mitrovic. Pushing the limits of self-supervised resnets: Can we outperform supervised learning without labels on imagenet? arXiv preprint arXiv:2201.05119, 2022. 3
  45. 45.Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv e-prints, pages arXiv–1807, 2018. 3, 6
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, volume 30, 2017. 5, 6
  47. 47.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. NeurIPS, 32, 2019. 2, 6, 7, 13
  48. 48.Luya Wang, Feng Liang, Yangguang Li, Wanli Ouyang, Honggang Zhang, and Jing Shao. Repre: Improving self-supervised vision transformer with reconstructive pre-training. arXiv preprint arXiv:2201.06857, 2022. 4, 6
  49. 49.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Dense contrastive learning for self-supervised visual pre-training. In CVPR, pages 3024–3033, 2021. 3, 7
  50. 50.Yizhou Wang, Shixiang Tang, Feng Zhu, Lei Bai, Rui Zhao, Donglian Qi, and Wanli Ouyang. Revisiting the transferability of supervised pretraining: an mlp perspective. arXiv preprint arXiv:2112.00496, 2021. 5
  51. 51.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. arXiv preprint arXiv:2112.09133, 2021. 3
  52. 52.Fangyun Wei, Yue Gao, Zhirong Wu, Han Hu, and Stephen Lin. Aligning pretraining for detection via object-level contrastive learning. NeurIPS, 34:22682–22694, 2021. 3
  53. 53.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, pages 3733–3742, 2018. 3
  54. 54.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, pages 418–434, 2018. 12
  55. 55.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In CVPR, pages 16684–16693, 2021. 3
  56. 56.Yuwen Xiong, Mengye Ren, Wenyuan Zeng, and Raquel Urtasun. Self-supervised representation learning from flow equivariance. In ICCV, pages 10191–10200, 2021. 3
  57. 57.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. In ICML, pages 12310–12320. PMLR, 2021. 1, 3
  58. 58.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, pages 633–641, 2017. 2, 6
  59. 59.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. arXiv preprint arXiv:2111.07832, 2021. 4, 6

Citation

MLA
Tao, C., et al. “Siamese Image Modeling for Self-Supervised Vision Representation Learning”. arXiv, 2022, http://arxiv.org/abs/2206.01204v3.
APA
Tao, C., Zhu, X., Su, W., Huang, G., Li, B., Zhou, J., Qiao, Y., Wang, X., & Dai, J. (2022). Siamese Image Modeling for Self-Supervised Vision Representation Learning. arXiv. http://arxiv.org/abs/2206.01204v3
Chicago
Tao, C., X. Zhu, W. Su, et al. 2022. “Siamese Image Modeling for Self-Supervised Vision Representation Learning”. arXiv. http://arxiv.org/abs/2206.01204v3.
Harvard
Tao, C. et al. (2022) “Siamese Image Modeling for Self-Supervised Vision Representation Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.01204v3.
Vancouver
1. Tao C, Zhu X, Su W, Huang G, Li B, Zhou J, Qiao Y, Wang X, Dai J (2022) Siamese Image Modeling for Self-Supervised Vision Representation Learning. arXiv

BibTeX

@article{tao2022siamese,
  title = {Siamese Image Modeling for Self-Supervised Vision Representation Learning},
  author = {Tao, Chenxin and Zhu, Xizhou and Su, Weijie and Huang, Gao and Li, Bin and Zhou, Jie and Qiao, Yu and Wang, Xiaogang and Dai, Jifeng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.01204v3},
  eprint = {2206.01204}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE