Learning Where to Learn in Cross-View Self-Supervised Learning

Lang HuangShan YouMingkai ZhengFei WangChen QianToshihiko Yamasaki

article2022CVPR51 citationsCVPR 2022 Best Paper Finalist

Proposes an adaptive self-supervised learning framework called LEWEL that reinterprets standard projection heads as per-pixel predictors to dynamically generate spatial alignment maps, correcting cross-view spatial misalignment and improving performance across image classification, object detection, and semantic segmentation.

Listen

Self-supervised visual representation learning allows artificial intelligence models to learn from massive amounts of unlabeled imagery, reducing reliance on costly manual data annotation. Standard contrastive techniques train models by comparing differently transformed crops of the same image. However, uniform pixel averaging across an image risks capturing distracting background noise and causes spatial misalignment across different views. Prior efforts to address this issue relied on rigid, task-specific geometric rules, which improved localized spatial detection but degraded overall image classification performance.

The main objective of the article is to demonstrate that an adaptive, end-to-end framework—termed Learning Where to Learn (LEWEL)—can automatically predict spatial alignment maps to align visual features dynamically, thereby enhancing both global image-level recognition and dense, localized prediction tasks simultaneously.

The authors implemented their approach by reinterpreting the model's standard global projection head as a per-pixel projector. This design generates dynamic spatial alignment heat-maps directly from visual features during training without requiring external supervision or handcrafted rules. The framework was evaluated across standard computer vision benchmarks—including ImageNet-1K, Pascal VOC, and MS-COCO—using a standard ResNet-50 architecture across linear classification, semi-supervised classification, object detection, and semantic segmentation protocols.

The findings confirm clear performance advantages across all evaluated domains. First, LEWEL consistently outperformed leading baselines, improving top-1 linear classification on ImageNet by 1.6 percentage points over MoCoV2 and 1.3 percentage points over BYOL. Second, in low-data regimes with only 1% labeled data, LEWEL achieved a 56.1% top-1 accuracy, outperforming prior methods by up to 1.3 percentage points and surpassing models trained for more than twice as many pre-training epochs. Third, in transfer learning benchmarks, LEWEL improved Pascal VOC object detection and semantic segmentation over baseline models by up to 0.6 and 0.7 percentage points, respectively. Finally, LEWEL delivered the top performance on MS-COCO object detection and instance segmentation, proving that performance gains stem from the adaptive alignment mechanism rather than increased model parameter size.

These results demonstrate that visual AI systems do not have to compromise between high-level image classification and fine-grained spatial localization. Because LEWEL attains superior accuracy in substantially fewer training cycles, it offers significant efficiency benefits, lowering computational costs and reducing model training timelines for production environments.

Organizations developing or deploying self-supervised computer vision models should consider integrating adaptive spatial alignment into their pre-training pipelines. For immediate adoption, teams can apply the default balanced weighting between global and aligned training objectives. Future work should evaluate the framework's scalability on larger vision transformer architectures and across diverse, specialized image domains beyond standard natural image benchmarks.

While confidence in the reported experimental results is high given the extensive comparative benchmarking, findings are primarily bounded by the standard ResNet-50 backbone and standard public datasets. Stakeholders should validate performance on proprietary or domain-specific datasets before large-scale production deployment.

arXiv: 2203.14898
Cover for Learning Where to Learn in Cross-View Self-Supervised Learning

Abstract

Self-supervised learning (SSL) has made enormous progress and largely narrowed the gap with the supervised ones, where the representation learning is mainly guided by a projection into an embedding space. During the projection, current methods simply adopt uniform aggregation of pixels for embedding; however, this risks involving object-irrelevant nuisances and spatial misalignment for different augmentations. In this paper, we present a new approach, Learning Where to Learn (LEWEL), to adaptively aggregate spatial information of features, so that the projected embeddings could be exactly aligned and thus guide the feature learning better. Concretely, we reinterpret the projection head in SSL as a per-pixel projection and predict a set of spatial alignment maps from the original features by this weight-sharing projection head. A spectrum of aligned embeddings is thus obtained by aggregating the features with spatial weighting according to these alignment maps. As a result of this adaptive alignment, we observe substantial improvements on both image-level prediction and dense prediction at the same time: LEWEL improves MoCov2 [15] by 1.6%/1.3%/0.5%/0.4% points, improves BYOL [14] by 1.3%/1.3%/0.7%/0.6% points, on ImageNet linear/semi-supervised classification, Pascal VOC semantic segmentation, and object detection, respectively.†

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Learning Where to Learn
  • 3.1 Generalized spatial aggregation
  • 3.2 Reinterpreting and coupling projection head
  • 3.3 Implementations
  • 3.4 Discussions
  • 4 Experiments
  • 4.1 Linear evaluation
  • 4.2 Semi-supervised classification
  • 4.3 Transfer learning to other tasks
  • 4.4 Ablations
  • 5 Conclusion
  • References
  • A Additional Illustration of LEWEL
  • A.1 Illustration of the channel grouping operation.
  • B Additional Experiment Results
  • B.1 Comparison with hand-crafted spatial alignment methods
  • C Additional Analyses on LEWEL
  • C.1 Visualization of the alignment maps.
  • C.2 Computational cost.
  • C.3 Limitations.

Knowls

  1. Knowl 1 — Learning Where to Learn (LEWEL) Framework and Generalized Spatial Aggregation

    model/method

    The Learning Where to Learn (LEWEL) framework adaptively aligns spatial representations across augmented views in cross-view self-supervised learning (SSL) without requiring task-specific spatial priors or bounding-box annotations.

    In standard SSL, global average pooling (GAP) aggregates a spatial feature map F′∈RD×H×WF' \in \mathbb{R}^{D \times H \times W} into a single vector y′=1HW∑i=1H∑j=1WF∗,i,j′∈RDy' = \frac{1}{HW}\sum_{i=1}^H \sum_{j=1}^W F'_{*, i, j} \in \mathbb{R}^D, which discards fine spatial information and introduces background nuisances. LEWEL generalizes spatial aggregation into an adaptive spatial weighting:

    y′=W′⊗F′=∑i=1H∑j=1WWi,j′F∗,i,j′y' = W' \otimes F' = \sum_{i=1}^H \sum_{j=1}^W W'_{i,j} F'_{*,i,j}

    where W′∈RH×WW' \in \mathbb{R}^{H \times W} is a non-negative spatial alignment map satisfying ∑i=1H∑j=1WWi,j′=1\sum_{i=1}^H \sum_{j=1}^W W'_{i,j} = 1 and Wi,j′≥0W'_{i,j} \ge 0. GAP is the uniform special case where Wi,j′=1HWW'_{i,j} = \frac{1}{HW}.

    To generate W′W' without supervision, LEWEL reinterprets the standard global projection head gθg_\theta as a per-pixel 1×11 \times 1 projection applied directly to the pre-pooling backbone feature F′F', producing unnormalized semantic heatmaps W~′=gθ(F′)∈Rd×H×W\widetilde{W}' = g_\theta(F') \in \mathbb{R}^{d \times H \times W}. Weight-sharing between the global projection head and alignment map prediction couples global invariant learning with local region alignment.

  2. Knowl 2 — Two-Stage Spatial Alignment Map Normalization in LEWEL

    equation

    Given unnormalized heatmaps W~′=gθ(F′)∈Rd×H×W\widetilde{W}' = g_\theta(F') \in \mathbb{R}^{d \times H \times W} obtained by applying the per-pixel projection head gθg_\theta to the spatial backbone feature map F′∈RD×H×WF' \in \mathbb{R}^{D \times H \times W}, LEWEL normalizes activations into valid spatial alignment maps using a two-stage normalization procedure:

    1. Per-pixel channel-wise ℓ2\ell_2-normalization:

    W∗,i,j′=W~∗,i,j′∥W~∗,i,j′∥2,∀i∈{1,…,H}, j∈{1,…,W}W'_{*, i, j} = \frac{\widetilde{W}'_{*, i, j}}{\|\widetilde{W}'_{*, i, j}\|_2}, \quad \forall i \in \{1, \dots, H\}, \, j \in \{1, \dots, W\}

    1. Per-channel spatial 2D softmax normalization across all spatial coordinates for each semantic channel k∈{1,…,d}k \in \{1, \dots, d\}:

    Wk,i,j′=exp⁡(Wk,i,j′)∑u=1H∑v=1Wexp⁡(Wk,u,v′),∀i∈{1,…,H}, j∈{1,…,W}W'_{k, i, j} = \frac{\exp(W'_{k, i, j})}{\sum_{u=1}^H \sum_{v=1}^W \exp(W'_{k, u, v})}, \quad \forall i \in \{1, \dots, H\}, \, j \in \{1, \dots, W\}

    This produces a collection of dd discrete spatial probability distributions {Wk′}k=1d\{W'_k\}_{k=1}^d, where each Wk′∈RH×WW'_k \in \mathbb{R}^{H \times W} integrates to 1 across the spatial grid.

  3. Knowl 3 — Channel Grouping and Aligned Embedding Extraction

    model/method

    To control the number of aligned embeddings and enrich the semantics encoded in each representation, LEWEL partitions the channel dimension DD of feature map F′∈RD×H×WF' \in \mathbb{R}^{D \times H \times W} into hh uniform, non-overlapping groups:

    F′=[F′(1),F′(2),…,F′(h)],F′(m)∈R(D/h)×H×WF' = [F'^{(1)}, F'^{(2)}, \dots, F'^{(h)}], \quad F'^{(m)} \in \mathbb{R}^{(D/h) \times H \times W}

    Using the dd spatial alignment maps {Wk′}k=1d\{W'_k\}_{k=1}^d, LEWEL computes d/hd/h aligned representations {yk′∈RD}k=1d/h\{y'_k \in \mathbb{R}^D\}_{k=1}^{d/h} via concatenated spatial aggregations:

    yk′=[W(k−1)h+1′⊗F′(1), W(k−1)h+2′⊗F′(2), …, Wkh′⊗F′(h)],∀k∈{1,…,d/h}y'_k = [W'_{(k-1)h + 1} \otimes F'^{(1)}, \, W'_{(k-1)h + 2} \otimes F'^{(2)}, \, \dots, \, W'_{kh} \otimes F'^{(h)}], \quad \forall k \in \{1, \dots, d/h\}

    where ⊗\otimes denotes spatial aggregation W⊗F=∑i=1H∑j=1WWi,jF∗,i,jW \otimes F = \sum_{i=1}^H \sum_{j=1}^W W_{i,j} F_{*, i, j}.

    A separate multi-layer perceptron projector pθp_\theta maps each aligned representation yk′y'_k to an aligned embedding:

    zk′=pθ(yk′)∈Rcz'_k = p_\theta(y'_k) \in \mathbb{R}^c

    where cc denotes the output dimensionality of pθp_\theta.

  4. Knowl 4 — Objective Functions and Instantiations of LEWEL

    model/method

    LEWEL trains representations using a joint objective balancing global image-level loss Lg\mathcal{L}_g and aligned local loss La\mathcal{L}_a:

    L=(1−β)Lg+βLa,where La=hd∑k=1d/hℓ(zk′,zk′′)\mathcal{L} = (1 - \beta)\mathcal{L}_g + \beta \mathcal{L}_a, \quad \text{where } \mathcal{L}_a = \frac{h}{d} \sum_{k=1}^{d/h} \ell(z'_k, z''_k)

    where β∈[0,1]\beta \in [0, 1] is a trade-off hyperparameter (default β=0.5\beta = 0.5), and zk′,zk′′∈Rcz'_k, z''_k \in \mathbb{R}^c are aligned embeddings corresponding to the kk-th semantic group from two transformed views.

    Two specific instantiations are defined:

    1. LEWELM\text{LEWEL}_M (InfoNCE-based): Uses contrastive InfoNCE loss:

    ℓInfoNCE(z′,z′′)=−log⁡exp⁡(sim(z′,z′′)/τ)exp⁡(sim(z′,z′′)/τ)+∑z−exp⁡(sim(z′,z−)/τ)\ell_{\text{InfoNCE}}(z', z'') = -\log \frac{\exp(\text{sim}(z', z'') / \tau)}{\exp(\text{sim}(z', z'') / \tau) + \sum_{z^-} \exp(\text{sim}(z', z^-) / \tau)}

    where sim(u,v)=uTv∥u∥2∥v∥2\text{sim}(u, v) = \frac{u^T v}{\|u\|_2 \|v\|_2} is cosine similarity and τ\tau is temperature. Negative samples z−z^- for Lg\mathcal{L}_g are retrieved from a first-in-first-out memory queue, while negative samples for La\mathcal{L}_a are drawn from other images in the current mini-batch.

    1. LEWELB\text{LEWEL}_B (MSE/BYOL-based): Uses normalized Mean Square Error without negative samples:

    ℓMSE(z′,z′′)=2−2×sim(qθ(z′),sg(z′′))\ell_{\text{MSE}}(z', z'') = 2 - 2 \times \text{sim}(q_\theta(z'), \text{sg}(z''))

    where sg\text{sg} is the stop-gradient operation, qθq_\theta is an extra prediction MLP for the global embedding, and an independent predictor sθs_\theta is applied to the online aligned embeddings zk′z'_k.

  5. Knowl 5 — Linear Classification Evaluation of LEWEL on ImageNet-1K

    data/table

    The linear evaluation protocol tests features by training a linear classifier atop fixed ResNet-50 backbones pre-trained on ImageNet-1K (IN-1K). LEWELM\text{LEWEL}_M and LEWELB\text{LEWEL}_B consistently improve upon their baseline counterparts MoCov2 and BYOL across various training epoch budgets.

    Method 100 Epochs 200 Epochs 400 Epochs
    Acc@1 Acc@5 Acc@1 Acc@5 Acc@1
    InstDisc - - 56.5 - -
    PCL - - 67.6 - -
    SimCLR 64.6 - 66.6 - -
    BYOL 66.5 - 70.6 - 73.2
    SwAV 66.5 - 69.1 - 70.7
    SimSiam 68.1 - 70.0 - 70.8
    MoCov2 64.5 86.1 67.5 88.1 -
    BYOL (reproduced) 70.6 89.9 71.9 90.4 -
    LEWELM_M 66.1 87.2 68.4 88.6 -
    LEWELB_B 71.9 90.5 72.8 91.0 73.8

    At 100 epochs, LEWELM\text{LEWEL}_M gains +1.6% Top-1 over MoCov2, and LEWELB\text{LEWEL}_B gains +1.3% Top-1 over BYOL. At 400 epochs, LEWELB\text{LEWEL}_B achieves 73.8% Top-1 accuracy.

  6. Knowl 6 — Semi-Supervised Fine-Tuning Performance on ImageNet-1K

    data/table

    ResNet-50 models pre-trained on ImageNet-1K are fine-tuned on standard 1% and 10% labeled subsets of ImageNet-1K for 50 epochs. LEWEL variants outperform baseline methods and achieve comparable or superior results with fewer pre-training epochs.

    Method Epochs 1% Labels 10% Labels
    Acc@1 Acc@5 Acc@1 Acc@5
    PCL 200 - 75.3 - 86.5
    MoCov2 200 43.8 72.3 61.9 84.6
    BYOL 200 54.8 78.8 68.0 88.5
    LEWELM_M 200 45.1 71.1 62.5 84.9
    LEWELB_B 200 56.1 79.9 68.7 88.9
    SimCLR 1000 48.3 75.5 65.6 87.8
    SwAV 800 53.9 78.5 70.2 89.9
    BYOL 800 53.2 78.4 68.8 89.0
    Barlow Twins 1000 55.0 79.2 69.7 89.3
    LEWELB_B 400 59.8 83.2 70.4 90.1

    Under 200 pre-training epochs with 1% labels, LEWELB\text{LEWEL}_B reaches 56.1% Top-1 accuracy (+1.3% over BYOL). With 400 epochs, LEWELB\text{LEWEL}_B achieves 59.8% Top-1 accuracy, outperforming 800-epoch BYOL (53.2%) and 1000-epoch Barlow Twins (55.0%).

  7. Knowl 7 — Transfer Learning to Pascal VOC Detection and Semantic Segmentation

    data/table

    Transfer performance is evaluated on Pascal VOC object detection (Faster R-CNN with ResNet-50 C4 backbone, trained on trainval07+12 and evaluated on test12) and semantic segmentation (dilated FCN with output stride 8, trained on VOC 2012 train+aug and evaluated on val).

    Method Epochs VOC 07+12 Det. 12 Seg.
    AP AP50_{50} AP75_{75} mIoU
    Supervised 90 53.5 81.3 58.8 67.7
    MoCov2 100 56.1 81.5 62.4 66.3
    BYOL 100 55.5 81.9 61.2 66.9
    LEWELM_M 100 56.5 82.1 63.0 66.8
    LEWELB_B 100 56.1 82.1 62.3 67.6
    SimCLR 200 55.5 81.8 61.4 -
    SwAV 200 55.4 81.5 61.4 -
    BYOL 200 55.8 81.6 61.6 67.2
    SimSiam 200 56.4 82.0 62.8 -
    MoCov2 200 57.0 82.2 63.4 66.7
    LEWELM_M 200 57.3 82.3 63.6 67.2
    LEWELB_B 200 56.5 82.6 63.7 67.8

    LEWEL consistently boosts dense transfer metrics over base models. At 200 pre-training epochs, LEWELM\text{LEWEL}_M attains 57.3 AP on object detection, and LEWELB\text{LEWEL}_B achieves 67.8 mIoU on semantic segmentation, outperforming supervised pre-training (67.7 mIoU).

  8. Knowl 8 — Transfer Learning to MS-COCO Object Detection and Instance Segmentation

    data/table

    Mask R-CNN models initialized with self-supervised ResNet-50 backbones (C4 and FPN variants) pre-trained on ImageNet-1K are fine-tuned on COCO 2017 train split (1×1\times schedule, 90k iterations) and evaluated on the validation set.

    Method Object Det. Instance Seg.
    AP AP50_{50} AP75_{75} AP AP50_{50} AP75_{75}
    ResNet50-C4 (200 epochs)
    Supervised 38.2 58.2 41.2 33.3 54.7 35.2
    MoCov2 38.8 58.0 42.0 34.0 55.2 36.3
    BYOL 38.1 58.4 40.9 33.3 55.0 35.3
    LEWELM_M 38.9 58.6 42.0 34.1 55.3 36.3
    LEWELB_B 38.5 58.9 41.2 33.7 55.5 35.5
    ResNet50-FPN (200 epochs)
    DenseCL 40.3 59.9 44.3 36.4 57.0 39.2
    ReSim 39.8 60.2 43.5 36.0 57.1 38.6
    LEWELM_M 40.0 59.8 43.7 36.1 57.0 38.7
    LEWELB_B 41.3 61.2 45.4 37.4 58.3 40.3
    ResNet50-FPN (400 epochs)
    PixelPro 41.4 61.6 45.4 37.4 - -
    LEWELB_B 41.9 62.4 46.0 37.9 59.3 40.7

    With ResNet50-FPN at 200 epochs, LEWELB\text{LEWEL}_B yields 41.3 box AP and 37.4 mask AP, outperforming dense contrastive methods that rely on spatial heuristics (DenseCL at 40.3 AP / 36.4 mask AP; ReSim at 39.8 AP / 36.0 mask AP). At 400 epochs, LEWELB\text{LEWEL}_B attains 41.9 box AP and 37.9 mask AP.

  9. Knowl 9 — Ablation of Component Contributions and Coupled Projection Head

    data/table

    Ablation experiments on ImageNet-100 (IN-100) linear classification accuracy and Pascal VOC 2012 semantic segmentation mIoU isolate the effects of the aligned loss La\mathcal{L}_a, coupled weight sharing in projector gθg_\theta, and parameter scaling.

    Method Global Lg\mathcal{L}_g Align La\mathcal{L}_a Coupled Head IN-100 Acc. VOC Seg. mIoU
    MoCov2 baseline ✓ × × 79.5 61.6
    Aligned only × ✓ × 80.0 62.6
    Uncoupled joint ✓ ✓ × 81.0 62.7
    LEWELM_M ✓ ✓ ✓ 82.1 63.4
    MoCov2 (d=256d=256) - - - 79.8 60.6
    MoCov2 (d=51d=512) - - - 80.2 61.5
    MoCov2 (+ extra pθp_\theta) - - - 79.9 62.2
    LEWELM_M w/ rand. W′W' - - - 79.8 61.9

    Key observations:

    1. The coupled projection head design provides an additional +1.1% linear accuracy and +0.7% mIoU improvement over the uncoupled joint formulation.
    2. Expanding MoCov2 projection head dimension or adding an extra projection head does not replicate these gains (reaching at most 80.2% linear accuracy and 62.2% mIoU).
    3. Replacing adaptive alignment maps W′W' with random maps causes accuracy to fall to 79.8% and mIoU to 61.9%, proving that the gains stem from semantic spatial alignment rather than extra regularization or parameters.
  10. Knowl 10 — Impact of Group Count $h$, Alignment Dimensionality $c$, and Loss Weight $\beta$

    empirical result

    Ablation experiments on ImageNet-100 and Pascal VOC 2012 reveal specific operational trade-offs for LEWEL hyperparameters:

    1. Channel Grouping Number (hh) vs. Output Dimension (dd): For LEWELB\text{LEWEL}_B with d=256d = 256, varying hh across {1,2,4,8,16}\{1, 2, 4, 8, 16\} changes the number of aligned embeddings (d/hd/h). Linear classification accuracy increases from 81.0% (h=1h=1) to a peak of 83.3% (h=4h=4), then slightly declines to 82.9% (h=16h=16). Conversely, semantic segmentation mIoU follows the opposite trend, dropping from 63.2% (h=1h=1) and 63.4% (h=4h=4) down to 61.4% (h=16h=16). This indicates that a smaller number of representations with grouped channels favors image-level global classification, whereas more fine-grained aligned representations benefit local spatial tasks.

    2. Aligned Embedding Dimensionality (cc): Varying c∈{8,32,64,128,256}c \in \{8, 32, 64, 128, 256\} in LEWELM\text{LEWEL}_M while fixing gθg_\theta shows that linear accuracy and segmentation mIoU improve sharply from c=8c=8 to c=64c=64 (reaching ∼82.1%\sim 82.1\% accuracy and ∼63.4%\sim 63.4\% mIoU) and plateau thereafter, indicating that modest dimensional capacity per semantic is sufficient to guide representation learning.

    3. Loss Weight (β\beta): When β≤0.3\beta \le 0.3, performance degrades significantly (<80.5%< 80.5\% linear accuracy, <62.5%< 62.5\% mIoU). Performance remains high and stable across β∈[0.5,0.9]\beta \in [0.5, 0.9], but drops when β=1.0\beta = 1.0 (aligned loss only), demonstrating that the global and aligned objectives mutually reinforce each other.

Coverage note — None was omitted; all primary methodological components, loss functions, architecture details, main benchmark results (linear eval, semi-supervised, VOC det/seg, COCO det/seg), and ablation studies are fully covered.

References

  1. 1.Yuki M. Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. In International Conference on Learning Representations, 2020.
  2. 2.Zhaowei Cai, Avinash Ravichandran, Subhransu Maji, Charless Fowlkes, Zhuowen Tu, and Stefano Soatto. Exponential moving average normalization for self-supervised and semi-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 194–203, 2021.
  3. 3.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In European Conference on Computer Vision, pages 132–149, 2018.
  4. 4.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems, 33, 2020.
  5. 5.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J  egou,  Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017.
  7. 7.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning, 2020.
  8. 8.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020.
  9. 9.Xinlei Chen and Kaiming He. Exploring simple Siamese representation learning. arXiv preprint arXiv:2011.10566, 2020.
  10. 10.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https : / / github . com / open -mmlab/mmsegmentation, 2020.
  11. 11.M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, Jan. 2015.
  12. 12.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. In International Conference on Learning Representation, 2018.
  13. 13.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672–2680, 2014.
  14. 14.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin  Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. Advances in Neural Information Processing Systems, 33, 2020.
  15. 15.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In IEEE Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020.
  16. 16.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir-  shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 37(9):1904–1916, 2015.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  19. 19.Olivier J Henaff, Skanda Koppula, Jean-Baptiste Alayrac,  Aaron van den Oord, Oriol Vinyals, and Joao Carreira. Ef-  ficient visual pretraining with contrastive detection. arXiv preprint arXiv:2103.10957, 2021.
  20. 20.Lang Huang, Chao Zhang, and Hongyang Zhang. Self-adaptive training: Bridging supervised and self-supervised learning. arXiv preprint arXiv:2101.08732, 2021.
  21. 21.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning, pages 448–456, 2015.
  22. 22.Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations, 2014.
  23. 23.Junnan Li, Pan Zhou, Caiming Xiong, Richard Socher, and Steven CH Hoi. Prototypical contrastive learning of unsupervised representations. In International Conference on Learning Representations, 2020.
  24. 24.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He,  Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  25. 25.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence  Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  26. 26.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015.
  27. 27.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations, 2017.
  28. 28.Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted Boltzmann machines. In International Conference on Machine Learning, 2010.
  29. 29.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In European Conference on Computer Vision, pages 69–84. Springer, 2016.
  30. 30.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  31. 31.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, pages 8024–8035, 2019.
  32. 32.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2536–2544, 2016.
  33. 33.Pedro O Pinheiro, Amjad Almahairi, Ryan Y Benmalek, Florian Golemo, and Aaron Courville. Unsupervised learning of dense visual representations. arXiv preprint arXiv:2011.05499, 2020.
  34. 34.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28:91–99, 2015.
  35. 35.Byungseok Roh, Wuhyun Shin, Ildoo Kim, and Sungwoong Kim. Spatially consistent representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1144–1153, 2021.
  36. 36.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  37. 37.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In European Conference on Computer Vision, 2020.
  38. 38.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning? Advances in Neural Information Processing Systems, 33:6827–6839, 2020.
  39. 39.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In International Conference on Machine Learning, pages 1096–1103, 2008.
  40. 40.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Dense contrastive learning for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3024–3033, 2021.
  41. 41.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019.
  42. 42.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3733–3742, 2018.
  43. 43.Tete Xiao, Colorado J Reed, Xiaolong Wang, Kurt Keutzer, and Trevor Darrell. Region similarity representation learning. arXiv preprint arXiv:2103.12902, 2021.
  44. 44.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16684–16693, 2021.
  45. 45.Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888, 2017.
  46. 46.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stephane Deny. Barlow twins: Self-supervised learning via  redundancy reduction. arXiv preprint arXiv:2103.03230, 2021.
  47. 47.Mingkai Zheng, Fei Wang, Shan You, Chen Qian, Changshui Zhang, Xiaogang Wang, and Chang Xu. Weakly supervised contrastive learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10042–10051, October 2021.
  48. 48.Mingkai Zheng, Shan You, Fei Wang, Chen Qian, Changshui Zhang, Xiaogang Wang, and Chang Xu. Ressl: Relational self-supervised learning with weak augmentation. Advances in Neural Information Processing Systems, 34, 2021.

Citation

MLA
Huang, L., et al. “Learning Where to Learn in Cross-View Self-Supervised Learning”. arXiv, 2022, http://arxiv.org/abs/2203.14898v1.
APA
Huang, L., You, S., Zheng, M., Wang, F., Qian, C., & Yamasaki, T. (2022). Learning Where to Learn in Cross-View Self-Supervised Learning. arXiv. http://arxiv.org/abs/2203.14898v1
Chicago
Huang, L., S. You, M. Zheng, F. Wang, C. Qian, and T. Yamasaki. 2022. “Learning Where to Learn in Cross-View Self-Supervised Learning”. arXiv. http://arxiv.org/abs/2203.14898v1.
Harvard
Huang, L. et al. (2022) “Learning Where to Learn in Cross-View Self-Supervised Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14898v1.
Vancouver
1. Huang L, You S, Zheng M, Wang F, Qian C, Yamasaki T (2022) Learning Where to Learn in Cross-View Self-Supervised Learning. arXiv

BibTeX

@article{huang2022learning,
  title = {Learning Where to Learn in Cross-View Self-Supervised Learning},
  author = {Huang, Lang and You, Shan and Zheng, Mingkai and Wang, Fei and Qian, Chen and Yamasaki, Toshihiko},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14898v1},
  eprint = {2203.14898}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE