Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning

Zhiqiang ShenZechun LiuZhuang LiuMarios SavvidesTrevor DarrellEric Poe Xing

article2022AAAI124 citations

Proposes Un-Mix, a simple plug-and-play image mixture and soft-label assignment strategy that prevents overconfident representations in self-supervised frameworks like SimCLR, MoCo, and BYOL to consistently boost downstream visual accuracy across multiple benchmarks.

Listen

Modern computer vision increasingly relies on unsupervised representation learning to train artificial intelligence models on vast amounts of unlabeled visual data without human annotation. Many leading frameworks train networks by comparing two transformed views of the same image. However, these systems often suffer from overfitting and over-confident predictions because typical training methods treat image pairs strictly as binary matches or non-matches. This rigid approach prevents the underlying models from learning subtle, fine-grained visual differences and smooth decision boundaries.

The article demonstrates that introducing soft similarity distances by mixing input images and softening target labels—a framework termed Unsupervised Image Mixtures (Un-Mix)—substantially improves model robustness, generalization, and visual representation quality.

The authors evaluated this concept by integrating image mixtures into five mainstream unsupervised learning architectures, including SimCLR, MoCo V1 and V2, BYOL, SwAV, and Whitening. The approach blends images within training mini-batches globally or regionally and adjusts the training loss functions proportionally to the mixture ratios. Extensive empirical tests were conducted across standard benchmarks, including CIFAR-10, CIFAR-100, STL-10, Tiny ImageNet, and ImageNet-1K, using ResNet backbones under fixed training hyperparameters.

The key findings show consistent performance gains across all evaluated settings. First, integrating Un-Mix yielded consistent linear classification improvements of 1% to 3% across small and medium benchmarks without altering baseline hyperparameters. Second, on the large-scale ImageNet-1K benchmark, combining Un-Mix with MoCo V2 increased top-1 accuracy by 1.1% over standard training, and by 2.3% when combined with multi-scale training. Third, models pre-trained with Un-Mix demonstrated superior transferability on downstream object detection tasks on PASCAL VOC and COCO. Notably, a model trained with Un-Mix for only 200 epochs outperformed a standard MoCo V2 baseline trained for 800 epochs on PASCAL VOC detection accuracy.

These findings indicate that simultaneously regularizing the input data and target label spaces prevents over-fitting and forces neural networks to capture richer visual features. For practitioners and engineering leaders, this method enhances model performance and transferability while requiring only a few lines of code to implement in existing training pipelines. It also delivers substantial efficiency gains, achieving higher downstream performance in significantly fewer training cycles.

The authors recommend adopting image mixtures within self-supervised training workflows, prioritizing region-level mixing for large-scale datasets like ImageNet and balanced mixtures for smaller datasets. The primary operational trade-off is an additional computational forward pass per training iteration, though the extra cost remains below one-third because reverse representations share the forward pass without requiring extra back-propagation passes. Overall confidence in the empirical results is high due to consistent verification across diverse architectures and datasets, though further fine-tuning of baseline hyperparameters could yield even larger performance improvements.

Cover for Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning

Abstract

The recently advanced unsupervised learning approaches use the siamese-like framework to compare two “views” from the same image for learning representations. Making the two views distinctive is a core to guarantee that unsupervised methods can learn meaningful information. However, such frameworks are sometimes fragile on overfitting if the augmentations used for generating two views are not strong enough, causing the over-confident issue on the training data. This drawback hinders the model from learning subtle variance and fine-grained information. To address this, in this work we aim to involve the soft distance concept on label space in the contrastive-based unsupervised learning task and let the model be aware of the soft degree of similarity between positive or negative pairs through mixing the input data space, to further work collaboratively for the input and loss spaces. Despite its conceptual simplicity, we show empirically that with the solution – Unsupervised image mixtures (Un-Mix), we can learn subtler, more robust and generalized representations from the transformed input and corresponding new label space. Extensive experiments are conducted on CIFAR-10, CIFAR-100, STL-10, Tiny ImageNet and standard ImageNet-1K with popular unsupervised methods SimCLR, BYOL, MoCo V1&V2, SwAV, etc. Our proposed image mixture and label assignment strategy can obtain consistent improvement by 1~3% following exactly the same hyperparameters and training procedures of the base methods. Code is publicly available at https://github.com/szq0214/Un-Mix.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Paradigms of Mixtures
  • 3.2. Image Mixture Strategies
  • 3.3. Loss Functions
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Baseline Approaches
  • 4.3. Implementation Details in Pre-training
  • 4.4. Linear Classification
  • 4.5. Downstream Tasks
  • 4.6. Visualization and Analysis
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Un-Mix couples image mixing with soft unsupervised targets

    model/method

    Un-Mix is a training modification for siamese-like unsupervised visual representation learning. Given two augmented views of an image, the main Un-Mix variant mixes only one view with another image from the current mini-batch and compares the mixed view with the corresponding unmixed view. The mixture proportion is used to soften the target distance or similarity: a mixed image is treated as partially associated with each source image rather than as an unmodified positive or negative instance.

    This differs from ordinary data augmentation, which changes the visual distance between views while leaving the training target unchanged. Un-Mix changes both the input space and the loss-space target, thereby exposing the learner to intermediate similarity levels. The same strategy can be applied to contrastive and non-contrastive methods, with or without memory banks.

    The authors also examine mixing both views, but this preserves the similarity between the two mixed views without fully using the mixture ratio in the loss. It is useful mainly on smaller datasets and gives little benefit on ImageNet-1K. Mixing only one view is the main strategy because it is more effective and requires one additional forward pass; the normal and reverse mixture orders can share that forward computation and require no extra back-propagation, making the reported extra cost less than one-third of the original training cost.

  2. Knowl 2 — Global and region-level image mixture operators

    equation

    Un-Mix uses two input-space mixture operators. For two images I1I_1 and I2I_2 with identical dimensions and pixel values, global Mixup forms an image-wide convex combination:

    Img=αI1+(1−α)I2,I_m^g = \alpha I_1 + (1-\alpha)I_2,

    where α∈[0,1]\alpha\in[0,1] is the global mixture coefficient and ImgI_m^g is the mixed image. Region-level CutMix uses a binary mask MbM_b with the same spatial dimensions as the images:

    Imr=Mb⊙I1+(1−Mb)⊙I2,I_m^r = M_b\odot I_1 + (\mathbf{1}-M_b)\odot I_2,

    where Mb∈{0,1}H×WM_b\in\{0,1\}^{H\times W}, HH and WW are the image height and width, 1\mathbf{1} is an all-one mask, and ⊙\odot denotes element-wise multiplication. The fraction of the selected region originating from I1I_1 supplies the corresponding mixture coefficient λ\lambda; for global mixing, λ=α\lambda=\alpha.

    Both operators act as regularizers that expose the network to intermediate images and encourage less-confident predictions, but Un-Mix additionally assigns a corresponding soft unsupervised target.

  3. Knowl 3 — Self-mixture construction within a mini-batch

    algorithm

    The main Un-Mix procedure uses only images from the current mini-batch, avoiding a separate image-sampling mechanism or a second memory bank. Let a mini-batch contain BB images, each with an unmixed target view. Pair the first image with the last, the second with the penultimate, and so on. Let PP be the probability of selecting global mixing, and sample a mixture coefficient from Beta⁡(γ,γ)\operatorname{Beta}(\gamma,\gamma).

    Input: Mini-batch of image views x1,…,xBx_1,\ldots,x_B, target views y1,…,yBy_1,\ldots,y_B, global-mixture probability PP, beta parameter γ\gamma
    Output: Original loss and mixture-weighted loss terms
    Sample λ∼Beta⁡(γ,γ)\lambda \sim \operatorname{Beta}(\gamma,\gamma)
    Choose global mixing with probability PP; otherwise choose region-level mixing
    for each index ii from 11 to BB do
        Set the paired index j=B+1−ij = B+1-i
        if global mixing was chosen then
            Construct normal mixture qi=λxi+(1−λ)xjq_i = \lambda x_i + (1-\lambda)x_j
            Construct reverse mixture qj=λxj+(1−λ)xiq_j = \lambda x_j + (1-\lambda)x_i
        else
            Construct a binary region mask MM with first-image area fraction λ\lambda
            Construct normal mixture qi=M⊙xi+(1−M)⊙xjq_i = M\odot x_i + (\mathbf{1}-M)\odot x_j
            Construct reverse mixture qj=M⊙xj+(1−M)⊙xiq_j = M\odot x_j + (\mathbf{1}-M)\odot x_i
        end if
    end for
    Forward the mixed batch q1,…,qBq_1,\ldots,q_B through the encoder and projection head
    Use the paired target views yiy_i with mixture weights λ\lambda and 1−λ1-\lambda
    Add the original base-method loss and the weighted normal- and reverse-order mixture losses
    return the combined objective

    The reverse-order mixtures are already obtained by the paired positions in the same mixed batch, so they do not require another independent forward pass. The experiments use γ=1.0\gamma=1.0, making λ\lambda uniform on [0,1][0,1].

  4. Knowl 4 — Soft distance assignment for mixed samples and memory-bank handling

    model/method

    For a mixed view IAM=λI1+(1−λ)I2I_A^M=\lambda I_1+(1-\lambda)I_2 and unmixed views I^1\hat I_1 and I^2\hat I_2 of the source images, Un-Mix assigns the mixed view a soft positive-pair distance according to the source selected by the target view:

    Ddis(IAM,I^A)={λ,if I^A=I^1,1−λ,if I^A=I^2.D_{\mathrm{dis}}(I_A^M,\hat I_A)= \begin{cases} \lambda, & \text{if }\hat I_A=\hat I_1,\\ 1-\lambda, & \text{if }\hat I_A=\hat I_2. \end{cases}

    Here DdisD_{\mathrm{dis}} is the distance scale used by the unsupervised loss, and λ\lambda is the fraction of the mixed image contributed by I1I_1. The same principle applies to a region mixture by using the fraction of the selected region contributed by the first source image.

    In methods without a memory bank, such as positive-pair or in-batch contrastive methods, this soft assignment modifies the positive-pair loss. In instance-classification methods with a memory bank, mixing changes the composition of negative pairs but does not change their original distance labels: original-original, original-mixed, and mixed-mixed negative pairs retain the base method's negative-pair values. Maintaining a memory bank containing representations of unmixed images is sufficient in the authors' experiments, although they report that this choice is not directly applicable to their multi-scale training variant.

  5. Knowl 5 — Mixture-weighted training objective

    equation

    Let Lori\mathcal{L}_{\mathrm{ori}} denote the original loss of the chosen unsupervised method, such as InfoNCE or an ℓ2\ell_2 representation loss. Let IAM(↓)I_A^M(\downarrow) and IAM(↑)I_A^M(\uparrow) denote the normal- and reverse-order mixed queries, and let I^A\hat I_A denote the corresponding unmixed target view. Un-Mix trains with

    Lfinal=Lori+λLm(IAM(↓),I^A)+(1−λ)Lm(IAM(↑),I^A),\mathcal{L}_{\mathrm{final}} = \mathcal{L}_{\mathrm{ori}} +\lambda\mathcal{L}_m\bigl(I_A^M(\downarrow),\hat I_A\bigr) +(1-\lambda)\mathcal{L}_m\bigl(I_A^M(\uparrow),\hat I_A\bigr),

    where Lm\mathcal{L}_m is the base method's loss applied to a mixed query and an unmixed target. If both branches are mixed, the corresponding objective instead adds a mixed-pair term, Lori+Lm(IAM,I^AM)\mathcal{L}_{\mathrm{ori}}+\mathcal{L}_m(I_A^M,\hat I_A^M).

    For a contrastive InfoNCE implementation, the two mixture terms are

    Lm(IAM(↓),I^A)=−log⁡exp⁡(qm⋅k∗/τ)∑i=0Kexp⁡(qm⋅ki/τ),\mathcal{L}_m\bigl(I_A^M(\downarrow),\hat I_A\bigr) =-\log\frac{\exp(q_m\cdot k^*/\tau)}{\sum_{i=0}^{K}\exp(q_m\cdot k_i/\tau)}, Lm(IAM(↑),I^A)=−log⁡exp⁡(qrm⋅k∗/τ)∑i=0Kexp⁡(qrm⋅ki/τ).\mathcal{L}_m\bigl(I_A^M(\uparrow),\hat I_A\bigr) =-\log\frac{\exp(q_{rm}\cdot k^*/\tau)}{\sum_{i=0}^{K}\exp(q_{rm}\cdot k_i/\tau)}.

    Here qmq_m and qrmq_{rm} are normal- and reverse-order mixed query representations, k∗k^* is the positive unmixed key, kik_i are the keys in the candidate set, KK is the number of negative keys, and τ>0\tau>0 is the temperature.

  6. Knowl 6 — Mutual-information rationale for adding mixtures

    theoretical result

    For latent representations zori=fθ1(Iori)z_{\mathrm{ori}}=f_{\theta_1}(I_{\mathrm{ori}}) and zmix=fθ2(Imix)z_{\mathrm{mix}}=f_{\theta_2}(I_{\mathrm{mix}}), the InfoNCE objective with NN candidate samples gives the mutual-information lower bound

    I(zori,zmix)≥log⁡(N)−LN,I(z_{\mathrm{ori}},z_{\mathrm{mix}})\ge \log(N)-\mathcal{L}_N,

    where I(⋅,⋅)I(\cdot,\cdot) is mutual information and LN\mathcal{L}_N is the InfoNCE loss with one positive and N−1N-1 negatives. Minimizing the loss increases this lower bound, while increasing the number of available relationships can also make the lower bound tighter.

    For a dataset with nn images, conventional pairing provides relationships associated with individual images, whereas mixing pairs of images introduces relationships for the (n2)\binom{n}{2} possible image pairs. The authors use this increase in mixture-derived relationships as a theoretical justification for why mixed representations can provide additional information beyond ordinary positive and negative pairs.

  7. Knowl 7 — Mixture-probability and coefficient ablations

    empirical result

    The authors evaluate the two mixture choices using PP, the probability of selecting global mixing rather than region-level mixing, and evaluate the coefficient distribution using λ∼Beta⁡(γ,γ)\lambda\sim\operatorname{Beta}(\gamma,\gamma). On CIFAR-10, the reported accuracy for γ=1.0\gamma=1.0, 0.80.8, and 0.50.5 is respectively 94.20%94.20\%, 94.12%94.12\%, and 93.93%93.93\%, so the experiments use γ=1.0\gamma=1.0, which produces a uniform coefficient on [0,1][0,1].

    On ImageNet-1K, the reported accuracy for P=1.0P=1.0, 0.50.5, and 0.00.0 is respectively 67.6%67.6\%, 68.3%68.3\%, and 68.6%68.6\%. Thus the selected configuration is P=0.5P=0.5 for the non-ImageNet experiments and P=0P=0—region-level mixing only—for ImageNet-1K.

    The non-ImageNet experiments use ResNet-18 and train for 1,000 epochs on CIFAR-10, CIFAR-100, and Tiny ImageNet, and for 2,000 epochs on STL-10. Learning rates are 3×10−33\times10^{-3} for CIFAR-10/100 and 2×10−32\times10^{-3} for Tiny ImageNet and STL-10. Training uses a 500-iteration warm-up and a learning-rate factor of 0.20.2 at 50 and 25 epochs before the end. ImageNet-1K uses the MoCo V2 configuration with ResNet-50, batch size 256 across eight NVIDIA V100 GPUs, γ=1.0\gamma=1.0, and P=0P=0.

  8. Knowl 8 — Consistent gains on CIFAR, STL-10, and Tiny ImageNet

    data/table

    The following results compare the original unsupervised method with its Un-Mix version using ResNet-18, without multi-scale training. Each dataset reports linear-classifier accuracy and 5-nearest-neighbor accuracy; the two entries in each pair are the baseline and Un-Mix results, respectively. The numbers show that Un-Mix improves nearly every baseline and generally yields gains of roughly 1–3 percentage points.

    Could not parse LaTeX table

    The asterisked MoCo values use symmetric loss, 1,000 training epochs, and a 200-epoch k-nearest-neighbor monitor. The strongest relative improvement is obtained by BYOL on CIFAR-100, where linear accuracy rises from 66.60%66.60\% to 71.50%71.50\% and 5-NN accuracy rises from 56.82%56.82\% to 63.83%63.83\%.

  9. Knowl 9 — ImageNet-1K linear evaluation improves over MoCo V2

    data/table

    On standard ImageNet-1K, the authors use ResNet-50 representations and keep the baseline MoCo V2 architecture, training settings, and hyperparameters. Un-Mix improves MoCo V2 at both 200- and 800-epoch budgets, despite not retuning the baseline for mixed samples. Multi-scale training gives an additional gain at 200 epochs.

    Could not parse LaTeX table

    The 200-epoch Un-Mix model reaches 68.6%68.6\% top-1 accuracy versus 67.5%67.5\% for MoCo V2, and the 800-epoch model reaches 71.8%71.8\% versus 71.1%71.1\%. The authors characterize these gains as conservative because all baseline hyperparameters are retained.

  10. Knowl 10 — Improved transfer to PASCAL VOC and COCO detection

    data/table

    The learned representations are evaluated by fine-tuning a ResNet-50 backbone for object detection with the same transfer-learning schedules used for the corresponding MoCo V2 baselines. PASCAL VOC uses Faster R-CNN R50-C4, the trainval07+12 split, VOC test2007 evaluation, and 24,000 fine-tuning iterations. COCO uses Mask R-CNN R50-C4 with the standard 2× schedule, train2017 for training, val2017 for evaluation, and 180,000 iterations.

    Could not parse LaTeX table

    Un-Mix improves every reported MoCo V2 detection metric at the matched training budgets. On VOC, the 200-epoch model improves AP50 by 0.60.6 percentage points and AP by 0.70.7 points; on COCO, it improves AP by 0.30.3 points.

Coverage note — Detailed t-SNE and convolutional-weight visual diagnostics, along with appendix-level multi-scale implementation details, are omitted because they are supplementary analyses rather than load-bearing method or quantitative results; the reported multi-scale results and the method's computational limitation are included.

References

  1. 1.Bachman, P.; Hjelm, R. D.; and Buchwalter, W. 2019. Learning representations by maximizing mutual information across views. In Advances in Neural Information Processing Systems, 15509–15519.
  2. 2.Caron, M.; Bojanowski, P.; Joulin, A.; and Douze, M. 2018. Deep clustering for unsupervised learning of visual features. In Proceedings of the European Conference on Computer Vision (ECCV), 132–149.
  3. 3.Caron, M.; Misra, I.; Mairal, J.; Goyal, P.; Bojanowski, P.; and Joulin, A. 2020. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882.
  4. 4.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020a. A Simple Framework for Contrastive Learning of Visual Representations. arXiv preprint arXiv:2002.05709.
  5. 5.Chen, X.; Fan, H.; Girshick, R.; and He, K. 2020b. Improved Baselines with Momentum Contrastive Learning. arXiv preprint arXiv:2003.04297.
  6. 6.Chen, X.; and He, K. 2020. Exploring Simple Siamese Representation Learning. arXiv preprint arXiv:2011.10566.
  7. 7.Chorowski, J.; and Jaitly, N. 2016. Towards better decoding and language model integration in sequence to sequence models. arXiv preprint arXiv:1612.02695.
  8. 8.Coates, A.; Ng, A.; and Lee, H. 2011. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, 215–223. JMLR Workshop and Conference Proceedings.
  9. 9.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee.
  10. 10.Donahue, J.; Krahenbühl, P.; and Darrell, T. 2016. Adversarial feature learning. arXiv preprint arXiv:1605.09782.
  11. 11.Donahue, J.; and Simonyan, K. 2019. Large scale adversarial representation learning. In Advances in Neural Information Processing Systems, 10541–10551.
  12. 12.Ermolov, A.; Siarohin, A.; Sangineto, E.; and Sebe, N. 2020a. https://github.com/htdt/self-supervised.
  13. 13.Ermolov, A.; Siarohin, A.; Sangineto, E.; and Sebe, N. 2020b. Whitening for self-supervised representation learning. arXiv preprint arXiv:2007.06346.
  14. 14.Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2): 303–338.
  15. 15.Gidaris, S.; Singh, P.; and Komodakis, N. 2018. Unsupervised Representation Learning by Predicting Image Rotations. In ICLR.
  16. 16.Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014. Generative adversarial nets. In Advances in neural information processing systems, 2672–2680.
  17. 17.Grill, J.-B.; Strub, F.; Altche, F.; Tallec, C.; Richemond, P. H.; Buchatskaya, E.; Doersch, C.; Pires, B. A.; Guo, Z. D.; Azar, M. G.; et al. 2020. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733.
  18. 18.Hadsell, R.; Chopra, S.; and LeCun, Y. 2006. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, 1735–1742. IEEE.
  19. 19.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2019. Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722.
  20. 20.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020a. https://colab.research.google.com/github/facebookresearch/moco/blob/colab-notebook/colab/moco cifar10 demo.ipynb.
  21. 21.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020b. https://github.com/facebookresearch/moco.
  22. 22.He, K.; Gkioxari, G.; Dollar, P.; and Girshick, R. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, 2961–2969.
  23. 23.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778.
  24. 24.Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2018. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670.
  25. 25.Kalantidis, Y.; Sariyildiz, M. B.; Pion, N.; Weinzaepfel, P.; and Larlus, D. 2020. Hard negative mixing for contrastive learning. arXiv preprint arXiv:2010.01028.
  26. 26.Kim, S.; Lee, G.; Bae, S.; and Yun, S.-Y. 2020. MixCo: Mixup Contrastive Learning for Visual Representation. arXiv preprint arXiv:2010.06300.
  27. 27.Krizhevsky, A.; and Hinton, G. 2009. Learning multiple layers of features from tiny images. Technical report, University of Toronto, Toronto, Ontario.
  28. 28.Krothapalli, U.; and Abbott, A. L. 2020. Adaptive Label Smoothing. arXiv:2009.06432.
  29. 29.Lee, K.; Zhu, Y.; Sohn, K.; Li, C.-L.; Shin, J.; and Lee, H. 2021. i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning. arXiv preprint arXiv:2010.08887.
  30. 30.Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer.
  31. 31.Masci, J.; Meier, U.; Cires¸an, D.; and Schmidhuber, J. 2011. Stacked convolutional auto-encoders for hierarchical feature extraction. In International conference on artificial neural networks, 52–59. Springer.
  32. 32.Misra, I.; and van der Maaten, L. 2019. Self-supervised learning of pretext-invariant representations. arXiv preprint arXiv:1912.01991.
  33. 33.Muller, R.; Kornblith, S.; and Hinton, G. E. 2019. When does label smoothing help? In Advances in Neural Information Processing Systems, 4696–4705.
  34. 34.Noroozi, M.; and Favaro, P. 2016. Unsupervised learning of visual representations by solving jigsaw puzzles. In European Conference on Computer Vision, 69–84. Springer.
  35. 35.Noroozi, M.; Pirsiavash, H.; and Favaro, P. 2017. Representation learning by learning to count. In Proceedings of the IEEE International Conference on Computer Vision, 5898–5906.
  36. 36.Olshausen, B. A.; and Field, D. J. 1996. Emergence of simple-cell receptive field properties by learning a sparse code for natural images. Nature, 381(6583): 607–609.
  37. 37.Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.
  38. 38.Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, 91–99.
  39. 39.Shen, Z.; Liu, Z.; Xu, D.; Chen, Z.; Cheng, K.-T.; and Savvides, M. 2021. Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study. In International Conference on Learning Representations.
  40. 40.Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2818–2826.
  41. 41.Tian, Y.; Krishnan, D.; and Isola, P. 2019. Contrastive multi-view coding. arXiv preprint arXiv:1906.05849.
  42. 42.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In Advances in neural information processing systems, 5998–6008.
  43. 43.Verma, V.; Lamb, A.; Beckham, C.; Najafi, A.; Mitliagkas, I.; Lopez-Paz, D.; and Bengio, Y. 2019. Manifold mixup: Better representations by interpolating hidden states. In International Conference on Machine Learning, 6438–6447. PMLR.
  44. 44.Vincent, P.; Larochelle, H.; Bengio, Y.; and Manzagol, P.-A. 2008. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, 1096–1103.
  45. 45.Vincent, P.; Larochelle, H.; Lajoie, I.; Bengio, Y.; and Manzagol, P.-A. 2010. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of machine learning research, 11(Dec): 3371–3408.
  46. 46.Wu, Y.; Kirillov, A.; Massa, F.; Lo, W.-Y.; and Girshick, R. 2019. Detectron2. https://github.com/facebookresearch/detectron2.
  47. 47.Wu, Z.; Xiong, Y.; Yu, S. X.; and Lin, D. 2018. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3733–3742.
  48. 48.Ye, M.; Zhang, X.; Yuen, P. C.; and Chang, S.-F. 2019. Unsupervised embedding learning via invariant and spreading instance feature. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6210–6219.
  49. 49.Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE International Conference on Computer Vision, 6023–6032.
  50. 50.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR.
  51. 51.Zhang, R.; Isola, P.; and Efros, A. A. 2016. Colorful image colorization. In European conference on computer vision, 649–666. Springer.
  52. 52.Zhang, R.; Isola, P.; and Efros, A. A. 2017. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1058–1067.

Citation

MLA
Shen, Z., et al. “Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, 2022, pp. 2216–24, https://doi.org/10.1609/AAAI.V36I2.20119.
APA
Shen, Z., Liu, Z., Liu, Z., Savvides, M., Darrell, T., & Xing, E. (2022). Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 36(2), 2216–2224. https://doi.org/10.1609/AAAI.V36I2.20119
Chicago
Shen, Z., Z. Liu, Z. Liu, M. Savvides, T. Darrell, and E. Xing. 2022. “Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (2): 2216–24. https://doi.org/10.1609/AAAI.V36I2.20119.
Harvard
Shen, Z. et al. (2022) “Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(2), pp. 2216–2224. Available at: https://doi.org/10.1609/AAAI.V36I2.20119.
Vancouver
1. Shen Z, Liu Z, Liu Z, Savvides M, Darrell T, Xing E (2022) Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning. Proceedings of the AAAI Conference on Artificial Intelligence 36:2216–2224

BibTeX

@article{Shen_2022, title={Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation Learning}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I2.20119}, DOI={10.1609/aaai.v36i2.20119}, number={2}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Shen, Zhiqiang and Liu, Zechun and Liu, Zhuang and Savvides, Marios and Darrell, Trevor and Xing, Eric}, year={2022}, month=June, pages={2216–2224} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF