Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey

Longlong JingYing-li Tian

article2019TPAMI2,068 citations

Synthesizes self-supervised visual feature learning techniques for images and videos, providing a structured breakdown of architectures, pretext tasks, and quantitative benchmark comparisons to aid development without human-annotated data.

Listen

Modern computer vision relies heavily on deep neural networks, but training these models requires massive datasets with human annotations. Manually labeling millions of images and videos is expensive, time-consuming, and difficult to scale, especially for multi-frame video data. In response, self-supervised visual feature learning has emerged as an unsupervised technique that extracts meaningful visual representations directly from unlabeled data without human supervision.

The article systematically reviews and evaluates the landscape of deep learning-based self-supervised feature learning across images, videos, audio, and 3D modalities. It aims to establish how effectively these methods perform compared to standard supervised learning and assess their readiness for practical deployment.

To conduct this evaluation, the article categorizes self-supervised pretext taskstasks where artificial pseudo labels are derived directly from the inherent structure of the datainto generation-based, context-based, free semantic label-based, and cross-modal approaches. The quality of representations learned from these pretext tasks is evaluated by transferring the pre-trained network weights to standard downstream benchmark tasks, including image classification, object detection, semantic segmentation, and video action recognition.

The analysis reveals several key findings. First, self-supervised learning nearly matches supervised pre-training on high-level spatial visual tasks; downstream performance differences on object detection and semantic segmentation are within a narrow margin of less than 3%. Second, context-based pretext tasks generally deliver higher downstream performance than generation-based or heuristic methods across image and video benchmarks. Third, downstream performance scales significantly with dataset size and neural network capacity. For example, pre-training across large unlabeled datasets allows cross-modal models to achieve an action recognition accuracy of 94.2% on the UCF101 benchmark, outperforming models supervised on Kinetics datasets.

These findings demonstrate that self-supervised pre-training is a viable, high-performance alternative to expensive human data annotation. Organizations can substantially lower the cost and overhead of labeling large datasets while mitigating overfitting risks when deploying models on smaller target tasks. Furthermore, these techniques enable organizations to unlock value from massive, untapped archives of raw, unannotated video, sensor, and web data.

Organizations developing computer vision systems should adopt context-based self-supervised learning as a standard pre-training step when labeled training data is scarce. Implementations should leverage larger backbone network architectures and combine multiple pretext tasks or data modalities whenever possible. Additionally, future development should explore synthetic game-engine data and multimodal sensor feeds, provided that domain differences are appropriately managed.

Confidence in these findings is high for 2D image domains due to consistent benchmarks and reproducible codebases across standard datasets. However, caution is warranted for video and 3D modalities, where evaluation benchmarks like UCF101 remain small, and standardized evaluation metrics beyond downstream task transfer are still developing.

Cover for Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey

Abstract

Large-scale labeled data are generally required to train deep neural networks in order to obtain better performance in visual feature learning from images or videos for computer vision applications. To avoid extensive cost of collecting and annotating large-scale datasets, as a subset of unsupervised learning methods, self-supervised learning methods are proposed to learn general image and video features from large-scale unlabeled data without using any human-annotated labels. This paper provides an extensive review of deep learning-based self-supervised general visual feature learning methods from images or videos. First, the motivation, general pipeline, and terminologies of this field are described. Then the common deep neural network architectures that used for self-supervised learning are summarized. Next, the main components and evaluation metrics of self-supervised learning methods are reviewed followed by the commonly used image and video datasets and the existing self-supervised visual feature learning methods. Finally, quantitative performance comparisons of the reviewed methods on benchmark datasets are summarized and discussed for both image and video feature learning. At last, this paper is concluded and lists a set of promising future directions for self-supervised visual feature learning.

Table of Contents

  • I Introduction
  • I-A Motivation
  • I-B Term Definition
  • II Formulation of Different Learning Schemas
  • II-A Supervised Learning Formulation
  • II-B Semi-Supervised Learning Formulation
  • II-C Weakly Supervised Learning Formulation
  • II-D Unsupervised Learning Formulation
  • II-D1 Self-supervised Learning
  • III Common Deep Network Architectures
  • III-A Architectures for Learning Image Features
  • III-A1 AlexNet
  • III-A2 VGG
  • III-A3 ResNet
  • III-A4 GoogLeNet
  • III-A5 DenseNet
  • III-B Architectures for Learning Video Features
  • III-B1 Two-Stream Network
  • III-B2 Spatiotemporal Convolutional Neural Network
  • III-B3 Recurrent Neural Network
  • III-C Summary of ConvNet Architectures
  • IV Commonly used Pretext and Downstream Tasks
  • IV-A Learning Visual Features from Pretext Tasks
  • IV-B Commonly Used Pretext Tasks
  • IV-C Commonly Used Downstream Tasks for Evaluation
  • IV-C1 Semantic Segmentation
  • IV-C2 Object Detection
  • IV-C3 Image Classification
  • IV-C4 Human Action Recognition
  • IV-C5 Qualitative Evaluation
  • V Datasets
  • V-A Image Datasets
  • V-B Video Datasets
  • VI Image Feature Learning
  • VI-A Generation-based Image Feature Learning
  • VI-A1 Image Generation with GAN
  • VI-A2 Image Generation with Inpainting
  • VI-A3 Image Generation with Super Resolution
  • VI-A4 Image Generation with Colorization
  • VI-B Context-Based Image Feature Learning
  • VI-B1 Learning with Context Similarity
  • VI-B2 Learning with Spatial Context Structure
  • VI-C Free Semantic Label-based Image Feature Learning
  • VI-C1 Learning with Labels Generated by Game Engines
  • VI-C2 Learning with Labels Generated by Hard-code programs
  • VII Video Feature Learning
  • VII-A Generation-based Video Feature Learning
  • VII-A1 Learning from Video Generation
  • VII-A2 Learning from Video Colorization
  • VII-A3 Learning from Video Prediction
  • VII-B Temporal Context-based Learning
  • VII-C Cross Modal-based Learning
  • VII-C1 Learning from RGB-Flow Correspondence
  • VII-C2 Learning from Visual-Audio Correspondence
  • VII-C3 Ego-motion
  • VIII Performance Comparison
  • VIII-A Performance of Image Feature Learning
  • VIII-B Performance of Video Feature Learning
  • VIII-C Summary
  • IX Future Directions
  • X Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Self-Supervised Pretext Tasks in Visual Feature Learning

    model/method

    Self-supervised visual feature learning methods design pretext tasks to extract supervisory signals directly from unlabeled data without human annotations. These pretext tasks are categorized into four primary classes based on the data attributes exploited:

    1. Generation-Based Methods: Features are learned by generating images, video frames, or missing data attributes. Subcategories include:

      • Image Generation: Generative Adversarial Networks (GANs), image super-resolution, image inpainting, and colorization.
      • Video Generation: Spatiotemporal GANs, video prediction of future frames/representations, and video colorization leveraging temporal color coherence.
    2. Context-Based Methods: Pretext tasks exploit spatial, temporal, or clustering context:

      • Context Similarity: Grouping samples via clustering (e.g., DeepCluster) or contrastive learning across augmented views (e.g., SimCLR, MoCo, CMC).
      • Spatial Context Structure: Predicting relative positions between image patches, solving jigsaw puzzles, or predicting 2D geometric transformations/rotations (e.g., RotNet).
      • Temporal Context Structure: Video frame or clip sequence verification (Shuffle and Learn), frame order recognition (OPN, VCOP), cubic puzzle solving, or identifying the arrow of time.
    3. Free Semantic Label-Based Methods: Training networks on pseudo-semantic annotations generated automatically without human annotators. These annotations are derived via:

      • Game Engines/Simulators: Rendering synthetic scenes with pixel-level ground truth (depth, surface normals, instance contours, optical flow) combined with feature-space domain adaptation (e.g., adversarial domain alignment).
      • Hard-Coded Algorithms: Mining motion boundaries, moving object masks from video optical flow fields, saliency maps, or relative depth maps.
    4. Cross-Modal Methods: Training networks to align or verify correspondence across distinct synchronized sensory streams, including visual-audio correspondence, RGB-optical flow correspondence, and egocentric video to ego-motion/odometry sensor signal transformations.

  2. Knowl 2 — Mathematical Formulations of Visual Learning Paradigms

    equation

    Visual feature learning frameworks can be formalized based on the supervision signals used to optimize network parameters θ\theta over training sets:

    1. Supervised Learning: Given a labeled dataset D={(Xi,Yi)}i=1N\mathcal{D} = \{(X_i, Y_i)\}_{i=1}^N with input data XiX_i and fine-grained human-annotated ground-truth labels YiY_i, the objective is: Lsup(D)=minθ1Ni=1Nloss(Xi,Yi)\mathcal{L}_{\text{sup}}(\mathcal{D}) = \min_\theta \frac{1}{N} \sum_{i=1}^N \text{loss}(X_i, Y_i)

    2. Semi-Supervised Learning: Given a small labeled dataset D1={(Xi,Yi)}i=1N\mathcal{D}_1 = \{(X_i, Y_i)\}_{i=1}^N and a large unlabeled dataset D2={Zi}i=1M\mathcal{D}_2 = \{Z_i\}_{i=1}^M, the objective is: Lsemi(D1,D2)=minθ(1Ni=1Nloss(Xi,Yi)+1Mi=1Mloss(Zi,R(Zi,D1)))\mathcal{L}_{\text{semi}}(\mathcal{D}_1, \mathcal{D}_2) = \min_\theta \left( \frac{1}{N} \sum_{i=1}^N \text{loss}(X_i, Y_i) + \frac{1}{M} \sum_{i=1}^M \text{loss}(Z_i, R(Z_i, \mathcal{D}_1)) \right) where R(Zi,D1)R(Z_i, \mathcal{D}_1) is a task-specific relation function between unlabeled sample ZiZ_i and the labeled set D1\mathcal{D}_1.

    3. Weakly Supervised Learning: Given training data D={(Xi,Ci)}i=1N\mathcal{D} = \{(X_i, C_i)\}_{i=1}^N where CiC_i denotes coarse-grained, noisy, or cheaper weak annotations (e.g., image tags or hashtags), the objective is: Lweak(D)=minθ1Ni=1Nloss(Xi,Ci)\mathcal{L}_{\text{weak}}(\mathcal{D}) = \min_\theta \frac{1}{N} \sum_{i=1}^N \text{loss}(X_i, C_i)

    4. Self-Supervised Learning: Given unlabeled data D={(Xi,Pi)}i=1N\mathcal{D} = \{(X_i, P_i)\}_{i=1}^N where PiP_i denotes pseudo-labels automatically derived from the intrinsic data attributes or structures without human annotation, the objective is: Lself(D)=minθ1Ni=1Nloss(Xi,Pi)\mathcal{L}_{\text{self}}(\mathcal{D}) = \min_\theta \frac{1}{N} \sum_{i=1}^N \text{loss}(X_i, P_i)

  3. Knowl 3 — Self-Supervised Visual Feature Learning and Transfer Framework

    model/method

    The self-supervised visual feature learning paradigm operates as a two-phase framework:

    1. Self-Supervised Pretext Training Phase: A neural network (e.g., 2D ConvNet, 3D ConvNet, or recurrent network) is trained on an unlabeled dataset by optimizing a loss function L(Oi,Pi)\mathcal{L}(O_i, P_i), where OiO_i is the network prediction for input XiX_i and PiP_i is a pseudo-label generated automatically from the inherent structure or metadata of the data (e.g., patch permutations, rotation angles, cluster assignments, future frames, or audio-visual synchrony).

    2. Supervised Downstream Transfer Phase: The trained parameters of the network serve as a pre-trained feature extractor or initialization point for target downstream vision tasks (such as object classification, object detection via Fast R-CNN, semantic segmentation via FCN, or action recognition). Typically, the learned weights from the initial convolutional layers (which capture low-level and mid-level general visual representations) are transferred and fine-tuned on labeled target datasets, avoiding over-fitting and accelerating convergence on small target datasets.

  4. Knowl 4 — Linear Probe Classification Performance on ImageNet and Places Across AlexNet Layers

    data/table

    The quality of representations learned by various self-supervised pretext tasks is quantitatively evaluated by freezing the weights of an AlexNet backbone pre-trained on ImageNet without labels, extracting feature activations from individual convolutional layers (conv1 through conv5), and training a linear classifier on top of each frozen layer for ImageNet and Places classification.

    ImageNet Places
    Method Pretext Tasks conv1 conv2 conv3 conv4 conv5 conv1 conv2 conv3 conv4 conv5
    Places labels 22.1 35.1 40.2 43.3 44.6
    ImageNet labels 19.3 36.3 44.2 48.3 50.5 22.7 34.8 38.4 39.4 38.7
    Random (Scratch) 11.6 17.1 16.9 16.3 14.1 15.7 20.3 19.8 19.1 17.5
    ColorfulColorization Generation 12.5 24.5 30.4 31.5 30.3 16.0 25.7 29.6 30.3 29.7
    BiGAN Generation 17.7 24.5 31.0 29.9 28.0 21.4 26.2 27.1 26.1 24.0
    SplitBrain Generation 17.7 29.3 35.4 35.2 32.8 21.3 30.7 34.0 34.1 32.5
    ContextEncoder Context 14.1 20.7 21.0 19.8 15.5 18.2 23.2 23.4 21.9 18.4
    ContextPrediction Context 16.2 23.3 30.2 31.7 29.6 19.7 26.7 31.9 32.7 30.9
    Jigsaw Context 18.2 28.8 34.0 33.9 27.1 23.0 32.1 35.5 34.8 31.3
    Learning2Count Context 18.0 30.6 34.3 32.5 25.7 23.3 33.9 36.3 34.7 29.6
    DeepClustering Context 13.4 32.3 41.0 39.6 38.2 19.6 33.2 39.2 39.8 34.7

    Key takeaways from these results:

    1. All self-supervised representations outperform the random initialization baseline across all layers on both datasets.
    2. Performance of self-supervised methods peaks at intermediate layers (conv3 and conv4), declining at conv5 due to specialization to the pretext task objective.
    3. On Places, where a domain gap exists for ImageNet-supervised features, top self-supervised models (e.g., DeepClustering at 39.8% on conv4) match or exceed ImageNet-supervised features (39.4%).
  5. Knowl 5 — Downstream Transfer Performance on PASCAL VOC Benchmarks

    data/table

    Representations learned with AlexNet on unlabeled ImageNet via various self-supervised pretext tasks are transferred and evaluated on PASCAL VOC benchmark tasks: classification (mean average precision on VOC 2007 test), object detection (mAP with Fast R-CNN on VOC 2007 test), and semantic segmentation (mean intersection over union with FCN on VOC 2012 validation).

    Method Pretext Tasks Classification (%) Detection (%) Segmentation (%)
    ImageNet Labels 79.9 56.8 48.0
    Random (Scratch) 57.0 44.5 30.1
    ContextEncoder Generation 56.5 44.5 29.7
    BiGAN Generation 60.1 46.9 35.2
    ColorfulColorization Generation 65.9 46.9 35.6
    SplitBrain Generation 67.1 46.7 36.0
    RankVideo Context 63.1 47.2 35.4
    PredictNoise Context 65.3 49.4 37.1
    JigsawPuzzle Context 67.6 53.2 37.6
    ContextPrediction Context 65.3 51.1
    Learning2Count Context 67.7 51.4 36.6
    DeepClustering Context 73.7 55.4 45.1
    WatchingVideo Free Semantic Label 61.0 52.2
    CrossDomain Free Semantic Label 68.0 52.6
    AmbientSound Cross Modal 61.3
    TiedToEgoMotion Cross Modal 41.7
    EgoMotion Cross Modal 54.2 43.9

    Key observations:

    1. Context-based methods (particularly DeepClustering) achieve the strongest results among pretext categories, reaching 55.4% detection mAP and 45.1% segmentation mIoU, within 1.4% and 2.9% of fully supervised ImageNet pretraining, respectively.
    2. Dense prediction tasks (detection and segmentation) benefit substantially from self-supervised pretraining over random initialization.
  6. Knowl 6 — Self-Supervised Video Representation Transfer on UCF101 and HMDB51 Action Recognition

    data/table

    The generalizability of self-supervised spatiotemporal features learned from video (or audio-visual) datasets is evaluated by fine-tuning backbones on action recognition on the UCF101 (split 1) and HMDB51 datasets.

    Method Architecture Pretraining Dataset UCF101 (%) HMDB51 (%)
    Fully supervised 3DResNet18 Kinetics 84.4 56.4
    Fully supervised R(2+1)D-18 Kinetics 93.1 63.6
    ShuffleLearn CaffeNet UCF/HMDB 50.2 18.1
    GeometryGuide CaffeNet UCF/HMDB 55.1 23.3
    RL CaffeNet UCF/HMDB 58.6 25.0
    CMC CaffeNet UCF/HMDB 59.1 26.7
    CrossLearn CaffeNet UCF/HMDB 58.7 27.2
    OPN VGG UCF/HMDB 59.8 23.8
    L3-Net VGG AudioSet 72.3 40.2
    MotionPred C3D Kinetics 61.2 33.4
    RotNet3D 3D-ResNet18 Kinetics 62.9 33.7
    DPC 3D-ResNet18 Kinetics 68.2 34.5
    ST-Puzzle 3D-ResNet18 Kinetics 65.8 33.7
    DPC 3D-ResNet34 Kinetics 75.7 35.7
    AVTS MC3-18 AudioSet 89.0 61.6
    VCOP R(2+1)D-18 Kinetics 72.4 30.9
    XDC R(2+1)D-18 Kinetics 84.2 47.1
    GDT R(2+1)D-18 Kinetics 88.7 57.8
    XDC R(2+1)D-18 IG65M 94.2 67.4

    Key observations:

    1. Backbone architecture capacity heavily impacts downstream action accuracy (e.g., DPC accuracy increases from 68.2% on 3D-ResNet18 to 75.7% on 3D-ResNet34 on UCF101).
    2. Scaling pretraining data volume dramatically improves downstream transfer: cross-modal deep clustering (XDC) scaled to 65 million videos (IG65M) achieves 94.2% on UCF101 and 67.4% on HMDB51, exceeding Kinetics-supervised pretraining with the same R(2+1)D-18 backbone.
  7. Knowl 7 — Self-Supervised Audio Classification Performance on DCASE and ESC50

    data/table

    Self-supervised representations learned jointly from audio-visual streams are evaluated on downstream acoustic classification benchmarks: DCASE (10-class acoustic scene classification) and ESC50 (50-class environmental sound classification).

    Method Pretraining Dataset # Training Data DCASE (%) ESC50 (%)
    Random Forest ESC50 1.6K 44.3
    Piczak ConvNet ESC50 1.6K 64.5
    ConvRBM ESC50 1.6K 86.5
    RNH DCASE 100 72.0
    Ensemble DCASE 100 77.0
    AVTS Kinetics-400 230K 91.0 76.7
    XDC Kinetics-400 230K 78.0
    GDT Kinetics-400 230K 94.0 78.6
    Autoencoder SoundNet 2M+ 39.9
    SoundNet SoundNet 2M+ 88.0 74.2
    L3-Net SoundNet 2M+ 93.0 79.3
    AVTS SoundNet 2M+ 94.0 82.3
    XDC AudioSet 1.8M 93.0 84.8

    Key observations:

    1. Cross-modal audio-visual self-supervision (e.g., AVTS, L3-Net, XDC, GDT) achieves competitive acoustic classification accuracy compared to supervised audio representations.
    2. Audio representations improve with larger unlabeled training datasets (e.g., AVTS achieves 94.0% on DCASE and 82.3% on ESC50 when pre-trained on SoundNet 2M+).
  8. Knowl 8 — Self-Supervised 3D Representation Learning on ModelNet40 Shape Classification

    data/table

    Self-supervised feature learning methods for 3D representations are evaluated on the ModelNet40 40-class 3D CAD object recognition benchmark using an SVM trained on frozen extracted features across point cloud, mesh, and multi-view image modalities.

    Method Pretext Task Modality Accuracy (%)
    PointNet* (Supervised) Point Cloud 89.2
    DGCNN* (Supervised) Point Cloud 92.2
    MeshNet* (Supervised) Mesh 91.9
    MVCNN* (Supervised) Images 90.1
    T-L Network Generation Point Cloud 74.4
    VConv-DAE Generation Point Cloud 75.5
    3D-GAN Generation Point Cloud 83.3
    Latent-GAN Generation Point Cloud 85.7
    MRTNet-VAE Generation Point Cloud 86.4
    Contrast-Cluster Context Point Cloud 86.8
    FoldingNet Generation Point Cloud 88.4
    PointCapsNet Generation Point Cloud 88.9
    3DMultiTask Multiple Point Cloud 89.1
    MVI Multiple Point Cloud 89.3
    XMV Multiple Point Cloud 89.8
    XMV Multiple Images 87.3
    MVI Multiple Image 88.2
    MVI Multiple Mesh 87.7

    Note: Methods marked with an asterisk () are trained using human-annotated category labels.*

    Key findings:

    1. Self-supervised multi-task and cross-modal contrasting methods (e.g., XMV at 89.8% on point clouds, MVI at 89.3%) match or surpass the supervised PointNet baseline (89.2%).
    2. Combining multi-view and cross-modality constraints yields stronger representations than single-task 3D generative auto-encoders.
  9. Knowl 9 — Representation Hierarchy and Layer Specialization in Pretext-Trained Networks

    empirical result

    Filter and representation analysis across deep convolutional networks pre-trained on self-supervised pretext tasks reveals a consistent structural hierarchy:

    1. Layer Hierarchy:

      • Shallow Layers (e.g., conv1, conv2 in AlexNet): Capture universal low-level visual features such as edges, corners, colors, and textures.
      • Middle Layers (e.g., conv3, conv4 in AlexNet): Capture general semantic mid-level parts and object patterns. Across benchmark downstream transfer evaluations (ImageNet, Places, PASCAL VOC), linear classifiers trained on conv3 and conv4 consistently outperform those trained on shallow layers or final layers.
      • Deep Layers (e.g., conv5 and fully connected layers): Specialize excessively toward solving the specific objective function of the pre-defined pretext task (e.g., predicting specific color distributions, rotations, or inpainting boundaries), which reduces their transferability to general downstream tasks.
    2. Filter Interpretability: Visual filter dissection reveals that self-supervised networks naturally discover and allocate units responding to objects, scenes, object parts, materials, textures, and colors, demonstrating that diverse pretext tasks induce representations functionally similar to supervised networks.

  10. Knowl 10 — Core Empirical Principles for Self-Supervised Visual Feature Learning

    empirical result

    Cross-benchmark analysis across image, video, audio, and 3D modalities yields the following foundational guidelines for self-supervised feature learning:

    1. Pretext Task Selection: Context-based methods (such as clustering-based pseudo-labeling and multi-view contrastive learning like SimCLR and MoCo) consistently achieve superior transfer performance compared to pure generation-based or low-level transformation tasks.
    2. Network Architecture Capacity: Increasing model capacity—including depth, width, kernel count, and feature embedding dimension—significantly boosts the quality of learned representations and accelerates downstream transfer.
    3. Pretraining Dataset Scale: Self-supervised representation quality scales directly with the volume of unlabeled pretraining data. Large-scale pretraining (e.g., 65M videos in XDC) allows self-supervised models to outperform fully supervised pretraining on standard benchmarks (UCF101 and HMDB51).
    4. Domain Alignment: Representation transfer improves when the visual domain gap between the pretext training set and downstream task dataset is minimized.
    5. Multi-Task and Multi-Modal Synergy: Formulating multiple complementary pretext objectives or leveraging cross-modal correspondences (e.g., visual-audio, optical flow-RGB, multi-view 3D) forces networks to capture modality-invariant semantic signals and yields more robust features than single-pretext models.

Coverage note — Standard deep network architectural summaries (AlexNet, VGG, ResNet, GoogLeNet, DenseNet, C3D, LRCN, Two-Stream) and descriptions of standard public datasets (CIFAR10, ImageNet, Places, Kinetics, AudioSet, UCF101, etc.) were omitted as standalone knowls because they represent established prior background rather than original contributions of this survey paper.

References

  1. 1.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR, pp. 580–587, 2014.
  2. 2.R. Girshick, “Fast R-CNN,” in ICCV, 2015.
  3. 3.S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NIPS, pp. 91–99, 2015.
  4. 4.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR, pp. 3431–3440, 2015.
  5. 5.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” TPAMI, 2018.
  6. 6.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR, pp. 2881–2890, 2017.
  7. 7.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR, pp. 3156–3164, 2015.
  8. 8.K. He, R. Girshick, and P. Dollar, “Rethinking imagenet pre-training,” ICCV, 2018.
  9. 9.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NIPS, pp. 1097–1105, 2012.
  10. 10.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” ICLR, pp. 1–14, 2015.
  11. 11.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR, pp. 1–9, 2015.
  12. 12.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, pp. 770–778, 2016.
  13. 13.G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten, “Densely connected convolutional networks,” in CVPR, p. 3, 2017.
  14. 14.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR, pp. 248–255, IEEE, 2009.
  15. 15.A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al., “The open images dataset v4,” IJCV, pp. 1–26, 2020.
  16. 16.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, A. Acosta, A. P. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi, “Photo-realistic single image super-resolution using a generative adversarial network,” in CVPR, pp. 105–114, 2017.
  17. 17.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3D convolutional networks,” in ICCV, 2015.
  18. 18.C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR, pp. 652–660, 2017.
  19. 19.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al., “The kinetics human action video dataset,” arXiv preprint arXiv:1705.06950, 2017.
  20. 20.R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in ECCV, pp. 649–666, Springer, 2016.
  21. 21.D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in CVPR, pp. 2536–2544, 2016.
  22. 22.M. Noroozi and P. Favaro, “Unsupervised learning of visual representions by solving jigsaw puzzles,” in ECCV, 2016.
  23. 23.D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba, “Network dissection: Quantifying interpretability of deep visual representations,” in CVPR, 2017.
  24. 24.K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS, pp. 568–576, 2014.
  25. 25.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR, 2015.
  26. 26.C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” ICLR, 2017.
  27. 27.D. Bau, J.-Y. Zhu, H. Strobelt, B. Zhou, J. B. Tenenbaum, W. T. Freeman, and A. Torralba, “Gan dissection: Visualizing and understanding generative adversarial networks,” ICLR, 2019.
  28. 28.M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV, pp. 818–833, Springer, 2014.
  29. 29.K. Soomro, A. R. Zamir, and M. Shah, “UCF101: A dataset of 101 human actions classes from videos in the wild,” CRCV-TR, 2012.
  30. 30.D. Pathak, R. Girshick, P. Dollar, T. Darrell, and B. Hariharan, “Learning features by watching objects move,” in CVPR, 2017.
  31. 31.M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in ECCV, 2018.
  32. 32.G. Larsson, M. Maire, and G. Shakhnarovich, “Colorization as a proxy task for visual understanding,” in CVPR, 2017.
  33. 33.I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and learn: unsupervised learning using temporal order verification,” in ECCV, pp. 527–544, Springer, 2016.
  34. 34.B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in NIPS, pp. 7773–7784, 2018.
  35. 35.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NIPS, pp. 2672–2680, 2014.
  36. 36.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV, 2017.
  37. 37.C. Vondrick, H. Pirsiavash, and A. Torralba, “Generating videos with scene dynamics,” in NIPS, pp. 613–621, 2016.
  38. 38.S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz, “Mocogan: Decomposing motion and content for video generation,” CVPR, 2018.
  39. 39.N. Srivastava, E. Mansimov, and R. Salakhutdinov, “Unsupervised Learning of Video Representations using LSTMs,” in ICML, 2015.
  40. 40.M. Noroozi, A. Vinjimoor, P. Favaro, and H. Pirsiavash, “Boosting self-supervised learning via knowledge transfer,” CVPR, 2018.
  41. 41.D. Li, W.-C. Hung, J.-B. Huang, S. Wang, N. Ahuja, and M.-H. Yang, “Unsupervised visual representation learning by graph-based consistent constraints,” in ECCV, 2016.
  42. 42.U. Ahsan, R. Madhok, and I. Essa, “Video jigsaw: Unsupervised learning of spatiotemporal context for video action recognition,” WACV, 2019.
  43. 43.C. Wei, L. Xie, X. Ren, Y. Xia, C. Su, J. Liu, Q. Tian, and A. L. Yuille, “Iterative reorganization with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning,” CVPR, 2019.
  44. 44.D. Kim, D. Cho, D. Yoo, and I. S. Kweon, “Learning image representations by completing damaged jigsaw puzzles,” WACV, 2018.
  45. 45.C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual representation learning by context prediction,” in ICCV, pp. 1422–1430, 2015.
  46. 46.S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in ICLR, 2018.
  47. 47.L. Jing and Y. Tian, “Self-supervised spatiotemporal feature learning by video geometric transformations,” arXiv preprint arXiv:1811.11387, 2018.
  48. 48.D. Wei, J. Lim, A. Zisserman, and W. T. Freeman, “Learning and using the arrow of time,” in CVPR, pp. 8052–8060, 2018.
  49. 49.H.-Y. Lee, J.-B. Huang, M. Singh, and M.-H. Yang, “Unsupervised representation learning by sorting sequences,” in ICCV, pp. 667–676, IEEE, 2017.
  50. 50.D. Xu, J. Xiao, Z. Zhao, J. Shao, D. Xie, and Y. Zhuang, “Self-supervised spatiotemporal learning via video clip order prediction,” in CVPR, pp. 10334–10343, 2019.
  51. 51.A. Faktor and M. Irani, “Video segmentation by non-local consensus voting.,” in BMVC, 2014.
  52. 52.O. Stretcu and M. Leordeanu, “Multiple frames matching for object discovery in video.,” in BMVC, 2015.
  53. 53.Z. Ren and Y. J. Lee, “Cross-domain self-supervised multi-task feature learning using synthetic imagery,” in CVPR, 2018.
  54. 54.P. Krhenbhl, “Free supervision from video games,” in CVPR, June 2018.
  55. 55.I. Croitoru, S.-V. Bogolin, and M. Leordeanu, “Unsupervised learning from video to detect foreground objects in single images,” ICCV, 2017.
  56. 56.Y. Li, M. Paluri, J. M. Rehg, and P. Dollar, “Unsupervised learning of edges,” CVPR, pp. 1619–1627, 2016.
  57. 57.H. Jiang, G. Larsson, M. Maire Greg Shakhnarovich, and E. Learned-Miller, “Self-supervised relative depth learning for urban scene understanding,” in ECCV, pp. 19–35, 2018.
  58. 58.R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in ICCV, pp. 609–617, IEEE, 2017.
  59. 59.N. Sayed, B. Brattoli, and B. Ommer, “Cross and learn: Cross-modal self-supervision,” GCPR, 2018.
  60. 60.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” arXiv preprint arXiv:2002.05709, 2020.
  61. 61.P. Agrawal, J. Carreira, and J. Malik, “Learning to see by moving,” in ICCV, pp. 37–45, 2015.
  62. 62.D. Jayaraman and K. Grauman, “Learning image representations tied to ego-motion,” in ICCV, pp. 1413–1421, 2015.
  63. 63.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” IJCV, pp. 303–338, 2010.
  64. 64.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR, pp. 3213–3223, 2016.
  65. 65.B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” IJCV, pp. 302–321, 2019.
  66. 66.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV, pp. 740–755, Springer, 2014.
  67. 67.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in CVPR, pp. 779–788, 2016.
  68. 68.J. Redmon and A. Farhadi, “Yolo9000: Better, faster, stronger,” in CVPR, 2017.
  69. 69.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in ECCV, pp. 21–37, Springer, 2016.
  70. 70.T.-Y. Lin, P. Dollar, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection.,” in CVPR, p. 4, 2017.
  71. 71.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, “Focal loss for dense object detection,” TPAMI, 2018.
  72. 72.J. A. Suykens and J. Vandewalle, “Least squares support vector machine classifiers,” Neural processing letters, pp. 293–300, 1999.
  73. 73.H. Kuehne, H. Jhuang, R. Stiefelhagen, and T. Serre, “Hmdb51: A large video database for human motion recognition,” in HPCSE, pp. 571–582, Springer, 2013.
  74. 74.A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” tech. rep., Citeseer, 2009.
  75. 75.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, pp. 2278–2324, 1998.
  76. 76.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS, pp. 487–495, 2014.
  77. 77.B. Zhou, A. Khosla, A. Lapedriza, A. Torralba, and A. Oliva, “Places: An image database for deep scene understanding,” arXiv preprint arXiv:1610.02055, 2016.
  78. 78.A. Coates, A. Ng, and H. Lee, “An analysis of single-layer networks in unsupervised feature learning,” in AISTATS, pp. 215–223, 2011.
  79. 79.S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser, “Semantic scene completion from a single depth image,” in CVPR, pp. 190–198, IEEE, 2017.
  80. 80.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in NIPSW, p. 5, 2011.
  81. 81.J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP, pp. 776–780, IEEE, 2017.
  82. 82.A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in CVPR, pp. 3354–3361, IEEE, 2012.
  83. 83.M. Monfort, B. Zhou, S. A. Bargal, T. Yan, A. Andonian, K. Ramakrishnan, L. Brown, Q. Fan, D. Gutfruend, C. Vondrick, et al., “Moments in time dataset: one million videos for event understanding,” TPAMI, pp. 502–508, 2019.
  84. 84.J. McCormac, A. Handa, S. Leutenegger, and A. J. Davison, “Scenenet rgb-d: Can 5m synthetic images beat generic imagenet pre-training on indoor segmentation,” in ICCV, 2017.
  85. 85.B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “The new data and new challenges in multimedia research,” arXiv preprint arXiv:1503.01817, 2015.
  86. 86.Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3d shapenets: A deep representation for volumetric shapes,” in CVPR, pp. 1912–1920, 2015.
  87. 87.A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al., “Shapenet: An information-rich 3d model repository,” arXiv preprint arXiv:1512.03012, 2015.
  88. 88.L. Yi, V. G. Kim, D. Ceylan, I.-C. Shen, M. Yan, H. Su, C. Lu, Q. Huang, A. Sheffer, and L. Guibas, “A scalable active framework for region annotation in 3d shape collections,” TOG, pp. 1–12, 2016.
  89. 89.D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia, pp. 1733–1746, 2015.
  90. 90.K. J. Piczak, “Esc: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM international conference on Multimedia, pp. 1015–1018, 2015.
  91. 91.L. Jing, Y. Chen, L. Zhang, M. He, and Y. Tian, “Self-supervised feature learning by cross-modality and cross-view correspondences,” arXiv preprint arXiv:2004.05749, 2020.
  92. 92.G. Larsson, M. Maire, and G. Shakhnarovich, “Learning representations for automatic colorization,” in ECCV, pp. 577–593, Springer, 2016.
  93. 93.J. Donahue, P. Krahenb¨uhl, and T. Darrell, “Adversarial feature learning,” ICLR, 2017.
  94. 94.S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Globally and Locally Consistent Image Completion,” SIGGRAPH, 2017.
  95. 95.A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” ICLR, 2016.
  96. 96.T. Chen, X. Zhai, M. Ritter, M. Lucic, and N. Houlsby, “Self-supervised gans via auxiliary rotation loss,” in CVPR, pp. 12154–12163, 2019.
  97. 97.R. Zhang, P. Isola, and A. A. Efros, “Split-brain autoencoders: Unsupervised learning by cross-channel prediction,” in CVPR, 2017.
  98. 98.S. Jenni and P. Favaro, “Self-supervised feature learning by learning to spot artifacts,” CVPR, 2018.
  99. 99.M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein gan,” ICML, 2017.
  100. 100.J. Xie, R. Girshick, and A. Farhadi, “Unsupervised deep embedding for clustering analysis,” in ICML, pp. 478–487, 2016.
  101. 101.Y. Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,” ICLR, 2020.
  102. 102.A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748, 2018.
  103. 103.R. Santa Cruz, B. Fernando, A. Cherian, and S. Gould, “Visual permutation learning,” TPAMI, 2018.
  104. 104.T. N. Mundhenk, D. Ho, and B. Y. Chen, “Improvements to context based self-supervised learning,” in CVPR, 2018.
  105. 105.J. Yang, D. Parikh, and D. Batra, “Joint unsupervised learning of deep representations and image clusters,” in CVPR, pp. 5147–5156, 2016.
  106. 106.M. Noroozi, H. Pirsiavash, and P. Favaro, “Representation learning by learning to count,” in ICCV, 2017.
  107. 107.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” CVPR, 2019.
  108. 108.C. Doersch and A. Zisserman, “Multi-task self-supervised visual learning,” in ICCV, 2017.
  109. 109.P. Bojanowski and A. Joulin, “Unsupervised learning by predicting noise,” ICML, 2017.
  110. 110.X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in ICCV, 2015.
  111. 111.G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science, pp. 504–507, 2006.
  112. 112.T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in NIPS, pp. 2234–2242, 2016.
  113. 113.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in NIPS, pp. 6626–6637, 2017.
  114. 114.A. Brock, J. Donahue, and K. Simonyan, “Large scale gan training for high fidelity natural image synthesis,” ICLR, 2019.
  115. 115.T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” CVPR, 2019.
  116. 116.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” CVPR, 2017.
  117. 117.R. Zhang, J.-Y. Zhu, P. Isola, X. Geng, A. S. Lin, T. Yu, and A. A. Efros, “Real-time user-guided image colorization with learned deep priors,” TOG, 2017.
  118. 118.S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Let there be color!: joint end-to-end learning of global and local image priors for automatic image colorization with simultaneous classification,” TOG, p. 110, 2016.
  119. 119.A. K. Jain, M. N. Murty, and P. J. Flynn, “Data clustering: a review,” CSUR, pp. 264–323, 1999.
  120. 120.N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR, pp. 886–893, IEEE, 2005.
  121. 121.D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” IJCV, pp. 91–110, 2004.
  122. 122.J. Sanchez, F. Perronnin, T. Mensink, and J. Verbeek, “Image classification with the fisher vector: Theory and practice,” IJCV, pp. 222–245, 2013.
  123. 123.M. Patrick, Y. M. Asano, R. Fong, J. F. Henriques, G. Zweig, and A. Vedaldi, “Multi-modal self-supervision from generalized data transformations,” arXiv preprint arXiv:2003.04298, 2020.
  124. 124.L. Jing, Y. Chen, L. Zhang, M. He, and Y. Tian, “Self-supervised modal and view invariant feature learning,” arXiv preprint, 2020.
  125. 125.D. Kim, D. Cho, and I. S. Kweon, “Self-supervised video representation learning with space-time cubic puzzles,” AAAI, 2019.
  126. 126.U. Buchler, B. Brattoli, and B. Ommer, “Improving spatiotemporal self-supervision by deep reinforcement learning,” in ECCV, pp. 770–786, 2018.
  127. 127.P. Goyal, D. Mahajan, A. Gupta, and I. Misra, “Scaling and benchmarking self-supervised visual representation learning,” ICCV, 2019.
  128. 128.S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and service robotics, pp. 621–635, Springer, 2018.
  129. 129.A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in CoRL, pp. 1–16, 2017.
  130. 130.S. Xingjian, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” in NIPS, pp. 802–810, 2015.
  131. 131.T. Han, W. Xie, and A. Zisserman, “Video representation learning by dense predictive coding,” in ICCVW, 2019.
  132. 132.Z. Luo, B. Peng, D.-A. Huang, A. Alahi, and L. Fei-Fei, “Unsupervised learning of long-term motion dynamics for videos,” in CVPR, 2017.
  133. 133.R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee, “Decomposing motion and content for natural video sequence prediction,” in ICLR, 2017.
  134. 134.M. Saito, E. Matsumoto, and S. Saito, “Temporal generative adversarial nets with singular value clipping,” in ICCV, p. 5, 2017.
  135. 135.C. Vondrick, A. Shrivastava, A. Fathi, S. Guadarrama, and K. Murphy, “Tracking emerges by colorizing videos,” in ECCV, 2018.
  136. 136.B. Brattoli, U. Buchler, A.-S. Wahl, M. E. Schwab, and B. Ommer, “Lstm self-supervision for detailed behavior analysis,” in CVPR, pp. 3747–3756, IEEE, 2017.
  137. 137.B. Fernando, H. Bilen, E. Gavves, and S. Gould, “Self-supervised video representation learning with odd-one-out networks,” in CVPR, 2017.
  138. 138.D. Jayaraman and K. Grauman, “Slow and steady feature analysis: higher order temporal coherence in video,” in CVPR, pp. 3852–3861, 2016.
  139. 139.X. Wang, K. He, and A. Gupta, “Transitive invariance for self-supervised visual representation learning,” in ICCV, 2017.
  140. 140.Y. Zhang, S. Khamis, C. Rhemann, J. Valentin, A. Kowdle, V. Tankovich, M. Schoenberg, S. Izadi, T. Funkhouser, and S. Fanello, “Activestereonet: End-to-end self-supervised learning for active stereo systems,” in ECCV, pp. 784–801, 2018.
  141. 141.A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in ECCV, pp. 801–816, Springer, 2016.
  142. 142.A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” ECCV, 2018.
  143. 143.A. Mahendran, J. Thewlis, and A. Vedaldi, “Cross pixel optical flow similarity for self-supervised learning,” ACCV, 2018.
  144. 144.Y. Zou, Z. Luo, and J.-B. Huang, “Df-net: Unsupervised joint learning of depth and flow using cross-task consistency,” in ECCV, pp. 38–55, Springer, 2018.
  145. 145.T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in CVPR, p. 7, 2017.
  146. 146.A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. Van Der Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” in ICCV, pp. 2758–2766, 2015.
  147. 147.E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, “Flownet 2.0: Evolution of optical flow estimation with deep networks,” in CVPR, pp. 1647–1655, IEEE, 2017.
  148. 148.Z. Yin and J. Shi, “Geonet: Unsupervised learning of dense depth, optical flow and camera pose,” in CVPR, 2018.
  149. 149.J. Wang, J. Jiao, L. Bao, S. He, Y. Liu, and W. Liu, “Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics,” in CVPR, pp. 4006–4015, 2019.
  150. 150.S. Meister, J. Hur, and S. Roth, “Unflow: Unsupervised learning of optical flow with a bidirectional census loss,” in AAAI, pp. 7251–7259, 2018.
  151. 151.G. Iyer, J. K. Murthy, G. Gupta, K. M. Krishna, and L. Paull, “Geometric consistency for self-supervised end-to-end visual odometry,” CVPRW, 2018.
  152. 152.H. Alwassel, D. Mahajan, L. Torresani, B. Ghanem, and D. Tran, “Self-supervised learning by cross-modal audio-video clustering,” arXiv preprint arXiv:1911.12667, 2019.
  153. 153.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Deep end2end voxel2voxel prediction,” in CVPRW, pp. 17–24, 2016.
  154. 154.M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” in ICLR, 2016.
  155. 155.F. A. Reda, G. Liu, K. J. Shih, R. Kirby, J. Barker, D. Tarjan, A. Tao, and B. Catanzaro, “Sdc-net: Video prediction using spatially-displaced convolution,” in ECCV, pp. 718–733, 2018.
  156. 156.M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine, “Stochastic variational video prediction,” ICLR, 2018.
  157. 157.X. Liang, L. Lee, W. Dai, and E. P. Xing, “Dual motion gan for future-flow embedded video prediction,” in ICCV, 2017.
  158. 158.C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” in NIPS, pp. 64–72, 2016.
  159. 159.K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet,” in CVPR, pp. 18–22, 2018.
  160. 160.D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in CVPR, pp. 6450–6459, 2018.
  161. 161.C. Gan, B. Gong, K. Liu, H. Su, and L. J. Guibas, “Geometry guided convolutional neural networks for self-supervised video representation learning,” in CVPR, pp. 5589–5597, 2018.
  162. 162.K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in MLSP, pp. 1–6, IEEE, 2015.
  163. 163.H. B. Sailor, D. M. Agrawal, and H. A. Patil, “Unsupervised filterbank learning using convolutional restricted boltzmann machine for environmental sound classification.,” in INTERSPEECH, pp. 3107–3111, 2017.
  164. 164.G. Roma, W. Nogueira, and P. Herrera, “Recurrence quantification analysis features for environmental sound recognition,” in WASPAA, pp. 1–4, IEEE, 2013.
  165. 165.Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in NIPS, pp. 892–900, 2016.
  166. 166.J. Wu, C. Zhang, T. Xue, B. Freeman, and J. Tenenbaum, “Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling,” in NIPS, pp. 82–90, 2016.
  167. 167.Y. Yang, C. Feng, Y. Shen, and D. Tian, “Foldingnet: Point cloud auto-encoder via deep grid deformation,” in CVPR, pp. 206–215, 2018.
  168. 168.P. Achlioptas, O. Diamanti, I. Mitliagkas, and L. Guibas, “Learning representations and generative models for 3d point clouds,” ICRLW, 2018.
  169. 169.M. Gadelha, R. Wang, and S. Maji, “Multiresolution tree networks for 3d point cloud processing,” in ECCV, pp. 103–118, 2018.
  170. 170.Y. Zhao, T. Birdal, H. Deng, and F. Tombari, “3d point capsule networks,” in CVPR, pp. 1009–1018, 2019.
  171. 171.R. Girdhar, D. F. Fouhey, M. Rodriguez, and A. Gupta, “Learning a predictable and generative vector representation for objects,” in ECCV, pp. 484–499, Springer, 2016.
  172. 172.A. Sharma, O. Grau, and M. Fritz, “Vconv-dae: Deep volumetric shape learning without object labels,” in ECCV, pp. 236–250, Springer, 2016.
  173. 173.L. Zhang and Z. Zhu, “Unsupervised feature learning for point cloud understanding by contrasting and clustering using graph convolutional neural networks,” in 3DV, pp. 395–404, IEEE, 2019.
  174. 174.K. Hassani and M. Haley, “Unsupervised multi-task feature learning on point clouds,” in ICCV, pp. 8160–8171, 2019.
  175. 175.Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon, “Dynamic graph cnn for learning on point clouds,” TOG, pp. 1–12, 2019.
  176. 176.Y. Feng, Y. Feng, H. You, X. Zhao, and Y. Gao, “Meshnet: mesh neural network for 3d shape representation,” in AAAI, pp. 8279–8286, 2019.
  177. 177.H. Su, S. Maji, E. Kalogerakis, and E. G. Learned-Miller, “Multi-view convolutional neural networks for 3d shape recognition,” in ICCV, 2015.
  178. 178.A. Kolesnikov, X. Zhai, and L. Beyer, “Revisiting self-supervised visual representation learning,” in CVPR, June 2019.
  179. 179.W. Hong, Z. Wang, M. Yang, and J. Yuan, “Conditional generative adversarial network for structured domain adaptation,” in CVPR, pp. 1335–1344, 2018.
  180. 180.Y. Cai, L. Ge, J. Cai, and J. Yuan, “Weakly-supervised 3d hand pose estimation from monocular rgb images,” in ECCV, pp. 666–682, 2018.
  181. 181.S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan, “Youtube-8m: A large-scale video classification benchmark,” arXiv preprint arXiv:1609.08675, 2016.
  182. 182.W. Li, L. Wang, W. Li, E. Agustsson, and L. Van Gool, “Webvision database: Visual learning and understanding from web data,” arXiv preprint arXiv:1708.02862, 2017.
  183. 183.L. Gomez, Y. Patel, M. Rusiñol, D. Karatzas, and C. Jawahar, “Self-supervised learning of visual features through embedding images into text topic spaces,” in CVPR, IEEE, 2017.
  184. 184.L. Jiang, Z. Zhou, T. Leung, L.-J. Li, and L. Fei-Fei, “Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels,” in ICML, 2018.
  185. 185.D. Ghadiyaram, D. Tran, and D. Mahajan, “Large-scale weakly-supervised pre-training for video action recognition,” in CVPR, pp. 12046–12055, 2019.
  186. 186.A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Evolving losses for unlabeled video representation learning,” CVPRW, 2019.

Citation

MLA
Jing, L., and Y. Tian. “Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey”. arXiv, 2019, http://arxiv.org/abs/1902.06162v1.
APA
Jing, L., & Tian, Y. (2019). Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey. arXiv. http://arxiv.org/abs/1902.06162v1
Chicago
Jing, L., and Y. Tian. 2019. “Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey”. arXiv. http://arxiv.org/abs/1902.06162v1.
Harvard
Jing, L. and Tian, Y. (2019) “Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1902.06162v1.
Vancouver
1. Jing L, Tian Y (2019) Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey. arXiv

BibTeX

@article{jing2019self,
  title = {Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey},
  author = {Jing, Longlong and Tian, Yingli},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1902.06162v1},
  eprint = {1902.06162}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF