Self-Supervised Learning: Generative or Contrastive

Xiao LiuFanjin ZhangZhenyu HouZhaoyu WangLi MianJing ZhangJie Tang

article2020TKDE2,245 citations

Classifies self-supervised representation learning across computer vision, natural language processing, and graph domains into generative, contrastive, and adversarial paradigms while connecting empirical architectures to their theoretical foundations.

Listen

Modern deep learning has achieved major breakthroughs across computer vision, natural language processing, and graph analytics, but traditional supervised approaches face critical bottlenecks. Supervised models require massive volumes of human-labeled data, which is slow and costly to producesuch as image segmentation labeling reaching millions of dollars for relatively small datasets. Furthermore, supervised models often suffer from poor generalization, vulnerability to adversarial disruptions, and brittle performance when encountering real-world data outside their training distribution. Self-supervised learning addresses these vulnerabilities by deriving supervisory signals directly from unlabeled data, allowing systems to learn robust representations by predicting or recovering hidden parts of an input from observed parts.

The article systematically evaluates empirical methodologies, theoretical foundations, and application domains across the self-supervised learning landscape. It categorizes existing approaches into three core frameworksgenerative, contrastive, and generative-contrastiveto evaluate their technical trade-offs and clarify why these models succeed across downstream operational tasks.

The review synthesizes findings across dozens of foundational and cutting-edge architectures spanning from 2012 through early 2021. The authors analyze auto-regressive, auto-encoding, and flow-based generative techniques, metric-learning contrastive approaches, and adversarial architectures, while evaluating mathematical frameworks such as mutual information maximization, variational lower bounds, and generalization bounds.

The analysis reveals several key findings. First, contrastive self-supervised learning has rapidly bridged the performance gap with fully supervised models on visual classification benchmarks, with advanced frameworks achieving near-supervised accuracy without requiring human annotations. Second, empirical and theoretical analyses demonstrate that mutual information maximization is only loosely tied to downstream success; rather, effective data augmentation, architecture design, and sampling strategies serve as the primary performance drivers. Third, although self-supervised pre-training does not increase the absolute peak accuracy ceiling when unlimited labels exist, it significantly improves data efficiencyallowing models using only 10% of labels to surpass fully supervised baselines when combined with semi-supervised self-training. Fourth, self-supervised models demonstrate superior operational robustness, showing heightened resistance to adversarial attacks, label corruption, and out-of-distribution shifts compared to standard supervised networks.

These findings have major strategic implications for enterprise AI deployment, project timelines, and development costs. Organizations can dramatically decrease manual data annotation expenses and mitigate operational risks stemming from brittle models in changing environments. While generative models remain indispensable for content creation and language modeling, contrastive frameworks provide lightweight, high-performing feature extractors tailored for classification and recognition tasks. Combining self-supervised pre-training with semi-supervised workflows provides a practical path to deploying high-performing machine learning systems with minimal labeled data.

Decision-makers and engineering teams should adopt self-supervised pre-training when developing classification pipelines with limited human labels, while carefully matching framework selection to specific operational requirements. Generative methods should be reserved for synthesis and sequence modeling, whereas contrastive paradigms should be deployed for discriminative classification tasks. Further applied research should focus on automating pretext task selection for downstream requirements, improving negative sampling efficiency, and establishing reliable data augmentation principles for discrete domains like natural language processing and structured graph learning.

While the underlying empirical evidence across computer vision benchmarks is highly robust, confidence should be tempered in certain domains. The mechanisms governing data augmentation remain theoretically underdeveloped, contrastive methods are vulnerable to early representation degeneration, and transferring learned representations across disparate graph structures remains an open challenge. Leaders should therefore validate self-supervised representations with targeted pilot evaluations before fully phasing out domain-specific supervised pipelines.

Cover for Self-Supervised Learning: Generative or Contrastive

Abstract

Deep supervised learning has achieved great success in the last decade. However, its deficiencies of dependence on manual labels and vulnerability to attacks have driven people to explore a better solution. As an alternative, self-supervised learning attracts many researchers for its soaring performance on representation learning in the last several years. Self-supervised representation learning leverages input data itself as supervision and benefits almost all types of downstream tasks. In this survey, we take a look into new self-supervised learning methods for representation in computer vision, natural language processing, and graph learning. We comprehensively review the existing empirical methods and summarize them into three main categories according to their objectives: generative, contrastive, and generative-contrastive (adversarial). We further investigate related theoretical analysis work to provide deeper thoughts on how self-supervised learning works. Finally, we briefly discuss open problems and future directions for self-supervised learning. An outline slide for the survey is provided.

Table of Contents

  • I Introduction
  • II Motivation of Self-supervised Learning
  • III Generative Self-supervised Learning
  • III-A Auto-regressive (AR) Model
  • III-B Flow-based Model
  • III-C Auto-encoding (AE) Model
  • III-C1 Basic AE Model
  • III-C2 Context Prediction Model (CPM)
  • III-C3 Denoising AE Model
  • III-C4 Variational AE Model
  • III-D Hybrid Generative Models
  • III-D1 Combining AR and AE Model.
  • III-D2 Combining AE and Flow-based Models
  • III-E Pros and Cons
  • IV Contrastive Self-supervised Learning
  • IV-A Context-Instance Contrast
  • IV-A1 Predict Relative Position
  • IV-A2 Maximize Mutual Information
  • IV-B Instance-Instance Contrast
  • IV-B1 Cluster Discrimination
  • IV-B2 Instance Discrimination
  • IV-C Self-supervised Contrastive Pre-training for Semi-supervised Self-training
  • IV-D Pros and Cons
  • V Generative-Contrastive (Adversarial) Self-supervised Learning
  • V-A Generate with Complete Input
  • V-B Recover with Partial Input
  • V-C Pre-trained Language Model
  • V-D Graph Learning
  • V-E Domain Adaptation and Multi-modality Representation
  • V-F Pros and Cons
  • VI Theory behind Self-supervised Learning
  • VI-A GAN
  • VI-A1 Divergence Matching
  • VI-A2 Disentangled Representation
  • VI-B Maximizing Lower Bound
  • VI-B1 Evidence Lower Bound
  • VI-B2 Mutual Information
  • VI-C Contrastive Self-supervised Representation Learning
  • VI-C1 Relationship with Supervised Learning
  • VI-C2 Understand Contrastive Loss
  • VI-C3 Generalization
  • VII Discussions and Future Directions
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomic and Conceptual Framework for Self-Supervised Learning

    model/method

    Self-supervised learning (SSL) algorithms leverage the inherent co-occurrence relationships and structure of unlabeled data to construct supervisory signals. All mainstream self-supervised representation learning frameworks can be unified into three primary categories based on their architectures and objectives:

    1. Generative SSL: Models an encoder fencf_{\text{enc}} to map input xx to an explicit latent representation zz, and a decoder fdecf_{\text{dec}} to reconstruct xx from zz. Training optimizes a point-wise reconstruction loss (such as mean squared error or negative log-likelihood under maximum likelihood estimation).
    2. Contrastive SSL: Employs an encoder to map input xx to an explicit representation zz, optimizing an objective that measures similarity in representation space (e.g., Noise Contrastive Estimation or InfoNCE). It discriminates positive pairs (similar views/contexts) from negative pairs without requiring a reconstruction decoder, utilizing lightweight projection heads/discriminators.
    3. Generative-Contrastive (Adversarial) SSL: Combines an encoder-decoder generator GG and a discriminator DD. Instead of sample-level point-wise reconstruction, the framework matches the generated data distribution to the empirical data distribution by minimizing statistical divergences (e.g., Jensen-Shannon divergence or Wasserstein distance) via minimax game objectives.

    Key architectural and operational differences among the three paradigms include:

    • Latent Representation zz: Explicitly modeled and directly extracted for downstream tasks in generative and contrastive paradigms; implicitly modeled in standard Generative Adversarial Networks (GANs), requiring bidirectional formulations (such as BiGAN/ALI) or auxiliary encoders to yield explicit representations.
    • Discriminator: Absent in pure generative methods, lightweight (e.g., 2–3 layer MLPs or inner products) in contrastive methods, and heavy (e.g., deep convolutional networks) in generative-contrastive methods.
    • Objective Type: Generative models minimize sample reconstruction loss; contrastive models optimize similarity bounds between instance/context representations; adversarial models minimize distributional divergences.
  2. Knowl 2 — Cross-Domain Categorization of Self-Supervised Representation Learning Methods

    data/table

    The landscape of self-supervised representation learning spans Computer Vision (CV), Natural Language Processing (NLP), and Graph Learning, differentiated by model type (Generative [G], Contrastive [C], Generative-Contrastive [G-C]), pretext task formulation, and positive/negative sampling strategies:

    Model Field Type Self-Supervision / Target Pretext Task Negative Sampling Strategy
    GPT / GPT-2 NLP G Following words Next word prediction -
    PixelCNN / PixelRNN CV G Following pixels Next pixel prediction -
    BERT / SpanBERT NLP G Masked words / spans Masked language model -
    VQ-VAE / VQ-VAE-2 CV G Quantized codebook / Whole image Image reconstruction -
    XLNet NLP G Permuted factorization Permutation language model -
    GraphAF Graph G Graph nodes edges Autoregressive flow generation -
    Deep InfoMax (DIM) CV C Global-local feature mutual info MI maximization End-to-end
    CPC Audio/CV C Context-future segment mutual info Contrastive predictive coding End-to-end
    DGI / InfoGraph Graph C Node/graph global-local MI Mutual info maximization Graph corruption / End-to-end
    DeepCluster / Local Agg. CV C Cluster assignments Pseudo-label discrimination K-means / Local soft-clustering
    SwAV CV C Swapped cluster prototype codes Swapped assignment prediction Online clustering / Multi-view
    MoCo / MoCo v2 CV C Instance identity Instance discrimination Momentum key encoder + Queue
    SimCLR / SimCLR v2 CV C Data augmentation views Pairwise instance discrimination Large minibatch (End-to-end)
    BYOL / SimSiam CV C Augmented views Representation regression No negative samples (Stop-gradient)
    GCC / GraphCL Graph C Subgraph instances / Augmentations Subgraph discrimination Momentum / End-to-end
    BiGAN / BigBiGAN CV G-C Joint distribution (x,z)(x, z) Adversarial feature learning Adversarial minimax
    ELECTRA / WKLM NLP G-C Corrupted tokens / entities Replaced token/entity detection Generator-discriminator (End-to-end)
    GraphGAN / GraphSGAN Graph G-C Connectivity / Density gaps Adversarial node/edge detection Policy gradient / Adversarial

    The table demonstrates that while generative methods rely on reconstruction likelihoods over discrete/continuous inputs without negative sampling, contrastive methods rely on instance discrimination across augmented views or global-local mutual information, utilizing memory banks, momentum queues, or large batches for negative sampling (with BYOL and SimSiam being notable exceptions avoiding negative samples via stop-gradient operations).

  3. Knowl 3 — Generative Self-Supervised Objectives: Autoregressive, Flow, and Masked Formulations

    equation

    Generative self-supervised learning models the input data distribution through explicit probabilistic factorizations and reconstruction objectives:

    1. Autoregressive (AR) Density Estimation: maxθpθ(x)=t=1Tlogpθ(xtx1:t1)\max_{\theta} p_\theta(x) = \sum_{t=1}^T \log p_\theta(x_t \mid x_{1:t-1}) where sequence tokens (or pixels) xtx_t are factorized sequentially along fixed directional orderings.

    2. Flow-Based Normalizing Flow Density Estimation: maxθilogpθ(x(i))=maxθi(logpZ(fθ(x(i)))+logdetfθ(x(i))x)\max_{\theta} \sum_i \log p_\theta(x^{(i)}) = \max_\theta \sum_i \left( \log p_Z(f_\theta(x^{(i)})) + \log \left| \det \frac{\partial f_\theta(x^{(i)})}{\partial x} \right| \right) where fθ:RdRdf_\theta: \mathbb{R}^d \to \mathbb{R}^d is an invertible and differentiable mapping from data space xx to a known base prior pZ(z)p_Z(z) (e.g., standard Gaussian).

    3. Variational Auto-Encoding (Evidence Lower Bound, ELBO): logp(x)Eqϕ(zx)[logpθ(xz)]DKL(qϕ(zx)p(z))=ELBO(θ,ϕ;x)\log p(x) \ge \mathbb{E}_{q_\phi(z \mid x)}[\log p_\theta(x \mid z)] - D_{\text{KL}}(q_\phi(z \mid x) \parallel p(z)) = \text{ELBO}(\theta, \phi; x) where qϕ(zx)q_\phi(z \mid x) approximates the true posterior pθ(zx)p_\theta(z \mid x), DKLD_{\text{KL}} regularizes the latent code against prior p(z)p(z), and E[logpθ(xz)]\mathbb{E}[\log p_\theta(x \mid z)] denotes reconstruction log-likelihood.

    4. Permutation Language Modeling (PLM): maxθEzZT[t=1Tlogpθ(xztxz<t)]\max_\theta \mathbb{E}_{z \sim \mathcal{Z}_T} \left[ \sum_{t=1}^T \log p_\theta(x_{z_t} \mid x_{z_{<t}}) \right] where ZT\mathcal{Z}_T denotes the set of all possible permutations of index sequence [1,2,,T][1, 2, \dots, T], permitting context conditioning from both directions without structural leakage.

  4. Knowl 4 — Contrastive Self-Supervised Learning Objectives and Taxonomy

    model/method

    Contrastive self-supervised learning optimizes representation spaces by contrasting positive (semantically related) sample pairs against negative (unrelated) sample pairs. The foundational objective is Noise Contrastive Estimation (NCE) and its multi-negative extension InfoNCE:

    LInfoNCE=Ex,x+,{xk}k=1K[logexp(sim(f(x),f(x+))τ)exp(sim(f(x),f(x+))τ)+k=1Kexp(sim(f(x),f(xk))τ)]\mathcal{L}_{\text{InfoNCE}} = \mathbb{E}_{x, x^+, \{x_k^-\}_{k=1}^K} \left[ -\log \frac{\exp\left(\frac{\text{sim}(f(x), f(x^+))}{\tau}\right)}{\exp\left(\frac{\text{sim}(f(x), f(x^+))}{\tau}\right) + \sum_{k=1}^K \exp\left(\frac{\text{sim}(f(x), f(x_k^-))}{\tau}\right)} \right]

    where f()f(\cdot) is an encoder mapping inputs to normalized embedding vectors, sim(u,v)=uTvu2v2\text{sim}(u, v) = \frac{u^T v}{\|u\|_2 \|v\|_2} is cosine similarity, τ>0\tau > 0 is a temperature hyperparameter, x+x^+ is a positive counterpart to anchor xx, and {xk}\{x_k^-\} are KK negative samples.

    Contrastive learning is categorized into two main paradigms:

    • Context-Instance (Global-Local) Contrast: Optimizes the associative belonging between a local feature vector vv and a global context summary vector s=g(f(x))s = g(f(x)). Subtypes include:
      • Predict Relative Position (PRP): Recovers spatial patch orders (jigsaw puzzles), rotation angles, or sentence ordering (e.g., Sentence Order Prediction).
      • Mutual Information (MI) Maximization: Maximizes local-to-global MI using summary readout functions (e.g., Deep InfoMax, Contrastive Predictive Coding [CPC], Deep Graph InfoMax [DGI]).
    • Instance-Instance Contrast: Discards global-local pooling and directly measures relationships between whole instance representations:
      • Cluster Discrimination: Generates pseudo-cluster labels to optimize cross-entropy or local soft-clustering objectives (e.g., DeepCluster, Local Aggregation, SwAV).
      • Instance Discrimination: Treats each individual sample as its own class, drawing augmented views of the same sample together while repelling distinct instance views (e.g., InstDisc, MoCo, SimCLR, BYOL, SimSiam).
  5. Knowl 5 — Generative-Contrastive (Adversarial) Representation Learning Framework

    model/method

    Generative-contrastive (adversarial) self-supervised learning incorporates adversarial game objectives to overcome the limitations of point-wise reconstruction losses. The core frameworks operate as follows:

    1. Bidirectional Adversarial Feature Learning (BiGAN / ALI): Unlike standard GANs that map noise zpz(z)z \sim p_z(z) to samples G(z)G(z) with an implicit latent space, BiGAN introduces an encoder E(x)E(x) alongside generator G(z)G(z). The discriminator D(x,z)D(x, z) is trained to distinguish joint data-latent pairs (x,E(x))(x, E(x)) from (G(z),z)(G(z), z): minG,EmaxDExpdata(x)[logD(x,E(x))]+Ezpz(z)[log(1D(G(z),z))]\min_{G, E} \max_D \mathbb{E}_{x \sim p_{\text{data}}(x)} [\log D(x, E(x))] + \mathbb{E}_{z \sim p_z(z)} [\log (1 - D(G(z), z))] At theoretical optimality, the encoder inverts the generator (E=G1E = G^{-1}), capturing global semantic representations.

    2. Partial-Input Adversarial Recovery: Instead of whole-image synthesis, models recover corrupted or masked input channels (e.g., colorization, inpainting, super-resolution) while a discriminator ensures the realism of the completed output.

    3. Replaced Token Detection (ELECTRA): For discrete language tokens where standard GAN backpropagation fails, a small masked language model serves as generator GG to predict masked words, while discriminator DD solves a binary classification task determining for each token in sequence xx whether it is original or replaced: minθG,θDxXLMLM(x;θG)+λLDisc(x;θD)\min_{\theta_G, \theta_D} \sum_{x \in \mathcal{X}} \mathcal{L}_{\text{MLM}}(x; \theta_G) + \lambda \mathcal{L}_{\text{Disc}}(x; \theta_D) This converts a kk-way categorical prediction into dense token-level binary classification over all positions.

  6. Knowl 6 — Information-Theoretic Lower Bounds and the Metric Learning Equivalence of InfoNCE

    theoretical result

    Contrastive self-supervised objectives rooted in mutual information maximization optimize a variational lower bound on the mutual information I(X;Y)I(X; Y) between two representations XX and YY:

    I(X;Y)=Ep(x,y)[logp(x,y)p(x)p(y)]log(K)LInfoNCEI(X; Y) = \mathbb{E}_{p(x, y)} \left[ \log \frac{p(x, y)}{p(x) p(y)} \right] \ge \log(K) - \mathcal{L}_{\text{InfoNCE}}

    where KK is the number of negative samples. As KK \to \infty, the bound becomes tighter.

    When the critic function f(x,y)f(x, y) in InfoNCE is constrained to the inner product of representations ϕ(x)Tϕ(y)\phi(x)^T \phi(y), the InfoNCE loss is mathematically equivalent to the multi-class kk-pair loss from deep metric learning:

    Lk-pair(ϕ)=1ki=1klog(1+jiexp(ϕ(xi)Tϕ(yj)ϕ(xi)Tϕ(yi)))\mathcal{L}_{k\text{-pair}}(\phi) = \frac{1}{k} \sum_{i=1}^k \log \left( 1 + \sum_{j \neq i} \exp\left( \phi(x_i)^T \phi(y_j) - \phi(x_i)^T \phi(y_i) \right) \right)

    Theoretical Limitations of MI/ELBO Maximization: Maximizing mutual information or evidence lower bounds (ELBO) does not strictly correlate with downstream task performance:

    • Representations achieving identical numerical lower bound values can exhibit vastly disparate downstream transfer accuracy.
    • Tighter bounds on mutual information do not necessarily produce better features; looser lower bounds frequently achieve higher classification accuracy on downstream evaluation tasks because the useful inductive bias arises from encoder architectures and sampling strategies rather than exact mutual information values.
  7. Knowl 7 — Decomposition of Contrastive Loss into Alignment and Uniformity

    theoretical result

    The contrastive loss asymptotically decomposes on the unit hypersphere into two distinct properties: alignment of positive pairs and uniformity of the induced feature distribution:

    Lcontrast=E(x,y)ppos[f(x)Tf(y)τ]+Expdata[log(ef(x)Tf(y)τ+i=1Kef(x)Tf(yi)τ)]\mathcal{L}_{\text{contrast}} = \mathbb{E}_{(x, y) \sim p_{\text{pos}}} \left[ -\frac{f(x)^T f(y)}{\tau} \right] + \mathbb{E}_{x \sim p_{\text{data}}} \left[ \log \left( e^{\frac{f(x)^T f(y)}{\tau}} + \sum_{i=1}^K e^{\frac{f(x)^T f(y_i^-)}{\tau}} \right) \right]

    Under spherical normalization, the two properties can be explicitly formalized and optimized as independent objectives:

    1. Alignment Loss (pulls positive representations close together): Lalign(f;α)E(x,y)ppos[f(x)f(y)2α],α>0\mathcal{L}_{\text{align}}(f; \alpha) \triangleq \mathbb{E}_{(x, y) \sim p_{\text{pos}}} \left[ \|f(x) - f(y)\|_2^\alpha \right], \quad \alpha > 0

    2. Uniformity Loss (encourages representations to be uniformly distributed on the unit hypersphere, maximizing information retention): Luniform(f;t)logEx,yi.i.d.pdata[etf(x)f(y)22],t>0\mathcal{L}_{\text{uniform}}(f; t) \triangleq \log \mathbb{E}_{x, y \stackrel{\text{i.i.d.}}{\sim} p_{\text{data}}} \left[ e^{-t \|f(x) - f(y)\|_2^2} \right], \quad t > 0

    Optimizing a weighted combination Lalign+λLuniform\mathcal{L}_{\text{align}} + \lambda \mathcal{L}_{\text{uniform}} directly matches or surpasses standard contrastive loss performance across architectures. Both alignment and uniformity are essential: if either loss is weighted excessively, representations either collapse to a single point (loss of uniformity) or fail to cluster semantically (loss of alignment).

  8. Knowl 8 — Generalization Guarantees for Contrastive Representation Learning on Downstream Tasks

    theoretical result

    Let similar pairs (x,x+)(x, x^+) be drawn from distribution Dsim\mathcal{D}_{\text{sim}} (sharing a latent class cCc \in \mathcal{C}) and negative samples xx^- drawn from marginal distribution Dneg\mathcal{D}_{\text{neg}}. The unsupervised contrastive loss is:

    Lun(f)=Ex,x+Dsim,xDneg[(f(x)T(f(x+)f(x)))]\mathcal{L}_{\text{un}}(f) = \mathbb{E}_{x, x^+ \sim \mathcal{D}_{\text{sim}}, x^- \sim \mathcal{D}_{\text{neg}}} \left[ \ell\left( f(x)^T (f(x^+) - f(x^-)) \right) \right]

    Decomposing the negative sampling expectation according to whether xx^- shares the latent class of xx yields Lun(f)=τLun=(f)+(1τ)Lun(f)\mathcal{L}_{\text{un}}(f) = \tau \mathcal{L}_{\text{un}}^{=} (f) + (1 - \tau) \mathcal{L}_{\text{un}}^{\neq}(f), where Lun\mathcal{L}_{\text{un}}^{\neq} denotes the loss over negative pairs drawn from differing classes and s(f)s(f) is the intraclass deviation.

    For a supervised mean classifier loss Lsupμ(f^)\mathcal{L}_{\text{sup}}^\mu(\hat{f}) on linear downstream tasks, minimizing the empirical unsupervised loss with empirical minimizer f^\hat{f} satisfies with probability at least 1δ1 - \delta:

    Lsup(f^)Lsupμ(f^)Lun(f)+βs(f)+ηGenM\mathcal{L}_{\text{sup}}(\hat{f}) \le \mathcal{L}_{\text{sup}}^\mu(\hat{f}) \le \mathcal{L}_{\text{un}}^{\neq}(f) + \beta s(f) + \eta \text{Gen}_M

    where GenM\text{Gen}_M is the generalization error for sample size MM, vanishing (GenM0,δ0\|\text{Gen}_M\| \to 0, \delta \to 0) as MM \to \infty and C|\mathcal{C}| \to \infty.

    When representations f(X)f(X) are σ2\sigma^2-sub-Gaussian per class with bounded norm R=maxxf(x)R = \max_{x} \|f(x)\|, downstream supervised error is bounded by:

    Lsupμ(f^)γ(f)Lγ(f),supμ(f)+βs(f)+ηGenM+ϵ\mathcal{L}_{\text{sup}}^\mu(\hat{f}) \le \gamma(f) \mathcal{L}_{\gamma(f), \text{sup}}^\mu(f) + \beta s(f) + \eta \text{Gen}_M + \epsilon

    with γ(f)=1+cRσlog(R/ϵ)\gamma(f) = 1 + c' R \sigma \sqrt{\log(R / \epsilon)}. Thus, optimizing the unsupervised contrastive surrogate provably minimizes downstream supervised classification error.

  9. Knowl 9 — Combined Self-Supervised Contrastive Pre-Training and Semi-Supervised Self-Training

    model/method

    Self-supervised pre-training and semi-supervised self-training provide complementary, orthogonal improvements for data-efficient learning. The unified three-stage paradigm (e.g., SimCLR v2) operates as follows:

    1. Unsupervised Pre-Training: Pre-train a deep neural network encoder (e.g., ResNet) on large-scale unlabeled data using contrastive instance discrimination loss (such as NT-Xent / InfoNCE) to learn general visual/structural representations.
    2. Supervised Fine-Tuning: Fine-tune the pre-trained network on a small subset of labeled data (e.g., 1% or 10% of class labels), adapting high-level projection features to task-specific label spaces.
    3. Self-Training with Pseudo-Label Distillation: Utilize the fine-tuned network as a frozen teacher model to generate pseudo-labels on the remaining large unlabeled dataset. Train a student model (which can be smaller or identical in capacity) to predict both real and teacher-generated pseudo-labels, injecting strong regularization noise (e.g., dropout, stochastic depth, RandAugment) exclusively into student training.

    Empirical findings demonstrate that while pure self-supervised pre-training plateaus in efficacy when downstream task supervision diverges (e.g., from instance discrimination to dense object detection), semi-supervised self-training steadily improves with unlabeled data volume, and their combination surpasses fully supervised models trained from scratch on the full label set.

  10. Knowl 10 — Pathologies and Open Challenges in Self-Supervised Representation Learning

    limitation

    Current self-supervised representation learning paradigms exhibit several fundamental theoretical and practical limitations:

    1. Early Degeneration in Contrastive Learning: Contrastive representations risk over-fitting prematurely to the discriminative pretext task, causing dimensional collapse in the embedding space. This optimization bias aids downstream linear classification but impairs transfer to structured dense prediction, sequence generation, or fine-grained entity extraction.
    2. Pathologies of Point-Wise Generative Objectives: Maximum Likelihood Estimation (MLE) LMLE=xlogp(xc)\mathcal{L}_{\text{MLE}} = -\sum_x \log p(x \mid c) suffers from extreme sensitivity to rare out-of-distribution samples (as p(xc)0p(x \mid c) \to 0, loss diverges), leading to overly conservative probability distributions. Furthermore, MLE models low-level point-wise details (pixels, raw tokens) rather than high-level semantic abstractions.
    3. Domain Transfer Discrepancies: While contrastive methods excel in continuous visual domains with smooth augmentations (cropping, color jittering), applying contrastive learning and data augmentations to discrete, abstract structures (NLP tokens, graph topologies) is constrained by semantic shifts.
    4. Asymptotic Diminishing Returns of Negative Sampling: In contrastive learning, increasing the number of negative samples KK reduces gradient variance up to a threshold, beyond which performance degrades or plateaus while imposing significant memory and compute overheads.

Coverage note — None was omitted. All primary contributions—including the three-way taxonomic categorization, unified generator-discriminator conceptual framework, mathematical formulations across generative, contrastive, and adversarial paradigms, theoretical proofs/bounds (InfoNCE-to-metric equivalence, alignment/uniformity decomposition, generalization guarantees), cross-domain method comparisons, semi-supervised integration, and fundamental limitations—have been comprehensively captured.

References

  1. 1.H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, and M. Marchand. Domain-adversarial neural networks. arXiv preprint arXiv:1412.4446, 2014.
  2. 2.F. Alam, S. Joty, and M. Imran. Domain adaptation with adversarial training and graph embeddings. arXiv preprint arXiv:1805.05151, 2018.
  3. 3.A. A. Alemi, B. Poole, I. Fischer, J. V. Dillon, R. A. Saurous, and K. Murphy. Fixing a broken elbo. arXiv preprint arXiv:1711.00464, 2017.
  4. 4.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214–223. PMLR, 2017.
  5. 5.S. Arora, H. Khandeparkar, M. Khodak, O. Plevrakis, and N. Saunshi. A theoretical analysis of contrastive unsupervised representation learning. arXiv preprint arXiv:1902.09229, 2019.
  6. 6.A. Asai, K. Hashimoto, H. Hajishirzi, R. Socher, and C. Xiong. Learning to retrieve reasoning paths over wikipedia graph for question answering. arXiv preprint arXiv:1911.10470, 2019.
  7. 7.P. Bachman, R. D. Hjelm, and W. Buchwalter. Learning representations by maximizing mutual information across views. In NIPS, pages 15509–15519, 2019.
  8. 8.Y. Bai, H. Ding, S. Bian, T. Chen, Y. Sun, and W. Wang. Simgnn: A neural network approach to fast graph similarity computation. In WSDM, pages 384–392, 2019.
  9. 9.D. H. Ballard. Modular learning in neural networks. In AAAI, pages 279–284, 1987.
  10. 10.D. Bau, J.-Y. Zhu, H. Strobelt, B. Zhou, J. B. Tenenbaum, W. T. Freeman, and A. Torralba. Gan dissection: Visualizing and understanding generative adversarial networks. arXiv preprint arXiv:1811.10597, 2018.
  11. 11.Y. Bengio, N. Leonard, and A. Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  12. 12.M. Besserve, R. Sun, and B. Scholkopf. Counterfactuals uncover the modular structure of deep generative models. arXiv preprint arXiv:1812.03253, 2018.
  13. 13.Y. Blau and T. Michaeli. Rethinking lossy compression: The rate-distortion-perception tradeoff. arXiv preprint arXiv:1901.07821, 2019.
  14. 14.P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146, 2017.
  15. 15.A. Brock, J. Donahue, and K. Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  16. 16.L. Cai and W. Y. Wang. Kbgan: Adversarial learning for knowledge graph embeddings. arXiv preprint arXiv:1711.04071, 2017.
  17. 17.M. Caron, P. Bojanowski, A. Joulin, and M. Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the ECCV (ECCV), pages 132–149, 2018.
  18. 18.M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882, 2020.
  19. 19.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020.
  20. 20.T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020.
  21. 21.T. Chen, Y. Sun, Y. Shi, and L. Hong. On sampling strategies for neural network-based collaborative filtering. In SIGKDD, pages 767–776, 2017.
  22. 22.X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In NIPS, pages 2172–2180, 2016.
  23. 23.X. Chen, H. Fan, R. Girshick, and K. He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020.
  24. 24.X. Chen and K. He. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020.
  25. 25.L. Chongxuan, T. Xu, J. Zhu, and B. Zhang. Triple generative adversarial nets. In NIPS, pages 4088–4098, 2017.
  26. 26.K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020.
  27. 27.A. Conneau, G. Lample, M. Ranzato, L. Denoyer, and H. Jegou. Word translation without parallel data. arXiv preprint arXiv:1710.04087, 2017.
  28. 28.Q. Dai, Q. Li, J. Tang, and D. Wang. Adversarial network embedding. In AAAI, 2018.
  29. 29.Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, 2019.
  30. 30.V. R. de Sa. Learning classification with unlabeled data. In NIPS, pages 112–119, 1994.
  31. 31.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009.
  32. 32.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, 2019.
  33. 33.M. Ding, J. Tang, and J. Zhang. Semi-supervised learning on graphs with generative adversarial nets. In Proceedings of the 27th ACM CIKM, pages 913–922, 2018.
  34. 34.M. Ding, C. Zhou, Q. Chen, H. Yang, and J. Tang. Cognitive graph for multi-hop reading comprehension at scale. arXiv preprint arXiv:1905.05460, 2019.
  35. 35.L. Dinh, D. Krueger, and Y. Bengio. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014.
  36. 36.L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using real nvp. arXiv preprint arXiv:1605.08803, 2016.
  37. 37.C. Doersch, A. Gupta, and A. A. Efros. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE ICCV, pages 1422–1430, 2015.
  38. 38.J. Donahue, P. Krahenb uhl, and T. Darrell. Adversarial feature learning. arXiv preprint arXiv:1605.09782, 2016.
  39. 39.J. Donahue and K. Simonyan. Large scale adversarial representation learning. In NIPS, pages 10541–10551, 2019.
  40. 40.C. Donnat, M. Zitnik, D. Hallac, and J. Leskovec. Learning structural node embeddings via diffusion wavelets. In SIGKDD, pages 1320–1329, 2018.
  41. 41.V. Dumoulin, I. Belghazi, B. Poole, O. Mastropietro, A. Lamb, M. Arjovsky, and A. Courville. Adversarially learned inference. arXiv preprint arXiv:1606.00704, 2016.
  42. 42.Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky. Domain-adversarial training of neural networks. The Journal of Machine Learning Research, 17(1):2096–2030, 2016.
  43. 43.S. Gidaris, P. Singh, and N. Komodakis. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728, 2018.
  44. 44.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, pages 580–587, 2014.
  45. 45.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, pages 2672–2680, 2014.
  46. 46.P. Goyal, M. Caron, B. Lefaudeux, M. Xu, P. Wang, V. Pai, M. Singh, V. Liptchinsky, I. Misra, A. Joulin, and P. Bojanowski. Self-supervised pretraining of visual features in the wild. arXiv preprint arXiv:2103.01988, 2021.
  47. 47.J.-B. Grill, F. Strub, F. Altche, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  48. 48.A. Grover and J. Leskovec. node2vec: Scalable feature learning for networks. In SIGKDD, pages 855–864, 2016.
  49. 49.M. Gutmann and A. Hyvarinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pages 297–304, 2010.
  50. 50.K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909, 2020.
  51. 51.K. Hassani and A. H. Khasahmadi. Contrastive multi-view representation learning on graphs. arXiv preprint arXiv:2006.05582, 2020.
  52. 52.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722, 2019.
  53. 53.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  54. 54.D. Hendrycks, M. Mazeika, S. Kadavath, and D. Song. Using self-supervised learning can improve model robustness and uncertainty. In NeurIPS, pages 15663–15674, 2019.
  55. 55.R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018.
  56. 56.J. Ho, X. Chen, A. Srinivas, Y. Duan, and P. Abbeel. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In ICML, pages 2722–2730, 2019.
  57. 57.W. Hu, B. Liu, J. Gomes, M. Zitnik, P. Liang, V. Pande, and J. Leskovec. Strategies for pre-training graph neural networks. In ICLR, 2019.
  58. 58.Z. Hu, Y. Dong, K. Wang, K.-W. Chang, and Y. Sun. Gpt-gnn: Generative pre-training of graph neural networks. arXiv preprint arXiv:2006.15437, 2020.
  59. 59.Z. Hu, Y. Dong, K. Wang, and Y. Sun. Heterogeneous graph transformer. arXiv preprint arXiv:2003.01332, 2020.
  60. 60.G. Huang, Z. Liu, and K. Q. Weinberger. Densely connected convolutional networks. 2017 IEEE CVPR, pages 2261–2269, 2017.
  61. 61.S. Iizuka, E. Simo-Serra, and H. Ishikawa. Globally and locally consistent image completion. ACM Transactions on Graphics (ToG), 36(4):1–14, 2017.
  62. 62.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In CVPR, pages 1125–1134, 2017.
  63. 63.L. Jing and Y. Tian. Self-supervised visual feature learning with deep neural networks: A survey. arXiv preprint arXiv:1902.06162, 2019.
  64. 64.M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77, 2020.
  65. 65.V. Karpukhin, B. Oguz, S. Min, L. Wu, S. Edunov, D. Chen, and W.-t. Yih. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906, 2020.
  66. 66.T. Karras, S. Laine, and T. Aila. A style-based generator architecture for generative adversarial networks. In CVPR, pages 4401–4410, 2019.
  67. 67.D. Kim, D. Cho, D. Yoo, and I. S. Kweon. Learning image representations by completing damaged jigsaw puzzles. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 793–802. IEEE, 2018.
  68. 68.D. P. Kingma and P. Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. In NIPS, pages 10215–10224, 2018.
  69. 69.D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  70. 70.T. N. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  71. 71.T. N. Kipf and M. Welling. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308, 2016.
  72. 72.L. Kong, C. d. M. d’Autume, W. Ling, L. Yu, Z. Dai, and D. Yogatama. A mutual information maximization perspective of language representation learning. arXiv preprint arXiv:1910.08350, 2019.
  73. 73.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, pages 1097–1105, 2012.
  74. 74.Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.
  75. 75.G. Larsson, M. Maire, and G. Shakhnarovich. Learning representations for automatic colorization. In ECCV, pages 577–593. Springer, 2016.
  76. 76.G. Larsson, M. Maire, and G. Shakhnarovich. Colorization as a proxy task for visual understanding. In CVPR, pages 6874–6883, 2017.
  77. 77.Y. LeCun, Y. Bengio, and G. Hinton. Deep learning. nature, 521(7553):436–444, 2015.
  78. 78.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photorealistic single image super-resolution using a generative adversarial network. In CVPR, pages 4681–4690, 2017.
  79. 79.D. Li, W.-C. Hung, J.-B. Huang, S. Wang, N. Ahuja, and M.-H. Yang. Unsupervised visual representation learning by graphbased consistent constraints. In ECCV, pages 678–694. Springer, 2016.
  80. 80.B. Liu. Sentiment analysis and opinion mining. Synthesis lectures on human language technologies, 5(1):1–167, 2012.
  81. 81.Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  82. 82.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015.
  83. 83.A. Makhzani, J. Shlens, N. Jaitly, I. Goodfellow, and B. Frey. Adversarial autoencoders. arXiv preprint arXiv:1511.05644, 2015.
  84. 84.M. Mathieu. Masked autoencoder for distribution estimation. 2015.
  85. 85.T. Mikolov, K. Chen, G. S. Corrado, and J. Dean. Efficient estimation of word representations in vector space. CoRR, abs/1301.3781, 2013.
  86. 86.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In NIPS’13, pages 3111–3119, 2013.
  87. 87.I. Misra and L. van der Maaten. Self-supervised learning of pretext-invariant representations. arXiv preprint arXiv:1912.01991, 2019.
  88. 88.J. Mitrovic, B. McWilliams, J. Walker, L. Buesing, and C. Blundell. Representation learning via invariant causal mechanisms. arXiv preprint arXiv:2010.07922, 2020.
  89. 89.T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957, 2018.
  90. 90.A. Newell and J. Deng. How useful is self-supervised pretraining for visual tasks? In CVPR, pages 7345–7354, 2020.
  91. 91.A. Ng et al. Sparse autoencoder. CS294A Lecture notes, 72(2011):1–19, 2011.
  92. 92.M. Noroozi and P. Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV, pages 69–84. Springer, 2016.
  93. 93.M. Noroozi, A. Vinjimoor, P. Favaro, and H. Pirsiavash. Boosting self-supervised learning via knowledge transfer. In CVPR, pages 9359–9367, 2018.
  94. 94.S. Nowozin, B. Cseke, and R. Tomioka. f-gan: Training generative neural samplers using variational divergence minimization. In NIPS, pages 271–279, 2016.
  95. 95.A. v. d. Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  96. 96.D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros. Context encoders: Feature learning by inpainting. In CVPR, pages 2536–2544, 2016.
  97. 97.Z. Peng, Y. Dong, M. Luo, X. ming Wu, and Q. Zheng. Self-supervised graph representation learning via global context prediction. ArXiv, abs/2003.01604, 2020.
  98. 98.Z. Peng, Y. Dong, M. Luo, X.-M. Wu, and Q. Zheng. Self-supervised graph representation learning via global context prediction. arXiv preprint arXiv:2003.01604, 2020.
  99. 99.B. Perozzi, R. Al-Rfou, and S. Skiena. Deepwalk: Online learning of social representations. In SIGKDD, pages 701–710, 2014.
  100. 100.M. Popova, M. Shvets, J. Oliva, and O. Isayev. Molecularrnn: Generating realistic molecular graphs with optimized properties. arXiv preprint arXiv:1905.13372, 2019.
  101. 101.J. Qiu, Q. Chen, Y. Dong, J. Zhang, H. Yang, M. Ding, K. Wang, and J. Tang. Gcc: Graph contrastive coding for graph neural network pre-training. arXiv preprint arXiv:2006.09963, 2020.
  102. 102.J. Qiu, J. Tang, H. Ma, Y. Dong, K. Wang, and J. Tang. Deepinf: Social influence prediction with deep learning. In KDD’18, pages 2110–2119. ACM, 2018.
  103. 103.X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang. Pre-trained models for natural language processing: A survey. arXiv preprint arXiv:2003.08271, 2020.
  104. 104.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.
  105. 105.A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever. Improving language understanding by generative pre-training.
  106. 106.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019.
  107. 107.P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  108. 108.A. Razavi, A. van den Oord, and O. Vinyals. Generating diverse high-fidelity images with vq-vae-2. In NIPS, pages 14837–14847, 2019.
  109. 109.N. Reimers and I. Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  110. 110.L. F. Ribeiro, P. H. Saverese, and D. R. Figueiredo. struc2vec: Learning node representations from structural identity. In SIGKDD, pages 385–394, 2017.
  111. 111.M. T. Ribeiro, S. Singh, and C. Guestrin. ” why should i trust you?” explaining the predictions of any classifier. In SIGKDD, pages 1135–1144, 2016.
  112. 112.N. Sarafianos, X. Xu, and I. A. Kakadiaris. Adversarial representation learning for text-to-image matching. In Proceedings of the IEEE ICCV, pages 5814–5824, 2019.
  113. 113.J. Shen, Y. Qu, W. Zhang, and Y. Yu. Adversarial representation learning for domain adaptation. stat, 1050:5, 2017.
  114. 114.T. Shen, T. Lei, R. Barzilay, and T. Jaakkola. Style transfer from non-parallel text by cross-alignment. In NIPS, pages 6830–6841, 2017.
  115. 115.C. Shi, M. Xu, Z. Zhu, W. Zhang, M. Zhang, and J. Tang. Graphaf: a flow-based autoregressive model for molecular graph generation. arXiv preprint arXiv:2001.09382, 2020.
  116. 116.A. Sinha, Z. Shen, Y. Song, H. Ma, D. Eide, B.-j. P. Hsu, and K. Wang. An overview of microsoft academic service (mas) and applications. In WWW’15, pages 243–246, 2015.
  117. 117.P. Smolensky. Information processing in dynamical systems: Foundations of harmony theory. Technical report, Colorado Univ at Boulder Dept of Computer Science, 1986.
  118. 118.F.-Y. Sun, J. Hoffmann, and J. Tang. Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. arXiv preprint arXiv:1908.01000, 2019.
  119. 119.F.-Y. Sun, M. Qu, J. Hoffmann, C.-W. Huang, and J. Tang. vgraph: A generative model for joint community detection and node representation learning. In NIPS, pages 512–522, 2019.
  120. 120.K. Sun, Z. Lin, and Z. Zhu. Multi-stage self-supervised learning for graph convolutional networks on graphs with few labeled nodes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 5892–5899, 2020.
  121. 121.K. Sun, Z. Zhu, and Z. Lin. Multi-stage self-supervised learning for graph convolutional networks. arXiv preprint arXiv:1902.11038, 2019.
  122. 122.Y. Sun, S. Wang, Y. Li, S. Feng, X. Chen, H. Zhang, X. Tian, D. Zhu, H. Tian, and H. Wu. Ernie: Enhanced representation through knowledge integration. arXiv preprint arXiv:1904.09223, 2019.
  123. 123.J. Tang, M. Qu, M. Wang, M. Zhang, J. Yan, and Q. Mei. Line: Large-scale information network embedding. In WWW’15, pages 1067–1077, 2015.
  124. 124.J. Tang, J. Zhang, L. Yao, J. Li, L. Zhang, and Z. Su. Arnetminer: extraction and mining of academic social networks. In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 990–998, 2008.
  125. 125.W. L. Taylor. “cloze procedure”: A new tool for measuring readability. Journalism quarterly, 30(4):415–433, 1953.
  126. 126.Y. Tian, D. Krishnan, and P. Isola. Contrastive multiview coding. arXiv preprint arXiv:1906.05849, 2019.
  127. 127.Y. Tian, C. Sun, B. Poole, D. Krishnan, C. Schmid, and P. Isola. What makes for good views for contrastive learning. arXiv preprint arXiv:2005.10243, 2020.
  128. 128.M. Tschannen, O. Bachem, and M. Lucic. Recent advances in autoencoder-based representation learning. arXiv preprint arXiv:1812.05069, 2018.
  129. 129.M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic. On mutual information maximization for representation learning. arXiv preprint arXiv:1907.13625, 2019.
  130. 130.A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu. Wavenet: A generative model for raw audio. In 9th ISCA Speech Synthesis Workshop, pages 125–125.
  131. 131.A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al. Conditional image generation with pixelcnn decoders. In NIPS, pages 4790–4798, 2016.
  132. 132.A. van den Oord, O. Vinyals, et al. Neural discrete representation learning. In NIPS, pages 6306–6315, 2017.
  133. 133.A. Van Oord, N. Kalchbrenner, and K. Kavukcuoglu. Pixel recurrent neural networks. In ICML, pages 1747–1756, 2016.
  134. 134.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In NIPS, pages 5998–6008, 2017.
  135. 135.P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  136. 136.P. Velickovic, W. Fedus, W. L. Hamilton, P. Li o, Y. Bengio, and ` R. D. Hjelm. Deep graph infomax. arXiv preprint arXiv:1809.10341, 2018.
  137. 137.H. Wang, J. Wang, J. Wang, M. Zhao, W. Zhang, F. Zhang, X. Xie, and M. Guo. Graphgan: Graph representation learning with generative adversarial nets. In AAAI, 2018.
  138. 138.P. Wang, S. Li, and R. Pan. Incorporating gan for negative sampling in knowledge representation learning. In AAAI, 2018.
  139. 139.T. Wang and P. Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. arXiv preprint arXiv:2005.10242, 2020.
  140. 140.Z. Wang, Q. She, and T. E. Ward. Generative adversarial networks: A survey and taxonomy. arXiv preprint arXiv:1906.01529, 2019.
  141. 141.C. Wei, L. Xie, X. Ren, Y. Xia, C. Su, J. Liu, Q. Tian, and A. L. Yuille. Iterative reorganization with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning. In CVPR, pages 1910–1919, 2019.
  142. 142.Z. Wu, Y. Xiong, S. X. Yu, and D. Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, pages 3733–3742, 2018.
  143. 143.Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le. Self-training with noisy student improves imagenet classification. In CVPR, pages 10687–10698, 2020.
  144. 144.W. Xiong, J. Du, W. Y. Wang, and V. Stoyanov. Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model. arXiv preprint arXiv:1912.09637, 2019.
  145. 145.K. Xu, J. Li, M. Zhang, S. S. Du, K.-i. Kawarabayashi, and S. Jegelka. How neural networks extrapolate: From feedforward to graph neural networks. arXiv preprint arXiv:2009.11848, 2020.
  146. 146.X. Yan, I. Misra, A. Gupta, D. Ghadiyaram, and D. Mahajan. Clusterfit: Improving generalization of visual representations. arXiv preprint arXiv:1912.03330, 2019.
  147. 147.J. Yang, D. Parikh, and D. Batra. Joint unsupervised learning of deep representations and image clusters. In CVPR, pages 5147–5156, 2016.
  148. 148.Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le. Xlnet: Generalized autoregressive pretraining for language understanding. In NIPS, pages 5754–5764, 2019.
  149. 149.Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600, 2018.
  150. 150.J. You, B. Liu, Z. Ying, V. Pande, and J. Leskovec. Graph convolutional policy network for goal-directed molecular graph generation. In NIPS, pages 6410–6421, 2018.
  151. 151.J. You, R. Ying, X. Ren, W. Hamilton, and J. Leskovec. Graphrnn: Generating realistic graphs with deep auto-regressive models. In ICML, pages 5708–5717, 2018.
  152. 152.Y. You, T. Chen, Y. Sui, T. Chen, Z. Wang, and Y. Shen. Graph contrastive learning with augmentations. arXiv preprint arXiv:2010.13902, 2020.
  153. 153.Y. You, T. Chen, Z. Wang, and Y. Shen. When does self-supervision help graph convolutional networks? arXiv preprint arXiv:2006.09136, 2020.
  154. 154.F. Zhang, X. Liu, J. Tang, Y. Dong, P. Yao, J. Zhang, X. Gu, Y. Wang, B. Shao, R. Li, and K. Wang. Oag: Toward linking large-scale heterogeneous entity graphs. In KDD’19, pages 2585–2595, 2019.
  155. 155.J. Zhang, Y. Dong, Y. Wang, J. Tang, and M. Ding. Prone: fast and scalable network representation learning. In IJCAI, pages 4278–4284, 2019.
  156. 156.M. Zhang, Z. Cui, M. Neumann, and Y. Chen. An end-to-end deep learning architecture for graph classification. In AAAI, 2018.
  157. 157.R. Zhang, P. Isola, and A. A. Efros. Colorful image colorization. In ECCV, pages 649–666. Springer, 2016.
  158. 158.R. Zhang, P. Isola, and A. A. Efros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In CVPR, pages 1058–1067, 2017.
  159. 159.Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu. Ernie: Enhanced language representation with informative entities. arXiv preprint arXiv:1905.07129, 2019.
  160. 160.D. Zhu, P. Cui, D. Wang, and W. Zhu. Deep variational network embedding in wasserstein space. In SIGKDD, pages 2827–2836, 2018.
  161. 161.J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, and E. Shechtman. Toward multimodal image-to-image translation. In NIPS, pages 465–476, 2017.
  162. 162.C. Zhuang, A. L. Zhai, and D. Yamins. Local aggregation for unsupervised learning of visual embeddings. In Proceedings of the IEEE ICCV, pages 6002–6012, 2019.
  163. 163.B. Zoph, G. Ghiasi, T.-Y. Lin, Y. Cui, H. Liu, E. D. Cubuk, and Q. V. Le. Rethinking pre-training and self-training. arXiv preprint arXiv:2006.06882, 2020.
  164. 164.B. Zoph and Q. V. Le. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016.

Citation

MLA
Liu, X., et al. “Self-supervised Learning: Generative or Contrastive”. IEEE Transactions on Knowledge and Data Engineering, 2021, pp. 1–1, https://doi.org/10.1109/TKDE.2021.3090866.
APA
Liu, X., Zhang, F., Hou, Z., Mian, L., Wang, Z., Zhang, J., & Tang, J. (2021). Self-supervised Learning: Generative or Contrastive. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2021.3090866
Chicago
Liu, X., F. Zhang, Z. Hou, et al. 2021. “Self-supervised Learning: Generative or Contrastive”. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2021.3090866.
Harvard
Liu, X. et al. (2021) “Self-supervised Learning: Generative or Contrastive”, IEEE Transactions on Knowledge and Data Engineering, pp. 1–1. Available at: https://doi.org/10.1109/TKDE.2021.3090866.
Vancouver
1. Liu X, Zhang F, Hou Z, Mian L, Wang Z, Zhang J, Tang J (2021) Self-supervised Learning: Generative or Contrastive. IEEE Transactions on Knowledge and Data Engineering 1–1

BibTeX

@article{Liu_2021, title={Self-supervised Learning: Generative or Contrastive}, ISSN={2326-3865}, url={http://dx.doi.org/10.1109/TKDE.2021.3090866}, DOI={10.1109/tkde.2021.3090866}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Liu, Xiao and Zhang, Fanjin and Hou, Zhenyu and Mian, Li and Wang, Zhaoyu and Zhang, Jing and Tang, Jie}, year={2021}, pages={1–1} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/