Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives

Yanzhao ZhangRichong ZhangSamuel MensahXudong LiuYongyi Mao

article2022AAAI70 citations

Proposes MixCSE, a contrastive sentence representation framework that overcomes vanishing gradient signals by continually generating artificial hard negative features through mixing positive and negative samples, achieving state-of-the-art results on semantic textual similarity and transfer tasks.

Listen

Modern natural language processing systems rely heavily on unsupervised sentence representation to convert text into numerical vectors without requiring costly human labeling. These vectors power key enterprise applications such as semantic search, information retrieval, and text classification. The current leading method, contrastive learning, trains models by pulling similar sentence representations together while pushing dissimilar representations apart. However, standard techniques rely on random sampling to pick dissimilar examples. This approach quickly stalls model training because randomly chosen examples become too easy to distinguish, causing learning signals to fade.

The article demonstrates why difficult, hard-to-distinguish negative examples are mathematically essential for effective contrastive learning in sentence representation. It introduces a new framework, named MixCSE, which continuously synthesizes challenging negative examples by blending positive and negative sentence features during training.

The researchers proved mathematically and empirically that randomly sampled negative examples fail to maintain adequate training momentum, especially when starting with narrow representation spaces from standard language models such as BERT. To solve this, MixCSE creates artificial "mix negatives" by combining features from the target sentence with randomly selected sentences while applying a stop-gradient safeguard to keep updates stable. The team trained the model on one million unlabeled Wikipedia sentences and evaluated performance across fourteen benchmark datasets covering Semantic Textual Similarity and downstream transfer tasks.

The findings show that MixCSE establishes new state-of-the-art performance. On semantic similarity benchmarks, MixCSE improved average correlation scores from 74.83 to 77.66 on standard base models and reached 78.80 on large models, consistently outperforming the prior benchmark, SimCSE. MixCSE also yielded superior results on downstream transfer tasks, achieving an average classification accuracy of 87.77 on the large architecture while exhibiting significantly lower performance variance across runs. Mathematical and empirical analyses confirmed that synthetic mixed negatives preserve strong training signals throughout the entire learning cycle, producing more uniformly distributed representations.

These results demonstrate that organizations can significantly improve the quality and stability of their language processing pipelines without collecting expensive labeled datasets or increasing computational hardware demands. By generating harder synthetic training examples on the fly, models converge faster and achieve higher accuracy across diverse text-processing tasks.

Organizations developing or deploying text embedding models should adopt synthetic negative mixing strategies in place of pure random sampling within their contrastive learning workflows. When implementing this approach, teams must maintain the stop-gradient operator and tune the mixing parameter cautiously to avoid confusing hard negatives with true positive matches. Future initiatives should evaluate how this mixing technique transfers to broader multilingual settings, domain-specific corpora, and larger language model architectures.

Cover for Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives

Abstract

Unsupervised sentence representation learning is a fundamental problem in natural language processing. Recently, contrastive learning has made great success on this task. Existing contrastive learning based models usually apply random sampling to select negative examples for training. Previous work in computer vision has shown that hard negative examples help contrastive learning to achieve faster convergence and better optimization for representation learning. However, the importance of hard negatives in contrastive learning for sentence representation is yet to be explored. In this study, we prove that hard negatives are essential for maintaining strong gradient signals in the training process while random sampling negative examples is ineffective for sentence representation. Accordingly, we present a contrastive model, MixCSE, that extends the current state-of-the-art SimCSE by continually constructing hard negatives via mixing both positive and negative features. The superior performance of the proposed approach is demonstrated via empirical studies on Semantic Textual Similarity datasets and Transfer task datasets.

Table of Contents

  • Introduction
  • Related Work
  • Unsupervised Sentence Representation
  • Contrastive Learning for Sentence Representation
  • Model
  • Contrastive Learning Framework
  • Distribution of BERT Embeddings
  • MixCSE
  • Experiment
  • Evaluation Setup
  • Implementation Details
  • Experiment Results
  • Ablation Study
  • Analysis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — MixCSE Contrastive Learning Framework with Hard Negative Mixing

    model/method

    MixCSE is an unsupervised contrastive sentence representation framework that synthesizes artificial hard negative features by interpolating positive representations with randomly sampled negative representations.

    Let D\mathcal{D} be a corpus of sentences. For an input sentence xi∈Dx_i \in \mathcal{D}, an encoder (such as BERT) and a data augmentation module (such as two distinct dropout masks) produce two ℓ2\ell_2-normalized feature vectors hi,hi′∈Sd−1⊂Rdh_i, h'_i \in \mathbb{S}^{d-1} \subset \mathbb{R}^d, forming a positive pair (hi,hi′)(h_i, h'_i) where hih_i acts as the anchor feature and hi′h'_i acts as the positive feature. For each randomly sampled negative sentence representation hj′h'_j (j∈{1,…,N}j \in \{1, \dots, N\}), an artificial mixed negative feature h~i,j′\tilde{h}'_{i,j} is generated via feature interpolation:

    h~i,j′=λhi′+(1−λ)hj′∥λhi′+(1−λ)hj′∥2\tilde{h}'_{i,j} = \frac{\lambda h'_i + (1 - \lambda)h'_j}{\|\lambda h'_i + (1 - \lambda)h'_j\|_2}

    where λ∈(0,1)\lambda \in (0, 1) is a scalar hyperparameter controlling the mixing ratio.

    The MixCSE contrastive loss for anchor hih_i incorporates these mixed negative features with a stop-gradient operator SG(⋅)\text{SG}(\cdot):

    Lmix=−log⁡exp⁡(hiThi′/τ)C+∑j=1Nexp⁡(hiTSG(h~i,j′)/τ)L_{\text{mix}} = -\log \frac{\exp(h_i^T h'_i / \tau)}{C + \sum_{j=1}^N \exp(h_i^T \text{SG}(\tilde{h}'_{i,j}) / \tau)}

    where τ>0\tau > 0 is a temperature hyperparameter and C=exp⁡(hiThi′/τ)+∑j=1Nexp⁡(hiThj′/τ)C = \exp(h_i^T h'_i / \tau) + \sum_{j=1}^N \exp(h_i^T h'_j / \tau).

    Symmetrically, when hi′h'_i serves as the anchor, mixed negative features h~i,j=λhi+(1−λ)hj∥λhi+(1−λ)hj∥2\tilde{h}_{i,j} = \frac{\lambda h_i + (1 - \lambda)h_j}{\|\lambda h_i + (1 - \lambda)h_j\|_2} are constructed from hih_i and hjh_j and included in the loss.

  2. Knowl 2 — Gradient Dynamics of Contrastive Loss and Non-Vanishing Gradients via Mixed Negatives

    theoretical result

    In standard InfoNCE contrastive learning on the unit sphere Sd−1\mathbb{S}^{d-1} with anchor hih_i, positive view hi′h'_i, and NN randomly sampled negative views {hj′}j=1N\{h'_j\}_{j=1}^N, the loss is:

    Lcl=−log⁡exp⁡(hiThi′/τ)exp⁡(hiThi′/τ)+∑j=1Nexp⁡(hiThj′/τ)L_{\text{cl}} = -\log \frac{\exp(h_i^T h'_i / \tau)}{\exp(h_i^T h'_i / \tau) + \sum_{j=1}^N \exp(h_i^T h'_j / \tau)}

    The derivative with respect to the anchor hih_i is:

    ∂Lcl∂hi=−1Cτ∑j=1Nexp⁡(hiThj′τ)(hi′−hj′)\frac{\partial L_{\text{cl}}}{\partial h_i} = -\frac{1}{C\tau} \sum_{j=1}^N \exp\left(\frac{h_i^T h'_j}{\tau}\right)(h'_i - h'_j)

    where C=exp⁡(hiThi′/τ)+∑j=1Nexp⁡(hiThj′/τ)C = \exp(h_i^T h'_i / \tau) + \sum_{j=1}^N \exp(h_i^T h'_j / \tau) and τ>0\tau > 0 is the temperature hyperparameter.

    The gradient updates hih_i in the direction of hi′−hj′=(hi′−hi)−(hj′−hi)h'_i - h'_j = (h'_i - h_i) - (h'_j - h_i), pulling hih_i toward the positive view hi′h'_i and pushing it away from each negative hj′h'_j. The magnitude of the gradient contribution for negative hj′h'_j scales exponentially with the inner product hiThj′h_i^T h'_j.

    As contrastive training pushes random negatives away, their cosine similarities approach zero (hiThj′≈0h_i^T h'_j \approx 0), whereas positive alignment increases (hiThi′≈1h_i^T h'_i \approx 1). Consequently, ∑j=1Nexp⁡(hiThj′/τ)\sum_{j=1}^N \exp(h_i^T h'_j / \tau) becomes negligible compared to exp⁡(hiThi′/τ)\exp(h_i^T h'_i / \tau), causing the gradient signal ∂Lcl/∂hi\partial L_{\text{cl}} / \partial h_i to diminish toward zero and stalling optimization.

    When synthetic mixed negatives h~i,j′=λhi′+(1−λ)hj′∥λhi′+(1−λ)hj′∥2\tilde{h}'_{i,j} = \frac{\lambda h'_i + (1-\lambda)h'_j}{\|\lambda h'_i + (1-\lambda)h'_j\|_2} are introduced, assuming positive alignment (hiThi′≈1h_i^T h'_i \approx 1) and orthogonal random negatives (hiThj′≈0h_i^T h'_j \approx 0), the inner product between anchor and mixed negative satisfies:

    hiTh~i,j′≈λλ2+(1−λ)2>0h_i^T \tilde{h}'_{i,j} \approx \frac{\lambda}{\sqrt{\lambda^2 + (1 - \lambda)^2}} > 0

    This strictly positive inner product ensures that mixed negatives continuously generate a non-vanishing gradient signal throughout the training process.

  3. Knowl 3 — Geometric Analysis of Anisotropic Pre-Trained Embeddings and Negative Cancellation

    theoretical result

    Pre-trained language model representations (such as those from BERT) exhibit anisotropy, confining ℓ2\ell_2-normalized sentence embeddings to a narrow spherical cap Oω={z∈Sd−1:∠(z,μ)≤ω}\mathcal{O}_\omega = \{z \in \mathbb{S}^{d-1} : \angle(z, \mu) \le \omega\} centered around a north pole μ=(1,0,…,0)∈Rd\mu = (1, 0, \dots, 0) \in \mathbb{R}^d with maximum angular radius ω∈(0,π)\omega \in (0, \pi).

    For two points h,h′∈Oωh, h' \in \mathcal{O}_\omega with ∠(h,μ)=ϕ1\angle(h, \mu) = \phi_1, ∠(h′,μ)=ϕ1′\angle(h', \mu) = \phi'_1, projection angle β=∠(proj(h),proj(h′))\beta = \angle(\text{proj}(h), \text{proj}(h')) on the equatorial plane, and pairwise angle θ=∠(h,h′)\theta = \angle(h, h'):

    cos⁡θ=cos⁡ϕ1cos⁡ϕ1′+sin⁡ϕ1sin⁡ϕ1′cos⁡β\cos \theta = \cos \phi_1 \cos \phi'_1 + \sin \phi_1 \sin \phi'_1 \cos \beta

    When h′h' is distributed uniformly over Oω\mathcal{O}_\omega, the variance of cos⁡θ\cos \theta monotonically decreases toward 0 as the embedding dimension dd increases. When ω≥π/2\omega \ge \pi/2 (as the representation space expands during contrastive learning), the mean of cos⁡θ\cos \theta approaches 0 regardless of ϕ1\phi_1, indicating that virtually all randomly sampled negatives become orthogonal to the anchor.

    Increasing the number of sampled negatives NN fails to produce effective gradient signals due to a cancellation effect: in the large NN limit, for every negative h′h' at angle ∠(h,h′)∈[0,ω−ϕ]\angle(h, h') \in [0, \omega - \phi], there exists a symmetrically located negative h′′h'' such that (h′+h′′)/2=ρh(h' + h'')/2 = \rho h for some ρ<1\rho < 1. The directional gradient contribution of h′′h'' cancels that of h′h', while increasing NN inflates the normalizing denominator C=exp⁡(hiThi′/τ)+∑j=1Nexp⁡(hiThj′/τ)C = \exp(h_i^T h'_i / \tau) + \sum_{j=1}^N \exp(h_i^T h'_j / \tau), further weakening the overall gradient signal.

  4. Knowl 4 — Stop-Gradient Necessity on Synthetic Mixed Negatives

    theoretical result

    In MixCSE, applying a stop-gradient operator SG(⋅)\text{SG}(\cdot) to synthetic mixed negatives h~i,j′=λhi′+(1−λ)hj′∥λhi′+(1−λ)hj′∥2\tilde{h}'_{i,j} = \frac{\lambda h'_i + (1 - \lambda)h'_j}{\|\lambda h'_i + (1 - \lambda)h'_j\|_2} is mathematically necessary to avoid degrading positive pair alignment.

    Let Lmixno-sgL_{\text{mix}}^{\text{no-sg}} denote the contrastive loss formulated without the stop-gradient operator. Its partial derivative with respect to the positive representation hi′h'_i is:

    ∂Lmixno-sg∂hi′=−1C′τ((∑j=1Nexp⁡(hiThj′τ)+exp⁡(hiTh~i,j′τ))hiT−∂h~i,j′∂hi′exp⁡(hiTh~i,j′τ)hiT)\frac{\partial L_{\text{mix}}^{\text{no-sg}}}{\partial h'_i} = -\frac{1}{C'\tau} \left( \left( \sum_{j=1}^N \exp\left(\frac{h_i^T h'_j}{\tau}\right) + \exp\left(\frac{h_i^T \tilde{h}'_{i,j}}{\tau}\right) \right) h_i^T - \frac{\partial \tilde{h}'_{i,j}}{\partial h'_i} \exp\left(\frac{h_i^T \tilde{h}'_{i,j}}{\tau}\right) h_i^T \right)

    where C′=exp⁡(hiThi′/τ)+∑j=1Nexp⁡(hiThj′/τ)+∑j=1Nexp⁡(hiTh~i,j′/τ)C' = \exp(h_i^T h'_i / \tau) + \sum_{j=1}^N \exp(h_i^T h'_j / \tau) + \sum_{j=1}^N \exp(h_i^T \tilde{h}'_{i,j} / \tau).

    The term −∂h~i,j′∂hi′exp⁡(hiTh~i,j′/τ)hiT-\frac{\partial \tilde{h}'_{i,j}}{\partial h'_i} \exp(h_i^T \tilde{h}'_{i,j} / \tau) h_i^T yields an update direction that pushes the positive representation hi′h'_i away from the anchor hih_i. By applying SG(h~i,j′)\text{SG}(\tilde{h}'_{i,j}), back-propagation through h~i,j′\tilde{h}'_{i,j} into hi′h'_i is blocked, ensuring the encoder exclusively pulls positive representations closer to the anchor.

  5. Knowl 5 — Theoretical Upper Bound and Selection of Mixing Parameter Lambda

    theoretical result

    To ensure that synthetic mixed features act as hard negatives rather than unintended pseudo-positives, the mixed negative h~i,j′=λhi′+(1−λ)hj′∥λhi′+(1−λ)hj′∥2\tilde{h}'_{i,j} = \frac{\lambda h'_i + (1 - \lambda)h'_j}{\|\lambda h'_i + (1 - \lambda)h'_j\|_2} must remain strictly farther from the anchor hih_i than the positive representation hi′h'_i.

    Let ∠(hi,hi′)=γ\angle(h_i, h'_i) = \gamma be the angle between the anchor and its positive representation. Enforcing ∠(hi,h~i,j′)≥γ\angle(h_i, \tilde{h}'_{i,j}) \ge \gamma leads to the condition:

    arccos⁡(λ+(1−λ)hiThj′∥λhi′+(1−λ)hj′∥2)≥γ\arccos\left( \frac{\lambda + (1 - \lambda)h_i^T h'_j}{\|\lambda h'_i + (1 - \lambda)h'_j\|_2} \right) \ge \gamma

    which imposes the theoretical upper bound on the mixing parameter λ\lambda:

    λ<∥λhi′+(1−λ)hj′∥2cos⁡(γ)−cos⁡(hiThj′)1−cos⁡(hiThj′)\lambda < \frac{\|\lambda h'_i + (1 - \lambda)h'_j\|_2 \cos(\gamma) - \cos(h_i^T h'_j)}{1 - \cos(h_i^T h'_j)}

    Because the true semantic boundary γ\gamma is unobserved in unsupervised training, setting λ\lambda to a small constant value (such as λ=0.2\lambda = 0.2) prevents generating false negative features that lie within the positive semantic neighborhood.

  6. Knowl 6 — Performance on Semantic Textual Similarity (STS) Benchmarks

    data/table

    Evaluation of unsupervised sentence embedding models across seven Semantic Textual Similarity (STS) test datasets using the SentEval toolkit. Results report the Spearman's rank correlation coefficient (×100\times 100) aggregated over 5 runs (mean ±\pm standard deviation where available).

    Model STS12 STS13 STS14 STS15 STS16 STS-B SICK-R Avg
    Avg.GloVe 55.14 70.66 59.73 68.25 63.66 58.02 53.76 61.32
    BERTbase_{\text{base}} 35.20 59.53 49.37 63.39 62.73 48.18 58.60 53.86
    BERTbase_{\text{base}}-flow 58.40 67.10 60.85 75.16 71.22 68.66 64.47 66.55
    BERTbase_{\text{base}}-whitening 57.83 66.90 60.90 75.08 71.31 68.24 63.73 66.28
    IS-BERTbase_{\text{base}} 56.77 69.24 61.21 75.23 70.16 69.21 64.25 66.58
    ConSBERTbase_{\text{base}} 64.64 78.49 69.07 79.72 75.95 73.97 67.31 72.74
    SimCSE-BERTbase_{\text{base}} 67.17±\pm9.61 79.79±\pm2.72 71.96±\pm3.97 80.21±\pm1.42 77.65±\pm1.24 76.46±\pm1.44 70.57±\pm1.25 74.83±\pm2.32
    MixCSE-BERTbase_{\text{base}} 71.71±\pm4.04 83.14±\pm0.72 75.49±\pm1.25 83.64±\pm2.32 79.00±\pm0.16 78.48±\pm0.82 72.19±\pm0.46 77.66±\pm0.61
    BERTlarge_{\text{large}} 33.06 57.64 47.95 55.83 62.42 49.66 53.87 51.49
    BERTlarge_{\text{large}}-flow 65.20 73.39 69.42 74.92 77.63 72.26 62.50 70.76
    BERTlarge_{\text{large}}-whitening 64.35 74.60 69.64 74.68 75.94 60.81 72.47 70.35
    ConSBERTlarge_{\text{large}} 70.69 82.96 74.13 82.78 76.66 77.53 70.37 76.45
    SimCSE-BERTlarge_{\text{large}} 70.21±\pm1.49 83.97±\pm1.18 75.92±\pm0.56 83.9±\pm0.49 78.87±\pm0.75 79.0±\pm1.0 73.89±\pm1.08 77.97±\pm0.7
    MixCSE-BERTlarge_{\text{large}} 72.55±\pm0.49 84.32±\pm0.53 76.69±\pm0.76 84.31±\pm0.10 79.67±\pm0.28 79.90±\pm0.18 74.07±\pm0.13 78.80±\pm0.09

    MixCSE achieves state-of-the-art performance, outperforming SimCSE by +2.83 points on BERTbase_{\text{base}} and +0.83 points on BERTlarge_{\text{large}} on average, while exhibiting lower variance across runs.

  7. Knowl 7 — Performance on Downstream Transfer Classification Tasks

    data/table

    Evaluation of unsupervised sentence representations on seven transfer classification datasets from the SentEval benchmark: Movie Reviews (MR), Customer Reviews (CR), Subjectivity (SUBJ), MPQA opinion polarity, Stanford Sentiment Treebank (SST-2), TREC question classification, and Microsoft Research Paraphrase Corpus (MRPC). For each method, a linear classifier is trained on fixed sentence embeddings, and accuracy is reported over 5 runs (mean ±\pm standard deviation).

    Model MR CR SUBJ MPQA SST TREC MRPC Avg
    GloVe.Avg 77.25 78.30 91.17 87.85 80.18 83.00 72.87 81.52
    BERTbase_{\text{base}} 78.66 86.25 94.37 88.66 84.40 92.80 69.54 84.94
    IS-BERTbase_{\text{base}} 81.09 87.18 94.96 88.75 85.96 88.64 74.24 85.83
    SimCSE-BERTbase_{\text{base}} 71.12±\pm5.94 85.92±\pm0.41 98.56±\pm2.05 88.61±\pm0.18 85.34±\pm0.37 88.4±\pm0.59 73.48±\pm1.26 84.49±\pm0.59
    MixCSE-BERTbase_{\text{base}} 81.3±\pm1.75 86.77±\pm0.4 99.64±\pm0.01 89.71±\pm0.17 85.87±\pm0.49 84.91±\pm0.24 76.08±\pm0.68 86.33±\pm0.26
    BERTlarge_{\text{large}} 60.89 90.15 99.62 86.04 89.95 93.00 69.86 84.22
    SimCSE-BERTlarge_{\text{large}} 73.93±\pm2.79 88.87±\pm0.75 99.6±\pm0 89.49±\pm0.16 90.59±\pm0.96 91.72±\pm0.95 75.49±\pm0.88 86.8±\pm0.42
    MixCSE-BERTlarge_{\text{large}} 82.95±\pm0.52 89.57±\pm0.05 99.67±\pm0.01 90.14±\pm0.02 89.17±\pm0.84 86.13±\pm0.58 76.74±\pm0.16 87.77±\pm0.11

    MixCSE-BERTbase_{\text{base}} achieves an average transfer accuracy of 86.33% (+1.84% over SimCSE-BERTbase_{\text{base}}), and MixCSE-BERTlarge_{\text{large}} achieves 87.77% (+0.97% over SimCSE-BERTlarge_{\text{large}}), demonstrating that mixing negatives produces representations that transfer effectively to downstream semantic and classification tasks.

  8. Knowl 8 — Ablation Analysis on Stop-Gradient and Bidirectional Negative Mixing

    data/table

    Ablation study evaluating the contributions of parallel bidirectional mixed negative generation and the stop-gradient operator across Semantic Textual Similarity (STS Avg) and Transfer classification tasks (TR Avg).

    • MixCSEsingle\text{MixCSE}_{\text{single}} constructs mixed negative representations only for the anchor hih_i rather than for both views hih_i and hi′h'_i.
    • MixCSEwo sg\text{MixCSE}_{\text{wo sg}} removes the stop-gradient operation SG(⋅)\text{SG}(\cdot), allowing gradients to back-propagate through the synthetic mixed negatives h~i,j′\tilde{h}'_{i,j}.
    Model STS(Avg) TR(Avg)
    MixCSE-BERTbase_{\text{base}} 77.66±\pm0.61 86.33±\pm0.26
    MixCSEsingle_{\text{single}}-BERTbase_{\text{base}} 76.43±\pm0.11 85.86±\pm0.04
    MixCSEwo sg_{\text{wo sg}}-BERTbase_{\text{base}} 75.74±\pm0.87 84.55±\pm0.27
    MixCSE-BERTlarge_{\text{large}} 78.80±\pm0.09 87.77±\pm0.11
    MixCSEsingle_{\text{single}}-BERTlarge_{\text{large}} 78.11±\pm0.20 87.21±\pm0.23
    MixCSEwo sg_{\text{wo sg}}-BERTlarge_{\text{large}} 77.61±\pm0.45 86.02±\pm0.27

    Removing bidirectional mixing (MixCSEsingle\text{MixCSE}_{\text{single}}) causes a 1.23-point STS drop on BERTbase_{\text{base}} and 0.69-point drop on BERTlarge_{\text{large}}. Removing the stop-gradient operator (MixCSEwo sg\text{MixCSE}_{\text{wo sg}}) causes a larger drop of 1.92 points on STS and 1.78 points on TR for BERTbase_{\text{base}}, confirming that both components are essential.

  9. Knowl 9 — Experimental Configuration for Unsupervised MixCSE Sentence Representation

    experimental setup

    MixCSE is trained in an unsupervised setting using the following setup:

    • Training corpus: 1,000,000 unlabeled sentences randomly sampled from English Wikipedia.
    • Encoder backbones: Pre-trained BERTbase\text{BERT}_{\text{base}} and BERTlarge\text{BERT}_{\text{large}} models, with sentence embeddings extracted from the [CLS] token representation.
    • Data augmentation: Two independent standard dropout masks applied in two forward passes of the same input sentence through the encoder to obtain positive pair (hi,hi′)(h_i, h'_i).
    • Hyperparameters: Temperature τ=0.05\tau = 0.05, mixing ratio λ=0.2\lambda = 0.2, batch size 64, trained for 1 epoch with early stopping.
    • Optimization: Adam optimizer with a learning rate of 3×10−53 \times 10^{-5} for BERTbase\text{BERT}_{\text{base}} and 1×10−51 \times 10^{-5} for BERTlarge\text{BERT}_{\text{large}}.
    • Implementation and Hardware: Implemented in Python 3.6 with PyTorch 1.6.0, executed on a single 32 GB NVIDIA A100 GPU.
  10. Knowl 10 — Empirical Dynamics of Feature Similarity, Embedding Distribution Expansion, and Alignment-Uniformity Trade-Off

    empirical result

    Analysis of representation distributions during training on the STS-B development set demonstrates the following properties of MixCSE:

    1. Similarity score dynamics: The positive cosine score (hiThi′/τh_i^T h'_i / \tau) remains high throughout training, whereas the random negative score (hiThj′/τh_i^T h'_j / \tau) rapidly decays toward zero (corresponding to an angle of ≈π/2\approx \pi/2). The mixed negative score (hiTh~i,j′/τh_i^T \tilde{h}'_{i,j} / \tau with τ=0.05\tau = 0.05) decreases slightly but stays significantly above the random negative score, confirming that mixed negatives maintain a steady hard negative signal.
    2. Spherical cap angle expansion: Tracking the maximum angle θ\theta between any two sentence embeddings in STS-B indicates that MixCSE expands the angular radius ω\omega of the embedding distribution faster than SimCSE and converges to a higher maximum angle, mitigating BERT's representational anisotropy more effectively.
    3. Alignment and uniformity metrics: Evaluating alignment Lalign=E(x,y)∼ppos[∥f(x)−f(y)∥22]\mathcal{L}_{\text{align}} = \mathbb{E}_{(x, y) \sim p_{\text{pos}}}[\|f(x) - f(y)\|_2^2] and uniformity Luniform=log⁡E(x,y)∼pdata[e−2∥f(x)−f(y)∥22]\mathcal{L}_{\text{uniform}} = \log \mathbb{E}_{(x, y) \sim p_{\text{data}}}[e^{-2\|f(x) - f(y)\|_2^2}] shows that MixCSE achieves substantially better (lower) uniformity than SimCSE and ConSBERT while preserving comparable positive alignment.

Coverage note — None was omitted. All primary theoretical derivations, methodology components, experimental protocols, benchmark results on STS and Transfer tasks, ablations, and analytical findings from the paper are fully covered.

References

  1. 1.Agirre, E.; Banea, C.; Cardie, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Guo, W.; Lopez-Gazpio, I.; Maritxalar, M.; Mihalcea, R.; Rigau, G.; Uria, L.; and Wiebe, J. 2015. SemEval-2015 Task 2: Semantic Textual Similarity, English, Spanish and Pilot on Interpretability. In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015), 252–263. Denver, Colorado: Association for Computational Linguistics.
  2. 2.Agirre, E.; Banea, C.; Cardie, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Guo, W.; Mihalcea, R.; Rigau, G.; and Wiebe, J. 2014. SemEval-2014 Task 10: Multilingual Semantic Textual Similarity. In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), 81–91. Dublin, Ireland: Association for Computational Linguistics.
  3. 3.Agirre, E.; Banea, C.; Cer, D.; Diab, M.; Gonzalez-Agirre, A.; Mihalcea, R.; Rigau, G.; and Wiebe, J. 2016. SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-Lingual Evaluation. In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), 497–511. San Diego, California: Association for Computational Linguistics.
  4. 4.Agirre, E.; Cer, D.; Diab, M.; and Gonzalez-Agirre, A. 2012. SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity. In *SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), 385–393. Montreal, Canada: Association for Computational Linguistics.
  5. 5.Arora, S.; Liang, Y.; and Ma, T. 2017. A Simple but Tough-to-Beat Baseline for Sentence Embeddings. In ICLR.
  6. 6.Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), 1–14. Vancouver, Canada: Association for Computational Linguistics.
  7. 7.Cer, D.; Yang, Y.; Kong, S.-y.; Hua, N.; Limtiaco, N.; John, R. S.; Constant, N.; Guajardo-Cespedes, M.; Yuan, S.; Tar, C.; et al. 2018. Universal sentence encoder. arXiv preprint arXiv:1803.11175.
  8. 8.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, 1597–1607. PMLR.
  9. 9.Conneau, A.; and Kiela, D. 2018. SentEval: An Evaluation Toolkit for Universal Sentence Representations. arXiv preprint arXiv:1803.05449.
  10. 10.Conneau, A.; Kiela, D.; Schwenk, H.; Barrault, L.; and Bordes, A. 2017. Supervised learning of universal sentence representations from natural language inference data. arXiv preprint arXiv:1705.02364.
  11. 11.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  12. 12.Dinh, L.; Krueger, D.; and Bengio, Y. 2014. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516.
  13. 13.Dolan, W. B.; and Brockett, C. 2005. Automatically Constructing a Corpus of Sentential Paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
  14. 14.Ethayarajh, K. 2018. Unsupervised Random Walk Sentence Embeddings: A Strong but Simple Baseline. In Proceedings of The Third Workshop on Representation Learning for NLP, 91–100. Melbourne, Australia: Association for Computational Linguistics.
  15. 15.Ethayarajh, K. 2019. How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings. arXiv preprint arXiv:1909.00512.
  16. 16.Gao, J.; He, D.; Tan, X.; Qin, T.; Wang, L.; and Liu, T.-Y. 2019. Representation degeneration problem in training natural language generation models. arXiv preprint arXiv:1907.12009.
  17. 17.Gao, T.; Yao, X.; and Chen, D. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. arXiv preprint arXiv:2104.08821.
  18. 18.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9729–9738.
  19. 19.Kalantidis, Y.; Sariyildiz, M. B.; Pion, N.; Weinzaepfel, P.; and Larlus, D. 2020. Hard negative mixing for contrastive learning. arXiv preprint arXiv:2010.01028.
  20. 20.Kifer, D.; Ben-David, S.; and Gehrke, J. 2004. Detecting change in data streams. In VLDB, volume 4, 180–191. Toronto, Canada.
  21. 21.Kim, T.; Yoo, K. M.; and Lee, S.-g. 2021. Self-Guided Contrastive Learning for BERT Sentence Representations. arXiv preprint arXiv:2106.07345.
  22. 22.Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  23. 23.Kiros, R.; Zhu, Y.; Salakhutdinov, R. R.; Zemel, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015. Skip-thought vectors. In Advances in neural information processing systems, 3294–3302.
  24. 24.Lee, D.-H.; et al. 2013. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3.
  25. 25.Li, B.; Zhou, H.; He, J.; Wang, M.; Yang, Y.; and Li, L. 2020. On the sentence embeddings from pre-trained language models. arXiv preprint arXiv:2011.05864.
  26. 26.Marelli, M.; Menini, S.; Baroni, M.; Bentivogli, L.; Bernardi, R.; and Zamparelli, R. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), 216–223. Reykjavik, Iceland: European Language Resources Association (ELRA).
  27. 27.Pang, B.; and Lee, L. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058.
  28. 28.Pang, B.; and Lee, L. 2005. Seeing Stars: Exploiting Class Relationships for Sentiment Categorization with Respect to Rating Scales. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), 115–124. Ann Arbor, Michigan: Association for Computational Linguistics.
  29. 29.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32: 8026–8037.
  30. 30.Qiao, Y.; Xiong, C.; Liu, Z.; and Liu, Z. 2019. Understanding the Behaviors of BERT in Ranking. arXiv preprint arXiv:1904.07531.
  31. 31.Robinson, J.; Chuang, C.-Y.; Sra, S.; and Jegelka, S. 2020. Contrastive learning with hard negative samples. arXiv preprint arXiv:2010.04592.
  32. 32.Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, 1631–1642.
  33. 33.Su, J.; Cao, J.; Liu, W.; and Ou, Y. 2021. Whitening sentence representations for better semantics and faster retrieval. arXiv preprint arXiv:2103.15316.
  34. 34.Voorhees, E. M.; and Tice, D. M. 2000. Building a question answering test collection. In Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, 200–207.
  35. 35.Wang, F.; and Liu, H. 2021. Understanding the behaviour of contrastive loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2495–2504.
  36. 36.Wang, T.; and Isola, P. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning, 9929–9939. PMLR.
  37. 37.Wiebe, J.; Wilson, T.; and Cardie, C. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39(2): 165–210.
  38. 38.Xuan, H.; Stylianou, A.; Liu, X.; and Pless, R. 2020. Hard negative examples are hard, but useful. In European Conference on Computer Vision, 126–142. Springer.
  39. 39.Yan, Y.; Li, R.; Wang, S.; Zhang, F.; Wu, W.; and Xu, W. 2021. ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer. arXiv preprint arXiv:2105.11741.
  40. 40.Zhang, Y.; He, R.; Liu, Z.; Lim, K. H.; and Bing, L. 2020. An Unsupervised Sentence Embedding Method by Mutual Information Maximization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1601–1610.

Citation

MLA
Zhang, Y., et al. “Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 10, 2022, pp. 11730–38, https://doi.org/10.1609/AAAI.V36I10.21428.
APA
Zhang, Y., Zhang, R., Mensah, S., Liu, X., & Mao, Y. (2022). Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10), 11730–11738. https://doi.org/10.1609/AAAI.V36I10.21428
Chicago
Zhang, Y., R. Zhang, S. Mensah, X. Liu, and Y. Mao. 2022. “Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (10): 11730–38. https://doi.org/10.1609/AAAI.V36I10.21428.
Harvard
Zhang, Y. et al. (2022) “Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(10), pp. 11730–11738. Available at: https://doi.org/10.1609/AAAI.V36I10.21428.
Vancouver
1. Zhang Y, Zhang R, Mensah S, Liu X, Mao Y (2022) Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives. Proceedings of the AAAI Conference on Artificial Intelligence 36:11730–11738

BibTeX

@article{Zhang_2022, title={Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I10.21428}, DOI={10.1609/aaai.v36i10.21428}, number={10}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhang, Yanzhao and Zhang, Richong and Mensah, Samuel and Liu, Xudong and Mao, Yongyi}, year={2022}, month=June, pages={11730–11738} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF