Debiased Contrastive Learning of Unsupervised Sentence Representations

Kun ZhouBeichen ZhangWayne Xin ZhaoJi-Rong Wen

article2022ACL126 citations

Proposes a debiased contrastive learning framework that improves unsupervised sentence embeddings by downweighting false negatives and generating optimized noise-based negative samples to overcome representation anisotropy.

Listen

Natural language processing systems rely on high-quality numerical representations of sentences to power critical applications, including search engines, document retrieval, and text matching. While modern pre-trained language models serve as strong foundational tools, their raw sentence representations suffer from severe geometric collapse, crowding into a narrow vector space rather than distributing evenly. Contrastive learning has emerged as the standard solution by pulling similar sentences together and pushing dissimilar ones apart. However, current methods randomly select negative examples from training batches, which introduces severe sampling bias: models frequently push away sentences that are actually semantically similar (false negatives), and the sampled negatives remain confined to the original narrow representation space, degrading model performance.

The article introduces and evaluates DCLR (Debiased Contrastive Learning of unsupervised sentence Representations), a framework designed to eliminate negative sampling bias. Its objective is to demonstrate that systematically penalizing false negatives and generating synthetic negatives across the entire vector space significantly enhances the quality and uniformity of sentence representations.

To accomplish this, the authors designed a two-part approach. First, the framework introduces synthetic noise vectors initialized from a Gaussian distribution and optimizes them using gradient steps to target non-uniform regions across the entire semantic space. Second, it employs an instance-weighting mechanism powered by a complementary model to score candidate negatives, completely zeroing out the influence of false negatives that exceed a similarity threshold. The researchers evaluated this approach by fine-tuning standard language model backbones (BERT and RoBERTa) on one million unlabeled Wikipedia sentences and benchmarking performance across seven standard Semantic Textual Similarity tasks.

The evaluation produced four key findings in order of significance. First, DCLR established new top performance benchmarks across all seven semantic similarity tasks, improving average correlation scores over the previous best-in-class baseline (SimCSE) across base and large variants of BERT and RoBERTa. Second, the framework significantly enhanced vector uniformity throughout training compared to baseline models, preventing representation collapse. Third, ablation testing revealed that both synthetic negative generation and instance weighting are critical, with the removal of instance weighting causing the steepest performance drop. Fourth, the model demonstrated exceptional data efficiency: when trained on an extreme subsample of only 0.3% of the training data, performance declined by only 4% to 9%, confirming remarkable stability in resource-constrained settings.

These findings indicate that addressing negative sampling bias delivers immediate accuracy and robustness gains for enterprise language models without requiring costly manual data labeling. The ability of the framework to maintain high accuracy under severe data scarcity reduces compute costs and training data acquisition bottlenecks, making high-performance text matching viable in low-resource environments.

Organizations deploying semantic search, retrieval, or text-matching systems should consider adopting debiased contrastive sampling strategies to refine their sentence encoders. Implementation teams should tune the weighting threshold and negative proportion parameters according to domain needs, as semantic sensitivity varies by application. Future engineering efforts should explore extending this debiasing methodology to multilingual models, multimodal architectures, and initial model pre-training pipelines.

Confidence in these findings is high given the consistent gains across seven standard benchmarks and multiple model architectures. However, practitioners should note limitations: the weighting mechanism relies on a reliable complementary model, and the underlying pre-trained models can still inherit social and contextual biases from their initial training corpora, warranting routine validation in production deployments.

arXiv: 2205.00656
Cover for Debiased Contrastive Learning of Unsupervised Sentence Representations

Abstract

Recently, contrastive learning has been shown to be effective in improving pre-trained language models (PLM) to derive high-quality sentence representations. It aims to pull close positive examples to enhance the alignment while push apart irrelevant negatives for the uniformity of the whole representation space. However, previous works mostly adopt in-batch negatives or sample from training data at random. Such a way may cause the sampling bias that improper negatives (e.g., false negatives and anisotropy representations) are used to learn sentence representations, which will hurt the uniformity of the representation space. To address it, we present a new framework DCLR (Debiased Contrastive Learning of unsupervised sentence Representations) to alleviate the influence of these improper negatives. In DCLR, we design an instance weighting method to punish false negatives and generate noise-based negatives to guarantee the uniformity of the representation space. Experiments on seven semantic textual similarity tasks show that our approach is more effective than competitive baselines. Our code and data are publicly available at the link: https://github.com/RUCAIBox/DCLR.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Approach
  • 4.1 Generating Noise-based Negatives
  • 4.2 Contrastive Learning with Instance Weighting
  • 4.3 Overview and Discussion
  • 4.3.1 Overview of DCLR
  • 4.3.2 Discussion
  • 5 Experiment - Main Results
  • 5.1 Experiment Setup
  • 5.2 Main Results
  • 6 Experiment - Analysis and Extension
  • 6.1 Debiased Contrastive Learning on Other Methods
  • 6.2 Ablation Study
  • 6.3 Uniformity Analysis
  • 6.4 Performance under Few-shot Settings
  • 6.5 Hyper-parameters Analysis
  • 7 Conclusion
  • Ethical Consideration
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — DCLR debiased contrastive learning framework

    model/method

    DCLR is an unsupervised sentence-representation framework designed to reduce two sources of bias in randomly sampled contrastive negatives: false negatives that are semantically close to the input, and anisotropic negatives concentrated in the narrow region occupied by pretrained language-model representations. For an unlabeled sentence xix_i, a BERT- or RoBERTa-based encoder produces a sentence vector hi∈Rdh_i\in\mathbb{R}^d from the final-layer [CLS] token. DCLR creates a positive vector hi+h_i^+ through dropout augmentation, combines ordinary in-batch negative representations with optimized Gaussian noise vectors, assigns each negative a weight using a complementary SimCSE model, and trains the encoder with a weighted contrastive objective. The two mechanisms are intended respectively to suppress semantically misleading negatives and to expose the encoder to points outside the pretrained model's anisotropic representation cone.

  2. Knowl 2 — Gradient-optimized noise-based negatives

    equation

    For each input sentence representation hi∈Rdh_i\in\mathbb{R}^d, DCLR initializes kk synthetic negative vectors independently from a zero-mean isotropic Gaussian distribution with standard deviation σ\sigma:

    {h^1,h^2,…,h^k}∼N(0,σ2Id),\{\hat h_1,\hat h_2,\ldots,\hat h_k\}\sim\mathcal{N}(0,\sigma^2 I_d),

    where IdI_d is the d×dd\times d identity matrix. Let hi+h_i^+ be a dropout-augmented positive representation, let sim⁡(a,b)=a⊤b/(∥a∥2∥b∥2)\operatorname{sim}(a,b)=a^\top b/(\lVert a\rVert_2\lVert b\rVert_2) be cosine similarity, and let τu>0\tau_u>0 be the temperature used for noise optimization. DCLR defines the non-uniformity objective

    LU(hi,hi+,{h^j}j=1k)=−log⁡exp⁡(sim⁡(hi,hi+)/τu)∑j=1kexp⁡(sim⁡(hi,h^j)/τu).L_U(h_i,h_i^+,\{\hat h_j\}_{j=1}^k)=-\log\frac{\exp(\operatorname{sim}(h_i,h_i^+)/\tau_u)}{\sum_{j=1}^{k}\exp(\operatorname{sim}(h_i,\hat h_j)/\tau_u)}.

    Each synthetic negative is then updated for tt gradient-ascent steps. At one step, with learning rate β>0\beta>0, its update is

    h^j←h^j+βg(h^j)∥g(h^j)∥2,g(h^j)=∇h^jLU(hi,hi+,{h^ℓ}ℓ=1k).\hat h_j\leftarrow\hat h_j+\beta\frac{g(\hat h_j)}{\lVert g(\hat h_j)\rVert_2},\qquad g(\hat h_j)=\nabla_{\hat h_j}L_U(h_i,h_i^+,\{\hat h_\ell\}_{\ell=1}^k).

    The optimized vectors are intended to move toward non-uniform points in the full semantic space, so contrasting with them encourages broader and more uniform sentence representations rather than relying only on anisotropic representations of real sentences.

  3. Knowl 3 — Instance weighting for false-negative suppression

    equation

    DCLR uses a separately trained complementary SimCSE encoder to estimate whether a negative is semantically too close to the original sentence. For any negative representation h−h^-—either an in-batch negative or a noise-based negative—and original-sentence representation hih_i, the complementary model computes cosine similarity sim⁡C(hi,h−)\operatorname{sim}_C(h_i,h^-). Given threshold ϕ\phi, DCLR assigns the binary weight

    αh−={0,sim⁡C(hi,h−)≥ϕ,1,sim⁡C(hi,h−)<ϕ.\alpha_{h^-}=\begin{cases} 0,&\operatorname{sim}_C(h_i,h^-)\geq\phi,\\ 1,&\operatorname{sim}_C(h_i,h^-)<\phi. \end{cases}

    Thus, a negative judged sufficiently similar to the input is treated as a false negative and removed from the contrastive denominator. With τ>0\tau>0 denoting the contrastive temperature, the encoder is trained using

    L=−log⁡exp⁡(sim⁡(hi,hi+)/τ)∑h−∈Hinoise∪Hibatchαh−exp⁡(sim⁡(hi,h−)/τ),L=-\log\frac{\exp(\operatorname{sim}(h_i,h_i^+)/\tau)}{\sum_{h^-\in\mathcal{H}_i^{\mathrm{noise}}\cup\mathcal{H}_i^{\mathrm{batch}}}\alpha_{h^-}\exp(\operatorname{sim}(h_i,h^-)/\tau)},

    where Hinoise\mathcal{H}_i^{\mathrm{noise}} is the set of optimized synthetic negatives and Hibatch\mathcal{H}_i^{\mathrm{batch}} is the set of randomly sampled in-batch negatives.

  4. Knowl 4 — Training and evaluation configuration

    experimental setup

    DCLR was trained on 1,000,000 sentences randomly sampled from Wikipedia and evaluated on seven semantic textual similarity datasets: STS 2012, STS 2013, STS 2014, STS 2015, STS 2016, STS Benchmark, and SICK-Relatedness. Each dataset contains sentence pairs with human similarity scores from 0 to 5. Evaluation used the Spearman correlation between human scores and cosine similarities of the learned sentence embeddings.

    The authors trained BERT-base, RoBERTa-base, BERT-large, and RoBERTa-large models for three epochs with Adam and contrastive temperature τ=0.05\tau=0.05. The batch size was 128 for the base models and 256 for the large models. The learning rate was 3×10−53\times10^{-5} for BERT-base, RoBERTa-base, and BERT-large, and 1×10−51\times10^{-5} for RoBERTa-large. The instance-weighting thresholds ϕ\phi were, respectively, 0.900.90, 0.850.85, 0.900.90, and 0.850.85. The ratio kk of synthetic negatives to batch size was 1, 2.5, 4, and 5 for the same four backbones. Synthetic negatives used standard deviation σ=1\sigma=1, were updated four times, and used noise-update learning rate β=10−3\beta=10^{-3}. Models were evaluated every 150 training steps on the STS-B and SICK-R development sets, and the best checkpoint was used for test evaluation.

  5. Knowl 5 — Main semantic textual similarity results

    data/table

    The main comparison on page 6 evaluates Spearman correlation on seven STS tasks. The table below reports DCLR against SimCSE, the strongest competing contrastive baseline by average score for each backbone. DCLR obtains the highest average in all four backbone settings, although it is slightly below SimCSE on a few individual datasets.

    Backbone Method STS12 STS13 STS14 STS15 STS16 STS-B SICK-R Avg.
    BERT-base SimCSE 68.40 82.41 74.38 80.91 78.56 76.85 72.23 76.25
    BERT-base DCLR 70.81 83.73 75.11 82.56 78.44 78.31 71.59 77.22
    BERT-large SimCSE 70.88 84.16 76.43 84.50 79.76 79.26 73.88 78.41
    BERT-large DCLR 71.87 84.83 77.37 84.70 79.81 79.55 74.19 78.90
    RoBERTa-base SimCSE 70.16 81.77 73.24 81.36 80.65 80.22 68.56 76.57
    RoBERTa-base DCLR 70.01 83.08 75.09 83.66 81.06 81.86 70.33 77.87
    RoBERTa-large SimCSE 72.86 83.99 75.62 84.77 81.80 81.98 71.26 78.90
    RoBERTa-large DCLR 73.09 84.57 76.13 85.15 81.99 82.35 71.80 79.30

    The average improvement over SimCSE is 0.97 points for BERT-base, 0.49 for BERT-large, 1.30 for RoBERTa-base, and 0.40 for RoBERTa-large. The full comparison also included GloVe, USE, native PLM pooling, Flow, Whitening, Contrastive-BT, ConSERT, and SG-OPT; DCLR generally outperformed these alternatives as well.

  6. Knowl 6 — Both debiasing components are necessary

    data/table

    The ablation study on the seven-task STS average, reported on page 7, shows that removing either DCLR component reduces performance. It also compares alternative implementations of the negative-generation or weighting mechanisms.

    Model STS-Avg.
    BERT-base + Ours 77.22
    Without noise-based negatives 76.17
    Without instance weighting 76.31
    BERT-base + random noise 75.22
    BERT-base + knowledge distillation 75.05
    BERT-base + self instance weighting 73.93

    Removing instance weighting causes a 0.91-point drop from the complete model, while removing noise-based negatives causes a 1.05-point drop. Directly using unoptimized random noise, distilling SimCSE into the model, or using the model itself as the complementary weighting model performs worse than the proposed DCLR design.

  7. Knowl 7 — DCLR transfers across positive augmentation strategies

    empirical result

    The paper tests DCLR while keeping its negative-sampling and weighting mechanisms fixed but replacing dropout positives with five other augmentation strategies: feature cutoff, span cutoff, token cutoff, token shuffling, and dropout. On the seven-task STS test average, DCLR improves the corresponding baseline for every augmentation strategy. The comparison chart on page 7 shows dropout producing the highest overall performance among the tested positive-construction methods, supporting the authors' choice to use dropout because it preserves the original sentence semantics more reliably than the other tested transformations.

  8. Knowl 8 — DCLR improves the measured uniformity of sentence embeddings

    equation

    The paper measures representation uniformity with

    ℓuniform=log⁡Exi,xj∼i.i.d.pdata[exp⁡(−2∥f(xi)−f(xj)∥22)],\ell_{\mathrm{uniform}}=\log\mathbb{E}_{x_i,x_j\overset{\mathrm{i.i.d.}}\sim p_{\mathrm{data}}}\left[\exp\left(-2\lVert f(x_i)-f(x_j)\rVert_2^2\right)\right],

    where xix_i and xjx_j are independently sampled sentences, pdatap_{\mathrm{data}} is the training-sentence distribution, and f(x)f(x) is the learned sentence representation. A smaller value indicates better uniformity. On BERT-base, the training curves on the STS-B validation setting show DCLR's uniformity loss below SimCSE's for nearly the entire training process; DCLR also decreases faster, whereas SimCSE shows no comparable downward trend. The result is consistent with the intended effect of adding optimized noise-based negatives outside the pretrained representation cone.

  9. Knowl 9 — Robustness under severe training-data reduction

    empirical result

    Using BERT-base, the authors retrain DCLR with training-data proportions ranging from 100% down to 0.3% and evaluate on STS-B and SICK-Relatedness. The performance curves on page 8 remain comparatively stable as the available data decreases. At the most extreme 0.3% setting, the reported performance drop relative to the full-data setting is only 9% on STS-B and 4% on SICK-Relatedness. The authors interpret this as evidence that DCLR remains effective in data-scarce settings, although the paper does not provide a theoretical guarantee of few-shot robustness.

  10. Knowl 10 — Sensitivity to weighting threshold and synthetic-negative proportion

    empirical result

    The paper analyzes two DCLR hyperparameters with BERT-base on STS-B and SICK-Relatedness: the false-negative weighting threshold ϕ\phi and the synthetic-negative proportion kk, where kk is the number of noise-based negatives relative to batch size. STS-B is sensitive to ϕ\phi: thresholds that are too high fail to suppress enough false negatives, while thresholds that are too low can incorrectly remove true negatives. SICK-Relatedness is comparatively insensitive to ϕ\phi. Performance is generally strongest when kk is close to 1, meaning that the number of synthetic negatives is close to the batch size; larger or smaller proportions can reduce the balance between uniformity improvement and alignment preservation.

Coverage note — No substantial technical contribution was omitted; background, related work, references, acknowledgements, and the non-technical ethical/future-work discussion were excluded.

References

  1. 1.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Iñigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe. 2015. Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability. In NAACL-HLT, pages 252–263.
  2. 2.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014. Semeval-2014 task 10: Multilingual semantic textual similarity. In COLING, pages 81–91.
  3. 3.Eneko Agirre, Carmen Banea, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016. Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation. In COLING, pages 497–511.
  4. 4.Eneko Agirre, Daniel M. Cer, Mona T. Diab, and Aitor Gonzalez-Agirre. 2012. Semeval-2012 task 6: A pilot on semantic textual similarity. In NAACL-HLT, pages 385–393.
  5. 5.Eneko Agirre, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. *sem 2013 shared task: Semantic textual similarity. In *SEM, pages 32–43.
  6. 6.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  7. 7.Shuqing Bian, Wayne Xin Zhao, Kun Zhou, Jing Cai, Yancheng He, Cunxiang Yin, and Ji-Rong Wen. 2021. Contrastive curriculum learning for sequential user behavior modeling via data augmentation. In CIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021, pages 3737–3746. ACM.
  8. 8.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In EMNLP, pages 632–642.
  9. 9.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018. Universal sentence encoder for english. In EMNLP, pages 169–174.
  10. 10.Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In ACL, pages 1–14.
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A simple framework for contrastive learning of visual representations. In ICML, volume 119 of Proceedings of Machine Learning Research, pages 1597–1607.
  12. 12.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. In LREC.
  13. 13.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. In EMNLP, pages 670–680.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pages 4171–4186.
  15. 15.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings. In EMNLP-IJCNLP, pages 55–65.
  16. 16.Hongchao Fang and Pengtao Xie. 2020. CERT: contrastive self-supervised learning for language understanding. CoRR, abs/2005.12766.
  17. 17.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In EMNLP, pages 6894–6910. Association for Computational Linguistics.
  18. 18.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. In CVPR, pages 1735–1742.
  19. 19.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In CVPR, pages 9726–9735.
  20. 20.Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016. Learning distributed representations of sentences from unlabelled data. In NAACL-HLT, pages 1367–1377.
  21. 21.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the knowledge in a neural network. CoRR, abs/1503.02531.
  22. 22.Junjie Huang, Duyu Tang, Wanjun Zhong, Shuai Lu, Linjun Shou, Ming Gong, Daxin Jiang, and Nan Duan. 2021. Whiteningbert: An easy unsupervised sentence embedding approach. CoRR, abs/2104.01767.
  23. 23.Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020. SMART: robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In ACL, pages 2177–2190.
  24. 24.Taeuk Kim, Kang Min Yoo, and Sang-goo Lee. 2021. Self-guided contrastive learning for BERT sentence representations. In ACL, pages 2528–2540.
  25. 25.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR.
  26. 26.Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 3294–3302.
  27. 27.Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. 2017. Adversarial examples in the physical world. In ICLR.
  28. 28.Quoc V. Le and Tomás Mikolov. 2014. Distributed representations of sentences and documents. In ICML, volume 32 of JMLR Workshop and Conference Proceedings, pages 1188–1196.
  29. 29.Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. In EMNLP, pages 9119–9130.
  30. 30.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  31. 31.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards deep learning models resistant to adversarial attacks. In ICLR.
  32. 32.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparrelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In LREC, pages 216–223.
  33. 33.Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pages 3111–3119.
  34. 34.Takeru Miyato, Andrew M. Dai, and Ian J. Goodfellow. 2017. Adversarial training methods for semi-supervised text classification. In ICLR.
  35. 35.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. 2019. Virtual adversarial training: A regularization method for supervised and semi-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell., 41(8):1979–1993.
  36. 36.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In EMNLP, pages 1532–1543.
  37. 37.Ruizhi Qiao, Lingqiao Liu, Chunhua Shen, and Anton Van Den Hengel. 2016. Less is more: zero-shot learning from online textual documents with noise suppression. In CVPR, pages 2249–2257.
  38. 38.Chongli Qin, James Martens, Sven Gowal, Dilip Krishnan, Krishnamurthy Dvijotham, Alhussein Fawzi, Soham De, Robert Stanforth, and Pushmeet Kohli. 2019. Adversarial robustness through local linearization. In NeurIPS, pages 13824–13833.
  39. 39.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In EMNLP-IJCNLP, pages 3980–3990.
  40. 40.Candace Ross, Boris Katz, and Andrei Barbu. 2020. Measuring social biases in grounded vision and language embeddings. arXiv preprint arXiv:2002.08911.
  41. 41.Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou. 2021. Whitening sentence representations for better semantics and faster retrieval. CoRR, abs/2103.15316.
  42. 42.Haipeng Sun, Rui Wang, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2020. Robust unsupervised neural machine translation with adversarial denoising training. In COLING, pages 4239–4250.
  43. 43.Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL-HLT, pages 1112–1122.
  44. 44.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In EMNLP - Demos, pages 38–45.
  45. 45.Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020. CLEAR: contrastive learning for sentence representation. CoRR, abs/2012.15466.
  46. 46.Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021. Consert: A contrastive framework for self-supervised sentence representation transfer. In ACL/IJCNLP, pages 5065–5075.
  47. 47.Yuanhang Zhou, Kun Zhou, Wayne Xin Zhao, Cheng Wang, Peng Jiang, and He Hu. 2022. C²-crs: Coarse-to-fine contrastive learning for conversational recommender system. In WSDM ’22: The Fifteenth ACM International Conference on Web Search and Data Mining, Virtual Event / Tempe, AZ, USA, February 21 - 25, 2022, pages 1488–1496. ACM.
  48. 48.Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020. Freelb: Enhanced adversarial training for natural language understanding. In ICLR.

Citation

MLA
Zhou, K., et al. “Debiased Contrastive Learning of Unsupervised Sentence Representations”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 6120–30, https://doi.org/10.18653/v1/2022.acl-long.423.
APA
Zhou, K., Zhang, B., Zhao, X., & Wen, J.-R. (2022). Debiased Contrastive Learning of Unsupervised Sentence Representations. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6120–6130. https://doi.org/10.18653/v1/2022.acl-long.423
Chicago
Zhou, K., B. Zhang, X. Zhao, and J.-R. Wen. 2022. “Debiased Contrastive Learning of Unsupervised Sentence Representations”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6120–30. https://doi.org/10.18653/v1/2022.acl-long.423.
Harvard
Zhou, K. et al. (2022) “Debiased Contrastive Learning of Unsupervised Sentence Representations”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6120–6130. Available at: https://doi.org/10.18653/v1/2022.acl-long.423.
Vancouver
1. Zhou K, Zhang B, Zhao X, Wen J-R (2022) Debiased Contrastive Learning of Unsupervised Sentence Representations. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6120–6130

BibTeX

@inproceedings{zhou-etal-2022-debiased,
    title = "Debiased Contrastive Learning of Unsupervised Sentence Representations",
    author = "Zhou, Kun  and
      Zhang, Beichen  and
      Zhao, Xin  and
      Wen, Ji-Rong",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.423/",
    doi = "10.18653/v1/2022.acl-long.423",
    pages = "6120--6130"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/