PromptBERT: Improving BERT Sentence Embeddings with Prompts

Ting JiangJian JiaoShaohan HuangZihan ZhangDeqing WangFuzhen ZhuangFuru WeiHaizhen HuangDenvy DengQi Zhang

article2022EMNLP158 citations

Introduces a prompt-based contrastive learning framework with template denoising that overcomes token embedding biases in pretrained transformers to significantly outperform SimCSE on unsupervised sentence representation benchmarks.

Listen

Generating high-quality sentence embeddings—numerical representations that capture the meaning of whole sentences—is essential for core natural language processing applications like search, document retrieval, and semantic text comparison. While advanced language models such as BERT and RoBERTa have driven significant progress, their original, off-the-shelf versions paradoxically perform poorly at generating sentence representations, often falling behind traditional, simpler word-embedding techniques. Past research primarily attributed this flaw to directional narrowness (anisotropy) in the embedding space, but this diagnosis fails to explain why standard model layers degrade sentence representations.

The article aims to uncover the true causes of BERT's poor sentence representation and proposes a prompt-based framework, named PromptBERT, to transform how pre-trained models generate and optimize sentence embeddings without requiring complex architectural overhauls.

The researchers evaluated model representations across standard semantic textual similarity benchmarks and transfer tasks. They analyzed the mathematical geometry and token distributions of standard models, tested prompt-based formatting (framing sentences into fill-in-the-blank templates such as "This sentence: '[X]' means [MASK]"), and implemented contrastive learning with a novel template-denoising objective across both unsupervised and supervised settings.

The investigation produced several critical findings. First, the article reveals that the primary causes of poor sentence representation are static word-embedding biases—driven by token frequency, capitalization, and subwords—combined with ineffective processing in original upper layers, rather than directional narrowness alone. Second, simply wrapping sentences in discrete or continuous prompts allows unmodified models to bypass these biases and outperform existing post-processing baselines, boosting performance on the semantic textual similarity benchmark from 52.57 up to 73.59. Third, in fine-tuned unsupervised settings, PromptBERT and PromptRoBERTa establish new state-of-the-art benchmarks, outperforming leading contrastive learning methods (such as SimCSE) by 2.29 and 2.58 points, respectively. Finally, the proposed template-denoising technique delivers substantially higher training stability, reducing performance variance across random runs by more than 80% compared to previous contrastive learning techniques.

These findings indicate that organizations can significantly boost semantic search and text retrieval accuracy without costly model pre-training from scratch or relying heavily on expensive human-labeled data. By using prompt templates and denoising contrastive objectives, unsupervised systems can achieve performance levels that closely rival supervised alternatives, lowering labeling costs and reducing production variance risks.

Organizations deploying transformer models for search and semantic tasks should adopt prompt-based extraction methods and contrastive template-denoising objectives. Engineering teams should prioritize prompt formulation and explore continuous prompt optimization during pipeline fine-tuning rather than relying on standard layer-averaging techniques.

The findings are supported by consistent evaluations across multiple benchmarks and random seeds; however, a noted limitation is that discrete prompt templates still require manual search and engineering, as fully automated text-generation methods tested in the article did not match human-crafted templates. Stakeholders can proceed with high confidence when implementing manual or continuous prompt strategies, while reserving further validation for domain-specific vocabulary and automated prompt-generation pipelines.

arXiv: 2201.04337kongds/Prompt-BERT
Cover for PromptBERT: Improving BERT Sentence Embeddings with Prompts

Abstract

We propose PromptBERT, a novel contrastive learning method for learning better sentence representation. We firstly analyze the drawback of current sentence embedding from original BERT and find that it is mainly due to the static token embedding bias and ineffective BERT layers. Then we propose the first prompt-based sentence embeddings method and discuss two prompt representing methods and three prompt searching methods to make BERT achieve better sentence embeddings. Moreover, we propose a novel unsupervised training objective by the technology of template denoising, which substantially shortens the performance gap between the supervised and unsupervised settings. Extensive experiments show the effectiveness of our method. Compared to SimCSE, PromptBert achieves 2.29 and 2.58 points of improvement based on BERT and RoBERTa in the unsupervised setting. Our code is available at https://github.com/kongds/Prompt-BERT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Rethinking the Sentence Embeddings of the Original BERT
  • 4 Prompt Based Sentence Embeddings
  • 4.1 Represent Sentence with the Prompt
  • 4.2 Prompt Search
  • 4.3 Prompt Based Contrastive Learning with Template Denoising
  • 5 Experiments
  • 5.1 Dataset
  • 5.2 Baselines
  • 5.3 Implementation Details
  • 5.4 Non Fine-Tuned BERT Results
  • 5.5 Fine-Tuned BERT Results
  • 5.6 Effectiveness of Prompt Based Contrastive Learning with Template Denoising
  • 6 Discussion
  • 6.1 Template Denoising
  • 6.2 Stability in Unsupervised Contrastive Learning
  • 7 Conclusion
  • 8 Limitation
  • 9 Acknowledgments
  • References
  • A Static Token Embeddings Biases
  • A.1 Eliminating Biases by Removing Tokens
  • A.2 Eliminating Biases by Pre-training
  • B Training Details
  • C Transfer Tasks

Knowls

  1. Knowl 1 — Prompt-Based Sentence Representation Using Masked Token Embeddings

    model/method

    PromptBERT represents a sentence by formulating sentence representation as a masked language modeling (MLM) task. Given an input sentence xinx_{\text{in}}, it is mapped to a prompt input xpromptx_{\text{prompt}} using a template containing a sentence placeholder [X][X] and a masked token [MASK][\text{MASK}], such as:

    This sentence : “[X]” means [MASK].\text{This sentence : ``}[X]\text{'' means }[\text{MASK}] .

    Feeding xpromptx_{\text{prompt}} into a pre-trained transformer language model (such as BERT or RoBERTa), the sentence representation vector h∈Rdh \in \mathbb{R}^d is defined directly as the output hidden vector of the [MASK][\text{MASK}] token:

    h=h[MASK]h = h_{[\text{MASK}]}

    This method leverages the contextual representations and large-scale knowledge of original pre-trained layers while circumventing token embedding biases (such as token frequency, subwords, and capitalization biases) that corrupt simple pooling of static token embeddings.

  2. Knowl 2 — Prompt-Based Contrastive Learning with Template Denoising

    model/method

    In PromptBERT, positive pairs for unsupervised contrastive learning are formed by processing the same input sentence xix_i through two distinct prompt templates, generating two semantically consistent representations from different viewpoints. Because prompt templates themselves introduce template-specific bias vectors into the output [MASK][\text{MASK}] representations, template denoising is used to subtract the template bias.

    For a sentence xix_i consisting of LiL_i tokens, the model computes the templated sentence embedding hih_i using a template. To compute the corresponding template bias h^i\hat{h}_i, the model evaluates the template alone without xix_i, shifting the position IDs of all template tokens following [X][X] by LiL_i so that their positional embeddings match those in the full templated sequence. The denoised sentence representation is then:

    hidenoised=hi−h^ih_i^{\text{denoised}} = h_i - \hat{h}_i

    During inference at evaluation time, a single template without denoising is used.

  3. Knowl 3 — Template Denoising Contrastive Loss Formulation

    equation

    The contrastive learning objective for a mini-batch of NN sentences {x1,…,xN}\{x_1, \dots, x_N\} under template denoising is formulated as:

    ℓi=−log⁡exp⁡(cos⁡(hi−h^i,hi′−h^i′)/τ)∑j=1Nexp⁡(cos⁡(hi−h^i,hj′−h^j′)/τ)\ell_i = -\log \frac{\exp\left(\cos\left(h_i - \hat{h}_i, h'_i - \hat{h}'_i\right) / \tau\right)}{\sum_{j=1}^N \exp\left(\cos\left(h_i - \hat{h}_i, h'_j - \hat{h}'_j\right) / \tau\right)}

    where:

    • hih_i and hi′h'_i are the [MASK][\text{MASK}] hidden embeddings of sentence xix_i obtained using two different prompt templates TT and T′T'.
    • h^i\hat{h}_i and h^i′\hat{h}'_i are the respective template bias embeddings computed by passing TT and T′T' alone with position IDs adjusted to match the length of xix_i.
    • cos⁡(u,v)=u⊤v∥u∥2∥v∥2\cos(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2} denotes the cosine similarity.
    • τ>0\tau > 0 is a temperature hyperparameter.
    • NN is the mini-batch size.

    In supervised fine-tuning where labeled positive and negative pairs are provided, template denoising is applied using the same template across instances.

  4. Knowl 4 — Identification and Analysis of Static Token Embedding Biases in Pre-trained Language Models

    empirical result

    Static token embeddings in pre-trained models (such as BERT-base-uncased, BERT-base-cased, and RoBERTa-base) exhibit three distinct distributional biases:

    1. Frequency Bias: High-frequency tokens cluster densely together, whereas low-frequency tokens are dispersed sparsely. In BERT, begin-of-word tokens are more sensitive to frequency than subwords, whereas in RoBERTa, subword tokens are more sensitive.
    2. Subword Bias: Token embeddings cluster into separate regions for whole/begin-of-word tokens versus WordPiece/BPE subwords.
    3. Case Bias: In cased models, uppercase begin-of-word tokens cluster separately from lowercase begin-of-word tokens.

    Measuring token-level anisotropy by the average pairwise cosine similarity between all static token embeddings yields 0.4445 for bert-base-uncased, 0.1465 for bert-base-cased, and 0.0235 for roberta-base. Even though roberta-base static token embeddings are almost isotropic (0.0235), they still suffer from frequency and subword biases, demonstrating that embedding bias is distinct from anisotropy.

  5. Knowl 5 — Performance Gain on Semantic Textual Similarity by Removing Biased Static Tokens

    empirical result

    Averaging static token embeddings after manually removing biased tokens substantially improves sentence embedding performance on Semantic Textual Similarity (STS) benchmarks without any model fine-tuning or transformer contextual layers:

    Embedding Configuration BERT-base-cased BERT-base-uncased RoBERTa-base
    Raw Static Token Avg. 56.93 56.02 55.88
    −- Top-36 Frequent Tokens 60.27 59.65 65.41
    −- Freq. Subwords 64.83 62.20 64.89
    −- Freq. Subwords Uppercase 65.07 – 65.06
    −- Freq. Subwords Case Punctuation 66.05 63.10 67.64

    Metric is average Spearman correlation (ρ×100\rho \times 100) over 7 STS datasets (STS 2012--2016, STS-B, and SICK-R). Removing biased tokens improves performance by +9.12 points on BERT-base-cased, +7.08 points on BERT-base-uncased, and +11.76 points on RoBERTa-base. Filtered RoBERTa-base static token averaging (67.64) outperforms post-processing methods such as BERT-flow (66.55) and BERT-whitening (66.28).

  6. Knowl 6 — MLM Classification Head Weight Tying as Root Cause of Static Token Embedding Bias

    empirical result

    In standard BERT pre-training, tying the input static token embedding weights with the output masked language modeling (MLM) classification head weight matrix transfers the output gradient bias into the input token embeddings.

    Comparing two identical BERT-like models trained with the MLM objective for 125,000 steps with batch size 2,000—one tying and one untying the static embedding and MLM head weights—shows that static token embeddings in the untied model exhibit substantially less clustering bias by frequency, subword, and case. On STS benchmark evaluations without fine-tuning, the average Spearman correlation is:

    • 49.41 for static token embeddings of the untied weights model
    • 45.68 for static token embeddings of the tied weights model
    • 43.33 for the MLM classification head weights of the untied model
  7. Knowl 7 — Prompt Search Strategies and Template Engineering for Sentence Embeddings

    empirical result

    Template formulation critically determines sentence embedding quality. Evaluating search methods on the STS-B development set using frozen bert-base-uncased demonstrates:

    1. Greedy Manual Search: Dividing templates into relationship tokens and prefix tokens wrapping [X][X] indicates that explicit relational phrasing and quotes improve correlation:
    • [X] [MASK] . achieves 39.34
    • [X] means [MASK] . achieves 63.56
    • This sentence : "[X]" means [MASK] . achieves 73.44 (an improvement of +34.10 over simple concatenation).
    1. T5 Automatic Template Generation: Generating templates from dictionary word definitions (e.g., matching words with dictionary definitions via T5) yielded a top template Also called [MASK]. [X], achieving a Spearman correlation of 64.75 on STS-B dev, falling short of manual templates due to the task gap between word definitions and sentence semantics.

    2. OptiPrompt (Continuous Prompt Tuning): Initializing continuous prompt vectors with manual template embeddings and optimizing them via unsupervised contrastive learning with frozen BERT parameters improves STS-B dev Spearman correlation from 73.44 to 80.90.

  8. Knowl 8 — Semantic Textual Similarity Benchmark Performance of PromptBERT

    data/table

    Evaluation of PromptBERT and PromptRoBERTa across 7 Semantic Textual Similarity (STS) test sets (STS12--STS16, STS-B, and SICK-R) via SentEval. Reported values are Spearman rank correlation ρ×100\rho \times 100:

    Method STS12 STS13 STS14 STS15 STS16 STS-B SICK-R Avg.
    Unsupervised Fine-tuned Models
    IS-BERTbase_{\text{base}} 56.77 69.24 61.21 75.23 70.16 69.21 64.25 66.58
    ConSERTbase_{\text{base}} 64.64 78.49 69.07 79.72 75.95 73.97 67.31 72.74
    SimCSE-BERTbase_{\text{base}} 68.40 82.41 74.38 80.91 78.56 76.85 72.23 76.25
    PromptBERTbase_{\text{base}} 71.56 84.58 76.98 84.47 80.60 81.60 69.87 78.54
    RoBERTabase_{\text{base}}-whitening 46.99 63.24 57.23 71.36 68.99 61.36 62.91 61.73
    SimCSE-RoBERTabase_{\text{base}} 70.16 81.77 73.24 81.36 80.65 80.22 68.56 76.57
    PromptRoBERTabase_{\text{base}} 73.94 84.74 77.28 84.99 81.74 81.88 69.50 79.15
    Supervised Fine-tuned Models
    InferSent-GloVe 52.86 66.75 62.15 72.77 66.87 68.03 65.65 65.01
    SBERTbase_{\text{base}} 70.97 76.53 73.19 79.09 74.30 77.03 72.91 74.89
    ConSERTbase_{\text{base}} 74.07 83.93 77.05 83.66 78.76 81.36 76.77 79.37
    SimCSE-BERTbase_{\text{base}} 75.30 84.67 80.19 85.40 80.82 84.25 80.39 81.57
    PromptBERTbase_{\text{base}} 75.48 85.59 80.57 85.99 81.08 84.56 80.52 81.97
    SimCSE-RoBERTabase_{\text{base}} 76.53 85.21 80.95 86.03 82.57 85.83 80.50 82.52
    PromptRoBERTabase_{\text{base}} 76.75 85.93 82.28 86.69 82.80 86.14 80.04 82.95

    PromptBERTbase_{\text{base}} outperforms SimCSE-BERTbase_{\text{base}} by +2.29 points in the unsupervised setting (78.54 vs. 76.25), and PromptRoBERTabase_{\text{base}} outperforms SimCSE-RoBERTabase_{\text{base}} by +2.58 points (79.15 vs. 76.57). In the non-fine-tuned setting on BERT-base-uncased, manual prompt BERT achieves 67.85 and manual + OptiPrompt achieves 73.59, outperforming BERT last-layer averaging (52.57) and BERT-flow (66.55).

  9. Knowl 9 — Training Stability and Ablation of Unsupervised Contrastive Objectives

    empirical result

    Ablation over unsupervised contrastive objectives across 10 random seeds highlights the advantage of pairing different templates with template denoising:

    Training Objective BERTbase_{\text{base}} Avg. STS RoBERTabase_{\text{base}} Avg. STS
    Same template (inner dropout noise) 78.16±0.1778.16 \pm 0.17 78.16±0.4478.16 \pm 0.44
    Different templates (no denoising) 78.19±0.2978.19 \pm 0.29 78.17±0.4478.17 \pm 0.44
    Different templates with denoising (PromptBERT) 78.54±0.15\mathbf{78.54} \pm \mathbf{0.15} 79.15±0.25\mathbf{79.15} \pm \mathbf{0.25}

    When evaluated over 10 random seeds on BERTbase_{\text{base}}:

    • SimCSE-BERTbase_{\text{base}} yields a mean of 75.42±0.8675.42 \pm 0.86 (maximum 76.64, minimum 73.50, range 3.14).
    • PromptBERTbase_{\text{base}} yields a mean of 78.54±0.1578.54 \pm 0.15 (maximum 78.86, minimum 78.33, range 0.53), demonstrating higher stability and reduced sensitivity to initialization seeds.
  10. Knowl 10 — Transfer Task Performance on SentEval Classification Benchmarks

    data/table

    Evaluation of frozen sentence representations on 7 transfer classification benchmarks from SentEval (MR, CR, SUBJ, MPQA, SST-2, TREC, and MRPC) using logistic regression classifiers:

    Method MR CR SUBJ MPQA SST-2 TREC MRPC Avg.
    Unsupervised models
    Avg. BERT embedding 78.66 86.25 94.37 88.66 84.40 92.80 69.54 84.94
    BERT-[CLS] embedding 78.68 84.85 94.21 88.23 84.13 91.40 71.13 84.66
    IS-BERT 81.09 87.18 94.96 88.75 85.96 88.64 74.24 85.83
    SimCSE-BERT 81.18 86.46 94.45 88.88 85.50 89.80 74.43 85.81
    PromptBERT 80.74 85.49 93.65 89.32 84.95 88.20 76.06 85.49
    SimCSE-RoBERTa 81.04 87.74 93.28 86.94 86.60 84.60 73.68 84.84
    PromptRoBERTa 83.82 88.72 93.19 90.36 88.08 90.60 76.75 87.36
    Supervised models
    InferSent-GloVe 81.57 86.54 92.50 90.38 84.18 88.20 75.77 85.59
    Universal Sentence Encoder 80.09 85.19 93.98 86.70 86.38 93.20 70.14 85.10
    SBERT 83.64 89.43 94.39 89.86 88.96 89.60 76.00 87.41
    SimCSE-BERT 82.69 89.25 94.81 89.59 87.31 88.40 73.51 86.51
    PromptBERT 83.14 89.38 94.49 89.93 87.37 87.40 76.58 86.90
    SRoBERTa 84.91 90.83 92.56 88.75 90.50 88.60 78.14 87.76
    SimCSE-RoBERTa 84.92 92.00 94.11 89.82 91.27 88.80 75.65 88.08
    PromptRoBERTa 85.74 91.47 94.81 90.93 92.53 90.40 77.10 89.00

    PromptRoBERTa achieves the highest overall performance on transfer tasks, reaching 87.36% in the unsupervised setting (+2.52 points over SimCSE-RoBERTa) and 89.00% in the supervised setting (+0.92 points over SimCSE-RoBERTa).

  11. Knowl 11 — Limitations of Prompt-Based Sentence Representation

    limitation

    PromptBERT's primary performance gains rely on manually hand-crafted discrete prompt templates. Automatic template generation based on T5 and dictionary word definitions underperformed manual templates (achieving 64.75 vs. 73.44 Spearman correlation on the STS-B development set) due to the semantic gap between word definitions and full sentence semantics. Although continuous templates via OptiPrompt demonstrate that continuous representations can overcome manual template limitations in frozen models, finding an effective, fully automated discrete template generation mechanism remains an unaddressed challenge.

Coverage note — None was omitted; all contributed methods, empirical findings on static token biases, prompt formulation and search techniques, contrastive loss equations, and benchmark evaluations are represented.

References

  1. 1.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. 2015. Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability. In Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015), pages 252–263.
  2. 2.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014. Semeval-2014 task 10: Multilingual semantic textual similarity. In Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014), pages 81–91.
  3. 3.Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez Agirre, Rada Mihalcea, German Rigau Claramunt, and Janyce Wiebe. 2016. Semeval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation. In SemEval-2016. 10th International Workshop on Semantic Evaluation; 2016 Jun 16-17; San Diego, CA. Stroudsburg (PA): ACL; 2016. p. 497-511. ACL (Association for Computational Linguistics).
  4. 4.Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012. Semeval-2012 task 6: A pilot on semantic textual similarity. In * SEM 2012: The First Joint Conference on Lexical and Computational Semantics–Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), pages 385–393.
  5. 5.Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. * sem 2013 shared task: Semantic textual similarity. In Second joint conference on lexical and computational semantics (* SEM), volume 1: proceedings of the Main conference and the shared task: semantic textual similarity, pages 32–43.
  6. 6.Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2017. A simple but tough-to-beat baseline for sentence embeddings. In International conference on learning representations.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  8. 8.Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055.
  9. 9.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018. Universal sentence encoder for english. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 169–174.
  10. 10.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. arXiv preprint arXiv:1803.05449.
  11. 11.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. arXiv preprint arXiv:1705.02364.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186.
  13. 13.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings. arXiv preprint arXiv:1909.00512.
  14. 14.Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019. Representation degeneration problem in training natural language generation models. arXiv preprint arXiv:1907.12009.
  15. 15.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021a. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830.
  16. 16.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021b. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821.
  17. 17.Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from bert for semantic textual similarity. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9119–9130.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  19. 19.Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, Roberto Zamparelli, et al. 2014. A sick cure for the evaluation of compositional distributional semantic models. In Lrec, pages 216–223. Reykjavik.
  20. 20.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543.
  21. 21.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992.
  22. 22.Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou. 2021. Whitening sentence representations for better semantics and faster retrieval. arXiv preprint arXiv:2103.15316.
  23. 23.Hayato Tsukagoshi, Ryohei Sasano, and Koichi Takeda. 2021. Defsent: Sentence embeddings using definition sentences. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 411–418.
  24. 24.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355.
  25. 25.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.
  26. 26.Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021. Consert: A contrastive framework for self-supervised sentence representation transfer. arXiv preprint arXiv:2105.11741.
  27. 27.Yan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim, and Lidong Bing. 2020. An unsupervised sentence embedding method by mutual information maximization. arXiv preprint arXiv:2009.12061.
  28. 28.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5017–5033.

Citation

MLA
Jiang, T., et al. “PromptBERT: Improving BERT Sentence Embeddings with Prompts”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 8826–37, https://doi.org/10.18653/v1/2022.emnlp-main.603.
APA
Jiang, T., Jiao, J., Huang, S., Zhang, Z., Wang, D., Zhuang, F., Wei, F., Huang, H., Deng, D., & Zhang, Q. (2022). PromptBERT: Improving BERT Sentence Embeddings with Prompts. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8826–8837. https://doi.org/10.18653/v1/2022.emnlp-main.603
Chicago
Jiang, T., J. Jiao, S. Huang, et al. 2022. “PromptBERT: Improving BERT Sentence Embeddings with Prompts”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8826–37. https://doi.org/10.18653/v1/2022.emnlp-main.603.
Harvard
Jiang, T. et al. (2022) “PromptBERT: Improving BERT Sentence Embeddings with Prompts”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8826–8837. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.603.
Vancouver
1. Jiang T, Jiao J, Huang S, Zhang Z, Wang D, Zhuang F, Wei F, Huang H, Deng D, Zhang Q (2022) PromptBERT: Improving BERT Sentence Embeddings with Prompts. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8826–8837

BibTeX

@inproceedings{jiang-etal-2022-promptbert,
    title = "{P}rompt{BERT}: Improving {BERT} Sentence Embeddings with Prompts",
    author = "Jiang, Ting  and
      Jiao, Jian  and
      Huang, Shaohan  and
      Zhang, Zihan  and
      Wang, Deqing  and
      Zhuang, Fuzhen  and
      Wei, Furu  and
      Huang, Haizhen  and
      Deng, Denvy  and
      Zhang, Qi",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.603/",
    doi = "10.18653/v1/2022.emnlp-main.603",
    pages = "8826--8837"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/