PromptBERT: Improving BERT Sentence Embeddings with Prompts
Ting JiangJian JiaoShaohan HuangZihan ZhangDeqing WangFuzhen ZhuangFuru WeiHaizhen HuangDenvy DengQi Zhang
Introduces a prompt-based contrastive learning framework with template denoising that overcomes token embedding biases in pretrained transformers to significantly outperform SimCSE on unsupervised sentence representation benchmarks.
Generating high-quality sentence embeddings—numerical representations that capture the meaning of whole sentences—is essential for core natural language processing applications like search, document retrieval, and semantic text comparison. While advanced language models such as BERT and RoBERTa have driven significant progress, their original, off-the-shelf versions paradoxically perform poorly at generating sentence representations, often falling behind traditional, simpler word-embedding techniques. Past research primarily attributed this flaw to directional narrowness (anisotropy) in the embedding space, but this diagnosis fails to explain why standard model layers degrade sentence representations.
The article aims to uncover the true causes of BERT's poor sentence representation and proposes a prompt-based framework, named PromptBERT, to transform how pre-trained models generate and optimize sentence embeddings without requiring complex architectural overhauls.
The researchers evaluated model representations across standard semantic textual similarity benchmarks and transfer tasks. They analyzed the mathematical geometry and token distributions of standard models, tested prompt-based formatting (framing sentences into fill-in-the-blank templates such as "This sentence: '[X]' means [MASK]"), and implemented contrastive learning with a novel template-denoising objective across both unsupervised and supervised settings.
The investigation produced several critical findings. First, the article reveals that the primary causes of poor sentence representation are static word-embedding biases—driven by token frequency, capitalization, and subwords—combined with ineffective processing in original upper layers, rather than directional narrowness alone. Second, simply wrapping sentences in discrete or continuous prompts allows unmodified models to bypass these biases and outperform existing post-processing baselines, boosting performance on the semantic textual similarity benchmark from 52.57 up to 73.59. Third, in fine-tuned unsupervised settings, PromptBERT and PromptRoBERTa establish new state-of-the-art benchmarks, outperforming leading contrastive learning methods (such as SimCSE) by 2.29 and 2.58 points, respectively. Finally, the proposed template-denoising technique delivers substantially higher training stability, reducing performance variance across random runs by more than 80% compared to previous contrastive learning techniques.
These findings indicate that organizations can significantly boost semantic search and text retrieval accuracy without costly model pre-training from scratch or relying heavily on expensive human-labeled data. By using prompt templates and denoising contrastive objectives, unsupervised systems can achieve performance levels that closely rival supervised alternatives, lowering labeling costs and reducing production variance risks.
Organizations deploying transformer models for search and semantic tasks should adopt prompt-based extraction methods and contrastive template-denoising objectives. Engineering teams should prioritize prompt formulation and explore continuous prompt optimization during pipeline fine-tuning rather than relying on standard layer-averaging techniques.
The findings are supported by consistent evaluations across multiple benchmarks and random seeds; however, a noted limitation is that discrete prompt templates still require manual search and engineering, as fully automated text-generation methods tested in the article did not match human-crafted templates. Stakeholders can proceed with high confidence when implementing manual or continuous prompt strategies, while reserving further validation for domain-specific vocabulary and automated prompt-generation pipelines.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). SimCSE establishes the foundational contrastive learning framework and benchmark baseline for unsupervised sentence embeddings that PromptBERT directly analyzes and surpasses.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). This paper analyzes the geometric anisotropy and layer-wise representations of BERT embeddings, providing the diagnostic foundation that PromptBERT investigates and refines.
- Paper: Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Nils Reimers et al. (2019). Sentence-BERT introduces standard methods and benchmarks for deriving sentence representations from BERT, establishing the core problem setting PromptBERT seeks to improve.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). LM-BFF demonstrates how cloze-style prompt formatting and template engineering elicit better representations from pre-trained language models, directly inspiring PromptBERT's prompt-based approach.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). P-Tuning introduces continuous prompt optimization to avoid discrete prompt brittleness, which PromptBERT adapts into its continuous prompt framework for sentence representation.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). BERT introduces the masked language model architecture whose token-level and upper-layer biases PromptBERT diagnoses and reforms for sentence embedding generation.
- Paper: RoBERTa: A Robustly Optimized BERT Pretraining Approach, Yinhan Liu et al. (2019). RoBERTa provides the pre-trained masked language model backbone evaluated alongside BERT in PromptBERT's comparative experiments.
- Paper: Supervised Learning of Universal Sentence Representations from Natural Language Inference Data, Alexis Conneau et al. (2017). This work establishes standard natural language inference supervised training techniques and transfer evaluation protocols for universal sentence representations.
- Paper: DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings, Yung-Sung Chuang et al. (2022). DiffCSE extends unsupervised contrastive sentence embedding techniques beyond standard dropout and prompt methods by incorporating difference-based token prediction.
- Paper: Debiased Contrastive Learning of Unsupervised Sentence Representations, Kun Zhou et al. (2022). DCLR builds on contrastive sentence representation methods by directly penalizing false negative samples and synthesizing negatives across the vector space.
- Paper: Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives, Yanzhao Zhang et al. (2022). MixCSE addresses representation degradation in contrastive sentence learning by continually synthesizing hard negatives through feature mixing.
- Paper: RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank, Jiduan Liu et al. (2023). RankCSE advances unsupervised sentence embeddings by replacing standard binary contrastive objectives with listwise ranking principles and ranking distillation.
- Paper: PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval, Shengyao Zhuang et al. (2024). PromptReps expands the concept of prompt-driven representations from encoder models to modern large language models for dual dense and sparse zero-shot retrieval.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This work generalizes prompt-based text embeddings by employing in-context learning demonstrations to adapt large language model embedders in few-shot settings.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Qwen3 Embedding scales foundation-model embedding paradigms to large multilingual datasets and modern instruction-aware representation architectures.
