A Simple but Tough-to-Beat Baseline for Sentence Embeddings
Sanjeev AroraYingyu LiangTengyu Ma
Proposes an unsupervised sentence embedding baseline combining smooth inverse frequency weighting with principal component removal that consistently outperforms complex neural network models on semantic similarity benchmarks.
Natural language processing systems often require converting full sentences into numerical vector representations, known as embeddings, to measure meaning and similarity. Standard industry practices have increasingly turned to complex, computationally heavy neural networks that demand extensive labeled training datasets. However, obtaining sufficient annotated data is expensive, and complex architectures often struggle to transfer effectively across different domains.
The article evaluates whether a lightweight, completely unsupervised mathematical approach can outperform resource-intensive supervised neural architectures when generating sentence representations. Its core objective is to formulate a principled sentence embedding technique that avoids complex model retraining while establishing a rigorous, highly competitive baseline for textual similarity and classification tasks.
The evaluated method operates in two straightforward steps: it calculates a weighted average of pre-existing word vectors using a smooth inverse frequency weighting scheme to downweight frequent words, and it removes the first principal component across sentence vectors to eliminate common syntactic noise. The researchers grounded this approach in a theoretical generative text model and validated it across 22 standard semantic textual similarity benchmark datasets, as well as downstream supervised tasks including textual entailment and sentiment analysis.
The findings show that this simple technique consistently matches or outperforms complex alternatives. On semantic textual similarity tasks, the unsupervised method improved performance over basic unweighted word averaging by 10% to 30%, surpassing supervised recurrent neural networks and long short-term memory models. Applying the method to semi-supervised embeddings achieved the highest performance across most benchmark datasets. The approach proved highly robust across varying weighting parameters and different background text sources, and ablation results confirmed that both the frequency weighting and the principal component removal independently drive significant performance gains.
These results demonstrate that organizations do not necessarily need expensive computing infrastructure, massive labeled training datasets, or complex neural architectures to achieve top-tier sentence representation. A simple linear algebraic adjustment can achieve equivalent or superior performance at a fraction of the cost, timeline, and operational risk, challenging the prevailing assumption that deeper neural models are inherently superior for semantic representation.
Teams working on text similarity and representation tasks should adopt this method as a mandatory, standard baseline before investing in computationally demanding architectures. In low-resource or out-of-domain environments, practitioners can deploy this unsupervised approach directly to reduce compute and data collection costs. Future engineering efforts should explore hybrid systems that combine this semantic weighting method with sequential models to capture word order effectively.
Decision-makers should note that the approach relies on word embeddings that ignore word order and syntax, which can lead to lower accuracy in tasks heavily dependent on negation or sentiment classification. Nevertheless, given the extensive evaluation across 22 benchmarks, confidence in the method’s efficacy for semantic similarity and general representation tasks remains exceptionally high.
- Paper: GloVe: Global Vectors for Word Representation, Jeffrey Pennington et al. (2014). It introduces GloVe embeddings, which serve as one of the primary foundational word vector representations averaged and evaluated in the baseline method.
- Paper: Distributed Representations of Words and Phrases and their Compositionality, Tomas Mikolov et al. (2013). It establishes the Skip-gram word representation framework that underlies word-vector averaging baselines for sentence semantics.
- Paper: Skip-Thought Vectors, Ryan Kiros et al. (2015). It introduces unsupervised encoder-decoder sentence embeddings (Skip-Thought), which the source paper directly benchmarks against and outperforms with its weighted averaging approach.
- Paper: Distributed Representations of Sentences and Documents, Quoc V. Le et al. (2014). It presents Paragraph Vector, an early standard unsupervised method for sentence and document representations that contextualizes simple composition baselines.
- Paper: Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks, Kai Sheng Tai et al. (2015). It develops Tree-LSTM architectures for sentence-level semantic representations, providing the complex neural baseline that the source's simpler method aims to challenge.
- Paper: Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank, Richard Socher et al. (2013). It demonstrates recursive composition models over parse trees, illustrating the structured neural alternatives compared against weighted bag-of-words averaging.
- Paper: Supervised Learning of Universal Sentence Representations from Natural Language Inference Data, Alexis Conneau et al. (2017). It introduces InferSent, demonstrating that supervised sentence embeddings trained on NLI data provide a strong alternative paradigm to unsupervised weighted-average baselines.
- Paper: Universal Sentence Encoder, Daniel Cer et al. (2018). It extends sentence representation learning to transformer and deep averaging architectures (Universal Sentence Encoder) trained across diverse web-scale corpora.
- Paper: Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Nils Reimers et al. (2019). It introduces Sentence-BERT to produce efficient, dense sentence embeddings via siamese transformer networks, heavily improving upon static vector averaging.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). It presents SimCSE, a contrastive learning method that addresses representation anisotropy in sentence embeddings to establish a modern state-of-the-art baseline.
- Paper: A Structured Self-attentive Sentence Embedding, Zhouhan Lin et al. (2017). It advances beyond 1D vector averaging by learning 2D matrix sentence embeddings through self-attention to capture multiple semantic aspects.
- Paper: Deep contextualized word representations, Matthew E. Peters et al. (2018). It establishes deep contextualized word representations (ELMo), moving NLP past the static word vectors utilized by the source's simple baseline.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). It builds general-purpose dense text embeddings (E5) via weakly-supervised contrastive pre-training across broad text retrieval benchmarks.
