A Simple but Tough-to-Beat Baseline for Sentence Embeddings

Sanjeev AroraYingyu LiangTengyu Ma

article2017ICLR1,431 citations

Proposes an unsupervised sentence embedding baseline combining smooth inverse frequency weighting with principal component removal that consistently outperforms complex neural network models on semantic similarity benchmarks.

Listen

Natural language processing systems often require converting full sentences into numerical vector representations, known as embeddings, to measure meaning and similarity. Standard industry practices have increasingly turned to complex, computationally heavy neural networks that demand extensive labeled training datasets. However, obtaining sufficient annotated data is expensive, and complex architectures often struggle to transfer effectively across different domains.

The article evaluates whether a lightweight, completely unsupervised mathematical approach can outperform resource-intensive supervised neural architectures when generating sentence representations. Its core objective is to formulate a principled sentence embedding technique that avoids complex model retraining while establishing a rigorous, highly competitive baseline for textual similarity and classification tasks.

The evaluated method operates in two straightforward steps: it calculates a weighted average of pre-existing word vectors using a smooth inverse frequency weighting scheme to downweight frequent words, and it removes the first principal component across sentence vectors to eliminate common syntactic noise. The researchers grounded this approach in a theoretical generative text model and validated it across 22 standard semantic textual similarity benchmark datasets, as well as downstream supervised tasks including textual entailment and sentiment analysis.

The findings show that this simple technique consistently matches or outperforms complex alternatives. On semantic textual similarity tasks, the unsupervised method improved performance over basic unweighted word averaging by 10% to 30%, surpassing supervised recurrent neural networks and long short-term memory models. Applying the method to semi-supervised embeddings achieved the highest performance across most benchmark datasets. The approach proved highly robust across varying weighting parameters and different background text sources, and ablation results confirmed that both the frequency weighting and the principal component removal independently drive significant performance gains.

These results demonstrate that organizations do not necessarily need expensive computing infrastructure, massive labeled training datasets, or complex neural architectures to achieve top-tier sentence representation. A simple linear algebraic adjustment can achieve equivalent or superior performance at a fraction of the cost, timeline, and operational risk, challenging the prevailing assumption that deeper neural models are inherently superior for semantic representation.

Teams working on text similarity and representation tasks should adopt this method as a mandatory, standard baseline before investing in computationally demanding architectures. In low-resource or out-of-domain environments, practitioners can deploy this unsupervised approach directly to reduce compute and data collection costs. Future engineering efforts should explore hybrid systems that combine this semantic weighting method with sequential models to capture word order effectively.

Decision-makers should note that the approach relies on word embeddings that ignore word order and syntax, which can lead to lower accuracy in tasks heavily dependent on negation or sentiment classification. Nevertheless, given the extensive evaluation across 22 benchmarks, confidence in the method’s efficacy for semantic similarity and general representation tasks remains exceptionally high.

Cover for A Simple but Tough-to-Beat Baseline for Sentence Embeddings

Abstract

The success of neural network methods for computing word embeddings has motivated methods for generating semantic embeddings of longer pieces of text, such as sentences and paragraphs. Surprisingly, Wieting et al (ICLR'16) showed that such complicated methods are outperformed, especially in out-of-domain (transfer learning) settings, by simpler methods involving mild retraining of word embeddings and basic linear regression. The method of Wieting et al. requires retraining with a substantial labeled dataset such as Paraphrase Database (Ganitkevitch et al., 2013).

The current paper goes further, showing that the following completely unsupervised sentence embedding is a formidable baseline: Use word embeddings computed using one of the popular methods on unlabeled corpus like Wikipedia, represent the sentence by a weighted average of the word vectors, and then modify them a bit using PCA/SVD. This weighting improves performance by about 10% to 30% in textual similarity tasks, and beats sophisticated supervised methods including RNN's and LSTM's. It even improves Wieting et al.'s embeddings. This simple method should be used as the baseline to beat in future, especially when labeled training data is scarce or nonexistent.

The paper also gives a theoretical explanation of the success of the above unsupervised method using a latent variable generative model for sentences, which is a simple extension of the model in Arora et al. (TACL'16) with new "smoothing" terms that allow for words occuring out of context, as well as high probabilities for words like and, not in all contexts.

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORK
  • 3 A SIMPLE METHOD FOR SENTENCE EMBEDDING
  • 3.1 CONNECTION TO SUBSAMPLING PROBABILITIES IN WORD2VEC
  • 4 EXPERIMENTS
  • 4.1 TEXTUAL SIMILARITY TASKS
  • 4.1.1 EFFECT OF WEIGHTING PARAMETER ON PERFORMANCE
  • 4.2 SUPERVISED TASKS
  • 4.3 THE EFFECT OF THE ORDER OF WORDS IN SENTENCES
  • 5 CONCLUSIONS
  • 6 ACKNOWLEDGEMENTS
  • REFERENCES
  • A DETAILS OF EXPERIMENTAL SETTING
  • A.1 UNSUPERVISED TASK: TEXTUAL SIMILARITY
  • A.2 SUPERVISED TASKS
  • A.3 ADDITIONAL SUPERVISED TASKS

Knowls

  1. Knowl 1 — Smooth Inverse Frequency (SIF) Sentence Embedding Algorithm

    algorithm

    The Smooth Inverse Frequency (SIF) sentence embedding algorithm computes an unsupervised vector representation for a sentence by calculating a frequency-weighted average of its constituent word vectors and subtracting the projection along the first principal component of the sentence vectors across the corpus.

    Input: Word embeddings {v_w in R^d : w in V}, sentence collection S, weighting parameter a > 0, word frequency distribution {p(w) : w in V}
    Output: Sentence embeddings {v_s in R^d : s in S}
    for each sentence s in S do
        v_s <- (1 / |s|) * sum_{w in s} (a / (a + p(w))) * v_w
    end for
    Compute the first principal component u of {v_s : s in S}
    for each sentence s in S do
        v_s <- v_s - u * (u^T * v_s)
    end for
    return {v_s : s in S}

    The parameter aa is a scalar hyperparameter typically set between 10−410^{-4} and 10−310^{-3}, and p(w)p(w) is the marginal unigram probability of word ww estimated from an unlabeled background corpus. The subtraction of uuopvsu u^ op v_s removes the dominant common component across sentence embeddings, which primarily reflects syntactic frequency artifacts rather than semantic discourse content.

  2. Knowl 2 — Generative Latent Variable Model for Sentences with Smoothing and Syntax Discourse

    model/method

    The generation of words in a sentence ss is modeled as a dynamic process driven by a sentence discourse vector c_s hicksim ext{Uniform}( ext{unit sphere}) ext{ in } ar{\mathbb{R}}^d. Words w∈Vw \in V possess static word vectors vw∈Rdv_w \in \mathbb{R}^d. To account for words occurring out of context and common stop words appearing regardless of discourse, the probability of emitting word ww in sentence ss given csc_s is modeled as:

    Pr⁡[w emitted in sentence s∣cs]=αp(w)+(1−α)exp⁡(⟨c~s,vw⟩)Zc~s\Pr[w \text{ emitted in sentence } s \mid c_s] = \alpha p(w) + (1 - \alpha) \frac{\exp(\langle \tilde{c}_s, v_w \rangle)}{Z_{\tilde{c}_s}}

    where:

    c~s=βc0+(1−β)cs,c0⊥cs\tilde{c}_s = \beta c_0 + (1 - \beta) c_s, \quad c_0 \perp c_s

    Here, α,β∈[0,1]\alpha, \beta \in [0,1] are scalar hyperparameters, p(w)p(w) is the unigram probability of word ww across the whole corpus, c0∈Rdc_0 \in \mathbb{R}^d is a common discourse vector representing frequent syntax-driven word co-occurrences, and Zc~s=∑w∈Vexp⁡(⟨c~s,vw⟩)Z_{\tilde{c}_s} = \sum_{w \in V} \exp(\langle \tilde{c}_s, v_w \rangle) is the partition function.

  3. Knowl 3 — Maximum Likelihood Estimation of Sentence Discourse Vectors

    theoretical result

    Assuming word vectors vwv_w are roughly uniformly dispersed in Rd\mathbb{R}^d, the partition function Zc~s=∑w∈Vexp⁡(⟨c~s,vw⟩)Z_{\tilde{c}_s} = \sum_{w \in V} \exp(\langle \tilde{c}_s, v_w \rangle) is approximately invariant across directions, behaving as a constant ZZ. Under this condition and by first-order Taylor expansion around c~s=0\tilde{c}_s = 0, the maximum likelihood estimator for the modified discourse vector c~s\tilde{c}_s on the unit sphere for sentence ss is given by:

    arg⁡max⁡∥c~s∥=1∑w∈slog⁡Pr⁡[w∣cs]∝∑w∈sap(w)+avw\arg\max_{\|\tilde{c}_s\| = 1} \sum_{w \in s} \log \Pr[w \mid c_s] \propto \sum_{w \in s} \frac{a}{p(w) + a} v_w

    where a=1−ααZa = \frac{1 - \alpha}{\alpha Z}.

    The unconstrained maximum likelihood estimate of c~s\tilde{c}_s is therefore proportional to the weighted average of the word vectors, where the smooth inverse frequency weight a/(p(w)+a)a / (p(w) + a) naturally downweights high-frequency words. The latent discourse vector csc_s is recovered by subtracting the orthogonal component along c0c_0, which corresponds to subtracting the projection on the first principal component of {c~s:s∈S}\{\tilde{c}_s : s \in S\}.

  4. Knowl 4 — Equivalence of Word2Vec Subsampling Heuristic to Expected SIF Gradients

    theoretical result

    In the Continuous Bag-of-Words (CBOW) architecture of Word2Vec, the vanilla gradient of the loss with respect to word vector vwv_w updates along the unweighted average direction of the context window vˉt=15∑i=15vwt−i\bar{v}_t = \frac{1}{5} \sum_{i=1}^5 v_{w_{t-i}}.

    The subsampling heuristic randomly keeps context word wt−kw_{t-k} with probability:

    q(wt−k)=min⁡(1,10−5p(wt−k))q(w_{t-k}) = \min\left(1, \sqrt{\frac{10^{-5}}{p(w_{t-k})}}\right)

    The expected value of the resulting sampled stochastic gradient is:

    E[∇~g(vw)]=α∑k=15q(wt−k)vwt−k\mathbb{E}[\tilde{\nabla} g(v_w)] = \alpha \sum_{k=1}^5 q(w_{t-k}) v_{w_{t-k}}

    The subsampling probability function q(w)q(w) closely matches the SIF weight curve aa+p(w)\frac{a}{a + p(w)} when a=10−4a = 10^{-4}. Consequently, Word2Vec with subsampling functions as a stochastic gradient update targeting the SIF-weighted discourse vector rather than an unweighted average.

  5. Knowl 5 — Textual Similarity Benchmark Results for SIF Sentence Embeddings

    data/table

    The table below evaluates sentence embeddings on Semantic Textual Similarity (STS 2012–2015), SICK 2014 Semantic Relatedness, and Twitter 2015 tasks using Pearson's correlation coefficient (r×100r \times 100). GloVe+WR denotes the fully unsupervised SIF approach (weighting + common component removal) on GloVe vectors, while PSL+WR applies the same method to PARAGRAM-SL999 (PSL) vectors.

    Supervision Supervised Unsupervised / Semi-sup. Our approach
    Tasks PP PP-proj. DAN RNN iRNN LSTM(no) LSTM(o.g.) ST avg-GloVe tfidf-GloVe GloVe+WR PSL+WR
    STS'12 58.7 60.0 56.0 48.1 58.4 51.0 46.4 30.8 52.5 58.7 56.2 59.5
    STS'13 55.8 56.8 54.2 44.7 56.7 45.2 41.5 24.8 42.3 52.1 56.6 61.8
    STS'14 70.9 71.3 69.5 57.7 70.9 59.8 51.5 31.4 54.2 63.8 68.5 73.5
    STS'15 75.8 74.8 72.7 57.2 75.6 63.9 56.0 31.0 52.7 60.6 71.7 76.3
    SICK'14 71.6 71.6 70.7 61.2 71.2 63.9 59.0 49.8 65.9 69.4 72.2 72.9
    Twitter'15 52.9 52.8 53.7 45.1 52.9 47.6 36.1 24.7 30.3 33.8 48.0 49.0

    GloVe+WR improves upon unweighted GloVe averaging (avg-GloVe) by 10% to 30% relative across tasks, outperforming supervised RNN and LSTM models. PSL+WR achieves the highest score across four of the six benchmark suites and outperforms all supervised models initialized with the same PSL embeddings.

  6. Knowl 6 — Ablation of SIF Weighting and Common Component Removal

    empirical result

    The individual and combined contributions of smooth inverse frequency weighting (W) and common component removal (R) were evaluated on textual similarity benchmarks:

    • On standard GloVe embeddings, applying SIF weighting alone (GloVe+W) improves Pearson's rr over unweighted averaging (avg-GloVe) by approximately 5% on average. Applying common component removal alone (GloVe+R) improves performance by approximately 10%. Combining both mechanisms (GloVe+WR) achieves an average improvement of 13%.
    • On PSL embeddings, PSL+W and PSL+R each achieve approximately a 10% improvement over avg-PSL independently, while the combined approach (PSL+WR) achieves a 13% improvement.

    Both the frequency-based downweighting and the geometric subtraction of the shared syntax direction c0c_0 are necessary to obtain maximum similarity correlation.

  7. Knowl 7 — Sensitivity of SIF Embeddings to Hyperparameter Choice and Corpus Statistics

    empirical result

    The SIF sentence embedding method exhibits high stability across hyperparameter selections and corpus sources for estimating word frequencies p(w)p(w):

    • Weight parameter aa: When evaluated across values a∈{10−i,3×10−i:1≤i≤5}a \in \{10^{-i}, 3 \times 10^{-i} : 1 \le i \le 5\}, performance remains consistently superior to unweighted averaging across a broad spectrum, with optimal performance attained for a∈[10−4,10−3]a \in [10^{-4}, 10^{-3}].
    • Reference corpus for p(w)p(w): Estimating unigram frequencies p(w)p(w) from diverse corpora—including English Wikipedia (enwiki, 3B tokens), political blogs (poliblogs, 5M tokens), Common Crawl (commoncrawl, 800B tokens), and text8 (1M tokens)—produces nearly identical Pearson correlation scores on STS 2012 benchmarks when a=10−3a = 10^{-3}.
  8. Knowl 8 — Performance of Frozen SIF Embeddings on Supervised Downstream Tasks

    data/table

    Sentence embeddings computed in an unsupervised manner using SIF were evaluated as frozen features for downstream supervised tasks. Classifiers were trained on top of linear projections of the embeddings to 2400 dimensions.

    Task PP DAN RNN LSTM(no) LSTM(o.g.) skip-thought Ours (SIF)
    Similarity (SICK) 84.9 85.96 73.13 85.45 83.41 85.8 86.03
    Entailment (SICK) 83.1 84.5 76.4 83.2 82.0 - 84.6
    Sentiment (SST) 79.4 83.4 86.5 86.6 89.2 - 82.2
    SNLI (3-class acc.) - - 72.2 77.6 - - 78.2

    For SICK similarity, metric is Pearson's r×100r \times 100; for SICK entailment, SST binary sentiment, and SNLI natural language inference, metric is accuracy (%). SIF embeddings outperform or match fully supervised neural architectures (DAN, RNN, LSTM) on SICK similarity, SICK entailment, and SNLI classification, despite the embeddings being constructed without labeled training data.

  9. Knowl 9 — Effect of Word Order Shuffling on Sentence Representation Models

    empirical result

    To analyze the importance of word order versus lexical semantics in sentence embeddings, supervised recurrent neural architectures were trained and tested on original benchmarks versus versions where the word order in every sentence was randomly shuffled:

    • On SICK similarity (r×100r \times 100), RNN dropped from 73.13 to 54.50, LSTM without output gates dropped from 85.45 to 77.24, and LSTM with output gates dropped from 83.41 to 79.39.
    • On SICK entailment (accuracy %), RNN dropped from 76.4 to 61.7, LSTM (no) dropped from 83.2 to 78.2, and LSTM (o.g.) dropped from 82.0 to 81.0.
    • On SST sentiment (accuracy %), RNN dropped from 86.5 to 84.2, LSTM (no) dropped from 86.6 to 82.9, and LSTM (o.g.) dropped from 89.2 to 84.1.

    While word order contributes to recurrent model accuracy, order-agnostic SIF embeddings match or exceed the performance of unshuffled recurrent models on similarity and entailment, indicating that simple weighted lexical composition captures semantic discourse content more effectively than sequence modeling in these benchmarks.

  10. Knowl 10 — Limitations of SIF Embeddings in Sentiment Classification

    limitation

    SIF sentence embeddings underperform supervised recurrent models (such as LSTMs with output gates) on sentiment classification tasks (e.g., scoring 82.2% vs 89.2% on SST binary classification). This limitation arises from two principal issues:

    1. The Antonym Problem in Distributional Word Vectors: Standard word embeddings map antonym pairs (such as 'good' and 'bad') to vectors with high cosine similarity because they share similar context distributions, obscuring sentiment polarity.
    2. Downweighting of Functional Sentiment Modifiers: High-frequency grammatical words such as negation particles (e.g., 'not') receive small weights under the smooth inverse frequency scheme aa+p(w)\frac{a}{a+p(w)}, suppressing critical sentiment switches in bag-of-words compositions.

Coverage note — None omitted; full individual per-subtask breakdowns from Table 5 and the small sample IMDB comparisons were summarized into the primary benchmark and downstream task knowls.

References

  1. 1.Eneko Agirre, Mona Diab, Daniel Cer, and Aitor Gonzalez-Agirre. Semeval-2012 task 6: A pilot on semantic textual similarity. In Proceedings of the First Joint Conference on Lexical and Computational Semantics-Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation, pp. 385–393. Association for Computational Linguistics, 2012.
  2. 2.Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. Sem 2013 shared task: Semantic textual similarity. in second joint conference on lexical and computational semantics. In Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity, 2013.
  3. 3.Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. Semeval-2014 task 10: Multilingual semantic textual similarity. In Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014), pp. 81–91, 2014.
  4. 4.Eneko Agirrea, Carmen Baneab, Claire Cardiec, Daniel Cerd, Mona Diabe, Aitor Gonzalez-Agirrea, Weiwei Guof, Inigo Lopez-Gazpioa, Montse Maritxalara, Rada Mihalceab, et al. Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability. In Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015), pp. 252–263, 2015.
  5. 5.Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. A latent variable model approach to PMI-based word embeddings. Transaction of Association for Computational Linguistics, 2016.
  6. 6.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin. A neural probabilistic language model. Journal of Machine Learning Research, 2003.
  7. 7.William Blacoe and Mirella Lapata. A comparison of vector-based representations for semantic composition. In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, 2012.
  8. 8.Phil Blunsom, Edward Grefenstette, and Nal Kalchbrenner. A convolutional neural network for modelling sentences. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, 2014.
  9. 9.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2015.
  10. 10.Christian Buck, Kenneth Heafield, and Bas van Ooyen. N-gram counts and language models from the common crawl. In Proceedings of the Language Resources and Evaluation Conference, 2014.
  11. 11.Ronan Collobert and Jason Weston. A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of the 25th International Conference on Machine Learning, 2008.
  12. 12.Scott C. Deerwester, Susan T Dumais, Thomas K. Landauer, George W. Furnas, and Richard A. Harshman. Indexing by latent semantic analysis. Journal of the American Society for Information Science, 1990.
  13. 13.Felix A Gers, Nicol N Schraudolph, and Jurgen Schmidhuber. Learning precise timing with lstm recurrent networks. Journal of machine learning research, 2002.
  14. 14.Tatsunori B. Hashimoto, David Alvarez-Melis, and Tommi S. Jaakkola. Word embeddings as metric recovery in semantic spaces. Transactions of the Association for Computational Linguistics, 2016.
  15. 15.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural computation, 1997.
  16. 16.Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daume III. Deep unordered composition rivals syntactic methods for text classification. In Proceedings of the Association for Computational Linguistics, 2015.
  17. 17.Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Skip-thought vectors. In Advances in neural information processing systems, 2015.
  18. 18.Quoc Le and Tomas Mikolov. Distributed representations of sentences and documents. In Proceedings of The 31st International Conference on Machine Learning, 2014.
  19. 19.J. Lei Ba, J. R. Kiros, and G. E. Hinton. Layer Normalization. ArXiv e-prints, 2016.
  20. 20.Omer Levy and Yoav Goldberg. Neural word embedding as implicit matrix factorization. In Advances in Neural Information Processing Systems, 2014.
  21. 21.Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, pp. 142–150. Association for Computational Linguistics, 2011.
  22. 22.Matt Mahoney. Wikipedia text preprocess script. http://mattmahoney.net/dc/textdata.html, 2008. Accessed Mar-2015.
  23. 23.Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. Semeval-2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment. SemEval-2014, 2014.
  24. 24.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems, 2013a.
  25. 25.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. Linguistic regularities in continuous space word representations. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2013b.
  26. 26.Jeff Mitchell and Mirella Lapata. Vector-based models of semantic composition. In Association for Computational Linguistics, 2008.
  27. 27.Jeff Mitchell and Mirella Lapata. Composition in distributional models of semantics. Cognitive science, 2010.
  28. 28.Ellie Pavlick, Pushpendre Rastogi, Juri Ganitkevitch, Benjamin Van Durme2, and Chris Callison-Burch. Ppdb 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2015.
  29. 29.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. Proceedings of the Empiricial Methods in Natural Language Processing, 2014.
  30. 30.Stephen Robertson. Understanding inverse document frequency: on theoretical arguments for idf. Journal of documentation, 2004.
  31. 31.Richard Socher, Eric H Huang, Jeffrey Pennin, Christopher D Manning, and Andrew Y Ng. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection. In Advances in Neural Information Processing Systems, 2011.
  32. 32.Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the conference on empirical methods in natural language processing (EMNLP), 2013.
  33. 33.Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng. Grounded compositional semantics for finding and describing images with sentences. Transactions of the Association for Computational Linguistics, 2014.
  34. 34.Karen Sparck Jones. A statistical interpretation of term specificity and its application in retrieval. Journal of documentation, 1972.
  35. 35.Kai Sheng Tai, Richard Socher, and Christopher D Manning. Improved semantic representations from tree-structured long short-term memory networks. arXiv preprint arXiv:1503.00075, 2015.
  36. 36.Sida Wang and Christopher D Manning. Baselines and bigrams: Simple, good sentiment and topic classification. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2, pp. 90–94. Association for Computational Linguistics, 2012.
  37. 37.Yashen Wang, Heyan Huang, Chong Feng, Qiang Zhou, Jiahui Gu, and Xiong Gao. Cse: Conceptual sentence embeddings based on attention model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2016.
  38. 38.John Wieting, Mohit Bansal, Kevin Gimpel, Karen Livescu, and Dan Roth. From paraphrase database to compositional paraphrase model and back. Transactions of the Association for Computational Linguistics, 2015.
  39. 39.John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. Towards universal paraphrastic sentence embeddings. In International Conference on Learning Representations, 2016.
  40. 40.Wikimedia. English Wikipedia dump. http://dumps.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2, 2012. Accessed Mar-2015.
  41. 41.Wei Xu, Chris Callison-Burch, and William B Dolan. Semeval-2015 task 1: Paraphrase and semantic similarity in twitter (pit). Proceedings of SemEval, 2015.
  42. 42.Tae Yano, William W Cohen, and Noah A Smith. Predicting response to political blog posts with topic models. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 2009.

Citation

MLA
Arora, S., et al. “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”. International Conference on Learning Representations, 2017, https://oar.princeton.edu/bitstream/88435/pr1rk2k/1/BaselineSentenceEmbedding.pdf.
APA
Arora, S., Liang, Y., & Ma, T. (2017). A Simple but Tough-to-Beat Baseline for Sentence Embeddings. International Conference on Learning Representations. https://oar.princeton.edu/bitstream/88435/pr1rk2k/1/BaselineSentenceEmbedding.pdf
Chicago
Arora, S., Y. Liang, and T. Ma. 2017. “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”. International Conference on Learning Representations. https://oar.princeton.edu/bitstream/88435/pr1rk2k/1/BaselineSentenceEmbedding.pdf.
Harvard
Arora, S., Liang, Y. and Ma, T. (2017) “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”, International Conference on Learning Representations [Preprint]. Available at: https://oar.princeton.edu/bitstream/88435/pr1rk2k/1/BaselineSentenceEmbedding.pdf.
Vancouver
1. Arora S, Liang Y, Ma T (2017) A Simple but Tough-to-Beat Baseline for Sentence Embeddings. International Conference on Learning Representations

BibTeX

@article{arora2017simple,
  title = {A Simple but Tough-to-Beat Baseline for Sentence Embeddings},
  author = {Arora, Sanjeev and Liang, Yingyu and Ma, Tengyu},
  year = {2017},
  journal = {International Conference on Learning Representations},
  url = {https://oar.princeton.edu/bitstream/88435/pr1rk2k/1/BaselineSentenceEmbedding.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors