Baselines and Bigrams: Simple, Good Sentiment and Topic Classification

Sida I. WangChristopher D. Manning

article2012ACL1,364 citations

Demonstrates that simple Naive Bayes and Support Vector Machine baselines with word bigrams—particularly the hybrid NBSVM model—can outperform complex, structure-sensitive methods across standard sentiment and topic classification benchmarks.

Listen

Organizations often deploy complex and computationally expensive machine learning models to analyze text sentiment and topic categories, under the assumption that basic baseline algorithms cannot handle nuanced language. However, standard baselines vary widely in performance depending on their specific mathematical formulation, the input features provided, and the length of the processed text. Without well-tuned baseline benchmarks, teams risk investing substantial resources into complex architectures that may deliver inferior performance compared to simpler alternatives.

The main objective of the article is to systematically evaluate simple linear classification methods—specifically variants of Naive Bayes and Support Vector Machines—across multiple text lengths and classification tasks to establish rigorous, lightweight performance baselines.

The researchers conducted an empirical evaluation using standard benchmark datasets covering short sentence snippets, full-length movie reviews up to 50,000 documents, and topic-based newsgroup classifications. The analysis evaluated simple feature representations (single words and two-word pairs) without relying on external linguistic lexicons, rule systems, or complex parsing structures, testing standard algorithms alongside a hybrid model that uses Naive Bayes log-count ratios as features within a Support Vector Machine.

The key findings demonstrate that baseline choice depends heavily on document length and task type. First, for short text snippets, Multinomial Naive Bayes consistently outperformed Support Vector Machines, achieving top accuracies between 79.0% and 93.6% and beating complex rule-based and deep neural models. Second, for long reviews, Support Vector Machines outperformed Naive Bayes, while the hybrid model delivered the strongest results (89.45% and 91.22% accuracy), matching or setting new state-of-the-art benchmarks. Third, adding two-word phrase features consistently improved sentiment classification across all datasets, whereas they offered little to no benefit for topic-based classification. Finally, the hybrid model proved exceptionally robust across all dataset lengths and task types.

These findings imply that engineering teams can achieve state-of-the-art text and sentiment classification without the computational cost, implementation complexity, or runtime overhead of highly specialized architectures. The results challenge the assumption that short text fragments require intricate, hand-crafted linguistic rules or heavy deep learning models, showing instead that simple, well-tuned linear classifiers capture sentiment patterns more effectively at a fraction of the operational cost.

Decision-makers and practitioners should adopt the hybrid model as the standard, default baseline before greenlighting more complex natural language processing models. Teams should ensure that text representations use binary word indicators rather than raw frequency counts, include two-word phrases for sentiment analysis tasks, and rely on Multinomial rather than multivariate Bernoulli Naive Bayes formulations. If deploying the hybrid model, teams should set its interpolation parameter within the recommended range of one-quarter to one-half for stable general performance.

Confidence in these findings is high across the evaluated English benchmark tasks, given the consistent empirical validation across varying sample sizes. However, readers should note that the evaluations were restricted to standard linear models on English text classifications, meaning validation on proprietary or multilingual domain data remains necessary before full operational rollout.

sidaw/nbsvmWang et al (2012).pdf
Cover for Baselines and Bigrams: Simple, Good Sentiment and Topic Classification

Abstract

Variants of Naive Bayes (NB) and Support Vector Machines (SVM) are often used as baseline methods for text classification, but their performance varies greatly depending on the model variant, features used and task/dataset. We show that: (i) the inclusion of word bigram features gives consistent gains on sentiment analysis tasks; (ii) for short snippet sentiment tasks, NB actually does better than SVMs (while for longer documents the opposite result holds); (iii) a simple but novel SVM variant using NB log-count ratios as feature values consistently performs well across tasks and datasets. Based on these observations, we identify simple NB and SVM variants which outperform most published results on sentiment analysis datasets, sometimes providing a new state-of-the-art performance level.

Table of Contents

  • 1 Introduction
  • 2 The Methods
  • 2.1 Multinomial Naive Bayes (MNB)
  • 2.2 Support Vector Machine (SVM)
  • 2.3 SVM with NB features (NBSVM)
  • 3 Datasets and Task
  • 4 Experiments and Results
  • 4.1 Experimental setup
  • 4.2 MNB is better at snippets
  • 4.3 SVM is better at full-length reviews
  • 4.4 Benefits of bigrams depends on the task
  • 4.5 NBSVM is a robust performer
  • 4.6 Other results
  • References

Knowls

  1. Knowl 1 — NBSVM: Support Vector Machine with Naive Bayes Feature Weights and Interpolation

    model/method

    NBSVM is a linear classification model that combines generative Multinomial Naive Bayes (MNB) log-count ratios as feature scalings with discriminative Support Vector Machine (SVM) training and weight interpolation.

    Let f^(i)∈{0,1}∣V∣\hat{f}^{(i)} \in \{0, 1\}^{|V|} denote the binarized feature count vector for training example ii with vocabulary VV and binary label y(i)∈{−1,+1}y^{(i)} \in \{-1, +1\}, where f^j(i)=1{fj(i)>0}\hat{f}^{(i)}_j = \mathbf{1}\{f^{(i)}_j > 0\} for raw word counts f(i)f^{(i)}. The smoothed positive and negative class count vectors are defined as:

    p=α+∑i:y(i)=1f^(i),q=α+∑i:y(i)=−1f^(i)p = \alpha + \sum_{i: y^{(i)}=1} \hat{f}^{(i)}, \quad q = \alpha + \sum_{i: y^{(i)}=-1} \hat{f}^{(i)}

    where α>0\alpha > 0 is an additive smoothing parameter (set to α=1\alpha = 1). The log-count ratio vector r^∈R∣V∣\hat{r} \in \mathbb{R}^{|V|} is:

    r^=log⁡(p/∥p∥1q/∥q∥1)\hat{r} = \log \left( \frac{p / \|p\|_1}{q / \|q\|_1} \right)

    Each input vector f^(k)\hat{f}^{(k)} is transformed into f~(k)=r^∘f^(k)\tilde{f}^{(k)} = \hat{r} \circ \hat{f}^{(k)}, where ∘\circ represents elementwise multiplication. The model learns parameters w∈R∣V∣w \in \mathbb{R}^{|V|} and b∈Rb \in \mathbb{R} by minimizing the L2L_2-regularized L2L_2-loss SVM objective:

    wTw+C∑i=1Nmax⁡(0,1−y(i)(wTf~(i)+b))2w^T w + C \sum_{i=1}^N \max\left(0, 1 - y^{(i)}(w^T \tilde{f}^{(i)} + b)\right)^2

    To regularize towards the Naive Bayes model, the weight vector is interpolated as:

    w′=(1−β)wˉ+βww' = (1 - \beta)\bar{w} + \beta w

    where wˉ=∥w∥1/∣V∣\bar{w} = \|w\|_1 / |V| is the mean magnitude of ww, and β∈[0,1]\beta \in [0, 1] is the interpolation parameter. The decision rule for a test instance kk is:

    y(k)=sign(w′Tf~(k)+b)y^{(k)} = \text{sign}\left(w'^T \tilde{f}^{(k)} + b\right)

    Standard default hyperparameters are α=1\alpha = 1, C=1C = 1, and β=0.25\beta = 0.25.

  2. Knowl 2 — Binarized Multinomial Naive Bayes Linear Formulation

    model/method

    Multinomial Naive Bayes (MNB) can be parameterized directly as a linear classifier of the form y(k)=sign(wTx(k)+b)y^{(k)} = \text{sign}(w^T x^{(k)} + b) over binarized feature vectors f^(k)=1{f(k)>0}\hat{f}^{(k)} = \mathbf{1}\{f^{(k)} > 0\}, where f(k)f^{(k)} is the raw term count vector for instance kk.

    Given training cases with labels y(i)∈{−1,+1}y^{(i)} \in \{-1, +1\}, let N+N_+ and N−N_- denote the number of positive and negative training examples, and let α>0\alpha > 0 be the smoothing hyperparameter (typically α=1\alpha = 1). Smoothed count vectors p^,q^∈R∣V∣\hat{p}, \hat{q} \in \mathbb{R}^{|V|} are computed as:

    p^=α+∑i:y(i)=1f^(i),q^=α+∑i:y(i)=−1f^(i)\hat{p} = \alpha + \sum_{i: y^{(i)}=1} \hat{f}^{(i)}, \quad \hat{q} = \alpha + \sum_{i: y^{(i)}=-1} \hat{f}^{(i)}

    The weight vector ww corresponds to the log-count ratio vector r^\hat{r}:

    w=r^=log⁡(p^/∥p^∥1q^/∥q^∥1)w = \hat{r} = \log \left( \frac{\hat{p} / \|\hat{p}\|_1}{\hat{q} / \|\hat{q}\|_1} \right)

    and the scalar bias is:

    b=log⁡(N+N−)b = \log\left(\frac{N_+}{N_-}\right)

    The resulting linear prediction for test case kk is:

    y(k)=sign(r^Tf^(k)+log⁡(N+N−))y^{(k)} = \text{sign}\left(\hat{r}^T \hat{f}^{(k)} + \log\left(\frac{N_+}{N_-}\right)\right)

  3. Knowl 3 — Sentiment and Subjectivity Classification on Snippet Datasets

    data/table

    On short snippet sentiment and subjectivity datasets, binarized Multinomial Naive Bayes (MNB) and NBSVM consistently outperform standard linear Support Vector Machines (SVM) and complex models based on parse trees, Recursive Autoencoders, or hand-coded polarity reversal rules.

    Method RT-s MPQA CR Subj.
    MNB-uni 77.9 85.3 79.8 92.6
    MNB-bi 79.0 86.3 80.0 93.6
    SVM-uni 76.2 86.1 79.0 90.8
    SVM-bi 77.7 86.7 80.8 91.7
    NBSVM-uni 78.1 85.3 80.5 92.4
    NBSVM-bi 79.4 86.3 81.8 93.2
    RAE 76.8 85.7 – –
    RAE-pretrain 77.7 86.4 – –
    Voting-w/Rev. 63.1 81.7 74.2 –
    Rule 62.9 81.8 74.3 –
    BoF-noDic. 75.7 81.8 79.3 –
    BoF-w/Rev. 76.4 84.1 81.4 –
    Tree-CRF 77.3 86.1 81.4 –
    BoWSVM – – – 90.0

    Values represent 10-fold cross-validation classification accuracy (%). The datasets are: RT-s (Rotten Tomatoes sentence snippet movie reviews, average length l=21l=21), MPQA (opinion polarity subtask, l=3l=3), CR (customer review sentences, l=20l=20), and Subj. (subjectivity vs. objectivity dataset, l=24l=24). uni denotes unigram features, bi denotes unigram plus bigram features, RAE denotes Recursive Autoencoders, and Tree-CRF denotes dependency tree-based conditional random fields.

    Linear SVM is a comparatively weak baseline for short snippet classification. MNB and NBSVM beat or match sophisticated models without relying on parsers, lexicons, or unsupervised pretraining.

  4. Knowl 4 — Sentiment Classification Performance on Full-Length Review Datasets

    data/table

    On full-length document review tasks, Support Vector Machines (SVM) outperform standard Multinomial Naive Bayes (MNB), while NBSVM achieves superior performance that matches or outperforms complex neural and specialized feature representations.

    Method RT-2k IMDB Subj.
    MNB-uni 83.45 83.55 92.58
    MNB-bi 85.85 86.59 93.56
    SVM-uni 86.25 86.95 90.84
    SVM-bi 87.40 89.16 91.74
    NBSVM-uni 87.80 88.29 92.40
    NBSVM-bi 89.45 91.22 93.18
    BoW (bnc) 85.45 87.80 87.77
    BoW (bΔ\Deltat'c) 85.80 88.23 85.65
    LDA 66.70 67.42 66.65
    Full+BoW 87.85 88.33 88.45
    Full+Unlab'd+BoW 88.90 88.89 88.13
    BoWSVM 87.15 – 90.00
    Valence Shifter 86.20 – –
    tf.Δ\Deltaidf 88.10 – –
    Appr. Taxonomy 90.20 – –
    WRRBM – 87.42 –
    WRRBM + BoW(bnc) – 89.23 –

    Classification accuracy (%) is evaluated on RT-2k (standard 2000 full-length movie reviews, average length l=787l=787, 10-fold CV), IMDB (50k full-length movie reviews, average length l=231l=231, standard train/test split), and the Subj. snippet dataset (l=24l=24, 10-fold CV) for comparison.

    On long documents, the independence assumptions of MNB cause it to lag behind SVM. However, combining NB log-count ratio features with an SVM in NBSVM-bi achieves 89.45% on RT-2k and 91.22% on IMDB, exceeding Word Representation Restricted Boltzmann Machines (WRRBM) and delta-idf baselines.

  5. Knowl 5 — Task-Dependent Utility of Word Bigram Features

    data/table

    The addition of word bigram features to unigram feature sets consistently improves classification performance on sentiment and subjectivity tasks, but offers negligible or negative utility on topical text classification tasks.

    Method AthR XGraph BbCrypt
    MNB-uni 85.0 90.0 99.3
    MNB-bi 85.1 (+0.1) 91.2 (+1.2) 99.4 (+0.1)
    SVM-uni 82.6 85.1 98.3
    SVM-bi 83.7 (+1.1) 86.2 (+0.9) 97.7 (-0.5)
    NBSVM-uni 87.9 91.2 99.7
    NBSVM-bi 87.7 (-0.2) 90.7 (-0.5) 99.5 (-0.2)
    ActiveSVM – 90.0 99.0
    DiscLDA 83.0 – –

    Accuracy (%) is reported across three 20-Newsgroups pairwise topic classification tasks: AthR (alt.atheism vs. religion.misc, l=345l=345), XGraph (comp.windows.x vs. comp.graphics, l=261l=261), and BbCrypt (rec.sport.baseball vs. sci.crypt, l=269l=269). uni indicates unigram features; bi indicates unigram + bigram features.

    On topical tasks, unigram keywords are sufficiently indicative on their own; adding bigrams to NBSVM yields a slight performance drop of 0.2% to 0.5%. In sentiment tasks, by contrast, bigrams provide substantial gains across all models by capturing local context and modifier structures (e.g., negated adjectives and modified verbs). Adding trigrams, however, slightly degrades performance across tasks.

  6. Knowl 6 — L2-Regularized L2-Loss Support Vector Machine Formulation

    model/method

    The baseline Support Vector Machine (SVM) classifier is formulated as an L2L_2-regularized linear model with squared hinge loss (L2L_2-loss) over binarized feature vectors f^(k)∈{0,1}∣V∣\hat{f}^{(k)} \in \{0, 1\}^{|V|}, where f^(k)=1{f(k)>0}\hat{f}^{(k)} = \mathbf{1}\{f^{(k)} > 0\}.

    Given training pairs (f^(i),y(i))(\hat{f}^{(i)}, y^{(i)}) with y(i)∈{−1,+1}y^{(i)} \in \{-1, +1\} for i=1,…,Ni = 1, \dots, N, the weight vector w∈R∣V∣w \in \mathbb{R}^{|V|} and bias scalar b∈Rb \in \mathbb{R} are obtained by minimizing:

    wTw+C∑i=1Nmax⁡(0,1−y(i)(wTf^(i)+b))2w^T w + C \sum_{i=1}^N \max\left(0, 1 - y^{(i)}(w^T \hat{f}^{(i)} + b)\right)^2

    where C>0C > 0 is the regularization parameter (set to C=0.1C = 0.1 for standard SVM). The L2L_2-regularized L2L_2-loss objective demonstrates superior numerical stability over L1L_1-loss SVM in text classification settings. Predictions for unseen examples kk are computed as y(k)=sign(wTf^(k)+b)y^{(k)} = \text{sign}(w^T \hat{f}^{(k)} + b).

  7. Knowl 7 — Superiority of Binarized Features and Multinomial over Bernoulli Naive Bayes

    empirical result

    For Multinomial Naive Bayes (MNB) and NBSVM, using binarized feature indicators f^(k)=1{f(k)>0}\hat{f}^{(k)} = \mathbf{1}\{f^{(k)} > 0\} yields approximately 1% higher classification accuracy than utilizing raw term counts f(k)f^{(k)} on longer documents, while showing negligible difference on short snippets.

    Furthermore, Multivariate Bernoulli Naive Bayes (BNB) performs substantially worse and is less stable than MNB across text classification benchmarks, lagging behind MNB by up to 10% accuracy and only matching MNB on short snippet tasks using unigram features.

  8. Knowl 8 — Regularization and Sensitivity of the NBSVM Interpolation Parameter

    empirical result

    In the NBSVM model, the weight vector is defined via interpolation between the mean weight magnitude and the discriminatively trained SVM weights:

    w′=(1−β)wˉ+βww' = (1 - \beta)\bar{w} + \beta w

    where wˉ=∥w∥1/∣V∣\bar{w} = \|w\|_1 / |V| and β∈[0,1]\beta \in [0, 1]. This formulation acts as a regularizer: it defaults to generative Naive Bayes weights unless the discriminative SVM is sufficiently confident.

    On long document tasks (such as RT-2k and IMDB), classification accuracy is stable within 0.1% for all β∈[0.25,1.0]\beta \in [0.25, 1.0]. On short snippet datasets, setting β=0.25\beta = 0.25 provides an average 0.5% accuracy improvement over β=1.0\beta = 1.0 (pure SVM weights). Selecting β∈[0.25,0.5]\beta \in [0.25, 0.5] provides a robust configuration across both snippet and document classification tasks.

Coverage note — None was omitted; all contributed model formulations (MNB, SVM, NBSVM), experimental benchmarks across snippet, full-length review, and topical datasets, and feature/hyperparameter analyses are included.

References

  1. 1.R. Collobert and J. Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of ICML.
  2. 2.George E. Dahl, Ryan P. Adams, and Hugo Larochelle. 2012. Training restricted boltzmann machines on word observations. arXiv:1202.5695v1 [cs.LG].
  3. 3.Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin. 2008. LIBLINEAR: A library for large linear classification. Journal of Machine Learning Research, 9:1871–1874, June.
  4. 4.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings ACM SIGKDD, pages 168–177.
  5. 5.Alistair Kennedy and Diana Inkpen. 2006. Sentiment classification of movie reviews using contextual valence shifters. Computational Intelligence, 22.
  6. 6.Simon Lacoste-Julien, Fei Sha, and Michael I. Jordan. 2008. DiscLDA: Discriminative learning for dimensionality reduction and classification. In Proceedings of NIPS, pages 897–904.
  7. 7.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of ACL.
  8. 8.Justin Martineau and Tim Finin. 2009. Delta tfidf: An improved feature space for sentiment analysis. In Proceedings of ICWSM.
  9. 9.Andrew McCallum and Kamal Nigam. 1998. A comparison of event models for naive bayes text classification. In AAAI-98 Workshop, pages 41–48.
  10. 10.Vangelis Metsis, Ion Androutsopoulos, and Georgios Paliouras. 2006. Spam filtering with naive bayes - which naive bayes? In Proceedings of CEAS.
  11. 11.Karo Moilanen and Stephen Pulman. 2007. Sentiment composition. In Proceedings of RANLP, pages 378–382, September 27-29.
  12. 12.Tetsuji Nakagawa, Kentaro Inui, and Sadao Kurohashi. 2010. Dependency tree-based sentiment classification using CRFs with hidden variables. In Proceedings of ACL:HLT.
  13. 13.Andrew Y Ng and Michael I Jordan. 2002. On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. In Proceedings of NIPS, volume 2, pages 841–848.
  14. 14.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of ACL.
  15. 15.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of ACL.
  16. 16.Jason D. Rennie, Lawrence Shih, Jaime Teevan, and David R. Karger. 2003. Tackling the poor assumptions of naive bayes text classifiers. In Proceedings of ICML, pages 616–623.
  17. 17.Greg Schohn and David Cohn. 2000. Less is more: Active learning with support vector machines. In Proceedings of ICML, pages 839–846.
  18. 18.Richard Socher, Jeffrey Pennington, Eric H. Huang, Andrew Y. Ng, and Christopher D. Manning. 2011. Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions. In Proceedings of EMNLP.
  19. 19.Casey Whitelaw, Navendu Garg, and Shlomo Argamon. 2005. Using appraisal taxonomies for sentiment analysis. In Proceedings of CIKM-05.
  20. 20.Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language Resources and Evaluation, 39(2-3):165–210.

Citation

MLA
Wang, S. I., and C. D. Manning. “Baselines and Bigrams: Simple, Good Sentiment and Topic Classification”. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2012, pp. 90–94, https://aclanthology.org/P12-2018/.
APA
Wang, S. I., & Manning, C. D. (2012). Baselines and Bigrams: Simple, Good Sentiment and Topic Classification. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 90–94. https://aclanthology.org/P12-2018/
Chicago
Wang, S. I., and C. D. Manning. 2012. “Baselines and Bigrams: Simple, Good Sentiment and Topic Classification”. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 90–94. https://aclanthology.org/P12-2018/.
Harvard
Wang, S.I. and Manning, C.D. (2012) “Baselines and Bigrams: Simple, Good Sentiment and Topic Classification”, Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 90–94. Available at: https://aclanthology.org/P12-2018/.
Vancouver
1. Wang SI, Manning CD (2012) Baselines and Bigrams: Simple, Good Sentiment and Topic Classification. In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 90–94

BibTeX

@inproceedings{wang-manning-2012-baselines,
    title = "Baselines and Bigrams: Simple, Good Sentiment and Topic Classification",
    author = "Wang, Sida  and
      Manning, Christopher",
    editor = "Li, Haizhou  and
      Lin, Chin-Yew  and
      Osborne, Miles  and
      Lee, Gary Geunbae  and
      Park, Jong C.",
    booktitle = "Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2012",
    address = "Jeju Island, Korea",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/P12-2018/",
    pages = "90--94"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/