A New ANEW: Evaluation of a Word List for Sentiment Analysis in Microblogs
Finn Årup Nielsen
Presents a sentiment lexicon specifically designed for microblogs, demonstrating that adapting affective word lists to informal Twitter text improves valence scoring compared to traditional resources like ANEW.
Analyzing sentiment on microblogging platforms like Twitter has become increasingly vital for understanding public mood, brand perception, and real-time social trends. However, conventional sentiment analysis lexicons—such as the Affective Norms for English Words (ANEW)—were created prior to the widespread adoption of microblogging and lack the informal slang, acronyms, and vulgarities typical of online communication. This raises questions about how accurately traditional word lists measure emotion in short, informal texts.
The article evaluates whether a newly constructed sentiment lexicon tailored specifically for microblogs outperforms established word lists in estimating the sentiment strength of Twitter posts. Specifically, it compares the author's 2,477-word list against ANEW, General Inquirer, OpinionFinder, and the specialized SentiStrength software.
To conduct the assessment, the author scored the new lexicon manually on a scale from minus five to plus five, incorporating internet slang and obscenities. The performance of all approaches was benchmarked against an independent ground-truth dataset of 1,000 tweets, each evaluated ten times by human annotators via Amazon Mechanical Turk. Lexicon performance was primarily measured by calculating the correlation between automated sentiment scores and the human benchmark ratings.
The evaluation revealed several key findings. First, the specialized SentiStrength software achieved the highest accuracy, with a correlation of 0.610. Second, among simple word-matching methods, the author's new lexicon outperformed ANEW, achieving a correlation of 0.564 compared to ANEW's 0.525. Third, polarity-only lexicons such as OpinionFinder and General Inquirer performed significantly worse, posting correlations of only 0.458 and 0.374, respectively. Finally, a direct comparison using only the overlapping words between the new list and ANEW showed that ANEW's ratings were marginally superior (0.52 versus 0.49), indicating that the advantage of the new lexicon stems from its broader microblog-focused vocabulary rather than better valence calibration.
These findings demonstrate that while sophisticated algorithmic tools like SentiStrength provide the highest accuracy, lexicon-based methods can achieve competitive results if they include platform-specific language, such as online acronyms and informal expressions. Furthermore, assigning graded sentiment strength rather than simple positive or negative classifications is essential for informal text analysis. Organizations seeking low-complexity, real-time sentiment tracking can use tailored lexicons as a cost-effective alternative to complex software pipelines without substantial performance loss.
Teams deploying sentiment analysis on social data should prioritize lexicons that incorporate internet slang and support continuous vocabulary expansion. When maximum precision is required, organizations should favor multi-rule tools like SentiStrength that account for spelling variations and grammatical context. Further work should explore whether adding automated negation handling and emoticon detection improves simple word lists without introducing undue computational complexity.
Confidence in these results is moderate to high for basic Twitter text, supported by human consensus ratings across 1,000 posts. However, users should remain cautious because the new word list relied on subjective ratings from a single annotator, and the benchmark dataset was limited to a single 1,000-tweet sample, which may not capture all regional dialects or emerging online jargon.
- Paper: Measuring praise and criticism: Inference of semantic orientation from association, Peter D. Turney et al. (2003). This seminal paper introduces methods for inferring the semantic orientation and sentiment valence of individual words, establishing foundational concepts evaluated in microblog lexicons.
- Paper: Thumbs up? Sentiment Classification using Machine Learning Techniques, Bo Pang et al. (2002). This benchmark work formalizes computational sentiment analysis and baseline word-matching strategies that motivate the need for specialized microblog lexicons.
- Paper: Predicting the Semantic Orientation of Adjectives, Vasileios Hatzivassiloglou et al. (1997). It provides the foundational framework for determining the semantic orientation and polarity of descriptive words used across affective word lists.
- Paper: A holistic lexicon-based approach to opinion mining, Xiaowen Ding et al. (2008). It establishes core lexicon-based sentiment analysis methodologies and word scoring techniques directly compared and evaluated in microblog contexts.
- Paper: Thumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews, Peter D. Turney (2002). This paper demonstrates how unsupervised semantic orientation of words can be aggregated to determine overall text sentiment, a direct precursor to lexicon-based tweet evaluation.
- Paper: What is Twitter, a social network or a news media?, Haewoon Kwak et al. (2010). It details the core linguistic and structural characteristics of Twitter communication that make standard pre-microblogging lexicons like ANEW less effective.
- Paper: Twitter mood predicts the stock market, Johan Bollen et al. (2010). It illustrates early applications and limitations of lexicon-based sentiment tracking over microblog streams like Twitter.
- Paper: CROWDSOURCING A WORD–EMOTION ASSOCIATION LEXICON, Saif M. Mohammad et al. (2013). It advances beyond simple microblog valence lists by using crowdsourcing to construct a comprehensive term-emotion and polarity association lexicon.
- Paper: Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank, Richard Socher et al. (2013). It moves past bag-of-words lexicon matching by modeling syntactic compositionality and phrase-level sentiment over treebanks.
- Paper: Convolutional Neural Networks for Sentence Classification, Yoon Kim (2014). It extends short-text sentiment classification from simple word-list matching to deep convolutional neural networks built on distributed word embeddings.
- Paper: Deep learning for sentiment analysis: A survey, Lei Zhang et al. (2018). It offers a comprehensive survey detailing the progression from rule- and lexicon-based microblog analysis to modern deep neural sentiment architectures.
- Paper: Automated Hate Speech Detection and the Problem of Offensive Language, Thomas Davidson et al. (2017). It applies Twitter lexical and sentiment features to fine-grained downstream classification tasks distinguishing offensive language from hate speech.
