A New ANEW: Evaluation of a Word List for Sentiment Analysis in Microblogs

Finn Årup Nielsen

article2011#MSM1,398 citations

Presents a sentiment lexicon specifically designed for microblogs, demonstrating that adapting affective word lists to informal Twitter text improves valence scoring compared to traditional resources like ANEW.

Listen

Analyzing sentiment on microblogging platforms like Twitter has become increasingly vital for understanding public mood, brand perception, and real-time social trends. However, conventional sentiment analysis lexicons—such as the Affective Norms for English Words (ANEW)—were created prior to the widespread adoption of microblogging and lack the informal slang, acronyms, and vulgarities typical of online communication. This raises questions about how accurately traditional word lists measure emotion in short, informal texts.

The article evaluates whether a newly constructed sentiment lexicon tailored specifically for microblogs outperforms established word lists in estimating the sentiment strength of Twitter posts. Specifically, it compares the author's 2,477-word list against ANEW, General Inquirer, OpinionFinder, and the specialized SentiStrength software.

To conduct the assessment, the author scored the new lexicon manually on a scale from minus five to plus five, incorporating internet slang and obscenities. The performance of all approaches was benchmarked against an independent ground-truth dataset of 1,000 tweets, each evaluated ten times by human annotators via Amazon Mechanical Turk. Lexicon performance was primarily measured by calculating the correlation between automated sentiment scores and the human benchmark ratings.

The evaluation revealed several key findings. First, the specialized SentiStrength software achieved the highest accuracy, with a correlation of 0.610. Second, among simple word-matching methods, the author's new lexicon outperformed ANEW, achieving a correlation of 0.564 compared to ANEW's 0.525. Third, polarity-only lexicons such as OpinionFinder and General Inquirer performed significantly worse, posting correlations of only 0.458 and 0.374, respectively. Finally, a direct comparison using only the overlapping words between the new list and ANEW showed that ANEW's ratings were marginally superior (0.52 versus 0.49), indicating that the advantage of the new lexicon stems from its broader microblog-focused vocabulary rather than better valence calibration.

These findings demonstrate that while sophisticated algorithmic tools like SentiStrength provide the highest accuracy, lexicon-based methods can achieve competitive results if they include platform-specific language, such as online acronyms and informal expressions. Furthermore, assigning graded sentiment strength rather than simple positive or negative classifications is essential for informal text analysis. Organizations seeking low-complexity, real-time sentiment tracking can use tailored lexicons as a cost-effective alternative to complex software pipelines without substantial performance loss.

Teams deploying sentiment analysis on social data should prioritize lexicons that incorporate internet slang and support continuous vocabulary expansion. When maximum precision is required, organizations should favor multi-rule tools like SentiStrength that account for spelling variations and grammatical context. Further work should explore whether adding automated negation handling and emoticon detection improves simple word lists without introducing undue computational complexity.

Confidence in these results is moderate to high for basic Twitter text, supported by human consensus ratings across 1,000 posts. However, users should remain cautious because the new word list relied on subjective ratings from a single annotator, and the benchmark dataset was limited to a single 1,000-tweet sample, which may not capture all regional dialects or emerging online jargon.

Cover for A New ANEW: Evaluation of a Word List for Sentiment Analysis in Microblogs

Abstract

Sentiment analysis of microblogs such as Twitter has recently gained a fair amount of attention. One of the simplest sentiment analysis approaches compares the words of a posting against a labeled word list, where each word has been scored for valence, -- a 'sentiment lexicon' or 'affective word lists'. There exist several affective word lists, e.g., ANEW (Affective Norms for English Words) developed before the advent of microblogging and sentiment analysis. I wanted to examine how well ANEW and other word lists performs for the detection of sentiment strength in microblog posts in comparison with a new word list specifically constructed for microblogs. I used manually labeled postings from Twitter scored for sentiment. Using a simple word matching I show that the new word list may perform better than ANEW, though not as good as the more elaborate approach found in SentiStrength.

Table of Contents

  • 1 Introduction
  • 2 Construction of word list
  • 3 Twitter data
  • 4 Results
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — AFINN Sentiment Lexicon Construction and Scoring Design

    model/method

    The AFINN sentiment lexicon is an affective word list specifically constructed for sentiment strength detection in microblogs and short informal text (such as Twitter). Each word or short phrase is scored manually for valence on an integer scale from −5-5 (very negative) to +5+5 (very positive).

    Lexicon characteristics and design choices include:

    • Vocabulary scale and expansion: An initial version (termed AFINN-96) contains 1,468 words. The expanded version comprises 2,477 unique entries (including 15 phrases).
    • Sources: The lexicon was initiated from sets of obscene words and positive seed words, and expanded using Twitter postings related to the COP15 climate conference, the Original Balanced Affective Word List by Greg Siegle, Steven J. DeRose's The Compass DeRose Guide to Emotion Words, Wiktionary synonym entries, Microsoft Web n-gram context similarity clusters, and Urban Dictionary slang and acronyms (e.g., WTF, LOL, ROFL).
    • Disambiguation and filtering: To prevent part-of-speech ambiguity in simple matching, words with variable polarity or category shifts (e.g., patient, firm, mean, power, frank) were excluded. High-arousal words with context-dependent sentiment (e.g., surprise) were also omitted.
    • Valence distribution: The lexicon exhibits an asymmetric bias toward negative terms, containing 1,598 negative entries (65%), 878 positive entries (35%), and 1 neutral phrase. The modal positive rating is +2+2, the modal negative rating is −2-2, and strong obscenities are assigned −4-4 or −5-5.
  2. Knowl 2 — Comparative Sentiment Strength Correlation on Microblog Posts

    data/table

    Evaluation of sentiment strength detection methods on a benchmark corpus of 1,000 English tweets. Each tweet was independently scored by 10 Amazon Mechanical Turk (AMT) raters on an integer scale from 1 (negative) to 9 (positive), with the mean score defining the ground truth.

    Five methods were compared: AFINN (2,477 words, valence [−5,+5][-5, +5]), ANEW (1,034 words, valence [1,9][1, 9]), General Inquirer (GI, binary polarity mapped to {−1,+1}\{-1, +1\}), OpinionFinder (OF, binary polarity mapped to {−1,+1}\{-1, +1\}), and SentiStrength (SS, rule-based microblog analyzer).

    Method AMT AFINN (My) ANEW GI OF
    AFINN (My) .564
    ANEW .525 .696
    GI .374 .525 .592
    OF .458 .675 .624 .705
    SS .610 .604 .546 .474 .512

    Spearman rank correlations against AMT ratings are: AFINN (ρ=0.596\rho = 0.596), ANEW (ρ=0.544\rho = 0.544), GI (ρ=0.422\rho = 0.422), OF (ρ=0.491\rho = 0.491), and SentiStrength (ρ=0.616\rho = 0.616). In terms of token coverage over the 4,095 unique words identified in the 1,000 tweets, AFINN matched 422 words, ANEW matched 398 words, GI matched 358 words (out of 3,392 non-zero words), and OF matched 562 words (out of 6,442 words).

  3. Knowl 3 — Lexicon-Based Tweet Sentiment Scoring and Aggregation Schemes

    model/method

    Let a tweet T=(w1,w2,…,wN)T = (w_1, w_2, \dots, w_N) be a sequence of NN tokenized words, and let v(w)∈Rv(w) \in \mathbb{R} denote the sentiment valence assigned to word ww by a sentiment lexicon LL (with v(w)=0v(w) = 0 if w∉Lw \notin L).

    The primary sentiment strength score S(T)S(T) computes the mean valence across all tokens in the tweet: S(T)=1N∑i=1Nv(wi)S(T) = \frac{1}{N} \sum_{i=1}^{N} v(w_i)

    Alternative aggregation formulations evaluated on the 1,000-tweet benchmark dataset include:

    • Unnormalized sum: Ssum(T)=∑i=1Nv(wi)S_{\text{sum}}(T) = \sum_{i=1}^N v(w_i), which yielded the highest Pearson correlation for AFINN (r=0.581r = 0.581).
    • Active-token normalization: Normalizing the sum by the count of tokens that have non-zero valence: Nactive=∣{wi:v(wi)≠0}∣N_{\text{active}} = |\{w_i : v(w_i) \neq 0\}|.
    • Extreme valence selection: Selecting the single valence with the largest absolute magnitude: arg⁡max⁡v(wi)∣v(wi)∣\arg\max_{v(w_i)} |v(w_i)|, which yielded the lowest Pearson correlation for AFINN (r=0.543r = 0.543).
    • Ternary quantization: Quantizing the aggregated score to {−1,0,+1}\{-1, 0, +1\}, which yielded r=0.548r = 0.548 for AFINN.

    For ANEW, Pearson correlation remained stable across all four scoring schemes (r∈[0.522,0.526]r \in [0.522, 0.526]). For AFINN, normalizing by total word count NN achieved the highest Spearman rank correlation (ρ=0.596\rho = 0.596).

  4. Knowl 4 — Vocabulary Coverage vs. Valence Scoring in AFINN and ANEW

    empirical result

    To isolate whether AFINN's performance advantage over ANEW on Twitter text (r=0.564r = 0.564 vs. r=0.525r = 0.525) originates from superior valence calibration or broader vocabulary coverage, both lexicons were evaluated strictly on their intersection.

    Direct exact matching between AFINN and ANEW yielded an overlapping vocabulary of 299 words. Two sentiment analyzers were constructed using only these 299 words: one using AFINN's [−5,+5][-5, +5] integer valences and the other using ANEW's [1,9][1, 9] mean valences.

    When evaluated on the 1,000-tweet benchmark dataset:

    • The ANEW valence scorer achieved a Pearson correlation of r=0.52r = 0.52.
    • The AFINN valence scorer achieved a Pearson correlation of r=0.49r = 0.49.

    Because ANEW outperforms AFINN when the lexicon is held constant, the overall superiority of AFINN on microblog data is attributable to its larger domain-appropriate vocabulary (incorporating slang and obscenities) rather than superior per-word valence scoring.

  5. Knowl 5 — Lexicon Correlation and Valence Discrepancies Between AFINN and ANEW

    empirical result

    A direct comparison between the scores in AFINN (scaled [−5,+5][-5, +5]) and the mean valence scores in ANEW (scaled [1,9][1, 9]) on their overlapping words shows substantial overall agreement:

    • Pearson correlation: r=0.91r = 0.91
    • Spearman rank correlation: ρ=0.81\rho = 0.81
    • Kendall tau correlation: τ=0.63\tau = 0.63

    Applying Porter word stemming to both lexicons or WordNet lemmatization to AFINN does not significantly change these correlation values.

    When splitting ANEW at neutral valence 5.0 and AFINN at neutral valence 0, qualitative polarity discrepancies emerge on specific terms: aggressive, mischief, ennui, hard, silly, alert, mischiefs, and noisy. Stemming introduces additional discrepancies, such as between alien and alienation, affection and affected, and profit and profiteer.

  6. Knowl 6 — Sentiment Detection Performance as a Function of Lexicon Size

    empirical result

    Resampling experiments measuring sentiment strength correlation against 1,000 AMT-labeled tweets as a function of AFINN lexicon size (evaluated at 5, 100, 300, 500, 1,000, 1,500, 2,000, and 2,477 words over 50 resamples per size) demonstrate:

    • Pearson correlation rises rapidly from ≈0.0\approx 0.0 at 5 words to ≈0.40\approx 0.40 at 300 words and ≈0.50\approx 0.50 at 1,000 words, before plateauing toward 0.5640.564 at 2,477 words.
    • Spearman rank correlation follows a parallel trajectory, rising from ≈0.35\approx 0.35 at 100 words to 0.5960.596 at 2,477 words.
    • Performance gains exhibit diminishing returns beyond 1,000 words, but continue to show positive slope up to the full 2,477-word list.
  7. Knowl 7 — Limitations of AFINN Relative to Psychometric Norms and Rule-Based Analyzers

    limitation

    The AFINN lexicon and its baseline word-matching implementation have two notable limitations:

    1. Single-annotator scoring: Valences were manually assigned by a single author without multi-rater averaging, inter-annotator agreement metrics, or standard deviations, whereas ANEW provides psycholinguistically validated norms across multiple human subjects. AFINN also excludes affective dimensions other than valence (such as arousal, dominance, and subjectivity/objectivity).
    2. Lack of rule-based syntactic and morphological processing: Unlike specialized microblog sentiment systems such as SentiStrength (which achieved superior performance with Pearson r=0.610r = 0.610 and Spearman ρ=0.616\rho = 0.616), baseline AFINN lookup does not incorporate negation handling (valence flipping), emoticon detection, booster word scaling, or spelling correction/elongation normalization.

Coverage note — None was omitted; all primary contributions, lexicon definitions, experimental comparisons, intersection analyses, and stated limitations are fully covered.

References

  1. 1.Pang, B., Lee, L.: Opinion mining and sentiment analysis. Foundations and Trends in Information Retrieval 2(1-2) (2008) 1–135
  2. 2.Thelwall, M., Buckley, K., Paltoglou, G., Cai, D., Kappas, A.: Sentiment strength detection in short informal text. Journal of the American Society for Information Science and Technology 61(12) (2010) 2544–2558
  3. 3.Bradley, M.M., Lang, P.J.: Affective norms for English words (ANEW): Instruction manual and affective ratings. Technical Report C-1, The Center for Research in Psychophysiology, University of Florida (1999)
  4. 4.Wilson, T., Wiebe, J., Hoffmann, P.: Recognizing contextual polarity in phrase-level sentiment analysis. In: Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA, Association for Computational Linguistics (2005)
  5. 5.Hansen, L.K., Arvidsson, A., Nielsen, F.Å., Colleoni, E., Etter, M.: Good friends, bad news — affect and virality in Twitter. Accepted for The 2011 International Workshop on Social Computing, Network, and Services (SocialComNet 2011) (2011)
  6. 6.Akkaya, C., Conrad, A., Wiebe, J., Mihalcea, R.: Amazon Mechanical Turk for subjectivity word sense disambiguation. In: Proceedings of the NAACL HLT 2010 Workshop on Creating, Speech and Language Data with Amazon’s Mechanical Turk, Association for Computational Linguistics (2010) 195–203
  7. 7.Baudhuin, E.S.: Obscene language and evaluative response: an empirical study. Psychological Reports 32 (1973)
  8. 8.Sapolsky, B.S., Shafer, D.M., Kaye, B.K.: Rating offensive words in three television program contexts. BEA 2008, Research Division (2008)
  9. 9.Bird, S., Klein, E., Loper, E.: Natural Language Processing with Python. O’Reilly, Sebastopol, California (June 2009)
  10. 10.Biever, C.: Twitter mood maps reveal emotional states of America. The New Scientist 207(2771) (July 2010) 14

Citation

MLA
Nielsen, F. Å. “A New ANEW: Evaluation of a Word List for Sentiment Analysis in Microblogs”. Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big Things Come in Small Packages (2011) 93-98, 2011, http://arxiv.org/abs/1103.2903v1.
APA
Nielsen, F. Å. (2011). A new ANEW: Evaluation of a word list for sentiment analysis in microblogs. Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big Things Come in Small Packages (2011) 93-98. http://arxiv.org/abs/1103.2903v1
Chicago
Nielsen, F. Å. 2011. “A New ANEW: Evaluation of a Word List for Sentiment Analysis in Microblogs”. Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big Things Come in Small Packages (2011) 93-98. http://arxiv.org/abs/1103.2903v1.
Harvard
Nielsen, F.Å. (2011) “A new ANEW: Evaluation of a word list for sentiment analysis in microblogs”, Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big things come in small packages (2011) 93-98 [Preprint]. Available at: http://arxiv.org/abs/1103.2903v1.
Vancouver
1. Nielsen FÅ (2011) A new ANEW: Evaluation of a word list for sentiment analysis in microblogs. Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big things come in small packages (2011) 93-98

BibTeX

@article{nielsen2011new,
  title = {A new ANEW: Evaluation of a word list for sentiment analysis in microblogs},
  author = {Nielsen, Finn Årup},
  year = {2011},
  journal = {Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big things come in small packages (2011) 93-98},
  url = {http://arxiv.org/abs/1103.2903v1},
  eprint = {1103.2903}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/