keyword
semantic orientation
Semantic orientation refers to the evaluative character or emotional valence of a word, phrase, or text, indicating whether the expressed sentiment conveys praise, favorability, criticism, or disapproval. This measure encompasses both direction, distinguishing between positive, neutral, and negative viewpoints, and magnitude, reflecting the intensity of the evaluation from mild to strong. In computational linguistics and sentiment analysis, determining semantic orientation is essential for tasks such as opinion mining, customer review classification, and social media monitoring. It is typically assessed by analyzing syntactic patterns, measuring statistical associations with established positive and negative reference words, or applying machine learning models to aggregate lexical sentiment into an overall assessment of subjective opinion.
9 items

Mining Opinion Features in Customer Reviews
Minqing Hu, Bing Liu
Why you should read this
Proposes an effective unsupervised framework combining part-of-speech tagging and association rule mining to extract product features from unstructured customer reviews for feature-based opinion summarization.
It is a common practice that merchants selling products on the Web ask their customers to review the products and associated services. As e-commerce is becoming more and more popular, the number of customer reviews that a product receives grows rapidly. For a popular product, the number of reviews can be in hundreds. This makes it difficult for a potential customer to read them in order to make a decision on whether to buy the product. In this project, we aim to summarize all the customer reviews of a product. This summarization task is different from traditional text summarization because we are only interested in the specific features of the product that customers have opinions on and also whether the opinions are positive or negative. We do not summarize the reviews by selecting or rewriting a subset of the original sentences from the reviews to capture their main points as in the classic text summarization. In this paper, we only focus on mining opinion/product features that the reviewers have commented on. A number of techniques are presented to mine such features. Our experimental results show that these techniques are highly effective.
Added
2026-09-25

Determining the Sentiment of Opinions
Soo-Min Kim, Eduard Hovy
Why you should read this
Develops a topic-oriented sentiment analysis framework that identifies opinion holders and computes sentence-level polarity by combining WordNet-derived word sentiment strengths within targeted syntactic regions.
Identifying sentiments (the affective parts of opinions) is a challenging problem. We present a system that, given a topic, automatically finds the people who hold opinions positive about that topic and the sentiment of each opinion. The system contains a module for determining word sentiment and another for combining sentiments within a sentence. We experiment with various models of classifying and combining sentiment at word and sentence levels, with promising results.
Added
2026-09-24

Measuring praise and criticism: Inference of semantic orientation from association
Peter D. Turney, Michael L. Littman
Why you should read this
Proposes a method for automatically determining the positive or negative semantic orientation of words across diverse parts of speech by measuring their statistical association with paradigm seed words using pointwise mutual information and latent semantic analysis.
The evaluative character of a word is called its semantic orientation. Positive semantic orientation indicates praise (e.g., "honest", "intrepid") and negative semantic orientation indicates criticism (e.g., "disturbing", "superfluous"). Semantic orientation varies in both direction (positive or negative) and degree (mild to strong). An automated system for measuring semantic orientation would have application in text classification, text filtering, tracking opinions in online discussions, analysis of survey responses, and automated chat systems (chatbots). This paper introduces a method for inferring the semantic orientation of a word from its statistical association with a set of positive and negative paradigm words. Two instances of this approach are evaluated, based on two different statistical measures of word association: pointwise mutual information (PMI) and latent semantic analysis (LSA). The method is experimentally tested with 3,596 words (including adjectives, adverbs, nouns, and verbs) that have been manually labeled positive (1,614 words) and negative (1,982 words). The method attains an accuracy of 82.8% on the full test set, but the accuracy rises above 95% when the algorithm is allowed to abstain from classifying mild words.
Added
2026-09-24

Predicting the Semantic Orientation of Adjectives
Vasileios Hatzivassiloglou, Kathleen McKeown
Why you should read this
Proposes a method to automatically determine the positive or negative sentiment of adjectives by leveraging syntactic constraints from conjunctions in a corpus and clustering them into distinct polarity groups.
We identify and validate from a large corpus constraints from conjunctions on the positive or negative semantic orientation of the conjoined adjectives. A log-linear regression model uses these constraints to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently. Combining the constraints across many adjectives, a clustering algorithm separates the adjectives into groups of different orientations, and finally, adjectives are labeled positive or negative. Evaluations on real data and simulation experiments indicate high levels of performance: classification precision is more than 90% for adjectives that occur in a modest number of conjunctions in the corpus.
Added
2026-09-16

Mining the peanut gallery: opinion extraction and semantic classification of product reviews
Kushal Dave, Steve Lawrence, David M. Pennock
Why you should read this
Proposes an opinion mining system that uses information retrieval scoring techniques and variable-length text patterns to classify review sentiment and synthesize unstructured web feedback into product attribute summaries.
The web contains a wealth of product reviews, but sifting through them is a daunting task. Ideally, an opinion mining tool would process a set of search results for a given item, generating a list of product attributes (quality, features, etc.) and aggregating opinions about each of them (poor, mixed, good). We begin by identifying the unique properties of this problem and develop a method for automatically distinguishing between positive and negative reviews. Our classifier draws on information retrieval techniques for feature extraction and scoring, and the results for various metrics and heuristics vary depending on the testing situation. The best methods work as well as or better than traditional machine learning. When operating on individual sentences collected from web searches, performance is limited due to noise and ambiguity. But in the context of a complete web-based tool and aided by a simple method for grouping sentences into attributes, the results are qualitatively quite useful. discussion boards and mailing list archives, as well as in Usenet via Google Groups. Users also comment on products in their personal web sites and blogs, which are then aggregated by sites such as Blogstreet.com, AllConsuming.net, and onfocus.com. When trying to locate information on a product, a general web search turns up several useful sites, but getting an overall sense of these reviews can be daunting or time-consuming. In the movie review domain, sites like Rottentomatoes.com have sprung up to try to impose some order on the void, providing ratings and brief quotes from numerous reviews and generating an aggregate opinion. Such sites even have their own category—“Review Hubs”—on Yahoo! On the commercial side, Internet clipping services like Webclipping.com, eWatch.com, and TracerLock.com watch news sites and discussion areas for mentions of a given company or product, trying to track “buzz.” Print clipping services have been providing competitive intelligence for some time. The ease of publishing on the web led to an explosion in content to be surveyed, but the same technology makes automation much more feasible. This paper describes a tool for sifting through and synthesizing product reviews, automating the sort of work done by aggregation sites or clipping services. We begin by using structured reviews for testing and training, identifying appropriate features and scoring methods from information retrieval for determining whether reviews are positive or negative. These results perform as well as traditional machine learning methods. We then use the classifier to identify and classify review sentences from the web, where classification is more difficult. However, a simple technique for identifying the relevant attributes of a product produces a subjectively useful summary.
Added
2026-09-14

CROWDSOURCING A WORD–EMOTION ASSOCIATION LEXICON
Saif M. Mohammad, Peter D. Turney
Why you should read this
Presents a practical crowdsourcing methodology for constructing large-scale word-emotion lexicons, proving that sense-verification questions and association-based framing substantially increase annotation quality and inter-annotator agreement.
Even though considerable attention has been given to the polarity of words (positive and negative) and the creation of large polarity lexicons, research in emotion analysis has had to rely on limited and small emotion lexicons. In this paper we show how the combined strength and wisdom of the crowds can be used to generate a large, high-quality, word-emotion and word-polarity association lexicon quickly and inexpensively. We enumerate the challenges in emotion annotation in a crowdsourcing scenario and propose solutions to address them. Most notably, in addition to questions about emotions associated with terms, we show how the inclusion of a word choice question can discourage malicious data entry, help identify instances where the annotator may not be familiar with the target term (allowing us to reject such annotations), and help obtain annotations at sense level (rather than at word level). We conducted experiments on how to formulate the emotion-annotation questions, and show that asking if a term is associated with an emotion leads to markedly higher inter-annotator agreement than that obtained by asking if a term evokes an emotion.
Added
2026-09-14

A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts
Bo Pang, Lillian Lee
Why you should read this
Proposes a graph min-cut method to isolate subjective sentences using cross-sentence contextual constraints, significantly improving document-level sentiment classification.
Sentiment analysis seeks to identify the viewpoint(s) underlying a text span; an example application is classifying a movie review as "thumbs up" or "thumbs down". To determine this sentiment polarity, we propose a novel machine-learning method that applies text-categorization techniques to just the subjective portions of the document. Extracting these portions can be implemented using efficient techniques for finding minimum cuts in graphs; this greatly facilitates incorporation of cross-sentence contextual constraints.
Added
2026-09-11

Thumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews
Peter D. Turney
Why you should read this
Proposes an unsupervised sentiment classification method that determines review polarity by measuring the semantic orientation of extracted phrases relative to positive and negative anchor words without requiring labeled training data.
This paper presents a simple unsupervised learning algorithm for classifying reviews as recommended (thumbs up) or not recommended (thumbs down). The classification of a review is predicted by the average semantic orientation of the phrases in the review that contain adjectives or adverbs. A phrase has a positive semantic orientation when it has good associations (e.g., "subtle nuances") and a negative semantic orientation when it has bad associations (e.g., "very cavalier"). In this paper, the semantic orientation of a phrase is calculated as the mutual information between the given phrase and the word "excellent" minus the mutual information between the given phrase and the word "poor". A review is classified as recommended if the average semantic orientation of its phrases is positive. The algorithm achieves an average accuracy of 74% when evaluated on 410 reviews from Epinions, sampled from four different domains (reviews of automobiles, banks, movies, and travel destinations). The accuracy ranges from 84% for automobile reviews to 66% for movie reviews.
Added
2026-09-09

Thumbs up? Sentiment Classification using Machine Learning Techniques
Bo Pang, Lillian Lee, Shivakumar Vaithyanathan
Why you should read this
Establishes machine learning as a viable framework for sentiment classification by demonstrating that standard classifiers outperform human baselines on movie reviews while identifying key challenges that make sentiment harder than topic-based categorization.
We consider the problem of classifying documents not by topic, but by overall sentiment, e.g., determining whether a review is positive or negative. Using movie reviews as data, we find that standard machine learning techniques definitively outperform human-produced baselines. However, the three machine learning methods we employed (Naive Bayes, maximum entropy classification, and support vector machines) do not perform as well on sentiment classification as on traditional topic-based categorization. We conclude by examining factors that make the sentiment classification problem more challenging.
Added
2026-09-07
