Opinion spam and analysis

Nitin JindalB. Liu

article2008WSDM1,675 citationsBest Paper Award

Establishes the foundational framework for detecting fake online reviews by formalizing opinion spam categories and exploiting duplicate content to train machine learning models across millions of Amazon transactions.

Listen

Online consumer reviews strongly influence purchasing choices and provide critical market intelligence for manufacturers. However, because review platforms lack quality control and financial incentives reward product promotion or defamation, deceptive and low-quality reviews have become widespread. This threatens consumer trust and degrades automated sentiment analysis systems. The article evaluates the prevalence of opinion spam and demonstrates methods to detect it using statistical and machine learning models.

To study this issue, the article analyzed a dataset crawled from Amazon.com in June 2006, covering 5.8 million reviews, 2.14 million reviewers, and 6.7 million products across four main categories. The analysis categorized spam into three types: untruthful fake reviews (Type 1), reviews discussing only brands or sellers rather than the specific product (Type 2), and non-reviews like advertisements or irrelevant text (Type 3). The authors applied supervised logistic regression to manually labeled samples to detect Types 2 and 3. Because humans cannot reliably label deceptive Type 1 reviews by reading them alone, the article identified duplicate and near-duplicate reviews as positive spam examples to train predictive models on manufactured products, assessing model accuracy through classification metrics and lift analysis.

The investigation revealed several key findings regarding review behavior and spam detection. First, logistic regression effectively identifies brand-only and non-review spam with a 98.7% classification accuracy score when combining text, reviewer, and product features. Second, duplicate and near-duplicate reviews are widespread, and using them as training proxies allows models to identify non-duplicate Type 1 spam with a 78% accuracy score. Third, spammers heavily target malicious negative reviews: reviewers who repeatedly post negative rating deviations across a brand show exceptionally high spam likelihood, catching 44% of such outliers in the top 10% of ranked cases. Fourth, products with only a single review are substantially more likely to be spam than first or second reviews on multi-review products, with the model capturing 52% of single-review cases in the top 10% pool. Finally, reviews from top-ranked reviewers are disproportionately likely to exhibit spam behaviors, while reader helpfulness feedback fails to separate spam from legitimate content.

These findings indicate that conventional platform trust metrics—specifically reviewer rankings and helpfulness voting—fail to protect consumers and may actively elevate deceptive content. User feedback can be manipulated or misled by well-crafted fake reviews, meaning systems relying purely on sentiment or text mining risk drawing false business conclusions. Furthermore, spam activities disproportionately affect low-selling items and products with only one review, distorting early market traction.

Platform operators should integrate multi-attribute machine learning models that combine reviewer behavior, rating deviations, and duplicate tracking rather than relying on manual flagging or helpfulness votes. Instead of making hard binary spam classifications, systems can use the model's calculated spam probabilities to down-weight suspicious reviews in aggregate scores and search ranks. Product teams must also re-evaluate reviewer leaderboard algorithms to prevent high-volume, biased accounts from gaining artificial prominence.

The findings are subject to certain limitations. While the overall dataset is massive, the detailed Type 1 predictive modeling was conducted primarily on manufactured products (228,422 reviews), and duplicate reviews served as a proxy for untruthful content rather than direct ground truth. Nevertheless, the underlying statistical patterns strongly indicate high confidence in the structural detection framework, offering a solid baseline for expanding spam defenses across other product categories, forums, and blogs.

Cover for Opinion spam and analysis

Abstract

Evaluative texts on the Web have become a valuable source of opinions on products, services, events, individuals, etc. Recently, many researchers have studied such opinion sources as product reviews, forum posts, and blogs. However, existing research has been focused on classification and summarization of opinions using natural language processing and data mining techniques. An important issue that has been neglected so far is opinion spam or trustworthiness of online opinions. In this paper, we study this issue in the context of product reviews, which are opinion rich and are widely used by consumers and product manufacturers. In the past two years, several startup companies also appeared which aggregate opinions from product reviews. It is thus high time to study spam in reviews. To the best of our knowledge, there is still no published study on this topic, although Web spam and email spam have been investigated extensively. We will see that opinion spam is quite different from Web spam and email spam, and thus requires different detection techniques. Based on the analysis of 5.8 million reviews and 2.14 million reviewers from amazon.com, we show that opinion spam in reviews is widespread. This paper analyzes such spam activities and presents some novel techniques to detect them.

Table of Contents

  • Categories and Subject Descriptors
  • General Terms
  • Keywords
  • 1. INTRODUCTION
  • 2. RELATED WORK
  • 3. OPINION DATA AND ANALYSIS
  • 3.1 Review Data from Amazon.com
  • 3.2 Reviews, Reviewers and Products
  • 3.3 Review Ratings and Feedbacks
  • 4. SPAM DETECTION
  • 4.1 Detection of Duplicate Reviews
  • 4.2 Detecting Type 2 & Type 3 Spam Reviews
  • 4.2.1 Model Building Using Logistic Regression
  • 4.2.2 Feature Identification and Construction
  • 4.2.2.1 Review Centric Features
  • 4.2.2.2 Reviewer Centric Features
  • 4.2.2.3 Product Centric Features
  • 4.2.3 Results of Type 1 and Type 2 Spam Detection
  • 5. ANALYSIS OF TYPE 1 SPAM REVIEWS
  • 5.1 Making Use of Duplicates
  • 5.1.1 Model Building Using Duplicates
  • 5.1.2 Predicting Outlier Reviews
  • 5.2 Some Other Interesting Reviews
  • 5.2.1 Only Reviews
  • 5.2.2 Reviews from Top-Ranked Reviewers
  • 5.2.3 Reviews with Different Levels of Feedbacks
  • 5.2.4 Reviews of Products with Varied Sales Ranks
  • 6. CONCLUSIONS
  • 7. ACKNOWLEDGEMENTS
  • 8. REFERENCES

Knowls

  1. Knowl 1 — Taxonomy of Product Review Spam

    definition

    Product review spam is categorized into three distinct types based on review content, object targeting, and malicious intent:

    • Type 1 (Untruthful opinions / Fake reviews): Reviews written to deliberately deceive readers or automated opinion mining systems. They comprise hyper spam (unmerited positive reviews that promote target products, typically written by product owners or interested parties) and defaming spam (unjust or malicious negative reviews aimed at damaging the reputation of competitor products). These reviews cannot be reliably identified by human reading alone.
    • Type 2 (Reviews on brands only): Reviews that do not evaluate the specific product features or performance, but comment exclusively on the manufacturer, brand, or seller. Although they may contain opinions, they are considered spam because they do not evaluate the specific product item and are often biased.
    • Type 3 (Non-reviews): Submissions that contain no product evaluation, divided into:
      1. Advertisements: promotional text or external commercial links.
      2. Irrelevant text: questions, answers, conversational exchanges, or random text without product opinions.
  2. Knowl 2 — Multi-Centric Feature Set for Review Spam Detection

    model/method

    Review spam classification employs a 36-feature representation categorized across review-centric, reviewer-centric, and product-centric dimensions:

    • Review-Centric Features (F1F_1 to F21F_{21}):
      • Feedback metrics: Total feedback count (F1F_1), helpful feedback count (F2F_2), and percentage of positive feedback (F3F_3).
      • Length: Word count of the review title (F4F_4) and review body (F5F_5).
      • Positioning: Date-sorted position in ascending (F6F_6) and descending (F7F_7) order; binary indicators for being the first review (F8F_8) or the sole review (F9F_9) on a product.
      • Textual features: Proportions of positive (F10F_{10}) and negative (F11F_{11}) opinion words, cosine similarity between review text and product feature descriptions (F12F_{12}), brand name mention percentage (F13F_{13}), and proportions of numerals (F14F_{14}), capitalized words (F15F_{15}), and all-caps words (F16F_{16}).
      • Rating features: Assigned star rating (F17∈{1,2,3,4,5}F_{17} \in \{1, 2, 3, 4, 5\}), absolute rating deviation from the product mean (F18=∣r−μproduct∣F_{18} = |r - \mu_{\text{product}}|), rating class (F19∈{Good (r≥4),Bad (r≤2.5),Average (2.5<r<4)}F_{19} \in \{\text{Good } (r \ge 4), \text{Bad } (r \le 2.5), \text{Average } (2.5 < r < 4)\}), and sequence inversion indicators (bad review immediately following a first good review (F20F_{20}) and vice versa (F21F_{21})).
    • Reviewer-Centric Features (F22F_{22} to F32F_{32}):
      • Early review ratio: Proportion of reviews written that were the first review of a product (F22F_{22}) and proportion where the reviewer was the sole reviewer (F23F_{23}).
      • Rating statistics: Mean rating assigned by the reviewer (F24F_{24}), standard deviation of reviewer ratings (F25F_{25}), and an indicator if the reviewer strictly assigns one rating class (F26F_{26}).
      • Rating variance flags: Binary indicators for posting good and bad ratings (F27F_{27}), good and average ratings (F28F_{28}), bad and average ratings (F29F_{29}), or all three rating types (F30F_{30}).
      • Damage control frequency: Percentage of reviews authored matching feature F20F_{20} (F31F_{31}) and feature F21F_{21} (F32F_{32}).
    • Product-Centric Features (F33F_{33} to F36F_{36}): Product price (F33F_{33}), sales rank (F34F_{34}), average star rating across all reviews (F35F_{35}), and standard deviation of star ratings (F36F_{36}).
  3. Knowl 3 — Duplicate Review Detection and Pseudo-Supervised Learning Strategy for Type 1 Spam

    model/method

    Because ground-truth labels for untruthful opinion spam (Type 1) cannot be acquired reliably via human inspection, duplicate and near-duplicate reviews serve as pseudo-positive training examples for supervised learning.

    Duplicate reviews are identified using 2-gram Jaccard similarity with a similarity threshold of at least 0.900.90:

    J(R1,R2)=∣N(R1)∩N(R2)∣∣N(R1)∪N(R2)∣≥0.90J(R_1, R_2) = \frac{|N(R_1) \cap N(R_2)|}{|N(R_1) \cup N(R_2)|} \ge 0.90

    where N(Ri)N(R_i) denotes the set of word 2-grams in review RiR_i.

    Duplicate review pairs are categorized into four types:

    1. Duplicates from different user IDs on the same product.
    2. Duplicates from the same user ID on different products.
    3. Duplicates from different user IDs on different products.
    4. Duplicates from the same user ID on the same product.

    Reviews under categories 1, 2, and 3 are identified as spam because legitimate users do not copy reviews across different products or accounts. Category 4 is excluded from spam labels because it predominantly arises from re-click submissions or error correction.

    In the pseudo-supervised training framework, all reviews belonging to duplicate categories 1–3 are labeled as positive spam examples, while non-duplicate reviews serve as negative examples. A logistic regression model trained on this dataset using multi-centric features is then used to predict the spam probability of non-duplicate Type 1 spam reviews.

  4. Knowl 4 — Classification Performance on Type 2 and Type 3 Review Spam

    data/table

    Supervised logistic regression models evaluated via 10-fold cross-validation on a manually labeled dataset of 470 spam reviews (221 Type 2 brand-only reviews and 249 Type 3 non-reviews) against non-spam reviews yield the following Area Under the ROC Curve (AUC) values:

    Spam Type Number of Spam Reviews AUC (All Features) AUC (Text Features Only) AUC (Without Feedbacks)
    Types 2 3 470 98.7% 90.0% 98.0%
    Type 2 only 221 98.5% 88.0% 98.0%
    Type 3 only 249 99.0% 92.0% 98.0%

    The classifier achieves near-perfect discrimination (>98%>98\% AUC) across both spam types. Restricting the feature set to textual features causes AUC to drop by 7.0 to 10.5 percentage points, demonstrating the importance of metadata features. In contrast, removing helpfulness feedback features (F1,F2,F3F_1, F_2, F_3) yields an AUC of 98.0%98.0\%, indicating that the model does not depend on user feedback to detect Type 2 and Type 3 spam.

  5. Knowl 5 — Classification Performance of Duplicate-Trained Spam Models

    data/table

    Logistic regression models trained on manufactured product (mProducts) reviews using 4,488 duplicate spam reviews as positive instances and 218,514 non-duplicate reviews as negative instances achieve the following 10-fold cross-validation Area Under the ROC Curve (AUC) scores:

    Features Used AUC
    All features 78.0%
    Without feedback features 77.0%
    Only review features 75.0%
    Only reviewer features 72.5%
    Only text features 63.0%

    Combining all multi-centric features yields the highest discrimination (78.0% AUC). Using textual features alone yields an AUC of only 63.0%, indicating that review text alone is insufficient to identify spam patterns and that reviewer behavioral metadata and product context are necessary. Removing feedback features results in a minor AUC decrease of 1.0 percentage point (to 77.0%), demonstrating that the model remains robust without feedback data.

  6. Knowl 6 — Detection Lift on Rating Outliers and Brand-Level Bias

    empirical result

    Evaluating the duplicate-trained logistic regression classifier on non-duplicate reviews using lift curves demonstrates how review spam correlates with product rating deviations (∣r−μproduct∣>1.0|r - \mu_{\text{product}}| > 1.0 on products with at least 5 reviews):

    • Negative deviation on the same brand: For reviewers who write multiple reviews for a brand where all reviews deviate negatively from the product mean by more than −1.0-1.0, the classifier captures 44% of these outlier reviews in the top 10% probability decile and 75% in the top 20% probability decile (compared to a 10% and 20% random baseline). This indicates that persistent negative rating bias against a brand is a strong indicator of malicious defaming spam.
    • Positive deviation on the same brand: The lift curve for persistent positive rating deviations on a single brand is significantly lower than that for negative deviations.
    • General positive deviations: For general positive rating deviations (r−μproduct>+1.0r - \mu_{\text{product}} > +1.0), the lift curve falls below the random baseline, indicating that generic positive rating outliers are less likely to be classified as spam. The only setting where positive deviations exceed the baseline is for bad products (mean rating ≤2.5\le 2.5) receiving average ratings (2.5<r<4.02.5 < r < 4.0), reflecting promotional spam that avoids extreme 5-star ratings to evade outlier detection.
  7. Knowl 7 — Spam Susceptibility of Single-Review Products

    empirical result

    In manufactured products, 46% of reviewed products have exactly one review. Applying the duplicate-trained logistic regression model (with review position and review count features excluded to avoid target leakage) to non-duplicate test reviews shows that sole reviews ("only reviews", N=7,330N = 7,330) exhibit a much higher spam probability than the first (N=9,855N = 9,855) or second (N=9,834N = 9,834) reviews of multi-review products:

    • The top 10% probability decile contains 52% of all "only reviews".
    • The top 20% probability decile captures 77% of all "only reviews".

    In contrast, the first and second reviews of products with multiple reviews demonstrate substantially lower lift curves. This indicates that unreviewed products are frequent targets for isolated, unverified spam submissions.

  8. Knowl 8 — High Spam Propensity Among Top-Ranked Reviewers

    empirical result

    Applying the duplicate-trained logistic regression model across Amazon reviewer ranking tiers (assigned by Amazon based on cumulative helpful feedbacks) indicates that top-ranked reviewers have a higher probability of writing spam reviews than low-ranked reviewers:

    • For Top 100 ranked reviewers (N=609N = 609 reviews), the top 10% spam probability decile captures 69% of their reviews, and the top 20% decile captures 87%.
    • Reviewers ranked between 100 and 1,000 (N=1,413N = 1,413 reviews) exhibit the second highest lift curve.
    • Reviewers ranked below 2,000,000 (N=1,771N = 1,771 reviews) exhibit the lowest spam lift, lying closest to the random baseline.

    Top-ranked reviewers score higher on spam indicators because they author unusually large volumes of reviews (hundreds or thousands across categories), deviate frequently from average product ratings, and write a disproportionately high number of sole reviews for products.

  9. Knowl 9 — Ineffectiveness of Review Helpfulness Feedbacks as Spam Indicators

    empirical result

    Evaluating reviews categorized by helpfulness feedback against the duplicate-trained spam prediction model demonstrates that user feedback is uninformative for spam filtering:

    • The lift curve for reviews with 100% positive feedback (N=19,237N = 19,237, minimum 5 total feedbacks) and the lift curve for reviews with 0% positive feedback (N=638N = 638, minimum 5 total feedbacks) both closely follow the 45-degree random distribution baseline.
    • Untruthful spam reviews and non-spam reviews receive positive feedback votes at similar rates, demonstrating that readers are often misled by well-crafted spam reviews.
    • Excluding feedback features (F1,F2,F3F_1, F_2, F_3) from model training causes minimal change in detection performance (AUC drops from 98.7% to 98.0% for Types 2 and 3 spam, and from 78.0% to 77.0% for duplicate spam), providing robustness against feedback manipulation and click fraud.
  10. Knowl 10 — Inverse Relationship Between Product Sales Rank and Spam Volume

    empirical result

    Evaluating the lift of the duplicate-trained spam model across products grouped by Amazon sales rank (with price and sales rank features excluded during model training to avoid feature leakage) indicates that review spam is concentrated on low-selling products:

    • High sales volume / low sales rank (Sales Rank ≤50\le 50, N=3,009N = 3,009): The lift curve lies strictly below the random baseline, showing minimal spam activity.
    • Low sales volume / high sales rank (Sales Rank ≥100,000\ge 100,000, N=3,751N = 3,751): The lift curve rises significantly above the baseline, capturing over 40% of reviews within the top 20% probability decile.

    This indicates that spammers concentrate their efforts on niche or low-selling products where individual fake reviews exert a substantial relative influence on aggregate consumer perception, whereas popular, top-selling products are difficult to manipulate with isolated spam reviews.

Coverage note — Standard descriptive power-law dataset statistics from Section 3 were omitted as they describe empirical corpus properties rather than detection methodologies or findings.

References

  1. 1.E. Amitay, D. Carmel, A. Darlow, R. Lempel & A. Soffer. The connectivity sonar: detecting site functionality by structural patterns. Hypertext’03, 2003.
  2. 2.M. Andreolini, A. Bulgarelli, M. Colajanni & F. Mazzoni. Honeyspam: Honeypots fighting spam at the source. In Proc. USENIX SRUTI 2005, Cambridge, MA, July 2005.
  3. 3.R. Baeza-Yates, C. Castillo & V. Lopez. PageRank increase under different collusion topologies. AIRWeb’05, 2005.
  4. 4.A. Z. Broder. On the resemblance and containment of documents. In Proceedings of Compression and Complexity of Sequences 1997, IEEE Computer Society, 1997.
  5. 5.C. Castillo, D. Donato, L. Becchetti, P. Boldi, S. Leonardi, M. Santini, S. Vigna. A reference collection for web spam, SIGIR Forum’06, 2006.
  6. 6.S. Chakrabarti. Mining the Web: discovering knowledge from hypertext data. Morgan Kaufmann, 2003.
  7. 7.K. Dave, S. Lawrence & D. Pennock. Mining the peanut gallery: opinion extraction and semantic classification of product reviews. WWW’2003.
  8. 8.I. Fette, N. Sadeh-Koniecpol, A. Tomasic. Learning to Detect Phishing Emails. WWW2007.
  9. 9.D. Fetterly, M. Manasse & M. Najork. Detecting phrase-level duplication on the World Wide Web. SIGIR’2005.
  10. 10.Z. Gyongyi & H. Garcia-Molina. Web Spam Taxonomy. Technical Report, Stanford University, 2004.
  11. 11.M. R. Henzinger: Finding near-duplicate web pages: a large-scale evaluation of algorithms. SIGIR’06, 2006.
  12. 12.M. Hu & B. Liu. Mining and summarizing customer reviews. KDD’2004.
  13. 13.N. Jindal and B. Liu. Product Review Analysis. Technical Report, UIC, 2007.
  14. 14.N. Jindal and B. Liu. Analyzing and Detecting Review Spam. ICDM2007.
  15. 15.W. Li, N. Zhong, C. Liu. Combining Multiple Email Filters Based on Multivariate Statistical Analysis. ISMIS 2006.
  16. 16.B. Liu. Web Data Mining: Exploring hyperlinks, contents and usage data. Springer, 2007.
  17. 17.A. Metwally, D. Agrawal, A. Abbadi. DETECTIVES: DETEcting Coalition hiT Inflation attacks in adVertising nEtworks Streams. WWW2007.
  18. 18.B. Mobasher, R. Burke & J. J Sandvig. Model-based collaborative filtering as a defense against profile injection attacks. AAAI'2006.
  19. 19.A. Ntoulas, M. Najork, M. Manasse & D. Fetterly. Detecting Spam Web Pages through Content Analysis. WWW’2006.
  20. 20.B. Pang, L. Lee & S. Vaithyanathan. Thumbs up? Sentiment classification using machine learning techniques. EMNLP’2002.
  21. 21.A-M. Popescu and O. Etzioni. Extracting Product Features and Opinions from Reviews. EMNLP’2005.
  22. 22.M. Sahami and S. Dumais and D. Heckerman and E. Horvitz. A Bayesian Approach to Filtering Junk {E}-Mail. AAAI Technical Report WS-98-05, 1998.
  23. 23.P. Turney. Thumbs up or thumbs down? semantic orientation applied to unsupervised classification of reviews. ACL’2002.
  24. 24.Y. Wang, M. Ma, Y. Niu, H. Chen. Spam Double-Funnel: Connecting Web Spammers with Advertisers. WWW2007.
  25. 25.B. Wu and B. D. Davison. Identifying link farm spam pages. WWW’06, 2006.
  26. 26.B. Wu, V. Goel & B. D. Davison. Topical TrustRank: using topicality to combat Web spam. WWW'2006.
  27. 27.S. Ye, R. Song, J.-R. Wen, W.-Y. Ma. A Query-dependent duplicate detection approach for large scale search engines. APWeb’04, 2004.
  28. 28.Z. Zhang & B. Varadarajan, Utility scoring of product reviews, CIKM‘2006.

Citation

MLA
Jindal, N., and B. Liu. “Opinion Spam and Analysis”. Proceedings of the International Conference on Web Search and Web Data Mining - WSDM '08, 2008, p. 219, https://doi.org/10.1145/1341531.1341560.
APA
Jindal, N., & Liu, B. (2008). Opinion spam and analysis. Proceedings of the International Conference on Web Search and Web Data Mining - WSDM '08, 219. https://doi.org/10.1145/1341531.1341560
Chicago
Jindal, N., and B. Liu. 2008. “Opinion Spam and Analysis”. Proceedings of the International Conference on Web Search and Web Data Mining - WSDM '08, 219. https://doi.org/10.1145/1341531.1341560.
Harvard
Jindal, N. and Liu, B. (2008) “Opinion spam and analysis”, Proceedings of the international conference on Web search and web data mining - WSDM '08. ACM Press, p. 219. Available at: https://doi.org/10.1145/1341531.1341560.
Vancouver
1. Jindal N, Liu B (2008) Opinion spam and analysis. In: Proceedings of the international conference on Web search and web data mining - WSDM '08. ACM Press, p 219

BibTeX

@inproceedings{Jindal_2008, series={WSDM ’08}, title={Opinion spam and analysis}, url={http://dx.doi.org/10.1145/1341531.1341560}, DOI={10.1145/1341531.1341560}, booktitle={Proceedings of the international conference on Web search and web data mining  - WSDM ’08}, publisher={ACM Press}, author={Jindal, Nitin and Liu, Bing}, year={2008}, pages={219}, collection={WSDM ’08} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF