Mining Opinion Features in Customer Reviews

Minqing HuBing Liu

article2004AAAI1,486 citations

Proposes an effective unsupervised framework combining part-of-speech tagging and association rule mining to extract product features from unstructured customer reviews for feature-based opinion summarization.

Listen

Online commerce platforms host vast numbers of customer reviews, often leaving prospective buyers overwhelmed by hundreds of opinions per product. Manually reading these reviews to inform purchasing decisions is impractical, and manufacturers cannot reliably provide comprehensive feature lists due to differing customer vocabulary, unexpected product uses, or omission of product flaws. The article aims to demonstrate an automated, domain-independent method for extracting specific product features that consumers comment on within unstructured customer reviews, serving as the foundational step for structured opinion summarization.

To achieve this, the authors developed a multi-stage approach combining natural language processing and data mining. The system parses customer reviews using part-of-speech tagging to identify nouns and noun phrases. It then applies association rule mining to discover frequently discussed candidate features, refines these candidates using compactness and redundancy pruning, extracts nearby opinion adjectives, and leverages those opinion words to detect less common, infrequent features. The framework was evaluated across a dataset of 500 reviews covering five consumer electronic devices from commercial retail websites against a human-tagged benchmark.

Key findings show that basic association mining alone captures candidate features with an average precision of only 56% at 68% recall, introducing substantial noise. Applying compactness pruning—which checks word proximity—improves precision to 66%. Adding pure support pruning—which removes single-word features subsumed by broader phrases—dramatically elevates precision to 79% while holding recall steady at 67%. Finally, using extracted opinion words to discover infrequent features boosts the overall recall significantly to 80% with an acceptable final precision of 72%. Infrequent features account for 15% to 20% of all discovered features.

These results demonstrate that structured, feature-based summaries can be generated automatically without requiring pre-labeled training corpora or domain-specific templates. For e-commerce operators and manufacturers, this reduces the cost and technical complexity of large-scale customer sentiment analysis. It enables platforms to present structured summaries that highlight both prevalent product attributes and niche feedback, improving the consumer purchasing experience and product performance tracking.

Organizations handling high-volume customer feedback should consider implementing this tiered feature extraction and pruning pipeline. Because the initial feature identification process can capture non-product nouns during the infrequent feature search, the system should retain frequency-based ranking to ensure irrelevant terms remain at the bottom of customer-facing displays. Future work must integrate semantic sentiment orientation to classify opinions as positive or negative, and refine methods to capture implicit features that are not explicitly stated as nouns.

Cover for Mining Opinion Features in Customer Reviews

Abstract

It is a common practice that merchants selling products on the Web ask their customers to review the products and associated services. As e-commerce is becoming more and more popular, the number of customer reviews that a product receives grows rapidly. For a popular product, the number of reviews can be in hundreds. This makes it difficult for a potential customer to read them in order to make a decision on whether to buy the product. In this project, we aim to summarize all the customer reviews of a product. This summarization task is different from traditional text summarization because we are only interested in the specific features of the product that customers have opinions on and also whether the opinions are positive or negative. We do not summarize the reviews by selecting or rewriting a subset of the original sentences from the reviews to capture their main points as in the classic text summarization. In this paper, we only focus on mining opinion/product features that the reviewers have commented on. A number of techniques are presented to mine such features. Our experimental results show that these techniques are highly effective.

Table of Contents

  • Introduction
  • Related Work
  • The Proposed Techniques
  • Part-of-Speech Tagging (POS)
  • Frequent Features Generation
  • Feature Pruning
  • Opinion Words Extraction
  • Infrequent Feature Identification
  • Experiments
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Feature-Based Opinion Summarization Pipeline for Customer Reviews

    model/method

    The feature-based opinion summarization system extracts product features commented on by online customers and summarizes positive and negative opinions associated with each feature. The feature extraction stage operates through a five-step pipeline:

    1. Part-of-Speech (POS) Tagging and Preprocessing: Sentences in raw reviews are parsed to assign POS tags and identify noun phrases/groups. Stopwords are removed, words are stemmed, and fuzzy string matching is applied to unify variant spellings and hyphenations (e.g., mapping "autofocus" to "auto-focus").
    2. Frequent Feature Generation: Association rule mining is applied to transaction representations of sentences containing only nouns and noun phrases to extract candidate frequent features (up to 3 words) appearing with at least 1% support.
    3. Feature Pruning: Compactness pruning eliminates multi-word candidate phrases whose words do not appear close together in sentences, and pure support (pp-support) pruning removes redundant single-word features that occur almost exclusively within longer feature phrases.
    4. Opinion Word Extraction: Adjectives located adjacent to and modifying validated frequent features are extracted to construct a domain opinion lexicon.
    5. Infrequent Feature Identification: Sentences containing no frequent features are scanned for extracted opinion words; the nearest noun or noun phrase modifying an opinion word is extracted as an infrequent candidate feature.
  2. Knowl 2 — Frequent Feature Generation via Association Rule Mining on Noun Transactions

    algorithm

    To identify product features that receive frequent customer commentary ("hot" features), review sentences are transformed into transaction records containing only preprocessed nouns and noun phrases, and candidate features are extracted via frequent itemset mining:

    Input: Set of review sentences SS, minimum support threshold smin⁡=1%s_{\min} = 1\%
    Output: Candidate frequent feature set FcandF_{\text{cand}}
    Initialize transaction set T←∅T \leftarrow \emptyset
    for each sentence s∈Ss \in S:
        Extract all nouns and noun phrases from ss
        Apply stopword deletion, stemming, and fuzzy matching to the extracted nouns
        Add the resulting list of noun terms as a transaction tt to TT
    Run Apriori frequent itemset generation on TT with support threshold smin⁡s_{\min}
    Initialize Fcand←∅F_{\text{cand}} \leftarrow \emptyset
    for each discovered frequent itemset II:
        if ∣I∣≤3|I| \le 3:
            Add II to FcandF_{\text{cand}}
    return FcandF_{\text{cand}}

    Candidate itemsets are restricted to a maximum length of 3 words because product feature expressions in natural language rarely exceed three words.

  3. Knowl 3 — Compact Feature Phrase and Compactness Pruning

    definition

    Because association rule mining treats transactions as unordered bags of words, it can generate candidate multi-word features whose words occur far apart in text and do not form genuine noun phrases. Compactness pruning removes such spurious phrases based on word distance.

    Let ff be a candidate frequent feature phrase containing nn words. Let ss be a sentence containing ff, in which the sequence of words belonging to ff appears as w1,w2,…,wnw_1, w_2, \dots, w_n. The phrase ff is compact in sentence ss if the word distance in ss between any two adjacent words wiw_i and wi+1w_{i+1} in the sequence is at most 3 words.

    A candidate phrase ff is defined as a compact feature phrase across the corpus if it occurs in mm sentences and is compact in at least 2 of those mm sentences. Candidate feature phrases containing two or more words that are compact in fewer than 2 sentences are pruned.

  4. Knowl 4 — Pure Support (p-support) and Redundancy Pruning

    definition

    Redundancy pruning removes uninteresting single-word features that appear primarily as constituents of longer, more descriptive multi-word feature phrases (for example, "life" appearing almost exclusively in "battery life").

    The pure support (p-supportp\text{-support}) of a candidate single-word feature ftrftr is the number of sentences in the review corpus where ftrftr appears as a noun or noun phrase and that do not contain any candidate feature phrase that is a proper superset of ftrftr.

    Let τp\tau_p be the minimum pure support threshold (set to τp=3\tau_p = 3). If a single-word feature ftrftr is a subphrase of at least one multi-word feature phrase and satisfies: p-support(ftr)<τpp\text{-support}(ftr) < \tau_p then ftrftr is pruned as redundant. If p-support(ftr)≥τpp\text{-support}(ftr) \ge \tau_p, ftrftr is retained as an independent feature (for instance, "manual" occurring both independently and within "manual mode").

  5. Knowl 5 — Opinion Word Extraction from Adjacent Frequent Features

    model/method

    Opinion words are evaluative terms used by customers to express positive or negative sentiments toward product features. Observing that opinion expressions appear in close syntactic proximity to the features they modify, opinion words are extracted using the pruned set of frequent features.

    For each sentence in the review database, if the sentence contains a recognized frequent feature, the algorithm identifies adjacent adjectives that modify the noun or noun phrase representing that feature. Each modifying adjective found is added to an opinion word lexicon. Stemming and fuzzy matching are applied to unify morphological variants and misspellings. The resulting lexicon is subsequently used to discover infrequent features and determine opinion orientations.

  6. Knowl 6 — Infrequent Feature Identification via Opinion Word Association

    algorithm

    Infrequent features are product attributes commented on by few reviewers that fail to meet the minimum support threshold of association mining. They are discovered by identifying noun phrases modified by known opinion words in sentences lacking frequent features:

    Input: Review database SS, validated frequent feature set FfreqF_{\text{freq}}, opinion word lexicon OO
    Output: Infrequent feature set FinfreqF_{\text{infreq}}
    Initialize Finfreq←∅F_{\text{infreq}} \leftarrow \emptyset
    for each sentence s∈Ss \in S:
        if ss does not contain any feature from FfreqF_{\text{freq}}:
            if ss contains at least one opinion word o∈Oo \in O:
                Find the nearest noun or noun phrase nn to oo in sentence ss
                Add nn to FinfreqF_{\text{infreq}}
    return FinfreqF_{\text{infreq}}

    Because users frequently describe different product features with the same set of evaluative adjectives (e.g., "easy", "clear"), searching for the nearest noun phrase to an opinion word captures low-frequency features without requiring complex semantic parsing. Non-feature nouns captured by this heuristic have low pp-support and rank at the bottom of the final feature summary.

  7. Knowl 7 — Experimental Setup for Opinion Feature Mining

    experimental setup

    The opinion feature extraction framework was evaluated on customer reviews for five consumer electronics products crawled from Amazon.com and C|net.com: two digital cameras (Digital camera1, Digital camera2), one cellular phone, one MP3 player, and one DVD player. For each product, the first 100 customer reviews were downloaded and cleaned of HTML markup.

    Sentences were tagged and chunked using the NLProcessor linguistic parser. Ground truth product features were manually identified by a human tagger who read all 100 reviews per product to compile comprehensive lists of explicit and implicit features (79 features for Digital camera1, 96 for Digital camera2, 67 for Cellular phone, 57 for MP3 player, and 49 for DVD player; average of 69 features per product). Evaluation measured precision and recall across each progressive step of the extraction pipeline.

  8. Knowl 8 — Step-by-Step Evaluation of Feature Extraction and Pruning

    data/table

    The precision and recall of the feature mining pipeline were evaluated across four progressive stages: initial frequent feature generation via association mining, compactness pruning, pp-support pruning, and infrequent feature identification.

    Product Name Frequent Features Compactness Pruning P-support Pruning Infrequent Features
    Recall Precision Recall Precision Recall Precision Recall Precision
    Digital camera1 0.671 0.552 0.658 0.634 0.658 0.825 0.822 0.747
    Digital camera2 0.594 0.594 0.594 0.679 0.594 0.781 0.792 0.710
    Cellular phone 0.731 0.563 0.716 0.676 0.716 0.828 0.761 0.718
    Mp3 player 0.652 0.573 0.652 0.683 0.652 0.754 0.818 0.692
    DVD player 0.754 0.531 0.754 0.634 0.754 0.765 0.797 0.743
    Average 0.68 0.56 0.67 0.66 0.67 0.79 0.80 0.72

    Association mining alone yields high noise, achieving 0.68 recall but only 0.56 precision on average. Compactness pruning increases average precision to 0.66 while maintaining recall at 0.67. pp-support pruning further raises precision to 0.79 with no loss in recall. Infrequent feature identification substantially improves recall to 0.80 while maintaining an overall precision of 0.72.

  9. Knowl 9 — Precision and Recall Dynamics in Opinion Feature Mining

    empirical result

    Empirical evaluation on five electronic product datasets demonstrates the specific impact of each component of the opinion feature mining system:

    • Pruning stages: Compactness pruning and pure support (pp-support) pruning collectively increase precision by 23 percentage points (from 0.56 to 0.79 on average) while recall remains virtually steady (dropping by only 1 percentage point from 0.68 to 0.67).
    • Infrequent feature identification: Incorporating opinion-word-based infrequent feature extraction increases average recall by 13 percentage points (from 0.67 to 0.80), with infrequent features constituting approximately 15% to 20% of all discovered features. Although average precision decreases from 0.79 to 0.72 due to occasional extraction of irrelevant nouns (e.g., "salesman"), the error is non-disruptive because infrequent features have low pp-support and appear at the bottom of the ranked feature summary.

Coverage note — Deliberately omitted the opinion sentence orientation determination method (using WordNet bootstrapping), as the paper presents it only in brief high-level outline and defers its detailed design and evaluation to a subsequent publication.

References

  1. 1.Agrawal, R. and Srikant, R. 1994. ‘‘Fast algorithm for mining association rules.’’ VLDB’94, 1994.
  2. 2.Barzilay, R., and Elhadad, M. 1997. Using lexical chains for text summarization. ACL Workshop on Intelligent, scalable text summarization.
  3. 3.Boguraev, B., and Kennedy, C. 1997. Salience-based content characterization of text documents. In Proceedings of the ACL Workshop on Intelligent Scalable Text Summarization.
  4. 4.Church, K. and Hanks, P. 1990. Word association norms, mutual information and lexicography. Computational Linguistics, 16(1) : 22-29.
  5. 5.Daille. 1996. Study and Implementation of Combined Techniques for Automatic Extraction of Terminology. The Balancing Act: Combining Symbolic and Statistical Approaches to Language Processing. MIT Press, Cambridge
  6. 6.Dave, K., Lawrence, S., and Pennock, D., 2003. Mining the Peanut Gallery: Opinion Extraction and Semantic Classification of Product Reviews. WWW-2003.
  7. 7.DeJong, G. 1982. An Overview of the FRUMP System. Strategies for Natural Language Parsing. 149-176
  8. 8.Hovy, E., and Lin, C.Y. 1997. Automated Text Summarization in SUMMARIST. ACL Workshop on Intelligent, Scalable Text Summarization
  9. 9.Jacquemin, C., and Bourigault, D. 2001. Term extraction and automatic indexing. In R. Mitkov, editor, Handbook of Computational Linguistics. Oxford University Press.
  10. 10.Jokinen P., and Ukkonen, E. 1991. Two algorithms for approximate string matching in static texts. In A. Tarlecki, (ed.), Mathematical Foundations of Computer Science.
  11. 11.Justeson, J. S., and Katz, S.M. 1995. Technical Terminology: some linguistic properties and an algorithm for identification in text. Natural Language Engineering 1(1):9-27.
  12. 12.Kan, M. and McKeown, K. 1999. Information Extraction and Summarization: Domain Independence through Focus Types. Columbia University Technical Report CUCS-030-99.
  13. 13.Kupiec, J., Pedersen, J., and Chen, F. 1995. A Trainable Document Summarizer. SIGIR-1995.
  14. 14.Liu, B., Hsu, W., Ma, Y. 1998. Integrating Classification and Association Rule Mining. KDD-98, 1998.
  15. 15.Mani, I., and Bloedorn, E., 1997. Multi-document Summarization by Graph Search and Matching. AAAI-97.
  16. 16.Manning, C. and Schütze, H. 1999. Foundations of Statistical Natural Language Processing, MIT Press. May 1999.
  17. 17.Marcu, D. 1997. From Discourse Structures to Text Summaries. ACL Workshop on Intelligent, Scalable Text Summarization.
  18. 18.Miller, G., Beckwith, R, Fellbaum, C., Gross, D., and Miller, K. 1990. Introduction to WordNet: An on-line lexical database. International Journal of Lexicography, 3(4):235-312.
  19. 19.NLProcessor – Text Analysis Toolkit. 2000. http://www.infogistics.com/textanalysis.html
  20. 20.Paice, C. D. 1990. Constructing Literature Abstracts by Computer: techniques and prospects. Information Processing and Management 26:171-186.
  21. 21.Radev, D. and McKeown, K. 1998. Generating natural language summaries from multiple on-line sources. Computational Linguistics, 24(3):469-500, September 1998.
  22. 22.Sparck-Jones, K. 1993a. Discourse Modeling for Automatic Text Summarizing. Technical Report 290, University of Cambridge.
  23. 23.Sparck-Jones, K. 1993b. What might be in a summary? Information Retrieval 93: 9-26.
  24. 24.Tait, J. 1983. Automatic Summarizing of English Texts. Ph.D. Dissertation, University of Cambridge.

Citation

MLA
Hu, M., and B. Liu. “Mining Opinion Features in Customer Reviews”. National Conference on Artificial Intelligence, 2004, pp. 755–60, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.329.8214.
APA
Hu, M., & Liu, B. (2004). Mining opinion features in customer reviews. National Conference on Artificial Intelligence, 755–760. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.329.8214
Chicago
Hu, M., and B. Liu. 2004. “Mining Opinion Features in Customer Reviews”. National Conference on Artificial Intelligence, 755–60. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.329.8214.
Harvard
Hu, M. and Liu, B. (2004) “Mining opinion features in customer reviews”, National Conference on Artificial Intelligence, pp. 755–760. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.329.8214.
Vancouver
1. Hu M, Liu B (2004) Mining opinion features in customer reviews. National Conference on Artificial Intelligence 755–760

BibTeX

@article{hu2004mining,
  title = {Mining opinion features in customer reviews},
  author = {Hu, Minqing and Liu, Bing},
  year = {2004},
  journal = {National Conference on Artificial Intelligence},
  pages = {755-760},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.329.8214}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF