Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Apriori algorithm

The Apriori algorithm is a fundamental data mining algorithm used to discover frequent itemsets and generate association rules in transactional databases. It operates on the principle that any subset of a frequent itemset must also be frequent, a concept known as the Apriori property or downward-closure property of support. The algorithm proceeds iteratively through a breadth-first search, first identifying individual items that meet a user-defined minimum support threshold and then extending them into progressively larger candidate itemsets one level at a time. Candidates containing any infrequent subsets are pruned before scanning the database to verify their support counts. Once all frequent itemsets are determined, the algorithm can derive association rules that satisfy a minimum confidence requirement, making it a foundational tool for market basket analysis, recommendation systems, and pattern discovery.

5 items

Mining Opinion Features in Customer Reviews

Mining Opinion Features in Customer Reviews

Minqing Hu, Bing Liu

OrganizationsUniversity of Illinois Chicago

Why you should read this

Proposes an effective unsupervised framework combining part-of-speech tagging and association rule mining to extract product features from unstructured customer reviews for feature-based opinion summarization.

It is a common practice that merchants selling products on the Web ask their customers to review the products and associated services. As e-commerce is becoming more and more popular, the number of customer reviews that a product receives grows rapidly. For a popular product, the number of reviews can be in hundreds. This makes it difficult for a potential customer to read them in order to make a decision on whether to buy the product. In this project, we aim to summarize all the customer reviews of a product. This summarization task is different from traditional text summarization because we are only interested in the specific features of the product that customers have opinions on and also whether the opinions are positive or negative. We do not summarize the reviews by selecting or rewriting a subset of the original sentences from the reviews to capture their main points as in the classic text summarization. In this paper, we only focus on mining opinion/product features that the reviewers have commented on. A number of techniques are presented to mine such features. Our experimental results show that these techniques are highly effective.

Added

2026-09-25

An effective hash-based algorithm for mining association rules

An effective hash-based algorithm for mining association rules

Jong Soo Park, Ming-Syan Chen, Philip S. Yu

OrganizationsIBMSungshin Women's University

Why you should read this

Introduces the Direct Hashing and Pruning (DHP) algorithm, which substantially speeds up association rule mining by using a hash technique to prune candidate 2-itemsets and progressively reduce transaction database sizes during early iterations.

In this paper, we examine the issue of mining association rules among items in a large database of sales transactions. The mining of association rules can be mapped into the problem of discovering large itemsets where a large itemset is a group of items which appear in a sufficient number of transactions. The problem of discovering large itemsets can be solved by constructing a candidate set of itemsets first and then, identifying, within this candidate set, those itemsets that meet the large itemset requirement. Generally this is done iteratively for each large k-itemset in increasing order of k where a large k-itemset is a large itemset with k items. To determine large itemsets from a huge number of candidate large itemsets in early iterations is usually the dominating factor for the overall data mining performance. To address this issue, we propose an effective hash-based algorithm for the candidate set generation. Explicitly, the number of candidate 2-itemsets generated by the proposed algorithm is, in orders of magnitude, smaller than that by previous methods, thus resolving the performance bottleneck. Note that the generation of smaller candidate sets enables us to effectively trim the transaction database size at a much earlier stage of the iterations, thereby reducing the computational cost for later iterations significantly. Extensive simulation study is conducted to evaluate performance of the proposed algorithm.

Added

2026-09-24