keyword
Apriori algorithm
The Apriori algorithm is a fundamental data mining algorithm used to discover frequent itemsets and generate association rules in transactional databases. It operates on the principle that any subset of a frequent itemset must also be frequent, a concept known as the Apriori property or downward-closure property of support. The algorithm proceeds iteratively through a breadth-first search, first identifying individual items that meet a user-defined minimum support threshold and then extending them into progressively larger candidate itemsets one level at a time. Candidates containing any infrequent subsets are pruned before scanning the database to verify their support counts. Once all frequent itemsets are determined, the algorithm can derive association rules that satisfy a minimum confidence requirement, making it a foundational tool for market basket analysis, recommendation systems, and pattern discovery.
5 items

Mining Opinion Features in Customer Reviews
Minqing Hu, Bing Liu
Why you should read this
Proposes an effective unsupervised framework combining part-of-speech tagging and association rule mining to extract product features from unstructured customer reviews for feature-based opinion summarization.
It is a common practice that merchants selling products on the Web ask their customers to review the products and associated services. As e-commerce is becoming more and more popular, the number of customer reviews that a product receives grows rapidly. For a popular product, the number of reviews can be in hundreds. This makes it difficult for a potential customer to read them in order to make a decision on whether to buy the product. In this project, we aim to summarize all the customer reviews of a product. This summarization task is different from traditional text summarization because we are only interested in the specific features of the product that customers have opinions on and also whether the opinions are positive or negative. We do not summarize the reviews by selecting or rewriting a subset of the original sentences from the reviews to capture their main points as in the classic text summarization. In this paper, we only focus on mining opinion/product features that the reviewers have commented on. A number of techniques are presented to mine such features. Our experimental results show that these techniques are highly effective.
Added
2026-09-25

Approximate Frequency Counts over Data Streams
G. Manku, R. Motwani
Why you should read this
Presents space-efficient, single-pass streaming algorithms that identify frequent items and itemsets with guaranteed bounds on approximation error and zero false negatives.
We present algorithms for computing frequency counts exceeding a user-specified threshold over data streams. Our algorithms are simple and have provably small memory footprints. Although the output is approximate, the error is guaranteed not to exceed a user-specified parameter. Our algorithms can easily be deployed for streams of singleton items like those found in IP network monitoring. We can also handle streams of variable sized sets of items exemplified by a sequence of market basket transactions at a retail store. For such streams, we describe an optimized implementation to compute frequent itemsets in a single pass.
Added
2026-09-24

An effective hash-based algorithm for mining association rules
Jong Soo Park, Ming-Syan Chen, Philip S. Yu
Why you should read this
Introduces the Direct Hashing and Pruning (DHP) algorithm, which substantially speeds up association rule mining by using a hash technique to prune candidate 2-itemsets and progressively reduce transaction database sizes during early iterations.
In this paper, we examine the issue of mining association rules among items in a large database of sales transactions. The mining of association rules can be mapped into the problem of discovering large itemsets where a large itemset is a group of items which appear in a sufficient number of transactions. The problem of discovering large itemsets can be solved by constructing a candidate set of itemsets first and then, identifying, within this candidate set, those itemsets that meet the large itemset requirement. Generally this is done iteratively for each large k-itemset in increasing order of k where a large k-itemset is a large itemset with k items. To determine large itemsets from a huge number of candidate large itemsets in early iterations is usually the dominating factor for the overall data mining performance. To address this issue, we propose an effective hash-based algorithm for the candidate set generation. Explicitly, the number of candidate 2-itemsets generated by the proposed algorithm is, in orders of magnitude, smaller than that by previous methods, thus resolving the performance bottleneck. Note that the generation of smaller candidate sets enables us to effectively trim the transaction database size at a much earlier stage of the iterations, thereby reducing the computational cost for later iterations significantly. Extensive simulation study is conducted to evaluate performance of the proposed algorithm.
Added
2026-09-24

An Efficient Algorithm for Mining Association Rules in Large Databases
Ashoka Savasere, Edward Omiecinski, Shamkant B. Navathe
Why you should read this
Proposes the Partition algorithm for association rule mining, which guarantees finding all large itemsets in at most two database scans to dramatically cut I/O and computation costs on massive transaction datasets.
Mining for association rules between items in a large database of sales transactions has been described as an important database mining problem. In this paper we present an efficient algorithm for mining association rules that is fundamentally different from known algorithms. Compared to previous algorithms, our algorithm not only reduces the I/O overhead significantly but also has lower CPU overhead for most cases. We have performed extensive experiments and compared the performance of our algorithm with one of the best existing algorithms. It was found that for large databases, the CPU overhead was reduced by as much as a factor of four and I/O was reduced by almost an order of magnitude. Hence this algorithm is especially suitable for very large size databases.
Added
2026-09-17

Integrating Classification and Association Rule Mining
Bing Liu, Wynne Hsu, Yiming Ma
Why you should read this
Introduces the CBA algorithm to integrate classification and association rule mining by discovering class association rules, building classifiers that outperform C4.5 while capturing interpretable patterns standard decision trees overlook.
Classification rule mining aims to discover a small set of rules in the database that forms an accurate classifier. Association rule mining finds all the rules existing in the database that satisfy some minimum support and minimum confidence constraints. For association rule mining, the target of discovery is not pre-determined, while for classification rule mining there is one and only one pre-determined target. In this paper, we propose to integrate these two mining techniques. The integration is done by focusing on mining a special subset of association rules, called class association rules (CARs). An efficient algorithm is also given for building a classifier based on the set of discovered CARs. Experimental results show that the classifier built this way is, in general, more accurate than that produced by the state-of-the-art classification system C4.5. In addition, this integration helps to solve a number of problems that exist in the current classification systems.
Added
2026-09-14
