Data Mining: An Overview from a Database Perspective

Ming-Syan ChenJiawei HanPhilip S. Yu

article1996TKDE2,630 citations

Categorizes foundational database-oriented data mining techniques across association rules, data cubes, classification, clustering, and sequence matching to evaluate algorithmic scalability and query efficiency on massive datasets.

Listen

The article surveys data mining and knowledge discovery techniques, with a primary focus on association rule mining from large transactional and relational datasets. It addresses the growing need for scalable methods to extract actionable patterns from massive data volumes generated by modern applications such as retail transactions, web logs, and scientific databases.

The work evaluates a range of algorithms and frameworks for discovering frequent itemsets and generating association rules, comparing their performance across different data characteristics and hardware constraints. Key aspects examined include candidate generation strategies, pruning techniques, parallel and distributed implementations, and handling of constraints such as minimum support and confidence thresholds.

Major findings indicate that Apriori-style breadth-first search remains foundational but can be outperformed by depth-first or hybrid approaches on dense datasets; vertical data formats and bitmap representations significantly reduce I/O costs; and incremental or online mining methods offer practical advantages when data evolves over time. Performance gains of 2- to 10-fold are reported for optimized implementations on benchmark datasets containing millions of transactions.

These results matter because effective pattern discovery directly supports business intelligence, fraud detection, recommendation systems, and scientific hypothesis generation. Organizations that adopt efficient mining pipelines can reduce decision latency and uncover revenue opportunities that simpler query-based analysis would miss.

Recommended next steps include tighter integration of mining with database engines, development of privacy-preserving variants, and extension of the methods to streaming, graph, and multi-relational data. Further empirical studies on real-world, high-dimensional datasets are needed before widespread deployment of the most advanced techniques.

Chen et al (1996).pdf
Cover for Data Mining: An Overview from a Database Perspective

Abstract

Mining information and knowledge from large databases has been recognized by many researchers as a key research topic in database systems and machine learning, and by many industrial companies as an important area with an opportunity of major revenues. Researchers in many different fields have shown great interest in data mining. Several emerging applications in information providing services, such as data warehousing and on-line services over the Internet, also call for various data mining techniques to better understand user behavior, to improve the service provided, and to increase the business opportunities. In response to such a demand, this article is to provide a survey, from a database researcher's point of view, on the data mining techniques developed recently. A classification of the available data mining techniques is provided and a comparative study of such techniques is presented.

Table of Contents

  • 1 Introduction
  • 1.1 Requirements and challenges of data mining
  • 2 An Overview of Data Mining Techniques
  • 2.1 Classifying data mining techniques
  • 2.2 Mining different kinds of knowledge from databases
  • 3 Mining Association Rules
  • 3.1 Algorithm Apriori and DHP
  • 3.2 Mining generalized and multiple-level association rules
  • 3.3 Interestingness of discovered association rules
  • 3.4 Improving the efficiency of mining association rules
  • 3.4.1 Database scan reduction
  • 3.4.2 Sampling: Mining with adjustable accuracy
  • 3.4.3 Incremental updating of discovered association rules
  • 3.4.4 Parallel data mining
  • 4 Multi-level Data Generalization, Summarization, and Characterization
  • 4.1 Data cube approach
  • 4.2 Attribute-oriented induction approach
  • 5 Data Classification
  • 5.1 Classification based on decision trees
  • 5.2 Methods for performance improvement
  • 6 Clustering analysis
  • 6.1 Clustering large applications based upon randomized search
  • 6.2 Focusing methods
  • 6.3 Clustering features and CF trees
  • 7 Pattern-based Similarity Search
  • 7.1 Similarity measures
  • 7.2 Alternative approaches
  • 8 Mining Path Traversal Patterns
  • 9 Summary
  • References

Knowls

  1. Knowl 1 — Formal Model of Association Rules in Transaction Databases

    definition

    In a transaction database, let I={i1,i2,,im}\mathcal{I} = \{i_1, i_2, \ldots, i_m\} denote a set of literal items. Let DD be a set of database transactions, where each transaction TT is an itemset such that TIT \subseteq \mathcal{I}, identified by a unique identifier (TID). A transaction TT is said to contain an itemset XIX \subseteq \mathcal{I} if and only if XTX \subseteq T.

    An association rule is an implication of the form: XYX \Longrightarrow Y where XIX \subset \mathcal{I}, YIY \subset \mathcal{I}, and XY=X \cap Y = \emptyset.

    The rule XYX \Longrightarrow Y holds in transaction database DD with:

    • Support ss, where ss is the proportion of transactions in DD containing XYX \cup Y: Support(XY)={TDXYT}D\text{Support}(X \Longrightarrow Y) = \frac{|\{T \in D \mid X \cup Y \subseteq T\}|}{|D|}
    • Confidence cc, where cc is the proportion of transactions in DD containing XX that also contain YY: Confidence(XY)={TDXYT}{TDXT}=Support(XY)Support(X)\text{Confidence}(X \Longrightarrow Y) = \frac{|\{T \in D \mid X \cup Y \subseteq T\}|}{|\{T \in D \mid X \subseteq T\}|} = \frac{\text{Support}(X \cup Y)}{\text{Support}(X)}

    A rule is defined as strong if its support and confidence meet or exceed pre-determined minimum support (smins_{\min}) and minimum confidence (cminc_{\min}) thresholds.

  2. Knowl 2 — Statistical Independence Filter for Rule Interestingness

    equation

    To filter out misleading association rules ABA \Longrightarrow B that pass minimum support and confidence thresholds simply because the consequent BB occurs with high overall background probability, tests of statistical independence are applied:

    P(AB)P(A)P(B)>d\frac{P(A \cap B)}{P(A)} - P(B) > d or equivalently, P(AB)P(A)P(B)>kP(A \cap B) - P(A)P(B) > k where P(A)P(A) is the probability (support) of itemset AA, P(B)P(B) is the probability of itemset BB, P(AB)P(A \cap B) is the joint support of AA and BB, and dd and kk are positive constant thresholds.

    When P(AB)P(A)P(B)0P(A \cap B) - P(A)P(B) \le 0, the presence of itemset AA decreases or does not change the likelihood of itemset BB (negative association or statistical independence), indicating that the implication is misleading regardless of high confidence.

  3. Knowl 3 — Scan-Reduction Optimization for Large Itemset Mining

    model/method

    In level-wise association rule mining, standard approaches compute candidate kk-itemsets CkC_k from verified large (k1)(k-1)-itemsets Lk1L_{k-1} (Lk1Lk1L_{k-1} * L_{k-1}), requiring one full database scan for every itemset size kk.

    The scan-reduction technique constructs candidate 3-itemsets C3C'_3 directly from candidate 2-itemsets (C2C2C_2 * C_2) rather than waiting for large 2-itemsets L2L_2. If C2C_2 and C3C'_3 fit in main memory, transaction support for candidates in both C2C_2 and C3C'_3 can be counted simultaneously during a single database scan, determining L2L_2 and L3L_3 in parallel.

    Generalizing this process by generating candidate sets CkC'_k (k3k \ge 3) from Ck1C'_{k-1} in memory allows all large itemsets LkL_k (k2k \ge 2) to be evaluated in as few as two total database scans: one initial scan to determine large 1-itemsets L1L_1, and a single final scan to count and determine all higher-order large itemsets.

  4. Knowl 4 — Attribute-Oriented Induction for Database Generalization

    algorithm

    Attribute-Oriented Induction (AOI) is an online data generalization and summarization algorithm that transforms relational data tuples into generalized concept relations or summary cubes using background concept hierarchies.

    Input: Target relation RR of nn tuples matching a query, concept hierarchies for active attributes, and attribute generalization thresholds TiT_i.
    Output: Generalized summary relation / cube with aggregated counts.
    for each attribute AiA_i in RR do
        Count the number of distinct values of AiA_i in RR.
        if the number of distinct values of AiA_i exceeds TiT_i then
            if a concept hierarchy exists for AiA_i then
                Generalize attribute values by ascending the concept hierarchy (concept-tree climbing).
            else
                Remove attribute AiA_i from the generalized relation (attribute removal).
    Replace each original tuple in RR with its generalized tuple.
    Merge identical generalized tuples and aggregate associated measure values (e.g., tuple counts, sums, averages).
    return the generalized summary relation.

    The algorithm scans the raw data relation once. The time complexity of AOI is O(n)O(n) when using a dense data cube structure, and O(nlogp)O(n \log p) when producing a generalized relation, where nn is the number of raw data tuples and pp is the number of distinct generalized summary tuples (pnp \ll n).

  5. Knowl 5 — Data Cube Greedy View Materialization Bound

    theoretical result

    In multidimensional data cubes organized as a lattice of aggregate views, selecting a subset of views to physically materialize under storage constraints to minimize average query response time is NP-hard.

    A greedy selection algorithm that iteratively materializes the view offering the highest incremental improvement in average query response time (given the views already selected) achieves a performance guarantee: the total benefit of the views selected by the greedy algorithm is always within (11/e)63%(1 - 1/e) \approx 63\% of the benefit generated by the optimal subset of views across all lattice configurations.

  6. Knowl 6 — Impurity and Splitting Metrics in Decision Tree Classification

    definition

    Decision tree induction constructs a classifier by recursively selecting an attribute split that maximizes partition purity across training instances. For a dataset TT containing instances partitioned into nn classes where pip_i denotes the relative frequency or probability of class ii:

    1. Information Measure / Entropy: i=i=1npiln(pi)i = -\sum_{i=1}^n p_i \ln(p_i) Attribute selection chooses the attribute that maximizes information gain (minimizing expected post-split entropy across child subtrees).

    2. Gini Impurity Index: gini(T)=1i=1npi2\text{gini}(T) = 1 - \sum_{i=1}^n p_i^2 An attribute split is chosen to minimize the weighted sum of the Gini impurity indices of the resulting partition subsets.

  7. Knowl 7 — CLARANS Randomized Search Clustering Algorithm

    algorithm

    CLARANS (Clustering Large Applications based upon RANdomized Search) identifies kk medoids among nn multidimensional objects by treating the search space as a graph where every node represents a set of kk medoids, and two nodes are adjacent if they differ by exactly one medoid.

    Input: Dataset of nn objects, cluster count kk, maximum neighbors examined maxneighbormaxneighbor, and local search trials numlocalnumlocal.
    Output: Optimal set of kk medoids minimizing total clustering distance error.
    for i=1i = 1 to numlocalnumlocal do
        Select an initial set of kk medoids (node CurrentCurrent) uniformly at random from the dataset.
        j=1j = 1
        while jmaxneighborj \le maxneighbor do
            Select a random neighbor node SS' by swapping one current medoid with a random non-medoid object.
            Compute the total distance cost differential ΔE=Cost(S)Cost(Current)\Delta E = \text{Cost}(S') - \text{Cost}(Current).
            if ΔE<0\Delta E < 0 then
                Current=SCurrent = S'
                j=1j = 1
            else
                j=j+1j = j + 1
        Record CurrentCurrent as a local minimum.
    return the local minimum node with the lowest overall cost across all numlocalnumlocal runs.

    The computational complexity of each iteration in CLARANS scales linearly (O(n)O(n)) with the number of objects nn, compared to O(k(nk)2)O(k(n-k)^2) per iteration in deterministic Partitioning Around Medoids (PAM).

  8. Knowl 8 — Clustering Feature and CF-Tree in BIRCH

    definition

    In the BIRCH clustering method, a Clustering Feature (CFCF) is a summary triplet representing the zero-th, first, and second moments of a subcluster of dd-dimensional points.

    Given a cluster of NN dd-dimensional vectors {Xi}i=1N\{\vec{X}_i\}_{i=1}^N, CFCF is defined as: CF=(N,LS,SS)CF = (N, \vec{LS}, SS) where:

    • NN is the number of points in the subcluster.
    • LS=i=1NXi\vec{LS} = \sum_{i=1}^N \vec{X}_i is the dd-dimensional linear sum of the points.
    • SS=i=1NXi2SS = \sum_{i=1}^N \|\vec{X}_i\|^2 is the scalar square sum of the points.

    A CF-Tree is a height-balanced search tree defined by a maximum branching factor BB and a cluster diameter threshold TT at leaf entries. Non-leaf nodes maintain CFCF vectors equal to the sum of their child CFCF vectors. Inserting NN points requires a single data pass with O(N)O(N) CPU and I/O complexity. If the tree exceeds memory limits, it is dynamically rebuilt with an increased threshold TT directly from leaf CFCF entries without rescanning raw points.

  9. Knowl 9 — Subsequence Matching via Normalized Correlation and DFT

    equation

    To match a target query sequence {xj}j=1n\{x_j\}_{j=1}^n against a database sequence {yj}j=1N\{y_j\}_{j=1}^N (nNn \le N) with invariance to scaling and translation offsets, the normalized linear cross-correlation coefficient cic_i at offset i{1,,Nn+1}i \in \{1, \ldots, N - n + 1\} is defined as:

    ci=j=1nxjyi+jj=1nxj2j=1nyi+j2c_i = \frac{\sum_{j=1}^n x_j y_{i+j}}{\sqrt{\sum_{j=1}^n x_j^2} \sqrt{\sum_{j=1}^n y_{i+j}^2}}

    By appending zero-padding to length l=N+n1l = N + n - 1 and applying the Discrete Fourier Transform (DFT), the convolution theorem and Parseval's theorem allow correlation across all offsets to be computed in the frequency domain via pointwise multiplication followed by the Inverse Discrete Fourier Transform (F1F^{-1}):

    ci=F1{XjYj}j=1nXj2j=1nYj2c_i = \frac{F^{-1}\{X_j^* Y_j\}}{\sqrt{\sum_{j=1}^n X_j^2} \sqrt{\sum_{j=1}^n Y_j^2}} where XjX_j and YjY_j are the DFT coefficients of the padded sequences, and XjX_j^* is the complex conjugate of XjX_j. The correlation coefficient satisfies 1ci1-1 \le c_i \le 1, where 11 denotes an exact shape match.

  10. Knowl 10 — Mining Path Traversal Patterns via Maximal Forward References

    model/method

    Mining user traversal patterns in linked hypertext/web environments identifies frequent navigation routes while filtering out backward references used solely for backtracking.

    1. Conversion to Maximal Forward References (MFRs): An access log sequence of user page requests containing forward links and backtracks (e.g., {A,B,C,D,C,B,E,G,H,G,W,A,O,U,O,V}\{A, B, C, D, C, B, E, G, H, G, W, A, O, U, O, V\}) is partitioned. A forward traversal continues until a backward reference occurs; the path from the starting page to the deepest point reached prior to backtracking is output as an MFR. For the example sequence, this generates the set of MFRs: {ABCD,ABEGH,ABEGW,AOU,AOV}\{ABCD, ABEGH, ABEGW, AOU, AOV\}
    2. Mining Large Reference Sequences: From the collection of MFRs across all user sessions, frequent consecutive subsequences (large reference sequences) meeting a minimum support threshold are mined. Large reference sequences that are not proper subsequences of any other large reference sequence are identified as maximal reference sequences, representing dominant navigation paths.

Coverage note — None was omitted; the top 10 knowls capture all primary contributed models, algorithms, definitions, theoretical bounds, and equations spanning association rules, generalization, classification, clustering, time-series similarity search, and path traversal pattern mining.

References

  1. 1.R. Agrawal, C. Faloutsos, and A. Swami. Efficient Similarity Search in Sequence Databases. Proceedings of the 4th Intl. conf. on Foundations of Data Organization and Algorithms, October, 1993.
  2. 2.R. Agrawal, S. Ghosh, T. Imielinski, B. Iyer, and A. Swami. An Interval Classifier for Database Mining Applications. Proceedings of the 18th International Conference on Very Large Data Bases, pages 560–573, August 1992.
  3. 3.R. Agrawal, T. Imielinski, and A. Swami. Database Mining: A Performance Perspective. IEEE Transactions on Knowledge and Data Engineering, pages 914–925, December 1993.
  4. 4.R. Agrawal, T. Imielinski, and A. Swami. Mining Association Rules between Sets of Items in Large Databases. Proceedings of ACM SIGMOD, pages 207–216, May 1993.
  5. 5.R. Agrawal, K.-I. Lin, H.S. Sawhney, and K. Shim. Fast Similarity Search in the Presence of Noise, Scaling, and Translation in Time-Series Databases. Proceedings of the 21th International Conference on Very Large Data Bases, pages 490–501, September 1995.
  6. 6.R. Agrawal, M. Mehta, J. Shafer, R. Srikant, A. Arning, and T. Bollinger. The Quest data mining system. In Proc. 1996 Int'l Conf. on Data Mining and Knowledge Discovery (KDD'96), Portland, Oregon, August 1996.
  7. 7.R. Agrawal and R. Srikant. Fast Algorithms for Mining Association Rules in Large Databases. Proceedings of the 20th International Conference on Very Large Data Bases, pages 478–499, September 1994.
  8. 8.R. Agrawal and R. Srikant. Mining Sequential Patterns. Proceedings of the 11th International Conference on Data Engineering, pages 3–14, March 1995.
  9. 9.K. K. Al-Taha, R. T. Snodgrass, and M. D. Soo. Bibliography on spatiotemporal databases. ACM SIGMOD Record, 22(1):59–67, March 1993.
  10. 10.T.M. Anwar, H.W. Beck, and S.B. Navathe. Knowledge Mining by Imprecise Querying: A Classification-Based Approach. Proceedings of the 8th International Conference on Data Engineering, pages 622–630, February 1992.
  11. 11.N. Beckmann, H.-P. Kriegel, R. Schneider, and B. Seeger. The R*-tree: An efficient and robust access method for points and rectangles. In Proc. 1990 ACM-SIGMOD Int. Conf. Management of Data, pages 322–331, Atlantic City, NJ, June 1990.
  12. 12.M. Bieber and J. Wan. Backtracking in a Multiple-Window Hypertext Environment. ACM European Conf. on Hypermedia Technology, pages 158–166, 1994.
  13. 13.R. Brachman and T. Anand. The process of knowledge discovery in databases: A human-centered approach. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 37–58. AAAI/MIT Press, 1996.
  14. 14.L. Breiman, J. Friedman, R. Olshen, and C. Stone. Classification of Regression Trees. Wadsworth, 1984.
  15. 15.E. Caramel, S. Crawford, and H. Chen. Browsing in Hypertext: A Cognitive Study. IEEE Transactions on Systems, Man and Cybernetics, 22(5):865–883, September 1992.
  16. 16.L. D. Catledge and J. E. Pitkow. Characterizing browsing strategies in the world-wide web. Proceedings of the 3rd WWW Conference, April 1995.
  17. 17.P. K. Chan and S. J. Stolfo. Learning arbiter and combiner trees from partitioned data for scaling machine learning. Proc. 1st Int. Conf. on Knowledge Discovery and Data Mining (KDD'95), pages 39–44, August 1995.
  18. 18.P. Cheeseman and J. Stutz. Bayesian classification (AutoClass): Theory and results. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 153–180. AAAI/MIT Press, 1996.
  19. 19.M.-S. Chen, J.-S. Park, and P. S. Yu. Data Mining for Path Traversal Patterns in a Web Environment. Proceedings of the 16th International Conference on Distributed Computing Systems, pages 385–392, May 27-30 1996.
  20. 20.M.-S. Chen and P. S. Yu. Using Multi-Attribute Predicates for Mining Classification Rules. IBM Research Report, 1995.
  21. 21.D.W. Cheung, J. Han, V. Ng, and C.Y. Wong. Maintenance of discovered association rules in large databases: An incremental updating technique. In Proc. 1996 Int'l Conf. on Data Engineering, New Orleans, Louisiana, Feb. 1996.
  22. 22.C. Clifton and D. Marks. Security and privacy implications of data mining. In Proc. 1996 SIGMOD'96 Workshop on Research Issues on Data Mining and Knowledge Discovery (DMKD'96), pages 15–20, Montreal, Canada, June 1996.
  23. 23.J. December and N. Randall. The World Wide Web Unleashed. SAMS Publishing, 1994.
  24. 24.V. Dhar and A. Tuzhilin. Abstract-Driven Pattern Discovery in Databases. IEEE Transactions on Knowledge and Data Engineering, pages 926–938, December 1993.
  25. 25.S. Dzeroski. Inductive logic programming and knowledge discovery. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 117–152. AAAI/MIT Press, 1996.
  26. 26.J. Elder IV and D. Pregibon. A statistical perspective on knowledge discovery in databases. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 83–115. AAAI/MIT Press, 1996.
  27. 27.M. Ester, H.-P. Kriegel, and X. Xu. Knowledge discovery in large spatial databases: Focusing techniques for efficient class identification. In Proc. 4th Int. Symp. on Large Spatial Databases (SSD'95), pages 67–82, Portland, Maine, August 1995.
  28. 28.C. Faloutsos and K.-I. Lin. FastMap: A Fast Algorithm for Indexing, Data-Mining and Visualization of Tranditional and Multimedia Datasets. Proceedings of ACM SIGMOD, pages 163–174, May, 1995.
  29. 29.C. Faloutsos, M. Ranganathan, and Y. Manolopoulos. Fast Subsequence Matching in Time-Series Databases. Proceedings of ACM SIGMOD, Minneapolis, MN, pages 419–429, May, 1994.
  30. 30.U. M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy. Advances in Knowledge Discovery and Data Mining. AAAI/MIT Press, 1996.
  31. 31.D. Fisher. Improving inference through conceptual clustering. In Proc. 1987 AAAI Conf., pages 461–465, Seattle, Washington, July 1987.
  32. 32.D. Fisher. Optimization and simplification of hierarchical clusterings. In Proc. 1st Int. Conf. on Knowledge Discovery and Data Mining (KDD'95), pages 118–123, Montreal, Canada, Aug. 1995.
  33. 33.Y. Fu and J. Han. Meta-rule-guided mining of association rules in relational databases. In Proc. 1st Int'l Workshop on Integration of Knowledge Discovery with Deductive and Object-Oriented Databases (KDOOD'95), pages 39–46, Singapore, Dec. 1995.
  34. 34.B. R. Gains. Tranforming rules and trees into comprehensive knowledge structures. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 205–228. AAAI/MIT Press, 1996.
  35. 35.A. Gupta, V. Harinarayan, and D. Quass. Aggregate-query processing in data warehousing environment. In Proc. 21st Int. Conf. Very Large Data Bases, pages 358–369, Zurich, Switzerland, Sept. 1995.
  36. 36.J. Han. Mining knowledge at multiple concept levels. In Proc. 4th Int. Conf. on Information and Knowledge Management, pages 19–24, Baltimore, Maryland, Nov. 1995.
  37. 37.J. Han, Y. Cai, and N. Cercone. Data-driven discovery of quantitative rules in relational databases. IEEE Trans. Knowledge and Data Engineering, 5:29–40, 1993.
  38. 38.J. Han and Y. Fu. Dynamic generation and refinement of concept hierarchies for knowledge discovery in databases. In Proc. AAAI'94 Workshop on Knowledge Discovery in Databases (KDD'94), pages 157–168, Seattle, WA, July 1994.
  39. 39.J. Han and Y. Fu. Discovery of Multiple-Level Association Rules from Large Databases. Proceedings of the 21th International Conference on Very Large Data Bases, pages 420–431, September 1995.
  40. 40.J. Han and Y. Fu. Exploration of the power of attribute-oriented induction in data mining. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 399–421. AAAI/MIT Press, 1996.
  41. 41.J. Han, S. Nishio, and H. Kawano. Knowledge discovery in object-oriented and active databases. In F. Fuchi and T. Yokoi, editors, Knowledge Building and Knowledge Sharing, pages 221–230. Ohmsha, Ltd. and IOS Press, 1994.
  42. 42.J. Han, Y. Fu, W. Wang, J. Chiang, W. Gong, K. Koperski, D. Li, Y. Lu, A. Rajan, N. Stefanovic, B. Xia, and O. R. Zaiane. DBMiner: A system for mining knowledge in large relational databases. In Proc. 1996 Int'l Conf. on Data Mining and Knowledge Discovery (KDD'96), Portland, Oregon, August 1996.
  43. 43.V. Harinarayan, J. D. Ullman, and A. Rajaraman. Implementing data cubes efficiently. In Proc. 1996 ACM-SIGMOD Int. Conf. Management of Data, Montreal, Canada, June 1996.
  44. 44.IBM. Scalable POWERparallel Systems. Technical Report GA23-2475-02, February 1995.
  45. 45.T. Imielinski and A. Virmani. DataMine – application programming interface and query language for kdd applications. In Proc. 1996 Int'l Conf. on Data Mining and Knowledge Discovery (KDD'96), Portland, Oregon, August 1996.
  46. 46.H.V. Jagadish. A Retrieval Technique for Similar Shapes. Proceedings of ACM SIGMOD, pages 208–217, 1991.
  47. 47.A. K. Jain and R. C. Dubes. Algorithms for Clustering Data. Printice Hall, 1988.
  48. 48.L. Kaufman and P. J. Rousseeuw. Finding Groups in Data: an Introduction to Cluster Analysis. John Wiley & Sons, 1990.
  49. 49.D. Keim, H. Kriegel, and T. Seidl. Supporting data mining of large databases by visual feedback queries. In Proc. 10th of Int. Conf. on Data Engineering, pages 302–313, Houston, TX, Feb. 1994.
  50. 50.W. Kim. Introduction to Objected-Oriented Databases. The MIT Press, Cambridge, Massachusetts, 1990.
  51. 51.M. Klemettinen, H. Mannila, P. Ronkainen, H. Toivonen, and A. I. Verkamo. Finding interesting rules from large sets of discovered association rules. In Proc. 3rd Int'l Conf. on Information and Knowledge Management, pages 401–408, Gaithersburg, Maryland, Nov. 1994.
  52. 52.W. Klösgen. Explora: a multipattern and multistrategy discovery assistant. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 249–271. AAAI/MIT Press, 1996.
  53. 53.K. Koperski and J. Han. Discovery of spatial association rules in geographic information databases. In Proc. 4th Int'l Symp. on Large Spatial Databases (SSD'95), pages 47–66, Portland, Maine, Aug. 1995.
  54. 54.C-S. Li, P.S. Yu, and V. Castelli. HierarchyScan: A Hierarchical Similarity Search Algorithm for Databases of Long Sequences. Proceedings of the 12th International Conference on Data Engineering, February 1996.
  55. 55.H. Lu, R. Setiono, and H. Liu. NeuroRule: A Connectionist Approach to Data Mining. Proceedings of the 21th International Conference on Very Large Data Bases, pages 478–489, September 1995.
  56. 56.W. Lu, J. Han, and B. C. Ooi. Knowledge discovery in large spatial databases. In Far East Workshop on Geographic Information Systems, pages 275–289, Singapore, June 1993.
  57. 57.H. Mannila, H. Toivonen, and A. Inkeri Verkamo. Efficient Algorithms for Discovering Association Rules. Proceedings of AAAI Workshop on Knowledge Discovery in Databases, pages 181–192, July, 1994.
  58. 58.C.J. Matheus, G. Piatetsky-Shapiro, and D. McNeil. Selecting and reporting what is interesting: The KEFIR application to healthcare data. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 495–516. AAAI/MIT Press, 1996.
  59. 59.M. Mehta, R. Agrawal, and J. Rissanen. SLIQ: A fast scalable classifier for data mining. In Proc. 1996 Int. Conference on Extending Database Technology (EDBT'96), Avignon, France, March 1996.
  60. 60.R. S. Michalski. A theory and methodology of inductive learning. In Michalski et al., editor, Machine Learning: An Artificial Intelligence Approach, Vol. 1, pages 83–134. Morgan Kaufmann, 1983.
  61. 61.R. S. Michalski, L. Kerschberg, K. A. Kaufman, and J.S. Ribeiro. Mining for knowledge in databases: The INLEN architecture, initial implementation and first results. J. Int. Info. Systems, 1:85–114, 1992.
  62. 62.R. Ng and J. Han. Efficient and effective clustering method for spatial data mining. In Proc. 1994 Int. Conf. Very Large Data Bases, pages 144–155, Santiago, Chile, September 1994.
  63. 63.D. E. O'Leary. Knowledge discovery as a threat to database security. In G. Piatetsky-Shapiro and W. J. Frawley, editors, Knowledge Discovery in Databases, pages 507–516. AAAI/MIT Press, 1991.
  64. 64.A. Papoulis. Probability, Random Variable, and Stochastic Process. McGraw Hills, New York, 1984.
  65. 65.J.-S. Park, M.-S. Chen, and P. S. Yu. Mining Association Rules with Adjustable Accuracy. IBM Research Report, 1995.
  66. 66.J.-S. Park, M.-S. Chen, and P. S. Yu. An Effective Hash Based Algorithm for Mining Association Rules. Proceedings of ACM SIGMOD, pages 175–186, May, 1995.
  67. 67.J.-S. Park, M.-S. Chen, and P. S. Yu. Efficient Parallel Data Mining for Association Rules. Proceedings of the 4th Intern'l Conf. on Information and Knowledge Management, pages 31–36, Nov. 29 - Dec. 3, 1995.
  68. 68.G. Piatetsky-Shapiro. Discovery, analysis, and presentation of strong rules. In G. Piatetsky-Shapiro and W. J. Frawley, editors, Knowledge Discovery in Databases, pages 229–238. AAAI/MIT Press, 1991.
  69. 69.G. Piatetsky-Shapiro, U. Fayyad, and P. Smyth. From data mining to knowledge discovery: An overview. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 1–35. AAAI/MIT Press, 1996.
  70. 70.G. Piatetsky-Shapiro and W. J. Frawley. Knowledge Discovery in Databases. AAAI/MIT Press, 1991.
  71. 71.J. R. Quinlan. Induction of decision trees. Machine Learning, 1:81–106, 1986.
  72. 72.J. R. Quinlan. C4.5: Programs for Machine Learning. Morgan Kaufmann, 1993.
  73. 73.A. Savasere, E. Omiecinski, and S. Navathe. An Efficient Algorithm for Mining Association Rules in Large Databases. Proceedings of the 21th International Conference on Very Large Data Bases, pages 432–444, September 1995.
  74. 74.P. G. Selfridge, D. Srivastava, and L. O. Wilson. IDEA: Interactive data exploration and analysis. In Proc. 1996 ACM-SIGMOD Int. Conf. Management of Data, Montreal, Canada, June 1996.
  75. 75.W. Shen, K. Ong, B. Mitbander, and C. Zaniolo. Metaqueries for data mining. In U.M. Fayyad, G. Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, editors, Advances in Knowledge Discovery and Data Mining, pages 375–398. AAAI/MIT Press, 1996.
  76. 76.A. Silberschatz, M. Stonebraker, and J. D. Ullman. Database research: Achievements and opportunities into the 21st century. In Report of an NSF Workshop on the Future of Database Systems Research, May 1995.
  77. 77.A. Silberschatz and A. Tuzhilin. On subjective measure of interestingness in knowledge discovery. In Proc. 1st Int. Conf. on Knowledge Discovery and Data Mining (KDD'95), pages 275–281, Montreal, Canada, Aug. 1995.
  78. 78.R. Srikant and R. Agrawal. Mining Generalized Association Rules. Proceedings of the 21th International Conference on Very Large Data Bases, pages 407–419, September 1995.
  79. 79.R. Srikant and R. Agrawal. Mining quantitative association rules in large relational tables. In Proc. 1996 ACM-SIGMOD Int. Conf. Management of Data, Montreal, Canada, June 1996.
  80. 80.R. Stam and R. Snodgrass. A bibliography on temporal databases. IEEE Bulletin on Data Engineering, 11(4), December 1988.
  81. 81.Y. Stettiner, D. Malah, and D. Chazan. Dynamic Time Warping with Path Control and Non local Cost. Proceedings of 12th IAPR Intern'l Conf. on Pattern Recognition, pages 174–177, October 1994.
  82. 82.S. M. Weiss and C. A. Kulikowski. Computer Systems that Learn: Classification and Prediction Methods from Statistics, Neural Nets, Machine Learning, and Expert Systems. Morgan Kaufman, 1991.
  83. 83.J. Widom. Research problems in data warehousing. In Proc. 4th Int. Conf. on Information and Knowledge Management, pages 25–30, Baltimore, Maryland, Nov. 1995.
  84. 84.W. P. Yan and P. Larson. Eager aggregation and lazy aggregation. In Proc. 21st Int. Conf. Very Large Data Bases, pages 345–357, Zurich, Switzerland, Sept. 1995.
  85. 85.T. Zhang, R. Ramakrishnan, and M. Livny. BIRCH: an efficient data clustering method for very large databases. In Proc. 1996 ACM-SIGMOD Int. Conf. Management of Data, Montreal, Canada, June 1996.
  86. 86.Y. Zhuge, H. Garcia-Molina, J. Hammer, and J. Widom. View maintenance in a warehousing environment. In Proc. 1995 ACM-SIGMOD Int. Conf. Management of Data, pages 316–327, San Jose, CA, May 1995.
  87. 87.W. Ziarko. Rough Sets, Fuzzy Sets and Knowledge Discovery. Springer-Verlag, 1994.

Citation

MLA
Ming-Syan Chen, et al. “Data Mining: An Overview from a Database Perspective”. IEEE Transactions on Knowledge and Data Engineering, vol. 8, no. 6, 1996, pp. 866–83, https://doi.org/10.1109/69.553155.
APA
Ming-Syan Chen, Jiawei Han, & Yu, P. S. (1996). Data mining: an overview from a database perspective. IEEE Transactions on Knowledge and Data Engineering, 8(6), 866–883. https://doi.org/10.1109/69.553155
Chicago
Ming-Syan Chen, Jiawei Han, and P. S. Yu. 1996. “Data Mining: An Overview from a Database Perspective”. IEEE Transactions on Knowledge and Data Engineering 8 (6): 866–83. https://doi.org/10.1109/69.553155.
Harvard
Ming-Syan Chen, Jiawei Han and Yu, P.S. (1996) “Data mining: an overview from a database perspective”, IEEE Transactions on Knowledge and Data Engineering, 8(6), pp. 866–883. Available at: https://doi.org/10.1109/69.553155.
Vancouver
1. Ming-Syan Chen, Jiawei Han, Yu PS (1996) Data mining: an overview from a database perspective. IEEE Transactions on Knowledge and Data Engineering 8:866–883

BibTeX

@article{Ming_Syan_Chen_1996, title={Data mining: an overview from a database perspective}, volume={8}, ISSN={1041-4347}, url={http://dx.doi.org/10.1109/69.553155}, DOI={10.1109/69.553155}, number={6}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Ming-Syan Chen and Jiawei Han and Yu, P.S.}, year={1996}, pages={866–883} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF