Data mining with big data

Xindong WuXingquan ZhuGong-Qing WuWei Ding

article2016TKDE1,916 citations

Presents the HACE theorem to define the fundamental characteristics of big data and establishes a structured processing framework for extracting actionable knowledge from decentralized, evolving, and heterogeneous information sources.

Listen

Modern organizations and scientific disciplines face an unprecedented surge in data generation, producing billions of data points daily across social media, sensory networks, and biomedical research. Traditional database tools and centralized computing methods can no longer capture, manage, or analyze this information within acceptable timeframes. The article sets out to define the core characteristics of massive data environments and propose a structured data-driven processing framework to extract meaningful, real-time knowledge.

To address this challenge, the article synthesizes findings from recent computing literature, national research initiatives, and industrial implementations. It conceptualizes the operational landscape through the HACE theorem—which characterizes Big Data by its Heterogeneous sources, Autonomous decentralized control, and Complex, Evolving relationships—and formulates a three-tiered processing architecture spanning computing platforms, domain semantics, and mining algorithms.

Key findings show that centralizing massive distributed data into a single memory repository is computationally infeasible and cost-prohibitive, necessitating cluster-based parallel programming frameworks such as MapReduce and cloud infrastructures. Furthermore, privacy and domain semantics present fundamental constraints; sharing data across distributed environments requires robust anonymization or secure protocols without sacrificing data utility. At the analytical tier, real-world data streams are inherently sparse, incomplete, and uncertain, requiring advanced preprocessing, local pattern mining, and model fusion rather than traditional direct modeling. Finally, network and relationship complexities scale non-linearly (for example, a one-million-node network entails up to a trillion potential connections), demonstrating that the primary business and scientific value lies in deciphering complex associations and dynamic shifts rather than managing raw volume alone.

These insights demonstrate that simply expanding physical storage capacity is an inadequate strategy. Organizations must adopt decentralized analytical models that process data locally and fuse the resulting models globally, thereby reducing network transmission costs, mitigating security risks, and enabling near real-time operational feedback. Decision-makers should prioritize investing in scalable, cluster-based processing architectures and privacy-preserving data sharing protocols rather than pursuing centralized data warehouses. Moving forward, technical teams should implement streaming analytics capable of handling concept drift and dynamic pattern evolution, supported by ongoing pilot validations in high-volume, multi-source operational domains.

Wu et al (2016).pdf
Cover for Data mining with big data

Abstract

Big Data concerns large-volume, complex, growing data sets with multiple, autonomous sources. With the fast development of networking, data storage, and the data collection capacity, Big Data is now rapidly expanding in all science and engineering domains, including physical, biological and bio-medical sciences. This article presents a HACE theorem that characterizes the features of the Big Data revolution, and proposes a Big Data processing model, from the data mining perspective. This data-driven model involves demand-driven aggregation of information sources, mining and analysis, user interest modeling, and security and privacy considerations. We analyze the challenging issues in the data-driven model and also in the Big Data revolution.

Table of Contents

  • 2. Big Data Characteristics: HACE Theorem
  • 2.1 Huge Data with Heterogeneous and Diverse Dimensionality
  • 2.2 Autonomous Sources with Distributed and Decentralized Control
  • 2.3 Complex and Evolving Relationships
  • 3. Data Mining Challenges with Big Data
  • 3.1 Tier I: Big Data Mining Platform
  • 3.2 Tier II: Big Data Semantics and Application Knowledge
  • 3.2.1 Information Sharing and Data Privacy
  • 3.2.2 Domain and Application Knowledge
  • 3.3 Tier III: Big Data Mining Algorithms
  • 3.3.2 Mining from Sparse, Uncertain, and Incomplete Data
  • 3.3.3 Mining Complex and Dynamic Data
  • 4. Research Initiatives and Projects
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — HACE Theorem for Big Data Characteristics

    theoretical result

    The HACE Theorem defines the foundational characteristics of Big Data environments from a data mining perspective:

    • H (Huge volume with Heterogeneous and diverse dimensionality): Big Data is characterized by massive scale collected via disparate data schemata and multimodal representations (e.g., structured records, images, video, text, genomic sequences) for the same underlying entities.
    • A (Autonomous sources with distributed and decentralized control): Data generation and collection sites function independently without centralized control or synchronization, making centralized aggregation impractical due to communication costs, storage limitations, and security policies.
    • CE (Complex and Evolving relationships): Data entities are linked by non-linear, multi-relational, and dynamic dependency networks (e.g., social ties, hyperlinks, temporal trajectories) where features, connectivity, and statistical distributions evolve over time.
  2. Knowl 2 — Three-Tier Big Data Processing and Mining Framework

    model/method

    The Big Data processing framework conceptualizes the end-to-end data mining workflow into three concentric tiers:

    • Tier I: Big Data Mining Platform (Low-Level Computing and Storage): Focuses on distributed data storage architectures, high-performance computing clusters, and parallel programming execution engines (e.g., MapReduce) designed to access and compute across petabyte- and exabyte-scale distributed datasets.
    • Tier II: Big Data Semantics and Application Knowledge (Context and Constraints): Incorporates high-level domain knowledge to guide feature selection and task specification, while enforcing data privacy policies, access control protocols, and anonymization mechanisms across multi-party environments.
    • Tier III: Big Data Mining Algorithms (Algorithmic Processing Cycle): Executes a three-stage mining pipeline consisting of:
      1. Preprocessing and fusing sparse, heterogeneous, uncertain, incomplete, and multi-source data.
      2. Mining complex, dynamic, and structured data.
      3. Synthesizing global knowledge via local learning and model fusion, with feedback loops to dynamically tune preprocessing parameters and models.
  3. Knowl 3 — Multi-Level Local Learning and Model Fusion Architecture

    model/method

    To mine distributed, autonomous data sources without centralizing raw data, the framework utilizes a two-step hierarchical mechanism consisting of local mining followed by global correlation across three levels:

    • Data Level: Distributed nodes compute localized summary statistics (e.g., feature means, variances, frequency counts) and exchange these statistical summaries to construct a global data distribution without transmitting raw sample instances.
    • Model/Pattern Level: Each site independently executes data mining algorithms on its local data partition to discover localized patterns or models. An aggregation mechanism then synthesizes a unified global pattern set from these distributed local patterns.
    • Knowledge Level: Model correlation analysis evaluates the semantic relevance and interdependencies among models generated by distinct sources, weighting and fusing decisions to produce global predictive outputs.
  4. Knowl 4 — Mining Sparse, Uncertain, and Incomplete Big Data

    model/method

    Big Data mining requires algorithmic adaptations to handle intrinsic imperfections in high-dimensional and decentralized data:

    • Sparse Data: When data resides in very high-dimensional feature spaces (e.g., >1000> 1000 dimensions), data point scarcity degrades model reliability. Mitigation techniques include dimensionality reduction, streaming feature selection, and sample augmentation via unsupervised learning.
    • Uncertain Data: Attribute measurements subject to sensor noise (e.g., GPS measurement bounds) or deliberate privacy perturbations are modeled as probability distributions rather than deterministic point values. Error-aware algorithms incorporate distribution parameters (such as mean and variance) directly into learning formulations (e.g., distribution-aware Naïve Bayes and decision trees).
    • Incomplete Data: Missing attribute values caused by node dropouts or intentional communication suppression are handled either via algorithms natively robust to missing inputs or through data imputation models that infer missing values from observed inter-attribute dependencies.
  5. Knowl 5 — Mining Complex Semantic Associations and Dynamic Networks

    model/method

    Extracting value from Big Data depends on modeling complex structural and semantic relationships:

    • Heterogeneous Data Models: Data spans structured tabular tables, semi-structured documents, and unstructured streams (text, audio, video). Standard relational structures are supplemented by key-value stores, bigtable structures, document stores, and graph databases.
    • Cross-Modal Semantic Associations: Entities appearing across distinct media types (e.g., news articles, social posts, images, and video feeds) are linked through semantic association models that bridge the semantic gap across modalities.
    • Dynamic Relationship Networks: Network connections scale quadratically (O(N2)O(N^2) for NN nodes), requiring scalable algorithms to detect community structures, track topology evolution, identify network outliers/spammers, and model information diffusion over time.
  6. Knowl 6 — Summation-Based Parallelization for Distributed Learning

    model/method

    In cluster-based Big Data platforms (Tier I), statistical machine learning algorithms (such as linear regression, kk-means, logistic regression, Naïve Bayes, support vector machines, and expectation-maximization) can be parallelized by restructuring their core optimization loops into independent summation operations:

    1. Map Step: The full dataset is partitioned across MM worker nodes. Each mapper independently computes partial summation operations (e.g., gradient accumulations, sufficient statistics, covariance terms) over its assigned data subset.
    2. Reduce Step: The master or reducer nodes sum the intermediate results across all mappers to update global model parameters.

    This transformation eliminates the need to load the complete dataset into the main memory of a single machine.

  7. Knowl 7 — Privacy Preservation and Auditing in Big Data Sharing

    model/method

    In the Big Data processing model (Tier II), sharing data across multiple autonomous entities requires mechanisms that protect sensitive attributes and access behaviors:

    • Access Control and Public Auditing: Secure access policies restrict raw data inspection. Third-party auditing (TPA) uses public-key cryptography to verify data integrity and compliance in large-scale remote cloud storage without downloading local copies or compromising data content.
    • Data Anonymization and Perturbation: Data is sanitized prior to release via suppression of sensitive fields, generalization (ensuring kk-anonymity where every record is indistinguishable from at least k−1k-1 others), random noise injection, and attribute permutation.
    • Access Pattern Obfuscation: Virtual disk interfaces and oblivious access protocols hide query patterns from storage servers to prevent adversaries from correlating co-located users or identifying common user interests.
  8. Knowl 8 — Taxonomy of Concept Drift in Dynamic Data Streams

    definition

    In continuous Big Data stream mining, concept drift describes the phenomenon where underlying target concepts, statistical distributions, or context change over time. Concept drift is categorized into three structural forms:

    • Mutation Drift: Abrupt, discontinuous shifts in the target concept or data distribution.
    • Progressive Drift: Gradual, incremental transitions of the concept over extended time intervals.
    • Data Distribution Drift: Systematic variations in the underlying marginal or conditional probability distributions of the data.

    These drift modes can manifest and require monitoring across single features, multiple joint features, or dynamic streaming feature spaces.

Coverage note — Summaries of national research funding initiatives, specific project administrative grant details (Section 4), and general background reviews of related literature (Section 5) were deliberately omitted as they do not constitute standalone technical contributions.

References

  1. 1.Ahmed and Karypis 2012, Rezwan Ahmed, George Karypis, Algorithms for mining the evolution of conserved relational states in dynamic networks, Knowledge and Information Systems, December 2012, Volume 33, Issue 3, pp 603-630
  2. 2.Alam et al. 2012, Md. Hijbul Alam, JongWoo Ha, SangKeun Lee, Novel approaches to crawling important pages early, Knowledge and Information Systems, December 2012, Volume 33, Issue 3, pp 707-734
  3. 3.Aral S. and Walker D. 2012, Identifying influential and susceptible members of social networks, Science, vol.337, pp.337-341.
  4. 4.Machanavajjhala and Reiter 2012, Ashwin Machanavajjhala, Jerome P. Reiter: Big privacy: protecting confidentiality in big data. ACM Crossroads, 19(1): 20-23, 2012.
  5. 5.Banerjee and Agarwal 2012, Soumya Banerjee, Nitin Agarwal, Analyzing collective behavior from blogs using swarm intelligence, Knowledge and Information Systems, December 2012, Volume 33, Issue 3, pp 523-547
  6. 6.Birney E. 2012, The making of ENCODE: Lessons for big-data projects, Nature, vol.489, pp.49-51.
  7. 7.Bollen et al. 2011, J. Bollen, H. Mao, and X. Zeng, Twitter Mood Predicts the Stock Market, Journal of Computational Science, 2(1):1-8, 2011.
  8. 8.Borgatti S., Mehra A., Brass D., and Labianca G. 2009, Network analysis in the social sciences, Science, vol. 323, pp.892-895.
  9. 9.Bughin et al. 2010, J Bughin, M Chui, J Manyika, Clouds, big data, and smart assets: Ten tech-enabled business trends to watch, McKinSey Quarterly, 2010.
  10. 10.Centola D. 2010, The spread of behavior in an online social network experiment, Science, vol.329, pp.1194-1197.
  11. 11.Chang et al., 2009, Chang E.Y., Bai H., and Zhu K., Parallel algorithms for mining large-scale rich-media data, In: Proceedings of the 17th ACM International Conference on Multimedia (MM '09), New York, NY, USA, 2009, pp. 917-918.
  12. 12.Chen et al. 2004, R. Chen, K. Sivakumar, and H. Kargupta, Collective Mining of Bayesian Networks from Distributed Heterogeneous Data, Knowledge and Information Systems, 6(2):164-187, 2004.
  13. 13.Chen et al. 2012, Yi-Cheng Chen, Wen-Chih Peng, Suh-Yin Lee, Efficient algorithms for influence maximization in social networks, Knowledge and Information Systems, December 2012, Volume 33, Issue 3, pp 577-601
  14. 14.Chu et al., 2006, Chu C.T., Kim S.K., Lin Y.A., Yu Y., Bradski G.R., Ng A.Y., Olukotun K., Map-reduce for machine learning on multicore, In: Proceedings of the 20th Annual Conference on Neural Information Processing Systems (NIPS '06), MIT Press, 2006, pp. 281-288.
  15. 15.Cormode G. and Srivastava D. 2009, Anonymized Data: Generation, Models, Usage, in Proc. of SIGMOD, 2009. pp. 1015-1018.
  16. 16.Das et al., 2010, Das S., Sismanis Y., Beyer K.S., Gemulla R., Haas P.J., McPherson J., Ricardo: Integrating R and Hadoop, In: Proceedings of the 2010 ACM SIGMOD International Conference on Management of data (SIGMOD '10), 2010, pp. 987-998.
  17. 17.Dewdney P., Hall P., Schilizzi R., and Lazio J. 2009, The square kilometre Array, Proc. of IEEE, vol.97, no.8.
  18. 18.Domingos and Hulten, 2000, Domingos P. and Hulten G., Mining high-speed data streams, In: Proceedings of the sixth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD΄00), 2000, pp. 71-80.
  19. 19.Duncan G. 2007, Privacy by design, Science, vol. 317, pp.1178-1179.
  20. 20.Efron B. 1994, Missing data, imputation, and the Bootstrap, Journal of the American Statistical Association, vol.89, no.426, pp.463-475.
  21. 21.Ghoting et al., 2009, Ghoting A., Pednault E., Hadoop-ML: An infrastructure for the rapid implementation of parallel reusable analytics, In: Proceedinds of the Large-Scale Machine Learning: Parallelism and Massive Datasets Workshop (NIPS-2009).
  22. 22.Gillick et al., 2006, Gillick D., Faria A., DeNero J., MapReduce: Distributed Computing for Machine Learning, Berkley, December 18, 2006.
  23. 23.Helft M. 2008, Google uses searches to track Flu’s spread, The New York Times, http://www.nytimes.com/2008/11/12/technology/internet/12flu.html.
  24. 24.Howe D. et al. 2008, Big data: the future of biocuration, Nature, 455, pp.47-50, Sept. 2008.
  25. 25.Huberman B. 2012, Sociology of science: Big data deserve a bigger audience, Nature, vol. 482, pp.308.
  26. 26.IBM 2012, What is big data: Bring big data to the enterprise, http://www-01.ibm.com/software/data/bigdata/, IBM.
  27. 27.Jacobs A. 2009, The pathologies of big data, Communication of the ACM, vol.52, no.8, pp.36-44.
  28. 28.Kopanas et al. 2002, I. Kopanas, N. Avouris, and S. Daskalaki, The Role of Domain Knowledge in a Large Scale Data Mining Project, in I.P Vlahavas, C.D. Spyropoulos (eds), Methods and Applications of Artificial Intelligence, Lecture Notes in AI, LNAI no. 2308, pp. 288-299, Springer-Verlag, Berlin, 2002.
  29. 29.Labrinidis and Jagadish 2012, A. Labrinidis and H. Jagadish, Challenges and Opportunities with Big Data, In Proc. of the VLDB Endowment, 5(12):2032-2033, 2012.
  30. 30.Lindell Y. and Pinkas B. 2000, Privacy Preserving Data Mining, Journal of Cryptology, pp.36-54.
  31. 31.Liu and Wang 2012, Wuying Liu, Ting Wang, Online active multi-field learning for efficient email spam filtering, Knowledge and Information Systems, October 2012, Volume 33, Issue 1, pp 117-136
  32. 32.Lorch et al, 2013, J. Lorch, B. Parno, J. Mickens, M. Raykova, and J. Schiffman, Shoroud: Ensuring Private Access to Large-Scale Data in the Data Center, In: Proc. of the 11th USENIX Conference on File and Storage Technologies (FAST’13), San Jose, CA, 2013.
  33. 33.Luo et al. 2012, Dijun Luo, Chris Ding, Heng Huang, Parallelization with Multiplicative Algorithms for Big Data Mining, In: Proc. of IEEE 12th International Conference on Data Mining, pp.489-498, 2012
  34. 34.Mervis J. 2012, U.S. SCIENCE POLICY: Agencies Rally to Tackle Big Data, Science, vol.336, no.6077, pp.22.
  35. 35.Michel F. 2012, How many photos are uploaded to Flickr every day and month? http://www.flickr.com/photos/franckmichel/6855169886/.
  36. 36.Mitchell T. 2009, Mining our reality, Science, vol.326, pp.1644-1645.
  37. 37.Nature Editorial 2008, Community cleverness required, Nature, Vol.455, no.7209, Sept. 4, 2008.
  38. 38.Papadimitriou and Sun, 2008, Papadimitriou S., Sun J., Disco: Distributed co-clustering with map-reduce: A case study towards petabyte-scale end-to-end mining. In: Proceedings of the 8th IEEE International Conference on Data Mining (ICDM '08), 2008, pp. 512-521.
  39. 39.Ranger et al., 2007, Ranger C., Raghuraman R., Penmetsa A., Bradski, G., and Kozyrakis C., Evaluating MapReduce for multi-core and multiprocessor systems, In: Proceedings of the 13th IEEE International Symposium on High Performance Computer Architecture (HPCA '07), 2007, pp. 13-24.
  40. 40.Rajaraman and Ullman, 2011, A. Rajaraman and J. Ullman, Mining of Massive Datasets, Cambridge University Press, 2011.
  41. 41.Reed C., Thompson D., Majid W., and Wagstaff K. 2011, Real time machine learning to find fast transient radio anomalies: A semi-supervised approach combining detection and RFI excision, Int’l Astronomical Union Sym. on Time Domain Astronomy, UK. Sept. 2011
  42. 42.Schadt E. 2012, The changing privacy landscape in the era of big data, Molecular Systems, 8, Article number 612.
  43. 43.Shafer et al. 1996, J. Shafer, R. Agrawal, and M. Mehta, SPRINT: A Scalable Parallel Classifier for Data Mining, In: Proc. of the 22nd VLDB Conference, Mumbai, India, 1996.
  44. 44.Silva et al. 2012, Alzennyr da Silva, Raja Chiky, Georges Hébrail, A clustering approach for sampling data streams in sensor networks, Knowledge and Information Systems, July 2012, Volume 32, Issue 1, pp 1-23
  45. 45.Su et al., 2006, Su K., Huang H., Wu X., and Zhang S., A logical framework for identifying quality knowledge from different data sources, Decision Support Systems, 2006, 42(3): 1673-1683.
  46. 46.Twitter Blog 2012, Dispatch from the Denver debate, http://blog.twitter.com/2012/10/dispatch-from-denver-debate.html, October 2012.
  47. 47.Wegener et al., 2009, Wegener D., Mock M., Adranale D., Wrobel S., Toolkit-Based high-performance data mining of large data on MapReduce clusters, In: Proceedings of the ICDM Workshop, 2009, pp. 296-301.
  48. 48.Wang et al. 2013, Qian Wang; Kui Ren; Wenjing Lou, Privacy-Preserving Public Auditing for Data Storage Security in Could Computing, IEEE Transactions on Computers, 62(2):362-375, 2013.
  49. 49.Wu X. and Zhu X. 2008, Mining with Noise Knowledge: Error-Aware Data Mining, IEEE Transactions on Systems, Man and Cybernetics, Part A, vol.38, no.4, pp.917-932.
  50. 50.Wu X. and Zhang S. 2003, Synthesizing High-Frequency Rules from Different Data Sources, IEEE Transactions on Knowledge and Data Engineering, vol.15, no.2, pp.353-367.
  51. 51.Wu et al., 2005, Wu X., Zhang C., and Zhang S., Database classification for multi-database mining, Information Systems, 2005, 30(1): 71-88.
  52. 52.Wu X. 2000, Building Intelligent Learning Database Systems, AI Magazine, vol.21, no.3, pp.61-67.
  53. 53.Wu et al., 2013, Wu X., Yu K., Ding W., Wang H., and Zhu X., Online feature selection with streaming features, IEEE Trans. on Pattern Analysis and Machine Intelligence, 35(5):1178-1192, 2013.
  54. 54.Yao A. 1986, How to generate and exchange secretes, in Proc. Of 27th FOCS Conference, pp.162-167.
  55. 55.Ye et al., 2013, Ye M., Wu X., Hu X., Hu D., Anonymizing classification data using rough set theory, Knowledge-Based Systems, 43: 82-94, 2013,.
  56. 56.Zhao et al. 2012, Jichang Zhao, Junjie Wu, Xu Feng, Hui Xiong, Ke Xu, Information propagation in online social networks: a tie-strength perspective, Knowledge and Information Systems, September 2012, Volume 32, Issue 3, pp 589-608.

Citation

MLA
Xindong Wu, et al. “Data Mining with Big Data”. IEEE Transactions on Knowledge and Data Engineering, vol. 26, no. 1, 2014, pp. 97–107, https://doi.org/10.1109/TKDE.2013.109.
APA
Xindong Wu, Xingquan Zhu, Gong-Qing Wu, & Wei Ding. (2014). Data mining with big data. IEEE Transactions on Knowledge and Data Engineering, 26(1), 97–107. https://doi.org/10.1109/TKDE.2013.109
Chicago
Xindong Wu, Xingquan Zhu, Gong-Qing Wu, and Wei Ding. 2014. “Data Mining with Big Data”. IEEE Transactions on Knowledge and Data Engineering 26 (1): 97–107. https://doi.org/10.1109/TKDE.2013.109.
Harvard
Xindong Wu et al. (2014) “Data mining with big data”, IEEE Transactions on Knowledge and Data Engineering, 26(1), pp. 97–107. Available at: https://doi.org/10.1109/TKDE.2013.109.
Vancouver
1. Xindong Wu, Xingquan Zhu, Gong-Qing Wu, Wei Ding (2014) Data mining with big data. IEEE Transactions on Knowledge and Data Engineering 26:97–107

BibTeX

@article{Xindong_Wu_2014, title={Data mining with big data}, volume={26}, ISSN={1041-4347}, url={http://dx.doi.org/10.1109/TKDE.2013.109}, DOI={10.1109/tkde.2013.109}, number={1}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Xindong Wu and Xingquan Zhu and Gong-Qing Wu and Wei Ding}, year={2014}, month=Jan, pages={97–107} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF