Learning under Concept Drift: A Review

Jie LuAnjin LiuFan DongFeng GuJoão GamaGuangquan Zhang

article2019TKDE1,960 citations

Establishes a unified structural framework covering concept drift detection, understanding, and adaptation while evaluating 24 benchmark datasets across more than 130 studies to guide machine learning on non-stationary data streams.

Listen

Modern predictive and decision-support systems increasingly rely on streaming data, but changes in underlying data distributions over time degrade model accuracy. This statistical change, known as concept drift, undermines automated decision-making across critical domains like fraud detection, early warning systems, and consumer behavioral analytics.

The article systematically reviews modern approaches to machine learning in non-stationary environments. It introduces an integrated architectural framework that organizes the handling of concept drift into detection, understanding, and adaptation.

To conduct this review, the authors analyzed more than 130 high-impact studies published primarily over a ten-year span and examined 10 synthetic and 14 real-world benchmark datasets. They categorized drift detection techniques into error-rate monitoring, data distribution metrics, and multiple hypothesis testing. Adaptation approaches were categorized into complete model retraining, ensemble updating, and regional model parameter adjustments.

The review outlines four principal findings. First, while nearly all detection algorithms can establish the timing of a drift, very few existing techniques can quantify its severity or map the specific spatial regions where changes occur. Second, modern adaptation research has shifted toward adaptive models and ensemble strategies, which handle sudden, gradual, and recurring changes more efficiently than simple full retraining. Third, multiple hypothesis testing architectures, especially hierarchical frameworks, are emerging as effective ways to balance detection sensitivity with low false-alarm rates. Fourth, a critical vulnerability across most state-of-the-art methods is an over-reliance on the assumption that ground-truth labels are instantly available, which rarely holds true in operational settings.

These findings have direct operational and economic implications for enterprise data systems. Implementing concept drift awareness protects organizations from silent model degradation, unmanaged operational risk, and misinformed strategic decisions. Furthermore, isolating localized regions of drift avoids the significant computing costs and operational delays associated with retraining large models from scratch.

To address these challenges, future system deployments should prioritize modular ensemble methods and partial update mechanisms over rigid retraining pipelines. Organizations should also invest in unsupervised and semi-supervised drift handling to reduce reliance on costly and delayed manual labeling. Finally, standardized evaluation benchmarks using realistic data streams are needed to rigorously test algorithms before deployment.

Confidence in the core findings remains high due to the extensive literature surveyed. However, because real-world benchmark datasets often lack precise ground-truth labels for drift points, practitioners should validate these methods in domain-specific pilot settings before committing to full-scale production rollouts.

  • Paper: Learning with Drift Detection, João Gama et al. (2004). This seminal work introduced the foundational statistical error-rate tracking framework (DDM) for detecting concept drift in streaming data, providing direct intellectual groundwork for the survey.
  • Paper: Learning from Time-Changing Data with Adaptive Windowing, Albert Bifet et al. (2007). This paper establishes the widely used ADWIN adaptive windowing algorithm with theoretical error guarantees, serving as a cornerstone drift detection and adaptation method reviewed in the survey.
  • Paper: Mining high-speed data streams, Pedro Domingos et al. (2000). This foundational paper introduces Hoeffding trees for high-speed streaming classification, establishing the core data-stream learning paradigm that concept drift algorithms build upon.
  • Paper: A Framework for Clustering Evolving Data Streams, Charu C. Aggarwal et al. (2003). This paper develops the two-phase CluStream framework for clustering evolving data streams, providing fundamental principles for tracking evolving data distributions over time.
  • Paper: A Survey on Transfer Learning, Sinno Jialin Pan et al. (2010). This classic survey establishes key taxonomies and formal definitions for distribution mismatch and transfer learning that inform concept drift characterization.
  • Paper: Online Passive-Aggressive Algorithms, K. Crammer et al. (2003). This work establishes the Passive-Aggressive online learning framework, which provides core adaptive update mechanisms for models responding to changing streaming data.
  • Paper: Detecting and Correcting for Label Shift with Black Box Predictors, Zachary C. Lipton et al. (2018). This paper formalizes black-box detection and correction techniques for label distribution shift, representing a key prerequisite methodology for understanding dataset and concept shift.
Cover for Learning under Concept Drift: A Review

Abstract

Concept drift describes unforeseeable changes in the underlying distribution of streaming data over time. Concept drift research involves the development of methodologies and techniques for drift detection, understanding and adaptation. Data analysis has revealed that machine learning in a concept drift environment will result in poor learning results if the drift is not addressed. To help researchers identify which research topics are significant and how to apply related techniques in data analysis tasks, it is necessary that a high quality, instructive review of current research developments and trends in the concept drift field is conducted. In addition, due to the rapid development of concept drift in recent years, the methodologies of learning under concept drift have become noticeably systematic, unveiling a framework which has not been mentioned in literature. This paper reviews over 130 high quality publications in concept drift related research areas, analyzes up-to-date developments in methodologies and techniques, and establishes a framework of learning under concept drift including three main components: concept drift detection, concept drift understanding, and concept drift adaptation. This paper lists and discusses 10 popular synthetic datasets and 14 publicly available benchmark datasets used for evaluating the performance of learning algorithms aiming at handling concept drift. Also, concept drift related research directions are covered and discussed. By providing state-of-the-art knowledge, this survey will directly support researchers in their understanding of research developments in the field of learning under concept drift.

Table of Contents

  • I Introduction
  • II Problem Description
  • II-A Concept drift definition and the sources
  • II-B The types of concept drift
  • III Concept Drift detection
  • III-A A general framework for drift detection
  • III-B Concept drift detection algorithms
  • III-B1 Error rate-based drift detection
  • III-B2 Data Distribution-based Drift Detection
  • III-B3 Multiple Hypothesis Test Drift Detection
  • III-C Summary of concept drift detection methods/algorithms
  • IV Concept Drift understanding
  • IV-A The time of concept drift occurs (When)
  • IV-B The severity of concept drift (How)
  • IV-C The drift regions of concept drift (Where)
  • IV-D Summary of drift understanding
  • V Drift adaptation
  • V-A Training new models for global drift
  • V-B Model ensemble for recurring drift
  • V-C Adjusting existing models for regional drift
  • VI Evaluation, Datasets and Benchmarks
  • VI-A Evaluation Systems
  • VI-B Synthetic datasets
  • VI-C Real-world datasets
  • VII The Concept Drift Problem in Other Research Areas
  • VII-A Class imbalance
  • VII-B Big data mining
  • VII-C Active learning and semi-supervised learning
  • VII-D Decision Rules
  • VIII Conclusions: findings and future directions
  • References

Knowls

  1. Knowl 1 — Formal Definition and Sources of Concept Drift

    definition

    In streaming data environments, data instances arrive sequentially over time. Let S0,t={d0,d1,…,dt}S_{0,t} = \{d_0, d_1, \dots, d_t\} denote a sequence of observations up to timestamp tt, where each observation di=(Xi,yi)d_i = (X_i, y_i) comprises a feature vector Xi∈RdX_i \in \mathbb{R}^d and a target label yiy_i. The sample sequence S0,tS_{0,t} is governed by a joint probability distribution F0,t(X,y)=Pt(X,y)F_{0,t}(X, y) = P_t(X, y).

    Concept drift occurs at timestamp t+1t+1 if the joint probability distribution governing the data stream changes between time periods, denoted formally as: ∃t:Pt(X,y)≠Pt+1(X,y)\exists t: P_t(X, y) \neq P_{t+1}(X, y)

    Applying the chain rule of probability, the joint distribution decomposes as Pt(X,y)=Pt(X)Pt(y∣X)=Pt(y)Pt(X∣y)P_t(X, y) = P_t(X) P_t(y|X) = P_t(y) P_t(X|y). Consequently, concept drift arises from three distinct sources:

    1. Source I (Virtual Drift / Feature Space Drift): The feature distribution changes while the conditional posterior probability of labels remains constant: Pt(X)≠Pt+1(X)andPt(y∣X)=Pt+1(y∣X)P_t(X) \neq P_{t+1}(X) \quad \text{and} \quad P_t(y|X) = P_{t+1}(y|X) This drift shifts data density in the input space without altering the true decision boundaries.

    2. Source II (Actual Drift / Decision Boundary Drift): The conditional posterior probability of labels changes while the feature distribution remains invariant: Pt(y∣X)≠Pt+1(y∣X)andPt(X)=Pt+1(X)P_t(y|X) \neq P_{t+1}(y|X) \quad \text{and} \quad P_t(X) = P_{t+1}(X) This drift fundamentally modifies the decision boundary and directly degrades the predictive accuracy of previously trained models.

    3. Source III (Combined Drift): Both the input feature distribution and the posterior class distribution change simultaneously: Pt(X)≠Pt+1(X)andPt(y∣X)≠Pt+1(y∣X)P_t(X) \neq P_{t+1}(X) \quad \text{and} \quad P_t(y|X) \neq P_{t+1}(y|X)

  2. Knowl 2 — Taxonomy of Concept Drift Temporal Patterns

    definition

    Concept drift manifests across data streams according to distinct temporal patterns and transition dynamics between a previous concept Pt(X,y)P_t(X,y) and an emerging concept Pt+1(X,y)P_{t+1}(X,y):

    • Sudden (Abrupt) Drift: An existing concept is abruptly and completely replaced by a new concept within a very short time interval.
    • Gradual Drift: A new concept gradually replaces an old concept over an extended transition window. During the transition, instances generated by the old concept and the new concept appear concurrently, with the proportion of instances from the new concept increasing over time.
    • Incremental Drift: The underlying data distribution shifts smoothly and continuously through a series of intermediate concepts over time, such that each intermediate state differs marginally from its predecessor while accumulating substantial drift over time.
    • Reoccurring Concepts: A concept that was active in a prior time window temporarily disappears and later becomes active again, potentially reappearing via sudden, gradual, or incremental shifts.
  3. Knowl 3 — Four-Stage Concept Drift Detection Framework

    model/method

    A general concept drift detection pipeline operates across four successive functional stages to identify distributional change points or change intervals in streaming data:

    1. Stage 1 (Data Retrieval): Collects data chunks from continuous streams via windowing mechanisms. Standard windowing strategies include:

      • Landmark window: The starting timestamp is fixed while the endpoint extends as new instances arrive.
      • Sliding window: A fixed-capacity window advances across time, discarding oldest instances as new ones arrive.
      • Adaptive window: Window capacity expands or contracts dynamically based on detected rates of distributional change.
    2. Stage 2 (Data Modeling - Optional): Transforms and abstracts retrieved data instances into lower-dimensional representations or statistical summaries (e.g., through Principal Component Analysis, decision tree structures, or kernel density estimators) to reduce storage and computational overhead.

    3. Stage 3 (Test Statistics Calculation): Computes a quantitative dissimilarity or divergence statistic θ\theta between historical data WhistW_{\text{hist}} and new data WnewW_{\text{new}}. Common metrics include differences in online classification error rates, statistical distance functions (e.g., Kullback-Leibler divergence, Total Variation distance, or Relativized Discrepancy), and density difference estimators.

    4. Stage 4 (Hypothesis Testing / Statistical Bounds): Evaluates whether the test statistic θ\theta computed in Stage 3 exhibits statistical significance against a null hypothesis of stationarity (H0:Phist=PnewH_0: P_{\text{hist}} = P_{\text{new}}). This stage establishes decision thresholds using distribution estimation (e.g., Gaussian approximations, EWMA control limits), non-parametric permutation/bootstrapping tests, or concentration inequalities (e.g., Hoeffding's inequality, McDiarmid's bounds).

  4. Knowl 4 — Categorization of Concept Drift Detection Methods

    model/method

    Concept drift detection algorithms are categorized into three primary classes based on the nature of their test statistics and operational architecture:

    1. Error Rate-Based Detection: Tracks the online classification error rate of a base learner. A statistically significant increase in error rate triggers a drift alert. Prominent examples include:

      • Drift Detection Method (DDM): Monitors online binomial error rate using the standard deviation bound pi+sip_i + s_i; triggers a warning level at pi+si≥pmin⁡+2smin⁡p_i + s_i \ge p_{\min} + 2s_{\min} and a drift level at pi+si≥pmin⁡+3smin⁡p_i + s_i \ge p_{\min} + 3s_{\min}.
      • Early Drift Detection Method (EDDM): Monitors the distance between consecutive classification errors to enhance sensitivity to slow/gradual drift.
      • Exponentially Weighted Moving Average (ECDD): Applies dynamic-mean EWMA control charts to online error rates.
      • Statistical Test of Equal Proportions (STEPD): Compares error proportions between a recent window and the cumulative historical window via standard normal approximation.
      • Adaptive Windowing (ADWIN): Dynamically adjusts window size by evaluating differences in sub-window means via Hoeffding bounds.
    2. Data Distribution-Based Detection: Directly evaluates statistical divergence between input distributions of historical and new sliding windows without relying solely on learner error:

      • kdqTree: Partitions multidimensional space using tree structures and applies Kullback-Leibler divergence combined with bootstrap hypothesis testing.
      • Relativized Discrepancy (RD): Computes Total Variation-inspired discrepancies over data stream distributions with VC-dimension bounds.
      • Least Squares Density Difference (LSDD-CDT / LSDD-INC): Employs direct density difference estimation across sliding windows.
      • Local Drift Degree (LDD-DSDA): Assesses localized density discrepancies across kk-nearest neighbor neighborhoods.
    3. Multiple Hypothesis Test Detection: Deploys multiple statistical tests to monitor multi-dimensional or multi-aspect changes:

      • Parallel Architecture: Runs concurrent hypothesis tests across individual features or metric components (e.g., Just-In-Time (JIT) classifiers using CI-CUSUM over PCA components, Linear Four Rate (LFR) monitoring TP, TN, FP, and FN rates, or Three-layer detection monitoring P(y)P(y), P(X)P(X), and P(X∣y)P(X|y)).
      • Hierarchical Architecture: Employs a low-complexity, low-delay detection layer to flag candidate drifts, followed by an activated validation layer (e.g., permutation test, two-dimensional Kolmogorov-Smirnov test) to confirm the drift and prevent false alarms (e.g., Hierarchical Change-Detection Tests (HCDTs), HLFR, HHT-CU, HHT-AG).
  5. Knowl 5 — Concept Drift Understanding: Diagnostic Dimensions (When, How, and Where)

    model/method

    Concept drift understanding extends standard detection by diagnosing the operational characteristics of the detected drift along three specific dimensions to inform downstream adaptation:

    1. When (Temporal Characterization): Determines the exact temporal properties of the drift, identifying:

      • Start time (tstartt_{\text{start}})
      • Detection alarm time (talarmt_{\text{alarm}})
      • Transition duration / change period
      • Stabilization / end time (tendt_{\text{end}}) Two-stage alert systems (warning and drift levels) define a transition window where instances accumulated between the warning and drift alarms serve as the initial training set for adapting the model.
    2. How (Severity Quantification): Quantifies the magnitude of distributional change using a distance or discrepancy function δ\delta: Δ=δ(Pt(X,y),Pt+1(X,y))\Delta = \delta(P_t(X, y), P_{t+1}(X, y)) where Δ≥0\Delta \ge 0 measures the severity of the drift. Severe drift indicates substantial decision boundary shifts requiring complete model retraining, whereas minor drift indicates localized boundary adjustments amenable to incremental parameter tuning.

    3. Where (Spatial/Regional Localization): Identifies the sub-regions in the feature space XX where distributional conflict occurs: {x∈X∣Pt(x,y)≠Pt+1(x,y)}\{x \in X \mid P_t(x, y) \neq P_{t+1}(x, y)\} Spatial localization techniques include tracking inner-node error rates in decision trees, computing Kulldorff's spatial scan statistics over kdqkdq-tree partitions, measuring competence model hypersphere distances, or calculating local drift degrees via nearest-neighbor density variations. Identifying stable versus drifting regions allows adaptation algorithms to retain valid historical sub-models for unaffected feature regions while updating only the affected sub-models.

  6. Knowl 6 — Concept Drift Adaptation Framework and Methodologies

    model/method

    Concept drift adaptation defines the mechanisms by which predictive learning systems modify their internal structures in response to identified distribution changes. Adaptation methods are grouped into three primary paradigms:

    1. Global Model Retraining: Retrains a new base model on the most recent data window to replace an obsolete model once drift is confirmed. Window management strategies (e.g., ADWIN cut-offs) eliminate historical data. Dual-model architectures (e.g., Paired Learners) run a conservative stable learner alongside an aggressive reactive learner, substituting the stable learner with the reactive learner when the latter demonstrates superior performance.

    2. Adaptive Ensemble Methods: Maintains a collection of heterogeneous or homogeneous base classifiers combined through adaptive voting rules:

      • Dynamic Weighting & Pruning: Adjusts base model weights according to recent batch/window accuracy and prunes underperforming classifiers (e.g., Dynamic Weighted Majority (DWM), Learn++.NSE, Accuracy Update Ensemble (AUE2)).
      • Online Ensemble Generation: Trains new classifiers on recent data upon drift warnings and incorporates them into the ensemble (e.g., Leveraging Bagging with ADWIN, Adaptive Random Forest (ARF)).
      • Memory and Concept Repositories: Stores historical base models to reactivate them when recurring concepts are recognized, bypassing the need to retrain from scratch.
    3. Regional Model Adjustment: Updates only the local sub-structures of a model corresponding to the feature space regions exhibiting drift, retaining unaffected sub-structures for unchanged regions. In tree-based stream learning (e.g., Concept-adapting Very Fast Decision Trees (CVFDT), VFDTc, IADEM-3), alternative sub-trees are grown concurrently for drifting nodes and swapped in when their accuracy surpasses the original branch, leaving the remaining tree intact.

  7. Knowl 7 — Summary of Drift Detection Algorithms by Framework Stage and Understanding Capabilities

    data/table

    The following table synthesizes the operational configurations of prominent concept drift detection algorithms across the four stages of the general detection framework, alongside their capabilities in diagnosing the understanding dimensions (When, How, Where):

    Category Algorithm Stage 1 (Retrieval) Stage 2 (Modeling) Stage 3 (Statistic) Stage 4 (Test) When How Where
    Error rate DDM Landmark Learner Online error rate Distribution estimation Yes No No
    Error rate EDDM Landmark Learner Online error rate Distribution estimation Yes No No
    Error rate FW-DDM Landmark Learner Online error rate Distribution estimation Yes No No
    Error rate DEML Landmark Learner Online error rate Distribution estimation Yes No No
    Error rate STEPD Predefined whist,wneww_{\text{hist}}, w_{\text{new}} Learner Error rate difference Distribution estimation Yes No No
    Error rate ADWIN Auto cut whist,wneww_{\text{hist}}, w_{\text{new}} Learner Error rate difference Hoeffding's Bound Yes No No
    Error rate ECDD Landmark Learner Online error rate EWMA Chart Yes No No
    Error rate HDDM Landmark Learner Online error rate Hoeffding's Bound Yes No No
    Error rate LLDD Landmark / sliding window Decision trees Tree node error rate Hoeffding's Bound Yes No Yes
    Data distribution kdqTree Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} kdqTree KL divergence Bootstrapping Yes Yes Yes
    Data distribution CM Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} Competence model Competence distance Permutation test Yes Yes Yes
    Data distribution RD Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} KS structure Relativized Discrepancy VC-Dimension Yes Yes No
    Data distribution SCD Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} Kernel density estimator Log-likelihood Distribution estimation Yes Yes No
    Data distribution EDE Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} Nearest neighbor Density scale Permutation test Yes No No
    Data distribution SyncStream Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} PCA P-Tree Wilcoxon test Yes Yes No
    Data distribution PCA-CD Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} PCA Change-Score Page-Hinkley test Yes Yes No
    Data distribution LSDD-CDT Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} Learner Relative difference Distribution estimation Yes No No
    Data distribution LSDD-INC Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} Learner Relative difference Distribution estimation Yes No No
    Data distribution LDD-DSDA Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} kk-nearest neighbor Local drift degree Distribution estimation Yes Yes Yes
    Multiple Hyp. JIT Landmark Selected features 4 configurations Distribution estimation Yes No No
    Multiple Hyp. LFR Landmark Learner TP, TN, FP, FN Distribution estimation Yes No No
    Multiple Hyp. Three-layer Sliding whist,wneww_{\text{hist}}, w_{\text{new}} Learner P(y),P(X),P(X∣y)P(y), P(X), P(X|y) Distribution estimation Yes No No
    Multiple Hyp. TSMSD-EWMA Landmark Learner Online error rate EWMA Chart Yes No No
    Multiple Hyp. HCDTs Landmark Layer-dependent Layer-dependent Layer-dependent Yes No No
    Multiple Hyp. HLFR Landmark Learner TP, TN, FP, FN Distribution estimation Yes No No
    Multiple Hyp. HHT-CU Landmark Learner Classification uncertainty Hoeffding / Permutation Yes No No
    Multiple Hyp. HHT-AG Fixed whistw_{\text{hist}}, Sliding wneww_{\text{new}} N/A KS statistic 1D KS / 2D KS test Yes No No
  8. Knowl 8 — Validation Methodologies and Specialized Performance Metrics for Drifting Data Streams

    model/method

    Evaluating predictive algorithms on non-stationary data streams requires dedicated validation protocols and metrics designed for temporal dependencies and concept drift:

    Validation Methodologies:

    1. Holdout: Evaluates performance against an independent test set matching the distribution at timestamp tt; primarily applicable to synthetic benchmarks with known ground-truth drift schedules.
    2. Prequential Evaluation (Interleaved Test-Then-Train): Each incoming streaming instance is first utilized as an unseen test sample to assess model performance and compute loss f(y^t,yt)f(\hat{y}_t, y_t), and is subsequently used to update/train the model. Cumulative prequential error is given by: S=∑t=1nf(y^t,yt)S = \sum_{t=1}^n f(\hat{y}_t, y_t) Prequential error rates are tracked via landmark windows, sliding windows, or exponential fading factors.
    3. Controlled Permutation: Evaluates models on localized permutations of data order to ensure temporal robustness without destroying the underlying continuous distribution shift.

    Specialized Evaluation Metrics:

    • Cohen's Kappa Statistic (κ\kappa): Corrects classification accuracy for class imbalance: κ=p−pran1−pran\kappa = \frac{p - p_{\text{ran}}}{1 - p_{\text{ran}}} where pp is the classifier accuracy and pranp_{\text{ran}} is the accuracy of a chance classifier.
    • Kappa-Temporal Statistic (κper\kappa_{\text{per}}): Corrects accuracy for temporal autocorrelation: κper=p−pper1−pper\kappa_{\text{per}} = \frac{p - p_{\text{per}}}{1 - p_{\text{per}}} where pperp_{\text{per}} is the accuracy of a persistent classifier predicting the most recently observed true label.
    • Combined Kappa Statistic (κ+\kappa^+): Takes the geometric mean of non-negative kappa metrics: κ+=max⁡(0,κ)max⁡(0,κper)\kappa^+ = \sqrt{\max(0, \kappa) \max(0, \kappa_{\text{per}})}
    • Detector Accuracy Metrics: Evaluated via True Detection Rate, False Alarm Rate, Miss Detection Rate, and Detection Delay (tdetected−tactualt_{\text{detected}} - t_{\text{actual}}).
  9. Knowl 9 — Statistical Hypothesis Tests for Algorithm Comparison under Concept Drift

    model/method

    Statistical hypothesis testing verifies whether observed performance differences between predictive models on non-stationary data streams are statistically significant:

    1. McNemar's Test: Compares two classification models evaluated on NN paired instances. Let aa denote instances misclassified by Model 1 but correctly classified by Model 2, and bb denote instances misclassified by Model 2 but correctly classified by Model 1. The test statistic is computed as: M=sign(a−b)(a−b)2a+bM = \text{sign}(a - b) \frac{(a - b)^2}{a + b} which asymptotically follows a χ2\chi^2 distribution with 1 degree of freedom.

    2. Sign Test: Evaluates two models across NN independent observations or benchmark points. Letting BB denote the number of instances where Model 2 outperforms Model 1, and TT denote the number of ties, the one-sided exact binomial pp-value is computed as: p=∑k=BN−T(N−Tk)(0.5)k(0.5)N−T−kp = \sum_{k=B}^{N-T} \binom{N-T}{k} (0.5)^k (0.5)^{N-T-k}

    3. Wilcoxon Signed-Rank Test: Evaluates paired performance metrics xi,1x_{i,1} and xi,2x_{i,2} across NN datasets. For non-zero differences Nr=N−TN_r = N - T, differences ∣xi,1−xi,2∣|x_{i,1} - x_{i,2}| are ranked as RiR_i. The test statistic is: W=∑i=1Nr(sign(xi,1−xi,2)×Ri)W = \sum_{i=1}^{N_r} \left(\text{sign}(x_{i,1} - x_{i,2}) \times R_i\right) Equivalence is rejected if ∣W∣>Wcritical,Nr|W| > W_{\text{critical}, N_r}.

    4. Nemenyi Post-hoc Test: Compares multiple learning algorithms across multiple non-stationary datasets. Two classifiers among kk total algorithms evaluated across NN datasets perform significantly differently if their average rank difference exceeds the Critical Difference (CDCD): CD=qαk(k+1)6NCD = q_\alpha \sqrt{\frac{k(k + 1)}{6N}} where qαq_\alpha represents the critical value derived from the Studentized range distribution divided by 2\sqrt{2}.

  10. Knowl 10 — Synthetic and Real-World Benchmark Datasets for Concept Drift

    data/table

    Evaluation of concept drift detection and adaptation methodologies relies on both parameterized synthetic data generators (with controlled drift injection) and real-world non-stationary benchmark datasets:

    Dataset Type / Source #Instances #Attributes #Classes Drift Characteristic
    STAGGER Synthetic / Source II Custom 3 2 Sudden
    SEA Synthetic / Source II Custom 3 2 Sudden
    Rotating Hyperplane Synthetic / Source II Custom 10 2 Gradual, Incremental
    Random RBF Synthetic / Source III Custom Custom Custom Sudden, Gradual, Incremental
    Random Tree Synthetic / Source II Custom Custom Custom Sudden, Reoccurring
    LED Synthetic / Source II Custom 24 10 Sudden
    Waveform Synthetic / Source II Custom 40 3 Sudden
    Sine Synthetic / Source II Custom 2 2 Sudden
    Circle Synthetic / Source III Custom 2 2 Gradual
    Rotating Chessboard Synthetic / Source II Custom 2 2 Gradual
    Airlines Real-world 539,384 7 2 Yearly temporal drift
    Covertype Real-world 581,012 54 7 Geographical / spatial drift
    Electricity Real-world 45,312 8 2 Seasonal / market drift
    Poker-Hand Real-world 1,025,010 10 10 Categorical drift
    NOAA Weather Real-world 18,159 8 2 Yearly weather drift
    Sensor Real-world 2,219,803 5 54 Daily stream drift
    KDDCup'99 Real-world 494,021 41 23 Network intrusions, novel classes
    Usenet (1 2) Real-world 1,500 99 2 Recurrent textual drift
    Email / Spam Real-world 1,500 / 9,324 913 / 499 2 Sudden textual drift
    Spam Assassin Real-world 9,324 39,916 2 Gradual spam drift
    ECUE (1 2) Real-world 10,983 / 11,905 287,034 / 166,047 2 Email drift, novel classes

    Benchmark Limitations: While synthetic datasets permit precise calculation of true detection rates and detection delays, real-world data streams typically lack precise ground truth regarding drift onset, transition duration, and region boundaries, often containing confounding mixtures of drift types.

Coverage note — None was omitted; the complete framework of concept drift definition, detection paradigms, understanding dimensions (when, how, where), adaptation strategies, evaluation protocols/metrics, statistical tests, benchmark datasets, and cross-cutting intersections is fully covered.

References

  1. 1.G. Widmer and M. Kubat, “Learning in the presence of concept drift and hidden contexts,” Machine Learning, vol. 23, no. 1, pp. 69–101, 1996.
  2. 2.N. Lu, J. Lu, G. Zhang, and R. Lopez de Mantaras, “A concept drift-tolerant case-base editing technique,” Artif. Intell., vol. 230, pp. 108–133, 2016.
  3. 3.N. Lu, G. Zhang, and J. Lu, “Concept drift detection via competence models,” Artif. Intell., vol. 209, pp. 11–28, 2014.
  4. 4.A. Liu, Y. Song, G. Zhang, and J. Lu, “Regional concept drift detection and density synchronized drift adaptation,” in Proc. 26th Int. Joint Conf. Artificial Intelligence. Accept, 2017, Conference Proceedings.
  5. 5.A. Liu, G. Zhang, and J. Lu, “Fuzzy time windowing for gradual concept drift adaptation,” in Proc. 26th IEEE Int. Conf. Fuzzy Systems. IEEE, 2017, Conference Proceedings.
  6. 6.B. Krawczyk, L. L. Minku, J. Gama, J. Stefanowski, and M. Woźniak, “Ensemble learning for data stream analysis: A survey,” Information Fusion, vol. 37, pp. 132–156, 2017.
  7. 7.S. Ramírez-Gallego, B. Krawczyk, S. García, M. Woźniak, and F. Herrera, “A survey on data preprocessing for data stream mining: Current status and future directions,” Neurocomputing, vol. 239, pp. 39–57, 2017.
  8. 8.J. Gama, I. Žliobaitė, A. Bifet, M. Pechenizkiy, and A. Bouchachia, “A survey on concept drift adaptation,” ACM Comput. Surv., vol. 46, no. 4, pp. 1–37, 2014.
  9. 9.G. Ditzler, M. Roveri, C. Alippi, and R. Polikar, “Learning in nonstationary environments: a survey,” IEEE Comput. Intell. Mag., vol. 10, no. 4, pp. 12–25, 2015.
  10. 10.J. Gama, “A survey on learning from data streams: current and future trends,” Progress in Artificial Intelligence, vol. 1, no. 1, pp. 45–55, 2012.
  11. 11.J. A. Silva, E. R. Faria, R. C. Barros, E. R. Hruschka, A. C. P. L. F. d. Carvalho, and J. Gama, “Data stream clustering: A survey,” ACM Comput. Surv., vol. 46, no. 1, pp. 1–31, 2013.
  12. 12.J. C. Schlimmer and R. H. Granger Jr, “Incremental learning from noisy data,” Machine learning, vol. 1, no. 3, pp. 317–354, 1986.
  13. 13.V. Losing, B. Hammer, and H. Wersing, “Knn classifier with self adjusting memory for heterogeneous concept drift,” in Proc. 16th Int. Conf. Data Mining, 2016, Conference Proceedings, pp. 291–300.
  14. 14.I. Žliobaitė and J. Hollmén, “Optimizing regression models for data streams with missing values,” Machine Learning, vol. 99, no. 1, pp. 47–73, 2014.
  15. 15.S. Amos, “When training and test sets are different: characterizing learning transfer,” Dataset Shift in Machine Learning, pp. 3–28, 2009.
  16. 16.J. G. Moreno-Torres, T. Raeder, R. Alaiz-Rodríguez, N. V. Chawla, and F. Herrera, “A unifying view on dataset shift in classification,” Pattern Recognit., vol. 45, no. 1, pp. 521–530, 2012.
  17. 17.M. Basseville and I. V. Nikiforov, Detection of abrupt changes: theory and application. Prentice Hall Englewood Cliffs, 1993, vol. 104.
  18. 18.A. Dries and U. Rückert, “Adaptive concept drift detection,” Statistical Analysis and Data Mining: The ASA Data Science Journal, vol. 2, no. 5–6, pp. 311–327, 2009.
  19. 19.C. Alippi and M. Roveri, “Just-in-time adaptive classifiers part i: Detecting nonstationary changes,” IEEE Trans. Neural Networks, vol. 19, no. 7, pp. 1145–1153, 2008.
  20. 20.J. Gama, P. Medas, G. Castillo, and P. Rodrigues, “Learning with drift detection,” in Proc. 17th Brazilian Symp. Artificial Intelligence, ser. Lecture Notes in Computer Science. Springer, 2004, Book Section, pp. 286–295.
  21. 21.L. Bu, C. Alippi, and D. Zhao, “A pdf-free change detection test based on density difference estimation,” IEEE Trans. Neural Networks Learn. Syst., vol. PP, no. 99, pp. 1–11, 2016.
  22. 22.T. Dasu, S. Krishnan, S. Venkatasubramanian, and K. Yi, “An information-theoretic approach to detecting changes in multi-dimensional data streams,” in Proc. Symp. the Interface of Statistics, Computing Science, and Applications. Citeseer, 2006, Conference Proceedings, pp. 1–24.
  23. 23.I. Frias-Blanco, J. d. Campo-Avila, G. Ramos-Jimenez, R. Morales-Bueno, A. Ortiz-Diaz, and Y. Caballero-Mota, “Online and non-parametric drift detection methods based on hoeffding’s bounds,” IEEE Trans. Knowl. Data Eng., vol. 27, no. 3, pp. 810–823, 2015.
  24. 24.M. Yamada, A. Kimura, F. Naya, and H. Sawada, “Change-point detection with feature selection in high-dimensional time-series data,” in Proc. 23rd Int. Joint Conf. Artificial Intelligence, 2013, Conference Proceedings, pp. 1827–1833.
  25. 25.J. Gama and G. Castillo, “Learning with local drift detection,” in Proc. 2nd Int. Conf. Advanced Data Mining and Applications. Springer, 2006, Conference Proceedings, pp. 42–55.
  26. 26.M. Baena-García, J. del Campo-Avila, R. Fidalgo, A. Bifet, R. Gavaldà, and R. Morales-Bueno, “Early drift detection method,” in Proc. 4th Int. Workshop Knowledge Discovery from Data Streams, 2006, Conference Paper.
  27. 27.S. Xu and J. Wang, “Dynamic extreme learning machine for data stream classification,” Neurocomputing, vol. 238, pp. 433–449, 2017.
  28. 28.G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew, “Extreme learning machine: Theory and applications,” Neurocomputing, vol. 70, no. 1–3, pp. 489–501, 2006.
  29. 29.G. J. Ross, N. M. Adams, D. K. Tasoulis, and D. J. Hand, “Exponentially weighted moving average charts for detecting concept drift,” Pattern Recognit. Lett., vol. 33, no. 2, pp. 191–198, 2012.
  30. 30.K. Nishida and K. Yamauchi, “Detecting concept drift using statistical testing,” in Proc. 10th Int. Conf. Discovery Science, V. Corruble, M. Takeda, and E. Suzuki, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2007, Conference Proceedings, pp. 264–269.
  31. 31.A. Bifet and R. Gavaldà, “Learning from time-changing data with adaptive windowing,” in Proc. 2007 SIAM Int. Conf. Data Mining, vol. 7. SIAM, 2007, Conference Proceedings, p. 2007.
  32. 32.——, “Adaptive learning from evolving data streams,” in Proc. 8th Int. Symp. Intelligent Data Analysis. Springer, 2009, Conference Proceedings, pp. 249–260.
  33. 33.A. Bifet, G. Holmes, B. Pfahringer, and R. Gavaldà, “Improving adaptive bagging methods for evolving data streams,” in Proc. 1st Asian Conf. Machine Learning, ser. Lecture Notes in Computer Science, Z.-H. Zhou and T. Washio, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2009, Book Section, pp. 23–37.
  34. 34.A. Bifet, G. Holmes, B. Pfahringer, R. Kirkby, and R. Gavaldà, “New ensemble methods for evolving data streams,” in Proc. 15th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. ACM, 2009, Conference Proceedings, pp. 139–148.
  35. 35.H. M. Gomes, A. Bifet, J. Read, J. P. Barddal, F. Enembreck, B. Pfharinger, G. Holmes, and T. Abdessalem, “Adaptive random forests for evolving data stream classification,” Machine Learning, 2017.
  36. 36.J. Shao, Z. Ahmadi, and S. Kramer, “Prototype-based learning on concept-drifting data streams,” in Proc. 20th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. 2623609: ACM, 2014, Conference Proceedings, pp. 412–421.
  37. 37.D. Kifer, S. Ben-David, and J. Gehrke, “Detecting change in data streams,” in Proc. 30th Int. Conf. Very Large Databases, vol. 30. VLDB Endowment, 2004, Conference Proceedings, pp. 180–191.
  38. 38.X. Song, M. Wu, C. Jermaine, and S. Ranka, “Statistical change detection for multi-dimensional data,” in Proc. 13th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. San Jose, California, USA: ACM, 2007, Conference Paper, pp. 667–676.
  39. 39.A. A. Qahtan, B. Alharbi, S. Wang, and X. Zhang, “A pca-based change detection framework for multidimensional data streams,” in Proc. 21th Int. Conf. on Knowledge Discovery and Data Mining. ACM, 2015, Conference Proceedings, pp. 935–944.
  40. 40.F. Gu, G. Zhang, J. Lu, and C.-T. Lin, “Concept drift detection based on equal density estimation,” in Proc. 2016 Int. Joint Conf. Neural Networks. IEEE, 2016, Conference Proceedings, pp. 24–30.
  41. 41.L. Bu, D. Zhao, and C. Alippi, “An incremental change detection test based on density difference estimation,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. PP, no. 99, pp. 1–13, 2017.
  42. 42.C. Alippi and M. Roveri, “Just-in-time adaptive classifiers part ii: designing the classifier,” IEEE Trans. Neural Networks, vol. 19, no. 12, pp. 2053–2064, 2008.
  43. 43.C. Alippi, G. Boracchi, and M. Roveri, “A just-in-time adaptive classification system based on the intersection of confidence intervals rule,” Neural Networks, vol. 24, no. 8, pp. 791–800, 2011.
  44. 44.——, “Just-in-time ensemble of classifiers,” in Proc. 2012 Int. Joint Conf. Neural Networks. IEEE, 2012, Conference Proceedings, pp. 1–8.
  45. 45.——, “Just-in-time classifiers for recurrent concepts,” IEEE Trans. Neural Networks Learn. Syst., vol. 24, no. 4, pp. 620–634, 2013.
  46. 46.W. Heng and Z. Abraham, “Concept drift detection for streaming data,” in Proc. 2015 Int. Joint Conf. Neural Networks, 2015, Conference Proceedings, pp. 1–9.
  47. 47.Y. Zhang, G. Chu, P. Li, X. Hu, and X. Wu, “Three-layer concept drifting detection in text data streams,” Neurocomputing, vol. 260, pp. 393–403, 2017.
  48. 48.L. Du, Q. Song, L. Zhu, and X. Zhu, “A selective detector ensemble for concept drift detection,” The Computer Journal, vol. 58, no. 3, pp. 457–471, 2014.
  49. 49.B. I. F. Maciel, S. G. T. C. Santos, and R. S. M. Barros, “A lightweight concept drift detection ensemble,” in Proc. 27th IEEE Int. Conf. on Tools with Artificial Intelligence. IEEE, 2015, pp. 1061–1068.
  50. 50.C. Alippi, G. Boracchi, and M. Roveri, “Hierarchical change-detection tests,” IEEE Trans. Neural Networks Learn. Syst., vol. 28, no. 2, pp. 246–258, 2017.
  51. 51.S. Yu and Z. Abraham, “Concept drift detection with hierarchical hypothesis testing,” in Proc. 2017 SIAM Int. Conf. Data Mining. SIAM, 2017, Conference Proceedings, pp. 768–776.
  52. 52.H. Raza, G. Prasad, and Y. Li, “Ewma model based shift-detection methods for detecting covariate shifts in non-stationary environments,” Pattern Recognit., vol. 48, no. 3, pp. 659–669, 2015.
  53. 53.S. Yu, X. Wang, and J. C. Principe, “Request-and-reverify: Hierarchical hypothesis testing for concept drift detection with expensive labels,” arXiv preprint arXiv:1806.10131, 2018.
  54. 54.P. M. Gonçalves Jr, S. G. de Carvalho Santos, R. S. Barros, and D. C. Vieira, “A comparative study on concept drift detectors,” Expert Systems with Applications, vol. 41, no. 18, pp. 8144–8156, 2014.
  55. 55.F. Pukelsheim, “The three sigma rule,” The American Statistician, vol. 48, no. 2, pp. 88–91, 1994.
  56. 56.A. Tsymbal, M. Pechenizkiy, P. Cunningham, and S. Puuronen, “Dynamic integration of classifiers for handling concept drift,” Information Fusion, vol. 9, no. 1, pp. 56–68, 2008.
  57. 57.S. H. Bach and M. Maloof, “Paired learners for concept drift,” in Proc. 8th Int. Conf. Data Mining, 2008, Conference Proceedings, pp. 23–32.
  58. 58.D. Liu, Y. Wu, and H. Jiang, “Fp-elm: An online sequential learning algorithm for dealing with concept drift,” Neurocomputing, vol. 207, pp. 322–334, 2016.
  59. 59.D. Han, C. Giraud-Carrier, and S. Li, “Efficient mining of high-speed uncertain data streams,” Applied Intelligence, vol. 43, no. 4, pp. 773–785, 2015.
  60. 60.S. G. Soares and R. Araújo, “An adaptive ensemble of on-line extreme learning machines with variable forgetting factor for dynamic system prediction,” Neurocomputing, vol. 171, pp. 693–707, 2016.
  61. 61.B. F. J. Manly and D. Mackenzie, “A cumulative sum type of method for environmental monitoring,” Environmetrics, vol. 11, no. 2, pp. 151–166, 2000.
  62. 62.N. C. Oza and S. Russell, “Experimental comparisons of online and batch versions of bagging and boosting,” in Proc. 7th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. 502565: ACM, 2001, Conference Proceedings, pp. 359–364.
  63. 63.A. Bifet, G. Holmes, and B. Pfahringer, “Leveraging bagging for evolving data streams,” in Proc. 2010 Joint European Conf. Machine Learning and Knowledge Discovery in Databases. Springer, 2010, Conference Proceedings, pp. 135–150.
  64. 64.F. Chu and C. Zaniolo, “Fast and light boosting for adaptive mining of data streams,” in Proc. 8th Pacific-Asia Conf. Knowledge Discovery and Data Mining, H. Dai, R. Srikant, and C. Zhang, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004, Book Section, pp. 282–292.
  65. 65.P. Li, X. Wu, X. Hu, and H. Wang, “Learning concept-drifting data streams with random ensemble decision trees,” Neurocomputing, vol. 166, pp. 68–83, 2015.
  66. 66.J. Z. Kolter and M. A. Maloof, “Dynamic weighted majority: An ensemble method for drifting concepts,” Journal of Machine Learning Research, 2007.
  67. 67.R. Elwell and R. Polikar, “Incremental learning of concept drift in nonstationary environments,” IEEE Trans. Neural Networks, vol. 22, no. 10, pp. 1517–31, 2011.
  68. 68.X.-C. Yin, K. Huang, and H.-W. Hao, “De2: Dynamic ensemble of ensembles for learning nonstationary data,” Neurocomputing, vol. 165, pp. 14–22, 2015.
  69. 69.P. Zhang, J. Li, P. Wang, B. J. Gao, X. Zhu, and L. Guo, “Enabling fast prediction for ensemble models on data streams,” in Proc. 17th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. San Diego, California, USA: ACM, 2011, Conference Paper, pp. 177–185.
  70. 70.Y. Xu, R. Xu, W. Yan, and P. Ardis, “Concept drift learning with alternating learners,” in Proc. 2017 Int. Joint Conf. Neural Networks, 2017, Conference Proceedings, pp. 2104–2111.
  71. 71.L. Pietruczuk, L. Rutkowski, M. Jaworski, and P. Duda, “A method for automatic adjustment of ensemble size in stream data mining,” in Proc. 2016 Int. Joint Conf. Neural Networks, 2016, Conference Proceedings, pp. 9–15.
  72. 72.S.-C. You and H.-T. Lin, “A simple unlearning framework for online learning under concept drifts,” in Proc. 20th Pacific-Asia Conf. Knowledge Discovery and Data Mining. Springer, 2016, Conference Proceedings, pp. 115–126.
  73. 73.D. Brzezinski and J. Stefanowski, “Reacting to different types of concept drift: The accuracy updated ensemble algorithm,” IEEE Trans. Neural Networks Learn. Syst., vol. 25, no. 1, pp. 81–94, 2014.
  74. 74.P. Zhang, X. Zhu, and Y. Shi, “Categorizing and mining concept drifting data streams,” in Proc. 14th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. Las Vegas, Nevada, USA: ACM, 2008, Conference Paper, pp. 812–820.
  75. 75.Y. Sun, K. Tang, L. L. Minku, S. Wang, and X. Yao, “Online ensemble learning of data streams with gradually evolved classes,” IEEE Trans. Knowl. Data Eng., vol. 28, no. 6, pp. 1532–1545, 2016.
  76. 76.J. Gama and P. Kosina, “Recurrent concepts in data streams classification,” Knowledge and Information Systems, vol. 40, no. 3, pp. 489–507, 2013.
  77. 77.J. B. Gomes, M. M. Gaber, P. A. Sousa, and E. Menasalvas, “Mining recurring concepts in a dynamic feature space,” IEEE Trans. Neural Networks Learn. Syst., vol. 25, no. 1, pp. 95–110, 2014.
  78. 78.Z. Ahmadi and S. Kramer, “Modeling recurring concepts in data streams: a graph-based framework,” Knowledge and Information Systems, 2017.
  79. 79.P. Domingos and G. Hulten, “Mining high-speed data streams,” in Proc. 6th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. ACM, 2000, Conference Proceedings, pp. 71–80.
  80. 80.G. Hulten, L. Spencer, and P. Domingos, “Mining time-changing data streams,” in Proc. 7th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. San Francisco, California: ACM, 2001, Conference Paper, pp. 97–106.
  81. 81.J. Gama, R. Rocha, and P. Medas, “Accurate decision trees for mining high-speed data streams,” in Proc. 9th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. ACM, 2003, Conference Proceedings, pp. 523–528.
  82. 82.H. Yang and S. Fong, “Incrementally optimized decision tree for noisy big data,” in Proc. 1st Int. Workshop Big Data, Streams and Heterogeneous Source Mining Algorithms, Systems, Programming Models and Applications. Beijing, China: ACM, 2012, Conference Paper, pp. 36–44.
  83. 83.——, “Countering the concept-drift problems in big data by an incrementally optimized stream mining model,” Journal of Systems and Software, vol. 102, pp. 158–166, 2015.
  84. 84.L. Rutkowski, M. Jaworski, L. Pietruczuk, and P. Duda, “Decision trees for mining data streams based on the gaussian approximation,” IEEE Trans. Knowl. Data Eng., vol. 26, no. 1, pp. 108–119, 2014.
  85. 85.L. Rutkowski, L. Pietruczuk, P. Duda, and M. Jaworski, “Decision trees for mining data streams based on the mcdiarmid’s bound,” IEEE Trans. Knowl. Data Eng., vol. 25, no. 6, pp. 1272–1279, 2013.
  86. 86.L. Rutkowski, M. Jaworski, L. Pietruczuk, and P. Duda, “A new method for data stream mining based on the misclassification error,” IEEE Trans. Neural Networks Learn. Syst., vol. 26, no. 5, pp. 1048–1059, 2015.
  87. 87.I. Frías-Blanco, J. d. Campo-Ávila, G. Ramos-Jiménez, A. C. P. L. F. Carvalho, A. Ortiz-Díaz, and R. Morales-Bueno, “Online adaptive decision trees based on concentration inequalities,” Knowledge-Based Systems, vol. 104, pp. 179–194, 2016.
  88. 88.J. Gama, R. Sebastiao, and P. P. Rodrigues, “On evaluating stream learning algorithms,” Machine Learning, vol. 90, no. 3, pp. 317–346, 2012.
  89. 89.I. Žliobaitė, “Controlled permutations for testing adaptive learning models,” Knowledge and Information Systems, vol. 39, no. 3, pp. 565–578, 2014.
  90. 90.A. Bifet, G. Holmes, B. Pfahringer, and E. Frank, “Fast perceptron decision tree learning from evolving data streams,” in Proc. 14th Pacific-Asia Conf. Knowledge Discovery and Data Mining, M. J. Zaki, J. X. Yu, B. Ravindran, and V. Pudi, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010, Book Section, pp. 299–310.
  91. 91.J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960.
  92. 92.I. Žliobaitė, A. Bifet, J. Read, B. Pfahringer, and G. Holmes, “Evaluation methods and decision theory for classification of streaming data with temporal dependence,” Machine Learning, vol. 98, no. 3, pp. 455–482, 2015.
  93. 93.D. Brzezinski and J. Stefanowski, “Prequential auc for classifier evaluation and drift detection in evolving data streams,” in Proc. 3rd Int. Workshop New Frontiers in Mining Complex Patterns, A. Appice, M. Ceci, C. Loglisci, G. Manco, E. Masciari, and Z. W. Ras, Eds. Cham: Springer International Publishing, 2014, Book Section, pp. 87–101.
  94. 94.A. Bifet, G. d. F. Morales, J. Read, G. Holmes, and B. Pfahringer, “Efficient online evaluation of big data stream classifiers,” in Proc. 21th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. Sydney, NSW, Australia: ACM, 2015, Conference Paper, pp. 59–68.
  95. 95.N. Japkowicz and M. Shah, Evaluating learning algorithms: a classification perspective. Cambridge University Press, 2011.
  96. 96.Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947.
  97. 97.J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” Journal of Machine Learning Research, vol. 7, no. Jan, pp. 1–30, 2006.
  98. 98.J. Z. Kolter and M. A. Maloof, “Using additive expert ensembles to cope with concept drift,” in Proc. 22nd Int. Conf. Machine Learning. Bonn, Germany: ACM, 2005, Conference Paper, pp. 449–456.
  99. 99.X. Wu, P. Li, and X. Hu, “Learning from concept drifting data streams with unlabeled data,” Neurocomputing, vol. 92, pp. 145–155, 2012.
  100. 100.W. N. Street and Y. Kim, “A streaming ensemble algorithm (sea) for large-scale classification,” in Proc. Seventh ACM Int. Conf. Knowledge Discovery and Data Mining. 502568: ACM, 2001, Conference Proceedings, pp. 377–382.
  101. 101.R. Fok, A. An, and X. Wang, “Mining evolving data streams with particle filters,” Comput. Intell., vol. 33, no. 2, pp. 147–180, 2017.
  102. 102.P. Kosina and J. Gama, “Very fast decision rules for classification in data streams,” Data Mining and Knowledge Discovery, vol. 29, no. 1, pp. 168–202, 2015.
  103. 103.H. Wang, W. Fan, P. S. Yu, and J. Han, “Mining concept-drifting data streams using ensemble classifiers,” in Proc. 9th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. Washington, D.C.: ACM, 2003, Conference Paper, pp. 226–235.
  104. 104.A. Bifet, G. Holmes, R. Kirkby, and B. Pfahringer, “Moa: Massive online analysis,” Journal of Machine Learning Research, vol. 99, pp. 1601–1604, 2010.
  105. 105.V. M. Souza, D. F. Silva, J. Gama, and G. E. Batista, “Data stream classification guided by clustering on nonstationary environments and extreme verification latency,” in Proceedings of the 2015 SIAM International Conference on Data Mining. SIAM, 2015, pp. 873–881.
  106. 106.M. Lichman, “UCI machine learning repository,” 2013. [Online]. Available: http://archive.ics.uci.edu/ml
  107. 107.M. Harel, S. Mannor, R. El-Yaniv, and K. Crammer, “Concept drift detection through resampling,” in Proc. 31st Int. Conf. Machine Learning, 2014, Conference Proceedings, pp. 1009–1017.
  108. 108.X. Zhu, “Stream data mining repository,” 2010. [Online]. Available: http://www.cse.fau.edu/~xqzhu/stream.html
  109. 109.M. Harries and N. S. Wales, “Splice-2 comparative evaluation: Electricity pricing,” 1999.
  110. 110.I. Katakis, G. Tsoumakas, and I. Vlahavas, “Tracking recurring contexts using ensemble classifiers: an application to email filtering,” Knowledge and Information Systems, vol. 22, no. 3, pp. 371–391, 2009.
  111. 111.I. Katakis, G. Tsoumakas, E. Banos, N. Bassiliades, and I. Vlahavas, “An adaptive personalized news dissemination system,” Journal of Intelligent Information Systems, vol. 32, no. 2, pp. 191–212, 2008.
  112. 112.I. Katakis, G. Tsoumakas, and I. P. Vlahavas, “An ensemble of classifiers for coping with recurring contexts in data streams,” in 18th European Conf. Artificial Intelligence, 2008, Conference Proceedings, pp. 763–764.
  113. 113.S. J. Delany, P. Cunningham, A. Tsymbal, and L. Coyle, “A case-based technique for tracking concept drift in spam filtering,” Knowledge-Based Systems, vol. 18, no. 4–5, pp. 187–195, 2005.
  114. 114.L.-Y. Wang, C. Park, K. Yeon, and H. Choi, “Tracking concept drift using a constrained penalized regression combiner,” Comput. Stat. Data Anal., vol. 108, pp. 52–69, 2017.
  115. 115.I. Žliobaitė, A. Bifet, B. Pfahringer, and G. Holmes, “Active learning with drifting streaming data,” IEEE Trans. Neural Networks Learn. Syst., vol. 25, no. 1, pp. 27–39, 2014.
  116. 116.G. Song, Y. Ye, H. Zhang, X. Xu, R. Y. K. Lau, and F. Liu, “Dynamic clustering forest: An ensemble framework to efficiently classify textual data stream with concept drift,” Information Sciences, vol. 357, pp. 125–143, 2016.
  117. 117.G. Ditzler and R. Polikar, “Incremental learning of concept drift from streaming imbalanced data,” IEEE Trans. Knowl. Data Eng., vol. 25, no. 10, pp. 2283–2301, 2013.
  118. 118.B. Mirza, Z. Lin, and N. Liu, “Ensemble of subset online sequential extreme learning machine for class imbalance and concept drift,” Neurocomputing, vol. 149, pp. 316–329, 2015.
  119. 119.B. Mirza and Z. Lin, “Meta-cognitive online sequential extreme learning machine for imbalanced and concept-drifting data classification,” Neural Networks, vol. 80, pp. 79–94, 2016.
  120. 120.S. Wang, L. L. Minku, and X. Yao, “Resampling-based ensemble methods for online class imbalance learning,” IEEE Trans. Knowl. Data Eng., vol. 27, no. 5, pp. 1356–1368, 2015.
  121. 121.E. Arabmakki and M. Kantardzic, “Som-based partial labeling of imbalanced data stream,” Neurocomputing, vol. 262, pp. 120–133, 2017.
  122. 122.A. Katal, M. Wazid, and R. H. Goudar, “Big data: Issues, challenges, tools and good practices,” in Proc. 6th Int. Conf. Contemporary Computing (IC3), 2013, Conference Proceedings, pp. 404–409.
  123. 123.A. Andrzejak and J. B. Gomes, “Parallel concept drift detection with online map-reduce,” in Proc. 12th Int. Conf. Data Mining Workshops, 2012, Conference Proceedings, pp. 402–407.
  124. 124.M. Tennant, F. Stahl, O. Rana, and J. B. Gomes, “Scalable real-time classification of data streams with concept drift,” Future Generation Computer Systems, vol. 75, pp. 187–199, 2017.
  125. 125.C. C. Aggarwal, J. Han, J. Wang, and P. S. Yu, “A framework for clustering evolving data streams,” in Proc. 29th Int. Conf. Very Large Databases, vol. 29. VLDB Endowment, 2003, Conference Proceedings, pp. 81–92.
  126. 126.X. Song, H. He, S. Niu, and J. Gao, “A data streams analysis strategy based on hoeffding tree with concept drift on hadoop system,” in Proc. 4th Int. Conf. Advanced Cloud and Big Data, 2016, Conference Proceedings, pp. 45–48.
  127. 127.V. Nguyen, T. D. Nguyen, T. Le, S. Venkatesh, and D. Phung, “One-pass logistic regression for label-drift and large-scale classification on distributed systems,” in Proc. 16th Int. Conf. Data Mining, 2016, Conference Proceedings, pp. 1113–1118.
  128. 128.W. Chu, M. Zinkevich, L. Li, A. Thomas, and B. Tseng, “Unbiased online active learning in data streams,” in Proc. 17th ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining. San Diego, California, USA: ACM, 2011, Conference Paper, pp. 195–203.
  129. 129.G. Ditzler and R. Polikar, “Semi-supervised learning in non-stationary environments,” in Proc. 2011 Int. Joint Conf. Neural Networks, 2011, Conference Proceedings, pp. 2741–2748.
  130. 130.M. J. Hosseini, A. Gholipour, and H. Beigy, “An ensemble of cluster-based classifiers for semi-supervised classification of non-stationary data streams,” Knowledge and Information Systems, vol. 46, no. 3, pp. 567–597, 2015.
  131. 131.P. Zhang, X. Zhu, J. Tan, and L. Guo, “Classifier and cluster ensembles for mining concept drifting data streams,” in Proc. 10th Int. Conf. Data Mining, 2010, Conference Proceedings, pp. 1175–1180.
  132. 132.S. Chandra, A. Haque, L. Khan, and C. Aggarwal, “An adaptive framework for multistream classification,” in Proc. 25th ACM Int. on Conf. Information and Knowledge Management. Indianapolis, Indiana, USA: ACM, 2016, Conference Paper, pp. 1181–1190.
  133. 133.A. Haque, L. Khan, M. Baron, B. Thuraisingham, and C. Aggarwal, “Efficient handling of concept drift and concept evolution over stream data,” in Proc. 32nd Int. Conf. Data Engineering, 2003, Conference Proceedings, pp. 481–492.
  134. 134.A. Haque, L. Khan, and M. Baron, “Sand: Semi-supervised adaptive novel class detection and classification over data stream,” in 30th AAAI Conf. Artificial Intelligence, 2016, Conference Proceedings, pp. 1652–1658.
  135. 135.T. Le, F. Stahl, M. M. Gaber, J. B. Gomes, and G. D. Fatta, “On expressiveness and uncertainty awareness in rule-based classification for data streams,” Neurocomputing, vol. 265, pp. 127–141, 2017.
  136. 136.J. Cendrowska, “Prism: An algorithm for inducing modular rules,” Int. J. Man Mach. Stud., vol. 27, no. 4, pp. 349–370, 1987.
  137. 137.M. Pratama, S. G. Anavatti, M. Joo, and E. D. Lughofer, “pclass: An effective classifier for streaming examples,” IEEE Trans. Fuzzy Syst., vol. 23, no. 2, pp. 369–386, 2015.
  138. 138.Y.-R. Yeh and Y.-C. F. Wang, “A rank-one update method for least squares linear discriminant analysis with concept drift,” Pattern Recognit., vol. 46, no. 5, pp. 1267–1276, 2013.
  139. 139.R. C. Cavalcante, L. L. Minku, and A. L. I. Oliveira, “Fedd: Feature extraction for explicit concept drift detection in time series,” in Proc. 2016 Int. Joint Conf. Neural Networks, 2016, Conference Proceedings, pp. 740–747.
  140. 140.M. Pratama, J. Lu, E. Lughofer, G. Zhang, and S. Anavatti, “Scaffolding type-2 classifier for incremental learning under concept drifts,” Neurocomputing, vol. 191, pp. 304–329, 2016.

Citation

MLA
Lu, J., et al. “Learning Under Concept Drift: A Review”. IEEE Transactions on Knowledge and Data Engineering, 2018, pp. 1–1, https://doi.org/10.1109/TKDE.2018.2876857.
APA
Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under Concept Drift: A Review. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2018.2876857
Chicago
Lu, J., A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang. 2018. “Learning Under Concept Drift: A Review”. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2018.2876857.
Harvard
Lu, J. et al. (2018) “Learning under Concept Drift: A Review”, IEEE Transactions on Knowledge and Data Engineering, pp. 1–1. Available at: https://doi.org/10.1109/TKDE.2018.2876857.
Vancouver
1. Lu J, Liu A, Dong F, Gu F, Gama J, Zhang G (2018) Learning under Concept Drift: A Review. IEEE Transactions on Knowledge and Data Engineering 1–1

BibTeX

@article{Lu_2018, title={Learning under Concept Drift: A Review}, ISSN={2326-3865}, url={http://dx.doi.org/10.1109/TKDE.2018.2876857}, DOI={10.1109/tkde.2018.2876857}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Lu, Jie and Liu, Anjin and Dong, Fan and Gu, Feng and Gama, Joao and Zhang, Guangquan}, year={2018}, pages={1–1} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF