SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary

Alberto FernandezSalvador GarciaFrancisco HerreraNitesh V. Chawla

article2018JAIR2,129 citations
Listen

Real-world machine learning systems frequently face severely skewed data distributions, where critical eventssuch as fraudulent transactions, rare medical diagnoses, or equipment failuresare heavily outnumbered by normal instances. Standard predictive algorithms trained on such data tend to favor the majority class, creating a misleadingly high overall accuracy while failing to detect the vital minority cases. In high-stakes environments, these false negatives can lead to severe operational, financial, and safety risks.

The article provides a comprehensive evaluation of the Synthetic Minority Oversampling Technique, commonly known as SMOTE, marking its fifteen-year anniversary. It examines the algorithm's foundational role in addressing class imbalance, catalogs over eighty-five algorithmic extensions and adaptations across machine learning paradigms, and evaluates remaining challenges for modern enterprise environments.

To perform this assessment, the authors conducted an extensive review of the literature spanning fifteen years of research, analyzing theoretical properties, practical implementations, empirical benchmark studies, and over eighty-five specific variations of the technique. The analysis systematically evaluates how synthetic data generation interacts with diverse learning structures, high-dimensional datasets, streaming environments, and distributed computing frameworks.

The review produced several key findings. First, basic oversampling by exact replication causes models to overfit, whereas SMOTE's interpolation between neighboring minority examples generates new feature patterns that significantly improve model generalization. Second, researchers have developed over 85 variations that enhance baseline sampling, primarily by adaptively targeting difficult-to-learn examples, restricting synthesis along class boundaries, and filtering out noisy artificial points. Third, class imbalance alone is rarely the primary cause of model failure; performance degrades most severely when minority data suffers from overlapping class boundaries, small disjuncts (isolated sub-clusters), or label noise. Fourth, while the technique has successfully adapted into ensemble methods, multi-label tasks, and streaming data with concept drift, it struggles with the curse of dimensionality and large-scale distributed architectures, where distributed data partitioning often degrades performance compared to simpler sampling methods.

These findings indicate that addressing skewed data requires moving beyond simple rebalancing to actively manage underlying data complexity. Generating synthetic points without accounting for noise or class overlap risks creating harmful data artifacts, ultimately reducing decision accuracy and increasing business risk. Furthermore, in high-dimensional and large-scale distributed workflows, applying standard synthetic sampling without dimensionality reduction or specialized hardware acceleration can introduce major computational bottlenecks and suboptimal predictions.

Organizations handling imbalanced data should adopt hybrid approaches that combine synthetic sampling with intelligent noise filtering, clustering, or ensemble boosting rather than relying on standard oversampling alone. For high-dimensional datasets, teams should pair resampling with feature selection or extraction. Before deploying models in production, practitioners should validate systems using distribution-preserving partitioning to prevent validation bias, and they should evaluate exact nearest-neighbor implementations or graphical processing units when scaling to massive data volumes.

While the review provides high confidence in the utility and adaptability of synthetic oversampling for traditional and moderate datasets, caution is warranted in massive distributed environments. The article notes that theoretical understanding and scalable implementations for large-scale distributed frameworks remain limited, necessitating thorough pilot testing before implementing these techniques across large enterprise data pipelines.

  • Paper: Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss, Kaidi Cao et al. (2019). This work advances beyond heuristic synthetic resampling by introducing label-distribution-aware margin loss to mathematically optimize class margins and re-balancing schedules for extreme imbalance in deep networks.
  • Paper: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, Stephan Rabanser et al. (2019). This empirical study investigates practical methods for detecting dataset and label shifts at test time, extending the challenge of evolving class distributions highlighted in the SMOTE retrospective.
  • Paper: Detecting and Correcting for Label Shift with Black Box Predictors, Zachary C. Lipton et al. (2018). This paper introduces rigorous black-box methods for estimating and correcting shifting label distributions, providing a modern alternative to synthetic resampling when class proportions change over time.
  • Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). This work explores learning from severe class rarity and out-of-distribution anomalies by exposing models to auxiliary outliers rather than interpolating synthetic minority samples.
  • Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). This paper formulates deep one-class classification to handle extreme anomaly detection and class imbalance by enclosing normal data within a hypersphere rather than oversampling rare instances.

Table of Contents

  • 1. Introduction
  • 2. Synthetic Minority Oversampling Technique
  • 2.1 Why to Propose SMOTE
  • 2.2 SMOTE Description
  • 3. Extensions to SMOTE
  • 3.1 Properties for Categorizing the SMOTE-Based Extensions
  • 3.2 SMOTE-Based Extensions for Oversampling
  • 3.3 SMOTE-Based Extensions for Ensembles
  • 3.4 Exhaustive Empirical Studies Involving SMOTE
  • 4. Variations of SMOTE to Other Learning Paradigms
  • 4.1 Streaming Data
  • 4.2 Semi-supervised and Active Learning
  • 4.3 Multi-class, Multi-instance and Multi-label Classification
  • 4.4 Regression
  • 4.5 Other and More Complex Prediction Problems
  • 5. Challenges in SMOTE-Based Algorithms
  • 5.1 Small Disjuncts, Noise and Lack of Data
  • 5.2 Overlapping or Class Separability
  • 5.3 Dataset Shift
  • 5.4 Curse of Dimensionality and Interpolation Mechanisms
  • 5.5 Real-Time Processing
  • 5.6 Imbalanced Classification in Big Data Problems
  • 6. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Synthetic Minority Oversampling Technique (SMOTE) Algorithm

    algorithm

    The Synthetic Minority Oversampling Technique (SMOTE) is a data preprocessing algorithm designed to mitigate class imbalance by generating synthetic examples along feature-space line segments connecting neighboring minority class instances, rather than replicating existing instances.

    function SMOTE(T, N, k)
        Input: T (array of minority class instances), N (amount of oversampling as a percentage), k (number of nearest neighbors)
        Output: Synthetic (array of synthetic minority class instances)
        
        if N < 100 then
            Randomize order of instances in T
            T = subset of T of size (N / 100) * |T|
            N = 100
        end if
        
        num_samples = (int)(N / 100)
        newindex = 0
        Synthetic = empty 2D array
        
        for i = 1 to |T| do
            nnarray = find indices of k nearest minority neighbors of T[i]
            count = num_samples
            while count > 0 do
                nn = random integer from 1 to k
                neighbor_index = nnarray[nn]
                for attr = 1 to number of attributes do
                    dif = T[neighbor_index][attr] - T[i][attr]
                    gap = random real number in [0, 1]
                    Synthetic[newindex][attr] = T[i][attr] + gap * dif
                end for
                newindex = newindex + 1
                count = count - 1
            end while
        end for
        
        return Synthetic
    end function

    For continuous attributes, interpolation selects a random point along the vector difference x^xi\hat{x} - x_i where x^\hat{x} is a chosen nearest neighbor of xix_i. For nominal attributes, one of the two attribute values between xix_i and x^\hat{x} is selected uniformly at random.

  2. Knowl 2 — Seven-Dimensional Categorization Framework for SMOTE Extensions

    model/method

    To classify and analyze the landscape of oversampling techniques derived from SMOTE, oversampling methods are categorized along seven functional properties:

    1. Initial selection of instances to be oversampled: Identifies candidate instances before data generation to reduce noise and overlap (e.g., choosing borderline instances, filtering noisy instances, or selecting based on Support Vector Machine support vectors or Learning Vector Quantization prototypes).
    2. Integration with undersampling: Incorporates removal of majority class instances either before or concurrently with synthetic oversampling.
    3. Type of interpolation: Replaces standard linear interpolation with range-restricted methods, multi-instance interpolations, geometric topologies (e.g., ellipses, Voronoi diagrams), graph-based paths, clustering centroids, or probability density sampling (e.g., Gaussian, Markov chains).
    4. Operation with dimensionality changes: Alters the feature space prior to or during oversampling via Principal Component Analysis, autoencoders, manifold learning, feature selection, or kernel embeddings.
    5. Adaptive generation of synthetic examples: Allocates the number of generated synthetic instances per minority seed according to learning difficulty or local density (e.g., ADASYN).
    6. Relabeling: Adjusts labels of majority class instances or newly generated samples in boundary regions.
    7. Filtering of noisy generated instances: Post-processes the oversampled dataset using noise filters (such as Tomek links, Edited Nearest Neighbors, or rough-set filters) to eliminate overlapping or mislabeled synthetic samples.
  3. Knowl 3 — Maximum Fisher's Discriminant Ratio for Quantifying Class Overlap

    equation

    In class-imbalanced learning, the degree of class overlap often dominates the imbalance ratio as the primary factor driving classifier performance degradation. The degree of overlap between two classes along a single continuous feature jj is computed using Fisher's discriminant ratio fjf_j:

    fj=(μ1,jμ2,j)2σ1,j2+σ2,j2f_j = \frac{(\mu_{1,j} - \mu_{2,j})^2}{\sigma_{1,j}^2 + \sigma_{2,j}^2}

    where μ1,j\mu_{1,j} and μ2,j\mu_{2,j} are the sample means of the minority and majority classes along feature dimension jj, and σ1,j2\sigma_{1,j}^2 and σ2,j2\sigma_{2,j}^2 are the corresponding class variances for feature dimension jj.

    The overall dataset class overlap metric F1F1 (maximum Fisher's discriminant ratio) across all DD features is given by:

    F1=maxj=1,,DfjF1 = \max_{j=1, \dots, D} f_j

    A smaller value of F1F1 indicates higher class overlap and poorer linear separability, leading to higher classification error regardless of rebalancing.

  4. Knowl 4 — Adaptation of SMOTE to Non-Standard Machine Learning Paradigms

    model/method

    SMOTE preprocessing has been adapted across multiple non-standard learning paradigms to resolve class imbalance in complex data environments:

    • Streaming Data & Concept Drift: Adapting online oversampling (e.g., Learn++.NSE-SMOTE, GOS-IL) to non-stationary environments where data streams exhibit shifting class boundaries and evolving balance ratios over time.
    • Semi-Supervised & Active Learning: Generating synthetic points in unlabeled or partially labeled spaces (e.g., VIRTUAL, GS4, SEG-SSC, INNO) by leveraging support vectors or common hidden representations to guide active sample acquisition.
    • Multi-Instance Learning: Resolving imbalance at the bag or instance level (e.g., Instance-SMOTE, Bag-SMOTE, Informative-Bag-SMOTE) by generating synthetic minority instances within existing bags or synthesizing entirely new minority bags based on models of the negative population.
    • Multi-Label Classification: Handling instances associated with multiple simultaneous labels (e.g., MLSMOTE) by interpolating feature vectors and generating corresponding labelsets using neighborhood label correlation measures.
    • Imbalanced Regression & Ordinal Classification: Resampling continuous target variables (e.g., SMOTER) by defining rare regions via relevance thresholds and generating interpolated continuous targets, or using graph paths across ordered classes in ordinal regression (e.g., OGO-NI, OGO-ISP).
  5. Knowl 5 — Structural Taxonomy of SMOTE-Based Ensemble Architectures

    model/method

    Ensemble methods that incorporate SMOTE or its derivatives to achieve diversity among base classifiers are structured into three main architectural archetypes:

    • Boosting-Based SMOTE Ensembles: Embed oversampling inside the iterative boosting process (e.g., SMOTEBoost, DataBoost-IM, RAMOBoost, SkewBoost). In each boosting iteration, synthetic minority samples are generated to adaptively alter the training distribution, forcing subsequent weak learners to focus on difficult minority regions.
    • Bagging-Based SMOTE Ensembles: Integrate SMOTE into bootstrap sampling (e.g., SMOTEBagging, IIvotes+SPIDER, BEBS). Each bootstrap replicate is independently balanced with synthetic instances before training individual base models, preserving majority class diversity while stabilizing minority boundary learning.
    • Multi-Class Decomposition Ensembles: Integrate SMOTE with decomposition strategies such as One-Versus-All (OVA) or One-Versus-One (OVO) (e.g., OAA-DB, SMOTE+OVA, BBO), balancing each decomposed sub-problem independently via synthetic oversampling.
  6. Knowl 6 — Impact of Small Disjuncts, Noise, and Data Fragmentation on Imbalanced Learning

    theoretical result

    The difficulty of learning from imbalanced data is heavily driven by within-class sub-concepts formed as small disjuncts—tiny, isolated clusters of minority instances in feature space:

    • Pseudo-Noise Misclassification: Generalization mechanisms (such as decision tree pruning) tend to treat small disjuncts as class noise because they are surrounded by the majority class distribution, erroneously discarding valid minority clusters.
    • Data Fragmentation: Divide-and-conquer learning algorithms (including top-down decision trees and distributed partitioning schemes like MapReduce) iteratively split feature space, fragmenting small disjuncts into sub-partitions that contain zero or near-zero samples.
    • Interaction with SMOTE: Standard SMOTE reinforces small disjuncts when kk-nearest neighbors are located within the same cluster. However, when the neighborhood spans across classes due to overlap, standard SMOTE generates noisy instances inside majority regions. Effective mitigation requires combining SMOTE with cluster density estimation, instance cleaning filters (e.g., Tomek links, Edited Nearest Neighbors), or boosting weight updates.
  7. Knowl 7 — Hubness Problem and Curse of Dimensionality in SMOTE Interpolation

    limitation

    In high-dimensional feature spaces, SMOTE experiences significant performance degradation due to the curse of dimensionality and the emergence of hubness:

    • Hubness Occurrence: As dimensionality increases, distance distributions concentrate, resulting in a small subset of points ("hubs") appearing disproportionately as the nearest neighbors of a large number of other points.
    • Effect on Synthetic Data: When hub instances are repeatedly chosen during the kk-nearest neighbor phase of SMOTE, the generated synthetic samples become skewed toward the hub locations rather than populating the true underlying minority distribution. This increases the variance of the synthetic data and distorts the resulting classifier decision boundaries.
    • Remediation Strategies: Overcoming high-dimensional distortion requires applying feature selection or dimensionality reduction (e.g., Principal Component Analysis, autoencoders) before or after oversampling, or replacing Euclidean distance with skew-insensitive metrics such as Hellinger distance or Mahalanobis-based elliptical neighborhoods.
  8. Knowl 8 — Dataset Shift Categories and Validation Schemes in Class-Imbalanced Domains

    model/method

    Dataset shift occurs when training and test sets follow different probability distributions, a problem that is especially detrimental in imbalanced domains where minority samples are scarce:

    1. Prior Probability Shift: The class prior P(y)P(y) differs between training and test sets while conditional probability P(xy)P(x \mid y) remains constant. This is managed using stratified cross-validation.
    2. Covariate Shift: The input distribution P(x)P(x) changes between training and test sets while P(yx)P(y \mid x) remains invariant. Standard random kk-fold cross-validation often induces covariate shift by scattering scarce minority instances unevenly across folds. To prevent this, Distribution-Optimally-Balanced Stratified Cross-Validation (DOB-SCV) assigns close-by instances to distinct folds to ensure every fold contains representative samples from all regions.
    3. Concept Shift (Concept Drift): The relationship P(yx)P(y \mid x) changes over time. When processing continuous data streams, SMOTE must incorporate adaptive mechanisms to update synthetic instance generation in response to evolving class boundaries.
  9. Knowl 9 — Distributed and GPU Bottlenecks for SMOTE in Big Data Architectures

    limitation

    Scaling SMOTE to Big Data environments using distributed frameworks (such as MapReduce/Hadoop) or streaming GPU processors involves key architectural and algorithmic bottlenecks:

    • Map-Level Data Fragmentation: In standard MapReduce implementations where each Map task generates synthetic data locally on its assigned data chunk, minority class instances are fragmented across nodes. This local subsampling leads to inaccurate kk-nearest neighbor computation and worsens the small disjuncts problem.
    • Exact Distributed Nearest Neighbor Overhead: Computing global, exact kk-nearest neighbors across all minority class instances distributed among cluster nodes requires heavy inter-node communication and data shuffling.
    • GPU Memory Hierarchy Constraints: Parallelizing kk-NN searches on Graphics Processing Units (GPUs) requires fitting minority instances and intermediate distance structures into constrained GPU memory, demanding custom memory management and efficient parallel nearest-neighbor implementations.

Coverage note — Specific experimental benchmark outcomes for individual third-party datasets and comprehensive citations of all 85+ individual historical SMOTE variants were synthesized into the taxonomy and paradigm frameworks rather than retained as standalone knowls.

References

  1. 1.Abdi, L., & Hashemi, S. (2016). To combat multi-class imbalanced problems by means of over-sampling techniques. IEEE Transactions on Knowledge and Data Engineering, 28(1), 238–251.
  2. 2.Alejo, R., García, V., & Pacheco-Sánchez, J. H. (2015). An efficient over-sampling approach based on mean square error back-propagation for dealing with the multi-class imbalance problem. Neural Processing Letters, 42(3), 603–617.
  3. 3.Almogahed, B. A., & Kakadiaris, I. A. (2015). NEATER: filtering of over-sampled data using non-cooperative game theory. Soft Computing, 19(11), 3301–3322.
  4. 4.Anand, R., Mehrotra, K. G., Mohan, C. K., & Ranka, S. (1993). An improved algorithm for neural network classification of imbalanced training sets. IEEE Transactions on Neural Networks, 4(6), 962–969.
  5. 5.Bach, M., Werner, A., Zywiec, J., & Pluskiewicz, W. (2017). The study of under- and over-sampling methods utility in analysis of highly imbalanced data on osteoporosis. Information Sciences, 384, 174–190.
  6. 6.Barua, S., Islam, M. M., & Murase, K. (2011). A novel synthetic minority oversampling technique for imbalanced data set learning. In Neural Information Processing - 18th International Conference (ICONIP), pp. 735–744.
  7. 7.Barua, S., Islam, M. M., & Murase, K. (2013). ProWSyn: Proximity weighted synthetic oversampling technique for imbalanced data set learning. In Advances in Knowledge Discovery and Data Mining, 17th Pacific-Asia Conference (PAKDD), pp. 317–328.
  8. 8.Barua, S., Islam, M. M., & Murase, K. (2015). GOS-IL: A generalized over-sampling based online imbalanced learning framework. In Neural Information Processing - 22nd International Conference (ICONIP), pp. 680–687.
  9. 9.Barua, S., Islam, M. M., Yao, X., & Murase, K. (2014). MWMOTE-Majority weighted minority oversampling technique for imbalanced data set learning. . IEEE Transactions on Knowledge Data Engineering, 26(2), 405–425.
  10. 10.Batista, G. E. A. P. A., Prati, R. C., & Monard, M. C. (2004). A study of the behaviour of several methods for balancing machine learning training data. SIGKDD Explorations, 6(1), 20–29.
  11. 11.Bellinger, C., Drummond, C., & Japkowicz, N. (2016). Beyond the boundaries of SMOTE - A framework for manifold-based synthetically oversampling. In Machine Learning and Knowledge Discovery in Databases - European Conference (ECML PKDD), pp. 248–263.
  12. 12.Bellinger, C., Japkowicz, N., & Drummond, C. (2015). Synthetic oversampling for advanced radioactive threat detection. In 14th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 948–953.
  13. 13.Bhagat, R. C., & Patil, S. S. (2015). Enhanced SMOTE algorithm for classification of imbalanced big-data using random forest. In Advance Computing Conference (IACC), 2015 IEEE International, pp. 403–408.
  14. 14.Blagus, R., & Lusa, L. (2013). SMOTE for high-dimensional class-imbalanced data. BMC Bioinformatics, 14, 106.
  15. 15.Blaszczynski, J., Deckert, M., Stefanowski, J., & Wilk, S. (2012). Iivotes ensemble for imbalanced data. Intelligent Data Analysis, 16(5), 777–801.
  16. 16.Blaszczynski, J., & Stefanowski, J. (2015). Neighbourhood sampling in bagging for imbalanced data. . Neurocomputing, 150, 529–542.
  17. 17.Borowska, K., & Stepaniuk, J. (2016). Imbalanced data classification: A novel re-sampling approach combining versatile improved SMOTE and rough sets. In Computer Information Systems and Industrial Management - 15th IFIP TC8 International Conference (CISIM), pp. 31–42.
  18. 18.Branco, P., Torgo, L., & Ribeiro, R. P. (2016). A survey of predictive modelling under imbalanced distributions. ACM Computing Surveys, 49(2), 31:1–31:50.
  19. 19.Bruzzone, L., & Serpico, S. B. (1997). Classification of imbalanced remote-sensing data by neural networks. Pattern Recognition Letters, 18(11-13), 1323–1328.
  20. 20.Brzezinski, D., & Stefanowski, J. (2017). Prequential AUC: Properties of the area under the roc curve for data streams with concept drift. Knowledge and Information Systems, 52(2), 531–562.
  21. 21.Bunkhumpornpat, C., Sinapiromsaran, K., & Lursinsap, C. (2009). Safe-level-SMOTE: Safe-level-synthetic minority over-sampling technique for handling the class imbalanced problem. In Proceedings of the 13th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, (PAKDD '09), pp. 475–482.
  22. 22.Bunkhumpornpat, C., Sinapiromsaran, K., & Lursinsap, C. (2012). DBSMOTE: Density-based synthetic minority over-sampling TEchnique. Applied Intelligence, 36(3), 664–684.
  23. 23.Bunkhumpornpat, C., & Subpaiboonkit, S. (2013). Safe level graph for synthetic minority over-sampling techniques. In 13th International Symposium on Communications and Information Technologies (ISCIT), pp. 570–575.
  24. 24.Cao, H., Li, X., Woon, D. Y., & Ng, S. (2011). SPO: structure preserving oversampling for imbalanced time series classification. In 11th IEEE International Conference on Data Mining (ICDM), pp. 1008–1013.
  25. 25.Cao, H., Li, X., Woon, D. Y., & Ng, S. (2013). Integrated oversampling for imbalanced time series classification. IEEE Transactions on Knowledge and Data Engineering, 25(12), 2809–2822.
  26. 26.Cao, P., Liu, X., Yang, J., & Zhao, D. (2017a). A multi-kernel based framework for heterogeneous feature selection and over-sampling for computer-aided detection of pulmonary nodules. Pattern Recognition, 64, 327–346.
  27. 27.Cao, P., Liu, X., Zhang, J., Zhao, D., Huang, M., & Zaïane, O. R. (2017b). l21 norm regularized multi-kernel based joint nonlinear feature selection and over-sampling for imbalanced data classification. Neurocomputing, 234, 38–57.
  28. 28.Cao, Q., & Wang, S. (2011). Applying over-sampling technique based on data density and cost-sensitive svm to imbalanced learning. In International Conference on Information Management, Innovation Management and Industrial Engineering, pp. 543–548.
  29. 29.Cateni, S., Colla, V., & Vannucci, M. (2011). Novel resampling method for the classification of imbalanced datasets for industrial and other real-world problems. In 11th International Conference on Intelligent Systems Design and Applications (ISDA), pp. 402–407.
  30. 30.Cervantes, J., García-Lamont, F., Rodríguez-Mazahua, L., Chau, A. L., Ruiz-Castilla, J. S., & Trueba, A. (2017). PSO-based method for SVM classification on skewed data sets. Neurocomputing, 228, 187–197.
  31. 31.Charte, F., Rivera, A. J., del Jesus, M. J., & Herrera, F. (2015). MLSMOTE: approaching imbalanced multilabel learning through synthetic instance generation. Knowledge-Based Systems, 89, 385–397.
  32. 32.Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over–sampling technique. Journal of Artificial Intelligent Research, 16, 321–357.
  33. 33.Chawla, N. V., Japkowicz, N., & Kolcz, A. (2004). Editorial: Special issue on class imbalances. SIGKDD Explorations, 6(1), 1–6.
  34. 34.Chawla, N. V., Cieslak, D. A., Hall, L. O., & Joshi, A. (2008). Automatically countering imbalance and its empirical relationship to cost. Data Minining and Knowledge Discovery, 17(2), 225–252.
  35. 35.Chawla, N. V., Lazarevic, A., Hall, L. O., & Bowyer, K. W. (2003). SMOTEBoost: Improving prediction of the minority class in boosting. In Proceedings of 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (PKDD'03), pp. 107–119.
  36. 36.Chen, L., Cai, Z., Chen, L., & Gu, Q. (2010a). A novel differential evolution-clustering hybrid resampling algorithm on imbalanced datasets. . In WKDD, pp. 81–85.
  37. 37.Chen, S., He, H., & Garcia, E. A. (2010b). RAMOBoost: ranked minority oversampling in boosting. IEEE Transactions on Neural Networks, 21(10), 1624–1642.
  38. 38.Chen, S., Guo, G., & Chen, L. (2010c). A new over-sampling method based on cluster ensembles. In 7th International Conference on Advanced Information Networking and Applications Workshops, pp. 599–604.
  39. 39.Cieslak, D. A., & Chawla, N. V. (2008). Start globally, optimize locally, predict globally: Improving performance on imbalanced data. In Data Mining, 2008. ICDM'08. Eighth IEEE International Conference on, pp. 143–152. IEEE.
  40. 40.Cieslak, D. A., Hoens, T. R., Chawla, N. V., & Kegelmeyer, W. P. (2012). Hellinger distance decision trees are robust and skew-insensitive. Data Mining and Knowledge Discovery, 24(1), 136–158.
  41. 41.Cohen, G., Hilario, M., Sax, H., Hugonnet, S., & Geissbóhler, A. (2006). Learning from imbalanced data in surveillance of nosocomial infection. Artificial Intelligence in Medicine, 37(1), 7–18.
  42. 42.Dal Pozzolo, A., Caelen, O., & Bontempi, G. (2015). When is undersampling effective in unbalanced classification tasks?. In Appice, A., Rodrigues, P. P., Costa, V. S., Soares, C., Gama, J., & Jorge, A. (Eds.), ECML/PKDD, Vol. 9284 of Lecture Notes in Computer Science, pp. 200–215. Springer.
  43. 43.Dang, X. T., Tran, D. H., Hirose, O., & Satou, K. (2015). SPY: A novel resampling method for improving classification performance in imbalanced data. In 2015 Seventh International Conference on Knowledge and Systems Engineering (KSE), pp. 280–285.
  44. 44.Das, B., Krishnan, N. C., & Cook, D. J. (2015). RACOG and wRACOG: Two probabilistic oversampling techniques. IEEE Transactions on Knowledge and Data Engineering, 27(1), 222–234.
  45. 45.de la Calleja, J., & Fuentes, O. (2007). A distance-based over-sampling method for learning from imbalanced data sets. In Proceedings of the Twentieth International Florida Artificial Intelligence, pp. 634–635.
  46. 46.de la Calleja, J., Fuentes, O., & González, J. (2008). Selecting minority examples from misclassified data for over-sampling. In Proceedings of the Twenty-First International Florida Artificial Intelligence, pp. 276–281.
  47. 47.Dean, J., & Ghemawat, S. (2008). MapReduce: simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
  48. 48.Deepa, T., & Punithavalli, M. (2011). An E-SMOTE technique for feature selection in high-dimensional imbalanced dataset. In 3rd International Conference on Electronics Computer Technology (ICECT), pp. 322–324.
  49. 49.Dietterich, T. G., Lathrop, R. H., & Lozano-Pérez, T. (1997). Solving the multiple instance problem with axis-parallel rectangles. Artificial intelligence, 89(1), 31–71.
  50. 50.Ditzler, G., & Polikar, R. (2013). Incremental learning of concept drift from streaming imbalanced data. IEEE Transactions on Knowledge Data Engineering, 25(10), 2283–2301.
  51. 51.Ditzler, G., Polikar, R., & Chawla, N. V. (2010). An incremental learning algorithm for non-stationary environments and class imbalance. In 20th International Conference on Pattern Recognition (ICPR), pp. 2997–3000.
  52. 52.Dong, A., Chung, F.-l., & Wang, S. (2016). Semi-supervised classification method through oversampling and common hidden space. Information Sciences, 349, 216–228.
  53. 53.Dong, Y., & Wang, X. (2011). A new over-sampling approach: Random-SMOTE for learning from imbalanced data sets. In Knowledge Science, Engineering and Management - 5th International Conference (KSEM), pp. 343–352.
  54. 54.Douzas, G., & Bacao, F. (2017). Self-Organizing Map Oversampling (SOMO) for imbalanced data set learning. Expert Systems with Applications, 82, 40–52.
  55. 55.Ertekin, S. (2013). Adaptive oversampling for imbalanced data classification. In Information Sciences and Systems 2013 - Proceedings of the 28th International Symposium on Computer and Information Sciences (ISCIS), pp. 261–269.
  56. 56.Estabrooks, A., Jo, T., & Japkowicz, N. (2004). A multiple resampling method for learning from imbalanced data sets. Computational Intelligence, 20(1), 18–36.
  57. 57.Fan, X., Tang, K., & Weise, T. (2011). Margin-based over-sampling method for learning from imbalanced datasets. In Advances in Knowledge Discovery and Data Mining - 15th Pacific-Asia Conference (PAKDD), pp. 309–320.
  58. 58.Farquad, M. A. H., & Bose, I. (2012). Preprocessing unbalanced data using support vector machine. Decision Support Systems, 53(1), 226–233.
  59. 59.Fernandez, A., del Rio, S., Chawla, N. V., & Herrera, F. (2017). An insight into imbalanced big data classification: Outcomes and challenges. Complex and Intelligent Systems, 3(2), 105–120.
  60. 60.Fernandez, A., Garcia, S., del Jesus, M. J., & Herrera, F. (2008). A study of the behaviour of linguistic fuzzy rule based classification systems in the framework of imbalanced data–sets. Fuzzy Sets and Systems, 159(18), 2378–2398.
  61. 61.Fernandez, A., Garcia, S., Luengo, J., Bernado-Mansilla, E., & Herrera, F. (2010). Genetics-based machine learning for rule induction: State of the art, taxonomy and comparative study. IEEE Transactions on Evolutionary Computation, 14(6), 913–941.
  62. 62.Fernández, A., López, V., Galar, M., Del Jesus, M. J., & Herrera, F. (2013). Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches. Knowledge-based systems, 42, 97–110.
  63. 63.Fernández, A., Río, S., López, V., Bawakid, A., del Jesus, M. J., Benítez, J. M., & Herrera, F. (2014). Big data with cloud computing: An information sciencesight on the computing environment, mapreduce and programming framework. WIREs Data Mining and Knowledge Discovery, 4(5), 380–409.
  64. 64.Fernandez, A., Carmona, C. J., del Jesus, M. J., & Herrera, F. (2017). A pareto based ensemble with feature and instance selection for learning from multi-class imbalanced datasets. International Journal of Neural Systems, 27(6), 1–21.
  65. 65.Fernández-Navarro, F., Hervás-Martínez, C., & Gutiérrez, P. A. (2011). A dynamic over-sampling procedure based on sensitivity for multi-class problems. Pattern Recognition, 44(8), 1821–1833.
  66. 66.Frank, E., & Pfahringer, B. (2006). Improving on bagging with input smearing. In Advances in Knowledge Discovery and Data Mining, 10th Pacific-Asia Conference (PAKDD), pp. 97–106.
  67. 67.Friedman, J. H. (1996). Another approach to polychotomous classification. Tech. rep., Department of Statistics, Stanford University.
  68. 68.Galar, M., Fernández, A., Barrenechea, E., Bustince, H., & Herrera, F. (2011). An overview of ensemble methods for binary classifiers in multi-class problems: Experimental study on one-vs-one and one-vs-all schemes. Pattern Recognition, 44(8), 1761–1776.
  69. 69.Galar, M., Fernandez, A., Barrenechea, E., Bustince, H., & Herrera, F. (2012). A review on ensembles for class imbalance problem: Bagging, boosting and hybrid based approaches. IEEE Transactions on System, Man and Cybernetics Part C: Applications and Reviews, 42(4), 463–484.
  70. 70.Galar, M., Fernández, A., Barrenechea, E., Bustince, H., & Herrera, F. (2016). Ordering-based pruning for improving the performance of ensembles of classifiers in the framework of imbalanced datasets. . Information Sciences, 354, 178–196.
  71. 71.Gao, K., Khosgoftaar, T. M., & Wald, R. (2014a). the use of under- and oversampling within ensemble feature selection and classification for software quality prediction. International Journal of Reliability, Quality and Safety Engineering, 21(1).
  72. 72.Gao, M., Hong, X., Chen, S., Harris, C. J., & Khalaf, E. (2014b). PDFOS: PDF estimation based over-sampling for imbalanced two-class problems. Neurocomputing, 138, 248–259.
  73. 73.García, S., Luengo, J., & Herrera, F. (2016). Tutorial on practical tips of the most influential data preprocessing algorithms in data mining. Knowledge-Based Systems, 98, 1–29.
  74. 74.García, V., Mollineda, R. A., & Sánchez, J. S. (2008). On the k-nn performance in a challenging scenario of imbalance and overlapping. Pattern Analysis and Applications, 11(3-4), 269–280.
  75. 75.García, V., Sánchez, J. S., de J. Ochoa Domínguez, H., & Cleofas-Sánchez, L. (2015). Dissimilarity-based learning from imbalanced data with small disjuncts and noise. . In Paredes, R., Cardoso, J. S., & Pardo, X. M. (Eds.), IbPRIA, Vol. 9117 of Lecture Notes in Computer Science, pp. 370–378. Springer.
  76. 76.Gazzah, S., & Amara, N. E. B. (2008). New oversampling approaches based on polynomial fitting for imbalanced data sets. In The Eighth IAPR International Workshop on Document Analysis Systems, pp. 677–684.
  77. 77.Gazzah, S., Hechkel, A., & Amara, N. E. B. (2015). A hybrid sampling method for imbalanced data. In 12th International Multi-Conference on Systems, Signals and Devices, pp. 1–6.
  78. 78.Ghazikhani, A., Monsefi, R., & Sadoghi Yazdi, H. (2013). Ensemble of online neural networks for non-stationary and imbalanced data streams. Neurocomputing, 122, 535–544.
  79. 79.Gong, C., & Gu, L. (2016). A novel SMOTE-based classification approach to online data imbalance problem. Mathematical Problems in Engineering, Article ID 5685970, 14.
  80. 80.Gong, J., & Kim, H. (2017). RHSBoost: improving classification performance in imbalance data. Computational Statistics and Data Analysis, 111(C), 1–13.
  81. 81.Gu, Q., Cai, Z., & Zhu, L. (2009). Classification of imbalanced data sets by using the hybrid re-sampling algorithm based on isomap. . In ISICA, Vol. 5821 of Lecture Notes in Computer Science, pp. 287–296.
  82. 82.Guo, H., & Viktor, H. L. (2004). Learning from imbalanced data sets with boosting and data generation: the DataBoost-IM approach. SIGKDD Explorations Newsletter, 6, 30–39.
  83. 83.Gutierrez, P. D., Lastra, M., Bacardit, J., Benitez, J. M., & Herrera, F. (2016). GPU-SME-kNN: Scalable and memory efficient kNN and lazy learning using GPUs. Information Sciences, 373, 165–182.
  84. 84.Gutierrez, P. D., Lastra, M., Benitez, J. M., & Herrera, F. (2017). SMOTE-GPU: Big data preprocessing on commodity hardware for imbalanced classification. Progress in Artificial Intelligence, 6(4), 347–354.
  85. 85.Haixiang, G., Yijing, L., Shang, J., Mingyun, G., Yuanyue, H., & Bing, G. (2017). Learning from class-imbalanced data: Review of methods and applications. Expert Syst. with Applicat. , 73, 220 – 239.
  86. 86.Hamid, Y., Sugumaran, M., & Journaux, L. (2016). A fusion of feature extraction and feature selection technique for network intrusion detection. International Journal of Security and its Applications, 10(8), 151–158.
  87. 87.Han, H., Wang, W. Y., & Mao, B. H. (2005). Borderline–SMOTE: A new over–sampling method in imbalanced data sets learning. In Proceedings of the 2005 International Conference on Intelligent Computing (ICIC'05), Vol. 3644 of Lecture Notes in Computer Science, pp. 878–887.
  88. 88.He, H., Bai, Y., Garcia, E. A., & Li, S. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. In Proceedings of the 2008 IEEE International Joint Conference Neural Networks (IJCNN'08), pp. 1322–1328.
  89. 89.He, H., & Chen, S. (2011). Towards incremental learning of nonstationary imbalanced data stream: A multiple selectively recursive approach. Evolving Systems, 2(1), 35–50.
  90. 90.He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263–1284.
  91. 91.Herrera, F., Charte, F., Rivera, A. J., & del Jesós, M. J. (2016a). Multilabel Classification - Problem Analysis, Metrics and Techniques. Springer.
  92. 92.Herrera, F., Ventura, S., Bello, R., Cornelis, C., Zafra, A., Tarragó, D. S., & Vluymans, S. (2016b). Multiple Instance Learning - Foundations and Algorithms. Springer.
  93. 93.Ho, T. K., & Basu, M. (2002). Complexity measures of supervised classification problems. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24(3), 289–300.
  94. 94.Hoens, T. R., & Chawla, N. V. (2010). Generating diverse ensembles to counter the problem of class imbalance. . In Zaki, M. J., Yu, J. X., Ravindran, B., & Pudi, V. (Eds.), Proceedings in Advances in Knowledge Discovery and Data Mining (PAKDD), Vol. 6119 of Lecture Notes in Computer Science, pp. 488–499. Springer.
  95. 95.Hoens, T. R., & Chawla, N. V. (2013). Imbalanced datasets: from sampling to classifiers, pp. 43–59. Wiley.
  96. 96.Hoens, T. R., Chawla, N. V., & Polikar, R. (2011). Heuristic updatable weighted random subspaces for non-stationary environments. In Data Mining (ICDM), 2011 IEEE 11th International Conference on, pp. 241–250. IEEE.
  97. 97.Hoens, T. R., Polikar, R., & Chawla, N. V. (2012a). Learning from streaming data with concept drift and imbalance: an overview. Progress in Artificial Intelligence, 1(1), 89–101.
  98. 98.Hoens, T. R., Qian, Q., Chawla, N. V., & Zhou, Z.-H. (2012b). Building decision trees for the multi-class imbalance problem. . In Tan, P.-N., Chawla, S., Ho, C. K., & Bailey, J. (Eds.), Proceedings in Advances in knowledge discovery and data mining (PAKDD), Vol. 7301 of Lecture Notes in Computer Science, pp. 122–134. Springer.
  99. 99.Hoens, T. R., & Chawla, N. V. (2012). Learning in non-stationary environments with class imbalance. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 168–176. ACM.
  100. 100.Hu, F., & Li, H. (2013). A novel boundary oversampling algorithm based on neighborhood rough set model: NRSBoundary-SMOTE. Mathematical Problems in Engineering, Article ID 694809, 10.
  101. 101.Hu, F., Li, H., Lou, H., & Dai, J. (2014). A parallel oversampling algorithm based on nrsboundary-smote. Journal of Information and Computational Science, 11(13), 4655–4665.
  102. 102.Hu, F., & Li, H. (2013). A novel boundary oversampling algorithm based on neighborhood rough set model: Nrsboundary-smote. Mathematical Problems in Engineering, 2013, 694809.
  103. 103.Hukerikar, S., Tumma, A., Nikam, A., & Attar, V. (2011). SkewBoost: An algorithm for classifying imbalanced datasets. In Proceedings of the International Conference on Computer Communication Technology (ICCCT), pp. 46–52.
  104. 104.Hwang, D., Fotouhi, F., Finley Jr., R. L., & Grosky, W. I. (2003). Predictive model for yeast protein functions using modular neural approach. . In Proceedings of the Third IEEE Symposium on BioInformatics and BioEngineering (BIBE03), pp. 436–441. IEEE Computer Society.
  105. 105.Iglesias, E. L., Vieira, A. S., & Borrajo, L. (2013). An HMM-based over-sampling technique to improve text classification. Expert Systems with Applications, 40(18), 7184–7192.
  106. 106.Japkowicz, N., & Holte, R. (2000). Workshop report: AAAI2000 workshop on learning from imbalanced data-sets. AI Magazine, 22(1), 127–136.
  107. 107.Jeatrakul, P., & Wong, K. W. (2012). Enhancing classification performance of multi-class imbalanced data using the OAA-DB algorithm. In The 2012 International Joint Conference on Neural Networks (IJCNN), pp. 1–8.
  108. 108.Jiang, K., Lu, J., & Xia, K. (2016). A novel algorithm for imbalance data classification based on genetic algorithm improved SMOTE. Arabian Journal for Science and Engineering, 41(1), 3255–3266.
  109. 109.Jiang, L., Qiu, C., & Li, C. (2015). A novel minority cloning technique for cost-sensitive learning. International Journal of Pattern Recognition and Artificial Intelligence, 29(4).
  110. 110.Jo, T., & Japkowicz, N. (2004). Class imbalances versus small disjuncts. ACM SIGKDD Explorations Newsletter, 6(1), 40–49.
  111. 111.Kang, Y.-I., & Won, S. (2010). Weight decision algorithm for oversampling technique on class-imbalanced learning. In ICCAS, pp. 182–186.
  112. 112.Khan, S. H., Bennamoun, M., Sohel, F. A., & Togneri, R. (2018). Cost sensitive learning of deep feature representations from imbalanced data. IEEE Transactions on Neural Networks and Learning Systems, in press, doi: 10.1109/TNNLS.2017.2732482, 1–15.
  113. 113.Khoshgoftaar, T. M., Hulse, J. V., & Napolitano, A. (2011). Comparing boosting and bagging techniques with noisy and imbalanced data. IEEE Transactions on Systems, Man, and Cybernetics, Part A, 41(3), 552–568.
  114. 114.Koto, F. (2014). SMOTE-out, SMOTE-cosine, and selected-SMOTE: An enhancement strategy to handle imbalance in data level. In 6th International Conference on Advanced Computer Science and Information Systems (ICACSIS), pp. 193–197.
  115. 115.Koziarski, M., Krawczyk, B., & Wozniak, M. (2017). Radial-based approach to imbalanced data oversampling. . In de Pisón, F. J. M., Urraca-Valle, R., Quintián, H., & Corchado, E. (Eds.), HAIS, Vol. 10334 of Lecture Notes in Computer Science, pp. 318–327. Springer.
  116. 116.Krawczyk, B., Galar, M., Jelen, L., & Herrera, F. (2016). Evolutionary undersampling boosting for imbalanced classification of breast cancer malignancy. Applied Soft Computing Journal, 38, 714–726.
  117. 117.Krawczyk, B. (2016). Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence, 5(4), 221–232.
  118. 118.Krawczyk, B., Minku, L. L., Gama, J. a., Stefanowski, J., & Woźniak, M. (2017). Ensemble learning for data stream analysis: a survey. Information Fusion, 37, 132–156.
  119. 119.Kubat, M., Holte, R. C., & Matwin, S. (1998). Machine learning for the detection of oil spills in satellite radar images. Machine Learning, 30(2-3), 195–215.
  120. 120.Kubat, M., & Matwin, S. (1997). Addressing the curse of imbalanced training sets: one–sided selection. In Proceedings of the 14th International Conference on Machine Learning (ICML'97), pp. 179–186.
  121. 121.Lachheta, P., & Bawa, S. (2016). Combining synthetic minority oversampling technique and subset feature selection technique for class imbalance problem. In ACM International Conference Proceeding Series, pp. 1–8.
  122. 122.Last, M. (2002). Online classification of nonstationary data streams. Intelligent Data Analysis, 6(2), 129–147.
  123. 123.Lee, J., Kim, N., & Lee, J. (2015). An over-sampling technique with rejection for imbalanced class learning. In Proceedings of the 9th International Conference on Ubiquitous Information Management and Communication (IMCOM), pp. 102:1–102:6.
  124. 124.Lemaitre, G., Nogueira, F., & Aridas, C. K. (2017). Imbalanced-learn: A python toolbox to tackle the curse of imbalanced datasets in machine learning. Journal of Machine Learning Research, 18, 1–5.
  125. 125.Li, F., Yu, C., Yang, N., Xia, F., Li, G., & Kaveh-Yazdy, F. (2013a). Iterative nearest neighborhood oversampling in semisupervised learning from imbalanced data. . The Scientific World Journal. , Article ID 875450.
  126. 126.Li, H., Zou, P., Wang, X., & Xia, R. (2013b). A new combination sampling method for imbalanced data. In Proceedings of 2013 Chinese Intelligent Automation Conference, pp. 547–554.
  127. 127.Li, J., Fong, S., Wong, R. K., & Chu, V. W. (201). Adaptive multi-objective swarm fusion for imbalanced data classification. Information Fusion, 39, 1–24.
  128. 128.Li, J., Fong, S., & Zhuang, Y. (2015). Optimizing SMOTE by metaheuristics with neural network and decision tree. In 2015 3rd International Symposium on Computational and Business Intelligence (ISCBI), pp. 26–32.
  129. 129.Li, K., Zhang, W., Lu, Q., & Fang, X. (2014). An improved SMOTE imbalanced data classification method based on support degree. In International Conference on Identification, Information and Knowledge in the Internet of Things (IIKI), pp. 34–38.
  130. 130.Liang, Y., Hu, S., Ma, L., & He, Y. (2009). MSMOTE: Improving classification performance when training data is imbalanced. In Computer Science and Engineering, International Workshop on, Vol. 2, pp. 13–17.
  131. 131.Lichtenwalter, R. N., Lussier, J. T., & Chawla, N. V. (2010). New perspectives and methods in link prediction. In Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 243–252. ACM.
  132. 132.Lin, W.-J., & Chen, J. J. (2013). Class-imbalanced classifiers for high-dimensional data. . Briefings in Bioinformatics, 14(1), 13–26.
  133. 133.López, V., Fernández, A., García, S., Palade, V., & Herrera, F. (2013). An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics. Information Sciences, 250, 113–141.
  134. 134.Lopez, V., Fernandez, A., & Herrera, F. (2014). On the importance of the validation technique for classification with imbalanced datasets: Addressing covariate shift when data is skewed. Information Sciences, 257, 1–13.
  135. 135.Lopez, V., Fernandez, A., Moreno-Torres, J. G., & Herrera, F. (2012). Analysis of preprocessing vs. cost-sensitive learning for imbalanced classification. Open problems on intrinsic data characteristics. Expert Systems with Applications, 39(7), 6585–6608.
  136. 136.López, V., Triguero, I., Carmona, C. J., García, S., & Herrera, F. (2014). Addressing imbalanced classification with instance generation techniques: Ipade-id. Neurocomputing, 126, 15–28.
  137. 137.Luengo, J., Fernández, A., García, S., & Herrera, F. (2011). Addressing data complexity for imbalanced data sets: analysis of SMOTE–based oversampling and evolutionary undersampling. Soft Computing, 15(10), 1909–1936.
  138. 138.Ma, L., & Fan, S. (2017). CURE-SMOTE algorithm and hybrid algorithm for feature selection and parameter optimization based on random forests. BMC Bioinformatics, 18, 169.
  139. 139.Maciejewski, T., & Stefanowski, J. (2011). Local neighbourhood extension of SMOTE for mining imbalanced data. . In CIDM, pp. 104–111.
  140. 140.Mahmoudi, S., Moradi, P., Ahklaghian, F., & Moradi, R. (2014). Diversity and separable metrics in over-sampling technique for imbalanced data classification. In 4th International eConference on Computer and Knowledge Engineering (ICCKE), pp. 152–158.
  141. 141.Mao, W., Wang, J., & Wang, L. (2015). Online sequential classification of imbalanced data by combining extreme learning machine and improved SMOTE algorithm. In 2015 International Joint Conference on Neural Networks (IJCNN), pp. 1–8.
  142. 142.Martín-Félez, R., & Mollineda, R. A. (2010). On the suitability of combining feature selection and resampling to manage data complexity. In Proceedings of the Conferencia de la Asociación Española de Inteligencia Artificial (CAEPIA'09), Vol. 5988 of Lecture Notes on Artificial Intelligence, pp. 141–150.
  143. 143.Mathew, J., Luo, M., Pang, C. K., & Chan, H. L. (2015). Kernel-based SMOTE for SVM classification of imbalanced datasets. In Industrial Electronics Society, IECON 2015- 41st Annual Conference of the IEEE, pp. 001127–001132.
  144. 144.Maua, G., & Galinac Grbac, T. (2017). Co-evolutionary multi-population genetic programming for classification in software defect prediction: An empirical case study. Applied Soft Computing Journal, 55, 331–351.
  145. 145.Mease, D., Wyner, A. J., & Buja, A. (2007). Boosted classification trees and class probability/quantile estimation. Journal of Machine Learning Research, 8, 409–439.
  146. 146.Menardi, G., & Torelli, N. (2014). Training and assessing classification rules with imbalanced data. Data Mining and Knowledge Discoveryd, 28(1), 92–122.
  147. 147.Mera, C., Arrieta, J., Orozco-Alzate, M., & Branch, J. (2015). A bag oversampling approach for class imbalance in multiple instance learning. In Progress in Pattern Recognition, Image Analysis, Computer Vision, and Applications - 20th Iberoamerican Congress (CIARP), pp. 724–731.
  148. 148.Mera, C., Orozco-Alzate, M., & Branch, J. (2014). Improving representation of the positive class in imbalanced multiple-instance learning. In Image Analysis and Recognition - 11th International Conference (ICIAR), pp. 266–273.
  149. 149.Mirza, B., Lin, Z., & Liu, N. (2015). Ensemble of subset online sequential extreme learning machine for class imbalance and concept drift. . Neurocomputing, 149, 316–329.
  150. 150.Moniz, N., Branco, P., & Torgo, L. (2016). Resampling strategies for imbalanced time series. In 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) 2016, pp. 282–291.
  151. 151.Moreno-Torres, J. G., Raeder, T., Alaiz-Rodriguez, R., Chawla, N. V., & Herrera, F. (2012a). A unifying view on dataset shift in classification. Pattern Recognition, 45(1), 521–530.
  152. 152.Moreno-Torres, J. G., Sáez, J. A., & Herrera, F. (2012b). Study on the impact of partition-induced dataset shift on-fold cross-validation. IEEE Transactions on Neural Networks and Learning Systems, 23(8), 1304–1312.
  153. 153.Moreno-Torres, J. G., & Herrera, F. (2010). A preliminary study on overlapping and data fracture in imbalanced domains by means of genetic programming-based feature extraction. In Proceedings of the 10th International Conference on Intelligent Systems Design and Applications (ISDA'10), pp. 501–506.
  154. 154.Moutafis, P., & Kakadiaris, I. A. (2014). GS4: generating synthetic samples for semi-supervised nearest neighbor classification. In Trends and Applications in Knowledge Discovery and Data Mining (PAKDD), pp. 393–403.
  155. 155.Nakamura, M., Kajiwara, Y., Otsuka, A., & Kimura, H. (2013). LVQ-SMOTE - learning vector quantization based synthetic minority over-sampling technique for biomedical data. BioData Mining, 6, 16.
  156. 156.Napierala, K., & Stefanowski, J. (2016). Types of minority class examples and their influence on learning classifiers from imbalanced data. Journal of Intelligent Information Systems, 46(3), 563–597.
  157. 157.Napierala, K., Stefanowski, J., & Wilk, S. (2010). Learning from imbalanced data in presence of noisy and borderline examples. In Proceedings of the 7th International Conference on Rough Sets and Current Trends in Computing (RSCTC'10), Vol. 6086 of Lecture Notes on Artificial Intelligence, pp. 158–167.
  158. 158.Nekooeimehr, I., & Lai-Yuen, S. K. (2016). Adaptive semi-unsupervised weighted oversampling (A-SUWO) for imbalanced datasets. Expert Systems with Applications, 46(C), 405–416.
  159. 159.Nguyen, H. M., Cooper, E. W., & Kamei, K. (2011). Online learning from imbalanced data streams. In Proceedings of the 2011 International Conference of Soft Computing and Pattern Recognition, SoCPaR 2011, pp. 347–352.
  160. 160.Orriols-Puig, A., Bernadó-Mansilla, E., Goldberg, D. E., Sastry, K., & Lanzi, P. L. (2009). Facetwise analysis of XCS for problems with class imbalances. IEEE Transactions on Evolutionary Computation, 13, 260–283.
  161. 161.Palacios, A. M., Sánchez, L., & Couso, I. (2012). Equalizing imbalanced imprecise datasets for genetic fuzzy classifiers. International Journal of Computational Intelligence Systems, 5(2), 276–296.
  162. 162.Pan, S., Wu, J., Zhu, X., & Zhang, C. (2015). Graph ensemble boosting for imbalanced noisy graph stream classification. IEEE Transactions on Cybernetics, 45(5), 940–954.
  163. 163.Park, Y., Qi, Z., Chari, S. N., & Molloy, I. (2014). PAKDD'12 best paper: generating balanced classifier-independent training samples from unlabeled data. Knowledge and Information Systems, 41(3), 871–892.
  164. 164.Pekalska, E., & Duin, R. P. W. (2005). The Dissimilarity Representation for Pattern Recognition - Foundations and Applications, Vol. 64 of Series in Machine Perception and Artificial Intelligence. World Scientific.
  165. 165.Peng, L., Zhang, H., Yang, B., Chen, Y., & Zhou, X. (2016). SMOTE-DGC: an imbalanced learning approach of data gravitation based classification. In Intelligent Computing Theories and Application - 12th International Conference (ICIC), pp. 133–144.
  166. 166.Peng, Y., & Yao, J. (2010). AdaOUBoost: adaptive over-sampling and under-sampling to boost the concept learning in large scale imbalanced data sets. In Multimedia Information Retrieval, pp. 111–118.
  167. 167.Pérez-Ortiz, M., Gutiérrez, P. A., & Hervás-Martínez, C. (2013). Borderline kernel based over-sampling. In Hybrid Artificial Intelligent Systems - 8th International Conference (HAIS), pp. 472–481.
  168. 168.Pérez-Ortiz, M., Gutiérrez, P. A., Hervás-Martínez, C., & Yao, X. (2015). Graph-based approaches for over-sampling in the context of ordinal regression. IEEE Transactions on Knowledge Data Engineering, 27(5), 1233–1245.
  169. 169.Pérez-Ortiz, M., Gutiérrez, P. A., Tiño, P., & Hervás-Martínez, C. (2016). Oversampling the minority class in the feature space. IEEE Transactions on Neural Networks and Learning Systems, 27(9), 1947–1961.
  170. 170.Piras, L., & Giacinto, G. (2012). Synthetic pattern generation for imbalanced learning in image retrieval. Pattern Recognition Letters, 33(16), 2198–2205.
  171. 171.Pourhabib, A., Mallick, B. K., & Ding, Y. (2015). Absent data generating classifier for imbalanced class sizes. Journal of Machine Learning Research, 16(1), 2695–2724.
  172. 172.Prati, R. C., & Batista, G. E. A. P. A. (2004). Class imbalances versus class overlapping: an analysis of a learning system behavior. In Proceedings of the 2004 Mexican International Conference on Artificial Intelligence (MICAI'04), pp. 312–321.
  173. 173.Prati, R. C., Batista, G. E. A. P. A., & Silva, D. F. (2015). Class imbalance revisited: a new experimental setup to assess the performance of treatment methods. Knowledge and Information Systems, 45(1), 247–270.
  174. 174.Puntumapon, K., & Waiyamai, K. (2012). A pruning-based approach for searching precise and generalized region for synthetic minority over-sampling. In Advances in Knowledge Discovery and Data Mining - 16th Pacific-Asia Conference (PAKDD), pp. 371–382.
  175. 175.Radovanovic, M., Nanopoulos, A., & Ivanovic, M. (2010). Hubs in space: Popular nearest neighbors in high-dimensional data. . Journal of Machine Learning Research, 11, 2487–2531.
  176. 176.Radtke, P. V. W., Granger, E., Sabourin, R., & Gorodnichy, D. O. (2014). Skew-sensitive boolean combination for adaptive ensembles - an application to face recognition in video surveillance. Information Fusion, 20(1), 31–48.
  177. 177.Ramentol, E., Caballero, Y., Bello, R., & Herrera, F. (2012). SMOTE-RSB*: a hybrid preprocessing approach based on oversampling and undersampling for high imbalanced data-sets using smote and rough sets theory. . Knowledge and Information Systems, 33(2), 245–265.
  178. 178.Ramentol, E., Gondres, I., Lajes, S., Bello, R., Caballero, Y., Cornelis, C., & Herrera, F. (2016). Fuzzy-rough imbalanced learning for the diagnosis of high voltage circuit breaker maintenance: The SMOTE-FRST-2T algorithm. Engineering Applications of AI, 48, 134–139.
  179. 179.Ramírez-Gallego, S., Fernández, A., García, S., Chen, M., & Herrera, F. (2018). Big data: Tutorial and guidelines on information and process fusion for analytics algorithms with mapreduce. Information Fusion, 42, 51–61.
  180. 180.Ramírez-Gallego, S., Krawczyk, B., García, S., Woźniak, M., & Herrera, F. (2017). A survey on data preprocessing for data stream mining: Current status and future directions. Neurocomputing, 239, 39–57.
  181. 181.Raudys, S. J., & Jain, A. K. (1991). Small sample size effects in statistical pattern recognition: Recommendations for practitioners. IEEE Transactions on Pattern Analysis and Machine Intelligence, 13(3), 252–264.
  182. 182.Río, S., López, V., Benítez, J. M., & Herrera, F. (2014). On the use of mapreduce for imbalanced big data using random forest. Information Sciences, 285, 112–137.
  183. 183.Rivera, W. A. (2017). Noise reduction a priori synthetic over-sampling for class imbalanced data sets. Information Sciences, 408, 146–161.
  184. 184.Rivera, W. A., & Xanthopoulos, P. (2016). A priori synthetic over-sampling methods for increasing classification sensitivity in imbalanced data sets. Expert Systems with Appications. , 66, 124–135.
  185. 185.Rokach, L. (2016). Decision forest: Twenty years of research. Information Fusion, 27, 111–125.
  186. 186.Rong, T., Gong, H., & Ng, W. W. Y. (2014). Stochastic sensitivity oversampling technique for imbalanced data. In ICMLC (CCIS volume), Vol. 481 of Communications in Computer and Information Science, pp. 161–171.
  187. 187.Sáez, J. A., Luengo, J., Stefanowski, J., & Herrera, F. (2015). SMOTE-IPF: addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering. Information Sciences, 291, 184–203.
  188. 188.Sánchez, A. I., Morales, E. F., & Gonzalez, J. A. (2013). Synthetic oversampling of instances using clustering. International Journal on Artificial Intelligence Tools, 22(2).
  189. 189.Sandhan, T., & Choi, J. Y. (2014). Handling imbalanced datasets by partially guided hybrid sampling for pattern recognition. In 22nd International Conference on Pattern Recognition (ICPR), pp. 1449–1453.
  190. 190.Schapire, R. E. (1999). A brief introduction to boosting. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI'99), pp. 1401–1406.
  191. 191.Seiffert, C., Khoshgoftaar, T. M., Hulse, J. V., & Folleco, A. (2014). An empirical study of the classification performance of learners on imbalanced and noisy software quality data. Information Sciences, 259, 571–595.
  192. 192.Sen, A., Islam, M. M., Murase, K., & Yao, X. (2016). Binarization with boosting and oversampling for multiclass classification. IEEE Transactions on Cybernetics, 46(5), 1078–1091.
  193. 193.Shimodaira, H. (2000). Improving predictive inference under Covariate Shift by Weighting the Log-likelihood Function. Journal of Statistical Planning and Inference, 90(2), 227–244.
  194. 194.Stefanowski, J. (2016). Dealing with data difficulty factors while learning from imbalanced data. . In Matwin, S., & Mielniczuk, J. (Eds.), Challenges in Computational Statistics and Data Mining, Vol. 605 of Studies in Computational Intelligence, pp. 333–363. Springer.
  195. 195.Stefanowski, J., & Wilk, S. (2008). Selective pre-processing of imbalanced data for improving classification performance. In Data Warehousing and Knowledge Discovery, 10th International Conference, pp. 283–292.
  196. 196.Storkey, A. (2009). When training and test sets are different: Characterizing learning transfer. In nonero Candela, J. Q., Sugiyama, M., Schwaighofer, A., & Lawrence, N. D. (Eds.), Dataset Shift in Machine Learning, pp. 3–28. MIT Press.
  197. 197.Sun, Y., Wong, A. K. C., & Kamel, M. S. (2009). Classification of imbalanced data: A review. International Journal of Pattern Recognition and Artificial Intelligence, 23(4), 687–719.
  198. 198.Tang, B., & He, H. (2015). KernelADASYN: Kernel based adaptive synthetic data generation for imbalanced learning. In IEEE Congress on Evolutionary Computation (CEC), pp. 664–671.
  199. 199.Tang, S., & Chen, S. (2008). The generation mechanism of synthetic minority class examples. In 5th International Conference on Information Technology and Applications in Biomedicine (ITAB), pp. 444–447.
  200. 200.Tang, Y., Zhang, Y., Chawla, N. V., & Krasser, S. (2009). Svms modeling for highly imbalanced classification. IEEE Transactions on Systems, Man, and Cybernetics, Part B, 39(1), 281–288.
  201. 201.Thanathamathee, P., & Lursinsap, C. (2013). Handling imbalanced data sets with synthetic boundary data generation using bootstrap re-sampling and adaboost techniques. Pattern Recognition Letters, 34(12), 1339–1347.
  202. 202.Tomasev, N., & Mladenic, D. (2013). Class imbalance and the curse of minority hubs. . Knowledge-Based Systems, 53, 157–172.
  203. 203.Torgo, L., Branco, P., Ribeiro, R. P., & Pfahringer, B. (2015). Resampling strategies for regression. Expert Systems, 32(3), 465–476.
  204. 204.Torres, F. R., Carrasco-Ochoa, J. A., & Martínez Trinidad, J. F. (2016). SMOTE-D a deterministic version of SMOTE. In Pattern Recognition - 8th Mexican Conference (MCPR), pp. 177–188.
  205. 205.Triguero, I., García, S., & Herrera, F. (2015). SEG-SSC: A framework based on synthetic examples generation for self-labeled semi-supervised classification. IEEE Transactions on Cybernetics, 45(4), 622–634.
  206. 206.Verbiest, N., Ramentol, E., Cornelis, C., & Herrera, F. (2014). Preprocessing noisy imbalanced datasets using smote enhanced with fuzzy rough prototype selection. . Applied Soft Computing, 22, 511–517.
  207. 207.Vorraboot, P., Rasmequan, S., Chinnasarn, K., & Lursinsap, C. (2015). Improving classification rate constrained to imbalanced data between overlapped and non-overlapped regions by hybrid algorithms. . Neurocomputing, 152, 429–443.
  208. 208.Wang, J., Xu, M., Wang, H., & Zhang, J. (2006). Classification of imbalanced data by using the SMOTE algorithm and locally linear embedding. In 8th International Conference on Signal Processing (ICSP), Vol. 3, pp. 1–6. IEEE.
  209. 209.Wang, J., Yun, B., li Huang, P., & ao Liu, Y. (2013a). Applying threshold smote algorithm with attribute bagging to imbalanced datasets. In International Conference on Rough Sets and Knowledge Technology, pp. 221–228.
  210. 210.Wang, J., Yao, Y., Zhou, H., Leng, M., & Chen, X. (2013b). A new over-sampling technique based onsvm for imbalanced diseases data. In International Conference on Mechatronic Sciences, Electric Engineering and Computer (MEC), pp. 1224–1228.
  211. 211.Wang, Q., Luo, Z., Huang, J., Feng, Y., & Liu, Z. (2017). A novel ensemble method for imbalanced data learning: Bagging of extrapolation-SMOTE SVM. Computational Intelligence and Neuroscience, 2017, 1827016:1–1827016:11.
  212. 212.Wang, S., Minku, L. L., & Yao, X. (2013). A learning framework for online class imbalance learning. In Proceedings of the 2013 IEEE Symposium on Computational Intelligence and Ensemble Learning, CIEL 2013 - 2013 IEEE Symposium Series on Computational Intelligence, SSCI 2013, pp. 36–45.
  213. 213.Wang, S., & Yao, X. (2012). Multiclass imbalance problems: Analysis and potential solutions. IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics, 42(4), 1119–1130.
  214. 214.Wang, S., Li, Z., Chao, W., & Cao, Q. (2012). Applying adaptive over-sampling technique based on data density and cost-sensitive SVM to imbalanced learning. In The 2012 International Joint Conference on Neural Networks (IJCNN),, pp. 1–8.
  215. 215.Wang, S., Minku, L. L., Ghezzi, D., Caltabiano, D., Tiño, P., & Yao, X. (2013a). Concept drift detection for online class imbalance learning. . In IJCNN, pp. 1–10. IEEE.
  216. 216.Wang, S., Minku, L. L., & Yao, X. (2013b). Online class imbalance learning and its applications in fault detection. . International Journal of Computational Intelligence and Applications, 12(4).
  217. 217.Wang, S., Minku, L. L., & Yao, X. (2015). Resampling-based ensemble methods for online class imbalance learning. . IEEE Transactions on Knowledge Data Engineering, 27(5), 1356–1368.
  218. 218.Wang, S., & Yao, X. (2009). Diversity analysis on imbalanced data sets by using ensemble models. In Proceedings of the IEEE Symposium on Computational Intelligence and Data Mining (CIDM), pp. 324–331.
  219. 219.Wang, X., Liu, X., Japkowicz, N., & Matwin, S. (2013). Resampling and cost-sensitive methods for imbalanced multi-instance learning. In 13th IEEE International Conference on Data Mining Workshops (ICDM), pp. 808–816.
  220. 220.Wasikowski, M., & Chen, X. W. (2010). Combating the small sample class imbalance problem using feature selection. IEEE Transactions on Knowledge and Data Engineering, 22(10), 1388–1400.
  221. 221.Webb, G. I., Hyde, R., Cao, H., Nguyen, H. L., & Petitjean, F. (2016). Characterizing concept drift. Data Mining and Knowledge Discovery, 30(4), 964–994.
  222. 222.Weiss, G. M., & Provost, F. J. (2003). Learning when training data are costly: The effect of class distribution on tree induction. Journal of Artificial Intelligence Research, 19, 315–354.
  223. 223.Wilson, D. R., & Martinez, T. R. (1997). Improved heterogeneous distance functions. Journal of Artificial Intelligence Research, 6, 1–34.
  224. 224.Xie, Z., Jiang, L., Ye, T., & Li, X. (2015). A synthetic minority oversampling method based on local densities in low-dimensional space for imbalanced learning. In Database Systems for Advanced Applications - 20th International Conference (DASFAA), pp. 3–18.
  225. 225.Xu, Y. H., Le, L. P., & Tian, X. Y. (2014). Neighborhood triangular synthetic minority over-sampling technique for imbalanced prediction on small samples of chinese tourism and hospitality firms. In Seventh International Joint Conference on Computational Sciences and Optimization, pp. 534–538.
  226. 226.Yamazaki, K., Kawanabe, M., Watanabe, S., Sugiyama, M., & Móller, K.-R. (2007). Asymptotic bayesian generalization error when training and test distributions are different. . In Ghahramani, Z. (Ed.), ICML, Vol. 227 of ACM International Conference Proceeding Series, pp. 1079–1086. ACM.
  227. 227.Yin, H., & Gai, K. (2015). An empirical study on preprocessing high-dimensional class-imbalanced data for classification. In High Performance Computing and Communications (HPCC), 2015 IEEE 7th International Symposium on Cyberspace Safety and Security (CSS), 2015 IEEE 12th International Conferen on Embedded Software and Systems (ICESS), 2015 IEEE 17th International Conference on, pp. 1314–1319.
  228. 228.Yin, L., Ge, Y., Xiao, K., Wang, X., & Quan, X. (2013). Feature selection for high-dimensional imbalanced data. Neurocomputing, 105, 3–11.
  229. 229.Yongqing, Z., Min, Z., Danling, Z., Gang, M., & Daichuan, M. (2013). Improved SMOTE-Bagging and its application in imbalanced data classification. In Conference Anthology, IEEE, pp. 1–6.
  230. 230.Young, W. A., Nykl, S. L., Weckman, G. R., & Chelberg, D. M. (2015). Using voronoi diagrams to improve classification performances when modeling imbalanced datasets. Neural Computing and Applications, 26(5), 1041–1054.
  231. 231.Yun, J., Ha, J., & Lee, J. (2016). Automatic determination of neighborhood size in SMOTE. In Proceedings of the 10th International Conference on Ubiquitous Information Management and Communication (IMCOM), pp. 100:1–100:8.
  232. 232.Zhai, J., Zhang, S., & Wang, C. (2017). The classification of imbalanced large data sets based on mapreduce and ensemble of elm classifiers. . International Journal of Machine Learning and Cybernetics, 8(3), 1009–1017.
  233. 233.Zhang, H., Yang, J., Xie, J., Qian, J., & Zhang, B. (2017). Weighted sparse coding regularized nonconvex matrix regression for robust face recognition. Information Sciences, 394-395, 1–17.
  234. 234.Zhang, H., & Li, M. (2014). RWO-Sampling: A random walk over-sampling approach to imbalanced data classification. Information Fusion, 20, 99–116.
  235. 235.Zhang, H., & Wang, Z. (2011a). A normal distribution-based over-sampling approach to imbalanced data classification. In Advanced Data Mining and Applications - 7th International Conference (ADMA), pp. 83–96.
  236. 236.Zhang, L., & Wang, W. (2011b). A re-sampling method for class imbalance learning with credit data. In International Conference on Information Technology, Computer Engineering and Management Sciences (ICM), pp. 393–397.
  237. 237.Zhou, A., Qu, B. Y., Li, H., Zhao, S. Z., Suganthan, P. N., & Zhangd, Q. (2011). Multiobjective evolutionary algorithms: A survey of the state of the art. Swarm and Evolutionary Computation, 1(1), 32–49.
  238. 238.Zhou, B., Yang, C., Guo, H., & Hu, J. (2013). A quasi-linear SVM combined with assembled SMOTE for imbalanced data classification. In The 2013 International Joint Conference on Neural Networks (IJCNN), pp. 1–7.
  239. 239.Zhou, Z., & Liu, X. (2006). Training cost-sensitive neural networks with methods addressing the class imbalance problem. IEEE Transactions on Knowledge Data Engineering, 18(1), 63–77.
  240. 240.Zhu, X., Goldberg, A. B., Brachman, R., & Dietterich, T. (2009). Introduction to Semi-Supervised Learning. Morgan and Claypool Publishers.
  241. 241.Zieba, M., Tomczak, J. M., & Gonczarek, A. (2015). RBM-SMOTE: restricted boltzmann machines for synthetic minority oversampling technique. In Intelligent Information and Database Systems - 7th Asian Conference (ACIIDS), pp. 377–386.
  242. 242.Zikopoulos, P. C., Eaton, C., deRoos, D., Deutsch, T., & Lapis, G. (2011). Understanding Big Data - Analytics for Enterprise Class Hadoop and Streaming Data (1st edition). McGraw-Hill Osborne Media.
  243. 243.Zuo, Y., Zhao, J., & Xu, K. (2016). Word network topic model: a simple but general solution for short and imbalanced texts. Knowledge and Information Systems, 48(2), 379–398.

Citation

MLA
Fernandez, A., et al. “SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary”. Journal of Artificial Intelligence Research, vol. 61, 2018, pp. 863–905, https://doi.org/10.1613/JAIR.1.11192.
APA
Fernandez, A., Garcia, S., Herrera, F., & Chawla, N. V. (2018). SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary. Journal of Artificial Intelligence Research, 61, 863–905. https://doi.org/10.1613/JAIR.1.11192
Chicago
Fernandez, A., S. Garcia, F. Herrera, and N. V. Chawla. 2018. “SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary”. Journal of Artificial Intelligence Research 61: 863–905. https://doi.org/10.1613/JAIR.1.11192.
Harvard
Fernandez, A. et al. (2018) “SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary”, Journal of Artificial Intelligence Research, 61, pp. 863–905. Available at: https://doi.org/10.1613/JAIR.1.11192.
Vancouver
1. Fernandez A, Garcia S, Herrera F, Chawla NV (2018) SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary. Journal of Artificial Intelligence Research 61:863–905

BibTeX

@article{Fernandez_2018, title={SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary}, volume={61}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/JAIR.1.11192}, DOI={10.1613/jair.1.11192}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Fernandez, Alberto and Garcia, Salvador and Herrera, Francisco and Chawla, Nitesh V.}, year={2018}, month=Apr, pages={863–905} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF