Benchmarking Attribute Selection Techniques for Discrete Class Data Mining

Mark A. HallGeoffrey Holmes

article2003TKDE1,314 citations

Evaluates six major attribute selection methods across multiple benchmark and high-dimensional datasets to determine how ranking-based feature selection impacts the classification performance and efficiency of C4.5 and naive Bayes.

Listen

Modern data mining applications frequently suffer from poor predictive accuracy and heavy computational overhead due to datasets contaminated with irrelevant, redundant, and noisy variables. Identifying a compact subset of highly predictive attributes is critical for building efficient, interpretable models, yet practitioners lack comprehensive benchmarks to guide the selection among numerous competing techniques.

The article systematically evaluates and compares six major attribute selection methods for supervised classification on discrete class data. The primary objective is to demonstrate how these ranking-based dimensionality reduction techniques affect predictive accuracy, model complexity, and computational runtime across different types of learning algorithms.

The authors conducted rigorous experimental evaluations using sixteen standard benchmark classification datasets alongside three large-scale datasets from the UCI repository containing up to 1,557 features. Six attribute selection techniques were implemented: Information Gain, ReliefF, Principal Component Analysis, Correlation-based Feature Selection (CFS), a Consistency-based evaluator, and a Wrapper method. These techniques produced ranked lists of variables that were cross-validated to identify optimal subsets for two fundamentally distinct learning schemes: a probabilistic model (naive Bayes) and a decision tree inducer (C4.5).

The benchmark revealed four key findings. First, no single attribute selection method was best across all situations, as performance strongly depended on the underlying inductive bias of the learning algorithm. Second, the Wrapper approach achieved the highest classification accuracy for naive Bayes (scoring 34 wins and 4 losses), but it was computationally prohibitive on large datasets, taking an estimated 140 days on the largest benchmark. Third, for decision trees (C4.5), ReliefF and the Consistency method delivered the top predictive performance because they effectively capture attribute interactions that decision trees exploit during early node splitting. Fourth, CFS consistently eliminated the most unnecessary data, retaining only 42% to 48% of the original attributes on average, while producing the smallest, most interpretable decision trees and executing significantly faster than wrapper or instance-based techniques. Conversely, unsupervised Principal Component Analysis consistently degraded predictive accuracy across models.

These findings indicate that reducing data dimensionality by roughly 50% generally maintains or improves classification performance while producing simpler, more explainable models. For decision-makers, choosing the right attribute selection technique directly affects operational efficiency, project timelines, and model interpretability. Implementing computationally heavy wrapper models introduces severe timeline risks on large datasets without guaranteeing better results than faster heuristic filters.

Organizations should align their feature selection method with their modeling approach and resource constraints. When computational budget allows, Wrapper methods remain the top choice for naive Bayes, whereas ReliefF or Consistency methods are recommended for tree-based models where feature interactions matter. When processing speed, scalability, and compact model size are essential, teams should deploy CFS as a fast, balanced default. Highly complex pipelines should avoid unsupervised Principal Component Analysis for supervised discrete classification.

Confidence in these findings is high for standard tabular classification tasks given the extensive cross-validation methodology. However, caution is warranted when extrapolating these results to non-discrete regression problems, extremely high-dimensional sparse datasets, or architectures beyond naive Bayes and decision trees, which were outside the scope of this evaluation.

Cover for Benchmarking Attribute Selection Techniques for Discrete Class Data Mining

Abstract

Data engineering is generally considered to be a central issue in the development of data mining applications. The success of many learning schemes, in their attempts to construct models of data, hinges on the reliable identification of a small set of highly predictive attributes. The inclusion of irrelevant, redundant and noisy attributes in the model building process phase can result in poor predictive performance and increased computation.

Attribute selection generally involves a combination of search and attribute utility estimation plus evaluation with respect to specific learning schemes. This leads to a large number of possible permutations and has led to a situation where very few benchmark studies have been conducted.

This paper presents a benchmark comparison of several attribute selection methods for supervised classification. All the methods produce an attribute ranking, a useful devise for isolating the individual merit of an attribute. Attribute selection is achieved by cross-validating the attribute rankings with respect to a classification learner to find the best attributes. Results are reported for a selection of standard data sets and two diverse learning schemes C4.5 and naive Bayes.

Table of Contents

  • 1 Introduction
  • 2 Attribute Selection Techniques
  • 2.1 Information Gain Attribute Ranking
  • 2.2 Relief
  • 2.3 Principal Components
  • 2.4 CFS
  • 2.5 Consistency-based Subset Evaluation
  • 2.6 Wrapper Subset Evaluation
  • 3 Experimental Methodology
  • 3.1 Weka Experiment Editor
  • 4 Results on Sixteen Benchmark Data Sets
  • 5 Results on Large Data Sets
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — Modified Forward Selection for Attribute Ranking

    algorithm

    The modified forward selection search procedure adapts heuristic subset evaluators to generate a complete ranked list of individual features by forcing greedy forward hill-climbing search to continue until the entire feature set is included.

    Input: Full feature set F={f1,f2,…,fd}F = \{f_1, f_2, \dots, f_d\}, subset evaluation function M(S)M(S)
    Output: Ranked list of attributes RR
    S←∅S \leftarrow \emptyset
    R←[]R \leftarrow []
    U←FU \leftarrow F
    while U≠∅U \neq \emptyset do
        best_f←argmax⁡f∈UM(S∪{f})best\_f \leftarrow \operatorname{argmax}_{f \in U} M(S \cup \{f\})
        S←S∪{best_f}S \leftarrow S \cup \{best\_f\}
        U←U∖{best_f}U \leftarrow U \setminus \{best\_f\}
        R←append⁡(R,best_f)R \leftarrow \operatorname{append}(R, best\_f)
    end while
    return RR

    Unlike standard forward hill-climbing search—which terminates as soon as no single feature addition strictly increases the subset evaluation score M(S)M(S)—the modified algorithm forces the search to continue until all attributes are exhausted. In each iteration ii, the remaining feature that yields the maximum subset score when combined with the currently selected i−1i-1 features is identified, added to the subset, and recorded at rank ii in the output list. This converts subset evaluators such as Correlation-based Feature Selection (CFS), Consistency-based evaluation (CNS), and Wrapper evaluation (WRP) into attribute rankers.

  2. Knowl 2 — Optimal Subset Determination via Cross-Validated Feature Ranking Evaluation

    experimental setup

    To identify the best attribute subset from an ordered attribute ranking R=(f(1),f(2),…,f(d))R = (f_{(1)}, f_{(2)}, \dots, f_{(d)}) generated on training data, a forward cross-validation procedure is executed strictly on the training partition.

    For each candidate subset size k∈{1,2,…,d}k \in \{1, 2, \dots, d\}, a feature subset Sk={f(1),…,f(k)}S_k = \{f_{(1)}, \dots, f_{(k)}\} containing the top kk ranked features is formed. A 10-fold cross-validation is then performed on the training split using the target induction algorithm (e.g., Naive Bayes or C4.5) trained only on SkS_k. The subset size k∗k^* that yields the highest cross-validated classification accuracy across the training folds is chosen as optimal:

    k∗=argmax⁡k∈{1,…,d}Accuracy⁡CV(Sk)k^* = \operatorname{argmax}_{k \in \{1, \dots, d\}} \operatorname{Accuracy}_{\text{CV}}(S_k)

    The final classifier is subsequently trained on the complete training split using only the top k∗k^* features and evaluated on the unseen test split. This wrapper evaluation isolates feature ranking from final evaluation and prevents information leakage between training and testing sets.

  3. Knowl 3 — Correlation-Based Feature Selection Metric and Search

    model/method

    Correlation-Based Feature Selection (CFS) evaluates an attribute subset SS containing kk features based on the principle that an optimal feature subset contains features highly correlated with the target class yet mutually uncorrelated with one another. The heuristic subset merit MeritSMerit_S is computed as:

    MeritS=k rcf‾k+k(k−1) rff‾Merit_S = \frac{k \, \overline{r_{cf}}}{\sqrt{k + k(k - 1) \, \overline{r_{ff}}}}

    where kk is the number of features in SS, rcf‾\overline{r_{cf}} is the average correlation between the features in SS and the discrete class variable CC, and rff‾\overline{r_{ff}} is the average pairwise correlation across all feature pairs within SS.

    For discrete nominal variables, correlation between any two variables XX and YY is measured using symmetrical uncertainty (SUSU):

    SU(X,Y)=2.0×[H(X)+H(Y)−H(X,Y)H(X)+H(Y)]SU(X, Y) = 2.0 \times \left[ \frac{H(X) + H(Y) - H(X, Y)}{H(X) + H(Y)} \right]

    where H(X)=−∑x∈Xp(x)log⁡2p(x)H(X) = -\sum_{x \in X} p(x) \log_2 p(x) is Shannon entropy and H(X,Y)H(X, Y) is the joint entropy of XX and YY. Continuous numeric attributes are discretized prior to correlation computation using the entropy-based Minimum Description Length (MDL) discretization method of Fayyad and Irani.

  4. Knowl 4 — Consistency-Based Attribute Subset Evaluation Metric

    model/method

    The consistency-based subset evaluator assesses a candidate attribute subset ss by measuring the degree to which the attribute value combinations of ss partition the training data into pure, single-class groups.

    For a dataset with NN instances and an attribute subset ss, instances are grouped by their distinct combinations of attribute values on ss. Let JJ denote the total number of distinct attribute value combinations present in the dataset. For the ii-th combination, let ∣Di∣|D_i| denote the total number of instances exhibiting that attribute combination, and let ∣Mi∣|M_i| denote the number of instances belonging to the majority class within that combination. The subset consistency score ConsistencysConsistency_s is given by:

    Consistencys=1−∑i=1J(∣Di∣−∣Mi∣)NConsistency_s = 1 - \frac{\sum_{i=1}^{J} (|D_i| - |M_i|)}{N}

    A subset achieves maximum consistency (Consistencys=1.0Consistency_s = 1.0) if and only if every unique combination of attribute values maps exclusively to instances of one class. Continuous attributes are converted to discrete intervals using Fayyad and Irani's MDL discretizer prior to consistency evaluation.

  5. Knowl 5 — ReliefF Attribute Ranking for Multi-Class Noisy Domains

    algorithm

    ReliefF is an instance-based attribute ranking algorithm that estimates feature importance by measuring how effectively feature values distinguish between nearest neighbor instances from the same versus different classes.

    Input: Training set DD, sampling iterations m=250m = 250, neighbor count k=10k = 10, feature set AA
    Output: Weight vector WW for all attributes in AA
    for each attribute a∈Aa \in A do
        W[a]←0.0W[a] \leftarrow 0.0
    end for
    for i=1i = 1 to mm do
        Randomly select an instance R∈DR \in D
        Find kk nearest instances HjH_j to RR from class (R)(R) (nearest hits)
        for each class C≠class⁡(R)C \neq \operatorname{class}(R) do
            Find kk nearest instances Mj(C)M_j(C) to RR from class CC (nearest misses)
        end for
        for each attribute a∈Aa \in A do
            W[a]←W[a]−∑j=1kdiff⁡(a,R,Hj)m⋅k+∑C≠class⁡(R)[P(C)1−P(class⁡(R))∑j=1kdiff⁡(a,R,Mj(C))m⋅k]W[a] \leftarrow W[a] - \sum_{j=1}^k \frac{\operatorname{diff}(a, R, H_j)}{m \cdot k} + \sum_{C \neq \operatorname{class}(R)} \left[ \frac{P(C)}{1 - P(\operatorname{class}(R))} \sum_{j=1}^k \frac{\operatorname{diff}(a, R, M_j(C))}{m \cdot k} \right]
        end for
    end for
    return WW

    The difference function diff⁡(a,I1,I2)\operatorname{diff}(a, I_1, I_2) between instances I1I_1 and I2I_2 on attribute aa is defined as:

    diff⁡(a,I1,I2)={0if I1[a]=I2[a] (for discrete a)1if I1[a]≠I2[a] (for discrete a)∣I1[a]−I2[a]∣max⁡(a)−min⁡(a)for continuous a\operatorname{diff}(a, I_1, I_2) = \begin{cases} 0 & \text{if } I_1[a] = I_2[a] \text{ (for discrete } a\text{)} \\ 1 & \text{if } I_1[a] \neq I_2[a] \text{ (for discrete } a\text{)} \\ \frac{|I_1[a] - I_2[a]|}{\max(a) - \min(a)} & \text{for continuous } a \end{cases}

    where P(C)P(C) is the prior probability of class CC estimated from the training data.

  6. Knowl 6 — Classification Accuracy Benchmark for Naive Bayes

    data/table

    Six attribute selection methods—Information Gain (IG), ReliefF (RLF), Consistency-based selection (CNS), Principal Components Analysis (PC), Correlation-based Feature Selection (CFS), and Wrapper selection (WRP)—were evaluated with Naive Bayes against the unselected baseline (NB) across 15 standard UCI datasets using 10 runs of 10-fold cross-validation.

    Data Set NB IG RLF CNS PC CFS WRP
    zoo 95.04 94.34 93.37 93.85 93.86 93.94 94.34
    heart-c 83.83 82.54 82.12 82.28 81.85 82.64 82.68
    ionosphere 82.60 88.78 89.52 89.95 90.72 89.75 91.28
    soybean 92.90 92.43 92.56 92.81 90.93 92.46 92.64
    glass2 62.33 67.42 63.83 68.31 66.74 71.08 75.06
    vote 90.19 95.63 95.33 95.82 92.32 95.63 95.93
    heart-stat 84.37 85.11 86.00 83.48 82.07 85.07 85.00
    lymph 83.24 82.63 81.47 82.55 79.67 82.35 84.11
    labor 93.93 89.17 90.97 92.00 89.77 89.20 85.77
    diabetes 75.73 76.24 75.95 75.64 74.42 76.19 76.12
    breast-c 73.12 72.84 70.99 71.79 73.54 73.01 72.28
    credit-g 74.98 74.36 74.49 74.06 73.30 74.33 74.35
    segment 80.10 87.17 86.97 85.98 90.03 89.03 89.57
    horse colic 78.28 83.20 82.58 82.77 78.56 83.01 82.61
    anneal 86.51 87.06 89.17 89.71 90.65 87.16 92.91

    Pairwise statistical significance was tested at the 1% level using a two-sided paired tt-test across all pairs of schemes. Ranking methods by net wins (Wins minus Losses) demonstrates that Wrapper selection and CFS achieve the highest performance when paired with Naive Bayes:

    Scheme Wins−-Losses Wins Losses
    WRP 30 34 4
    CFS 7 21 14
    CNS 2 21 19
    IG -2 17 19
    RLF -3 19 22
    NB -7 28 35
    PC -27 17 44
  7. Knowl 7 — Classification Accuracy Benchmark for C4.5 Decision Trees

    data/table

    Six attribute selection methods—Information Gain (IG), Correlation-based Feature Selection (CFS), Consistency-based selection (CNS), ReliefF (RLF), Wrapper (WRP), and Principal Components Analysis (PC)—were benchmarked with C4.5 decision trees (release 8) against the unselected C4.5 baseline across 15 standard UCI datasets using 10 runs of 10-fold cross-validation.

    Data Set C4.5 IG CFS CNS RLF WRP PC
    zoo 92.26 91.65 91.06 93.65 92.95 90.45 91.49
    heart-stat 78.67 84.52 85.33 84.11 82.00 82.11 82.22
    ionosphere 89.74 89.40 91.09 91.05 91.43 91.80 88.80
    diabetes 73.74 73.92 73.67 73.71 73.58 73.50 71.51
    vote 96.46 95.84 95.65 95.98 95.79 95.74 92.07
    credit-g 71.18 72.72 72.99 72.20 71.63 72.23 69.34
    soybean 92.48 92.40 91.14 92.43 92.43 92.19 83.75
    heart-c 76.64 78.95 79.11 80.23 80.40 77.00 82.65
    glass2 77.97 78.35 78.53 77.05 79.53 76.53 66.41
    labor 80.20 80.60 81.00 79.73 79.53 78.33 88.60
    lymph 75.50 73.09 73.41 75.43 76.83 76.63 74.60
    breast-c 73.87 73.75 73.70 72.24 72.77 73.43 70.62
    segment 96.90 96.81 96.94 96.87 96.89 96.92 93.95
    anneal 98.58 98.72 98.47 98.65 98.73 98.66 96.26
    horse colic 85.44 84.18 83.94 84.00 84.90 84.14 78.18

    Pairwise comparison (two-sided paired tt-test at the 1% level) demonstrates that ReliefF and Consistency-based selection achieve the highest net accuracy rankings when paired with C4.5:

    Scheme Wins−-Losses Wins Losses
    RLF 15 22 7
    CNS 12 20 8
    C4.5 7 23 16
    CFS 5 21 16
    WRP 5 17 12
    IG 2 15 13
    PC -46 12 58
  8. Knowl 8 — Inducer-Specific Biases in Attribute Selection Effectiveness

    empirical result

    The relative effectiveness of attribute selection techniques is strongly dependent on the learning algorithm's inductive bias:

    1. Naive Bayes: Naive Bayes relies on the conditional independence assumption among attributes given the class. As a result, Wrapper selection (which tailors the subset directly to Naive Bayes) and Correlation-based Feature Selection (CFS, which explicitly penalizes redundant, inter-correlated features) achieve the best performance, yielding +30+30 and +7+7 net wins respectively over other schemes across benchmark datasets.
    2. C4.5 Decision Trees: C4.5 constructs univariate hierarchical splits and can exploit multi-feature interactions if relevant interactive features are selected together. Methods capable of detecting attribute interactions—specifically ReliefF (+15+15 net wins) and Consistency-based selection (+12+12 net wins)—outperform Wrapper selection (+5+5 net wins) and the unselected C4.5 baseline (+7+7 net wins).
    3. Wrapper Limitations with C4.5: The Wrapper underperforms ReliefF and Consistency on C4.5 for two primary reasons:
      • Greedy forward search: Forward selection evaluates attributes incrementally and can fail to identify strongly interacting attribute groups during early iterations.
      • Sample size reduction in internal CV: The Wrapper's internal 5-fold cross-validation on training data reduces the number of instances available to construct trees during evaluation, leading to noisy and unreliable subset evaluations for sample-sensitive decision tree learners.
  9. Knowl 9 — Decision Tree Complexity Reduction via Attribute Selection

    data/table

    Applying attribute selection prior to decision tree induction with C4.5 significantly reduces final tree complexity (measured by total node count) across 15 standard UCI datasets.

    Data Set C4.5 IG CFS CNS RLF WRP PC
    zoo 15.64 13.22 13.74 13.44 13.04 13.98 13.02
    heart-stat 34.84 12.12 11.98 13.52 13.66 14.92 4.82
    ionosphere 26.58 21.84 16.64 17.14 17.22 13.90 20.04
    diabetes 41.54 14.62 15.92 16.54 16.74 17.06 30.52
    vote 10.64 9.44 8.64 9.92 9.00 9.72 20.44
    credit-g 125.05 57.34 60.39 61.82 68.52 63.48 10.98
    soybean 92.27 86.50 88.29 92.25 91.21 90.75 88.84
    heart-c 42.34 19.72 19.45 22.48 23.17 24.20 8.16
    glass2 23.78 14.88 16.28 16.26 17.12 16.22 11.18
    labor 6.96 6.22 6.10 6.18 5.48 6.13 5.88
    lymph 27.41 14.71 14.35 12.26 14.56 14.43 18.18
    breast-c 12.38 10.47 10.26 15.09 11.80 8.42 7.72
    segment 81.86 80.82 80.26 79.44 80.96 79.50 119.00
    anneal 49.75 48.45 50.06 46.83 46.73 48.63 38.94
    horse colic 8.57 21.18 25.75 8.81 20.64 20.90 6.42

    Pairwise comparison across all pairs of schemes ranks the methods by net tree size reduction (Wins minus Losses):

    Scheme Wins−-Losses Wins Losses
    PC 21 47 26
    CFS 15 30 15
    IG 13 29 16
    RLF 7 25 18
    CNS 6 26 20
    WRP 0 22 22
    C4.5 -62 6 68

    Although Principal Components yields the smallest trees, it causes severe degradation in predictive accuracy. CFS is the most effective selector for reducing tree size while maintaining high classification accuracy, producing significantly smaller trees than baseline C4.5 on 11 datasets.

  10. Knowl 10 — Dimensionality Reduction and Computational Speed Trade-Offs

    empirical result

    Systematic evaluation of attribute selectors across benchmark datasets highlights distinct trade-offs between feature reduction rate, computational runtime, and scaling behavior:

    1. Degree of Feature Reduction: Correlation-based Feature Selection (CFS) consistently selects the most compact attribute subsets, retaining on average ≈48%\approx 48\% of attributes for Naive Bayes and ≈42%\approx 42\% for C4.5 (achieving +24+24 net wins in feature reduction rankings for both learners). The Wrapper retains ≈50%\approx 50\% for Naive Bayes and ≈42%\approx 42\% for C4.5. ReliefF retains slightly larger feature subsets (≈52%\approx 52\% for C4.5, ranking lowest in feature reduction among filters), which enables it to maintain superior classification accuracy on interaction-heavy datasets.
    2. Performance on Large Datasets: On large datasets (arrhythmia with 227 attributes, anonymous with 293 attributes, internet-ads with 1557 attributes), CFS reduces the feature space down to 1%–22%1\%\text{--}22\% of original attributes (e.g., selecting 6 of 227 features on arrhythmia and 11 of 293 on anonymous for Naive Bayes) while matching or improving upon baseline accuracy.
    3. Computational Efficiency: CFS and Information Gain are the fastest selectors overall (+50+50 and +49+49 net wins in execution time for Naive Bayes). Consistency-based selection is the fastest method for C4.5 (+34+34 net wins) because identifying high-quality attribute rankings early accelerates the internal cross-validation loop by producing smaller, faster-to-build decision trees. The Wrapper is by far the slowest scheme (−66-66 net wins for both learners) and becomes computationally intractable on high-dimensional domains (e.g., requiring an estimated 140 days of CPU time to complete forward selection on internet-ads).

Coverage note — None was omitted; all key methods, equations, algorithms, experimental designs, benchmark data tables, and empirical analyses from the paper are fully represented.

References

  1. 1.H. Almuallim and T. G. Dietterich. Learning with many irrelevant features. In Proceedings of the Ninth National Conference on Artificial Intelligence, pages 547–552. AAAI Press, 1991.
  2. 2.C. Blake, E. Keogh, and C. J. Merz. UCI Repository of Machine Learning Data Bases. University of California, Department of Information and Computer Science, Irvine, CA, 1998. [http://www.ics.uci.edu/~mlearn/MLRepository.html].
  3. 3.Avrim Blum and Pat Langley. Selection of relevant features and examples in machine learning. Artificial Intelligence, 97(1-2):245–271, 1997.
  4. 4.M. Dash and H. Liu. Feature selection for classification. Intelligent Data Analysis, 1(3), 1997.
  5. 5.S. Dumais, J. Platt, D. Heckerman, and M. Sahami. Inductive learning algorithms and representations for text categorization. In Proceedings of the International Conference on Information and Knowledge Management, pages 148–155, 1998.
  6. 6.U. M. Fayyad and K. B. Irani. Multi-interval discretisation of continuous-valued attributes. In Proceedings of the Thirteenth International Joint Conference on Artificial Intelligence, pages 1022–1027. Morgan Kaufmann, 1993.
  7. 7.M. A. Hall. Correlation-based feature selection for machine learning. PhD thesis, Department of Computer Science, University of Waikato, Hamilton, New Zealand, 1998.
  8. 8.Mark Hall. Correlation-based feature selection for discrete and numeric class machine learning. In Proc. of the 17th International Conference on Machine Learning (ICML2000), 2000.
  9. 9.K. Kira and L. Rendell. A practical approach to feature selection. In Proceedings of the Ninth International Conference on Machine Learning, pages 249–256. Morgan Kaufmann, 1992.
  10. 10.Ron Kohavi and George H. John. Wrappers for feature subset selection. Artificial Intelligence, 97:273–324, 1997.
  11. 11.I. Kononenko. Estimating attributes: Analysis and extensions of relief. In Proceedings of the Seventh European Conference on Machine Learning, pages 171–182. Springer-Verlag, 1994.
  12. 12.P. Langley, W. Iba, and K. Thompson. An analysis of Bayesian classifiers. In Proc. of the Tenth National Conference on Artificial Intelligence, pages 223–228, San Jose, CA, 1992. AAAI Press. [Langley92.ps.gz, from http://www.isle.org/~langley/papers/bayes.aaai92.ps].
  13. 13.H. Liu and R. Setiono. A probabilistic approach to feature selection: A filter solution. In Proceedings of the 13th International Conference on Machine Learning, pages 319–327. Morgan Kaufmann, 1996.
  14. 14.J. R. Quinlan. C4.5: Programs for Machine Learning. Morgan Kaufmann, San Mateo, CA., 1993.
  15. 15.M. Sikonja and I. Kononenko. An adaptation of relief for attribute estimation in regression. In Proceedings of the Fourteenth International Conference (ICML’97), pages 296–304. Morgan Kaufmann, 1997.
  16. 16.Yiming Yang and Jan O. Pedersen. A comparative study on feature selection in text categorization. In International Conference on Machine Learning, pages 412–420, 1997.

Citation

MLA
Hall, M. A., and G. Holmes. “Benchmarking Attribute Selection Techniques for Discrete Class Data Mining”. IEEE Transactions on Knowledge and Data Engineering, vol. 15, no. 6, 2003, pp. 1437–47, https://doi.org/10.1109/TKDE.2003.1245283.
APA
Hall, M. A., & Holmes, G. (2003). Benchmarking attribute selection techniques for discrete class data mining. IEEE Transactions on Knowledge and Data Engineering, 15(6), 1437–1447. https://doi.org/10.1109/TKDE.2003.1245283
Chicago
Hall, M. A., and G. Holmes. 2003. “Benchmarking Attribute Selection Techniques for Discrete Class Data Mining”. IEEE Transactions on Knowledge and Data Engineering 15 (6): 1437–47. https://doi.org/10.1109/TKDE.2003.1245283.
Harvard
Hall, M.A. and Holmes, G. (2003) “Benchmarking attribute selection techniques for discrete class data mining”, IEEE Transactions on Knowledge and Data Engineering, 15(6), pp. 1437–1447. Available at: https://doi.org/10.1109/TKDE.2003.1245283.
Vancouver
1. Hall MA, Holmes G (2003) Benchmarking attribute selection techniques for discrete class data mining. IEEE Transactions on Knowledge and Data Engineering 15:1437–1447

BibTeX

@article{Hall_2003, title={Benchmarking attribute selection techniques for discrete class data mining}, volume={15}, ISSN={1041-4347}, url={http://dx.doi.org/10.1109/TKDE.2003.1245283}, DOI={10.1109/tkde.2003.1245283}, number={6}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Hall, M.A. and Holmes, G.}, year={2003}, month=Nov, pages={1437–1447} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF