keyword
classification accuracy
Classification accuracy is a performance evaluation metric in machine learning and statistics that measures the proportion of correctly predicted instances out of the total number of evaluated instances. Calculated as the count of correct classifications divided by the total number of predictions, it quantifies how frequently a model assigns examples to their true categorical classes. While it serves as a straightforward and common measure for assessing the overall effectiveness of a classifier across a dataset, accuracy can be less informative when classes are heavily imbalanced or when different types of misclassification errors carry unequal costs.
5 items

An Analysis of Bayesian Classifiers
Pat Langley, Wayne Iba, Kevin Thompson
Why you should read this
Presents an average-case theoretical analysis explaining why simple Bayesian classifiers achieve high accuracy across diverse learning domains despite their strong attribute independence assumptions.
In this paper we present an average-case analysis of the Bayesian classifier, a simple induction algorithm that fares remarkably well on many learning tasks. Our analysis assumes a monotone conjunctive target concept, and independent, noise-free Boolean attributes. We calculate the probability that the algorithm will induce an arbitrary pair of concept descriptions and then use this to compute the probability of correct classification over the instance space. The analysis takes into account the number of training instances, the number of attributes, the distribution of these attributes, and the level of class noise. We also explore the behavioral implications of the analysis by presenting predicted learning curves for artificial domains, and give experimental results on these domains as a check on our reasoning.
Added
2026-09-25

The CN2 Induction Algorithm
Peter Clark, T. Niblett
Why you should read this
Introduces the CN2 rule induction algorithm, which combines the noise-handling capabilities and efficiency of decision trees with the flexible if-then representation of the AQ algorithm to learn accurate, interpretable classification rules from imperfect data.
Systems for inducing concept descriptions from examples are valuable tools for assisting in the task of knowledge acquisition for expert systems. This paper presents a description and empirical evaluation of a new induction system, CN2, designed for the efficient induction of simple, comprehensible production rules in domains where problems of poor description language and/or noise may be present. Implementations of the CN2, ID3, and AQ algorithms are compared on three medical classification tasks.
Added
2026-09-24

Very Simple Classification Rules Perform Well on Most Commonly Used Datasets
ROBERT C. HOLTE
Why you should read this
Demonstrates that extremely simple one-attribute classification rules achieve accuracy nearly competitive with complex decision tree algorithms across standard benchmarks, providing an essential baseline and questioning the assumed necessity of complex machine learning models.
This article reports an empirical investigation of the accuracy of rules that classify examples on the basis of a single attribute. On most datasets studied, the best of these very simple rules is as accurate as the rules induced by the majority of machine learning systems. The article explores the implications of this finding for machine learning research and applications.
Added
2026-09-18

Rotation Forest: A New Classifier Ensemble Method
Juan J. Rodríguez, Ludmila I. Kuncheva, Carlos J. Alonso
Why you should read this
Introduces Rotation Forest, a classifier ensemble technique that applies Principal Component Analysis to random feature subsets to simultaneously boost individual decision tree accuracy and ensemble diversity, consistently outperforming Bagging, AdaBoost, and Random Forest across 33 benchmark datasets.
We propose a method for generating classifier ensembles based on feature extraction. To create the training data for a base classifier, the feature set is randomly split into K subsets (K is a parameter of the algorithm) and Principal Component Analysis (PCA) is applied to each subset. All principal components are retained in order to preserve the variability information in the data. Thus, K axis rotations take place to form the new features for a base classifier. The idea of the rotation approach is to encourage simultaneously individual accuracy and diversity within the ensemble. Diversity is promoted through the feature extraction for each base classifier. Decision trees were chosen here because they are sensitive to rotation of the feature axes, hence the name "forest." Accuracy is sought by keeping all principal components and also using the whole data set to train each base classifier. Using WEKA, we examined the Rotation Forest ensemble on a random selection of 33 benchmark data sets from the UCI repository and compared it with Bagging, AdaBoost, and Random Forest. The results were favorable to Rotation Forest and prompted an investigation into the diversity-accuracy landscape of the ensemble models. Diversity-error diagrams revealed that Rotation Forest ensembles construct individual classifiers which are more accurate than these in AdaBoost and Random Forest, and more diverse than these in Bagging, sometimes more accurate as well.
Added
2026-09-18

Using AUC and accuracy in evaluating learning algorithms
Jin Huang, Charles X. Ling
Why you should read this
Establishes formal theoretical criteria alongside empirical evidence proving that AUC is a more consistent and discriminating evaluation metric than classification accuracy for machine learning algorithms.
The area under the ROC (Receiver Operating Characteristics) curve, or simply AUC, has been recently proposed as an alternative single-number measure for evaluating the predictive ability of learning algorithms. However, no formal arguments were given as to why AUC should be preferred over accuracy. In this paper, we establish formal criteria for comparing two different measures for learning algorithms, and we show theoretically and empirically that AUC is, in general, a better measure (defined precisely) than accuracy. We then reevaluate well-established claims in machine learning based on accuracy using AUC, and obtain interesting and surprising new results. We also show that AUC is more directly associated with the net profit than accuracy in direct marketing, suggesting that learning algorithms should optimize AUC instead of accuracy in real-world applications. The conclusions drawn in this paper may make a significant impact to machine learning and data mining applications.
Added
2026-09-16
