keyword
binary classification
Binary classification is a supervised machine learning task that involves categorizing data instances into one of two mutually exclusive classes or outcomes. Typically formulated to predict discrete conditions such as positive or negative, true or false, or the presence versus absence of a specific trait, a model learns a decision rule from training features to assign new observations to the correct class. It is distinct from multi-class classification, which assigns inputs to one of three or more categories, although multi-class tasks can often be decomposed into multiple binary subproblems. Common algorithms designed or adapted for binary classification include logistic regression, support vector machines, decision trees, neural networks, and gradient boosting, with performance typically evaluated using metrics such as accuracy, precision, recall, F-score, and the area under the receiver operating characteristic curve.
23 items

Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, Neel Nanda
Why you should read this
Demonstrates across 113 classification tasks that sparse autoencoder probes fail to outperform standard linear baselines in challenging regimes such as data scarcity and distribution shift, critically questioning the practical downstream utility of sparse dictionary learning in language models.
Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the validity of their interpretations due to the lack of a ground truth for the concepts used by an LLM, and a growing number of works have presented problems with current SAEs. One alternative source of evidence would be demonstrating that SAEs improve performance on downstream tasks beyond existing baselines. We test this by applying SAEs to the real-world task of LLM activation probing in four regimes: data scarcity, class imbalance, label noise, and covariate shift. Due to the difficulty of detecting concepts in these challenging settings, we hypothesize that SAEs’ basis of interpretable, concept-level latents should provide a useful inductive bias. However, although SAEs occasionally perform better than baselines on individual datasets, we are unable to ensemble SAEs and baselines to consistently improve over just baseline methods. Additionally, although SAEs initially appear promising for identifying spurious correlations, detecting poor dataset quality, and training multi-token probes, we are able to achieve similar results with simple non-SAE baselines as well. Though we cannot discount SAEs’ utility on other tasks, our findings highlight the shortcomings of current SAEs and the need to rigorously evaluate interpretability methods on downstream tasks with strong baselines.
Added
2026-10-02

Fair and Optimal Classification via Post-Processing
Ruicheng Xian, Lang Yin, Han Zhao
Why you should read this
Characterizes the fundamental accuracy-fairness tradeoff under demographic parity using optimal transport and presents an optimal post-processing algorithm for multi-group, multi-class classification with rigorous sample complexity bounds.
To mitigate the bias exhibited by machine learning models, fairness criteria can be integrated into the training process to ensure fair treatment across all demographics, but it often comes at the expense of model performance. Understanding such tradeoffs, therefore, underlies the design of fair algorithms. To this end, this paper provides a complete characterization of the inherent tradeoff of demographic parity on classification problems, under the most general multi-group, multi-class, and noisy setting. Specifically, we show that the minimum error rate achievable by randomized and attribute-aware fair classifiers is given by the optimal value of a Wasserstein-barycenter problem. On the practical side, our findings lead to a simple post-processing algorithm that derives fair classifiers from score functions, which yields the optimal fair classifier when the score is Bayes optimal. We provide suboptimality analysis and sample complexity for our algorithm, and demonstrate its effectiveness on benchmark datasets.
Added
2026-09-26

Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective
Yunrui Zhao, Qianqian Xu, Yangbangyan Jiang, Peisong Wen, Qingming Huang
Why you should read this
Proposes Dist-PU, a positive-unlabeled learning framework that aligns predicted and ground-truth label distributions while utilizing entropy minimization and Mixup regularization to eliminate negative-prediction bias in deep classifiers.
Positive-Unlabeled (PU) learning tries to learn binary classifiers from a few labeled positive examples with many unlabeled ones. Compared with ordinary semi-supervised learning, this task is much more challenging due to the absence of any known negative labels. While existing cost-sensitive-based methods have achieved state-of-the-art performances, they explicitly minimize the risk of classifying unlabeled data as negative samples, which might result in a negative-prediction preference of the classifier. To alleviate this issue, we resort to a label distribution perspective for PU learning in this paper. Noticing that the label distribution of unlabeled data is fixed when the class prior is known, it can be naturally used as learning supervision for the model. Motivated by this, we propose to pursue the label distribution consistency between predicted and ground-truth label distributions, which is formulated by aligning their expectations. Moreover, we further adopt the entropy minimization and Mixup regularization to avoid the trivial solution of the label distribution consistency on unlabeled data and mitigate the consequent confirmation bias. Experiments on three benchmark datasets validate the effectiveness of the proposed method.
Added
2026-09-26

Learning with Noisy Labels
Nagarajan Natarajan, I. Dhillon, Pradeep Ravikumar, Ambuj Tewari
Why you should read this
Establishes theoretical guarantees and practical surrogate loss modifications that enable standard classifiers like biased support vector machines and weighted logistic regression to learn effectively from class-conditional noisy labels.
In this paper, we theoretically study the problem of binary classification in the presence of random classification noise — the learner, instead of seeing the true labels, sees labels that have independently been flipped with some small probability. Moreover, random label noise is class-conditional — the flip probability depends on the class. We provide two approaches to suitably modify any given surrogate loss function. First, we provide a simple unbiased estimator of any loss, and obtain performance bounds for empirical risk minimization in the presence of iid data with noisy labels. If the loss function satisfies a simple symmetry condition, we show that the method leads to an efficient algorithm for empirical minimization. Second, by leveraging a reduction of risk minimization under noisy labels to classification with weighted 0-1 loss, we suggest the use of a simple weighted surrogate loss, for which we are able to obtain strong empirical risk bounds. This approach has a very remarkable consequence — methods used in practice such as biased SVM and weighted logistic regression are provably noise-tolerant. On a synthetic non-separable dataset, our methods achieve over 88% accuracy even when 40% of the labels are corrupted, and are competitive with respect to recently proposed methods for dealing with label noise in several benchmark datasets.
Added
2026-09-25

Metrics for Multi-Class Classification: an Overview
Margherita Grandini, Enrico Bagli, Giorgio Visani
Why you should read this
Analyzes key multi-class classification metrics by detailing their individual strengths, limitations, and practical applications across model selection and hyperparameter tuning.
Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.
Added
2026-09-25

Soft Margins for AdaBoost
Gunnar Rätsch, T. Onoda, Klaus-Robert Müller
Why you should read this
Explains why AdaBoost overfits on noisy data by linking its asymptotic behavior to hard-margin maximization, and proposes soft-margin regularization techniques via gradient descent and mathematical programming to prevent outliers from degrading classification performance.
Recently ensemble methods like ADABOOST have been applied successfully in many problems, while seemingly defying the problems of overfitting. ADABOOST rarely overfits in the low noise regime, however, we show that it clearly does so for higher noise levels. Central to the understanding of this fact is the margin distribution. ADABOOST can be viewed as a constraint gradient descent in an error function with respect to the margin. We find that ADABOOST asymptotically achieves a hard margin distribution, i.e. the algorithm concentrates its resources on a few hard-to-learn patterns that are interestingly very similar to Support Vectors. A hard margin is clearly a sub-optimal strategy in the noisy case, and regularization, in our case a “mistrust” in the data, must be introduced in the algorithm to alleviate the distortions that single difficult patterns (e.g. outliers) can cause to the margin distribution. We propose several regularization methods and generalizations of the original ADABOOST algorithm to achieve a soft margin. In particular we suggest (1) regularized ADABOOSTREG where the gradient decent is done directly with respect to the soft margin and (2) regularized linear and quadratic programming (LP/QP-) ADABOOST, where the soft margin is attained by introducing slack variables. Extensive simulations demonstrate that the proposed regularized ADABOOST-type algorithms are useful and yield competitive results for noisy data.
Added
2026-09-25

Toward Open Set Recognition
W. Scheirer, A. Rocha, Archana Sapkota, T. Boult
Why you should read this
Formalizes the open set recognition problem and introduces the 1-vs-Set Machine to limit classification risk in unconstrained spaces where unseen classes emerge at test time.
To date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This article explores the nature of open set recognition, and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step towards a solution, we introduce a novel "1-vs-Set Machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale experiments performed over data from the Caltech 256, ImageNet, and Labeled Faces in the Wild sets. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks.
Added
2026-09-24

Learning From Crowds
V. Raykar, Shipeng Yu, Linda H. Zhao, Gerardo Hermosillo Valadez, Charles Florin, L. Bogoni, Linda Moy
Why you should read this
Presents an expectation-maximization framework that jointly estimates underlying ground-truth labels, measures individual annotator reliability, and trains predictive classifiers from noisy, crowdsourced multi-annotator data without requiring a gold standard.
For many supervised learning tasks it may be infeasible (or very expensive) to obtain objective and reliable labels. Instead, we can collect subjective (possibly noisy) labels from multiple experts or annotators. In practice, there is a substantial amount of disagreement among the annotators, and hence it is of great practical interest to address conventional supervised learning problems in this scenario. In this paper we describe a probabilistic approach for supervised learning when we have multiple annotators providing (possibly noisy) labels but no absolute gold standard. The proposed algorithm evaluates the different experts and also gives an estimate of the actual hidden labels. Experimental results indicate that the proposed method is superior to the commonly used majority voting baseline.
Added
2026-09-24

Ensemble Tracking
S. Avidan
Why you should read this
Proposes a visual tracking framework that treats tracking as an online binary classification problem by combining AdaBoost and mean-shift optimization to adaptively distinguish target objects from complex backgrounds across video frames.
We consider tracking as a binary classification problem, where an ensemble of weak classifiers is trained on-line to distinguish between the object and the background. The ensemble of weak classifiers is combined into a strong classifier using AdaBoost. The strong classifier is then used to label pixels in the next frame as either belonging to the object or the background, giving a confidence map. The peak of the map, and hence the new position of the object, is found using mean shift. Temporal coherence is maintained by updating the ensemble with new weak classifiers that are trained on-line during tracking. We show a realization of this method and demonstrate it on several video sequences.
Added
2026-09-24

Self-Paced Learning for Latent Variable Models
M. P. Kumar, Ben Packer, D. Koller
Why you should read this
Introduces a self-paced learning framework that dynamically selects training samples from easiest to hardest to prevent latent variable models from getting trapped in poor local optima.
Latent variable models are a powerful tool for addressing several tasks in machine learning. However, the algorithms for learning the parameters of latent variable models are prone to getting stuck in a bad local optimum. To alleviate this problem, we build on the intuition that, rather than considering all samples simultaneously, the algorithm should be presented with the training data in a meaningful order that facilitates learning. The order of the samples is determined by how easy they are. The main challenge is that often we are not provided with a readily computable measure of the easiness of samples. We address this issue by proposing a novel, iterative self-paced learning algorithm where each iteration simultaneously selects easy samples and learns a new parameter vector. The number of samples selected is governed by a weight that is annealed until the entire training data has been considered. We empirically demonstrate that the self-paced learning algorithm outperforms the state of the art method for learning a latent structural SVM on four applications: object localization, noun phrase coreference, motif finding and handwritten digit recognition.
Added
2026-09-24

In Defense of One-Vs-All Classification
Ryan Rifkin, Aldebaro Klautau
Why you should read this
Demonstrates through rigorous empirical evaluations and theoretical analysis that simple one-vs-all multiclass strategies match the accuracy of far more complex methods when using well-tuned regularized binary classifiers.
We consider the problem of multiclass classification. Our main thesis is that a simple “one-vs-all” scheme is as accurate as any other approach, assuming that the underlying binary classifiers are well-tuned regularized classifiers such as support vector machines. This thesis is interesting in that it disagrees with a large body of recent published work on multiclass classification. We support our position by means of a critical review of the existing literature, a substantial collection of carefully controlled experimental work, and theoretical arguments.
Added
2026-09-16

Training linear SVMs in linear time
Thorsten Joachims
Why you should read this
Presents a cutting-plane algorithm based on an equivalent structural optimization formulation that provably trains linear Support Vector Machines in linear time, dramatically outperforming traditional decomposition methods on large, high-dimensional datasets.
Linear Support Vector Machines (SVMs) have become one of the most prominent machine learning techniques for high-dimensional sparse data commonly encountered in applications like text classification, word-sense disambiguation, and drug design. These applications involve a large number of examples n as well as a large number of features N, while each example has only s << N non-zero features. This paper presents a Cutting-Plane Algorithm for training linear SVMs that provably has training time O(sn) for classification problems and O(s n log(n)) for ordinal regression problems. The algorithm is based on an alternative, but equivalent formulation of the SVM optimization problem. Empirically, the Cutting-Plane Algorithm is several orders of magnitude faster than decomposition methods like SVM-Light for large datasets.
Source
https://web.engr.oregonstate.edu/~huanlian/teaching/ML/2023fall/extra/linear-svm-linear-time.pdfAdded
2026-09-16

Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil, Rich Caruana
Why you should read this
Demonstrates how common supervised learning models systematically distort probability estimates and establishes that applying Platt Scaling or Isotonic Regression transforms boosted trees, random forests, and support vector machines into the most accurate probability predictors.
We examine the relationship between the predictions made by different learning algorithms and true posterior probabilities. We show that maximum margin methods such as boosted trees and boosted stumps push probability mass away from 0 and 1 yielding a characteristic sigmoid shaped distortion in the predicted probabilities. Models such as Naive Bayes, which make unrealistic independence assumptions, push probabilities toward 0 and 1. Other models such as neural nets and bagged trees do not have these biases and predict well calibrated probabilities. We experiment with two ways of correcting the biased probabilities predicted by some learning methods: Platt Scaling and Isotonic Regression. We qualitatively examine what kinds of distortions these calibration methods are suitable for and quantitatively examine how much data they need to be effective. The empirical results show that after calibration boosted trees, random forests, and SVMs predict the best probabilities.
Added
2026-09-16

Mining the peanut gallery: opinion extraction and semantic classification of product reviews
Kushal Dave, Steve Lawrence, David M. Pennock
Why you should read this
Proposes an opinion mining system that uses information retrieval scoring techniques and variable-length text patterns to classify review sentiment and synthesize unstructured web feedback into product attribute summaries.
The web contains a wealth of product reviews, but sifting through them is a daunting task. Ideally, an opinion mining tool would process a set of search results for a given item, generating a list of product attributes (quality, features, etc.) and aggregating opinions about each of them (poor, mixed, good). We begin by identifying the unique properties of this problem and develop a method for automatically distinguishing between positive and negative reviews. Our classifier draws on information retrieval techniques for feature extraction and scoring, and the results for various metrics and heuristics vary depending on the testing situation. The best methods work as well as or better than traditional machine learning. When operating on individual sentences collected from web searches, performance is limited due to noise and ambiguity. But in the context of a complete web-based tool and aided by a simple method for grouping sentences into attributes, the results are qualitatively quite useful. discussion boards and mailing list archives, as well as in Usenet via Google Groups. Users also comment on products in their personal web sites and blogs, which are then aggregated by sites such as Blogstreet.com, AllConsuming.net, and onfocus.com. When trying to locate information on a product, a general web search turns up several useful sites, but getting an overall sense of these reviews can be daunting or time-consuming. In the movie review domain, sites like Rottentomatoes.com have sprung up to try to impose some order on the void, providing ratings and brief quotes from numerous reviews and generating an aggregate opinion. Such sites even have their own category—“Review Hubs”—on Yahoo! On the commercial side, Internet clipping services like Webclipping.com, eWatch.com, and TracerLock.com watch news sites and discussion areas for mentions of a given company or product, trying to track “buzz.” Print clipping services have been providing competitive intelligence for some time. The ease of publishing on the web led to an explosion in content to be surveyed, but the same technology makes automation much more feasible. This paper describes a tool for sifting through and synthesizing product reviews, automating the sort of work done by aggregation sites or clipping services. We begin by using structured reviews for testing and training, identifying appropriate features and scoring methods from information retrieval for determining whether reviews are positive or negative. These results perform as well as traditional machine learning methods. We then use the classifier to identify and classify review sentences from the web, where classification is more difficult. However, a simple technique for identifying the relevant attributes of a product produces a subjectively useful summary.
Added
2026-09-14
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of
Hung-Hsuan Chen
Why you should read this
Derives the terminal node update formula in Gradient Boosting Classifiers through Taylor series approximation, providing complete pseudocode and mathematical intuition for optimizing logistic loss during binary classification.
The Gradient Boosting Classifier (GBC) is a widely used machine learning algorithm for binary classification, which builds decision trees iteratively to minimize prediction errors. This document explains the GBC's training and prediction processes, focusing on the computation of terminal node values , which are crucial to optimizing the logistic loss function. We derive through a Taylor series approximation and provide a step-by-step pseudocode for the algorithm's implementation. The guide explains the theory of GBC and its practical application, demonstrating its effectiveness in binary classification tasks. We provide a step-by-step example in the appendix to help readers understand.
Added
2026-09-14


Supervised learning with quantum-enhanced feature spaces
Vojtech Havlicek, Antonio D. Córcoles, Kristan Temme, Aram W. Harrow, Abhinav Kandala, Jerry M. Chow, Jay M. Gambetta
Why you should read this
Demonstrates experimental supervised classification on a superconducting quantum processor by mapping data into exponentially large Hilbert spaces via variational circuits and direct quantum kernel estimation.
Machine learning and quantum computing are two technologies each with the potential for altering how computation is performed to address previously untenable problems. Kernel methods for machine learning are ubiquitous for pattern recognition, with support vector machines (SVMs) being the most well-known method for classification problems. However, there are limitations to the successful solution to such problems when the feature space becomes large, and the kernel functions become computationally expensive to estimate. A core element to computational speed-ups afforded by quantum algorithms is the exploitation of an exponentially large quantum state space through controllable entanglement and interference. Here, we propose and experimentally implement two novel methods on a superconducting processor. Both methods represent the feature space of a classification problem by a quantum state, taking advantage of the large dimensionality of quantum Hilbert space to obtain an enhanced solution. One method, the quantum variational classifier builds on [1,2] and operates through using a variational quantum circuit to classify a training set in direct analogy to conventional SVMs. In the second, a quantum kernel estimator, we estimate the kernel function and optimize the classifier directly. The two methods present a new class of tools for exploring the applications of noisy intermediate scale quantum computers [3] to machine learning.
Added
2026-09-14

An empirical comparison of supervised learning algorithms
Rich Caruana, Alexandru Niculescu-Mizil
Why you should read this
Presents a comprehensive benchmark of ten supervised learning algorithms across nine performance metrics, revealing how post-training probability calibration enables boosted decision trees and random forests to outperform methods like neural networks and support vector machines.
We present results from a large-scale empirical comparison between ten learning methods: SVMs, neural nets, logistic regression, naive bayes, memory-based learning, random forests, decision trees, bagged trees, boosted trees, and boosted stumps. We evaluate the methods on binary classification problems using nine performance criteria: accuracy, squared error, cross-entropy, ROC Area, F-score, precision/recall break-even point, average precision, lift, and calibration. Because some models (e.g. SVMs and boosted trees) do not predict well-calibrated probabilities, we compare the performance of the algorithms both before and after calibrating their predictions with Platt Scaling and Isotonic Regression. Before scaling, the models with the best overall performance are neural nets, bagged trees, and random forests. After scaling, the best models are boosted trees, random forests, and unscaled neural nets.
Added
2026-09-14

Text Classification from Labeled and Unlabeled Documents using EM
Kamal Nigam, Andrew Kachites Mccallum, Sebastian Thrun, Tom Mitchell
Why you should read this
Demonstrates how combining Expectation-Maximization with naive Bayes leverages abundant unlabeled text to significantly reduce classification error and labeled data requirements, while introducing practical extensions to address violated generative model assumptions.
This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents. This is important because in many text classification problems obtaining training labels is expensive, while large quantities of unlabeled documents are readily available. We introduce an algorithm for learning from labeled and unlabeled documents based on the combination of Expectation-Maximization (EM) and a naive Bayes classifier. The algorithm first trains a classifier using the available labeled documents, and probabilistically labels the unlabeled documents. It then trains a new classifier using the labels for all the documents, and iterates to convergence. This basic EM procedure works well when the data conform to the generative assumptions of the model. However these assumptions are often violated in practice, and poor performance can result. We present two extensions to the algorithm that improve classification accuracy under these conditions: (1) a weighting factor to modulate the contribution of the unlabeled data, and (2) the use of multiple mixture components per class. Experimental results, obtained using text from three different real-world tasks, show that the use of unlabeled data reduces classification error by up to 30%.
Added
2026-09-12

Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
Zhilu Zhang, Mert R. Sabuncu
Why you should read this
Proposes a generalized cross entropy loss function that unifies categorical cross entropy and mean absolute error, enabling deep neural networks to learn effectively from datasets with severe label noise without sacrificing convergence speed or accuracy.
Deep neural networks (DNNs) have achieved tremendous success in a variety of applications across many disciplines. Yet, their superior performance comes with the expensive cost of requiring correctly annotated large-scale datasets. Moreover, due to DNNs' rich capacity, errors in training labels can hamper performance. To combat this problem, mean absolute error (MAE) has recently been proposed as a noise-robust alternative to the commonly-used categorical cross entropy (CCE) loss. However, as we show in this paper, MAE can perform poorly with DNNs and challenging datasets. Here, we present a theoretically grounded set of noise-robust loss functions that can be seen as a generalization of MAE and CCE. Proposed loss functions can be readily applied with any existing DNN architecture and algorithm, while yielding good performance in a wide range of noisy label scenarios. We report results from experiments conducted with CIFAR-10, CIFAR-100 and FASHION-MNIST datasets and synthetically generated noisy labels.
Added
2026-09-11

A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, Jennifer Wortman Vaughan
Why you should read this
Establishes a foundational theoretical framework for domain adaptation by bounding target generalization error using an estimable divergence metric on unlabeled data and deriving the optimal balance between source and target training samples.
Discriminative learning methods for classification perform well when training and test data are drawn from the same distribution. Often, however, we have plentiful labeled training data from a source domain but wish to learn a classifier which performs well on a target domain with a different distribution and little or no labeled training data. In this work we investigate two questions. First, under what conditions can a classifier trained from source data be expected to perform well on target data? Second, given a small amount of labeled target data, how should we combine it during training with the large amount of labeled source data to achieve the lowest target error at test time?
Added
2026-09-11
