Built independently by an author, for readers. Read the story and support ChapterPal

keyword

binary classification

Binary classification is a supervised machine learning task that involves categorizing data instances into one of two mutually exclusive classes or outcomes. Typically formulated to predict discrete conditions such as positive or negative, true or false, or the presence versus absence of a specific trait, a model learns a decision rule from training features to assign new observations to the correct class. It is distinct from multi-class classification, which assigns inputs to one of three or more categories, although multi-class tasks can often be decomposed into multiple binary subproblems. Common algorithms designed or adapted for binary classification include logistic regression, support vector machines, decision trees, neural networks, and gradient boosting, with performance typically evaluated using metrics such as accuracy, precision, recall, F-score, and the area under the receiver operating characteristic curve.

23 items

Are Sparse Autoencoders Useful? A Case Study in Sparse Probing

Are Sparse Autoencoders Useful? A Case Study in Sparse Probing

Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, Neel Nanda

Why you should read this

Demonstrates across 113 classification tasks that sparse autoencoder probes fail to outperform standard linear baselines in challenging regimes such as data scarcity and distribution shift, critically questioning the practical downstream utility of sparse dictionary learning in language models.

Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the validity of their interpretations due to the lack of a ground truth for the concepts used by an LLM, and a growing number of works have presented problems with current SAEs. One alternative source of evidence would be demonstrating that SAEs improve performance on downstream tasks beyond existing baselines. We test this by applying SAEs to the real-world task of LLM activation probing in four regimes: data scarcity, class imbalance, label noise, and covariate shift. Due to the difficulty of detecting concepts in these challenging settings, we hypothesize that SAEs’ basis of interpretable, concept-level latents should provide a useful inductive bias. However, although SAEs occasionally perform better than baselines on individual datasets, we are unable to ensemble SAEs and baselines to consistently improve over just baseline methods. Additionally, although SAEs initially appear promising for identifying spurious correlations, detecting poor dataset quality, and training multi-token probes, we are able to achieve similar results with simple non-SAE baselines as well. Though we cannot discount SAEs’ utility on other tasks, our findings highlight the shortcomings of current SAEs and the need to rigorously evaluate interpretability methods on downstream tasks with strong baselines.

Added

2026-10-02

Fair and Optimal Classification via Post-Processing

Fair and Optimal Classification via Post-Processing

Ruicheng Xian, Lang Yin, Han Zhao

OrganizationsUniversity of Illinois Urbana-Champaign

Why you should read this

Characterizes the fundamental accuracy-fairness tradeoff under demographic parity using optimal transport and presents an optimal post-processing algorithm for multi-group, multi-class classification with rigorous sample complexity bounds.

To mitigate the bias exhibited by machine learning models, fairness criteria can be integrated into the training process to ensure fair treatment across all demographics, but it often comes at the expense of model performance. Understanding such tradeoffs, therefore, underlies the design of fair algorithms. To this end, this paper provides a complete characterization of the inherent tradeoff of demographic parity on classification problems, under the most general multi-group, multi-class, and noisy setting. Specifically, we show that the minimum error rate achievable by randomized and attribute-aware fair classifiers is given by the optimal value of a Wasserstein-barycenter problem. On the practical side, our findings lead to a simple post-processing algorithm that derives fair classifiers from score functions, which yields the optimal fair classifier when the score is Bayes optimal. We provide suboptimality analysis and sample complexity for our algorithm, and demonstrate its effectiveness on benchmark datasets.

Added

2026-09-26

Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective

Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective

Yunrui Zhao, Qianqian Xu, Yangbangyan Jiang, Peisong Wen, Qingming Huang

OrganizationsInstitute of Computing Technology, Chinese Academy of SciencesInstitute of Information Engineering, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Proposes Dist-PU, a positive-unlabeled learning framework that aligns predicted and ground-truth label distributions while utilizing entropy minimization and Mixup regularization to eliminate negative-prediction bias in deep classifiers.

Positive-Unlabeled (PU) learning tries to learn binary classifiers from a few labeled positive examples with many unlabeled ones. Compared with ordinary semi-supervised learning, this task is much more challenging due to the absence of any known negative labels. While existing cost-sensitive-based methods have achieved state-of-the-art performances, they explicitly minimize the risk of classifying unlabeled data as negative samples, which might result in a negative-prediction preference of the classifier. To alleviate this issue, we resort to a label distribution perspective for PU learning in this paper. Noticing that the label distribution of unlabeled data is fixed when the class prior is known, it can be naturally used as learning supervision for the model. Motivated by this, we propose to pursue the label distribution consistency between predicted and ground-truth label distributions, which is formulated by aligning their expectations. Moreover, we further adopt the entropy minimization and Mixup regularization to avoid the trivial solution of the label distribution consistency on unlabeled data and mitigate the consequent confirmation bias. Experiments on three benchmark datasets validate the effectiveness of the proposed method.

Added

2026-09-26

Learning with Noisy Labels

Learning with Noisy Labels

Nagarajan Natarajan, I. Dhillon, Pradeep Ravikumar, Ambuj Tewari

OrganizationsUniversity of MichiganUniversity of Texas at Austin

Why you should read this

Establishes theoretical guarantees and practical surrogate loss modifications that enable standard classifiers like biased support vector machines and weighted logistic regression to learn effectively from class-conditional noisy labels.

In this paper, we theoretically study the problem of binary classification in the presence of random classification noise — the learner, instead of seeing the true labels, sees labels that have independently been flipped with some small probability. Moreover, random label noise is class-conditional — the flip probability depends on the class. We provide two approaches to suitably modify any given surrogate loss function. First, we provide a simple unbiased estimator of any loss, and obtain performance bounds for empirical risk minimization in the presence of iid data with noisy labels. If the loss function satisfies a simple symmetry condition, we show that the method leads to an efficient algorithm for empirical minimization. Second, by leveraging a reduction of risk minimization under noisy labels to classification with weighted 0-1 loss, we suggest the use of a simple weighted surrogate loss, for which we are able to obtain strong empirical risk bounds. This approach has a very remarkable consequence — methods used in practice such as biased SVM and weighted logistic regression are provably noise-tolerant. On a synthetic non-separable dataset, our methods achieve over 88% accuracy even when 40% of the labels are corrupted, and are competitive with respect to recently proposed methods for dealing with label noise in several benchmark datasets.

Added

2026-09-25

Soft Margins for AdaBoost

Soft Margins for AdaBoost

Gunnar Rätsch, T. Onoda, Klaus-Robert Müller

OrganizationsCentral Research Institute of Electric Power IndustryGMD FIRSTUniversity of Potsdam

Why you should read this

Explains why AdaBoost overfits on noisy data by linking its asymptotic behavior to hard-margin maximization, and proposes soft-margin regularization techniques via gradient descent and mathematical programming to prevent outliers from degrading classification performance.

Recently ensemble methods like ADABOOST have been applied successfully in many problems, while seemingly defying the problems of overfitting. ADABOOST rarely overfits in the low noise regime, however, we show that it clearly does so for higher noise levels. Central to the understanding of this fact is the margin distribution. ADABOOST can be viewed as a constraint gradient descent in an error function with respect to the margin. We find that ADABOOST asymptotically achieves a hard margin distribution, i.e. the algorithm concentrates its resources on a few hard-to-learn patterns that are interestingly very similar to Support Vectors. A hard margin is clearly a sub-optimal strategy in the noisy case, and regularization, in our case a “mistrust” in the data, must be introduced in the algorithm to alleviate the distortions that single difficult patterns (e.g. outliers) can cause to the margin distribution. We propose several regularization methods and generalizations of the original ADABOOST algorithm to achieve a soft margin. In particular we suggest (1) regularized ADABOOSTREG where the gradient decent is done directly with respect to the soft margin and (2) regularized linear and quadratic programming (LP/QP-) ADABOOST, where the soft margin is attained by introducing slack variables. Extensive simulations demonstrate that the proposed regularized ADABOOST-type algorithms are useful and yield competitive results for noisy data.

Added

2026-09-25

Toward Open Set Recognition

Toward Open Set Recognition

W. Scheirer, A. Rocha, Archana Sapkota, T. Boult

OrganizationsUniversity of CampinasUniversity of Colorado Colorado Springs

Why you should read this

Formalizes the open set recognition problem and introduces the 1-vs-Set Machine to limit classification risk in unconstrained spaces where unseen classes emerge at test time.

To date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This article explores the nature of open set recognition, and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step towards a solution, we introduce a novel "1-vs-Set Machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale experiments performed over data from the Caltech 256, ImageNet, and Labeled Faces in the Wild sets. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks.

Added

2026-09-24

Self-Paced Learning for Latent Variable Models

Self-Paced Learning for Latent Variable Models

M. P. Kumar, Ben Packer, D. Koller

OrganizationsStanford University

Why you should read this

Introduces a self-paced learning framework that dynamically selects training samples from easiest to hardest to prevent latent variable models from getting trapped in poor local optima.

Latent variable models are a powerful tool for addressing several tasks in machine learning. However, the algorithms for learning the parameters of latent variable models are prone to getting stuck in a bad local optimum. To alleviate this problem, we build on the intuition that, rather than considering all samples simultaneously, the algorithm should be presented with the training data in a meaningful order that facilitates learning. The order of the samples is determined by how easy they are. The main challenge is that often we are not provided with a readily computable measure of the easiness of samples. We address this issue by proposing a novel, iterative self-paced learning algorithm where each iteration simultaneously selects easy samples and learns a new parameter vector. The number of samples selected is governed by a weight that is annealed until the entire training data has been considered. We empirically demonstrate that the self-paced learning algorithm outperforms the state of the art method for learning a latent structural SVM on four applications: object localization, noun phrase coreference, motif finding and handwritten digit recognition.

Added

2026-09-24

Predicting good probabilities with supervised learning

Predicting good probabilities with supervised learning

Alexandru Niculescu-Mizil, Rich Caruana

OrganizationsCornell University

Why you should read this

Demonstrates how common supervised learning models systematically distort probability estimates and establishes that applying Platt Scaling or Isotonic Regression transforms boosted trees, random forests, and support vector machines into the most accurate probability predictors.

We examine the relationship between the predictions made by different learning algorithms and true posterior probabilities. We show that maximum margin methods such as boosted trees and boosted stumps push probability mass away from 0 and 1 yielding a characteristic sigmoid shaped distortion in the predicted probabilities. Models such as Naive Bayes, which make unrealistic independence assumptions, push probabilities toward 0 and 1. Other models such as neural nets and bagged trees do not have these biases and predict well calibrated probabilities. We experiment with two ways of correcting the biased probabilities predicted by some learning methods: Platt Scaling and Isotonic Regression. We qualitatively examine what kinds of distortions these calibration methods are suitable for and quantitatively examine how much data they need to be effective. The empirical results show that after calibration boosted trees, random forests, and SVMs predict the best probabilities.

Added

2026-09-16

Mining the peanut gallery: opinion extraction and semantic classification of product reviews

Mining the peanut gallery: opinion extraction and semantic classification of product reviews

Kushal Dave, Steve Lawrence, David M. Pennock

OrganizationsNEC Laboratories America, Inc.Overture Services, Inc.

Why you should read this

Proposes an opinion mining system that uses information retrieval scoring techniques and variable-length text patterns to classify review sentiment and synthesize unstructured web feedback into product attribute summaries.

The web contains a wealth of product reviews, but sifting through them is a daunting task. Ideally, an opinion mining tool would process a set of search results for a given item, generating a list of product attributes (quality, features, etc.) and aggregating opinions about each of them (poor, mixed, good). We begin by identifying the unique properties of this problem and develop a method for automatically distinguishing between positive and negative reviews. Our classifier draws on information retrieval techniques for feature extraction and scoring, and the results for various metrics and heuristics vary depending on the testing situation. The best methods work as well as or better than traditional machine learning. When operating on individual sentences collected from web searches, performance is limited due to noise and ambiguity. But in the context of a complete web-based tool and aided by a simple method for grouping sentences into attributes, the results are qualitatively quite useful. discussion boards and mailing list archives, as well as in Usenet via Google Groups. Users also comment on products in their personal web sites and blogs, which are then aggregated by sites such as Blogstreet.com, AllConsuming.net, and onfocus.com. When trying to locate information on a product, a general web search turns up several useful sites, but getting an overall sense of these reviews can be daunting or time-consuming. In the movie review domain, sites like Rottentomatoes.com have sprung up to try to impose some order on the void, providing ratings and brief quotes from numerous reviews and generating an aggregate opinion. Such sites even have their own category—“Review Hubs”—on Yahoo! On the commercial side, Internet clipping services like Webclipping.com, eWatch.com, and TracerLock.com watch news sites and discussion areas for mentions of a given company or product, trying to track “buzz.” Print clipping services have been providing competitive intelligence for some time. The ease of publishing on the web led to an explosion in content to be surveyed, but the same technology makes automation much more feasible. This paper describes a tool for sifting through and synthesizing product reviews, automating the sort of work done by aggregation sites or clipping services. We begin by using structured reviews for testing and training, identifying appropriate features and scoring methods from information retrieval for determining whether reviews are positive or negative. These results perform as well as traditional machine learning methods. We then use the classifier to identify and classify review sentences from the web, where classification is more difficult. However, a simple technique for identifying the relevant attributes of a product produces a subjectively useful summary.

Added

2026-09-14

Supervised learning with quantum-enhanced feature spaces

Supervised learning with quantum-enhanced feature spaces

Vojtech Havlicek, Antonio D. Córcoles, Kristan Temme, Aram W. Harrow, Abhinav Kandala, Jerry M. Chow, Jay M. Gambetta

Why you should read this

Demonstrates experimental supervised classification on a superconducting quantum processor by mapping data into exponentially large Hilbert spaces via variational circuits and direct quantum kernel estimation.

Machine learning and quantum computing are two technologies each with the potential for altering how computation is performed to address previously untenable problems. Kernel methods for machine learning are ubiquitous for pattern recognition, with support vector machines (SVMs) being the most well-known method for classification problems. However, there are limitations to the successful solution to such problems when the feature space becomes large, and the kernel functions become computationally expensive to estimate. A core element to computational speed-ups afforded by quantum algorithms is the exploitation of an exponentially large quantum state space through controllable entanglement and interference. Here, we propose and experimentally implement two novel methods on a superconducting processor. Both methods represent the feature space of a classification problem by a quantum state, taking advantage of the large dimensionality of quantum Hilbert space to obtain an enhanced solution. One method, the quantum variational classifier builds on [1,2] and operates through using a variational quantum circuit to classify a training set in direct analogy to conventional SVMs. In the second, a quantum kernel estimator, we estimate the kernel function and optimize the classifier directly. The two methods present a new class of tools for exploring the applications of noisy intermediate scale quantum computers [3] to machine learning.

Added

2026-09-14

Text Classification from Labeled and Unlabeled Documents using EM

Text Classification from Labeled and Unlabeled Documents using EM

Kamal Nigam, Andrew Kachites Mccallum, Sebastian Thrun, Tom Mitchell

OrganizationsCarnegie Mellon UniversityJust Research

Why you should read this

Demonstrates how combining Expectation-Maximization with naive Bayes leverages abundant unlabeled text to significantly reduce classification error and labeled data requirements, while introducing practical extensions to address violated generative model assumptions.

This paper shows that the accuracy of learned text classifiers can be improved by augmenting a small number of labeled training documents with a large pool of unlabeled documents. This is important because in many text classification problems obtaining training labels is expensive, while large quantities of unlabeled documents are readily available. We introduce an algorithm for learning from labeled and unlabeled documents based on the combination of Expectation-Maximization (EM) and a naive Bayes classifier. The algorithm first trains a classifier using the available labeled documents, and probabilistically labels the unlabeled documents. It then trains a new classifier using the labels for all the documents, and iterates to convergence. This basic EM procedure works well when the data conform to the generative assumptions of the model. However these assumptions are often violated in practice, and poor performance can result. We present two extensions to the algorithm that improve classification accuracy under these conditions: (1) a weighting factor to modulate the contribution of the unlabeled data, and (2) the use of multiple mixture components per class. Experimental results, obtained using text from three different real-world tasks, show that the use of unlabeled data reduces classification error by up to 30%.

Added

2026-09-12

Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels

Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels

Zhilu Zhang, Mert R. Sabuncu

OrganizationsCornell University

Why you should read this

Proposes a generalized cross entropy loss function that unifies categorical cross entropy and mean absolute error, enabling deep neural networks to learn effectively from datasets with severe label noise without sacrificing convergence speed or accuracy.

Deep neural networks (DNNs) have achieved tremendous success in a variety of applications across many disciplines. Yet, their superior performance comes with the expensive cost of requiring correctly annotated large-scale datasets. Moreover, due to DNNs' rich capacity, errors in training labels can hamper performance. To combat this problem, mean absolute error (MAE) has recently been proposed as a noise-robust alternative to the commonly-used categorical cross entropy (CCE) loss. However, as we show in this paper, MAE can perform poorly with DNNs and challenging datasets. Here, we present a theoretically grounded set of noise-robust loss functions that can be seen as a generalization of MAE and CCE. Proposed loss functions can be readily applied with any existing DNN architecture and algorithm, while yielding good performance in a wide range of noisy label scenarios. We report results from experiments conducted with CIFAR-10, CIFAR-100 and FASHION-MNIST datasets and synthetically generated noisy labels.

Added

2026-09-11