Built independently by an author, for readers. Read the story and support ChapterPal

keyword

linear classifiers

A linear classifier is a machine learning model that determines the category or class label of an input data point based on a linear combination of its features. Geometrically, it separates distinct classes by establishing a flat decision boundary, which corresponds to a line in two dimensions, a plane in three dimensions, or a hyperplane in higher-dimensional vector spaces. A prediction is made by computing the dot product of the input feature vector with a learned weight vector, adding a bias term, and applying a threshold or decision rule to the resulting score. Prominent examples of linear classifiers include logistic regression, linear support vector machines, and perceptrons. These models are widely utilized across pattern recognition, natural language processing, and computer vision due to their computational efficiency, simplicity, convex optimization properties, and strong performance when operating on high-dimensional feature representations.

10 items

Distinguishing the Knowable from the Unknowable with Language Models

Distinguishing the Knowable from the Unknowable with Language Models

Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, Benjamin L. Edelman

OrganizationsHarvard University

Why you should read this

Demonstrates that language models internally separate reducible epistemic uncertainty from inherent aleatoric entropy, enabling simple probes and unsupervised methods to accurately detect when an uncertain prediction is caused by a lack of knowledge rather than inherent ambiguity.

We study the feasibility of identifying epistemic uncertainty (reflecting a lack of knowledge), as opposed to aleatoric uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over free-form text. In the absence of ground-truth probabilities, we explore a setting where, in order to (approximately) disentangle a given LLM’s uncertainty, a significantly larger model stands in as a proxy for the ground truth. We show that small linear probes trained on the embeddings of frozen, pretrained models accurately predict when larger models will be more confident at the token level and that probes trained on one text domain generalize to others. Going further, we propose a fully unsupervised method that achieves non-trivial accuracy on the same task. Taken together, we interpret these results as evidence that LLMs naturally contain internal representations of different types of uncertainty that could potentially be leveraged to devise more informative indicators of model confidence in diverse practical settings. Code can be found at: https://github.com/KempnerInstitute/llm_uncertainty

Added

2026-10-05

PAC-learning for Strategic Classification

PAC-learning for Strategic Classification

Ravi Sundaram, Anil Vullikanti, Haifeng Xu, Fan Yao

OrganizationsBiocomplexity InstituteDepartment of Computer ScienceNortheastern UniversityUniversity of ChicagoUniversity of Virginia

Why you should read this

Establishes a unified PAC-learning framework for strategic classification under heterogeneous agent preferences by introducing the strategic VC-dimension and fully characterizing both the statistical and computational limits of linear classifiers.

The study of strategic or adversarial manipulation of testing data to fool a classifier has attracted much recent attention. Most previous works have focused on two extreme situations where any testing data point either is completely adversarial or always equally prefers the positive label. In this paper, we generalize both of these through a unified framework by considering strategic agents with heterogenous preferences, and introduce the notion of strategic VC-dimension (SVC) to capture the PAC-learnability in our general strategic setup. SVC provably generalizes the recent concept of adversarial VC-dimension (AVC) introduced by Cullina et al. (2018). We instantiate our framework for the fundamental strategic linear classification problem. We fully characterize: (1) the statistical learnability of linear classifiers by pinning down its SVC; (2) its computational tractability by pinning down the complexity of the empirical risk minimization problem. Interestingly, the SVC of linear classifiers is always upper bounded by its standard VC-dimension. This characterization also strictly generalizes the AVC bound for linear classifiers in (Cullina et al., 2018). Finally, we briefly investigate the power of randomization in our strategic classification setup. We show that randomization may strictly increase the accuracy in general, but will not help in the special case of adversarial classification with zero-manipulation-cost.

Added

2026-10-03

Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets

Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets

Guy Hacohen, Avihu Dekel, Daphna Weinshall

OrganizationsThe Hebrew University of Jerusalem

Why you should read this

Reveals that active learning undergoes a phase transition requiring representative samples at low annotation budgets and atypical samples at high budgets, and introduces TypiClust to drastically improve low-budget model accuracy through typicality-based clustering.

Investigating active learning, we focus on the relation between the number of labeled examples (budget size), and suitable querying strategies. Our theoretical analysis shows a behavior reminiscent of phase transition: typical examples are best queried when the budget is low, while unrepresentative examples are best queried when the budget is large. Combined evidence shows that a similar phenomenon occurs in common classification models. Accordingly, we propose TypiClust – a deep active learning strategy suited for low budgets. In a comparative empirical investigation of supervised learning, using a variety of architectures and image datasets, TypiClust outperforms all other active learning strategies in the low-budget regime. Using TypiClust in the semi-supervised framework, performance gets an even more significant boost. In particular, state-of-the-art semi-supervised methods trained on CIFAR-10 with 10 labeled examples selected by TypiClust, reach 93.2% accuracy – an improvement of 39.4% over random selection. Code is available at https://github.com/avihu111/TypiClust.

Added

2026-09-30

Meta-Learning With Differentiable Convex Optimization

Meta-Learning With Differentiable Convex Optimization

Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, Stefano Soatto

OrganizationsAmazon Web ServicesUniversity of California, Los AngelesUniversity of California, San DiegoUniversity of Massachusetts Amherst

Why you should read this

Proposes MetaOptNet, an end-to-end meta-learning method that incorporates differentiable convex optimization as a base learner via implicit differentiation, outperforming standard nearest-neighbor approaches across standard few-shot classification benchmarks.

Many meta-learning approaches for few-shot learning rely on simple base learners such as nearest-neighbor classifiers. However, even in the few-shot regime, discriminatively trained linear predictors can offer better generalization. We propose to use these predictors as base learners to learn representations for few-shot learning and show they offer better tradeoffs between feature size and performance across a range of few-shot recognition benchmarks. Our objective is to learn feature embeddings that generalize well under a linear classification rule for novel categories. To efficiently solve the objective, we exploit two properties of linear classifiers: implicit differentiation of the optimality conditions of the convex problem and the dual formulation of the optimization problem. This allows us to use high-dimensional embeddings with improved generalization at a modest increase in computational overhead. Our approach, named MetaOptNet, achieves state-of-the-art performance on miniImageNet, tieredImageNet, CIFAR-FS, and FC100 few-shot learning benchmarks. Our code is available at this https URL.

Added

2026-09-25

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners

S. Raudys, Anil K. Jain

OrganizationsInstitute of Mathematics and Cybernetics, Lithuanian Academy of SciencesMichigan State University

Why you should read this

Presents practical guidelines and quantitative analyses to help practitioners choose appropriate training and test sample sizes, avoid small-sample bias in classifier design and feature selection, and accurately estimate classification error rates.

During the last two decades a considerable amount of effort has been devoted to the analysis of the influence of both training and testing sample size on the design and performance of pattern recognition systems. These questions are interesting to practitioners as well as theoreticians, because the small-sample effects can easily contaminate the design and evaluation of a proposed system. For applications with a large number of features and a complex classification rule, the training sample size must be quite large. A large test sample is required to accurately evaluate a classifier with a low error rate. The design of a pattern recognition system consists of several stages: data collection, formation of the pattern classes, feature selection, specification of the classification algorithm, and estimation of the classification error. In this paper, we will discuss the effects of sample size on feature selection and error estimation for several types of classifier. In addition to surveying prior work in this area, our emphasis is on giving practical advice to today's designers and users of statistical pattern recognition systems.

Added

2026-09-25

Improving the Fisher Kernel for Large-Scale Image Classification

Improving the Fisher Kernel for Large-Scale Image Classification

Florent Perronnin, Jorge Sánchez, Thomas Mensink

OrganizationsXerox

Why you should read this

Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.

The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.

Added

2026-09-14