keyword
linear classifiers
A linear classifier is a machine learning model that determines the category or class label of an input data point based on a linear combination of its features. Geometrically, it separates distinct classes by establishing a flat decision boundary, which corresponds to a line in two dimensions, a plane in three dimensions, or a hyperplane in higher-dimensional vector spaces. A prediction is made by computing the dot product of the input feature vector with a learned weight vector, adding a bias term, and applying a threshold or decision rule to the resulting score. Prominent examples of linear classifiers include logistic regression, linear support vector machines, and perceptrons. These models are widely utilized across pattern recognition, natural language processing, and computer vision due to their computational efficiency, simplicity, convex optimization properties, and strong performance when operating on high-dimensional feature representations.
10 items

Distinguishing the Knowable from the Unknowable with Language Models
Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, Benjamin L. Edelman
Why you should read this
Demonstrates that language models internally separate reducible epistemic uncertainty from inherent aleatoric entropy, enabling simple probes and unsupervised methods to accurately detect when an uncertain prediction is caused by a lack of knowledge rather than inherent ambiguity.
We study the feasibility of identifying epistemic uncertainty (reflecting a lack of knowledge), as opposed to aleatoric uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over free-form text. In the absence of ground-truth probabilities, we explore a setting where, in order to (approximately) disentangle a given LLM’s uncertainty, a significantly larger model stands in as a proxy for the ground truth. We show that small linear probes trained on the embeddings of frozen, pretrained models accurately predict when larger models will be more confident at the token level and that probes trained on one text domain generalize to others. Going further, we propose a fully unsupervised method that achieves non-trivial accuracy on the same task. Taken together, we interpret these results as evidence that LLMs naturally contain internal representations of different types of uncertainty that could potentially be leveraged to devise more informative indicators of model confidence in diverse practical settings. Code can be found at: https://github.com/KempnerInstitute/llm_uncertainty
Added
2026-10-05

PAC-learning for Strategic Classification
Ravi Sundaram, Anil Vullikanti, Haifeng Xu, Fan Yao
Why you should read this
Establishes a unified PAC-learning framework for strategic classification under heterogeneous agent preferences by introducing the strategic VC-dimension and fully characterizing both the statistical and computational limits of linear classifiers.
The study of strategic or adversarial manipulation of testing data to fool a classifier has attracted much recent attention. Most previous works have focused on two extreme situations where any testing data point either is completely adversarial or always equally prefers the positive label. In this paper, we generalize both of these through a unified framework by considering strategic agents with heterogenous preferences, and introduce the notion of strategic VC-dimension (SVC) to capture the PAC-learnability in our general strategic setup. SVC provably generalizes the recent concept of adversarial VC-dimension (AVC) introduced by Cullina et al. (2018). We instantiate our framework for the fundamental strategic linear classification problem. We fully characterize: (1) the statistical learnability of linear classifiers by pinning down its SVC; (2) its computational tractability by pinning down the complexity of the empirical risk minimization problem. Interestingly, the SVC of linear classifiers is always upper bounded by its standard VC-dimension. This characterization also strictly generalizes the AVC bound for linear classifiers in (Cullina et al., 2018). Finally, we briefly investigate the power of randomization in our strategic classification setup. We show that randomization may strictly increase the accuracy in general, but will not help in the special case of adversarial classification with zero-manipulation-cost.
Added
2026-10-03

Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets
Guy Hacohen, Avihu Dekel, Daphna Weinshall
Why you should read this
Reveals that active learning undergoes a phase transition requiring representative samples at low annotation budgets and atypical samples at high budgets, and introduces TypiClust to drastically improve low-budget model accuracy through typicality-based clustering.
Investigating active learning, we focus on the relation between the number of labeled examples (budget size), and suitable querying strategies. Our theoretical analysis shows a behavior reminiscent of phase transition: typical examples are best queried when the budget is low, while unrepresentative examples are best queried when the budget is large. Combined evidence shows that a similar phenomenon occurs in common classification models. Accordingly, we propose TypiClust – a deep active learning strategy suited for low budgets. In a comparative empirical investigation of supervised learning, using a variety of architectures and image datasets, TypiClust outperforms all other active learning strategies in the low-budget regime. Using TypiClust in the semi-supervised framework, performance gets an even more significant boost. In particular, state-of-the-art semi-supervised methods trained on CIFAR-10 with 10 labeled examples selected by TypiClust, reach 93.2% accuracy – an improvement of 39.4% over random selection. Code is available at https://github.com/avihu111/TypiClust.
Added
2026-09-30

FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hervé Jégou, Tomas Mikolov
Why you should read this
Proposes a product quantization approach to compress text classification models by two orders of magnitude with minimal accuracy loss, enabling fastText deployment on memory-constrained devices.
We consider the problem of producing compact architectures for text classification, such that the full model fits in a limited amount of memory. After considering different solutions inspired by the hashing literature, we propose a method built upon product quantization to store word embeddings. While the original technique leads to a loss in accuracy, we adapt this method to circumvent quantization artefacts. Our experiments carried out on several benchmarks show that our approach typically requires two orders of magnitude less memory than fastText while being only slightly inferior with respect to accuracy. As a result, it outperforms the state of the art by a good margin in terms of the compromise between memory usage and accuracy.
Added
2026-09-25

Meta-Learning With Differentiable Convex Optimization
Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, Stefano Soatto
Why you should read this
Proposes MetaOptNet, an end-to-end meta-learning method that incorporates differentiable convex optimization as a base learner via implicit differentiation, outperforming standard nearest-neighbor approaches across standard few-shot classification benchmarks.
Many meta-learning approaches for few-shot learning rely on simple base learners such as nearest-neighbor classifiers. However, even in the few-shot regime, discriminatively trained linear predictors can offer better generalization. We propose to use these predictors as base learners to learn representations for few-shot learning and show they offer better tradeoffs between feature size and performance across a range of few-shot recognition benchmarks. Our objective is to learn feature embeddings that generalize well under a linear classification rule for novel categories. To efficiently solve the objective, we exploit two properties of linear classifiers: implicit differentiation of the optimality conditions of the convex problem and the dual formulation of the optimization problem. This allows us to use high-dimensional embeddings with improved generalization at a modest increase in computational overhead. Our approach, named MetaOptNet, achieves state-of-the-art performance on miniImageNet, tieredImageNet, CIFAR-FS, and FC100 few-shot learning benchmarks. Our code is available at this https URL.
Added
2026-09-25

Small Sample Size Effects in Statistical Pattern Recognition: Recommendations for Practitioners
S. Raudys, Anil K. Jain
Why you should read this
Presents practical guidelines and quantitative analyses to help practitioners choose appropriate training and test sample sizes, avoid small-sample bias in classifier design and feature selection, and accurately estimate classification error rates.
During the last two decades a considerable amount of effort has been devoted to the analysis of the influence of both training and testing sample size on the design and performance of pattern recognition systems. These questions are interesting to practitioners as well as theoreticians, because the small-sample effects can easily contaminate the design and evaluation of a proposed system. For applications with a large number of features and a complex classification rule, the training sample size must be quite large. A large test sample is required to accurately evaluate a classifier with a low error rate. The design of a pattern recognition system consists of several stages: data collection, formation of the pattern classes, feature selection, specification of the classification algorithm, and estimation of the classification error. In this paper, we will discuss the effects of sample size on feature selection and error estimation for several types of classifier. In addition to surveying prior work in this area, our emphasis is on giving practical advice to today's designers and users of statistical pattern recognition systems.
Added
2026-09-25

On Discriminative vs. Generative Classifiers: A comparison of logistic regression and naive Bayes
A. Ng, Michael I. Jordan
Why you should read this
Demonstrates that while discriminative classifiers achieve lower asymptotic error, generative models require only logarithmic sample complexity to approach their asymptotic performance, explaining why naive Bayes frequently outperforms logistic regression on smaller training sets.
We compare discriminative and generative learning as typified by logistic regression and naive Bayes. We show, contrary to a widely-held belief that discriminative classifiers are almost always to be preferred, that there can often be two distinct regimes of performance as the training set size is increased, one in which each algorithm does better. This stems from the observation—which is borne out in repeated experiments—that while discriminative learning has lower asymptotic error, a generative classifier may also approach its (higher) asymptotic error much faster.
Added
2026-09-14

Improving the Fisher Kernel for Large-Scale Image Classification
Florent Perronnin, Jorge Sánchez, Thomas Mensink
Why you should read this
Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.
The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.
Added
2026-09-14

Bag of Tricks for Efficient Text Classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, Tomas Mikolov
Why you should read this
Presents fastText, a lightweight linear model that matches the accuracy of deep neural networks on text classification benchmarks while training orders of magnitude faster on standard CPU hardware.
This paper explores a simple and efficient baseline for text classification. Our experiments show that our fast text classifier fastText is often on par with deep learning classifiers in terms of accuracy, and many orders of magnitude faster for training and evaluation. We can train fastText on more than one billion words in less than ten minutes using a standard multicore~CPU, and classify half a million sentences among~312K classes in less than a minute.
Added
2026-09-10


Machine learning in automated text categorization
Fabrizio Sebastiani
Why you should read this
Synthesizes the core machine learning paradigm for automated text categorization by examining document representation, inductive classifier construction, and performance evaluation metrics.
The automated categorization (or classification) of texts into predefined categories has witnessed a booming interest in the last ten years, due to the increased availability of documents in digital form and the ensuing need to organize them. In the research community the dominant approach to this problem is based on machine learning techniques: a general inductive process automatically builds a classifier by learning, from a set of preclassified documents, the characteristics of the categories. The advantages of this approach over the knowledge engineering approach (consisting in the manual definition of a classifier by domain experts) are a very good effectiveness, considerable savings in terms of expert manpower, and straightforward portability to different domains. This survey discusses the main approaches to text categorization that fall within the machine learning paradigm. We will discuss in detail issues pertaining to three different problems, namely document representation, classifier construction, and classifier evaluation.
Added
2026-09-09
