keyword
outlier detection
Outlier detection, also known as anomaly detection, is the process of identifying data points, events, or observations in a dataset that deviate significantly from the expected distribution or norm. In data mining, statistics, and machine learning, this task focuses on discovering rare instances or unexpected patterns that may signify operational faults, security breaches, data corruption, or genuine novelties. Techniques for outlier detection encompass traditional statistical tests, distance- and density-based metrics that evaluate the isolation of points relative to their neighbors, and deep learning architectures such as autoencoders and one-class classifiers. These methods can operate under supervised, semi-supervised, or unsupervised paradigms, supporting critical applications such as financial fraud prevention, industrial defect inspection, medical diagnostics, and network intrusion monitoring.
12 items

Causal structure-based root cause analysis of outliers
Kailash Budhathoki, Lenon Minorics, Patrick Blöbaum, Dominik Janzing
Why you should read this
Develops a principled framework that combines functional causal models, information-theoretic score calibration, and Shapley values to pinpoint and quantify the root causes behind detected outliers.
Current techniques for explaining outliers cannot tell what caused the outliers. We present a formal method to identify “root causes” of outliers, amongst variables. The method requires a causal graph of the variables along with the functional causal model. It quantifies the contribution of each variable to the target outlier score, which explains to what extent each variable is a “root cause” of the target outlier. We study the empirical performance of the method through simulations and present a real-world case study identifying “root causes” of extreme river flows.
Added
2026-09-26

Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection
Choubo Ding, Guansong Pang, Chunhua Shen
Why you should read this
Proposes a multi-head framework that disentangles known, pseudo, and latent residual abnormalities to effectively detect both seen and novel anomaly classes using only a small set of labeled anomaly examples.
Despite most existing anomaly detection studies assume the availability of normal training samples only, a few labeled anomaly examples are often available in many real-world applications, such as defect samples identified during random quality inspection, lesion images confirmed by radiologists in daily medical screening, etc. These anomaly examples provide valuable knowledge about the application-specific abnormality, enabling significantly improved detection of similar anomalies in some recent models. However, those anomalies seen during training often do not illustrate every possible class of anomaly, rendering these models ineffective in generalizing to unseen anomaly classes. This paper tackles open-set supervised anomaly detection, in which we learn detection models using the anomaly examples with the objective to detect both seen anomalies ('gray swans') and unseen anomalies ('black swans'). We propose a novel approach that learns disentangled representations of abnormalities illustrated by seen anomalies, pseudo anomalies, and latent residual anomalies (i.e., samples that have unusual residuals compared to the normal data in a latent space), with the last two abnormalities designed to detect unseen anomalies. Extensive experiments on nine real-world anomaly detection datasets show superior performance of our model in detecting seen and unseen anomalies under diverse settings. Code and data are available at: https://github.com/choubo/DRA
Added
2026-09-26

Anomaly Detection with Robust Deep Autoencoders
Chong Zhou, R. Paffenroth
Why you should read this
Proposes a matrix-splitting framework combining deep autoencoders with sparse and group-sparse regularization to isolate complex noise and detect anomalies without requiring clean training data.
Deep autoencoders, and other deep neural networks, have demonstrated their effectiveness in discovering non-linear features across many problem domains. However, in many real-world problems, large outliers and pervasive noise are commonplace, and one may not have access to clean training data as required by standard deep denoising autoencoders. Herein, we demonstrate novel extensions to deep autoencoders which not only maintain a deep autoencoders’ ability to discover high quality, non-linear features but can also eliminate outliers and noise without access to any clean training data. Our model is inspired by Robust Principal Component Analysis, and we split the input data X into two parts, X = L_D + S, where L_D can be effectively reconstructed by a deep autoencoder and S contains the outliers and noise in the original data X. Since such splitting increases the robustness of standard deep autoencoders, we name our model a “Robust Deep Autoencoder (RDA)”. Further, we present generalizations of our results to grouped sparsity norms which allow one to distinguish random anomalies from other types of structured corruptions, such as a collection of features being corrupted across many instances or a collection of instances having more corruptions than their fellows. Such “Group Robust Deep Autoencoders (GRDA)” give rise to novel anomaly detection approaches whose superior performance we demonstrate on a selection of benchmark problems.
Added
2026-09-24

Deep Learning for Anomaly Detection: A Survey
Raghavendra Chalapathy, Sanjay Chawla
Why you should read this
Classifies deep learning anomaly detection methods across diverse application domains, evaluating their underlying assumptions, computational complexities, and practical limitations to guide model selection and identify critical research challenges.
Anomaly detection is an important problem that has been well-studied within diverse research areas and application domains. The aim of this survey is two-fold, firstly we present a structured and comprehensive overview of research methods in deep learning-based anomaly detection. Furthermore, we review the adoption of these methods for anomaly across various application domains and assess their effectiveness. We have grouped state-of-the-art research techniques into different categories based on the underlying assumptions and approach adopted. Within each category we outline the basic anomaly detection technique, along with its variants and present key assumptions, to differentiate between normal and anomalous behavior. For each category, we present we also present the advantages and limitations and discuss the computational complexity of the techniques in real application domains. Finally, we outline open issues in research and challenges faced while adopting these techniques.
Added
2026-09-18

Algorithms for Mining Distance-Based Outliers in Large Datasets
Edwin M. Knorr, Raymond T. Ng
Why you should read this
Proposes the concept of distance-based outliers alongside scalable nested-loop and cell-based mining algorithms that efficiently detect multidimensional anomalies in large disk-resident datasets without requiring prior knowledge of underlying data distributions.
This paper deals with finding outliers (exceptions) in large, multidimensional datasets. The identification of outliers can lead to the discovery of truly unexpected knowledge in areas such as electronic commerce, credit card fraud, and even the analysis of performance statistics of professional athletes. Existing methods that we have seen for finding outliers in large datasets can only deal efficiently with two dimensions/attributes of a dataset. Here, we study the notion of DB- (Distance-Based) outliers. While we provide formal and empirical evidence showing the usefulness of DB-outliers, we focus on the development of algorithms for computing such outliers. First, we present two simple algorithms, both having a complexity of O(k N²), k being the dimensionality and N being the number of objects in the dataset. These algorithms readily support datasets with many more than two attributes. Second, we present an optimized cell-based algorithm that has a complexity that is linear wrt N, but exponential wrt k. Third, for datasets that are mainly disk-resident, we present another version of the cell-based algorithm that guarantees at most 3 passes over a dataset. We provide experimental results showing that these cell-based algorithms are by far the best for k ≤ 4.
Added
2026-09-18

Quantile Regression Forests
Nicolai Meinshausen
Why you should read this
Introduces quantile regression forests, a consistent non-parametric method extending random forests to estimate full conditional distributions and construct accurate prediction intervals for high-dimensional data.
Random forests were introduced as a machine learning tool in Breiman (2001) and have since proven to be very popular and powerful for high-dimensional regression and classification. For regression, random forests give an accurate approximation of the conditional mean of a response variable. It is shown here that random forests provide information about the full conditional distribution of the response variable, not only about the conditional mean. Conditional quantiles can be inferred with quantile regression forests, a generalisation of random forests. Quantile regression forests give a non-parametric and accurate way of estimating conditional quantiles for high-dimensional predictor variables. The algorithm is shown to be consistent. Numerical examples suggest that the algorithm is competitive in terms of predictive power.
Added
2026-09-17

Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection
Bo Zong, Qi Song, Martin Renqiang Min, Wei Cheng, Cristian Lumezanu, Daeki Cho, Haifeng Chen
Why you should read this
Proposes an end-to-end unsupervised anomaly detection framework that jointly optimizes deep autoencoding reconstruction and Gaussian mixture density estimation, eliminating decoupled two-stage training to significantly improve detection accuracy on high-dimensional data.
Unsupervised anomaly detection on multi- or high-dimensional data is of great importance in both fundamental machine learning research and industrial applications, for which density estimation lies at the core. Although previous approaches based on dimensionality reduction followed by density estimation have made fruitful progress, they mainly suffer from decoupled model learning with inconsistent optimization goals and incapability of preserving essential information in the low-dimensional space. In this paper, we present a Deep Autoencoding Gaussian Mixture Model (DAGMM) for unsupervised anomaly detection. Our model utilizes a deep autoencoder to generate a low-dimensional representation and reconstruction error for each input data point, which is further fed into a Gaussian Mixture Model (GMM). Instead of using decoupled two-stage training and the standard Expectation-Maximization (EM) algorithm, DAGMM jointly optimizes the parameters of the deep autoencoder and the mixture model simultaneously in an end-to-end fashion, leveraging a separate estimation network to facilitate the parameter learning of the mixture model. The joint optimization, which well balances autoencoding reconstruction, density estimation of latent representation, and regularization, helps the autoencoder escape from less attractive local optima and further reduce reconstruction errors, avoiding the need of pre-training. Experimental results on several public benchmark datasets show that, DAGMM significantly outperforms state-of-the-art anomaly detection techniques, and achieves up to 14% improvement based on the standard F1 score.
Added
2026-09-16

MVTec AD — A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, C. Steger
Why you should read this
Introduces the first comprehensive real-world industrial inspection dataset with pixel-accurate defect annotations across fifteen object and texture categories, establishing a rigorous benchmark that reveals key limitations in current unsupervised anomaly detection and localization methods.
The detection of anomalous structures in natural image data is of utmost importance for numerous tasks in the field of computer vision. The development of methods for unsupervised anomaly detection requires data on which to train and evaluate new approaches and ideas. We introduce the MVTec Anomaly Detection (MVTec AD) dataset containing 5354 high-resolution color images of different object and texture categories. It contains normal, i.e., defect-free, images intended for training and images with anomalies intended for testing. The anomalies manifest themselves in the form of over 70 different types of defects such as scratches, dents, contaminations, and various structural changes. In addition, we provide pixel-precise ground truth regions for all anomalies. We also conduct a thorough evaluation of current state-of-the-art unsupervised anomaly detection methods based on deep architectures such as convolutional autoencoders, generative adversarial networks, and feature descriptors using pre-trained convolutional neural networks, as well as classical computer vision methods. This initial benchmark indicates that there is considerable room for improvement. To the best of our knowledge, this is the first comprehensive, multi-object, multi-defect dataset for anomaly detection that provides pixel-accurate ground truth regions and focuses on real-world applications.
Added
2026-09-15

Efficient algorithms for mining outliers from large data sets
Sridhar Ramaswamy, Rajeev Rastogi, Kyuseok Shim
Why you should read this
Proposes a k-nearest-neighbor distance formulation for ranking outliers alongside an efficient partition-based pruning algorithm that scales to large, high-dimensional datasets and outperforms traditional nested-loop and index joins by orders of magnitude.
In this paper, we propose a novel formulation for distance-based outliers that is based on the distance of a point from its k^th nearest neighbor. We rank each point on the basis of its distance to its k^th nearest neighbor and declare the top n points in this ranking to be outliers. In addition to developing relatively straightforward solutions to finding such outliers based on the classical nested-loop join and index join algorithms, we develop a highly efficient partition-based algorithm for mining outliers. This algorithm first partitions the input data set into disjoint subsets, and then prunes entire partitions as soon as it is determined that they cannot contain outliers. This results in substantial savings in computation. We present the results of an extensive experimental study on real-life and synthetic data sets. The results from a real-life NBA database highlight and reveal several expected and unexpected aspects of the database. The results from a study on synthetic data sets demonstrate that the partition-based algorithm scales well with respect to both data set size and data set dimensionality. above description of outliers, it may seem that outliers are a nuisance—impeding the inference process—and must be quickly identified and eliminated so that they do not interfere with the data analysis. However, this viewpoint is often too narrow since outliers contain useful information. Mining for outliers has a number of useful applications in telecom and credit card fraud, loan approval, pharmaceutical research, weather prediction, financial applications, marketing and customer segmentation. For instance, consider the problem of detecting credit card fraud. A major problem that credit card companies face is the illegal use of lost or stolen credit cards. Detecting and preventing such use is critical since credit card companies assume liability for unauthorized expenses on lost or stolen cards. Since the usage pattern for a stolen card is unlikely to be similar to its usage prior to being stolen, the new usage points are probably outliers (in an intuitive sense) with respect to the old usage pattern. Detecting these outliers is clearly an important task. The problem of detecting outliers has been extensively studied in the statistics community (see [BL94] for a good survey of statistical techniques). Typically, the user has to model the data points using a statistical distribution, and points are determined to be outliers depending on how they appear in relation to the postulated model. The main problem with these approaches is that in a number of situations, the user might simply not have enough knowledge about the underlying data distribution. In order to overcome this problem, Knorr and Ng [KN98] propose the following distance-based definition for outliers that is both simple and intuitive: A point p in a data set is an outlier with respect to parameters k and d if no more than k points in the data set are at a distance of d or less from p^1. The distance function can be any metric distance function^2. The main benefit of the approach in [KN98] is that it does not require any apriori knowledge of data distributions that the statistical methods do. Additionally, the definition of outliers considered is general enough to model statistical
Added
2026-09-14

Deep One-Class Classification
Lukas Ruff, Nico Görnitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Robert A. Vandermeulen, Alexander Binder, Emmanuel Müller, Marius Kloft
Why you should read this
Introduces Deep Support Vector Data Description (Deep SVDD), a direct one-class classification objective that trains neural networks to enclose normal representations within a minimal-volume hypersphere while establishing architectural principles to prevent hypersphere collapse.
Despite the great advances made by deep learning in many machine learning problems, there is a relative dearth of deep learning approaches for anomaly detection. Those approaches which do exist involve networks trained to perform a task other than anomaly detection, namely generative models or compression, which are in turn adapted for use in anomaly detection; they are not trained on an anomaly detection based objective. In this paper we introduce a new anomaly detection method—Deep Support Vector Data Description—, which is trained on an anomaly detection based objective. The adaptation to the deep regime necessitates that our neural network and training procedure satisfy certain properties, which we demonstrate theoretically. We show the effectiveness of our method on MNIST and CIFAR-10 image benchmark datasets as well as on the detection of adversarial examples of GTSRB stop signs.
Added
2026-09-14

LOF: identifying density-based local outliers
Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, Jörg Sander
Why you should read this
Introduces the Local Outlier Factor (LOF) algorithm, a foundational density-based approach that assigns each data point a continuous outlier score relative to its surrounding neighborhood to detect anomalies across regions of varying density.
For many KDD applications, such as detecting criminal activities in E-commerce, finding the rare instances or the outliers, can be more interesting than finding the common patterns. Existing work in outlier detection regards being an outlier as a binary property. In this paper, we contend that for many scenarios, it is more meaningful to assign to each object a degree of being an outlier. This degree is called the local outlier factor (LOF) of an object. It is local in that the degree depends on how isolated the object is with respect to the surrounding neighborhood. We give a detailed formal analysis showing that LOF enjoys many desirable properties. Using real-world datasets, we demonstrate that LOF can be used to find outliers which appear to be meaningful, but can otherwise not be identified with existing approaches. Finally, a careful performance evaluation of our algorithm confirms we show that our approach of finding local outliers can be practical.
Added
2026-09-07

Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift
Stephan Rabanser, Stephan Günnemann, Zachary C. Lipton
Why you should read this
Demonstrates that a two-sample-testing-based approach, leveraging pre-trained classifiers for dimensionality reduction, effectively detects and characterizes dataset shift, enabling ML systems to "fail loudly" rather than silently.
We might hope that when faced with unexpected inputs, well-designed software systems would fire off warnings. Machine learning (ML) systems, however, which depend strongly on properties of their inputs (e.g. the i.i.d. assumption), tend to fail silently. This paper explores the problem of building ML systems that fail loudly, investigating methods for detecting dataset shift, identifying exemplars that most typify the shift, and quantifying shift malignancy. We focus on several datasets and various perturbations to both covariates and label distributions with varying magnitudes and fractions of data affected. Interestingly, we show that across the dataset shifts that we explore, a two-sample-testing-based approach, using pre-trained classifiers for dimensionality reduction, performs best. Moreover, we demonstrate that domain-discriminating approaches tend to be helpful for characterizing shifts qualitatively and determining if they are harmful.
Added
2026-04-27
