Built independently by an author, for readers. Read the story and support ChapterPal

keyword

local outlier factor

The local outlier factor is an unsupervised, density-based anomaly detection algorithm that measures the degree to which a data point is an outlier relative to its surrounding neighborhood. It operates by computing the local reachability density of a given point and comparing it to the average local densities of its k-nearest neighbors. Data points situated in regions with densities comparable to their neighbors receive a score near one, signifying normal inlier behavior, while points located in substantially sparser regions compared to their neighbors yield significantly higher scores, identifying them as local outliers. By evaluating density locally rather than relying on global distance thresholds, this technique effectively identifies anomalies within complex datasets that feature clusters of varying shapes and densities.

3 items

Unsupervised Anomaly Detection Algorithms on Real-world Data: How Many Do We Need?

Unsupervised Anomaly Detection Algorithms on Real-world Data: How Many Do We Need?

Roel Bouman, Zaharah Bukhsh, Tom Heskes

OrganizationsEindhoven University of TechnologyInformation Systems, Industrial Engineering and Innovation SciencesInstitute for Computing and Information SciencesRadboud University

Why you should read this

Demonstrates through an evaluation of 33 algorithms across 52 real-world tabular datasets that practitioners only need Extended Isolation Forest and k-nearest neighbors to effectively detect both global and local anomalies.

In this study we evaluate 33 unsupervised anomaly detection algorithms on 52 real-world multivariate tabular data sets, performing the largest comparison of unsupervised anomaly detection algorithms to date. On this collection of data sets, the EIF (Extended Isolation Forest) algorithm significantly outperforms the most other algorithms. Visualizing and then clustering the relative performance of the considered algorithms on all data sets, we identify two clear clusters: one with “local” data sets, and another with “global” data sets. “Local” anomalies occupy a region with low density when compared to nearby samples, while “global” occupy an overall low density region in the feature space. On the local data sets the kNN (k-nearest neighbor) algorithm comes out on top. On the global data sets, the EIF (extended isolation forest) algorithm performs the best. Also taking into consideration the algorithms’ computational complexity, a toolbox with these two unsupervised anomaly detection algorithms suffices for finding anomalies in this representative collection of multivariate data sets. By providing access to code and data sets, our study can be easily reproduced and extended with more algorithms and/or data sets.

Added

2026-10-05

SoftPatch: Unsupervised Anomaly Detection with Noisy Data

SoftPatch: Unsupervised Anomaly Detection with Noisy Data

Xi Jiang, Jianlin Liu, Jinbao Wang, Qiang Nie, Kai Wu, Yong Liu, Chengjie Wang, Feng Zheng

OrganizationsDepartment of Computer Science and EngineeringSouthern University of Science and TechnologyTencent

Why you should read this

Proposes SoftPatch, a patch-level denoising and memory re-weighting method that prevents defective training samples from distorting decision boundaries in real-world unsupervised anomaly detection.

Although mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection but is seldom discussed. This paper considers label-level noise in image sensory anomaly detection for the first time. To solve this problem, we proposed a memory-based unsupervised AD method, SoftPatch, which efficiently denoises the data at the patch level. Noise discriminators are utilized to generate outlier scores for patch-level noise elimination before coreset construction. The scores are then stored in the memory bank to soften the anomaly detection boundary. Compared with existing methods, SoftPatch maintains a strong modeling ability of normal data and alleviates the overconfidence problem in coreset. Comprehensive experiments in various noise scenes demonstrate that SoftPatch outperforms the state-of-the-art AD methods on the MVTecAD and BTAD benchmarks and is comparable to those methods under the setting without noise.

Added

2026-09-26

LOF: identifying density-based local outliers

LOF: identifying density-based local outliers

Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, Jörg Sander

OrganizationsLudwig Maximilian University of MunichUniversity of British Columbia

Why you should read this

Introduces the Local Outlier Factor (LOF) algorithm, a foundational density-based approach that assigns each data point a continuous outlier score relative to its surrounding neighborhood to detect anomalies across regions of varying density.

For many KDD applications, such as detecting criminal activities in E-commerce, finding the rare instances or the outliers, can be more interesting than finding the common patterns. Existing work in outlier detection regards being an outlier as a binary property. In this paper, we contend that for many scenarios, it is more meaningful to assign to each object a degree of being an outlier. This degree is called the local outlier factor (LOF) of an object. It is local in that the degree depends on how isolated the object is with respect to the surrounding neighborhood. We give a detailed formal analysis showing that LOF enjoys many desirable properties. Using real-world datasets, we demonstrate that LOF can be used to find outliers which appear to be meaningful, but can otherwise not be identified with existing approaches. Finally, a careful performance evaluation of our algorithm confirms we show that our approach of finding local outliers can be practical.

Added

2026-09-07