keyword
instance-based learning
Instance-based learning is a family of machine learning algorithms that compares new problem instances directly with instances seen during training, which are stored in memory, rather than constructing an explicit, generalized model. Often referred to as lazy learning or memory-based learning, this approach delays computation until a prediction or classification request is made on unseen data. When a query instance is evaluated, the algorithm measures its similarity to the stored examples using a defined distance metric, such as Euclidean distance, and determines the output based on the closest matching instances. Because these methods rely directly on stored historical data, they can naturally adapt to new incoming data points without full model retraining, although their prediction phase can be computationally intensive and sensitive to irrelevant or noisy features.
4 items

Machine Learning for the Detection of Oil Spills in Satellite Radar Images
MIROSLAV KUBAT, ROBERT C. HOLTE, STAN MATWIN
Why you should read this
Presents practical machine learning solutions for detecting oil spills in satellite radar imagery while addressing real-world challenges such as extreme class imbalance, data preparation, and appropriate evaluation metrics.
During a project examining the use of machine learning techniques for oil spill detection, we encountered several essential questions that we believe deserve the attention of the research community. We use our particular case study to illustrate such issues as problem formulation, selection of evaluation measures, and data preparation. We relate these issues to properties of the oil spill application, such as its imbalanced class distribution, that are shown to be common to many applications. Our solutions to these issues are implemented in the Canadian Environmental Hazards Detection System (CEHDS), which is about to undergo field testing.
Added
2026-09-25

Learning in the Presence of Concept Drift and Hidden Contexts
G. Widmer, M. Kubát
Why you should read this
Presents the FLORA framework of incremental learning algorithms that dynamically adjust sample windows and reuse past concept descriptions to handle recurring hidden contexts and concept drift in continuous data streams.
On-line learning in domains where the target concept depends on some hidden context poses serious problems. A changing context can induce changes in the target concepts, producing what is known as concept drift. We describe a family of learning algorithms that flexibly react to concept drift and can take advantage of situations where contexts reappear. The general approach underlying all these algorithms consists of (1) keeping only a window of currently trusted examples and hypotheses; (2) storing concept descriptions and re-using them when a previous context re-appears; and (3) controlling both of these functions by a heuristic that constantly monitors the system's behavior. The paper reports on experiments that test the systems' performance under various conditions such as different levels of noise and different extent and rate of concept drift.
Added
2026-09-24

The Feature Selection Problem: Traditional Methods and a New Algorithm
Kenji Kira, Larry Rendell
Why you should read this
Introduces the Relief algorithm, a noise-tolerant feature selection method that identifies relevant attributes in linear time without relying on heuristics, even when strong feature interactions are present.
For real-world concept learning problems, feature selection is important to speed up learning and to improve concept quality. We review and analyze past approaches to feature selection and note their strengths and weaknesses. We then introduce and theoretically examine a new algorithm Relief which selects relevant features using a statistical method. Relief does not depend on heuristics, is accurate even if features interact, and is noise-tolerant. It requires only linear time in the number of given features and the number of training instances, regardless of the target concept complexity. The algorithm also has certain limitations such as non-optimal feature set size. Ways to overcome the limitations are suggested. We also report the test results of comparison between Relief and other feature selection algorithms. The empirical results support the theoretical analysis, suggesting a practical approach to feature selection for real-world problems.
Added
2026-09-16

Machine Learning Engineering
Andriy Burkov
Why you should read this
Covers the complete lifecycle of production ML systems, including best practices for monitoring, maintenance, fallback strategies, handling adversaries, and all the practical challenges that arise when deploying machine learning at scale.
Source
https://www.mlebook.com/Added
2025-09-09
License
Read first, buy later
