keyword
ensemble learning
Ensemble learning is a machine learning technique in which multiple models, often referred to as base learners, are trained and combined to solve a computational problem, producing higher predictive accuracy, generalization, and robustness than any single individual model could achieve alone. Rather than relying on the output of a single estimator, ensemble methods aggregate diverse predictions through strategies such as weighted averaging, majority voting, bagging, boosting, or stacking. By integrating multiple hypotheses, these methods systematically reduce error stemming from model variance, bias, or noise in the training data, making them widely applied across supervised classification, regression, time series forecasting, and complex multi-model artificial intelligence systems.
11 items

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra
Why you should read this
Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.
The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.
Added
2026-09-30

LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Dongfu Jiang, Xiang Ren, Bill Yuchen Lin
Why you should read this
Proposes LLM-Blender, an ensembling framework that combines multiple open-source language models by using cross-attention pairwise ranking to select top candidate outputs and a generative fusion module to merge them into superior responses.
We present LLM-BLENDER, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs). Our framework consists of two modules: PAIRRANKER and GENFUSER, addressing the observation that optimal LLMs for different examples can significantly vary. PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs. It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one. Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking. Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses. To facilitate large-scale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons. Our LLM-BLENDER significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap.
Added
2026-09-28

Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting
Yuwei Fu, Di Wu, Benoit Boulet
Why you should read this
Proposes a reinforcement learning framework that dynamically assigns ensemble weights to base models over time, effectively adapting forecasts to non-stationary distributions across real-world time series benchmarks.
Time series data appears in many real-world fields such as energy, transportation, communication systems. Accurate modelling and forecasting of time series data can be of significant importance to improve the efficiency of these systems. Extensive research efforts have been taken for time series problems. Different types of approaches, including both statistical-based methods and machine learning-based methods, have been investigated. Among these methods, ensemble learning has shown to be effective and robust. However, it is still an open question that how we should determine weights for base models in the ensemble. Sub-optimal weights may prevent the final model from reaching its full potential. To deal with this challenge, we propose a reinforcement learning (RL) based model combination (RLMC) framework for determining model weights in an ensemble for time series forecasting tasks. By formulating model selection as a sequential decision-making problem, RLMC learns a deterministic policy to output dynamic model weights for non-stationary time series data. RLMC further leverages deep learning to learn hidden features from raw time series data to adapt fast to the changing data distribution. Extensive experiments on multiple real-world datasets have been implemented to showcase the effectiveness of the proposed method.
Added
2026-09-26

Ensemble Distillation for Robust Model Fusion in Federated Learning
Tao Lin, Lingjing Kong, Sebastian U. Stich, Martin Jaggi
Why you should read this
Proposes an ensemble distillation framework that aggregates heterogeneous client architectures using unlabeled data in federated learning, significantly cutting communication rounds and training time compared to standard parameter-averaging methods.
Federated Learning (FL) is a machine learning setting where many devices collaboratively train a machine learning model while keeping the training data decentralized. In most of the current training schemes the central model is refined by averaging the parameters of the server model and the updated parameters from the client side. However, directly averaging model parameters is only possible if all models have the same structure and size, which could be a restrictive constraint in many scenarios. In this work we investigate more powerful and more flexible aggregation schemes for FL. Specifically, we propose ensemble distillation for model fusion, i.e. training the central classifier through unlabeled data on the outputs of the models from the clients. This knowledge distillation technique mitigates privacy risk and cost to the same extent as the baseline FL algorithms, but allows flexible aggregation over heterogeneous client models that can differ e.g. in size, numerical precision or structure. We show in extensive empirical experiments on various CV/NLP datasets (CIFAR-10/100, ImageNet, AG News, SST2) and settings (heterogeneous models/data) that the server model can be trained much faster, requiring fewer communication rounds than any existing FL technique so far.
Added
2026-09-24

Learning under Concept Drift: A Review
Jie Lu, Anjin Liu, Fan Dong, Feng Gu, João Gama, Guangquan Zhang
Why you should read this
Establishes a unified structural framework covering concept drift detection, understanding, and adaptation while evaluating 24 benchmark datasets across more than 130 studies to guide machine learning on non-stationary data streams.
Concept drift describes unforeseeable changes in the underlying distribution of streaming data over time. Concept drift research involves the development of methodologies and techniques for drift detection, understanding and adaptation. Data analysis has revealed that machine learning in a concept drift environment will result in poor learning results if the drift is not addressed. To help researchers identify which research topics are significant and how to apply related techniques in data analysis tasks, it is necessary that a high quality, instructive review of current research developments and trends in the concept drift field is conducted. In addition, due to the rapid development of concept drift in recent years, the methodologies of learning under concept drift have become noticeably systematic, unveiling a framework which has not been mentioned in literature. This paper reviews over 130 high quality publications in concept drift related research areas, analyzes up-to-date developments in methodologies and techniques, and establishes a framework of learning under concept drift including three main components: concept drift detection, concept drift understanding, and concept drift adaptation. This paper lists and discusses 10 popular synthetic datasets and 14 publicly available benchmark datasets used for evaluating the performance of learning algorithms aiming at handling concept drift. Also, concept drift related research directions are covered and discussed. By providing state-of-the-art knowledge, this survey will directly support researchers in their understanding of research developments in the field of learning under concept drift.
Added
2026-09-18

SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary
Alberto Fernandez, Salvador Garcia, Francisco Herrera, Nitesh V. Chawla
Why you should read this
Synthesizes fifteen years of research on the Synthetic Minority Oversampling Technique (SMOTE), detailing its numerous algorithm variants, cross-paradigm applications, and key open challenges in scaling to big data and complex class distributions.
The Synthetic Minority Oversampling Technique (SMOTE) preprocessing algorithm is considered “de facto” standard in the framework of learning from imbalanced data. This is due to its simplicity in the design of the procedure, as well as its robustness when applied to different type of problems. Since its publication in 2002, SMOTE has proven successful in a variety of applications from several different domains. SMOTE has also inspired several approaches to counter the issue of class imbalance, and has also significantly contributed to new supervised learning paradigms, including multilabel classification, incremental learning, semi-supervised learning, multi-instance learning, among others. It is standard benchmark for learning from imbalanced data. It is also featured in a number of different software packages — from open source to commercial. In this paper, marking the fifteen year anniversary of SMOTE, we reflect on the SMOTE journey, discuss the current state of affairs with SMOTE, its applications, and also identify the next set of challenges to extend SMOTE for Big Data problems.
Added
2026-09-16

Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning
Guillaume Lemaitre, Fernando Nogueira, Christos K. Aridas
Why you should read this
Presents imbalanced-learn, a Python library integrated with scikit-learn that provides standardized implementations of over-sampling, under-sampling, and ensemble algorithms to effectively train machine learning models on skewed class distributions.
Imbalanced-learn is an open-source python toolbox aiming at providing a wide range of methods to cope with the problem of imbalanced dataset frequently encountered in machine learning and pattern recognition. The implemented state-of-the-art methods can be categorized into 4 groups: (i) under-sampling, (ii) over-sampling, (iii) combination of over- and under-sampling, and (iv) ensemble learning methods. The proposed toolbox only depends on numpy, scipy, and scikit-learn and is distributed under MIT license. Furthermore, it is fully compatible with scikit-learn and is part of the scikit-learn-contrib supported project. Documentation, unit tests as well as integration tests are provided to ease usage and contribution. The toolbox is publicly available in GitHub: this https URL.
Added
2026-09-14

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization
Thomas G. Dietterich
Why you should read this
Demonstrates across 33 benchmark datasets that while boosting generates the most accurate decision tree ensembles on clean data, bagging remains superior under classification noise because boosting assigns excessive weight to mislabeled training instances.
Bagging and boosting are methods that generate a diverse ensemble of classifiers by manipulating the training data given to a "base" learning algorithm. Breiman has pointed out that they rely for their effectiveness on the instability of the base learning algorithm. An alternative approach to generating an ensemble is to randomize the internal decisions made by the base algorithm. This general approach has been studied previously by Ali and Pazzani and by Dietterich and Kong. This paper compares the effectiveness of randomization, bagging, and boosting for improving the performance of the decision-tree algorithm C4.5. The experiments show that in situations with little or no classification noise, randomization is competitive with (and perhaps slightly superior to) bagging but not as accurate as boosting. In situations with substantial classification noise, bagging is much better than boosting, and sometimes better than randomization.
Added
2026-09-12

Extremely randomized trees
Pierre Geurts, Damien Ernst, Louis Wehenkel
Why you should read this
Introduces the Extra-Trees ensemble method, which randomizes both feature and cut-point selection during node splitting to reduce variance and achieve high predictive accuracy with significantly faster training speeds than standard tree ensembles.
This paper proposes a new tree-based ensemble method for supervised classification and regression problems. It essentially consists of randomizing strongly both attribute and cut-point choice while splitting a tree node. In the extreme case, it builds totally randomized trees whose structures are independent of the output values of the learning sample. The strength of the randomization can be tuned to problem specifics by the appropriate choice of a parameter. We evaluate the robustness of the default choice of this parameter, and we also provide insight on how to adjust it in particular situations. Besides accuracy, the main strength of the resulting algorithm is computational efficiency. A bias/variance analysis of the Extra-Trees algorithm is also provided as well as a geometrical and a kernel characterization of the models induced.
Added
2026-09-07

Machine Learning Engineering
Andriy Burkov
Why you should read this
Covers the complete lifecycle of production ML systems, including best practices for monitoring, maintenance, fallback strategies, handling adversaries, and all the practical challenges that arise when deploying machine learning at scale.
Source
https://www.mlebook.com/Added
2025-09-09
License
Read first, buy later

The Hundred-Page Machine Learning Book
Andriy Burkov
Why you should read this
Distills the essential concepts of machine learning—from fundamental algorithms through neural networks and ensemble methods—into a concise, mathematically grounded introduction that balances theory with practical understanding in just 100 pages.
Added
2025-06-03
License
Read first, buy later
