Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ensemble learning

Ensemble learning is a machine learning technique in which multiple models, often referred to as base learners, are trained and combined to solve a computational problem, producing higher predictive accuracy, generalization, and robustness than any single individual model could achieve alone. Rather than relying on the output of a single estimator, ensemble methods aggregate diverse predictions through strategies such as weighted averaging, majority voting, bagging, boosting, or stacking. By integrating multiple hypotheses, these methods systematically reduce error stemming from model variance, bias, or noise in the training data, making them widely applied across supervised classification, regression, time series forecasting, and complex multi-model artificial intelligence systems.

11 items

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra

OrganizationsAmazonMicrosoftThe Alan Turing InstituteUniversity College London

Why you should read this

Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.

The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.

Added

2026-09-30

LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Dongfu Jiang, Xiang Ren, Bill Yuchen Lin

OrganizationsAllen Institute for AIUniversity of Southern CaliforniaZhejiang University

Why you should read this

Proposes LLM-Blender, an ensembling framework that combines multiple open-source language models by using cross-attention pairwise ranking to select top candidate outputs and a generative fusion module to merge them into superior responses.

We present LLM-BLENDER, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs). Our framework consists of two modules: PAIRRANKER and GENFUSER, addressing the observation that optimal LLMs for different examples can significantly vary. PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs. It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one. Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking. Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses. To facilitate large-scale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons. Our LLM-BLENDER significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap.

Added

2026-09-28

Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting

Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting

Yuwei Fu, Di Wu, Benoit Boulet

OrganizationsMcGill University

Why you should read this

Proposes a reinforcement learning framework that dynamically assigns ensemble weights to base models over time, effectively adapting forecasts to non-stationary distributions across real-world time series benchmarks.

Time series data appears in many real-world fields such as energy, transportation, communication systems. Accurate modelling and forecasting of time series data can be of significant importance to improve the efficiency of these systems. Extensive research efforts have been taken for time series problems. Different types of approaches, including both statistical-based methods and machine learning-based methods, have been investigated. Among these methods, ensemble learning has shown to be effective and robust. However, it is still an open question that how we should determine weights for base models in the ensemble. Sub-optimal weights may prevent the final model from reaching its full potential. To deal with this challenge, we propose a reinforcement learning (RL) based model combination (RLMC) framework for determining model weights in an ensemble for time series forecasting tasks. By formulating model selection as a sequential decision-making problem, RLMC learns a deterministic policy to output dynamic model weights for non-stationary time series data. RLMC further leverages deep learning to learn hidden features from raw time series data to adapt fast to the changing data distribution. Extensive experiments on multiple real-world datasets have been implemented to showcase the effectiveness of the proposed method.

Added

2026-09-26

Ensemble Distillation for Robust Model Fusion in Federated Learning

Ensemble Distillation for Robust Model Fusion in Federated Learning

Tao Lin, Lingjing Kong, Sebastian U. Stich, Martin Jaggi

OrganizationsÉcole Polytechnique Fédérale de LausanneMLO

Why you should read this

Proposes an ensemble distillation framework that aggregates heterogeneous client architectures using unlabeled data in federated learning, significantly cutting communication rounds and training time compared to standard parameter-averaging methods.

Federated Learning (FL) is a machine learning setting where many devices collaboratively train a machine learning model while keeping the training data decentralized. In most of the current training schemes the central model is refined by averaging the parameters of the server model and the updated parameters from the client side. However, directly averaging model parameters is only possible if all models have the same structure and size, which could be a restrictive constraint in many scenarios. In this work we investigate more powerful and more flexible aggregation schemes for FL. Specifically, we propose ensemble distillation for model fusion, i.e. training the central classifier through unlabeled data on the outputs of the models from the clients. This knowledge distillation technique mitigates privacy risk and cost to the same extent as the baseline FL algorithms, but allows flexible aggregation over heterogeneous client models that can differ e.g. in size, numerical precision or structure. We show in extensive empirical experiments on various CV/NLP datasets (CIFAR-10/100, ImageNet, AG News, SST2) and settings (heterogeneous models/data) that the server model can be trained much faster, requiring fewer communication rounds than any existing FL technique so far.

Added

2026-09-24

Learning under Concept Drift: A Review

Learning under Concept Drift: A Review

Jie Lu, Anjin Liu, Fan Dong, Feng Gu, João Gama, Guangquan Zhang

Why you should read this

Establishes a unified structural framework covering concept drift detection, understanding, and adaptation while evaluating 24 benchmark datasets across more than 130 studies to guide machine learning on non-stationary data streams.

Concept drift describes unforeseeable changes in the underlying distribution of streaming data over time. Concept drift research involves the development of methodologies and techniques for drift detection, understanding and adaptation. Data analysis has revealed that machine learning in a concept drift environment will result in poor learning results if the drift is not addressed. To help researchers identify which research topics are significant and how to apply related techniques in data analysis tasks, it is necessary that a high quality, instructive review of current research developments and trends in the concept drift field is conducted. In addition, due to the rapid development of concept drift in recent years, the methodologies of learning under concept drift have become noticeably systematic, unveiling a framework which has not been mentioned in literature. This paper reviews over 130 high quality publications in concept drift related research areas, analyzes up-to-date developments in methodologies and techniques, and establishes a framework of learning under concept drift including three main components: concept drift detection, concept drift understanding, and concept drift adaptation. This paper lists and discusses 10 popular synthetic datasets and 14 publicly available benchmark datasets used for evaluating the performance of learning algorithms aiming at handling concept drift. Also, concept drift related research directions are covered and discussed. By providing state-of-the-art knowledge, this survey will directly support researchers in their understanding of research developments in the field of learning under concept drift.

Added

2026-09-18

SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary

SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary

Alberto Fernandez, Salvador Garcia, Francisco Herrera, Nitesh V. Chawla

OrganizationsUniversity of GranadaUniversity of Notre Dame

Why you should read this

Synthesizes fifteen years of research on the Synthetic Minority Oversampling Technique (SMOTE), detailing its numerous algorithm variants, cross-paradigm applications, and key open challenges in scaling to big data and complex class distributions.

The Synthetic Minority Oversampling Technique (SMOTE) preprocessing algorithm is considered “de facto” standard in the framework of learning from imbalanced data. This is due to its simplicity in the design of the procedure, as well as its robustness when applied to different type of problems. Since its publication in 2002, SMOTE has proven successful in a variety of applications from several different domains. SMOTE has also inspired several approaches to counter the issue of class imbalance, and has also significantly contributed to new supervised learning paradigms, including multilabel classification, incremental learning, semi-supervised learning, multi-instance learning, among others. It is standard benchmark for learning from imbalanced data. It is also featured in a number of different software packages — from open source to commercial. In this paper, marking the fifteen year anniversary of SMOTE, we reflect on the SMOTE journey, discuss the current state of affairs with SMOTE, its applications, and also identify the next set of challenges to extend SMOTE for Big Data problems.

Added

2026-09-16

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization

Thomas G. Dietterich

OrganizationsOregon State University

Why you should read this

Demonstrates across 33 benchmark datasets that while boosting generates the most accurate decision tree ensembles on clean data, bagging remains superior under classification noise because boosting assigns excessive weight to mislabeled training instances.

Bagging and boosting are methods that generate a diverse ensemble of classifiers by manipulating the training data given to a "base" learning algorithm. Breiman has pointed out that they rely for their effectiveness on the instability of the base learning algorithm. An alternative approach to generating an ensemble is to randomize the internal decisions made by the base algorithm. This general approach has been studied previously by Ali and Pazzani and by Dietterich and Kong. This paper compares the effectiveness of randomization, bagging, and boosting for improving the performance of the decision-tree algorithm C4.5. The experiments show that in situations with little or no classification noise, randomization is competitive with (and perhaps slightly superior to) bagging but not as accurate as boosting. In situations with substantial classification noise, bagging is much better than boosting, and sometimes better than randomization.

Added

2026-09-12