Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Data Distribution Generation

Data distribution generation is a machine learning approach that models and forecasts the future evolution of a data distribution over time and synthesizes prospective training samples according to that anticipated distribution. In non-stationary environments where continuous streaming data experiences concept drift, this technique captures predictable underlying trends in environmental changes rather than solely reacting to shifts after they are detected. By forecasting upcoming distribution states and proactively generating representative synthetic data, predictive models can be adapted in advance to maintain performance and generalization across evolving real-world scenarios.

1 item

DDG-DA: Data Distribution Generation for Predictable Concept Drift Adaptation

DDG-DA: Data Distribution Generation for Predictable Concept Drift Adaptation

Wendi Li, Xiao Yang, Weiqing Liu, Yingce Xia, Jiang Bian

OrganizationsMicrosoftUniversity of Wisconsin Madison

Why you should read this

Proposes DDG-DA, a proactive adaptation framework that forecasts future streaming data distributions and resamples historical samples via a differentiable distribution distance to train predictive models before concept drift occurs.

In many real-world scenarios, we often deal with streaming data that is sequentially collected over time. Due to the non-stationary nature of the environment, the streaming data distribution may change in unpredictable ways, which is known as concept drift. To handle concept drift, previous methods first detect when/where the concept drift happens and then adapt models to fit the distribution of the latest data. However, there are still many cases that some underlying factors of environment evolution are predictable, making it possible to model the future concept drift trend of the streaming data, while such cases are not fully explored in previous work. In this paper, we propose a novel method DDG-DA, that can effectively forecast the evolution of data distribution and improve the performance of models. Specifically, we first train a predictor to estimate the future data distribution, then leverage it to generate training samples, and finally train models on the generated data. We conduct experiments on three real-world tasks (forecasting on stock price trend, electricity load and solar irradiance) and obtain significant improvement on multiple widely-used models.

Added

2026-09-26