Transformed Distribution Matching for Missing Value Imputation
He ZhaoKe SunAmir DezfouliEdwin V. Bonilla
Proposes a missing value imputation framework that matches the empirical distributions of data batches in a learned latent space via deep invertible transformations, overcoming the geometric limitations of raw-data optimal transport while achieving state-of-the-art results across diverse missingness mechanisms.
Real-world datasets in critical fields like healthcare and operations frequently contain incomplete records with substantial missing values. Imputing these missing entries accurately without knowing the underlying true values or the specific downstream application is vital for reliable data analysis, predictive modeling, and informed decision-making.
The article introduces and evaluates Transformed Distribution Matching, a new unsupervised framework designed to impute missing continuous numerical values. The method aims to accurately capture complex data geometries by transforming incomplete data into a learned latent space and aligning sample distributions simultaneously.
The authors develop a method combining optimal transport—a mathematical framework for comparing probability distributions—with invertible neural networks. By mapping data batches into a transformed latent representation, standard distance metrics better reflect true data relationships while the invertible architecture prevents information loss and model collapse. To assess credibility and performance, the authors conducted extensive empirical evaluations across twelve standard real-world benchmark datasets, spanning varying sample sizes and feature dimensions. The method was benchmarked against ten baseline techniques across four distinct missing-data patterns with 30% missing rates.
The findings show that Transformed Distribution Matching consistently outperforms existing imputation approaches across standard evaluation metrics, including mean absolute error, root-mean-square error, and distribution distance. Downstream machine learning classification models trained on data imputed by this framework achieved the highest overall accuracy, demonstrating practical performance gains. The method maintained robust stability across training iterations without severe overfitting. However, learning the deep transformation layers requires approximately two to three times more computational runtime per iteration compared to simpler distribution-matching baselines.
These results demonstrate that transforming data into an invertible latent space substantially reduces imputation errors when handling data with complex geometrical structures. For organizations relying on machine learning pipelines over incomplete data, adopting this approach reduces the risk of biased analyses and poor downstream model performance, requiring minimal hyperparameter tuning compared to traditional multi-stage generative models.
Organizations handling incomplete continuous tabular data should consider deploying this transformed distribution matching approach in critical analytical pipelines where accuracy is paramount. Because the framework introduces a trade-off between higher imputation fidelity and increased computational training time, technical teams should evaluate compute resource budgets on large-scale datasets before full production integration.
The current framework is limited to real-valued continuous features and cannot directly process categorical variables without further adaptation. Confidence in the empirical results is high given the breadth of datasets and missing-data mechanisms tested, though practitioners should exercise caution when working with large-scale categorical records or latency-constrained environments until further algorithmic extensions are developed.
- Paper: GAIN: Missing Data Imputation using Generative Adversarial Nets, Jinsung Yoon et al. (2018). Read GAIN first to see an influential learned imputation method and baseline against which the source’s distribution-matching approach is positioned.
- Paper: MissForest - non-parametric missing value imputation for mixed-type data, Daniel J. Stekhoven et al. (2011). MissForest establishes a widely used nonlinear imputation approach for mixed data, clarifying the conventional benchmark that the source’s continuous-feature method advances beyond.
- Paper: Optimal Transport for Domain Adaptation, Nicolas Courty et al. (2014). This paper introduces optimal transport as a way to align data distributions, providing useful mathematical context for the transport-based matching at the heart of the source.
No sufficiently relevant recommendations were found.
