Transformed Distribution Matching for Missing Value Imputation

He ZhaoKe SunAmir DezfouliEdwin V. Bonilla

article2023ICML45 citations

Proposes a missing value imputation framework that matches the empirical distributions of data batches in a learned latent space via deep invertible transformations, overcoming the geometric limitations of raw-data optimal transport while achieving state-of-the-art results across diverse missingness mechanisms.

Listen

Real-world datasets in critical fields like healthcare and operations frequently contain incomplete records with substantial missing values. Imputing these missing entries accurately without knowing the underlying true values or the specific downstream application is vital for reliable data analysis, predictive modeling, and informed decision-making.

The article introduces and evaluates Transformed Distribution Matching, a new unsupervised framework designed to impute missing continuous numerical values. The method aims to accurately capture complex data geometries by transforming incomplete data into a learned latent space and aligning sample distributions simultaneously.

The authors develop a method combining optimal transport—a mathematical framework for comparing probability distributions—with invertible neural networks. By mapping data batches into a transformed latent representation, standard distance metrics better reflect true data relationships while the invertible architecture prevents information loss and model collapse. To assess credibility and performance, the authors conducted extensive empirical evaluations across twelve standard real-world benchmark datasets, spanning varying sample sizes and feature dimensions. The method was benchmarked against ten baseline techniques across four distinct missing-data patterns with 30% missing rates.

The findings show that Transformed Distribution Matching consistently outperforms existing imputation approaches across standard evaluation metrics, including mean absolute error, root-mean-square error, and distribution distance. Downstream machine learning classification models trained on data imputed by this framework achieved the highest overall accuracy, demonstrating practical performance gains. The method maintained robust stability across training iterations without severe overfitting. However, learning the deep transformation layers requires approximately two to three times more computational runtime per iteration compared to simpler distribution-matching baselines.

These results demonstrate that transforming data into an invertible latent space substantially reduces imputation errors when handling data with complex geometrical structures. For organizations relying on machine learning pipelines over incomplete data, adopting this approach reduces the risk of biased analyses and poor downstream model performance, requiring minimal hyperparameter tuning compared to traditional multi-stage generative models.

Organizations handling incomplete continuous tabular data should consider deploying this transformed distribution matching approach in critical analytical pipelines where accuracy is paramount. Because the framework introduces a trade-off between higher imputation fidelity and increased computational training time, technical teams should evaluate compute resource budgets on large-scale datasets before full production integration.

The current framework is limited to real-valued continuous features and cannot directly process categorical variables without further adaptation. Confidence in the empirical results is high given the breadth of datasets and missing-data mechanisms tested, though practitioners should exercise caution when working with large-scale categorical records or latency-constrained environments until further algorithmic extensions are developed.

No sufficiently relevant recommendations were found.

Cover for Transformed Distribution Matching for Missing Value Imputation

Abstract

We study the problem of imputing missing values in a dataset, which has important applications in many domains. The key to missing value imputation is to capture the data distribution with incomplete samples and impute the missing values accordingly. In this paper, by leveraging the fact that any two batches of data with missing values come from the same data distribution, we propose to impute the missing values of two batches of samples by transforming them into a latent space through deep invertible functions and matching them distributionally. To learn the transformations and impute the missing values simultaneously, a simple and well-motivated algorithm is proposed. Our algorithm has fewer hyperparameters to fine-tune and generates high-quality imputations regardless of how missing values are generated. Extensive experiments over a large number of datasets and competing benchmark algorithms show that our method achieves state-of-the-art performance¹.

Knowls

  1. Knowl 1 — TDM jointly imputes missing entries and learns a transport space

    model/method

    Let X∈RN×DX\in\mathbb{R}^{N\times D} be a dataset with NN samples and DD real-valued features, and let MM mark its missing entries. TDM maintains candidate values for X[M]X[M] and applies a learned invertible transformation fθ:RD→RDf_\theta:\mathbb{R}^{D}\to\mathbb{R}^{D} to completed samples. For two independently sampled batches X1X_1 and X2X_2 of size BB, let μ(Xb)=B−1∑i=1BδXb[i,:]\mu(X_b)=B^{-1}\sum_{i=1}^{B}\delta_{X_b[i,:]} be the batch’s empirical distribution, and let fθ#μ(Xb)f_{\theta\#}\mu(X_b) denote its transformed distribution. TDM minimizes

    min⁡X[M], θ  W22 ⁣(fθ#μ(X1),fθ#μ(X2)),\min_{X[M],\,\theta}\; W_2^2\!\left(f_{\theta\#}\mu(X_1),f_{\theta\#}\mu(X_2)\right),

    where W22W_2^2 is the squared 2-Wasserstein distance with squared Euclidean cost in the transformed space. Thus the missing values and transformation parameters are learned together by making the transformed batch distributions agree. The method does not explicitly fit a data distribution or presume a downstream task.

  2. Knowl 2 — Invertibility prevents collapse by maximizing empirical mutual information

    theoretical result

    For an empirical random variable XX with finite support, the paper establishes that a smooth invertible map fθf_\theta maximizes mutual information between the input and its representation among candidate transformations f′f':

    I(X;fθ(X))=H(X)≥I(X;f′(X)).I(X;f_\theta(X))=H(X)\ge I(X;f'(X)).

    Here H(X)H(X) is the entropy of the empirical random variable, and II denotes mutual information. The result applies when the transformation is smooth and invertible on the same-dimensional input and output space. Because an invertible map retains all information in XX, it rules out the constant representation that would otherwise trivially reduce TDM’s distribution-matching loss to zero regardless of the imputations. TDM enforces this property through invertible neural networks instead of adding an explicit mutual-information estimation loss.

  3. Knowl 3 — Affine coupling blocks implement the invertible transformation

    model/method

    TDM builds fθf_\theta as a composition of TT invertible blocks, each mapping RD\mathbb{R}^{D} to RD\mathbb{R}^{D}. A block splits its input yiny^{\mathrm{in}} into disjoint subvectors y1:diny^{\mathrm{in}}_{1:d} and yd+1:Diny^{\mathrm{in}}_{d+1:D}, with d=⌊D/2⌋d=\lfloor D/2\rfloor, and applies two complementary affine coupling transformations:

    y1:dout=y1:din⊙exp⁡ ⁣(g1(yd+1:Din))+h1(yd+1:Din),y^{\mathrm{out}}_{1:d}=y^{\mathrm{in}}_{1:d}\odot\exp\!\left(g_1(y^{\mathrm{in}}_{d+1:D})\right)+h_1(y^{\mathrm{in}}_{d+1:D}), yd+1:Dout=yd+1:Din⊙exp⁡ ⁣(g2(y1:dout))+h2(y1:dout).y^{\mathrm{out}}_{d+1:D}=y^{\mathrm{in}}_{d+1:D}\odot\exp\!\left(g_2(y^{\mathrm{out}}_{1:d})\right)+h_2(y^{\mathrm{out}}_{1:d}).

    The operator ⊙\odot is elementwise multiplication; g1,g2,h1,h2g_1,g_2,h_1,h_2 are neural networks, and the exponential and additive terms act elementwise. Each of these four networks is a succession of fully connected layers with SELU activations; the outputs of g1g_1 and g2g_2 are clamped using the arctan function. The sequential coupling structure makes each block invertible, so their composition is invertible. The paper’s experiments use three fully connected layers per network, each of width KDK D.

  4. Knowl 4 — Training procedure optimizes imputations and transformation parameters together

    algorithm

    The TDM training procedure takes a real-valued data matrix with missing-entry mask MM and returns completed values and a learned transformation. It initializes each missing cell to its feature’s observed-value mean plus noise sampled from N(0,0.1)\mathcal{N}(0,0.1), initializes the parameters θ\theta of the invertible network, and then repeatedly samples two batches of size BB. For the experiments, B=512B=512 when the dataset is large enough; for N<512N<512, the paper uses 2⌊N/2⌋2\lfloor N/2\rfloor. It computes transformed pairwise squared Euclidean costs, solves the equal-weight 2-Wasserstein transport problem with the network simplex method, and uses the resulting loss to update both the batch’s missing cells and θ\theta with RMSprop. The experiments use learning rate 10−210^{-2}, T=3T=3 coupling blocks, K=2K=2, and 10,000 iterations. The network-simplex solver avoids the entropic-regularization parameter required by Sinkhorn iterations.

    Input: Real-valued data X, missing-entry mask M
    Output: Completed data X and learned invertible map f_theta
    Initialize each missing cell with its feature's observed mean plus noise from Normal(0, 0.1)
    Initialize transformation parameters theta
    Set batch size B to 512, or to 2 floor(N/2) when N is less than 512
    For 10,000 iterations:
        Sample two batches X1 and X2, each of size B
        Transform both batches using f_theta
        Form pairwise costs G[i,j] = squared Euclidean distance between transformed samples
        Solve the equal-weight optimal transport problem for G using network simplex
        Compute the squared 2-Wasserstein loss from the optimal transport plan
        Use RMSprop with learning rate 0.01 to update theta and the missing entries in X1 and X2
    Return the completed X and f_theta
  5. Knowl 5 — Expected minibatch matching loss has global-distribution and batch-size properties

    theoretical result

    Let X~∈RN×D\widetilde X\in\mathbb{R}^{N\times D} be a fixed, completed dataset, let fθf_\theta be a fixed transformation, and define its transformed full-data empirical distribution as ν=N−1∑i=1Nδfθ(X~[i,:])\nu=N^{-1}\sum_{i=1}^{N}\delta_{f_\theta(\widetilde X[i,:])}. A minibatch is formed by independent uniform draws from the NN samples, and QbQ_b denotes the transformed empirical distribution of a batch of size bb. For two independent batches Qb,Qb′Q_b,Q'_b, the paper establishes

    E W22(Qb,Qb′)≥E W22(Qb,ν).\mathbb{E}\,W_2^2(Q_b,Q'_b)\ge \mathbb{E}\,W_2^2(Q_b,\nu).

    Thus the expected batch-to-batch loss lower-bounds the expected distance from a sampled batch to the full-data distribution. It also establishes that doubling the batch size cannot increase the expected batch-to-batch loss:

    E W22(Q2b,Q2b′)≤E W22(Qb,Qb′).\mathbb{E}\,W_2^2(Q_{2b},Q'_{2b})\le \mathbb{E}\,W_2^2(Q_b,Q'_b).

    For batch size one, the loss is the expected squared distance between two independently drawn transformed samples, E ∥fθ(X~[I,:])−fθ(X~[J,:])∥22\mathbb{E}\,\|f_\theta(\widetilde X[I,:])-f_\theta(\widetilde X[J,:])\|_2^2, where II and JJ are independent uniform sample indices. These properties describe the loss under fixed imputations and fixed transformation parameters; during training, both can change.

  6. Knowl 6 — Synthetic examples show the benefit of learning a geometry-aware space

    empirical result

    In two-dimensional synthetic examples with 500 samples per dataset, 60% of points are fully observed and 40% have one coordinate missing under MCAR. On examples with complex shapes, TDM’s imputations align with the underlying data structure more closely than imputations from OTImputer, which matches samples using quadratic distance directly in the original feature space and produces visibly misplaced values. In a further visualization, out-of-distribution points drawn from a two-dimensional normal distribution—excluded from training—are separated from in-domain points by TDM’s learned transformation. This is consistent with the method’s intended geometry: distances in the transformed space distinguish the in-domain structure from unrelated points. On simpler synthetic shapes, both methods impute well, and TDM’s learned transformation remains close to an isometry.

  7. Knowl 7 — Benchmark protocol covers four missingness settings and multiple metrics

    experimental setup

    The experiments use 12 standardized UCI datasets: california, qsar biodegradation, blood transfusion, wine quality, parkinsons, yacht hydrodynamics, seeds, glass, planning relax, concrete slump, anuran calls, and letter. Each dataset is evaluated with a 30% missing rate in four settings: MCAR, generated independently of the data; MAR, generated by a logistic model using a sampled subset of fully observed features; MNARL, generated with a logistic model whose input is masked by MCAR; and MNARQ, generated by randomly selecting missing values from ranges defined by lower and upper percentiles. Ten masks with different random seeds are used per dataset and setting, and means and standard deviations are reported.

    Imputation quality is assessed using mean absolute error (MAE), root-mean-square error (RMSE), and squared 2-Wasserstein distance between imputed and ground-truth distributions in the original data space; lower is better. For downstream classification, the experiments use an RBF-kernel support vector machine with an automatically selected kernel coefficient and report mean five-fold cross-validation accuracy over 10 runs. The comparison includes iterative imputation, generative models, OT-based methods, matrix completion, and causal-refinement methods, including ICE, MissForest, GAIN, MIWAE, MCFlow, EMFlow, OTImputer, SoftImpute, and MIRACLE.

  8. Knowl 8 — TDM improves imputation and downstream classification across the benchmark

    empirical result

    Across the four missingness settings, the 12 UCI datasets, and the reported imputation metrics, the paper reports that TDM achieves the best results in almost all comparisons. Against the two OTImputer variants, TDM’s learned transformations improve performance substantially; using network simplex rather than Sinkhorn yields a marginal but generally consistent improvement for OTImputer. Although TDM optimizes Wasserstein distance in the transformed space rather than directly minimizing data-space distance, it also outperforms OTImputer on the reported data-space Wasserstein evaluation. MCFlow and EMFlow, which also use invertible networks, do not match TDM’s reported imputation performance. In downstream classification, TDM generally has the highest accuracy among the imputation methods, though the paper does not report exact numerical values in its plots.

    TDM requires more computation per iteration than OTImputer in the four reported timing comparisons: on glass, 3.20 versus 1.14 seconds; seeds, 3.19 versus 1.03 seconds; blood transfusion, 3.35 versus 2.39 seconds; and anuran calls, 4.06 versus 2.33 seconds. The paper describes this as roughly two to three times the iteration time, with a smaller apparent runtime gap on larger datasets. It also reports that OTImputer can overfit on some datasets, such as glass, while TDM is more stable during training.

  9. Knowl 9 — Performance is relatively insensitive to coupling-network width

    empirical result

    In sensitivity experiments varying the number of invertible blocks TT and network-width multiplier KK from 1 to 4, increasing TT generally improves imputation performance, but the gain from T=3T=3 to T=4T=4 is marginal and overfitting appears on some datasets, including qsar biodegradation and anuran calls. Changing KK has no significant performance impact in the tested range. Separate comparisons using batch sizes 128, 256, and 512 show that larger batches generally improve both TDM and OTImputer. Based on these experiments, the paper uses T=3T=3 and K=2K=2 as practical settings.

  10. Knowl 10 — The method is limited to real-valued data and is slower than OTImputer

    limitation

    The paper’s implementation and experiments address real-valued features; TDM, like OTImputer, does not work with categorical data. Learning the neural transformations also makes TDM slower per iteration than OTImputer: in four reported dataset comparisons, its iteration time is about two to three times larger, although the runtime gap appears smaller on larger datasets. The paper identifies developing more efficient and broadly applicable versions as future work.

Coverage note — The OTImputer permutation characterization and proof-only intermediate steps are omitted because they motivate or support the contributions rather than add standalone contributed knowledge; exact per-dataset metric values are also omitted because the paper presents them as plots rather than numerical tables.

References

  1. 1.Ahuja, R. K., Magnanti, T. L., Orlin, J. B., and Reddy, M. Applications of network optimization. Handbooks in Operations Research and Management Science, 7:1–83, 1995.
  2. 2.Altschuler, J., Niles-Weed, J., and Rigollet, P. Near-linear time approximation algorithms for optimal transport via sinkhorn iteration. In NeurIPS, 2017.
  3. 3.Ardizzone, L., Kruse, J., Rother, C., and Kothe, U. Analyzing inverse problems with invertible neural networks. In ICLR, 2019.
  4. 4.Bachman, P., Hjelm, R. D., and Buchwalter, W. Learning representations by maximizing mutual information across views. In NeurIPS, 2019.
  5. 5.Barnard, J. and Meng, X.-L. Applications of multiple imputation in medical studies: from AIDS to NHANES. Statistical methods in medical research, 8(1):17–36, 1999.
  6. 6.Barthe, F. and Bordenave, C. Combinatorial Optimization Over Two Random Point Sets, pp. 483–535. Springer International Publishing, Heidelberg, 2013.
  7. 7.Bonneel, N., Van De Panne, M., Paris, S., and Heidrich, W. Displacement interpolation using Lagrangian mass transport. In SIGGRAPH Asia, pp. 1–12, 2011.
  8. 8.Bui, A. T., Le, T., Tran, Q. H., Zhao, H., and Phung, D. A unified Wasserstein distributional robustness framework for adversarial training. In ICLR, 2022.
  9. 9.Burda, Y., Grosse, R. B., and Salakhutdinov, R. Importance weighted autoencoders. In ICLR, 2016.
  10. 10.Chen, K., Liang, X., Zhang, Z., and Ma, Z. GEDI: A graph-based end-to-end data imputation framework. arXiv preprint arXiv:2208.06573, 2022.
  11. 11.Coeurdoux, F., Dobigeon, N., and Chainais, P. Learning optimal transport between two empirical distributions with normalizing flows. In ECML PKDD, 2022.
  12. 12.Cuturi, M. Sinkhorn distances: Lightspeed computation of optimal transport. In NeurIPS, 2013.
  13. 13.Cuturi, M. and Doucet, A. Fast computation of Wasserstein barycenters. In ICML, pp. 685–693, 2014.
  14. 14.Dai, Z., Bu, Z., and Long, Q. Multiple imputation via generative adversarial network for high-dimensional blockwise missing value problems. In 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 791–798, 2021.
  15. 15.Dai, Z., Bu, Z., and Long, Q. Multiple imputation with neural network Gaussian process for high-dimensional incomplete data. In ACML, 2022.
  16. 16.Dinh, L., Krueger, D., and Bengio, Y. NICE: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014.
  17. 17.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using Real NVP. In ICLR, 2017.
  18. 18.Dvurechensky, P., Gasnikov, A., and Kroshnin, A. Computational optimal transport: Complexity by accelerated gradient descent is better than by Sinkhorn’s algorithm. In ICML, pp. 1367–1376, 2018.
  19. 19.Fang, F. and Bao, S. FragmGAN: Generative adversarial nets for fragmentary data imputation and prediction. arXiv preprint arXiv:2203.04692, 2022.
  20. 20.Feydy, J., Sejourne, T., Vialard, F.-X., Amari, S.-i., Trouve, A., and Peyre, G. Interpolating between optimal transport and MMD using Sinkhorn divergences. In AISTATS, pp. 2681–2690, 2019.
  21. 21.Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Bois-bunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., Gautheron, L., Gayraud, N. T., Janati, H., Rakotomamonjy, A., Redko, I., Rolet, A., Schutz, A., Seguy, V., Sutherland, D. J., Tavenard, R., Tong, A., and Vayer, T. POT: Python optimal transport. JMLR, 22(78):1–8, 2021.
  22. 22.Gao, Z., Niu, Y., Cheng, J., Tang, J., Xu, T., Zhao, P., Li, L., Tsung, F., and Li, J. Handling missing data via max-entropy regularized graph autoencoder. In AAAI, 2023.
  23. 23.Ge, Z., Liu, S., Li, Z., Yoshie, O., and Sun, J. OTA: Optimal transport assignment for object detection. In CVPR, pp. 303–312, 2021.
  24. 24.Gelman, A. Parameterization and Bayesian modeling. Journal of the American Statistical Association, 99(466):537–545, 2004.
  25. 25.Genevay, A., Peyre, G., and Cuturi, M. Learning generative models with Sinkhorn divergences. In AISTATS, pp. 1608–1617, 2018.
  26. 26.Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B. The reversible residual network: Backpropagation without storing activations. In NeurIPS, 2017.
  27. 27.Gondara, L. and Wang, K. Multiple imputation using deep denoising autoencoders. arXiv preprint arXiv:1705.02737, 280, 2017.
  28. 28.Gong, Y., Hajimirsadeghi, H., He, J., Durand, T., and Mori, G. Variational selective autoencoder: Learning from partially-observed heterogeneous data. In AISTATS, pp. 2377–2385, 2021.
  29. 29.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020.
  30. 30.Guo, D., Tian, L., Zhang, M., Zhou, M., and Zha, H. Learning prototype-oriented set representations for meta-learning. In ICLR, 2021.
  31. 31.Guo, D., Li, Z., Zhao, H., Zhou, M., and Zha, H. Learning to re-weight examples with optimal transport for imbalanced classification. In NeurIPS, 2022a.
  32. 32.Guo, D., Tian, L., Zhao, H., Zhou, M., and Zha, H. Adaptive distribution calibration for few-shot learning with hierarchical optimal transport. In NeurIPS, 2022b.
  33. 33.Guo, D., Zhao, H., Zheng, H., Tanwisuth, K., Chen, B., Zhou, M., et al. Representing mixtures of word embeddings with mixtures of topic embeddings. In ICLR, 2022c.
  34. 34.Hastie, T., Mazumder, R., Lee, J. D., and Zadeh, R. Matrix completion and low-rank SVD via fast alternating least squares. JMLR, 16(1):3367–3402, 2015.
  35. 35.Heckerman, D., Chickering, D. M., Meek, C., Rounthwaite, R., and Kadie, C. Dependency networks for inference, collaborative filtering, and data visualization. JMLR, 1 (Oct):49–75, 2000.
  36. 36.Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. Learning deep representations by mutual information estimation and maximization. In ICLR, 2019.
  37. 37.Huang, B., Zhu, Y., Usman, M., Zhou, X., and Chen, H. Graph neural networks for missing value classification in a task-driven metric space. TKDE, 2022.
  38. 38.Huynh, V., Zhao, H., and Phung, D. OTLDA: A geometry-aware optimal transport approach for topic modeling. In NeurIPS, volume 33, pp. 18573–18582, 2020.
  39. 39.Ivanov, O., Figurnov, M., and Vetrov, D. Variational autoencoder with arbitrary conditioning. In ICLR, 2018.
  40. 40.Jacobsen, J.-H., Behrmann, J., Zemel, R., and Bethge, M. Excessive invariance causes adversarial vulnerability. In ICLR, 2019.
  41. 41.Jarrett, D., Cebere, B. C., Liu, T., Curth, A., and van der Schaar, M. HyperImpute: Generalized iterative imputation with automatic model selection. In ICML, pp. 9916–9937, 2022.
  42. 42.Kingma, D. P. and Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. In NeurIPS, 2018.
  43. 43.Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S. Self-normalizing neural networks. NeurIPS, 30, 2017.
  44. 44.Kobyzev, I., Prince, S. J., and Brubaker, M. A. Normalizing flows: An introduction and review of current methods. IEEE TPAMI, 43(11):3964–3979, 2020.
  45. 45.Kraskov, A., Stogbauer, H., and Grassberger, P. Estimating mutual information. Physical review E, 69(6):066138, 2004.
  46. 46.Kyono, T., Zhang, Y., Bellot, A., and van der Schaar, M. MIRACLE: Causally-aware imputation via learning missing data mechanisms. In NeurIPS, volume 34, pp. 23806–23817, 2021.
  47. 47.Li, S. C.-X., Jiang, B., and Marlin, B. MisGAN: Learning from incomplete data with generative adversarial networks. In ICLR, 2018.
  48. 48.Linsker, R. Self-organization in a perceptual network. Computer, 21(3):105–117, 1988.
  49. 49.Little, R. J. and Rubin, D. B. Statistical analysis with missing data, volume 793. John Wiley & Sons, 2019.
  50. 50.Liu, J., Gelman, A., Hill, J., Su, Y.-S., and Kropko, J. On the stationary distribution of iterative imputations. Biometrika, 101(1):155–173, 2014.
  51. 51.Ma, Q. and Ghosh, S. K. EMFlow: Data imputation in latent space via em and deep flow models. arXiv preprint arXiv:2106.04804, 2021.
  52. 52.Mattei, P.-A. and Frellsen, J. MIWAE: Deep generative modelling and imputation of incomplete data sets. In ICML, pp. 4413–4423, 2019.
  53. 53.Mayer, I., Sportisse, A., Josse, J., Tierney, N., and Vialaneix, N. R-miss-tastic: A unified platform for missing values methods and workflows. arXiv preprint arXiv:1908.04822, 2019.
  54. 54.Mazumder, R., Hastie, T., and Tibshirani, R. Spectral regularization algorithms for learning large incomplete matrices. JMLR, 11:2287–2322, 2010.
  55. 55.Mohan, K., Pearl, J., and Tian, J. Graphical models for inference with missing data. In NeurIPS, volume 26, 2013.
  56. 56.Morales-Alvarez, P., Gong, W., Lamb, A., Woodhead, S., Jones, S. P., Pawlowski, N., Allamanis, M., and Zhang, C. Simultaneous missing value imputation and structure learning with groups. In NeurIPS, 2022.
  57. 57.Muzellec, B., Josse, J., Boyer, C., and Cuturi, M. Missing data imputation using optimal transport. In ICML, pp. 7130–7140, 2020.
  58. 58.Nazabal, A., Olmos, P. M., Ghahramani, Z., and Valera, I. Handling incomplete heterogeneous data using VAEs. Pattern Recognition, 107:107501, 2020.
  59. 59.Nguyen, T., Le, T., Zhao, H., Tran, Q. H., Nguyen, T., and Phung, D. Most: Multi-source domain adaptation via optimal transport for student-teacher learning. In UAI, pp. 225–235, 2021.
  60. 60.Nguyen, T., Nguyen, V., Le, T., Zhao, H., Tran, Q. H., and Phung, D. Cycle class consistency with distributional optimal transport and knowledge distillation for unsupervised domain adaptation. In UAI, pp. 1519–1529, 2022.
  61. 61.Nguyen, X. Wasserstein distances for discrete measures and convergence in nonparametric mixture models. arXiv preprint arXiv:1109.3250v1, 2011.
  62. 62.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  63. 63.Papamakarios, G., Nalisnick, E. T., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. Normalizing flows for probabilistic modeling and inference. JMLR, 22(57):1–64, 2021.
  64. 64.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. Scikit-learn: Machine learning in python. JMLR, 12:2825–2830, 2011.
  65. 65.Peis, I., Ma, C., and Hernandez-Lobato, J. M. Missing data imputation and acquisition with deep hierarchical models and Hamiltonian Monte Carlo. In NeurIPS, 2022.
  66. 66.Peyre, G., Cuturi, M., et al. Computational optimal transport: With applications to data science. Foundations and Trends® in Machine Learning, 11(5-6):355–607, 2019.
  67. 67.Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G. On variational bounds of mutual information. In ICML, pp. 5171–5180, 2019.
  68. 68.Raghunathan, T. E., Lepkowski, J. M., Van Hoewyk, J., Solenberger, P., et al. A multivariate technique for multiply imputing missing values using a sequence of regression models. Survey methodology, 27(1):85–96, 2001.
  69. 69.Richardson, T. W., Wu, W., Lin, L., Xu, B., and Bernal, E. A. MCFLOW: Monte Carlo flow models for data imputation. In CVPR, pp. 14205–14214, 2020.
  70. 70.Rubin, D. B. Inference and missing data. Biometrika, 63(3):581–592, 1976.
  71. 71.Rubin, D. B. Multiple imputation for nonresponse in surveys, volume 81. John Wiley & Sons, 2004.
  72. 72.Seaman, S., Galati, J., Jackson, D., and Carlin, J. What is meant by “missing at random”? Statistical Science, 28 (2):257–268, 2013.
  73. 73.Stekhoven, D. J. and Bühlmann, P. MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112–118, 2012.
  74. 74.Tieleman, T., Hinton, G., et al. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4 (2):26–31, 2012.
  75. 75.Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S., and Lucic, M. On mutual information maximization for representation learning. In ICLR, 2020.
  76. 76.Van Buuren, S. Flexible imputation of missing data. CRC press, 2018.
  77. 77.Van Buuren, S. and Groothuis-Oudshoorn, K. MICE: Multivariate imputation by chained equations in r. Journal of statistical software, 45:1–67, 2011.
  78. 78.Van Buuren, S., Brand, J. P., Groothuis-Oudshoorn, C. G., and Rubin, D. B. Fully conditional specification in multivariate imputation. Journal of statistical computation and simulation, 76(12):1049–1064, 2006.
  79. 79.Vinas, R., Zheng, X., and Hayes, J. A graph-based imputation method for sparse medical records. arXiv preprint arXiv:2111.09084, 2021.
  80. 80.Vo, V., Le, T., Vuong, L.-T., Zhao, H., Bonilla, E., and Phung, D. Learning directed graphical models with optimal transport. arXiv preprint arXiv:2305.15927, 2023.
  81. 81.Vuong, T.-L., Le, T., Zhao, H., Zheng, C., Harandi, M., Cai, J., and Phung, D. Vector quantized Wasserstein auto-encoder. arXiv preprint arXiv:2302.05917, 2023.
  82. 82.Wang, S., Li, J., Miao, H., Zhang, J., Zhu, J., and Wang, J. Generative-free urban flow imputation. In CIKM, pp. 2028–2037, 2022.
  83. 83.Wang, Y., Li, D., Xu, C., and Yang, M. Missingness augmentation: A general approach for improving generative imputation models. arXiv preprint arXiv:2108.02566, 2021.
  84. 84.Yoon, J., Jordon, J., and Schaar, M. GAIN: Missing data imputation using generative adversarial nets. In ICML, pp. 5689–5698, 2018.
  85. 85.Yoon, S. and Sull, S. GAMIN: Generative adversarial multiple imputation network for highly missing data. In CVPR, pp. 8456–8464, 2020.
  86. 86.You, J., Ma, X., Ding, Y., Kochenderfer, M. J., and Leskovec, J. Handling missing data with graph representation learning. In NeurIPS, volume 33, pp. 19075–19087, 2020.
  87. 87.Zhang, C., Cai, Y., Lin, G., and Shen, C. Deepemd: Differentiable earth mover’s distance for few-shot learning. TPAMI, 2022.
  88. 88.Zhao, H., Phung, D., Huynh, V., Le, T., and Buntine, W. Neural topic model via optimal transport. In ICLR, 2021.
  89. 89.Zhu, J. and Raghunathan, T. E. Convergence properties of a sequential regression multiple imputation algorithm. Journal of the American Statistical Association, 110(511):1112–1124, 2015.

Citation

MLA
Zhao, H., et al. “Transformed Distribution Matching for Missing Value Imputation”. International Conference on Machine Learning, vol. 202, 2023, pp. 42159–86, https://proceedings.mlr.press/v202/zhao23h.html.
APA
Zhao, H., Sun, K., Dezfouli, A., & Bonilla, E. V. (2023). Transformed Distribution Matching for Missing Value Imputation. International Conference on Machine Learning, 202, 42159–42186. https://proceedings.mlr.press/v202/zhao23h.html
Chicago
Zhao, H., K. Sun, A. Dezfouli, and E. V. Bonilla. 2023. “Transformed Distribution Matching for Missing Value Imputation”. International Conference on Machine Learning 202: 42159–86. https://proceedings.mlr.press/v202/zhao23h.html.
Harvard
Zhao, H. et al. (2023) “Transformed Distribution Matching for Missing Value Imputation”, International Conference on Machine Learning. PMLR, pp. 42159–42186. Available at: https://proceedings.mlr.press/v202/zhao23h.html.
Vancouver
1. Zhao H, Sun K, Dezfouli A, Bonilla EV (2023) Transformed Distribution Matching for Missing Value Imputation. In: International Conference on Machine Learning. PMLR, pp 42159–42186

BibTeX

@InProceedings{pmlr-v202-zhao23h,
  title = 	 {Transformed Distribution Matching for Missing Value Imputation},
  author =       {Zhao, He and Sun, Ke and Dezfouli, Amir and Bonilla, Edwin V.},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {42159--42186},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/zhao23h/zhao23h.pdf},
  url = 	 {https://proceedings.mlr.press/v202/zhao23h.html},
  abstract = 	 {We study the problem of imputing missing values in a dataset, which has important applications in many domains. The key to missing value imputation is to capture the data distribution with incomplete samples and impute the missing values accordingly. In this paper, by leveraging the fact that any two batches of data with missing values come from the same data distribution, we propose to impute the missing values of two batches of samples by transforming them into a latent space through deep invertible functions and matching them distributionally. To learn the transformations and impute the missing values simultaneously, a simple and well-motivated algorithm is proposed. Our algorithm has fewer hyperparameters to fine-tune and generates high-quality imputations regardless of how missing values are generated. Extensive experiments over a large number of datasets and competing benchmark algorithms show that our method achieves state-of-the-art performance.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/