GAIN: Missing Data Imputation using Generative Adversarial Nets

Jinsung YoonJames JordonMihaela van der Schaar

article2018ICML1,456 citations

Presents Generative Adversarial Imputation Nets (GAIN), a framework that adapts GANs with a hint-conditioned discriminator to accurately model the underlying data distribution and outperform standard missing data imputation methods.

Listen

Missing data is a common challenge across healthcare, finance, and industrial analytics, occurring whenever measurements are lost, omitted, or too costly or risky to collect. Conventional imputation techniques often fail to capture complex distributions or require complete datasets during training, which is unrealistic for many real-world environments.

The article introduces and evaluates Generative Adversarial Imputation Nets (GAIN), a machine learning framework designed to infer missing values accurately even when only incomplete datasets are available for training.

The proposed method adapts generative adversarial frameworks by having a generator fill in missing values while a discriminator determines which individual components within a record are real and which are generated. To ensure the system learns the true underlying data distribution, the architecture introduces a hint mechanism that provides partial information about missingness to the discriminator. The authors evaluated the approach using five benchmark datasets spanning medical, language, and financial domains, assessing reconstruction error, downstream predictive accuracy, and parameter bias under varying missingness rates.

The evaluation produced four primary findings. First, GAIN consistently outperformed leading imputation baselines across all benchmark datasets, reducing imputation root mean square error by approximately 10% to 28% compared to standard alternatives. Second, ablation analysis confirmed that the adversarial architecture and the hint mechanism provide distinct performance boosts, yielding roughly 15% and 10% improvements respectively over simpler auto-encoder designs. Third, the framework remained robust under high data loss, demonstrating superior predictive performance over benchmark methods as missing data rates approached 80% to 90%. Fourth, the model preserved feature-label relationships more accurately than existing methods, achieving an 8.9% to 79.2% reduction in downstream parameter bias.

These results indicate that adopting this framework can enhance downstream decision-making and operational reliability by significantly reducing errors caused by incomplete data. The ability to train directly on incomplete records reduces data preparation costs and avoids discarding valuable partial information in data-constrained domains.

Organizations handling incomplete tabular datasets should consider integrating this framework into their data-processing pipelines, especially where high missingness degrades downstream model performance. Future work recommended by the source includes expanding evaluations into active sensing, recommender systems, and image error concealment.

Confidence in these findings is supported by rigorous cross-validated benchmarking across multiple domains. However, the theoretical guarantees presented are primarily established under the assumption that data is missing completely at random, meaning operational teams should exercise appropriate caution and validate results when working with non-random missingness patterns.

arXiv: 1806.02920
  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). It introduces the core Generative Adversarial Networks (GAN) framework of competing generator and discriminator networks that GAIN directly adapts for missing data imputation.
  • Paper: Conditional Generative Adversarial Nets, Mehdi Mirza et al. (2014). It formalizes conditioning in adversarial nets, providing the foundational mechanism GAIN relies upon to generate missing components conditioned on observed data and hint vectors.
  • Paper: Globally and locally consistent image completion, SATOSHI IIZUKA et al. (2017). It establishes the paradigm of using adversarial training for data completion and inpainting across arbitrary missing regions.
  • Paper: Generative Adversarial Networks: An Overview, Antonia Creswell et al. (2017). It provides a foundational overview of GAN training principles and architectures necessary for understanding adversarial generative modeling.
Cover for GAIN: Missing Data Imputation using Generative Adversarial Nets

Abstract

We propose a novel method for imputing missing data by adapting the well-known Generative Adversarial Nets (GAN) framework. Accordingly, we call our method Generative Adversarial Imputation Nets (GAIN). The generator (G) observes some components of a real data vector, imputes the missing components conditioned on what is actually observed, and outputs a completed vector. The discriminator (D) then takes a completed vector and attempts to determine which components were actually observed and which were imputed. To ensure that D forces G to learn the desired distribution, we provide D with some additional information in the form of a hint vector. The hint reveals to D partial information about the missingness of the original sample, which is used by D to focus its attention on the imputation quality of particular components. This hint ensures that G does in fact learn to generate according to the true data distribution. We tested our method on various datasets and found that GAIN significantly outperforms state-of-the-art imputation methods.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 2.1 Imputation
  • 3 Generative Adversarial Imputation Nets
  • 3.1 Generator
  • 3.2 Discriminator
  • 3.3 Hint
  • 3.4 Objective
  • 4 Theoretical Analysis
  • 5 GAIN Algorithm
  • 6 Experiments
  • 6.1 Source of gain
  • 6.2 Quantitative analysis of GAIN
  • 6.3 GAIN in different settings
  • 6.4 Prediction Performance
  • 6.5 Congeniality of GAIN
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Generative Adversarial Imputation Nets Architecture

    model/method

    Generative Adversarial Imputation Nets (GAIN) adapts the generative adversarial framework to impute missing values in data vectors conditioned on observed entries.

    Let X=X1×⋯×Xd\mathcal{X} = \mathcal{X}_1 \times \dots \times \mathcal{X}_d be a dd-dimensional feature space. A complete data vector is represented by a random variable X=(X1,…,Xd)∈XX = (X_1, \dots, X_d) \in \mathcal{X}, and a mask vector is represented by M=(M1,…,Md)∈{0,1}dM = (M_1, \dots, M_d) \in \{0, 1\}^d, where Mi=1M_i = 1 indicates that XiX_i is observed and Mi=0M_i = 0 indicates that XiX_i is missing. The partially observed data vector is defined as X~=(X~1,…,X~d)∈X~=X~1×⋯×X~d\tilde{X} = (\tilde{X}_1, \dots, \tilde{X}_d) \in \tilde{\mathcal{X}} = \tilde{\mathcal{X}}_1 \times \dots \times \tilde{\mathcal{X}}_d, where X~i=Xi∪{∗}\tilde{\mathcal{X}}_i = \mathcal{X}_i \cup \{*\} and:

    X~i={Xi,if Mi=1∗,otherwise\tilde{X}_i = \begin{cases} X_i, & \text{if } M_i = 1 \\ *, & \text{otherwise} \end{cases}

    GAIN consists of two neural networks:

    1. Generator (GG): Takes as input the observed vector X~\tilde{X} (with unobserved entries replaced by fixed values), the mask vector MM, and a masked noise vector (1−M)⊙Z(1 - M) \odot Z, where Z=(Z1,…,Zd)∈[0,1]dZ = (Z_1, \dots, Z_d) \in [0, 1]^d is independent random noise and ⊙\odot denotes element-wise multiplication. The generator outputs a full imputation vector Xˉ=G(X~,M,(1−M)⊙Z)∈X\bar{X} = G(\tilde{X}, M, (1 - M) \odot Z) \in \mathcal{X}. The completed data vector X^∈X\hat{X} \in \mathcal{X} is then constructed by replacing only the unobserved entries with the generator's outputs:

    X^=M⊙X~+(1−M)⊙Xˉ\hat{X} = M \odot \tilde{X} + (1 - M) \odot \bar{X}

    1. Discriminator (DD): Unlike standard GAN discriminators that predict whether an entire sample is real or fake, the GAIN discriminator D:X×H→[0,1]dD: \mathcal{X} \times \mathcal{H} \to [0, 1]^d assesses component-wise authenticity. Given the completed vector X^\hat{X} and a hint vector H∈HH \in \mathcal{H}, the ii-th output D(X^,H)iD(\hat{X}, H)_i predicts the probability that the ii-th component of X^\hat{X} was genuinely observed (Mi=1M_i = 1) rather than imputed (Mi=0M_i = 0).
  2. Knowl 2 — GAIN Hint Mechanism

    model/method

    The hint mechanism in GAIN provides the discriminator with partial information about the original mask vector M∈{0,1}dM \in \{0, 1\}^d, ensuring that the adversarial training process forces the generator to replicate the true underlying data distribution rather than arbitrary candidate distributions.

    A random variable B=(B1,…,Bd)∈{0,1}dB = (B_1, \dots, B_d) \in \{0, 1\}^d is generated by sampling an index k∈{1,…,d}k \in \{1, \dots, d\} uniformly at random and defining:

    Bj={1,if j≠k0,if j=kB_j = \begin{cases} 1, & \text{if } j \neq k \\ 0, & \text{if } j = k \end{cases}

    The hint vector H∈H={0,0.5,1}dH \in \mathcal{H} = \{0, 0.5, 1\}^d is then defined conditioned on MM and BB as:

    H=B⊙M+0.5(1−B)H = B \odot M + 0.5(1 - B)

    where ⊙\odot denotes element-wise multiplication. For each feature index j≠kj \neq k, Hj=MjH_j = M_j, explicitly revealing whether feature jj was observed or imputed. For index kk, Hk=0.5H_k = 0.5, providing no direct disclosure of MkM_k. For the optimal discriminator D∗D^*, D∗(x,h)i=hiD^*(x, h)_i = h_i when hi∈{0,1}h_i \in \{0, 1\}, so adversarial gradients only evaluate the ambiguous component where hi=0.5h_i = 0.5 (bi=0b_i = 0).

  3. Knowl 3 — Theoretical Convergence to the True Conditional Distribution under MCAR

    theoretical result

    Let data vector X∈XX \in \mathcal{X} and mask vector M∈{0,1}dM \in \{0, 1\}^d satisfy the Missing Completely at Random (MCAR) assumption, such that MM is statistically independent of XX. Let X^=M⊙X~+(1−M)⊙G(X~,M,(1−M)⊙Z)\hat{X} = M \odot \tilde{X} + (1 - M) \odot G(\tilde{X}, M, (1 - M) \odot Z), let hint vector H=B⊙M+0.5(1−B)H = B \odot M + 0.5(1 - B) with Bj=1B_j = 1 for j≠kj \neq k and Bk=0B_k = 0 for uniformly sampled k∈{1,…,d}k \in \{1, \dots, d\}, and let p^\hat{p} denote the joint density of (X^,M,H)(\hat{X}, M, H).

    1. Optimal Discriminator: For any fixed generator GG, the optimal discriminator D∗D^* satisfies:

    D∗(x,h)i=pm(Mi=1∣X=x,H=h)D^*(x, h)_i = p_m(M_i = 1 \mid X = x, H = h)

    1. Global Minimum of the Generator Criterion: Substituting D∗D^* into the minimax objective yields the generator criterion C(G)=EX^,M,H[∑i:Mi=1log⁡pm(Mi=1∣X^,H)+∑i:Mi=0log⁡pm(Mi=0∣X^,H)]C(G) = \mathbb{E}_{\hat{X}, M, H} \left[ \sum_{i: M_i=1} \log p_m(M_i = 1 \mid \hat{X}, H) + \sum_{i: M_i=0} \log p_m(M_i = 0 \mid \hat{X}, H) \right]. A global minimum of C(G)C(G) is achieved if and only if:

    p^(x∣h,Mi=t)=p^(x∣h)\hat{p}(x \mid h, M_i = t) = \hat{p}(x \mid h)

    for all i∈{1,…,d}i \in \{1, \dots, d\}, x∈Xx \in \mathcal{X}, t∈{0,1}t \in \{0, 1\}, and h∈Hh \in \mathcal{H} with ph(h∣Mi=t)>0p_h(h \mid M_i = t) > 0.

    1. Distribution Recovery: Under the specified hint mechanism H=B⊙M+0.5(1−B)H = B \odot M + 0.5(1 - B), the solution to the global minimum condition is unique and satisfies:

    p^(x∣m1)=p^(x∣m2)=p^(x∣1)=P(X)\hat{p}(x \mid m_1) = \hat{p}(x \mid m_2) = \hat{p}(x \mid \mathbf{1}) = P(X)

    for all m1,m2∈{0,1}dm_1, m_2 \in \{0, 1\}^d. Consequently, the distribution of completed samples X^\hat{X} coincides exactly with the true underlying data distribution P(X)P(X). If HH is independent of MM, the solution is non-unique and distribution recovery is not guaranteed.

  4. Knowl 4 — GAIN Loss Functions and Objectives

    equation

    The discriminator and generator in GAIN are trained using modified cross-entropy losses conditioned on the hint-mask B∈{0,1}dB \in \{0, 1\}^d, complemented by a reconstruction loss on observed components for the generator.

    Let m∈{0,1}dm \in \{0, 1\}^d be the true mask, m^=D(x^,h)∈[0,1]d\hat{m} = D(\hat{x}, h) \in [0, 1]^d be the discriminator's output, and b∈{0,1}db \in \{0, 1\}^d indicate revealed hint entries. The discriminator loss LD:{0,1}d×[0,1]d×{0,1}d→R\mathcal{L}_D: \{0, 1\}^d \times [0, 1]^d \times \{0, 1\}^d \to \mathbb{R} evaluates only the unrevealed component (bi=0b_i = 0):

    LD(m,m^,b)=∑i:bi=0[milog⁡(m^i)+(1−mi)log⁡(1−m^i)]\mathcal{L}_D(m, \hat{m}, b) = \sum_{i: b_i = 0} \left[ m_i \log(\hat{m}_i) + (1 - m_i) \log(1 - \hat{m}_i) \right]

    The generator's adversarial loss LG:{0,1}d×[0,1]d×{0,1}d→R\mathcal{L}_G: \{0, 1\}^d \times [0, 1]^d \times \{0, 1\}^d \to \mathbb{R} encourages the generator to fool the discriminator on imputed entries (mi=0m_i = 0) where bi=0b_i = 0:

    LG(m,m^,b)=−∑i:bi=0(1−mi)log⁡(m^i)\mathcal{L}_G(m, \hat{m}, b) = -\sum_{i: b_i = 0} (1 - m_i) \log(\hat{m}_i)

    The reconstruction loss LM:Rd×Rd→R\mathcal{L}_M: \mathbb{R}^d \times \mathbb{R}^d \to \mathbb{R} measures the discrepancy between the observed features xx and generator outputs xˉ\bar{x} on observed entries (mi=1m_i = 1):

    LM(x,xˉ)=∑i=1dmiLM(xi,xˉi),where LM(xi,xˉi)={(xˉi−xi)2,if xi is continuous−xilog⁡(xˉi),if xi is binary\mathcal{L}_M(x, \bar{x}) = \sum_{i=1}^d m_i L_M(x_i, \bar{x}_i), \quad \text{where } L_M(x_i, \bar{x}_i) = \begin{cases} (\bar{x}_i - x_i)^2, & \text{if } x_i \text{ is continuous} \\ -x_i \log(\bar{x}_i), & \text{if } x_i \text{ is binary} \end{cases}

    The generator is trained to minimize the combined objective with hyperparameter α>0\alpha > 0:

    min⁡GLG(m,D(x^,h),b)+αLM(x~,xˉ)\min_G \mathcal{L}_G(m, D(\hat{x}, h), b) + \alpha \mathcal{L}_M(\tilde{x}, \bar{x})

  5. Knowl 5 — Generative Adversarial Imputation Net Training Algorithm

    algorithm

    GAIN alternates between updating the discriminator DD and updating the generator GG using stochastic gradient descent on mini-batches of incomplete data samples.

    Input: Dataset D = {(x_tilde^(j), m^(j))}_{j=1}^N, batch sizes k_D, k_G, hyperparameter alpha
    Output: Trained generator G and discriminator D
    while training loss has not converged do
        // (1) Discriminator optimization
        Draw k_D samples {(x_tilde^(j), m^(j))}_{j=1}^(k_D) from D
        Draw k_D independent noise vectors {z^(j)}_{j=1}^(k_D) from Z ~ [0, 1]^d
        Draw k_D random hint masks {b^(j)}_{j=1}^(k_D) from B in {0, 1}^d
        for j = 1 to k_D do
            x_bar^(j) = G(x_tilde^(j), m^(j), (1 - m^(j)) * z^(j))
            x_hat^(j) = m^(j) * x_tilde^(j) + (1 - m^(j)) * x_bar^(j)
            h^(j) = b^(j) * m^(j) + 0.5 * (1 - b^(j))
        Update D using stochastic gradient descent on:
            grad_D -sum_{j=1}^(k_D) L_D(m^(j), D(x_hat^(j), h^(j)), b^(j))
        // (2) Generator optimization
        Draw k_G samples {(x_tilde^(j), m^(j))}_{j=1}^(k_G) from D
        Draw k_G independent noise vectors {z^(j)}_{j=1}^(k_G) from Z ~ [0, 1]^d
        Draw k_G random hint masks {b^(j)}_{j=1}^(k_G) from B in {0, 1}^d
        for j = 1 to k_G do
            x_bar^(j) = G(x_tilde^(j), m^(j), (1 - m^(j)) * z^(j))
            x_hat^(j) = m^(j) * x_tilde^(j) + (1 - m^(j)) * x_bar^(j)
            h^(j) = b^(j) * m^(j) + 0.5 * (1 - b^(j))
        Update G using stochastic gradient descent on:
            grad_G sum_{j=1}^(k_G) [ L_G(m^(j), D(x_hat^(j), h^(j)), b^(j)) + alpha * L_M(x_tilde^(j), x_bar^(j)) ]
    end while
  6. Knowl 6 — Imputation Error Across Benchmark Datasets

    data/table

    The imputation performance of GAIN was evaluated against five baseline methods (MICE, MissForest, Matrix completion, Auto-encoder, and Expectation-Maximization) across five UCI Machine Learning Repository datasets (Breast, Spam, Letter, Credit, and News) under a 20% Missing Completely at Random (MCAR) protocol. Experiments were conducted with 10 repetitions of 5-fold cross-validation, measuring Root Mean Squared Error (RMSE).

    Algorithm Breast Spam Letter Credit News
    GAIN .0546 ±\pm .0006 .0513 ±\pm .0016 .1198 ±\pm .0005 .1858 ±\pm .0010 .1441 ±\pm .0007
    MICE .0646 ±\pm .0028 .0699 ±\pm .0010 .1537 ±\pm .0006 .2585 ±\pm .0011 .1763 ±\pm .0007
    MissForest .0608 ±\pm .0013 .0553 ±\pm .0013 .1605 ±\pm .0004 .1976 ±\pm .0015 .1623 ±\pm .0120
    Matrix .0946 ±\pm .0020 .0542 ±\pm .0006 .1442 ±\pm .0006 .2602 ±\pm .0073 .2282 ±\pm .0005
    Auto-encoder .0697 ±\pm .0018 .0670 ±\pm .0030 .1351 ±\pm .0009 .2388 ±\pm .0005 .1667 ±\pm .0014
    EM .0634 ±\pm .0021 .0712 ±\pm .0012 .1563 ±\pm .0012 .2604 ±\pm .0015 .1912 ±\pm .0011

    GAIN achieved the lowest RMSE across all five benchmark datasets, outperforming both discriminative (MICE, MissForest, Matrix completion) and generative (Auto-encoder, EM) benchmarks.

  7. Knowl 7 — Ablation Study of GAIN Architectural Components

    data/table

    An ablation study evaluated the relative contribution of each architectural component in GAIN: the adversarial generator loss LG\mathcal{L}_G, the reconstruction loss on observed features LM\mathcal{L}_M, and the discriminator hint mechanism HH. Performance is measured in terms of mean ±\pm standard deviation of RMSE and percentage performance gain relative to ablated versions under 20% MCAR missingness across five UCI datasets.

    Algorithm Breast Spam Letter Credit News
    GAIN .0546 ±\pm .0006 .0513 ±\pm .0016 .1198 ±\pm .0005 .1858 ±\pm .0010 .1441 ±\pm .0007
    GAIN w/o LG\mathcal{L}_G .0701 ±\pm .0021 .0676 ±\pm .0029 .1344 ±\pm .0012 .2436 ±\pm .0012 .1612 ±\pm .0024
    (22.1%) (24.1%) (10.9%) (23.7%) (10.6%)
    GAIN w/o LM\mathcal{L}_M .0767 ±\pm .0015 .0672 ±\pm .0036 .1586 ±\pm .0024 .2533 ±\pm .0048 .2522 ±\pm .0042
    (28.9%) (23.7%) (24.4%) (26.7%) (42.9%)
    GAIN w/o Hint .0639 ±\pm .0018 .0582 ±\pm .0008 .1249 ±\pm .0011 .2173 ±\pm .0052 .1521 ±\pm .0008
    (14.6%) (11.9%) (4.1%) (14.5%) (5.3%)
    GAIN w/o Hint LM\mathcal{L}_M .0782 ±\pm .0016 .0700 ±\pm .0064 .1671 ±\pm .0052 .2789 ±\pm .0071 .2527 ±\pm .0052
    (30.1%) (26.7%) (28.3%) (33.4%) (43.0%)

    The full GAIN framework improves RMSE by approximately 15% over the standard auto-encoder baseline (GAIN w/o LG\mathcal{L}_G), and incorporating the hint mechanism provides an additional improvement of approximately 10% over the unhinted model.

  8. Knowl 8 — Post-Imputation Downstream Classification Performance

    data/table

    The utility of imputed datasets was evaluated by training a standard downstream classifier (logistic regression) on data completed by each imputation method under 20% MCAR missingness. Performance was assessed using the Area Under the Receiver Operating Characteristic Curve (AUROC) across four binary classification UCI datasets over 10 independent experiments with 5-fold cross-validation.

    Algorithm Breast Spam Credit News
    GAIN .9930 ±\pm .0073 .9529 ±\pm .0023 .7527 ±\pm .0031 .9711 ±\pm .0027
    MICE .9914 ±\pm .0034 .9495 ±\pm .0031 .7427 ±\pm .0026 .9451 ±\pm .0037
    MissForest .9860 ±\pm .0112 .9520 ±\pm .0061 .7498 ±\pm .0047 .9597 ±\pm .0043
    Matrix .9897 ±\pm .0042 .8639 ±\pm .0055 .7059 ±\pm .0150 .8578 ±\pm .0125
    Auto-encoder .9916 ±\pm .0059 .9403 ±\pm .0051 .7485 ±\pm .0031 .9321 ±\pm .0058
    EM .9899 ±\pm .0147 .9217 ±\pm .0093 .7390 ±\pm .0079 .8987 ±\pm .0157

    GAIN consistently achieves the highest post-imputation AUROC across all tested datasets. While downstream prediction differences are modest when 80% of data is observed, the superior imputation quality provides a consistent positive margin over all baseline methods.

  9. Knowl 9 — Congeniality of Imputation Models

    data/table

    Congeniality evaluates how faithfully an imputation model preserves the true underlying feature-label relationship. It is measured by calculating the discrepancy between logistic regression parameters ww fit on the complete Credit dataset and parameters w^\hat{w} fit on the Credit dataset after introducing missing values and imputing them. Discrepancies are measured by mean bias (∥w−w^∥1)(\|w - \hat{w}\|_1) and mean square error (∥w−w^∥2)(\|w - \hat{w}\|_2).

    Algorithm Mean Bias (∥w−w^∥1)(\|w - \hat{w}\|_1) MSE (∥w−w^∥2)(\|w - \hat{w}\|_2)
    GAIN 0.3163 ±\pm 0.0887 0.5078 ±\pm 0.1137
    MICE 0.8315 ±\pm 0.2293 0.9467 ±\pm 0.2083
    MissForest 0.6730 ±\pm 0.1937 0.7081 ±\pm 0.1625
    Matrix completion 1.5321 ±\pm 0.0017 1.6660 ±\pm 0.0015
    Auto-encoder 0.3500 ±\pm 0.1503 0.5608 ±\pm 0.1697
    EM 0.8418 ±\pm 0.2675 0.9369 ±\pm 0.2296

    GAIN achieves lower mean bias and mean square error in estimated downstream model parameters than all benchmark imputation algorithms, demonstrating performance improvements between 8.9% and 79.2% in preserving the true feature-label relationships.

  10. Knowl 10 — Empirical Robustness to Missing Rates, Sample Size, and Dimensionality

    empirical result

    Empirical sensitivity evaluations on the Credit dataset demonstrate the robustness of GAIN across varying data characteristics compared to its closest competitors, MissForest and Auto-encoder:

    1. Missing Rate Scaling: As the proportion of missing data increases from 10% to 80% (and up to 90% for post-imputation AUROC), GAIN maintains lower RMSE and higher AUROC across the entire range. The AUROC performance advantage of GAIN over MissForest and Auto-encoder widens significantly at higher missing rates (e.g., 60%--90%), because high missingness makes downstream predictions increasingly reliant on the generative quality of the imputed features.

    2. Sample Size Scaling: As sample size increases from 0.5×1040.5 \times 10^4 to 4.0×1044.0 \times 10^4, the RMSE of GAIN decreases rapidly, and the performance margin over MissForest and Auto-encoder widens due to GAIN's capacity to optimize its neural network parameters with larger training sample sizes.

    3. Feature Dimensionality: GAIN maintains stable imputation RMSE across feature dimensions ranging from 5 to 25. In contrast, discriminative models such as MissForest suffer substantial degradation in RMSE when the number of available feature dimensions is small.

Coverage note — Omitted qualitative image error concealment figures on MNIST mentioned in Section 7 as preliminary material relegated to the supplementary materials.

References

  1. 1.Alaa, A. M., Yoon, J., Hu, S., and van der Schaar, M. Personalized risk scoring for critical care prognosis using mixtures of gaussian processes. IEEE Transactions on Biomedical Engineering, 65(1):207–218, 2018.
  2. 2.Allen, A. and Li, W. Generative Adversarial Denoising Autoencoder for Face Completion, 2016. URL https://www.cc.gatech.edu/~hays/7476/projects/Avery_Wenchen/.
  3. 3.Barnard, J. and Meng, X.-L. Applications of multiple imputation in medical studies: from aids to nhanes. Statistical methods in medical research, 8(1):17–36, 1999.
  4. 4.Burgess, S., White, I. R., Resche-Rigon, M., and Wood, A. M. Combining multiple imputation and meta-analysis with individual participant data. Statistics in medicine, 32(26):4499–4514, 2013.
  5. 5.Buuren, S. and Groothuis-Oudshoorn, K. mice: Multivariate imputation by chained equations in r. Journal of statistical software, 45(3), 2011.
  6. 6.Buuren, S. v. and Oudshoorn, C. Multivariate imputation by chained equations: Mice v1. 0 user’s manual. Technical report, TNO, 2000.
  7. 7.Deng, Y., Chang, C., Ido, M. S., and Long, Q. Multiple imputation for general missing data patterns in the presence of high-dimensional data. Scientific reports, 6:21689, 2016.
  8. 8.Garc´ıa-Laencina, P. J., Sancho-Gomez, J.-L., and Figueiras-Vidal, A. R. Pattern classification with missing data: a review. Neural Computing and Applications, 19(2): 263–282, 2010.
  9. 9.Gondara, L. and Wang, K. Multiple imputation using deep denoising autoencoders. arXiv preprint arXiv:1705.02737, 2017.
  10. 10.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In Advances in Neural information processing systems, pp. 2672–2680, 2014.
  11. 11.Kreindler, D. M. and Lumsden, C. J. The effects of the irregular sample and missing data in time series analysis. Nonlinear Dynamical Systems Analysis for the Behavioral Sciences Using Real Data, pp. 135, 2012.
  12. 12.LeCun, Y. and Cortes, C. MNIST handwritten digit database. 2010. URL http://yann.lecun.com/exdb/mnist/.
  13. 13.Lichman, M. UCI machine learning repository, 2013. URL http://archive.ics.uci.edu/ml.
  14. 14.Mackinnon, A. The use and reporting of multiple imputation in medical research–a review. Journal of internal medicine, 268(6):586–593, 2010.
  15. 15.Mazumder, R., Hastie, T., and Tibshirani, R. Spectral regularization algorithms for learning large incomplete matrices. Journal of machine learning research, 11(Aug): 2287–2322, 2010a.
  16. 16.Mazumder, R., Hastie, T., and Tibshirani, R. Spectral regularization algorithms for learning large incomplete matrices. Journal of machine learning research, 11(Aug): 2287–2322, 2010b.
  17. 17.Meng, X.-L. Multiple-imputation inferences with uncongenial sources of input. Statistical Science, pp. 538–558, 1994.
  18. 18.Purwar, A. and Singh, S. K. Hybrid prediction model with missing value imputation for medical data. Expert Systems with Applications, 42(13):5621–5631, 2015.
  19. 19.Rubin, D. B. Multiple imputation for nonresponse in surveys, volume 81. John Wiley & Sons, 2004.
  20. 20.Schnabel, T., Swaminatan, A., Singh, A., Chandak, N., and Joachims, T. Recommendations as treatments: debiasing learning and evolution. ICML, 2016.
  21. 21.Stekhoven, D. J. and Buhlmann, P. Missforestnon-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112–118, 2011.
  22. 22.Sterne, J. A., White, I. R., Carlin, J. B., Spratt, M., Royston, P., Kenward, M. G., Wood, A. M., and Carpenter, J. R. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ, 338: b2393, 2009.
  23. 23.Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th International conference on Machine learning, pp. 1096–1103. ACM, 2008.
  24. 24.Yoon, J., Davtyan, C., and van der Schaar, M. Discovery and clinical decision support for personalized healthcare. IEEE journal of biomedical and health informatics, 21 (4):1133–1145, 2017.
  25. 25.Yoon, J., Jordon, J., and van der Schaar, M. GANITE: Estimation of individualized treatment effects using generative adversarial nets. In International Conference on Learning Representations, 2018a. URL https://openreview.net/forum?id=ByKWUeWA-.
  26. 26.Yoon, J., Zame, W. R., Banerjee, A., Cadeiras, M., Alaa, A. M., and van der Schaar, M. Personalized survival predictions via trees of predictors: An application to cardiac transplantation. PloS one, 13(3):e0194985, 2018b.
  27. 27.Yoon, J., Zame, W. R., and van der Schaar, M. Deep sensing: Active sensing using multi-directional recurrent neural networks. In International Conference on Learning Representations, 2018c. URL https://openreview.net/forum?id=r1SnX5xCb.
  28. 28.Yu, H.-F., Rao, H., and Dhillon, I. S. Temporal regularized matrix factorization for high-dimensional time series prediction. NIPS, 2016.
  29. 29.Yu, S., Krishnapuram, B., Rosales, R., and Rao, R. B. Active sensing. In Artificial Intelligence and Statistics, pp. 639–646, 2009.

Citation

MLA
Yoon, J., et al. “GAIN: Missing Data Imputation Using Generative Adversarial Nets”. arXiv, 2018, http://arxiv.org/abs/1806.02920v1.
APA
Yoon, J., Jordon, J., & Schaar, M. van . der . (2018). GAIN: Missing Data Imputation using Generative Adversarial Nets. arXiv. http://arxiv.org/abs/1806.02920v1
Chicago
Yoon, J., J. Jordon, and M. van . der . Schaar. 2018. “GAIN: Missing Data Imputation Using Generative Adversarial Nets”. arXiv. http://arxiv.org/abs/1806.02920v1.
Harvard
Yoon, J., Jordon, J. and Schaar, M. van . der . (2018) “GAIN: Missing Data Imputation using Generative Adversarial Nets”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1806.02920v1.
Vancouver
1. Yoon J, Jordon J, Schaar M van der (2018) GAIN: Missing Data Imputation using Generative Adversarial Nets. arXiv

BibTeX

@article{yoon2018gain,
  title = {GAIN: Missing Data Imputation using Generative Adversarial Nets},
  author = {Yoon, Jinsung and Jordon, James and Schaar, Mihaela van der},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1806.02920v1},
  eprint = {1806.02920}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/