TabDDPM: Modelling Tabular Data with Diffusion Models

Akim KotelnikovDmitry BaranchukIvan RubachevArtem Babenko

article2023ICML706 citations

Proposes TabDDPM, a versatile diffusion-based generative framework that simultaneously handles mixed continuous and categorical features to consistently outperform standard GAN and VAE baselines on tabular synthetic data generation.

Listen

Generating realistic synthetic tabular data has become critical for modern enterprises seeking to share data and develop machine learning models under strict privacy regulations such as GDPR. Tabular datasets present unique modeling difficulties because they combine mixed numerical and categorical features and are often limited in sample size. While diffusion models have demonstrated remarkable success in image and text generation, their application to structured tabular domains has remained largely unexplored.

The article introduces and evaluates TabDDPM, a versatile diffusion-based generative model tailored specifically for heterogeneous tabular data. The objective is to demonstrate that TabDDPM can reliably approximate complex tabular distributions, outperform existing deep generative baselines, and provide a strong balance between synthetic data utility and privacy preservation.

To establish credibility across varied conditions, the authors evaluated TabDDPM on 15 real-world benchmark datasets spanning binary classification, multiclass classification, and regression tasks. The framework models continuous numerical features using Gaussian diffusion and categorical variables using multinomial diffusion within a unified multi-layer neural network. TabDDPM was benchmarked against leading generative adversarial networks (CTGAN, CTABGAN, CTABGAN+), a variational autoencoder (TVAE), and an interpolation technique (SMOTE). Evaluation centered on synthetic data utility—measured by training machine learning models on synthetic data and testing them on real data—as well as statistical fidelity and empirical privacy safeguards against data leakage.

The evaluation revealed several key findings:

  1. TabDDPM achieved superior statistical fidelity, capturing individual feature distributions and pairwise correlations more accurately than existing deep generative models, earning the top average rank for correlation matrix alignment and categorical fidelity.
  2. In machine learning efficiency evaluations using state-of-the-art gradient boosting (CatBoost), TabDDPM consistently outperformed GAN- and VAE-based alternatives, matching or approaching the performance of models trained on real data.
  3. The simple interpolation baseline, SMOTE, demonstrated unexpectedly strong utility, performing competitively with TabDDPM and surpassing GAN and VAE baselines.
  4. TabDDPM demonstrated substantially stronger privacy protection than SMOTE, exhibiting greater distance to real records and showing robust resistance against full black-box membership inference attacks, where SMOTE exhibited vulnerability rates near 1.0 on several datasets.

These findings indicate that diffusion models provide a highly effective paradigm for tabular data synthesis, overcoming the historical stability and fidelity limitations of tabular GANs and VAEs. Furthermore, the results highlight a crucial trade-off: while simple interpolation techniques offer high utility, they carry unacceptable privacy risks by duplicating or interpolating close to original records. TabDDPM resolves this trade-off by generating truly novel synthetic records that maintain high utility without memorizing proprietary or sensitive training points.

For practical implementation, organizations should adopt TabDDPM when high-utility, privacy-conscious synthetic tabular data is required for external sharing or model development. When evaluating synthetic data generators, practitioners should benchmark against strong, tuned gradient-boosted models rather than weak classifiers, which can create misleading impressions of data quality. Before full-scale deployment in high-risk compliance environments, organizations should conduct tailored privacy audits, as distance metrics do not account for feature-specific sensitivities or partial attribute leakage.

arXiv: 2209.15421
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This paper establishes the core mathematical foundation of denoising diffusion probabilistic models (DDPM) that TabDDPM adapts and extends to mixed-type tabular datasets.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This work introduces discrete denoising diffusion processes with structured transition matrices, providing the exact theoretical formulation TabDDPM uses to model categorical tabular columns.
  • Paper: Modeling Tabular data using Conditional GAN, Lei Xu et al. (2019). This paper presents CTGAN and TVAE alongside standard benchmark protocols and mode-specific normalization strategies that serve as the primary generative baselines evaluated by TabDDPM.
  • Paper: Revisiting Deep Learning Models for Tabular Data, Yury Gorishniy et al. (2021). This study rigorously evaluates deep tabular architectures against gradient-boosted decision trees, establishing the empirical context and CatBoost evaluation methodologies adopted in TabDDPM.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This paper introduces non-Markovian deterministic sampling for diffusion models, which enables faster inference and serves as a fundamental acceleration technique for diffusion-based generators.
  • Paper: Deep Neural Networks and Tabular Data: A Survey, Vadim Borisov et al. (2021). This comprehensive survey categorizes the challenges of heterogeneous tabular data and deep learning architectures, contextualizing why standard deep generative models struggle on structured tables.
Cover for TabDDPM: Modelling Tabular Data with Diffusion Models

Abstract

Denoising diffusion probabilistic models are becoming the leading generative modeling paradigm for many important data modalities. Being the most prevalent in the computer vision community, diffusion models have recently gained some attention in other domains, including speech, NLP, and graph-like data. In this work, we investigate if the framework of diffusion models can be advantageous for general tabular problems, where data points are typically represented by vectors of heterogeneous features. The inherent heterogeneity of tabular data makes it quite challenging for accurate modeling since the individual features can be of a completely different nature, i.e., some of them can be continuous and some can be discrete. To address such data types, we introduce TabDDPM — a diffusion model that can be universally applied to any tabular dataset and handles any feature types. We extensively evaluate TabDDPM on a wide set of benchmarks and demonstrate its superiority over existing GAN/VAE alternatives, which is consistent with the advantage of diffusion models in other fields. The source code of TabDDPM is available at GitHub.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background
  • 4. TabDDPM
  • 5. Experiments
  • 5.1. Qualitative comparison
  • 5.2. Machine Learning efficiency
  • 5.3. Privacy
  • Limitations and discussion
  • 6. Conclusion
  • References
  • Appendix
  • A. MLP evaluation and tuning
  • B. Additional results
  • C. Additional visualizations
  • D. Distance to Closest Record using pretrained MLP features
  • E. Hyperparameters Search Spaces
  • F. Datasets
  • G. Environment and Runtime

Knowls

  1. Knowl 1 — TabDDPM Framework for Mixed Tabular Data

    model/method

    TabDDPM (Tabular Denoising Diffusion Probabilistic Model) is a generative framework designed to model heterogeneous tabular datasets containing both continuous numerical features and discrete categorical or binary features.

    Input Preprocessing and Representation

    A tabular sample x=[xnum,xcat1,…,xcatC]x = [x_{\text{num}}, x_{\text{cat}_1}, \dots, x_{\text{cat}_C}] consists of NnumN_{\text{num}} continuous numerical features xnum∈RNnumx_{\text{num}} \in \mathbb{R}^{N_{\text{num}}} and CC categorical features xcatix_{\text{cat}_i}, each having KiK_i possible categories. Numerical features are normalized via a Gaussian quantile transformation. Categorical features are converted into one-hot vectors xcatiohe∈{0,1}Kix_{\text{cat}_i}^{\text{ohe}} \in \{0, 1\}^{K_i}. The concatenated input vector x0=[xnumnorm,xcat1ohe,…,xcatCohe]x_0 = [x_{\text{num}}^{\text{norm}}, x_{\text{cat}_1}^{\text{ohe}}, \dots, x_{\text{cat}_C}^{\text{ohe}}] has total dimension Nnum+∑i=1CKiN_{\text{num}} + \sum_{i=1}^C K_i.

    Forward and Reverse Diffusion Mechanisms

    TabDDPM combines two diffusion processes over discrete timesteps t∈{1,…,T}t \in \{1, \dots, T\}:

    1. Gaussian Diffusion for Numerical Features: Continuous features undergo a standard forward Gaussian diffusion process: q(xt,num∣xt−1,num)=N(xt,num;1−βtxt−1,num,βtI)q(x_{t,\text{num}} \mid x_{t-1,\text{num}}) = \mathcal{N}\left(x_{t,\text{num}}; \sqrt{1-\beta_t}x_{t-1,\text{num}}, \beta_t I\right) where βt∈(0,1)\beta_t \in (0, 1) follows a cosine noise variance schedule. The neural network predicts the added noise component ϵθ(xt,t)\epsilon_\theta(x_t, t).
    2. Multinomial Diffusion for Categorical Features: Each categorical feature i∈{1,…,C}i \in \{1, \dots, C\} is corrupted independently via a categorical distribution with uniform class corruption: q(xt,cati∣xt−1,cati)=Cat(xt,cati;(1−βt)xt−1,cati+βtKi1)q(x_{t,\text{cat}_i} \mid x_{t-1,\text{cat}_i}) = \text{Cat}\left(x_{t,\text{cat}_i}; (1-\beta_t)x_{t-1,\text{cat}_i} + \frac{\beta_t}{K_i}\mathbf{1}\right) The network directly predicts the uncorrupted categorical one-hot probabilities x^0,cati\hat{x}_{0,\text{cat}_i}.

    Neural Network Architecture

    The reverse process model is a multi-layer perceptron (MLP). Given raw tabular input xinx_{\text{in}}, diffusion step tt, and class label yy (for classification tasks), the input representation is constructed as: temb=Linear(SiLU(Linear(SinTimeEmb(t))))t_{\text{emb}} = \text{Linear}(\text{SiLU}(\text{Linear}(\text{SinTimeEmb}(t)))) yemb=Embedding(y)y_{\text{emb}} = \text{Embedding}(y) x(0)=Linear(xin)+temb+yembx^{(0)} = \text{Linear}(x_{\text{in}}) + t_{\text{emb}} + y_{\text{emb}} where SinTimeEmb(t)\text{SinTimeEmb}(t) is a 128-dimensional sinusoidal time embedding, and all projection linear layers have dimension 128. The hidden layers follow: MLP(x)=Linear(MLPBlock(…(MLPBlock(x(0)))))\text{MLP}(x) = \text{Linear}(\text{MLPBlock}(\dots(\text{MLPBlock}(x^{(0)})))) MLPBlock(h)=Dropout(ReLU(Linear(h)))\text{MLPBlock}(h) = \text{Dropout}(\text{ReLU}(\text{Linear}(h))) For regression problems, the target variable is treated as an additional numerical continuous feature, and the joint distribution p(x,y)p(x, y) is learned unconditionally.

  2. Knowl 2 — TabDDPM Training Objective for Heterogeneous Data

    equation

    TabDDPM is trained by jointly optimizing the Gaussian diffusion loss for numerical features and the variational lower bound KL-divergences for categorical features at each diffusion timestep tt:

    LtTabDDPM=Ltsimple+1C∑i=1CLtiL_t^{\text{TabDDPM}} = L_t^{\text{simple}} + \frac{1}{C} \sum_{i=1}^C L_t^i

    where:

    • Ltsimple=Ex0,ϵ,t∥ϵ−ϵθ(xt,t)∥22L_t^{\text{simple}} = \mathbb{E}_{x_0, \epsilon, t} \|\epsilon - \epsilon_\theta(x_t, t)\|_2^2 is the mean-squared error between the true sampled Gaussian noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) and the model's predicted noise ϵθ(xt,t)\epsilon_\theta(x_t, t) for the continuous numerical features.
    • CC is the total number of categorical features in the tabular dataset.
    • Lti=KL(q(xt−1,cati∣xt,cati,x0,cati)∥pθ(xt−1,cati∣xt,cati))L_t^i = \text{KL}\left(q(x_{t-1,\text{cat}_i} \mid x_{t,\text{cat}_i}, x_{0,\text{cat}_i}) \parallel p_\theta(x_{t-1,\text{cat}_i} \mid x_{t,\text{cat}_i})\right) is the Kullback-Leibler divergence between the true posterior transition distribution and the reverse transition parameterization for the ii-th categorical feature at timestep tt.
    • The reverse posterior for categorical feature ii is parameterized as pθ(xt−1∣xt)=q(xt−1∣xt,x^0(xt,t))p_\theta(x_{t-1} \mid x_t) = q(x_{t-1} \mid x_t, \hat{x}_0(x_t, t)), where x^0(xt,t)\hat{x}_0(x_t, t) represents the logits predicted by the MLP backbone and normalized by a softmax function.
  3. Knowl 3 — Machine Learning Utility of TabDDPM with Tuned GBDT Evaluators

    data/table

    Machine learning (ML) efficiency measures the performance of a supervised model trained on synthetic tabular data and tested on real test data. Evaluated on tuned CatBoost models across 15 diverse real-world benchmarks, TabDDPM consistently outperforms deep generative models (CTGAN, TVAE, CTABGAN, CTABGAN+) and matches or exceeds real-data training utility across most datasets.

    Dataset Metric CTGAN TVAE CTABGAN CTABGAN+ SMOTE TabDDPM Real
    Abalone (AB) R2R^2 0.420±.0040.420_{\pm .004} 0.433±.0080.433_{\pm .008} – 0.467±.0040.467_{\pm .004} 0.549±.0050.549_{\pm .005} 0.550±.0100.550_{\pm .010} 0.556±.0040.556_{\pm .004}
    Adult (AD) F1F_1 0.789±.0010.789_{\pm .001} 0.781±.0020.781_{\pm .002} 0.783±.0020.783_{\pm .002} 0.772±.0030.772_{\pm .003} 0.791±.0020.791_{\pm .002} 0.795±.0010.795_{\pm .001} 0.815±.0020.815_{\pm .002}
    Buddy (BU) F1F_1 0.867±.0030.867_{\pm .003} 0.864±.0050.864_{\pm .005} 0.855±.0050.855_{\pm .005} 0.884±.0050.884_{\pm .005} 0.891±.0030.891_{\pm .003} 0.906±.0030.906_{\pm .003} 0.906±.0020.906_{\pm .002}
    Calif. Housing (CA) R2R^2 0.686±.0030.686_{\pm .003} 0.752±.0010.752_{\pm .001} – 0.525±.0040.525_{\pm .004} 0.840±.0010.840_{\pm .001} 0.836±.0020.836_{\pm .002} 0.857±.0010.857_{\pm .001}
    Cardio (CAR) F1F_1 0.730±.0010.730_{\pm .001} 0.717±.0010.717_{\pm .001} 0.717±.0010.717_{\pm .001} 0.733±.0010.733_{\pm .001} 0.732±.0010.732_{\pm .001} 0.737±.0010.737_{\pm .001} 0.738±.0010.738_{\pm .001}
    Churn (CH) F1F_1 0.723±.0060.723_{\pm .006} 0.732±.0060.732_{\pm .006} 0.688±.0060.688_{\pm .006} 0.702±.0120.702_{\pm .012} 0.743±.0050.743_{\pm .005} 0.755±.0060.755_{\pm .006} 0.740±.0090.740_{\pm .009}
    Default (DE) F1F_1 0.699±.0020.699_{\pm .002} 0.656±.0070.656_{\pm .007} 0.644±.0110.644_{\pm .011} 0.686±.0040.686_{\pm .004} 0.693±.0030.693_{\pm .003} 0.691±.0040.691_{\pm .004} 0.688±.0030.688_{\pm .003}
    Diabetes (DI) F1F_1 0.459±.0960.459_{\pm .096} 0.714±.0390.714_{\pm .039} 0.731±.0220.731_{\pm .022} 0.734±.0200.734_{\pm .020} 0.683±.0370.683_{\pm .037} 0.740±.0200.740_{\pm .020} 0.785±.0130.785_{\pm .013}
    Facebook (FB) R2R^2 0.443±.0050.443_{\pm .005} 0.685±.0030.685_{\pm .003} – 0.509±.0110.509_{\pm .011} 0.803±.0020.803_{\pm .002} 0.713±.0020.713_{\pm .002} 0.837±.0010.837_{\pm .001}
    Gesture Phase (GE) F1F_1 0.333±.0130.333_{\pm .013} 0.434±.0060.434_{\pm .006} 0.392±.0060.392_{\pm .006} 0.406±.0090.406_{\pm .009} 0.658±.0070.658_{\pm .007} 0.597±.0060.597_{\pm .006} 0.636±.0070.636_{\pm .007}
    Higgs Small (HI) F1F_1 0.575±.0060.575_{\pm .006} 0.638±.0030.638_{\pm .003} 0.575±.0040.575_{\pm .004} 0.664±.0020.664_{\pm .002} 0.722±.0010.722_{\pm .001} 0.722±.0010.722_{\pm .001} 0.724±.0010.724_{\pm .001}
    House 16H (HO) R2R^2 0.433±.0050.433_{\pm .005} 0.493±.0060.493_{\pm .006} – 0.504±.0050.504_{\pm .005} 0.662±.0040.662_{\pm .004} 0.677±.0100.677_{\pm .010} 0.662±.0030.662_{\pm .003}
    Insurance (IN) R2R^2 0.745±.0090.745_{\pm .009} 0.784±.0100.784_{\pm .010} – 0.797±.0050.797_{\pm .005} 0.812±.0020.812_{\pm .002} 0.809±.0020.809_{\pm .002} 0.814±.0010.814_{\pm .001}
    King (KI) R2R^2 0.772±.0050.772_{\pm .005} 0.824±.0030.824_{\pm .003} – 0.444±.0140.444_{\pm .014} 0.842±.0040.842_{\pm .004} 0.833±.0140.833_{\pm .014} 0.907±.0020.907_{\pm .002}
    MiniBooNE (MI) F1F_1 0.783±.0050.783_{\pm .005} 0.912±.0010.912_{\pm .001} 0.889±.0020.889_{\pm .002} 0.892±.0020.892_{\pm .002} 0.932±.0010.932_{\pm .001} 0.936±.0010.936_{\pm .001} 0.934±.0000.934_{\pm .000}
    Wilt (WI) F1F_1 0.749±.0150.749_{\pm .015} 0.501±.0120.501_{\pm .012} 0.906±.0190.906_{\pm .019} 0.798±.0210.798_{\pm .021} 0.913±.0070.913_{\pm .007} 0.904±.0090.904_{\pm .009} 0.898±.0060.898_{\pm .006}

    Scores are reported as mean ±\pm standard deviation computed over 5 random seeds for synthetic generation, each evaluated across 10 independent training seeds for the CatBoost estimator. CTABGAN cannot handle regression tasks (marked with --).

  4. Knowl 4 — Machine Learning Efficiency under Multi-Model Average Protocol

    data/table

    In the standard literature protocol, synthetic data utility is measured by averaging performance over an ensemble of simpler, default models (Decision Tree with max depth 28, Random Forest with max depth 28, Logistic/Ridge Regression with 500 iterations, and standard MLP with 100 iterations).

    Dataset TVAE CTABGAN CTABGAN+ SMOTE TabDDPM Real
    Abalone (AB, R2R^2) 0.238±.0120.238_{\pm .012} – 0.316±.0240.316_{\pm .024} 0.400±.0090.400_{\pm .009} 0.392±.0090.392_{\pm .009} 0.423±.0090.423_{\pm .009}
    Adult (AD, F1F_1) 0.742±.0010.742_{\pm .001} 0.737±.0070.737_{\pm .007} 0.730±.0070.730_{\pm .007} 0.750±.0040.750_{\pm .004} 0.758±.0050.758_{\pm .005} 0.750±.0060.750_{\pm .006}
    Buddy (BU, F1F_1) 0.779±.0040.779_{\pm .004} 0.786±.0080.786_{\pm .008} 0.837±.0060.837_{\pm .006} 0.842±.0030.842_{\pm .003} 0.851±.0030.851_{\pm .003} 0.845±.0040.845_{\pm .004}
    Calif. Housing (CA, R2R^2) −13.0±1.51-13.0_{\pm 1.51} – −7.59±.645-7.59_{\pm .645} 0.667±.0060.667_{\pm .006} 0.695±.0020.695_{\pm .002} 0.663±.0020.663_{\pm .002}
    Cardio (CAR, F1F_1) 0.693±.0020.693_{\pm .002} 0.684±.0030.684_{\pm .003} 0.708±.0020.708_{\pm .002} 0.693±.0010.693_{\pm .001} 0.696±.0010.696_{\pm .001} 0.683±.0020.683_{\pm .002}
    Churn (CH, F1F_1) 0.684±.0030.684_{\pm .003} 0.636±.0100.636_{\pm .010} 0.650±.0080.650_{\pm .008} 0.690±.0030.690_{\pm .003} 0.693±.0030.693_{\pm .003} 0.679±.0030.679_{\pm .003}
    Default (DE, F1F_1) 0.643±.0030.643_{\pm .003} 0.614±.0070.614_{\pm .007} 0.648±.0080.648_{\pm .008} 0.649±.0030.649_{\pm .003} 0.659±.0030.659_{\pm .003} 0.648±.0030.648_{\pm .003}
    Diabetes (DI, F1F_1) 0.712±.0100.712_{\pm .010} 0.655±.0150.655_{\pm .015} 0.727±.0230.727_{\pm .023} 0.677±.0130.677_{\pm .013} 0.675±.0110.675_{\pm .011} 0.699±.0120.699_{\pm .012}
    Facebook (FB, R2R^2) ≪0\ll 0 – ≪0\ll 0 0.651±.0020.651_{\pm .002} 0.527±.0050.527_{\pm .005} 0.645±.0050.645_{\pm .005}
    Gesture Phase (GE, F1F_1) 0.372±.0060.372_{\pm .006} 0.339±.0090.339_{\pm .009} 0.373±.0090.373_{\pm .009} 0.478±.0050.478_{\pm .005} 0.462±.0050.462_{\pm .005} 0.431±.0050.431_{\pm .005}
    Higgs Small (HI, F1F_1) 0.590±.0040.590_{\pm .004} 0.539±.0060.539_{\pm .006} 0.598±.0040.598_{\pm .004} 0.664±.0030.664_{\pm .003} 0.670±.0020.670_{\pm .002} 0.663±.0020.663_{\pm .002}
    House 16H (HO, R2R^2) 0.174±.0120.174_{\pm .012} – 0.222±.0420.222_{\pm .042} 0.394±.0060.394_{\pm .006} 0.426±.0070.426_{\pm .007} 0.415±.0070.415_{\pm .007}
    Insurance (IN, R2R^2) 0.470±.0240.470_{\pm .024} – 0.669±.0180.669_{\pm .018} 0.709±.0080.709_{\pm .008} 0.734±.0070.734_{\pm .007} 0.708±.0070.708_{\pm .007}
    King (KI, R2R^2) 0.666±.0060.666_{\pm .006} – 0.197±.0510.197_{\pm .051} 0.751±.0050.751_{\pm .005} 0.611±.0130.611_{\pm .013} 0.768±.0130.768_{\pm .013}
    MiniBooNE (MI, F1F_1) 0.880±.0020.880_{\pm .002} 0.856±.0030.856_{\pm .003} 0.867±.0020.867_{\pm .002} 0.860±.0010.860_{\pm .001} 0.850±.0040.850_{\pm .004} 0.850±.0040.850_{\pm .004}
    Wilt (WI, F1F_1) 0.497±.0010.497_{\pm .001} 0.656±.0110.656_{\pm .011} 0.653±.0270.653_{\pm .027} 0.793±.0040.793_{\pm .004} 0.792±.0040.792_{\pm .004} 0.684±.0040.684_{\pm .004}

    Negative R2R^2 scores indicate predictions worse than the optimal constant mean. Under this weak-classifier protocol, synthetic data generated by TabDDPM and SMOTE frequently outperforms real training data (e.g., on AD, BU, CA, CAR, CH, DE, GE, HI, HO, IN, and WI), creating a misleading appearance that synthetic data provides superior training utility compared to real data.

  5. Knowl 5 — Protocol Discrepancy in Tabular Generative Model Evaluation

    empirical result

    Evaluating synthetic tabular data via weak, default classifiers/regressors (Decision Trees, Logistic Regression, basic MLPs) produces two major methodological flaws:

    1. Severe Suboptimality: Weak models achieve absolute performance far below state-of-the-art tabular architectures (e.g., tuned CatBoost or tuned MLPs). For example, on California Housing, weak regressors score R2≈0.66R^2 \approx 0.66 while tuned CatBoost achieves R2=0.857R^2 = 0.857.
    2. False Advantage Over Real Data: Under weak models, classifiers trained on synthetic data (e.g., TabDDPM or SMOTE) frequently achieve higher F1F_1 or R2R^2 than those trained on real data (observed across 11 out of 16 benchmarks). This occurs because synthetic generation can smooth boundaries and regularize under-parameterized or untuned models. When evaluated with properly tuned models (CatBoost or tuned MLPs), real data performance remains equal to or higher than synthetic data in virtually all cases.

    Consequently, benchmarking tabular generative models requires evaluating on tuned state-of-the-art estimators (such as CatBoost) to reflect practical utility accurately.

  6. Knowl 6 — Distributional and Correlation Alignment across Generative Paradigms

    data/table

    To evaluate statistical fidelity, generated synthetic data is compared to ground-truth real data using:

    1. Wasserstein Distance (WD) on continuous numerical features.
    2. Jensen-Shannon (JS) Divergence on categorical distributions.
    3. L2L_2 Distance between Correlation Matrices (using Pearson correlation for numerical-numerical pairs, correlation ratio for categorical-numerical pairs, and Theil's UU statistic for categorical-categorical pairs).
    Model WD (Num.) JS (Cat.) L2L_2 (Corr. matrix)
    CTGAN 3.33 4.77 3.47
    TVAE 4.20 3.92 4.40
    CTABGAN+ 3.87 2.54 3.40
    SMOTE 1.67 2.15 2.00
    TabDDPM 1.93 1.62 1.73

    Values represent average ranks across all 15 benchmark datasets (lower rank is better). TabDDPM achieves the best rank for categorical feature fidelity (JS rank 1.62) and pairwise feature correlation preservation (L2L_2 rank 1.73), while performing comparably to SMOTE on continuous numerical feature distance.

  7. Knowl 7 — Privacy Preservation via Distance to Closest Record

    data/table

    Privacy of synthetic tabular data is quantified using the mean Distance to Closest Record (DCR): DCR(xsyn)=min⁡xreal∈Dtrain∥xsyn−xreal∥2\text{DCR}(x_{\text{syn}}) = \min_{x_{\text{real}} \in \mathcal{D}_{\text{train}}} \|x_{\text{syn}} - x_{\text{real}}\|_2 where mean DCR averages these minimal Euclidean distances across all generated samples. Higher DCR values indicate that the model produces novel samples rather than memorized near-duplicates of individual training records.

    Dataset AB AD BU CA CAR CH DE DI
    TVAE 0.088 0.220 0.226 0.056 0.010 0.241 0.096 0.146
    CTABGAN+ 0.081 0.400 0.242 0.070 0.020 0.235 0.131 0.204
    SMOTE 0.018 0.082 0.080 0.016 0.007 0.099 0.054 0.074
    TabDDPM 0.061 0.295 0.168 0.045 0.016 0.166 0.061 0.308
    Dataset FB GE HI HO IN KI MI WI
    TVAE 1.418 0.171 0.497 0.127 0.102 0.200 0.025 0.020
    CTABGAN+ 0.666 0.169 0.533 0.129 0.124 0.390 10.761 0.027
    SMOTE 0.264 0.041 0.209 0.066 0.050 0.090 0.012 0.009
    TabDDPM 0.785 0.076 0.473 0.096 0.050 0.252 0.574 0.023

    SMOTE yields distance distributions heavily concentrated near zero because it synthesizes data by convex linear interpolation along nearest-neighbor line segments. TabDDPM produces significantly higher DCR values than SMOTE across all datasets (and similarly when measured in the latent space of a pretrained MLP), confirming that TabDDPM creates truly novel samples rather than local interpolations.

  8. Knowl 8 — Robustness against Black-Box Membership Inference Attacks

    data/table

    A full black-box membership inference attack assesses whether a specific real record was present in the generative model's training set. Attack success is evaluated using the ROC-AUC score, where 0.500.50 corresponds to random guessing (optimal privacy defense) and higher values indicate privacy leakage.

    Dataset SMOTE Attack ROC-AUC TabDDPM Attack ROC-AUC
    Abalone (AB) 0.967 0.505
    Adult (AD) 0.619 0.511
    Buddy (BU) 0.710 0.569
    California Housing (CA) 0.986 0.516
    Cardio (CAR) 0.721 0.506
    Churn Modeling (CH) 0.891 0.721
    Default (DE) 0.679 0.497
    Diabetes (DI) 0.610 0.510
    Gesture Phase (GE) 0.864 0.533
    Higgs Small (HI) 0.999 0.527
    House 16H (HO) 0.826 0.546
    Insurance (IN) 0.712 0.868
    King (KI) 0.748 0.517
    MiniBooNE (MI) 0.990 0.500
    Wilt (WI) 0.954 0.516

    SMOTE synthetic samples are vulnerable to membership inference attacks, reaching ROC-AUC values up to 0.9990.999 on Higgs and 0.9900.990 on MiniBooNE. In contrast, TabDDPM maintains ROC-AUC values close to 0.500.50 on 13 of the 15 benchmarks, demonstrating that diffusion models do not leak individual training member identities.

  9. Knowl 9 — SMOTE Adaptation for Full Tabular Synthesis

    model/method

    SMOTE (Synthetic Minority Over-sampling Technique), originally developed for imbalanced class oversampling, can be generalized as a non-deep interpolation baseline for full tabular synthetic dataset generation.

    Procedure

    1. Classification Tasks: Synthetic points are generated by randomly selecting an original data sample xix_i and forming a convex combination with its kk-th nearest neighbor xjx_j belonging to the same class: xsyn=xi+λ(xj−xi)x_{\text{syn}} = x_i + \lambda (x_j - x_i) where λ∼Uniform[0,1]\lambda \sim \text{Uniform}[0, 1] and k∈[5,20]k \in [5, 20] is tuned via hyperparameter search.
    2. Regression Tasks: Continuous targets are partitioned into two balanced classes by the median of the target variable yy. Samples are then interpolated exclusively within the resulting binary partitions.

    Although SMOTE achieves machine learning utility competitive with TabDDPM and superior to GAN/VAE tabular generators, it acts as a local interpolator whose samples concentrate around true data points, severely compromising privacy.

  10. Knowl 10 — Hyperparameter Search Space and Tuning Protocol for TabDDPM

    experimental setup

    TabDDPM performance is sensitive to hyperparameter optimization. Tuning is conducted using the Optuna framework over 50 trials, guided by the ML efficiency score (w.r.t. CatBoost) on a hold-out validation set, averaged over 5 sampling seeds.

    TabDDPM Recommended Search Spaces

    • Learning rate: LogUniform[10−5,3×10−3]\text{LogUniform}[10^{-5}, 3 \times 10^{-3}]
    • Batch size: {256,4096}\{256, 4096\}
    • Diffusion timesteps (TT): {100,1000}\{100, 1000\}
    • Training iterations: {5000,10000,20000}\{5000, 10000, 20000\}
    • Number of MLP layers: {2,4,6,8}\{2, 4, 6, 8\}
    • MLP layer width: {128,256,512,1024}\{128, 256, 512, 1024\}
    • Proportion of synthetic samples relative to real training set size: {0.25,0.5,1.0,2.0,4.0,8.0}\{0.25, 0.5, 1.0, 2.0, 4.0, 8.0\}
    • Dropout: 0.00.0
    • Noise schedule: Cosine schedule
    • Gaussian loss: Mean-squared error (MSE)

    Hyperparameter tuning guided by CatBoost does not induce CatBoost-specific bias; synthetics generated by CatBoost-guided TabDDPM achieve equally optimal utility when evaluated using tuned MLPs.

  11. Knowl 11 — Privacy and Representation Limitations of TabDDPM

    limitation

    While TabDDPM demonstrates empirical privacy advantages over interpolation methods such as SMOTE, several limitations remain:

    1. Lack of Formal Differential Privacy: TabDDPM does not provide mathematical differential privacy guarantees (ϵ,δ\epsilon, \delta). Empirical metrics like Distance to Closest Record (DCR) compute isotropic Euclidean distances across all features and do not detect information leakage when a few specific sensitive features coincide exactly.
    2. Discrete Diffusion Alternatives: TabDDPM implements categorical variables via multinomial diffusion with uniform corruption. Continuous-time discrete diffusions, analog bits, or alternative categorical noise mechanisms remain unaddressed.
    3. Variable Type Subcategories: TabDDPM treats all numerical features as unconstrained real values after quantile transformation, without explicit structural adaptations for positive-only real numbers, heavy-tailed counts, or ordinal features.

Coverage note — None was omitted; all primary methodological components (hybrid diffusion, loss, MLP architecture), baseline adaptations (SMOTE), evaluation protocol insights, statistical fidelity metrics, ML efficiency benchmarks, privacy/attack analyses, hyperparameter search spaces, and stated limitations are fully represented.

References

  1. 1.Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 2623–2631, 2019.
  2. 2.Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R. Structured denoising diffusion models in discrete state-spaces. Advances in Neural Information Processing Systems, 34:17981–17993, 2021.
  3. 3.Baldi, P., Sadowski, P., and Whiteson, D. Searching for exotic particles in high-energy physics with deep learning. Nature Communications, 5, 2014.
  4. 4.Baranchuk, D., Rubachev, I., Voynov, A., Khrulkov, V., and Babenko, A. Label-efficient semantic segmentation with diffusion models. arXiv preprint arXiv:2112.03126, 2021.
  5. 5.Camino, R. D., Hammerschmidt, C. A., et al. Oversampling tabular data with deep generative models: Is it worth the effort? 2020.
  6. 6.Campbell, A., Benton, J., De Bortoli, V., Rainforth, T., Deligiannidis, G., and Doucet, A. A continuous time framework for discrete denoising models. Advances in Neural Information Processing Systems, 35:28266–28279, 2022.
  7. 7.Chawla, N. V., Bowyer, K. W., Hall, L. O., and Kegelmeyer, W. P. Smote: synthetic minority over-sampling technique. Journal of artificial intelligence research, 16:321–357, 2002.
  8. 8.Chen, D., Yu, N., Zhang, Y., and Fritz, M. Gan-leaks: A taxonomy of membership inference attacks against generative models. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, pp. 343–362, 2020a.
  9. 9.Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W. Wavegrad: Estimating gradients for waveform generation. arXiv preprint arXiv:2009.00713, 2020b.
  10. 10.Chen, T., Zhang, R., and Hinton, G. Analog bits: Generating discrete data using diffusion models with self-conditioning. arXiv preprint arXiv:2208.04202, 2022.
  11. 11.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. 2021.
  12. 12.Engelmann, J. and Lessmann, S. Conditional wasserstein gan-based oversampling of tabular data for imbalanced learning. Expert Systems with Applications, 174:114582, 2021.
  13. 13.Fan, J., Liu, T., Li, G., Chen, J., Shen, Y., and Du, X. Relational data synthesis using generative adversarial networks: A design space exploration. arXiv preprint arXiv:2008.12763, 2020.
  14. 14.Gorishniy, Y., Rubachev, I., Khrulkov, V., and Babenko, A. Revisiting deep learning models for tabular data. Advances in Neural Information Processing Systems, 34: 18932–18943, 2021.
  15. 15.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. 2020.
  16. 16.Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., and Welling, M. Argmax flows and multinomial diffusion: Learning categorical distributions. Advances in Neural Information Processing Systems, 34:12454–12465, 2021.
  17. 17.Hoogeboom, E., Satorras, V. G., Vignac, C., and Welling, M. Equivariant diffusion for molecule generation in 3d. In International Conference on Machine Learning, pp. 8867–8887. PMLR, 2022.
  18. 18.Jing, B., Corso, G., Chang, J., Barzilay, R., and Jaakkola, T. Torsional diffusion for molecular conformer generation. arXiv preprint arXiv:2206.01729, 2022.
  19. 19.Jordon, J., Yoon, J., and Van Der Schaar, M. Pate-gan: Generating synthetic data with differential privacy guarantees. In International conference on learning representations, 2018.
  20. 20.Kelley Pace, R. and Barry, R. Sparse spatial autoregressions. Statistics & Probability Letters, 33(3):291–297, 1997.
  21. 21.Kim, J., Jeon, J., Lee, J., Hyeong, J., and Park, N. Oct-gan: Neural ode-based conditional tabular gans. In Proceedings of the Web Conference 2021, pp. 1506–1515, 2021.
  22. 22.Kim, J., Lee, C., Shin, Y., Park, S., Kim, M., Park, N., and Cho, J. Sos: Score-based oversampling for tabular data. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 762–772, 2022.
  23. 23.Kohavi, R. Scaling up the accuracy of naive-bayes classifiers: a decision-tree hybrid. In KDD, 1996.
  24. 24.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:2009.09761, 2020.
  25. 25.Lee, J., Hyeong, J., Jeon, J., Park, N., and Cho, J. Invertible tabular gans: Killing two birds with one stone for tabular data synthesis. Advances in Neural Information Processing Systems, 34:4263–4273, 2021.
  26. 26.Li, H., Yang, Y., Chang, M., Feng, H., Xu, Z., Li, Q., and Chen, Y. Srdiff: Single image super-resolution with diffusion probabilistic models. 2021.
  27. 27.Li, X. L., Thickstun, J., Gulrajani, I., Liang, P., and Hashimoto, T. B. Diffusion-lm improves controllable text generation. arXiv preprint arXiv:2205.14217, 2022.
  28. 28.Madeo, R. C. B., Lima, C. A. M., and Peres, S. M. Gesture unit segmentation using support vector machines: segmenting gestures from rest positions. In Proceedings of the 28th Annual ACM Symposium on Applied Computing, SAC, 2013.
  29. 29.Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S. Sdedit: Image synthesis and editing with stochastic differential equations. 2021.
  30. 30.Naeem, M. F., Oh, S. J., Uh, Y., Choi, Y., and Yoo, J. Reliable fidelity and diversity metrics for generative models. In International Conference on Machine Learning, pp. 7176–7185. PMLR, 2020.
  31. 31.Nazabal, A., Olmos, P. M., Ghahramani, Z., and Valera, I. Handling incomplete heterogeneous data using vaes. Pattern Recognition, 107:107501, 2020.
  32. 32.Nichol, Alex & Dhariwal, P. Improved denoising diffusion probabilistic models. ICML, 2021.
  33. 33.Nock, R. and Guillame-Bert, M. Generative trees: Adversarial and copycat. ICML, 2022.
  34. 34.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  35. 35.Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., and Gulin, A. Catboost: unbiased boosting with categorical features. Advances in neural information processing systems, 31, 2018.
  36. 36.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  37. 37.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. 2021.
  38. 38.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  39. 39.Singh, K., Sandhu, R. K., and Kumar, D. Comment volume prediction using neural networks and decision trees. In IEEE UKSim-AMSS 17th International Conference on Computer Modelling and Simulation, UKSim, 2015.
  40. 40.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015.
  41. 41.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. In NeurIPS, 2019.
  42. 42.Song, Y. and Ermon, S. Improved techniques for training score-based generative models. NeurIPS, 2020.
  43. 43.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. 2021.
  44. 44.Tashiro, Y., Song, J., Song, Y., and Ermon, S. Csdi: Conditional score-based diffusion models for probabilistic time series imputation. Advances in Neural Information Processing Systems, 34:24804–24816, 2021.
  45. 45.Torfi, A., Fox, E. A., and Reddy, C. K. Differentially private synthetic medical data generation using convolutional gans. Information Sciences, 586:485–500, 2022.
  46. 46.Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L. Openml: networked science in machine learning. arXiv, 1407.7722v1, 2014.
  47. 47.Wen, B., Cao, Y., Yang, F., Subbalakshmi, K., and Chandramouli, R. Causal-tgan: Modeling tabular data using causally-aware gan. In ICLR Workshop on Deep Generative Models for Highly Structured Data, 2022.
  48. 48.Xu, L., Skoularidou, M., Cuesta-Infante, A., and Veeramachaneni, K. Modeling tabular data using conditional gan. Advances in Neural Information Processing Systems, 32, 2019.
  49. 49.Zhang, Y., Zaidi, N. A., Zhou, J., and Li, G. Ganblr: a tabular data generation model. In 2021 IEEE International Conference on Data Mining (ICDM), pp. 181–190. IEEE, 2021.
  50. 50.Zhao, Z., Kunar, A., Birke, R., and Chen, L. Y. Ctab-gan: Effective table data synthesizing. In Asian Conference on Machine Learning, pp. 97–112. PMLR, 2021.
  51. 51.Zhao, Z., Kunar, A., Birke, R., and Chen, L. Y. Ctab-gan+: Enhancing tabular data synthesis. arXiv preprint arXiv:2204.00401, 2022.
  52. 52.Zheng, S. and Charoenphakdee, N. Diffusion models for missing value imputation in tabular data. arXiv preprint arXiv:2210.17128, 2022.

Citation

MLA
Kotelnikov, A., et al. “TabDDPM: Modelling Tabular Data with Diffusion Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 17564–79, https://proceedings.mlr.press/v202/kotelnikov23a.html.
APA
Kotelnikov, A., Baranchuk, D., Rubachev, I., & Babenko, A. (2023). TabDDPM: Modelling Tabular Data with Diffusion Models. International Conference on Machine Learning, 202, 17564–17579. https://proceedings.mlr.press/v202/kotelnikov23a.html
Chicago
Kotelnikov, A., D. Baranchuk, I. Rubachev, and A. Babenko. 2023. “TabDDPM: Modelling Tabular Data with Diffusion Models”. International Conference on Machine Learning 202: 17564–79. https://proceedings.mlr.press/v202/kotelnikov23a.html.
Harvard
Kotelnikov, A. et al. (2023) “TabDDPM: Modelling Tabular Data with Diffusion Models”, International Conference on Machine Learning. PMLR, pp. 17564–17579. Available at: https://proceedings.mlr.press/v202/kotelnikov23a.html.
Vancouver
1. Kotelnikov A, Baranchuk D, Rubachev I, Babenko A (2023) TabDDPM: Modelling Tabular Data with Diffusion Models. In: International Conference on Machine Learning. PMLR, pp 17564–17579

BibTeX

@InProceedings{pmlr-v202-kotelnikov23a,
  title = 	 {{T}ab{DDPM}: Modelling Tabular Data with Diffusion Models},
  author =       {Kotelnikov, Akim and Baranchuk, Dmitry and Rubachev, Ivan and Babenko, Artem},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {17564--17579},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/kotelnikov23a/kotelnikov23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/kotelnikov23a.html},
  abstract = 	 {Denoising diffusion probabilistic models are becoming the leading generative modeling paradigm for many important data modalities. Being the most prevalent in the computer vision community, diffusion models have recently gained some attention in other domains, including speech, NLP, and graph-like data. In this work, we investigate if the framework of diffusion models can be advantageous for general tabular problems, where data points are typically represented by vectors of heterogeneous features. The inherent heterogeneity of tabular data makes it quite challenging for accurate modeling since the individual features can be of a completely different nature, i.e., some of them can be continuous and some can be discrete. To address such data types, we introduce TabDDPM — a diffusion model that can be universally applied to any tabular dataset and handles any feature types. We extensively evaluate TabDDPM on a wide set of benchmarks and demonstrate its superiority over existing GAN/VAE alternatives, which is consistent with the advantage of diffusion models in other fields.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/