TabDDPM: Modelling Tabular Data with Diffusion Models
Akim KotelnikovDmitry BaranchukIvan RubachevArtem Babenko
Proposes TabDDPM, a versatile diffusion-based generative framework that simultaneously handles mixed continuous and categorical features to consistently outperform standard GAN and VAE baselines on tabular synthetic data generation.
Generating realistic synthetic tabular data has become critical for modern enterprises seeking to share data and develop machine learning models under strict privacy regulations such as GDPR. Tabular datasets present unique modeling difficulties because they combine mixed numerical and categorical features and are often limited in sample size. While diffusion models have demonstrated remarkable success in image and text generation, their application to structured tabular domains has remained largely unexplored.
The article introduces and evaluates TabDDPM, a versatile diffusion-based generative model tailored specifically for heterogeneous tabular data. The objective is to demonstrate that TabDDPM can reliably approximate complex tabular distributions, outperform existing deep generative baselines, and provide a strong balance between synthetic data utility and privacy preservation.
To establish credibility across varied conditions, the authors evaluated TabDDPM on 15 real-world benchmark datasets spanning binary classification, multiclass classification, and regression tasks. The framework models continuous numerical features using Gaussian diffusion and categorical variables using multinomial diffusion within a unified multi-layer neural network. TabDDPM was benchmarked against leading generative adversarial networks (CTGAN, CTABGAN, CTABGAN+), a variational autoencoder (TVAE), and an interpolation technique (SMOTE). Evaluation centered on synthetic data utility—measured by training machine learning models on synthetic data and testing them on real data—as well as statistical fidelity and empirical privacy safeguards against data leakage.
The evaluation revealed several key findings:
- TabDDPM achieved superior statistical fidelity, capturing individual feature distributions and pairwise correlations more accurately than existing deep generative models, earning the top average rank for correlation matrix alignment and categorical fidelity.
- In machine learning efficiency evaluations using state-of-the-art gradient boosting (CatBoost), TabDDPM consistently outperformed GAN- and VAE-based alternatives, matching or approaching the performance of models trained on real data.
- The simple interpolation baseline, SMOTE, demonstrated unexpectedly strong utility, performing competitively with TabDDPM and surpassing GAN and VAE baselines.
- TabDDPM demonstrated substantially stronger privacy protection than SMOTE, exhibiting greater distance to real records and showing robust resistance against full black-box membership inference attacks, where SMOTE exhibited vulnerability rates near 1.0 on several datasets.
These findings indicate that diffusion models provide a highly effective paradigm for tabular data synthesis, overcoming the historical stability and fidelity limitations of tabular GANs and VAEs. Furthermore, the results highlight a crucial trade-off: while simple interpolation techniques offer high utility, they carry unacceptable privacy risks by duplicating or interpolating close to original records. TabDDPM resolves this trade-off by generating truly novel synthetic records that maintain high utility without memorizing proprietary or sensitive training points.
For practical implementation, organizations should adopt TabDDPM when high-utility, privacy-conscious synthetic tabular data is required for external sharing or model development. When evaluating synthetic data generators, practitioners should benchmark against strong, tuned gradient-boosted models rather than weak classifiers, which can create misleading impressions of data quality. Before full-scale deployment in high-risk compliance environments, organizations should conduct tailored privacy audits, as distance metrics do not account for feature-specific sensitivities or partial attribute leakage.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This paper establishes the core mathematical foundation of denoising diffusion probabilistic models (DDPM) that TabDDPM adapts and extends to mixed-type tabular datasets.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This work introduces discrete denoising diffusion processes with structured transition matrices, providing the exact theoretical formulation TabDDPM uses to model categorical tabular columns.
- Paper: Modeling Tabular data using Conditional GAN, Lei Xu et al. (2019). This paper presents CTGAN and TVAE alongside standard benchmark protocols and mode-specific normalization strategies that serve as the primary generative baselines evaluated by TabDDPM.
- Paper: Revisiting Deep Learning Models for Tabular Data, Yury Gorishniy et al. (2021). This study rigorously evaluates deep tabular architectures against gradient-boosted decision trees, establishing the empirical context and CatBoost evaluation methodologies adopted in TabDDPM.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This paper introduces non-Markovian deterministic sampling for diffusion models, which enables faster inference and serves as a fundamental acceleration technique for diffusion-based generators.
- Paper: Deep Neural Networks and Tabular Data: A Survey, Vadim Borisov et al. (2021). This comprehensive survey categorizes the challenges of heterogeneous tabular data and deep learning architectures, contextualizing why standard deep generative models struggle on structured tables.
- Paper: Accurate predictions on small data with a tabular foundation model, Noah Hollmann et al. (2025). This book introduces TabPFN, advancing tabular representation learning and synthetic data generation beyond diffusion models via in-context learning with tabular foundation models.
- Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). This work formulates stochastic interpolants to unify continuous diffusion and flow matching in finite time, providing a mathematical generalization relevant to the continuous diffusion mechanisms in TabDDPM.
- Paper: Generalization in diffusion models arises from geometry-adaptive harmonic representations, Zahra Kadkhodaie et al. (2024). This paper analyzes how neural denoisers achieve generalization versus memorization in diffusion models, offering foundational theoretical insights directly applicable to the privacy and replication trade-offs studied in TabDDPM.
- Paper: Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution, Aaron Lou et al. (2024). This work develops score entropy discrete diffusion for categorical data, presenting an alternative discrete diffusion framework to the multinomial diffusion used for categorical features in TabDDPM.
