CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis
Chaejeong LeeJayoung KimNoseong Park
Proposes CoDi, a tabular synthesis framework that splits continuous and discrete variables into two mutually conditioned diffusion models paired with contrastive learning to generate realistic mixed-type tables.
Generating high-quality synthetic tabular data is increasingly vital for real-world enterprise applications, data sharing, and machine learning workflows where data access or privacy is constrained. However, tabular data presents a unique challenge because it routinely mixes continuous numerical measurements and discrete categorical variables. Conventional generative approaches often map categorical variables into continuous spaces, causing sampling errors and distorting critical correlations between variable types. The article addresses this fundamental bottleneck by proposing CoDi, a generative architecture that processes continuous and discrete variables in their native mathematical spaces.
The main objective of the article is to evaluate and demonstrate whether decoupling mixed-type variables into two distinct, conditionally linked diffusion models—strengthened by contrastive learning—outperforms existing methods across data quality, diversity, and generation runtime. The approach employs two physically separated models: a continuous diffusion model operating in continuous space and a discrete diffusion model operating on categorical distributions. These models are co-evolved by continuously feeding each other's states as conditioning inputs throughout the forward noising and reverse denoising steps. Furthermore, the framework applies a contrastive learning objective with a specialized negative sampling strategy that shuffles discrete and continuous variable blocks across records, training the network to accurately preserve cross-type relationships. The approach was systematically evaluated against eight baseline generative models across 11 diverse real-world datasets spanning binary classification, multi-class classification, and regression tasks.
The experimental findings indicate substantial improvements across key performance metrics. First, CoDi achieved superior sampling quality; downstream machine learning models trained on CoDi's synthetic data scored higher average classification metrics (0.4726 binary F1 and 0.6221 multi-class macro F1) and were the only generative method to achieve a practical, positive predictive regression fit (average R-squared of 0.4794, compared to negative values for all eight baselines). Second, CoDi demonstrated higher sample diversity, attaining an average distribution coverage score of 0.6931, noticeably outperforming leading methods such as STaSy (0.5771) and TableGAN (0.5759). Third, the architecture achieved these gains with exceptional computational efficiency, generating samples nearly nine times faster than previous diffusion-based models (an average runtime of 0.5187 seconds versus 4.6417 seconds for STaSy to produce 10,000 records).
These results demonstrate that treating continuous and categorical variables in their proper mathematical spaces eliminates structural artifacts, such as over-represented categories or blurred decision boundaries. For organizations utilizing synthetic data, this translates directly to more reliable analytic models, lower downstream operational risk, and reduced computational costs. When selecting synthetic data solutions, decision-makers face trade-offs between rapid but low-fidelity generators and accurate but slow models. CoDi largely resolves this compromise by providing production-grade data fidelity and diversity at runtimes competitive with simpler models.
Organizations evaluating synthetic data pipelines should consider adopting dual-model architectures that process data types natively rather than forcing categorical variables into continuous spaces. Before deployment in production, teams should run pilot evaluations on their proprietary data schema, as data distributions vary widely across business domains. Although confidence in the experimental results is high across the 11 benchmark datasets, the primary limitation noted in the article is that CoDi is specialized for mixed-type tables and is not designed for purely continuous or purely discrete datasets. Additionally, as synthetic data quality approaches parity with genuine records, stakeholders must implement appropriate privacy protections to safeguard against potential downstream misuse.
- Paper: TabDDPM: Modelling Tabular Data with Diffusion Models, Akim Kotelnikov et al. (2023). Read TabDDPM first to understand the tabular diffusion baseline CoDi builds on and the mixed numerical–categorical generation problem it seeks to improve.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Ho et al.’s DDPM paper establishes the Gaussian forward-noising and learned reverse-denoising framework underlying CoDi’s continuous-variable diffusion model.
- Paper: Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions, Emiel Hoogeboom et al. (2021). Its multinomial diffusion formulation provides the discrete-variable denoising background needed to follow CoDi’s categorical diffusion model.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). D3PM explains how structured discrete-state diffusion can model categorical data, preparing readers for CoDi’s treatment of discrete table columns.
- Paper: Modeling Tabular data using Conditional GAN, Lei Xu et al. (2019). CTGAN introduces the mixed-type tabular synthesis challenges and benchmark tradition that frame CoDi’s comparisons with earlier generators.
No sufficiently relevant recommendations were found.
