CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis

Chaejeong LeeJayoung KimNoseong Park

article2023ICML93 citations

Proposes CoDi, a tabular synthesis framework that splits continuous and discrete variables into two mutually conditioned diffusion models paired with contrastive learning to generate realistic mixed-type tables.

Listen

Generating high-quality synthetic tabular data is increasingly vital for real-world enterprise applications, data sharing, and machine learning workflows where data access or privacy is constrained. However, tabular data presents a unique challenge because it routinely mixes continuous numerical measurements and discrete categorical variables. Conventional generative approaches often map categorical variables into continuous spaces, causing sampling errors and distorting critical correlations between variable types. The article addresses this fundamental bottleneck by proposing CoDi, a generative architecture that processes continuous and discrete variables in their native mathematical spaces.

The main objective of the article is to evaluate and demonstrate whether decoupling mixed-type variables into two distinct, conditionally linked diffusion models—strengthened by contrastive learning—outperforms existing methods across data quality, diversity, and generation runtime. The approach employs two physically separated models: a continuous diffusion model operating in continuous space and a discrete diffusion model operating on categorical distributions. These models are co-evolved by continuously feeding each other's states as conditioning inputs throughout the forward noising and reverse denoising steps. Furthermore, the framework applies a contrastive learning objective with a specialized negative sampling strategy that shuffles discrete and continuous variable blocks across records, training the network to accurately preserve cross-type relationships. The approach was systematically evaluated against eight baseline generative models across 11 diverse real-world datasets spanning binary classification, multi-class classification, and regression tasks.

The experimental findings indicate substantial improvements across key performance metrics. First, CoDi achieved superior sampling quality; downstream machine learning models trained on CoDi's synthetic data scored higher average classification metrics (0.4726 binary F1 and 0.6221 multi-class macro F1) and were the only generative method to achieve a practical, positive predictive regression fit (average R-squared of 0.4794, compared to negative values for all eight baselines). Second, CoDi demonstrated higher sample diversity, attaining an average distribution coverage score of 0.6931, noticeably outperforming leading methods such as STaSy (0.5771) and TableGAN (0.5759). Third, the architecture achieved these gains with exceptional computational efficiency, generating samples nearly nine times faster than previous diffusion-based models (an average runtime of 0.5187 seconds versus 4.6417 seconds for STaSy to produce 10,000 records).

These results demonstrate that treating continuous and categorical variables in their proper mathematical spaces eliminates structural artifacts, such as over-represented categories or blurred decision boundaries. For organizations utilizing synthetic data, this translates directly to more reliable analytic models, lower downstream operational risk, and reduced computational costs. When selecting synthetic data solutions, decision-makers face trade-offs between rapid but low-fidelity generators and accurate but slow models. CoDi largely resolves this compromise by providing production-grade data fidelity and diversity at runtimes competitive with simpler models.

Organizations evaluating synthetic data pipelines should consider adopting dual-model architectures that process data types natively rather than forcing categorical variables into continuous spaces. Before deployment in production, teams should run pilot evaluations on their proprietary data schema, as data distributions vary widely across business domains. Although confidence in the experimental results is high across the 11 benchmark datasets, the primary limitation noted in the article is that CoDi is specialized for mixed-type tables and is not designed for purely continuous or purely discrete datasets. Additionally, as synthetic data quality approaches parity with genuine records, stakeholders must implement appropriate privacy protections to safeguard against potential downstream misuse.

No sufficiently relevant recommendations were found.

Cover for CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis

Abstract

With growing attention to tabular data these days, the attempt to apply a synthetic table to various tasks has been expanded toward various scenarios. Owing to the recent advances in generative modeling, fake data generated by tabular data synthesis models become sophisticated and realistic. However, there still exists a difficulty in modeling discrete variables (columns) of tabular data. In this work, we propose to process continuous and discrete variables separately (but being conditioned on each other) by two diffusion models. The two diffusion models are co-evolved during training by reading conditions from each other. In order to further bind the diffusion models, moreover, we introduce a contrastive learning method with a negative sampling method. In our experiments with 11 real-world tabular datasets and 8 baseline methods, we prove the efficacy of the proposed method, called CoDi. Our code is available at https://github.com/ChaejeongLee/CoDi.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Diffusion Probabilistic Models
  • 2.1.1. Continuous Spaces
  • 2.1.2. Discrete Spaces
  • 2.2. Tabular Data Synthesis
  • 3. Proposed Method
  • 3.1. Co-evolving Conditional Diffusion Models
  • 3.2. Contrastive Learning
  • 3.3. Training & Sampling Algorithms
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.1.1. Datasets & Baselines
  • 4.1.2. Evaluation Methods
  • 4.2. Experimental Results
  • 4.2.1. Sampling Quality
  • 4.2.2. Sampling Diversity
  • 4.2.3. Sampling Time
  • 4.3. Continuous vs. Discrete Spaces for Discrete Data
  • 4.4. Negative Sampling Methods
  • 4.5. Ablation Study on Contrastive Learning
  • 5. Conclusions
  • Acknowledgements
  • References
  • A. Preliminary Experiment
  • B. Network Architecture
  • C. The Reverse Transition Probabilities for the Co-evolving Conditional Diffusion Models
  • D. Detailed Experimental Settings for Reproducibility
  • D.1. Experimental Environments
  • D.2. Hyperparameter Settings for CoDi
  • D.3. Datasets
  • D.4. Baselines
  • E. Additional Experimental Results
  • E.1. Sampling Quality
  • E.2. Sampling Diversity
  • E.3. Sampling Time

Knowls

  1. Knowl 1 — Paper content unavailable for knowl extraction

    limitation

    The source document (PDF file 25e0c3fc-f0d4-44cf-b88d-b411ba50c3f5.pdf) could not be read, so no methods, results, equations, or other contributed content from the paper are represented in this knowl set.

Coverage note — No knowls were extracted: the attached PDF's text content was not accessible in this session, so no contributed material could be identified or summarized. This is a completeness failure of the extraction, not a deliberate omission of a specific paper section.

Citation

MLA
Lee, C., et al. “CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis”. International Conference on Machine Learning, vol. 202, 2023, pp. 18940–56, https://proceedings.mlr.press/v202/lee23i.html.
APA
Lee, C., Kim, J., & Park, N. (2023). CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis. International Conference on Machine Learning, 202, 18940–18956. https://proceedings.mlr.press/v202/lee23i.html
Chicago
Lee, C., J. Kim, and N. Park. 2023. “CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis”. International Conference on Machine Learning 202: 18940–56. https://proceedings.mlr.press/v202/lee23i.html.
Harvard
Lee, C., Kim, J. and Park, N. (2023) “CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis”, International Conference on Machine Learning. PMLR, pp. 18940–18956. Available at: https://proceedings.mlr.press/v202/lee23i.html.
Vancouver
1. Lee C, Kim J, Park N (2023) CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis. In: International Conference on Machine Learning. PMLR, pp 18940–18956

BibTeX

@InProceedings{pmlr-v202-lee23i,
  title = 	 {{C}o{D}i: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular Synthesis},
  author =       {Lee, Chaejeong and Kim, Jayoung and Park, Noseong},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {18940--18956},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/lee23i/lee23i.pdf},
  url = 	 {https://proceedings.mlr.press/v202/lee23i.html},
  abstract = 	 {With growing attention to tabular data these days, the attempt to apply a synthetic table to various tasks has been expanded toward various scenarios. Owing to the recent advances in generative modeling, fake data generated by tabular data synthesis models become sophisticated and realistic. However, there still exists a difficulty in modeling discrete variables (columns) of tabular data. In this work, we propose to process continuous and discrete variables separately (but being conditioned on each other) by two diffusion models. The two diffusion models are co-evolved during training by reading conditions from each other. In order to further bind the diffusion models, moreover, we introduce a contrastive learning method with a negative sampling method. In our experiments with 11 real-world tabular datasets and 8 baseline methods, we prove the efficacy of the proposed method, called $\texttt{CoDi}$. Our code is available at https://github.com/ChaejeongLee/CoDi.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/