TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation

Bowen DengChang XuHao LiYuhao HuangMin HouJiang Bian

article2025KDD19 citations

Introduces TarDiff, a diffusion framework that incorporates influence-function gradients into synthetic electronic health record generation to directly maximize downstream clinical model performance instead of merely mimicking observed data distributions.

Listen

Clinical machine learning models have become vital for disease diagnosis, prognosis prediction, and treatment planning, yet their development is constrained by medical data scarcity, privacy regulations, and severe class imbalances. While generative models can create synthetic electronic health record time-series data to expand training sets, conventional approaches focus strictly on mimicking real-world data distributions. This traditional focus creates a critical bottleneck: synthetic samples often replicate majority-class patterns while failing to adequately represent rare but life-threatening clinical conditions, leaving downstream diagnostic models vulnerable to diagnostic errors.

The article demonstrates and evaluates TarDiff, a target-oriented diffusion framework that optimizes synthetic time-series generation specifically to improve the performance of downstream clinical prediction tasks. Rather than simply reproducing statistical patterns, the article evaluates whether steering synthetic generation using statistical influence functions—which measure how a generated sample directly reduces predictive error on target outcomes—yields higher clinical utility.

To achieve this, the authors developed a three-stage method that integrates gradient-based influence guidance into a conditional diffusion process. A downstream predictive model is first pre-trained on available data to establish task-specific parameters. Gradient information from a representative guidance set is then computed and cached to quantify how changes in training data affect predictive loss. Finally, these aggregated influence gradients are embedded into the reverse generation process to actively guide time-series generation toward high-utility samples. The framework was evaluated across six benchmark datasets, including large-scale critical care records (MIMIC-III and eICU) and physiological signal datasets covering electrocardiography and electroencephalography (such as APAVA, ADFTD, PTB, and TDBRAIN) for tasks including mortality and length-of-stay predictions.

The findings show that TarDiff consistently establishes state-of-the-art performance across all tested datasets. When predictive models were trained purely on synthetic data and evaluated on real-world test sets, TarDiff outperformed competing generative baselines by up to 20.4% in precision-recall area under the curve and up to 18.4% in receiver operating characteristic area under the curve. When used to augment real datasets, TarDiff delivered sustained performance gains across multiple synthetic-to-real ratios, whereas competing baselines degraded at higher mix levels. Crucially, the approach resolved class imbalances by naturally assigning larger gradient signals to minority cases, doubling the minority-class F1 metric (+93%) on MIMIC-III and boosting it by 44% on eICU compared to training on real data alone. Furthermore, the method added minimal computational overhead, requiring only a one-time gradient caching step lasting between 10 and 168 seconds and achieving faster per-sample diffusion runtime than existing diffusion alternatives.

These results indicate that generative frameworks in healthcare should shift from passive statistical fidelity toward active utility optimization. By prioritizing the reduction of downstream diagnostic errors, synthetic data generation can directly enhance clinical decision support, mitigate the risk of overlooked rare conditions, and facilitate privacy-conscious research collaboration without compromising predictive performance.

Healthcare technology leaders and data science teams should consider piloting influence-guided synthetic data pipelines to expand constrained clinical datasets, particularly in diagnostic workflows hindered by severe class imbalance. Initial deployments should utilize held-out validation data as guidance sets while systematically tuning the influence scaling factor to balance sample diversity against targeted task optimization.

Confidence in these findings is supported by consistent empirical results across diverse clinical modalities and scales. However, stakeholders should note that the approach relies on having an initial representative guidance set to compute accurate influence gradients. Organizations with limited or poorly annotated initial data should exercise caution, as low-quality guidance sets may reduce the accuracy of the steering signals.

arXiv: 2504.17613
  • Paper: Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation, Jiaming Song et al. (2023). This paper establishes the foundational framework for guiding diffusion models using arbitrary task loss gradients during reverse sampling, directly underpinning TarDiff's task-specific influence guidance mechanism.
  • Paper: Non-autoregressive Conditional Diffusion Models for Time Series Prediction, Lifeng Shen et al. (2023). This work introduces non-autoregressive conditional diffusion architectures for multivariate time series, providing key architectural and sampling concepts for time-series diffusion generation.
  • Paper: Time-series Generative Adversarial Networks, Jinsung Yoon et al. (2019). This seminal paper formulates synthetic time-series generation for medical and sequential datasets, establishing standard downstream utility metrics and baseline paradigms that TarDiff seeks to improve.
  • Paper: Generating High Fidelity Data from Low-density Regions using Diffusion Models, Vikash Sehwag et al. (2022). This paper details guided sampling techniques in diffusion models to actively steer generation toward underrepresented low-density regions, addressing class imbalance concepts central to TarDiff.
  • Paper: TabDDPM: Modelling Tabular Data with Diffusion Models, Akim Kotelnikov et al. (2023). This work explores diffusion probabilistic modeling on mixed structured data and evaluates downstream machine learning utility versus statistical fidelity, a core comparison perspective in TarDiff.
  • Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). This work introduces gradient-based classifier guidance during the reverse diffusion sampling process, providing the mathematical foundation adapted by influence-guided diffusion methods.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational paper establishes the standard formulation of denoising diffusion probabilistic models and reverse denoising trajectories utilized across diffusion-based generative modeling.

No sufficiently relevant recommendations were found.

Cover for TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation

Abstract

Synthetic Electronic Health Record (EHR) time-series generation is crucial for advancing clinical machine learning models, as it helps address data scarcity by providing more training data. However, most existing approaches focus primarily on replicating statistical distributions and temporal dependencies of real-world data. We argue that fidelity to observed data alone does not guarantee better model performance, as common patterns may dominate, limiting the representation of rare but important conditions. This highlights the need for generate synthetic samples to improve performance of specific clinical models to fulfill their target outcomes. To address this, we propose TarDiff, a novel target-oriented diffusion framework that integrates task-specific influence guidance into the synthetic data generation process. Unlike conventional approaches that mimic training data distributions, TarDiff optimizes synthetic samples by quantifying their expected contribution to improving downstream model performance through influence functions. Specifically, we measure the reduction in task-specific loss induced by synthetic samples and embed this influence gradient into the reverse diffusion process, thereby steering the generation towards utility-optimized data. Evaluated on six publicly available EHR datasets, TarDiff achieves state-of-the-art performance, outperforming existing methods by up to 20.4% in AUPRC and 18.4% in AUROC. Our results demonstrate that TarDiff not only preserves temporal fidelity but also enhances downstream model performance, offering a robust solution to data scarcity and class imbalance in healthcare analytics.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 2.1 Diffusion Models
  • 2.2 Task Formulation
  • 3 Methodology
  • 3.1 Influence Formulation
  • 3.2 Influence Guidance Diffusion
  • 3.3 Estimates of Influence
  • 3.4 TarDiff Pipeline
  • 4 Experiment Setup and Result Analysis
  • 4.1 Datasets, Baselines, and Evaluation Metrics
  • 4.2 Train on Synthetic, Test on Real (TSTR)
  • 4.3 Train on Synthetic and Real, Test on Real (TSRTR)
  • 4.4 Influence Guidance under Class Imbalance
  • 4.5 Complexity Analysis.
  • 4.6 Sample Influence with Performance
  • 5 Related work
  • 6 Conclusion.
  • References
  • A Dataset Details
  • A.1 Critical Care EHR Datasets (MIMIC-III and eICU)
  • A.2 Specialized Physiological Signal Datasets (EEG and ECG)
  • A.3 Scalability Across Different Dataset Sizes
  • A.4 Controllability of Guidance Set Scale
  • B Evaluation Metrics Details
  • C Runtime and Overhead Comparison
  • C.1 One-time Overhead
  • C.2 Sampling Runtime Comparison
  • D Model Structure and Implementation Details
  • E Additional Results on Fidelity & Privacy
  • F Visualization of Synthetic Data
  • F.1 More TSRTS Results

Knowls

  1. Knowl 1 — Task-Oriented Influence Formulation for Synthetic Data Generation

    model/method

    TarDiff formulates the objective of synthetic time-series generation not merely as matching empirical data distributions, but as maximizing the utility of generated samples for downstream clinical tasks.

    Let Dtrain={(xi,yi)}i=1n\mathcal{D}_{\text{train}} = \{(x_i, y_i)\}_{i=1}^n denote a training dataset, and let ℓT(x,y;ϕ)\ell_T(x, y; \phi) be a task-specific loss parameterized by ϕ\phi. The optimal parameter vector on the original training set is: ϕT∗=arg⁡min⁡ϕ∑(xi,yi)∈DtrainℓT(xi,yi;ϕ)\phi^*_T = \arg\min_\phi \sum_{(x_i, y_i) \in \mathcal{D}_{\text{train}}} \ell_T(x_i, y_i; \phi)

    When a candidate synthetic sample z^=(x,y)\hat{z} = (x, y) is added to Dtrain\mathcal{D}_{\text{train}}, the updated optimal parameter vector becomes: ϕTz^=arg⁡min⁡ϕ∑(xi,yi)∈Dtrain∪{z^}ℓT(xi,yi;ϕ)\phi^{\hat{z}}_T = \arg\min_\phi \sum_{(x_i, y_i) \in \mathcal{D}_{\text{train}} \cup \{\hat{z}\}} \ell_T(x_i, y_i; \phi)

    The pointwise change in downstream loss on an evaluation instance (x′,y′)(x', y') caused by adding (x,y)(x, y) to the training data is: Hϕ(x,y,x′,y′)=ℓT(x′,y′;ϕTz^)−ℓT(x′,y′;ϕT∗)H_\phi(x, y, x', y') = \ell_T(x', y'; \phi^{\hat{z}}_T) - \ell_T(x', y'; \phi^*_T)

    For an underlying data-generating distribution P\mathcal{P}, the task influence ΔLT(z^)\Delta\mathcal{L}_T(\hat{z}) of synthetic sample z^\hat{z} is defined as the expected reduction in downstream loss over unseen test instances: ΔLT(z^)≜−E(x′,y′)∼P[Hϕ(x,y,x′,y′)]=−E(x′,y′)∼P[ℓT(x′,y′;ϕTz^)−ℓT(x′,y′;ϕT∗)]\Delta\mathcal{L}_T(\hat{z}) \triangleq -\mathbb{E}_{(x', y') \sim \mathcal{P}} \left[ H_\phi(x, y, x', y') \right] = -\mathbb{E}_{(x', y') \sim \mathcal{P}} \left[ \ell_T(x', y'; \phi^{\hat{z}}_T) - \ell_T(x', y'; \phi^*_T) \right]

    The target synthetic sample z^∗\hat{z}^* is the one that maximizes this expected loss reduction: z^∗=arg⁡max⁡z^ΔLT(z^)\hat{z}^* = \arg\max_{\hat{z}} \Delta\mathcal{L}_T(\hat{z})

  2. Knowl 2 — First-Order Gradient Approximation of Influence and Diffusion Guidance

    model/method

    To avoid retraining downstream model parameters ϕ\phi for every candidate synthetic sample z^=(x,y)\hat{z} = (x, y), TarDiff computes parameter shift δϕ\delta\phi via first-order gradient analysis over an independent guidance set Dguide={(xj′,yj′)}j=1Ng\mathcal{D}_{\text{guide}} = \{(x'_j, y'_j)\}_{j=1}^{N_g} drawn i.i.d. from the target data distribution.

    For a small loss modification scale ε\varepsilon, the parameter adjustment induced by a sample z^\hat{z} is approximated as δϕ=ε∇ϕℓT(z^;ϕ)∥∇ϕℓT(z^;ϕ)∥2\delta\phi = \varepsilon \frac{\nabla_\phi \ell_T(\hat{z}; \phi)}{\|\nabla_\phi \ell_T(\hat{z}; \phi)\|^2}. The resulting aggregated reduction in task loss across Dguide\mathcal{D}_{\text{guide}} is: ΔLT(z^)=∑(x′,y′)∈Dguide[ℓT(x′,y′;ϕ)−ℓT(x′,y′;ϕ+δϕ)]≈−∑(x′,y′)∈Dguideε∇ϕℓT(x′,y′;ϕ)⋅∇ϕℓT(z^;ϕ)∥∇ϕℓT(z^;ϕ)∥2\Delta\mathcal{L}_T(\hat{z}) = \sum_{(x', y') \in \mathcal{D}_{\text{guide}}} [\ell_T(x', y'; \phi) - \ell_T(x', y'; \phi + \delta\phi)] \approx -\sum_{(x', y') \in \mathcal{D}_{\text{guide}}} \varepsilon \frac{\nabla_\phi \ell_T(x', y'; \phi) \cdot \nabla_\phi \ell_T(\hat{z}; \phi)}{\|\nabla_\phi \ell_T(\hat{z}; \phi)\|^2}

    Defining the cumulative guidance gradient vector GG over Dguide\mathcal{D}_{\text{guide}} as: G=−∑(x′,y′)∈Dguideε∇ϕℓT(x′,y′;ϕ)∥∇ϕℓT(z^;ϕ)∥2G = -\sum_{(x', y') \in \mathcal{D}_{\text{guide}}} \varepsilon \frac{\nabla_\phi \ell_T(x', y'; \phi)}{\|\nabla_\phi \ell_T(\hat{z}; \phi)\|^2} the sample influence simplifies to: ΔLT(z^)=∇ϕℓT(z^;ϕ)⋅G\Delta\mathcal{L}_T(\hat{z}) = \nabla_\phi \ell_T(\hat{z}; \phi) \cdot G

    During reverse diffusion sampling at timestep tt, given the current noisy sample state xtx_t and condition label yy, the transition mean μθ(xt,y,t)\mu_\theta(x_t, y, t) is shifted along the influence gradient: μ~θ(xt,y,t)=μθ(xt,y,t)+w∇xtΔLT(z^t)=μθ(xt,y,t)+w∇xt[G⋅∇ϕℓT(xt,y;ϕ∗)]\tilde{\mu}_\theta(x_t, y, t) = \mu_\theta(x_t, y, t) + w \nabla_{x_t}\Delta\mathcal{L}_T(\hat{z}_t) = \mu_\theta(x_t, y, t) + w \nabla_{x_t} \left[ G \cdot \nabla_\phi \ell_T(x_t, y; \phi^*) \right] where ww controls the influence guidance strength and ϕ∗\phi^* is the pre-trained downstream model parameter vector.

  3. Knowl 3 — TarDiff Generation Pipeline

    algorithm

    TarDiff executes a three-step generative workflow: (1) pre-training the downstream task model on real data, (2) computing and caching guidance dataset influence gradients, and (3) running reverse diffusion guided by the influence vector.

    Input: Training set Dtrain\mathcal{D}_{\text{train}}, guidance set Dguide\mathcal{D}_{\text{guide}}, conditional diffusion model (μθ,Σθ)(\mu_\theta, \Sigma_\theta), downstream model fϕf_\phi with loss ℓ\ell, diffusion timesteps TT, guidance scale ww, condition label yy
    Output: Synthetic sample z^=(x0,y)\hat{z} = (x_0, y)
    Step 1: Pre-train downstream model
    ϕ∗←arg⁡min⁡ϕ∑(x,y)∈Dtrainℓ(x,y;ϕ)\phi^* \leftarrow \arg\min_\phi \sum_{(x, y) \in \mathcal{D}_{\text{train}}} \ell(x, y; \phi)
    Step 2: Compute and cache data influence
    G←0G \leftarrow 0
    for each (xi,yi)∈Dguide(x_i, y_i) \in \mathcal{D}_{\text{guide}} do
        gi←∇ϕℓ(xi,yi;ϕ∗)g_i \leftarrow \nabla_\phi \ell(x_i, y_i; \phi^*)
        G←G+giG \leftarrow G + g_i
    end for
    G←1∣Dtrain∣GG \leftarrow \frac{1}{|\mathcal{D}_{\text{train}}|} G
    Step 3: Influence-guided reverse diffusion sampling
    xT∼N(0,I)x_T \sim \mathcal{N}(0, I)
    for t=Tt = T down to 11 do
        μt←μθ(xt,y,t)\mu_t \leftarrow \mu_\theta(x_t, y, t)
        J←∇xt[G⋅∇ϕℓ(xt,y;ϕ∗)]J \leftarrow \nabla_{x_t} [G \cdot \nabla_\phi \ell(x_t, y; \phi^*)]
        μ~t←μt+w⋅J\tilde{\mu}_t \leftarrow \mu_t + w \cdot J
        xt−1∼N(μ~t,Σθ(t))x_{t-1} \sim \mathcal{N}(\tilde{\mu}_t, \Sigma_\theta(t))
    end for
    return z^←(x0,y)\hat{z} \leftarrow (x_0, y)
  4. Knowl 4 — Train on Synthetic, Test on Real (TSTR) Clinical Performance

    data/table

    Under the Train on Synthetic, Test on Real (TSTR) evaluation protocol, a downstream classifier (TimesNet) is trained purely on synthetic data produced by a given generative method and evaluated on held-out real clinical test sets. TarDiff was evaluated against generative baselines across two critical care EHR datasets (MIMIC-III and eICU, evaluating in-hospital Mortality and ICU Stay > 3 days) and four physiological signal datasets (APAVA for Alzheimer's EEG, ADFD for dementia EEG, PTB for myocardial infarction ECG, and TDBRAIN for Parkinson's EEG).

    Dataset / Task Metric Real Data Best Baseline TarDiff
    MIMIC-III Mortality AUPRC 0.1736 0.1402 (TimeGAN) 0.1799
    MIMIC-III Mortality AUROC 0.6350 0.5392 (TimeVAE) 0.6373
    MIMIC-III ICU Stay AUPRC 0.4618 0.3939 (TimeVAE) 0.4183
    MIMIC-III ICU Stay AUROC 0.6282 0.5656 (TimeVAE) 0.5800
    eICU Mortality AUPRC 0.2072 0.1592 (TimeGAN) 0.1698
    eICU Mortality AUROC 0.6869 0.6238 (TimeGAN) 0.6308
    eICU ICU Stay AUPRC 0.6004 0.4753 (TimeVAE) 0.5583
    eICU ICU Stay AUROC 0.6615 0.5323 (BioDiffusion) 0.6184
    APAVA (EEG) AUPRC 0.7669 0.7427 (TimeVAE) 0.7652
    APAVA (EEG) AUROC 0.7206 0.6857 (TimeVAE) 0.7710
    ADFD (EEG) AUPRC 0.4335 0.4604 (BioDiffusion) 0.4795
    ADFD (EEG) AUROC 0.6239 0.6375 (BioDiffusion) 0.6443
    PTB (ECG) AUPRC 0.9677 0.9509 (TimeVAE) 0.9544
    PTB (ECG) AUROC 0.9331 0.8984 (TimeVQVAE) 0.9053
    TDBRAIN (EEG) AUPRC 0.9642 0.6093 (TimeGAN) 0.6542
    TDBRAIN (EEG) AUROC 0.9615 0.6278 (TimeGAN) 0.6444

    TarDiff achieves state-of-the-art synthetic data utility across all evaluated clinical and physiological tasks, outperforming baseline models (TimeGAN, TimeVAE, TimeVQVAE, DiffusionTS, BioDiffusion) in AUROC and AUPRC, and in several cases matching or exceeding the performance of models trained directly on the original real training sets.

  5. Knowl 5 — Train on Synthetic and Real, Test on Real (TSRTR) Augmentation Dynamics

    empirical result

    When real clinical training data is augmented with synthetic time series across varying synthetic-to-real mixing ratios α∈{0.2,0.4,0.6,0.8,1.0}\alpha \in \{0.2, 0.4, 0.6, 0.8, 1.0\}, where Dtrain(α)=Dreal∪Dsynthetic(α)\mathcal{D}_{\text{train}}(\alpha) = \mathcal{D}_{\text{real}} \cup \mathcal{D}_{\text{synthetic}}(\alpha):

    1. Baseline generative models (TimeGAN, TimeVAE, TimeVQVAE, DiffusionTS, BioDiffusion) display plateauing, erratic, or declining AUROC scores on downstream TimesNet classification as the synthetic fraction α\alpha increases from 0.20.2 to 1.01.0, indicating that unguided synthetic generation introduces confounding statistical artifacts at higher volumes.

    2. In contrast, TarDiff produces a monotonic upward trend in AUROC across both MIMIC-III and eICU tasks (Mortality prediction and ICU Stay prediction) as the mixing ratio α\alpha scales from 0.20.2 up to 1.01.0. The performance advantage over baselines widens substantially at higher synthetic ratios (0.80.8 and 1.01.0), demonstrating that target-oriented influence guidance generates samples that complement rather than dilute the real dataset's task-relevant signal.

  6. Knowl 6 — Class Imbalance Mitigation via Gradient Magnitude Disparities

    empirical result

    TarDiff inherently alleviates severe class imbalance in clinical time series without requiring explicit class weighting or specialized resampling techniques.

    1. Gradient Norm Disparities: In pre-trained TimesNet classifiers on mortality prediction datasets where minority positive instances comprise only ≈9–11%\approx 9\text{--}11\% of admissions, minority instances produce significantly larger loss gradient ℓ2\ell_2-norms than majority instances:
    • MIMIC-III Mortality: Majority gradient norm is 1.06±1.231.06 \pm 1.23, whereas minority gradient norm is 16.85±2.4816.85 \pm 2.48.
    • eICU Mortality: Majority gradient norm is 5.41±5.055.41 \pm 5.05, whereas minority gradient norm is 37.86±8.7337.86 \pm 8.73.
    1. Minority Class Classification Gains: Under the TSRTR protocol, augmenting real data with TarDiff-generated samples substantially boosts minority-class F1F_1 scores relative to the real-only baseline (TRTR):
    • On MIMIC-III, minority F1F_1 increases from 0.0560.056 to 0.1080.108 (+92.9%+92.9\% gain).
    • On eICU, minority F1F_1 increases from 0.0130.013 to 0.0180.018 (+38.5%+38.5\% gain).
    1. Guidance Ablation by Class: Isolating the gradient cache GG to minority-only samples achieves the highest minority F1F_1 (0.1630.163 on MIMIC-III, 0.0250.025 on eICU), while majority-only gradient guidance degrades minority performance (0.0660.066 on MIMIC-III, 0.0120.012 on eICU). Because harder, under-represented minority cases generate higher gradient magnitudes, unified influence guidance naturally biases the diffusion trajectory toward synthesizing high-utility minority-class dynamics.
  7. Knowl 7 — Computational Overhead and Sampling Efficiency of TarDiff

    empirical result

    The computational overhead of TarDiff comprises two components:

    1. One-Time Gradient Caching Overhead: Pre-training the downstream model and performing forward-backward passes over the guidance set Dguide\mathcal{D}_{\text{guide}} of size NtN_t takes O(Nt⋅fT(L,D))\mathcal{O}(N_t \cdot f_T(L, D)) operations, where fT(L,D)f_T(L, D) is the downstream backpropagation complexity for sequence length LL and feature dimension DD. Measured one-time gradient caching overhead across eight clinical/physiological datasets ranges from 10.5110.51 seconds (TDBRAIN) to 167.81167.81 seconds (ADFD); for MIMIC-III and eICU tasks it is under 3535 seconds.

    2. Per-Step Reverse Sampling Cost: Sampling across TT diffusion steps with batch size BsampleB_{\text{sample}} incurs O(T⋅Bsample⋅g(L,D))\mathcal{O}(T \cdot B_{\text{sample}} \cdot g(L, D)) overhead, where g(L,D)g(L, D) represents the lightweight dot-product and gradient pass through the fixed downstream model. Compared to standard diffusion cost O(T⋅Bsample⋅h(L,D))\mathcal{O}(T \cdot B_{\text{sample}} \cdot h(L, D)), the overhead ratio g(L,D)h(L,D)\frac{g(L, D)}{h(L, D)} is negligible.

    Sampling runtime per sample:

    • TimeGAN (GAN): 0.0005 s/sample0.0005\text{ s/sample}
    • TimeVQE (VAE): 0.0006 s/sample0.0006\text{ s/sample}
    • TimeVQVAE (VAE): 0.0047 s/sample0.0047\text{ s/sample}
    • TarDiff (Diffusion): 0.0259 s/sample0.0259\text{ s/sample}
    • DiffusionTS (Diffusion): 0.1340 s/sample0.1340\text{ s/sample}
    • BioDiffusion (Diffusion): 0.3008 s/sample0.3008\text{ s/sample}

    TarDiff samples substantially faster than comparable diffusion models while incorporating full downstream gradient guidance.

  8. Knowl 8 — Influence Scaling and Guidance Set Size Robustness

    empirical result

    Empirical evaluations demonstrate that TarDiff reliably modulates sample influence and is robust to the scale of the reference guidance set:

    1. Influence Scale Modulation: Sweeping the guidance scaling weight ww from −1000-1000 to +1000+1000 on MIMIC-III validation splits alters the estimated sample influence ΔLT(z^)\Delta\mathcal{L}_T(\hat{z}) and drives downstream AUROC. In both Mortality and ICU Stay prediction tasks, increasing guidance scale ww toward +1000+1000 yields corresponding monotonic improvements in validation AUROC, with downstream model performance peaking at moderate-to-high influence scales (w≈+1000w \approx +1000).

    2. Guidance Set Scale Invariance: Varying the size of Dguide\mathcal{D}_{\text{guide}} across sizes from 100100 to 450450 samples on the PTB dataset produces consistently stable downstream classification AUROCs (remaining strictly within the narrow band of 0.63830.6383 to 0.64160.6416). This indicates that TarDiff does not require large guidance subsets to construct reliable gradient caches.

  9. Knowl 9 — Distribution Fidelity, Membership Privacy, and Predictive Scores

    data/table

    TarDiff balances distribution fidelity, privacy preservation against membership inference attacks, and predictive utility on clinical EHR datasets. Evaluation metrics include Distribution Similarity (DS, measuring classifier discriminability between real and synthetic data; lower is better), Membership Inference Risk (MIR, measuring vulnerability to membership inference attacks; lower is better), and Predictive Score (PS, mean prediction error when models trained on synthetic data predict real outcomes; lower is better).

    MIMIC-III eICU
    Method Mortality ICU Stay Mortality ICU Stay
    Distribution Similarity (DS) ↓\downarrow
    TarDiff 0.000201 0.000000 0.1810 0.1488
    TimeGAN 0.000201 0.000000 0.2471 0.3781
    TimeVAE 0.000000 0.030400 0.1709 0.3668
    TimeVQ-VAE 0.032500 0.001500 0.3663 0.0000
    DiffusionTS 0.000101 0.445100 0.4931 0.5000
    BioDiffusion 0.000000 0.498900 0.5000 0.4968
    Membership Inference Risk (MIR) ↓\downarrow
    TarDiff 0.6761 0.6787 0.6667 0.6667
    BioDiffusion 0.7316 0.8114 0.8199 0.7736
    TimeVQ-VAE 0.6792 0.6818 0.6668 0.6667
    TimeGAN 0.6949 0.6762 0.6668 0.6668
    TimeVAE 0.6788 0.9169 0.6668 0.6667
    DiffusionTS 0.9683 0.6811 0.9976 0.9349
    Privacy Score (PS) ↓\downarrow
    TarDiff 0.5819 0.5669 0.5770 0.5795
    BioDiffusion 0.6226 0.6656 0.5744 0.5954
    TimeVQ-VAE 1.2898 1.2115 0.5671 0.5006
    TimeGAN 0.7035 0.8837 0.6420 0.5198
    TimeVAE 0.9856 0.9616 0.5028 0.5071
    DiffusionTS 0.8671 0.9097 0.6821 0.6558

    TarDiff maintains low distribution similarity discrepancy (matching real-data distributions), achieves minimal membership risk close to random guessing baseline (MIR≈0.667\text{MIR} \approx 0.667), and achieves lower predictive error than diffusion and GAN baselines.

  10. Knowl 10 — Denoising Network Architecture and Training Specification

    experimental setup

    The core generative component of TarDiff uses a 1D U-Net denoising architecture tailored for multi-channel medical time-series data:

    1. Backbone Configuration:
    • Base channels: 64 initial feature channels.
    • Multi-resolution levels: Four hierarchical levels with channel multipliers [1,2,4,4][1, 2, 4, 4].
    • Residual units: Three residual blocks per level, utilizing residual connections and scale-shift normalization during up-sampling and down-sampling operations.
    • Temporal Attention: Multi-head self-attention mechanisms with 88 attention heads placed at resolution scales with multipliers 11, 22, and 44 to capture long-range temporal dependencies.
    • Conditioning: Contextual label embeddings are projected into a 3232-dimensional latent space, supporting classifier-free guidance.
    1. Training and Hardware Details:
    • Batch size: 256256, trained for 20,00020,000 iterations.
    • Learning rate: Constant learning rate of 1×10−41 \times 10^{-4}.
    • Sampling guidance: Default guidance scale w=100w = 100.
    • Compute environment: A single NVIDIA A100 GPU with 80GB VRAM.

Coverage note — None was omitted; all core contributions, theoretical formulations, algorithms, architectures, and empirical evaluations from the paper are represented.

References

  1. 1.Nikhil Anand, Joshua Tan, and Maria Minakova. Influence scores at scale for efficient language data sampling. arXiv preprint arXiv:2311.16298, 2023.
  2. 2.Guillaume Charpiat, Nicolas Girard, Loris Felardos, and Yuliya Tarabalka. Input similarity from the neural network perspective. Advances in Neural Information Processing Systems, 32, 2019.
  3. 3.Jiabo Chen, Yongfan Lai, Deyun Zhang, Yue Wang, Shijia Geng, Hongyan Li, and Shenda Hong. Diffusets: 12-lead ecg generation conditioned on clinical text reports and patient-specific information. In Artificial Intelligence and Data Science for Healthcare: Bridging Data-Centric AI and People-Centric Healthcare, 2024.
  4. 4.Jintai Chen, Kuanlun Liao, Kun Wei, Haochao Ying, Danny Z Chen, and Jian Wu. Me-gan: Learning panoptic electrocardio representations for multi-view ecg synthesis conditioned on heart diseases. In International Conference on Machine Learning, pages 3360–3370. PMLR, 2022.
  5. 5.Edward Choi, Siddharth Biswal, Bradley A. Malin, Jon Duke, Walter F. Stewart, and Jimeng Sun. Generating multi-label discrete electronic health records using generative adversarial networks. CoRR, abs/1703.06490, 2017.
  6. 6.R Dennis Cook. Detection of influential observation in linear regression. Technometrics, 19(1):15–18, 1977.
  7. 7.Abhyuday Desai, Cynthia Freeman, Zuhui Wang, and Ian Beaver. Timevae: A variational auto-encoder for multivariate time series generation. arXiv preprint arXiv:2111.08095, 2021.
  8. 8.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021.
  9. 9.J Escudero, Daniel Abásolo, Roberto Hornero, Pedro Espino, and Miguel López. Analysis of electroencephalograms in alzheimer’s disease patients with multiscale entropy. Physiological measurement, 27(11):1091, 2006.
  10. 10.Xinyao Fan, Yueying Wu, Chang Xu, Yuhao Huang, Weiqing Liu, and Jiang Bian. MG-TSD: multigranularity time series diffusion models with guided learning process. In ICLR. OpenReview.net, 2024.
  11. 11.A Goldberger, L Amaral, L Glass, J Hausdorff, PC Ivanov, R Mark, HE Stanley, and PhysioToolkit PhysioBank. Physionet: Components of a new research resource for complex physiologic signals components of a new research resource for complex physiologic signals. Circulation, 101:e215–e220, 2000.
  12. 12.Benjamin A Goldstein, Ann Marie Navar, Michael J Pencina, and John PA Ioannidis. Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic review. Journal of the American Medical Informatics Association: JAMIA, 24(1):198, 2016.
  13. 13.Aman Gupta, Deepak Bhatt, and Anubha Pandey. Transitioning from real to synthetic data: Quantifying the bias in model. arXiv preprint arXiv:2105.04144, 2021.
  14. 14.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  15. 15.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  16. 16.Min Hou, Yueying Wu, Chang Xu, Yu-Hao Huang, Chenxi Bai, Le Wu, and Jiang Bian. Invdiff: Invariant guidance for bias mitigation in diffusion models. CoRR, abs/2412.08480, 2024.
  17. 17.Yu-Hao Huang, Chang Xu, Yang Liu, Weiqing Liu, Wu-Jun Li, and Jiang Bian. Controllable financial market generation with diffusion guided meta agent. arXiv preprint arXiv:2408.12991, 2024.
  18. 18.Yu-Hao Huang, Chang Xu, Yueying Wu, Wu-Jun Li, and Jiang Bian. Timedp: Learning to generate multi-domain time series with domain prompts. arXiv preprint arXiv:2501.05403, 2025.
  19. 19.Zepeng Huo, Xiaoning Qian, Shuai Huang, Zhangyang Wang, and Bobak J Mortazavi. Density-aware personalized training for risk prediction in imbalanced medical data. In Machine Learning for Healthcare Conference, pages 101–122. PMLR, 2022.
  20. 20.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016.
  21. 21.Hojjat Karami, Mary-Anne Hartley, David Atienza, and Anisoara Ionescu. Timehr: Image-based time series generation for electronic health records. CoRR, abs/2402.06318, 2024.
  22. 22.Shruti Kaushik, Abhinav Choudhury, Pankaj Kumar Sheron, Nataraj Dasgupta, Sayee Natarajan, Larry A. Pickett, and Varun Dutt. AI in healthcare: Time-series forecasting using statistical, neural, and ensemble architectures. Frontiers Big Data, 3:4, 2020.
  23. 23.Diederik P Kingma. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  24. 24.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. In ICML, volume 70 of Proceedings of Machine Learning Research, pages 1885–1894. PMLR, 2017.
  25. 25.Daesoo Lee, Sara Malacarne, and Erlend Aune. Vector quantized time series generation with a bidirectional prior model. arXiv preprint arXiv:2303.04743, 2023.
  26. 26.Hao Li, Yuping Wu, Viktor Schlegel, Riza Batista-Navarro, Thanh-Tung Nguyen, Abhinav Ramesh Kashyap, Xiao-Jun Zeng, Daniel Beck, Stefan Winkler, and Goran Nenadic. Team: PULSAR at probsum 2023: PULSAR: pre-training with extracted healthcare terms for summarising patients’ problems and data augmentation with black-box large language models. In BioNLP@ACL, pages 503–509. Association for Computational Linguistics, 2023.
  27. 27.Hao Li, Yu-Hao Huang, Chang Xu, Viktor Schlegel, Ren-He Jiang, Riza Batista-Navarro, Goran Nenadic, and Jiang Bian. Bridge: Bootstrapping text to control time-series generation via multiagent iterative optimization and diffusion modelling. arXiv preprint arXiv:2503.02445, 2025.
  28. 28.Xiaomin Li, Mykhailo Sakevych, Gentry Atkinson, and Vangelis Metsis. Biodiffusion: A versatile diffusion model for biomedical signal synthesis. CoRR, abs/2401.10282, 2024a.
  29. 29.Yizhi Li, Ge Zhang, Xingwei Qu, Jiali Li, Zhaoqun Li, Noah Wang, Hao Li, Ruibin Yuan, Yinghao Ma, Kai Zhang, Wangchunshu Zhou, Yiming Liang, Lei Zhang, Lei Ma, Jiajun Zhang, Zuowen Li, Wenhao Huang, Chenghua Lin, and Jie Fu. Cif-bench: A chinese instruction-following benchmark for evaluating the generalizability of large language models. In ACL (Findings), pages 12431–12446. Association for Computational Linguistics, 2024b.
  30. 30.Yuanyuan Liang, Yanbing Ju, Xiao-Jun Zeng, Hao Li, Peiwu Dong, and Tian Ju. A user-generated content-based social network large-scale group decision-making approach in healthcare service: Case study of general practitioners selection in uk. Expert Systems with Applications, page 125542, 2024.
  31. 31.Barbara Mukami Maweu, Rittika Shamsuddin, Sagnik Dakshit, and Balakrishnan Prabhakaran. Generating healthcare time series data for improving diagnostic accuracy of deep neural networks. IEEE Trans. Instrum. Meas., 70:1–15, 2021.
  32. 32.Andreas Miltiadous, Katerina D Tzimourta, Theodora Afrantou, Panagiotis Ioannidis, Nikolaos Grigoriadis, Dimitrios G Tsalikakis, Pantelis Angelidis, Markos G Tsipouras, Euripidis Glavas, Nikolaos Giannakeas, et al. A dataset of scalp eeg recordings of alzheimer’s disease, frontotemporal dementia and healthy subjects from routine eeg. Data, 8(6):95, 2023.
  33. 33.Aishik Nagar, Viktor Schlegel, Thanh-Tung Nguyen, Hao Li, Yuping Wu, Kuluhan Binici, and Stefan Winkler. Llms are not zero-shot reasoners for biomedical information extraction. CoRR, abs/2408.12249, 2024.
  34. 34.Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. The eicu collaborative research database, a freely available multi-center database for critical care research. Scientific data, 5(1):1–13, 2018.
  35. 35.Viktor Schlegel, Hao Li, Yuping Wu, Anand Subramanian, Thanh-Tung Nguyen, Abhinav Ramesh Kashyap, Daniel Beck, Xiao-Jun Zeng, Riza Theresa Batista-Navarro, Stefan Winkler, and Goran Nenadic. PULSAR at mediqa-sum 2023: Large language models augmented by synthetic dialogue convert patient dialogues to medical records. In CLEF (Working Notes), volume 3497 of CEUR Workshop Proceedings, pages 1668–1679. CEUR-WS.org, 2023.
  36. 36.Julian Schön, Raghavendra Selvan, Lotte Nygård, Ivan Richter Vogelius, and Jens Petersen. Explicit temporal embedding in deep generative latent models for longitudinal medical image synthesis. CoRR, abs/2301.05465, 2023.
  37. 37.Benjamin Shickel, Patrick James Tighe, Azra Bihorac, and Parisa Rashidi. Deep ehr: a survey of recent advances in deep learning techniques for electronic health record (ehr) analysis. IEEE journal of biomedical and health informatics, 22(5):1589–1604, 2017.
  38. 38.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  39. 39.Muhang Tian, Bernie Chen, Allan Guo, Shiyi Jiang, and Anru R Zhang. Reliable generation of privacy-preserving synthetic electronic health record time series via diffusion models. Journal of the American Medical Informatics Association, 31(11):2529–2539, 2024.
  40. 40.Tzu-Wei Tseng, Chang-Fu Su, and Feipei Lai. Fast healthcare interoperability resources for inpatient deterioration detection with time-series vital signs: Design and implementation study. JMIR Medical Informatics, 10(10):e42429, 2022.
  41. 41.Hanneke Van Dijk, Guido Van Wingen, Damiaan Denys, Sebastian Olbrich, Rosalinde Van Ruth, and Martijn Arns. The two decades brainclinics research archive for insights in neurophysiology (tdbrain) database. Scientific data, 9(1):333, 2022.
  42. 42.Yihe Wang, Nan Huang, Taida Li, Yujun Yan, and Xiang Zhang. Medformer: A multi-granularity patching transformer for medical time-series classification. arXiv preprint arXiv:2405.19363, 2024.
  43. 43.Qingsong Wen, Liang Sun, Fan Yang, Xiaomin Song, Jingkun Gao, Xue Wang, and Huan Xu. Time series data augmentation for deep learning: A survey. arXiv preprint arXiv:2002.12478, 2020.
  44. 44.Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. Timesnet: Temporal 2d-variation modeling for general time series analysis. arXiv preprint arXiv:2210.02186, 2022.
  45. 45.Jinsung Yoon, Daniel Jarrett, and Mihaela Van der Schaar. Time-series generative adversarial networks. Advances in neural information processing systems, 32, 2019.
  46. 46.Xinyu Yuan and Yan Qiao. Diffusion-ts: Interpretable diffusion for general time series generation. arXiv preprint arXiv:2403.01742, 2024.

Citation

MLA
Deng, B., et al. “TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation”. arXiv, 2025, http://arxiv.org/abs/2504.17613v1.
APA
Deng, B., Xu, C., Li, H., Huang, Y., Hou, M., & Bian, J. (2025). TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation. arXiv. http://arxiv.org/abs/2504.17613v1
Chicago
Deng, B., C. Xu, H. Li, Y. Huang, M. Hou, and J. Bian. 2025. “TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation”. arXiv. http://arxiv.org/abs/2504.17613v1.
Harvard
Deng, B. et al. (2025) “TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2504.17613v1.
Vancouver
1. Deng B, Xu C, Li H, Huang Y, Hou M, Bian J (2025) TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation. arXiv

BibTeX

@article{deng2025tardiff,
  title = {TarDiff: Target-Oriented Diffusion Guidance for Synthetic Electronic Health Record Time Series Generation},
  author = {Deng, Bowen and Xu, Chang and Li, Hao and Huang, Yuhao and Hou, Min and Bian, Jiang},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2504.17613v1},
  eprint = {2504.17613}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/