Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations

Nabeel SeedatFergus ImrieAlexis BellotZhaozhi QianMihaela van der Schaar

article2022ICML78 citations

Proposes a continuous-time causal inference framework using neural controlled differential equations and adversarial training to reliably estimate individual treatment effects from irregularly sampled longitudinal data subject to time-dependent confounding.

Listen

Real-world medical decision-making relies heavily on answering "what-if" questions to determine the optimal type and timing of treatments. While observational healthcare data offers a rich resource for estimating counterfactual outcomes over time, most existing causal inference methods assume data arrives at fixed, regular intervals across all patients. In clinical practice, however, patient records are sampled irregularly due to missed visits, differing monitoring protocols, and varying disease severity. Furthermore, past treatments and patient health trajectories create time-dependent confounding, a major source of bias that standard time-series techniques cannot resolve.

The main objective of the article is to develop and evaluate a causal inference framework that can reliably estimate counterfactual outcomes at any continuous point in time from irregularly sampled observational data with time-dependent confounding.

To achieve this, the article introduces the Treatment Effect Neural Controlled Differential Equation (TE-CDE). Unlike traditional recurrent neural network models that force irregular observations into artificial discrete bins, TE-CDE models a patient's underlying latent health trajectory as a continuous process governed by controlled differential equations. The framework incorporates domain adversarial training to remove time-dependent confounding bias by making the latent representations invariant to treatment assignments. To validate performance, the approach was tested in a controllable simulation environment modeling lung cancer tumor growth under chemotherapy and radiotherapy regimes, using a Hawkes point process to emulate realistic, state-dependent clinical sampling patterns across datasets of 10,000 patients.

The analysis yielded several critical findings. First, TE-CDE consistently achieved the lowest counterfactual estimation error across all irregular sampling regimes and levels of time-dependent confounding, reducing root mean square error by approximately 36% compared to the strongest discrete baseline at high confounding levels. Second, when forecasting multiple steps into the future, TE-CDE demonstrated a 40% reduction in estimation error over competing methods at high confounding. Third, improved estimation accuracy translated directly into superior decision-making, increasing optimal treatment selection accuracy by about 4% under moderate confounding and up to roughly 10% under high confounding. Fourth, incorporating domain adversarial training proved essential, as removing it significantly degraded accuracy under severe confounding. Finally, TE-CDE demonstrated superior data efficiency, experiencing only a 17.1% error degradation when training data was reduced tenfold, compared to 18.9% and 57.2% for baseline methods.

These findings imply that treating patient health as an inherently continuous trajectory avoids the errors introduced by discrete interpolation and data imputation. This shift provides more dependable decision support for individualized treatment planning, helping clinicians avoid suboptimal therapy choices and potentially reducing treatment costs and adverse health risks. Notably, the results demonstrate that standard Gaussian process and recurrent models struggle significantly in irregular longitudinal settings with high confounding, establishing continuous-time neural differential equations as a more appropriate paradigm.

Organizations evaluating continuous-time causal models should implement them within "human-in-the-loop" clinical workflows rather than fully automated pipelines. Specifically, teams can leverage TE-CDE's internal uncertainty estimates to flag and defer ambiguous patient trajectories—such as severe or volatile cases—to human experts, since a small fraction of uncertain cases accounts for the majority of overall estimation errors. Next steps should focus on pilot testing the framework in clinical decision-support environments and conducting further methodological research into complex observational challenges such as informative missingness and unobserved hidden confounders.

Readers should interpret these results with the understanding that evaluations were primarily conducted in synthetic and semi-synthetic tumor growth environments because counterfactual outcomes cannot be observed directly in real-world clinical datasets. The model also operates under standard causal assumptions, including unconfoundedness and positivity. While confidence in TE-CDE's mathematical stability and relative performance is high, real-world deployment requires cautious calibration against potential unmeasured confounders and broader clinical factors like drug toxicity.

Seedat et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations

Abstract

Estimating the effect of interventions – alternatively referred to as treatments or actions – is central to decision-making in many sequential settings, ranging from healthcare to economics to robot control. Motivated primarily by the analysis of longitudinal data arising in medicine, we here consider continuous-time 'treatment' trajectories which are observed intermittently over time: Each observation thus jointly reflects two ongoing processes of interest - the evolution of individual patient covariates over time, and delays in the measurement process driven by events (e.g., visits doctors). Our contributions are threefold: First, we formalize 'continuous-time potential outcomes', extending counterfactual theory traditionally developed for fixed treatment times; Second, we frame causal inference in terms of differential equations whose solution defines outcome paths; Third, we introduce Neural CDEs as powerful means towards their estimation.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Formulation
  • 4. Treatment Effect Controlled Differential Equation
  • 5. Experiments
  • 5.1. Modeling tumor growth under general observation patterns
  • 5.2. Impact of time-dependent confounding across varying sampling intensities
  • 5.3. Treatment-conditioned sampling
  • 5.4. Forecasting at additional time horizons
  • 5.5. Treatment selection
  • 5.6. Additional experiments
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Notation Summary
  • B. Extended Related Work
  • Treatment effect estimation with static data
  • Treatment effect estimation with longitudinal, time-varying data
  • Neural differential equations
  • C. Simulation Environment
  • C.1. Pharmacokinetic-Pharmacodynamic Model
  • C.2. Cancer Staging
  • D. TE-CDE
  • D.1. Additional Methodological Details
  • D.2. TE-CDE Implementation Details
  • E. Implementation details: Benchmark algorithms
  • E.1. Counterfactual Recurrent Network (CRN)
  • E.2. Recurrent Marginal Structural Network (RMSN)
  • E.3. Gaussian Process (GP)
  • F. Additional Experiments
  • F.1. Uncertainty Based Exclusion of Estimates
  • F.2. Data Efficiency Benchmark for Different Methods
  • F.3. Latent Representation and Discovery
  • F.3.1. TREATMENT INVARIANT LATENT REPRESENTATION
  • F.3.2. DISCOVERY AND INSIGHTS FROM LATENT REPRESENTATIONS
  • F.4. Linear f and g
  • F.5. Other assessment
  • F.6. Impact of Time-Dependent Confounding Across Varying Sampling Intensities
  • F.7. Treatment-Conditioned Sampling
  • F.8. Forecasting at Additional Time Horizons

Knowls

  1. Knowl 1 — TE-CDE represents patient history and hypothetical treatment paths continuously

    model/method

    Treatment Effect Neural Controlled Differential Equation (TE-CDE) estimates future potential outcomes from irregularly observed histories. For one patient, let Xs∈RdX_s\in\mathbb{R}^d be covariates, As∈{0,1}A_s\in\{0,1\} treatment, and Ys∈RY_s\in\mathbb{R} outcome at time ss; let the observed history be available through time tt, and let the latent state be Zs∈RlZ_s\in\mathbb{R}^l. The encoder embeds the initial observation at time t0t_0 and evolves the latent state continuously, driven by the interpolated covariate, treatment, and outcome paths:

    Zt0=gη(Xt0,At0,Yt0),dZs=fθ(Zs) dUs,Us=(Xs,As,Ys),t0<s≤t.Z_{t_0}=g_\eta(X_{t_0},A_{t_0},Y_{t_0}),\qquad dZ_s=f_\theta(Z_s)\,dU_s,\qquad U_s=(X_s,A_s,Y_s),\quad t_0<s\leq t.

    Here gηg_\eta is a neural network from Rd+2\mathbb{R}^{d+2} to Rl\mathbb{R}^l, and fθf_\theta is a learned vector field mapping latent states to matrices that act on the increments of the control path UU. For a user-specified future treatment path AsaA^a_s over t<s≤t′t<s\leq t', the decoder evolves a counterfactual latent path and maps it to the predicted outcome:

    dZsa=fϕ(Zsa,Asa) dAsa,Zta=Zt,Y^sa=hν(Zsa),t<s≤t′.dZ_s^a=f_\phi(Z_s^a,A_s^a)\,dA_s^a,\qquad Z_t^a=Z_t,\qquad \widehat{Y}_s^a=h_\nu(Z_s^a),\quad t<s\leq t'.

    The decoder vector field fϕf_\phi and outcome head hνh_\nu are neural networks; hνh_\nu outputs the outcome prediction. This construction yields latent states and potential-outcome predictions at arbitrary times, rather than only at a shared discrete grid. TE-CDE uses causal linear interpolation in its implementation; the paper also identifies rectilinear interpolation as causal. Its encoder and decoder use two-layer vector-field networks with hidden size 128 and latent dimension 8, and an adaptive Dormand–Prince (dopri5) solver. The solver is informed of treatment-path jumps so integration can align with discontinuities.

  2. Knowl 2 — Causal identification requires consistency, treatment overlap, and sequential randomization

    assumption

    For a patient with observed history Ft\mathcal{F}_t, let Yt′(a)Y_{t'}(a) denote the outcome at future time t′>tt'>t under a specified treatment path aa on [t,t′][t,t']. TE-CDE targets the conditional potential outcome E[Yt′(a)∣Ft]\mathbb{E}[Y_{t'}(a)\mid\mathcal{F}_t]. The paper states that potential outcomes are identifiable from observational data under these three assumptions:

    • Consistency: when the observed treatment path is A=aA=a, the factual outcome path equals the potential outcome path under that treatment, Y(a)=YY(a)=Y.
    • Overlap: the treatment-assignment intensity given observed history is stochastic and satisfies 0<λ(t∣Ft)<10<\lambda(t\mid\mathcal{F}_t)<1 at each relevant time. The intensity is the limiting conditional rate of a treatment-process increment: λ(t∣Ft)=lim⁡δt→0P(At+δt−At=1∣Ft)/δt\lambda(t\mid\mathcal{F}_t)=\lim_{\delta t\to0}\mathbb{P}(A_{t+\delta t}-A_t=1\mid\mathcal{F}_t)/\delta t.
    • Continuous-time sequential randomization: the treatment intensity conditional on the observed filtration is unchanged when that filtration is augmented with future outcomes: λ(t∣Ft)=λ(t∣Ft∪{σ(Ys):s>t})\lambda(t\mid\mathcal{F}_t)=\lambda(t\mid\mathcal{F}_t\cup\{\sigma(Y_s):s>t\}). In the paper’s interpretation, current observed information is sufficient to estimate future counterfactuals without bias from future outcomes.

    Here Ft\mathcal{F}_t contains the individual’s observable covariate, treatment, and outcome events up to time tt, and σ(Ys)\sigma(Y_s) denotes the information generated by the outcome at future time ss.

  3. Knowl 3 — Adversarial treatment prediction is used to balance the latent history representation

    model/method

    TE-CDE addresses time-dependent confounding by encouraging its latent state ZsZ_s not to reveal the observed treatment assignment at time ss. An outcome head predicts Y^s=hν(Zs)\widehat{Y}_s=h_\nu(Z_s), while a treatment classifier predicts p^s=hα(Zs)\widehat{p}_s=h_\alpha(Z_s), the probability of treatment As=1A_s=1. For kk observed times sjs_j in a prediction window, the outcome loss is mean squared error and the treatment loss is binary cross-entropy:

    L(y)=1k∑j=1k(Ysj−Y^sj)2,L(a)=−1k∑j=1k[Asjlog⁡p^sj+(1−Asj)log⁡(1−p^sj)].\mathcal{L}^{(y)}=\frac{1}{k}\sum_{j=1}^k(Y_{s_j}-\widehat{Y}_{s_j})^2,\qquad \mathcal{L}^{(a)}=-\frac{1}{k}\sum_{j=1}^k\left[A_{s_j}\log\widehat{p}_{s_j}+(1-A_{s_j})\log(1-\widehat{p}_{s_j})\right].

    The representation is trained using the objective L=1n∑i=1n(Li(y)−μLi(a))\mathcal{L}=\frac{1}{n}\sum_{i=1}^n(\mathcal{L}^{(y)}_i-\mu\mathcal{L}^{(a)}_i), where nn is the number of individuals and μ>0\mu>0 controls the trade-off. The treatment classifier learns to predict assignment, while the negative treatment-loss term makes the representation difficult to classify by treatment; the intended result is a representation balanced across treatment groups, ideally with P(Zs∣As=0)=P(Zs∣As=1)P(Z_s\mid A_s=0)=P(Z_s\mid A_s=1). In the experiments, μ\mu starts at zero and increases exponentially by epoch over the range [0,1][0,1]. The paper reports that omitting this adversarial term degrades counterfactual accuracy, increasingly so at higher confounding levels.

  4. Knowl 4 — The simulation combines tumor-growth dynamics, confounded treatment assignment, and irregular observation

    model/method

    The paper’s controllable synthetic environment generates lung-cancer tumor-volume trajectories under chemotherapy and radiotherapy. Tumor volume V(t)V(t) changes according to a pharmacokinetic–pharmacodynamic model:

    dV(t)dt=[ρlog⁡(KV(t))−βcC(t)−(αrd(t)+βrd(t)2)+et]V(t).\frac{dV(t)}{dt}=\left[\rho\log\left(\frac{K}{V(t)}\right)-\beta_c C(t)-\left(\alpha_r d(t)+\beta_r d(t)^2\right)+e_t\right]V(t).

    Here tt is time after diagnosis; KK is carrying capacity; ρ\rho is the growth parameter; C(t)C(t) is chemotherapy concentration; d(t)d(t) is radiotherapy dose; βc,αr,βr\beta_c,\alpha_r,\beta_r are treatment-effect parameters; and ete_t is Gaussian growth noise. Chemotherapy concentration decays with a one-day half-life, and administered doses are 5.0 mg/m3^3 of vinblastine; radiotherapy is given in 2.0 Gy fractions. Chemotherapy and radiotherapy assignments are Bernoulli variables with probabilities pc(t)p_c(t) and pr(t)p_r(t) that depend on average tumor diameter Dˉ(t)\bar D(t):

    pc(t)=σ(γcDmax⁡(Dˉ(t)−θc)),pr(t)=σ(γrDmax⁡(Dˉ(t)−θr)),p_c(t)=\sigma\left(\frac{\gamma_c}{D_{\max}}(\bar D(t)-\theta_c)\right),\qquad p_r(t)=\sigma\left(\frac{\gamma_r}{D_{\max}}(\bar D(t)-\theta_r)\right),

    where σ\sigma is the logistic function, Dmax⁡=13D_{\max}=13 cm, and θc=θr=Dmax⁡/2\theta_c=\theta_r=D_{\max}/2. Increasing γc\gamma_c or γr\gamma_r makes treatment assignment more dependent on tumor diameter and therefore increases time-dependent confounding. The possible regimens are no treatment, chemotherapy, radiotherapy, and both treatments.

    Observation times are generated by a self-exciting Hawkes process whose intensity depends on cancer stage ii and past observation times since the most recent stage change τ\tau:

    λ(t,st=i)=0.01κi+∑τ<tm<te−2(t−tm).\lambda(t,s_t=i)=0.01\kappa^i+\sum_{\tau<t_m<t}e^{-2(t-t_m)}.

    Here st=is_t=i is the patient’s stage, tmt_m are prior observation times, and κ≥1\kappa\geq1 controls the increase in baseline observation intensity with stage. The simulated stages map S1A, S1B, S2, and S3 to i=0,1,2,3i=0,1,2,3. Thus observation frequency can vary both within a patient’s trajectory as the disease stage changes and between patients; the experiments also vary κ\kappa by treatment group. Tumor diameter is obtained from volume under a spherical-tumor assumption, V=π6D3V=\frac{\pi}{6}D^3.

  5. Knowl 5 — TE-CDE has the lowest reported counterfactual error across confounding and sampling settings

    empirical result

    The primary simulation comparison evaluates normalized RMSE for tumor-volume counterfactuals, normalizing RMSE by the maximum tumor volume Vmax⁡=1150V_{\max}=1150 cm3^3. Treatment confounding was varied through γc=γr=γ∈{2,4,6,8,10}\gamma_c=\gamma_r=\gamma\in\{2,4,6,8,10\}, and stage-dependent sampling through κ∈{1,5,10}\kappa\in\{1,5,10\}. Each experiment used 10,000 training patients, 1,000 validation patients, and 10,000 test patients. CRN and RMSN, which require regular discrete inputs, were evaluated after equal-grid interpolation and imputation. The table gives the reported normalized RMSE values (mean ±\pm reported uncertainty); lower is better. TE-CDE is best in every displayed setting, and its advantage over CRN widens with confounding. At γ=10\gamma=10, the paper reports a 36% RMSE reduction versus CRN, the next-best method. Removing adversarial training also worsens TE-CDE performance as confounding increases.

    κ\kappa Model γ=2\gamma=2 γ=4\gamma=4 γ=6\gamma=6 γ=8\gamma=8 γ=10\gamma=10
    1 TE-CDE 0.96±0.010.96\pm0.01 1.55±0.021.55\pm0.02 1.83±0.081.83\pm0.08 2.26±0.062.26\pm0.06 3.07±0.353.07\pm0.35
    1 CRN 1.07±0.091.07\pm0.09 1.63±0.101.63\pm0.10 2.11±0.292.11\pm0.29 3.23±0.363.23\pm0.36 4.49±0.384.49\pm0.38
    1 RMSN 1.92±0.071.92\pm0.07 2.42±0.112.42\pm0.11 2.66±0.152.66\pm0.15 4.00±0.124.00\pm0.12 4.16±0.114.16\pm0.11
    1 TE-CDE (μ=0\mu=0) 1.04±0.161.04\pm0.16 1.77±0.111.77\pm0.11 2.14±0.342.14\pm0.34 3.00±0.213.00\pm0.21 4.26±0.514.26\pm0.51
    1 GP 1.74±0.761.74\pm0.76 2.80±0.932.80\pm0.93 3.74±0.403.74\pm0.40 6.15±1.146.15\pm1.14 6.97±1.026.97\pm1.02
    5 TE-CDE 0.88±0.010.88\pm0.01 1.31±0.021.31\pm0.02 2.32±0.082.32\pm0.08 2.62±0.062.62\pm0.06 3.02±0.353.02\pm0.35
    5 CRN 1.27±0.151.27\pm0.15 1.78±0.141.78\pm0.14 3.08±0.213.08\pm0.21 4.11±0.384.11\pm0.38 4.90±0.314.90\pm0.31
    5 RMSN 2.69±0.122.69\pm0.12 2.92±0.242.92\pm0.24 3.16±0.183.16\pm0.18 4.265±0.134.265\pm0.13 5.40±0.165.40\pm0.16
    5 TE-CDE (μ=0\mu=0) 0.88±0.070.88\pm0.07 1.68±0.051.68\pm0.05 2.97±0.202.97\pm0.20 4.00±0.224.00\pm0.22 4.64±0.594.64\pm0.59
    5 GP 9.08±1.539.08\pm1.53 14.19±1.7314.19\pm1.73 24.47±2.8924.47\pm2.89 39.89±1.0239.89\pm1.02 42.99±2.7642.99\pm2.76
    10 TE-CDE 0.78±0.010.78\pm0.01 1.16±0.021.16\pm0.02 1.93±0.081.93\pm0.08 2.67±0.062.67\pm0.06 3.03±0.353.03\pm0.35
    10 CRN 1.31±0.101.31\pm0.10 1.51±0.111.51\pm0.11 3.10±0.263.10\pm0.26 4.28±0.294.28\pm0.29 4.84±0.314.84\pm0.31
    10 RMSN 1.60±0.111.60\pm0.11 2.04±0.122.04\pm0.12 3.67±0.223.67\pm0.22 4.62±0.274.62\pm0.27 5.41±0.315.41\pm0.31
    10 TE-CDE (μ=0\mu=0) 0.85±0.060.85\pm0.06 1.23±0.071.23\pm0.07 2.48±0.062.48\pm0.06 3.49±0.513.49\pm0.51 4.94±0.594.94\pm0.59
    10 GP 9.41±0.989.41\pm0.98 14.43±0.9514.43\pm0.95 24.74±1.9724.74\pm1.97 40.21±0.9340.21\pm0.93 43.78±3.1143.78\pm3.11

    The GP is a continuous-time comparator; it performs especially poorly at larger κ\kappa and confounding in these experiments. The table also shows that the no-adversary variant is less accurate than full TE-CDE in most higher-confounding settings.

  6. Knowl 6 — TE-CDE retains advantages for treatment-conditioned sampling, longer forecasts, and treatment choice

    empirical result

    In a treatment-conditioned sampling experiment, the observation-intensity setting was κ=10\kappa=10 for treated patients and κ=1\kappa=1 for untreated patients, with confounding fixed at γ=4\gamma=4. The table reports normalized RMSE (%) overall and separately by treatment group. TE-CDE has the lowest error in all three comparisons. The authors note that 77% of untreated patients remained at cancer stage S1A, contributing to lower errors across methods for that subgroup.

    Model Overall Treated Untreated
    TE-CDE 1.18±0.051.18\pm0.05 1.56±0.061.56\pm0.06 0.22±0.020.22\pm0.02
    CRN 1.57±0.061.57\pm0.06 1.97±0.051.97\pm0.05 0.64±0.080.64\pm0.08
    RMSN 3.06±0.093.06\pm0.09 3.12±0.073.12\pm0.07 2.83±0.082.83\pm0.08

    For forecasts five observation steps ahead with κ=10\kappa=10, TE-CDE also outperformed CRN and RMSN at every tested confounding level; at γ=10\gamma=10, its RMSE was 40% lower than CRN’s. In a treatment-selection evaluation using the same five-step horizon, the chosen regimen was defined as the one minimizing predicted tumor volume. TE-CDE selected the optimal regimen more accurately than both recurrent benchmarks at all tested γ\gamma values; its absolute accuracy advantage was about 4 percentage points at γ=4\gamma=4 and about 10 points at γ=10\gamma=10.

  7. Knowl 7 — TE-CDE’s counterfactual error degrades less with reduced training data

    data/table

    The data-efficiency experiment compared each method’s RMSE at smaller training-set sizes with its own RMSE when trained on 10,000 patients. Confounding was fixed at γ=2\gamma=2 and sampling intensity at κ=10\kappa=10. The values are percent performance reduction relative to the 10,000-patient result, so lower indicates greater data efficiency. TE-CDE has the smallest degradation at both sample sizes, with the clearest advantage over RMSN.

    Training patients Model RMSE increase relative to 10,000 patients
    5,000 TE-CDE 4.8%
    5,000 CRN 6.4%
    5,000 RMSN 25.6%
    1,000 TE-CDE 17.1%
    1,000 CRN 18.9%
    1,000 RMSN 57.2%

    The authors suggest that a continuous latent trajectory may require fewer examples than discrete recurrent models, which must learn dynamics at discrete time steps.

  8. Knowl 8 — MC dropout provides a useful ranking of counterfactual uncertainty

    empirical result

    The paper estimates epistemic uncertainty in TE-CDE by applying Monte Carlo dropout to the trained model. For NN dropout draws, let y^(i)\widehat y^{(i)} be the counterfactual prediction from draw ii; the estimated prediction and uncertainty are the sample mean and variance:

    y~=1N∑i=1Ny^(i),σ^2=Var⁡i=1,…,N(y^(i)).\widetilde y=\frac{1}{N}\sum_{i=1}^{N}\widehat y^{(i)},\qquad \widehat{\sigma}^{2}=\operatorname{Var}_{i=1,\ldots,N}(\widehat y^{(i)}).

    The variance over predictions is used to rank estimates for possible review or deferral. In the simulation, excluding predictions with the greatest estimated uncertainty lowered RMSE on the retained cases, unlike random exclusion, and the uncertainty-based sparsification curve was close to the oracle ordering based on true errors. The reported area under the sparsification error curve was 2.17 for TE-CDE uncertainty and 455.92 for random ordering. The paper reports that fewer than 10% of samples accounted for the majority of total estimation error in this analysis.

  9. Knowl 9 — Counterfactual validity is limited by assumptions and the synthetic evaluation setting

    limitation

    TE-CDE’s causal interpretation relies on the stated consistency, overlap, and continuous-time sequential-randomization assumptions; the paper’s formulation assumes no hidden confounding. The authors also caution that the method only partially addresses the complexities of irregular clinical sampling, particularly informative sampling. The principal performance evaluation uses synthetic counterfactuals generated by a tumor-growth model. The MIMIC-III experiment predicts factual outcomes and therefore cannot establish accuracy for unobserved counterfactual outcomes. Finally, the predictions do not account for trade-offs involving treatment side effects, and the authors recommend clinical use within a human-in-the-loop process rather than relying on estimates alone.

Coverage note — Deliberately omitted the appendix’s latent-space visualization and subgroup-discovery analyses, the linear-vector-field ablation, and the MIMIC-III factual-only comparison because they are secondary exploratory analyses or do not evaluate counterfactual accuracy.

References

  1. 1.Alaa, A. M. and van der Schaar, M. Bayesian inference of individualized treatment effects using multi-task gaussian processes. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 3427–3435, 2017.
  2. 2.Alaa, A. M., Hu, S., and Schaar, M. Learning from clinical judgments: Semi-markov-modulated marked hawkes processes for risk prognosis. In International Conference on Machine Learning, pp. 60–69. PMLR, 2017.
  3. 3.Bao, Y., Kuang, Z., Peissig, P., Page, D., and Willett, R. Hawkes process modeling of adverse drug reactions with longitudinal observational data. In Machine learning for healthcare conference, pp. 177–190. PMLR, 2017.
  4. 4.Bartsch, H., Dally, H., Popanda, O., Risch, A., and Schmezer, P. Genetic risk profiles for cancer susceptibility and therapy response. Cancer Prevention, pp. 19–36, 2007.
  5. 5.Bellot, A. and van der Schaar, M. Policy analysis using synthetic controls in continuous-time. In Proceedings of 38th International Conference on Machine Learning (ICML 2021), 2021.
  6. 6.Bica, I., Alaa, A., and Van Der Schaar, M. Time series deconfounder: Estimating treatment effects over time in the presence of hidden confounders. In Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research. PMLR, 2020a.
  7. 7.Bica, I., Alaa, A. M., Jordon, J., and van der Schaar, M. Estimating counterfactual treatment outcomes over time through adversarially balanced representations. In Proceedings of 8th International Conference on Learning Representations (ICLR 2020), 2020b.
  8. 8.Bica, I., Alaa, A. M., Lambert, C., and Van Der Schaar, M. From real-world patient data to individualized treatment effects using machine learning: current and future methods to address underlying challenges. Clinical Pharmacology & Therapeutics, 109(1):87–100, 2021.
  9. 9.Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. Neural ordinary differential equations. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 6572–6583, 2018.
  10. 10.De Brouwer, E., Simm, J., Arany, A., and Moreau, Y. Gru-ode-bayes: Continuous modeling of sporadically-observed time series. In 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019.
  11. 11.Gal, Y. and Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pp. 1050–1059. PMLR, 2016.
  12. 12.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1):2096–2030, 2016.
  13. 13.Geng, C., Paganetti, H., and Grassberger, C. Prediction of treatment response for combined chemo-and radiation therapy for non-small cell lung cancer patients using a bio-mathematical model. Scientific reports, 7(1):1–12, 2017.
  14. 14.Gershenwald, J. E., Scolyer, R. A., Hess, K. R., Sondak, V. K., Long, G. V., Ross, M. I., Lazar, A. J., Faries, M. B., Kirkwood, J. M., McArthur, G. A., et al. Melanoma staging: evidence-based changes in the american joint committee on cancer eighth edition cancer staging manual. CA: a cancer journal for clinicians, 67(6):472–492, 2017.
  15. 15.GPy. GPy: A gaussian process framework in python. http://github.com/SheffieldML/GPy, since 2012.
  16. 16.Gwak, D., Sim, G., Poli, M., Massaroli, S., Choo, J., and Choi, E. Neural ordinary differential equations for intervention modeling. arXiv preprint arXiv:2010.08304, 2020.
  17. 17.Hawkes, A. G. Spectra of some self-exciting and mutually exciting point processes. Biometrika, 58(1):83–90, 04 1971. doi: 10.1093/biomet/58.1.83. URL https://doi.org/10.1093/biomet/58.1.83.
  18. 18.Hawkes, A. G. and Oakes, D. A cluster process representation of a self-exciting process. Journal of Applied Probability, 11(3):493–503, 1974.
  19. 19.Hernán, M. Á., Brumback, B., and Robins, J. M. Marginal structural models to estimate the causal effect of zidovudine on the survival of hiv-positive men. Epidemiology, pp. 561–570, 2000.
  20. 20.Hill, J. L. Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics, 20(1):217–240, 2011.
  21. 21.Ilg, E., Cicek, O., Galesso, S., Klein, A., Makansi, O., Hutter, F., and Brox, T. Uncertainty estimates and multi-hypotheses networks for optical flow. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 652–667, 2018.
  22. 22.Jesson, A., Mindermann, S., Shalit, U., and Gal, Y. Identifying causal-effect inference failure with uncertainty-aware models. Advances in Neural Information Processing Systems, 33, 2020.
  23. 23.Johansson, F., Shalit, U., and Sontag, D. Learning representations for counterfactual inference. In International conference on machine learning, pp. 3020–3029. PMLR, 2016.
  24. 24.Johansson, F. D., Sontag, D., and Ranganath, R. Support and invertibility in domain-invariant representations. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 527–536. PMLR, 2019.
  25. 25.Johansson, F. D., Shalit, U., Kallus, N., and Sontag, D. Generalization bounds and representation learning for estimation of potential outcomes and causal effects. arXiv preprint arXiv:2001.07426, 2020.
  26. 26.Kidger, P., Morrill, J., Foster, J., and Lyons, T. Neural controlled differential equations for irregular time series. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), 2020.
  27. 27.Lee, Y., Lim, K. W., and Ong, C. S. Hawkes processes with stochastic excitations. In International Conference on Machine Learning, pp. 79–88. PMLR, 2016.
  28. 28.Li, S. and Fu, Y. Matching on balanced nonlinear representations for treatment effects estimation. In NIPS, 2017.
  29. 29.Lim, B., Alaa, A. M., and van der Schaar, M. Forecasting treatment responses over time using recurrent marginal structural networks. NeurIPS, 18:7483–7493, 2018.
  30. 30.Lok, J. Statistical modeling of causal effects in continuous time. Annals of statistics, 36(3):1464–1507, 2008.
  31. 31.Lyons, T. J., Caruana, M., and Lévy, T. Differential equations driven by rough paths. Springer, 2007.
  32. 32.Mansournia, M. A., Danaei, G., Forouzanfar, M. H., Mahmoodi, M., Jamali, M., Mansournia, N., and Mohammad, K. Effect of physical activity on functional performance and knee pain in patients with osteoarthritis: analysis with marginal structural models. Epidemiology, pp. 631–640, 2012.
  33. 33.Mansournia, M. A., Etminan, M., Danaei, G., Kaufman, J. S., and Collins, G. Handling time varying confounding in observational research. bmj, 359, 2017.
  34. 34.Morrill, J., Kidger, P., Yang, L., and Lyons, T. Neural controlled differential equations for online prediction tasks. arXiv preprint arXiv:2106.11028, 2021.
  35. 35.Platt, R. W., Schisterman, E. F., and Cole, S. R. Time-modified confounding. American journal of epidemiology, 170(6):687–694, 2009.
  36. 36.Robins, J. A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9-12):1393–1512, 1986.
  37. 37.Robins, J., Hernán, M., and Brumback, B. Marginal structural models and causal inference in epidemiology. Epidemiology (Cambridge, Mass.), 11(5):550–560, 2000.
  38. 38.Robins, J. M. Correcting for non-compliance in randomized trials using structural nested mean models. Communications in Statistics-Theory and methods, 23(8):2379–2412, 1994.
  39. 39.Robins, J. M. Causal inference from complex longitudinal data. In Latent variable modeling and applications to causality, pp. 69–117. Springer, 1997.
  40. 40.Rosenbaum, P. R. and Rubin, D. B. The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1):41–55, 1983.
  41. 41.Roueff, F., Von Sachs, R., and Sansonnet, L. Locally stationary hawkes processes. Stochastic Processes and their Applications, 126(6):1710–1743, 2016.
  42. 42.Rubanova, Y., Chen, R. T., and Duvenaud, D. Latent ODEs for irregularly-sampled time series. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pp. 5320–5330, 2019.
  43. 43.Ryalen, P. C., Stensrud, M. J., and Røysland, K. The additive hazard estimator is consistent for continuous-time marginal structural models. Lifetime data analysis, 25(4):611–638, 2019.
  44. 44.Saarela, O. and Liu, Z. A flexible parametric approach for estimating continuous-time inverse probability of treatment and censoring weights. Statistics in medicine, 35(23):4238–4251, 2016.
  45. 45.Schisterman, E. F., Cole, S. R., and Platt, R. W. Overadjustment bias and unnecessary adjustment in epidemiologic studies. Epidemiology (Cambridge, Mass.), 20(4):488, 2009.
  46. 46.Schulam, P. and Saria, S. Reliable decision support using counterfactual models. Advances in Neural Information Processing Systems, 30:1697–1708, 2017.
  47. 47.Shalit, U., Johansson, F. D., and Sontag, D. Estimating individual treatment effect: generalization bounds and algorithms. In International Conference on Machine Learning, pp. 3076–3085. PMLR, 2017.
  48. 48.Tatekawa, K., Iwata, H., Kawaguchi, T., Ishikura, S., Baba, F., Otsuka, S., Miyakawa, A., Iwana, M., and Shibamoto, Y. Changes in volume of stage i non-small-cell lung cancer during stereotactic body radiotherapy. Radiation oncology, 9(1):1–5, 2014.
  49. 49.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  50. 50.Wager, S. and Athey, S. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113(523):1228–1242, 2018.
  51. 51.Wang, Y. and Blei, D. M. The blessings of multiple causes. Journal of the American Statistical Association, 114(528):1574–1596, 2019.
  52. 52.Yao, L., Li, S., Li, Y., Huai, M., Gao, J., and Zhang, A. Representation learning for treatment effect estimation from observational data. Advances in Neural Information Processing Systems, 31, 2018.
  53. 53.Yoon, J., Jordon, J., and Van Der Schaar, M. Ganite: Estimation of individualized treatment effects using generative adversarial nets. In International Conference on Learning Representations, 2018.
  54. 54.Zhang, H., Gao, X., Unterman, J., and Arodz, T. Approximation capabilities of neural ODEs and invertible residual networks. In International Conference on Machine Learning, pp. 11086–11095. PMLR, 2020.

Citation

MLA
Seedat, N., et al. “Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations”. International Conference on Machine Learning, vol. 162, 2022, pp. 19497–521, https://proceedings.mlr.press/v162/seedat22b.html.
APA
Seedat, N., Imrie, F., Bellot, A., Qian, Z., & Schaar, M. van . der . (2022). Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations. International Conference on Machine Learning, 162, 19497–19521. https://proceedings.mlr.press/v162/seedat22b.html
Chicago
Seedat, N., F. Imrie, A. Bellot, Z. Qian, and M. van . der . Schaar. 2022. “Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations”. International Conference on Machine Learning 162: 19497–521. https://proceedings.mlr.press/v162/seedat22b.html.
Harvard
Seedat, N. et al. (2022) “Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations”, International Conference on Machine Learning. PMLR, pp. 19497–19521. Available at: https://proceedings.mlr.press/v162/seedat22b.html.
Vancouver
1. Seedat N, Imrie F, Bellot A, Qian Z, Schaar M van der (2022) Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations. In: International Conference on Machine Learning. PMLR, pp 19497–19521

BibTeX

@InProceedings{pmlr-v162-seedat22b,
  title = 	 {Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations},
  author =       {Seedat, Nabeel and Imrie, Fergus and Bellot, Alexis and Qian, Zhaozhi and van der Schaar, Mihaela},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {19497--19521},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/seedat22b/seedat22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/seedat22b.html},
  abstract = 	 {Estimating counterfactual outcomes over time has the potential to unlock personalized healthcare by assisting decision-makers to answer "what-if" questions. Existing causal inference approaches typically consider regular, discrete-time intervals between observations and treatment decisions and hence are unable to naturally model irregularly sampled data, which is the common setting in practice. To handle arbitrary observation patterns, we interpret the data as samples from an underlying continuous-time process and propose to model its latent trajectory explicitly using the mathematics of controlled differential equations. This leads to a new approach, the Treatment Effect Neural Controlled Differential Equation (TE-CDE), that allows the potential outcomes to be evaluated at any time point. In addition, adversarial training is used to adjust for time-dependent confounding which is critical in longitudinal settings and is an added challenge not encountered in conventional time series. To assess solutions to this problem, we propose a controllable simulation environment based on a model of tumor growth for a range of scenarios with irregular sampling reflective of a variety of clinical scenarios. TE-CDE consistently outperforms existing approaches in all scenarios with irregular sampling.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/