Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations
Nabeel SeedatFergus ImrieAlexis BellotZhaozhi QianMihaela van der Schaar
Proposes a continuous-time causal inference framework using neural controlled differential equations and adversarial training to reliably estimate individual treatment effects from irregularly sampled longitudinal data subject to time-dependent confounding.
Real-world medical decision-making relies heavily on answering "what-if" questions to determine the optimal type and timing of treatments. While observational healthcare data offers a rich resource for estimating counterfactual outcomes over time, most existing causal inference methods assume data arrives at fixed, regular intervals across all patients. In clinical practice, however, patient records are sampled irregularly due to missed visits, differing monitoring protocols, and varying disease severity. Furthermore, past treatments and patient health trajectories create time-dependent confounding, a major source of bias that standard time-series techniques cannot resolve.
The main objective of the article is to develop and evaluate a causal inference framework that can reliably estimate counterfactual outcomes at any continuous point in time from irregularly sampled observational data with time-dependent confounding.
To achieve this, the article introduces the Treatment Effect Neural Controlled Differential Equation (TE-CDE). Unlike traditional recurrent neural network models that force irregular observations into artificial discrete bins, TE-CDE models a patient's underlying latent health trajectory as a continuous process governed by controlled differential equations. The framework incorporates domain adversarial training to remove time-dependent confounding bias by making the latent representations invariant to treatment assignments. To validate performance, the approach was tested in a controllable simulation environment modeling lung cancer tumor growth under chemotherapy and radiotherapy regimes, using a Hawkes point process to emulate realistic, state-dependent clinical sampling patterns across datasets of 10,000 patients.
The analysis yielded several critical findings. First, TE-CDE consistently achieved the lowest counterfactual estimation error across all irregular sampling regimes and levels of time-dependent confounding, reducing root mean square error by approximately 36% compared to the strongest discrete baseline at high confounding levels. Second, when forecasting multiple steps into the future, TE-CDE demonstrated a 40% reduction in estimation error over competing methods at high confounding. Third, improved estimation accuracy translated directly into superior decision-making, increasing optimal treatment selection accuracy by about 4% under moderate confounding and up to roughly 10% under high confounding. Fourth, incorporating domain adversarial training proved essential, as removing it significantly degraded accuracy under severe confounding. Finally, TE-CDE demonstrated superior data efficiency, experiencing only a 17.1% error degradation when training data was reduced tenfold, compared to 18.9% and 57.2% for baseline methods.
These findings imply that treating patient health as an inherently continuous trajectory avoids the errors introduced by discrete interpolation and data imputation. This shift provides more dependable decision support for individualized treatment planning, helping clinicians avoid suboptimal therapy choices and potentially reducing treatment costs and adverse health risks. Notably, the results demonstrate that standard Gaussian process and recurrent models struggle significantly in irregular longitudinal settings with high confounding, establishing continuous-time neural differential equations as a more appropriate paradigm.
Organizations evaluating continuous-time causal models should implement them within "human-in-the-loop" clinical workflows rather than fully automated pipelines. Specifically, teams can leverage TE-CDE's internal uncertainty estimates to flag and defer ambiguous patient trajectories—such as severe or volatile cases—to human experts, since a small fraction of uncertain cases accounts for the majority of overall estimation errors. Next steps should focus on pilot testing the framework in clinical decision-support environments and conducting further methodological research into complex observational challenges such as informative missingness and unobserved hidden confounders.
Readers should interpret these results with the understanding that evaluations were primarily conducted in synthetic and semi-synthetic tumor growth environments because counterfactual outcomes cannot be observed directly in real-world clinical datasets. The model also operates under standard causal assumptions, including unconfoundedness and positivity. While confidence in TE-CDE's mathematical stability and relative performance is high, real-world deployment requires cautious calibration against potential unmeasured confounders and broader clinical factors like drug toxicity.
- Paper: Estimating individual treatment effect: generalization bounds and algorithms, Uri Shalit et al. (2016). Its Counterfactual Regression framework formalizes balancing treatment groups in learned representations, a key precursor to the source’s adversarial approach to time-dependent confounding.
No sufficiently relevant recommendations were found.
