Estimating individual treatment effect: generalization bounds and algorithms

Uri ShalitFredrik D. JohanssonDavid Sontag

article2016ICML1,311 citations

Establishes generalization error bounds for individual treatment effect estimation and introduces algorithms that learn balanced representations by minimizing distribution discrepancies between treated and control groups.

Listen

Estimating individual-level causal effects from observational data is essential for high-stakes decision-making in domains such as precision medicine, public policy, and personalized education. However, observational data suffers from confounding and imbalance: we only observe the outcome for the action actually taken, leaving the counterfactual outcome unobserved. When treated and control populations differ substantially, standard machine learning methods fail to generalize across groups, leading to high-variance and biased predictions for individual treatment effects.

The article aims to establish a theoretical error bound for estimating individual treatment effects and to introduce a regularized representation-learning framework that minimizes this error in observational settings.

The authors developed a mathematical framework that upper-bounds individual treatment effect error by combining standard supervised prediction loss with a distributional distance penalty between treated and control groups. They implemented this approach using Counterfactual Regression, a deep neural network architecture that learns a shared, balanced data representation with separate prediction heads for treated and control outcomes. The framework incorporates distribution balancing metrics based on Integral Probability Metrics, specifically Wasserstein distance and Maximum Mean Discrepancy. The authors evaluated their method across within-sample and out-of-sample tasks using a benchmark semi-synthetic dataset based on infant health (747 subjects, 25 covariates) and a real-world employment dataset (3,212 subjects) combining randomized and observational data.

The evaluation revealed several key findings:

  1. The Counterfactual Regression framework substantially outperformed standard baselines on individual treatment effect estimation. On the infant health benchmark, the proposed models reduced out-of-sample error to approximately 0.76–0.78, compared to 2.1–2.3 for traditional balancing neural networks and tree-based methods, and 4.1–6.6 for nearest neighbors and random forests (a reduction of roughly 60% to over 80%).
  2. Incorporating distributional balancing penalties consistently improved performance over unpenalized neural models, maintaining a distinct performance advantage even as data imbalance between treatment groups intensified.
  3. On the real-world employment dataset, non-linear Counterfactual Regression models achieved the lowest out-of-sample policy risk (0.21) alongside causal forests, whereas linear models could only recommend uniform, one-size-fits-all treatment policies.

These findings demonstrate that optimizing for average treatment effects differs fundamentally from estimating individual-level effects. Standard models that perform adequately on population averages often fail to provide reliable personalized recommendations due to hidden group imbalances. Enforcing distribution balance at the representation level provides a principled regularization mechanism, reducing prediction risk when guiding targeted interventions in clinical care and public resource allocation.

Organizations implementing predictive models for personalized interventions should adopt balanced representation methods when training on observational logs. Practitioners should use nearest-neighbor approximations or policy risk metrics on validation data to select regularization strengths, ensuring models balance factual predictive accuracy against cross-group distribution distance.

The methodology assumes strong ignorability (no unobserved confounding variables), which cannot be verified directly from data and requires strong domain knowledge. Decision-makers should exercise caution when unmeasured confounders are likely present, and future work should focus on establishing automated tuning mechanisms for balancing weights and extending bounds to instrumental variable settings.

arXiv: 1606.03976
  • Paper: Correcting Sample Selection Bias by Unlabeled Data, Jiayuan Huang et al. (2006). Introduces non-parametric distribution matching via kernel mean matching to correct for covariate shift, establishing the fundamental mechanism of balancing sample distributions that underlies the representation learning approach in the source.
  • Paper: Optimal Transport for Domain Adaptation, Nicolas Courty et al. (2014). Develops optimal transport and Wasserstein distance formulations for domain adaptation, providing the core mathematical framework used by the source to measure and minimize distribution divergence.
  • Paper: Recursive partitioning for heterogeneous causal effects, Susan Athey et al. (2015). Establishes tree-based heterogeneous causal effect estimation under the unconfoundedness assumption, laying key foundational groundwork for individual-level treatment effect prediction from observational data.
  • Paper: Stability and Generalization, Olivier Bousquet et al. (2002). Provides fundamental theoretical tools for deriving statistical generalization error bounds in predictive machine learning models.
  • Paper: Estimation and Inference of Heterogeneous Treatment Effects using Random Forests, Stefan Wager et al. (2018). Builds on the estimation of heterogeneous treatment effects under unconfoundedness by developing asymptotic normality and confidence interval guarantees via causal random forests.
  • Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Extends the principle of learning representations that balance distributions across environments to general out-of-distribution generalization via invariant risk minimization.
  • Paper: Generalized random forests, Susan Athey et al. (2016). Generalizes heterogeneous causal effect and local parameter estimation to broad statistical estimating equations within a unified random forest framework.
  • Paper: Recommendations as Treatments: Debiasing Learning and Evaluation, Tobias Schnabel et al. (2016). Applies observational causal inference and propensity-based debiasing concepts to the domain of recommender systems and user feedback evaluation.
  • Paper: Unbiased Learning-to-Rank with Biased Feedback, Thorsten Joachims et al. (2017). Adapts counterfactual risk minimization to correct for presentation bias and unobserved counterfactuals in learning-to-rank systems.
Cover for Estimating individual treatment effect: generalization bounds and algorithms

Abstract

There is intense interest in applying machine learning to problems of causal inference in fields such as healthcare, economics and education. In particular, individual-level causal inference has important applications such as precision medicine. We give a new theoretical analysis and family of algorithms for predicting individual treatment effect (ITE) from observational data, under the assumption known as strong ignorability. The algorithms learn a "balanced" representation such that the induced treated and control distributions look similar. We give a novel, simple and intuitive generalization-error bound showing that the expected ITE estimation error of a representation is bounded by a sum of the standard generalization-error of that representation and the distance between the treated and control distributions induced by the representation. We use Integral Probability Metrics to measure distances between distributions, deriving explicit bounds for the Wasserstein and Maximum Mean Discrepancy (MMD) distances. Experiments on real and simulated data show the new algorithms match or outperform the state-of-the-art.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Estimating ITE: Error bounds
  • 3.1 Problem setup
  • 3.2 Bounds
  • 4 Algorithm for estimating ITE
  • 5 Experiments
  • 5.1 Simulated outcome: IHDP
  • 5.2 Real-world outcome: Jobs
  • 5.3 Results
  • 6 Conclusion
  • References
  • A Proofs
  • A.1 Definitions, assumptions, and auxiliary lemmas
  • A.2 General IPM bound
  • A.3 The family of 11-Lipschitz functions
  • A.4 Functions in the unit ball of a RKHS
  • B Algorithmic details
  • B.1 Minimizing the Wasserstein distance
  • B.2 Minimizing the maximum mean discrepancy
  • C Experimental details
  • C.1 Hyperparameter selection
  • C.2 Learned representations
  • C.3 Absolute error for increasingly imbalanced data

Knowls

  1. Knowl 1 — Generalization Error Bound for Individual Treatment Effect Estimation via Integral Probability Metrics

    theoretical result

    Let X⊂Rd\mathcal{X} \subset \mathbb{R}^d be a bounded covariate space, Y⊂R\mathcal{Y} \subset \mathbb{R} the outcome space, and t∈{0,1}t \in \{0, 1\} a binary treatment indicator. Assume strong ignorability holds, meaning the potential outcomes satisfy (Y0,Y1)⊥t∣x(Y_0, Y_1) \perp t \mid x and the propensity score satisfies 0<p(t=1∣x)<10 < p(t=1 \mid x) < 1 for all x∈Xx \in \mathcal{X}. Let mt(x)=E[Yt∣x]m_t(x) = \mathbb{E}[Y_t \mid x] denote the conditional expectation of potential outcome YtY_t, and let τ(x)=m1(x)−m0(x)\tau(x) = m_1(x) - m_0(x) be the true Individual Treatment Effect (ITE).

    Let Φ:X→R\Phi: \mathcal{X} \to \mathcal{R} be a one-to-one, differentiable representation mapping with inverse Ψ:R→X\Psi: \mathcal{R} \to \mathcal{X}, and let h:R×{0,1}→Yh: \mathcal{R} \times \{0, 1\} \to \mathcal{Y} be a hypothesis predicting the outcome such that f(x,t)=h(Φ(x),t)f(x, t) = h(\Phi(x), t). The estimated treatment effect is τ^f(x)=f(x,1)−f(x,0)\hat{\tau}_f(x) = f(x, 1) - f(x, 0). Under the squared loss L(y,y^)=(y−y^)2L(y, \hat{y}) = (y - \hat{y})^2, the Precision in Estimation of Heterogeneous Effect (PEHE) risk is:

    ϵPEHE(h,Φ)=∫X(τ^f(x)−τ(x))2p(x) dx.\epsilon_{\text{PEHE}}(h, \Phi) = \int_{\mathcal{X}} (\hat{\tau}_f(x) - \tau(x))^2 p(x)\, dx.

    Let pΦt=1(r)=pΦ(r∣t=1)p_\Phi^{t=1}(r) = p_\Phi(r \mid t=1) and pΦt=0(r)=pΦ(r∣t=0)p_\Phi^{t=0}(r) = p_\Phi(r \mid t=0) be the push-forward distributions induced by Φ\Phi over the representation space R\mathcal{R}. Let the expected factual losses on treated and control populations be:

    ϵFt=1(h,Φ)=∫X∫YL(Y1,h(Φ(x),1))p(Y1∣x)p(x∣t=1) dY1dx,\epsilon_F^{t=1}(h, \Phi) = \int_{\mathcal{X}} \int_{\mathcal{Y}} L(Y_1, h(\Phi(x), 1)) p(Y_1 \mid x) p(x \mid t=1)\, dY_1 dx, ϵFt=0(h,Φ)=∫X∫YL(Y0,h(Φ(x),0))p(Y0∣x)p(x∣t=0) dY0dx.\epsilon_F^{t=0}(h, \Phi) = \int_{\mathcal{X}} \int_{\mathcal{Y}} L(Y_0, h(\Phi(x), 0)) p(Y_0 \mid x) p(x \mid t=0)\, dY_0 dx.

    Let σY2=min⁡{σY02,σY12}\sigma_Y^2 = \min\{\sigma_{Y_0}^2, \sigma_{Y_1}^2\}, where σYt2=min⁡{Ep(x,t)[(Yt−mt(x))2],Ep(x,1−t)[(Yt−mt(x))2]}\sigma_{Y_t}^2 = \min\{\mathbb{E}_{p(x,t)}[(Y_t - m_t(x))^2], \mathbb{E}_{p(x,1-t)}[(Y_t - m_t(x))^2]\}. If there exists a constant BΦ>0B_\Phi > 0 and a function family G\mathcal{G} such that 1BΦ∫YL(Yt,h(r,t))p(Yt∣Ψ(r)) dYt∈G\frac{1}{B_\Phi} \int_{\mathcal{Y}} L(Y_t, h(r, t)) p(Y_t \mid \Psi(r))\, dY_t \in \mathcal{G} for t∈{0,1}t \in \{0, 1\}, then the expected ITE estimation error is upper bounded by:

    ϵPEHE(h,Φ)≤2(ϵFt=0(h,Φ)+ϵFt=1(h,Φ)+BΦ⋅IPMG(pΦt=1,pΦt=0)−2σY2),\epsilon_{\text{PEHE}}(h, \Phi) \le 2\left(\epsilon_F^{t=0}(h, \Phi) + \epsilon_F^{t=1}(h, \Phi) + B_\Phi \cdot \text{IPM}_\mathcal{G}\left(p_\Phi^{t=1}, p_\Phi^{t=0}\right) - 2\sigma_Y^2\right),

    where IPMG(p,q)=sup⁡g∈G∣∫g(s)(p(s)−q(s)) ds∣\text{IPM}_\mathcal{G}(p, q) = \sup_{g \in \mathcal{G}} \left|\int g(s)(p(s) - q(s))\, ds\right| is the Integral Probability Metric between the treated and control representations.

  2. Knowl 2 — Counterfactual Regression (CFR) Algorithm for Individual Treatment Effect Estimation

    algorithm

    Counterfactual Regression (CFR) is an end-to-end representation learning algorithm that jointly trains a shared representation network ΦW:X→R\Phi_W: \mathcal{X} \to \mathcal{R} parameterized by weights WW and two hypothesis heads hV=(hV0,hV1):R×{0,1}→Yh_V = (h_{V_0}, h_{V_1}): \mathcal{R} \times \{0, 1\} \to \mathcal{Y} parameterized by weights VV. Training minimizes a weighted factual loss combined with an Integral Probability Metric (IPM) penalty that enforces distribution balance between treated and control representations:

    min⁡hV,ΦW,∥Φ∥=11n∑i=1nwiL(hV(ΦW(xi),ti),yi)+λR(hV)+α⋅IPMG({ΦW(xi)}i:ti=0,{ΦW(xi)}i:ti=1),\min_{h_V, \Phi_W, \|\Phi\|=1} \frac{1}{n} \sum_{i=1}^n w_i L\left(h_V(\Phi_W(x_i), t_i), y_i\right) + \lambda \mathcal{R}(h_V) + \alpha \cdot \text{IPM}_\mathcal{G}\left(\{\Phi_W(x_i)\}_{i:t_i=0}, \{\Phi_W(x_i)\}_{i:t_i=1}\right),

    where u=1n∑i=1ntiu = \frac{1}{n} \sum_{i=1}^n t_i is the empirical treatment proportion, wi=ti2u+1−ti2(1−u)w_i = \frac{t_i}{2u} + \frac{1-t_i}{2(1-u)} reweights samples to match population proportions, R(h)\mathcal{R}(h) is a weight-decay regularizer on the hypothesis heads, and α≥0\alpha \ge 0 is the imbalance penalty hyperparameter. Setting α=0\alpha = 0 yields the Treatment-Agnostic Representation Network (TARNet).

    Input: Factual dataset (x1,t1,y1),…,(xn,tn,yn)(x_1, t_1, y_1), \dots, (x_n, t_n, y_n), scaling parameter α>0\alpha > 0, loss function L(⋅,⋅)L(\cdot, \cdot), representation network ΦW\Phi_W with initial weights WW, hypothesis heads hVh_V with initial weights VV, IPM function family G\mathcal{G}, weight decay parameter λ\lambda.
    Output: Optimized network parameters WW and VV.
    Compute u=1n∑i=1ntiu = \frac{1}{n} \sum_{i=1}^n t_i
    Compute wi=ti2u+1−ti2(1−u)w_i = \frac{t_i}{2u} + \frac{1-t_i}{2(1-u)} for i=1,…,ni = 1, \dots, n
    while not converged do
        Sample mini-batch indices {i1,i2,…,im}⊂{1,…,n}\{i_1, i_2, \dots, i_m\} \subset \{1, \dots, n\}
        Calculate gradient of the empirical IPM term:
            g1=∇WIPMG({ΦW(xij)}tij=0,{ΦW(xik)}tik=1)g_1 = \nabla_W \text{IPM}_\mathcal{G}\left(\{\Phi_W(x_{i_j})\}_{t_{i_j}=0}, \{\Phi_W(x_{i_k})\}_{t_{i_k}=1}\right)
        Calculate gradients of the empirical prediction loss:
            g2=∇V1m∑j=1mwijL(hV(ΦW(xij),tij),yij)g_2 = \nabla_V \frac{1}{m} \sum_{j=1}^m w_{i_j} L\left(h_V(\Phi_W(x_{i_j}), t_{i_j}), y_{i_j}\right)
            g3=∇W1m∑j=1mwijL(hV(ΦW(xij),tij),yij)g_3 = \nabla_W \frac{1}{m} \sum_{j=1}^m w_{i_j} L\left(h_V(\Phi_W(x_{i_j}), t_{i_j}), y_{i_j}\right)
        Obtain step sizes η\eta via an optimization algorithm such as Adam
        Update parameters:
            W←W−η(αg1+g3)W \leftarrow W - \eta (\alpha g_1 + g_3)
            V←V−η(g2+2λV)V \leftarrow V - \eta (g_2 + 2\lambda V)
        Check convergence criterion
    end while
  3. Knowl 3 — PEHE Generalization Error Bound with Wasserstein Distance

    theoretical result

    Let Φ:X→R\Phi: \mathcal{X} \to \mathcal{R} be a one-to-one, differentiable representation that is Jacobian-normalized, meaning sup⁡x∈Xσmax⁡(∂Φ(x)∂x)=1\sup_{x \in \mathcal{X}} \sigma_{\max}\left(\frac{\partial \Phi(x)}{\partial x}\right) = 1, where σmax⁡(A)\sigma_{\max}(A) and σmin⁡(A)\sigma_{\min}(A) denote the largest and smallest singular values of matrix AA. Define the Jacobian condition number ratio ρ(Φ)=sup⁡x∈Xσmax⁡(∂Φ(x)/∂x)σmin⁡(∂Φ(x)/∂x)≥1\rho(\Phi) = \sup_{x \in \mathcal{X}} \frac{\sigma_{\max}(\partial \Phi(x) / \partial x)}{\sigma_{\min}(\partial \Phi(x) / \partial x)} \ge 1.

    Assume that:

    1. The potential outcome conditional density p(Yt∣x)p(Y_t \mid x) is differentiable with ∥∂p(Yt∣x)∂x∥≤K\left\|\frac{\partial p(Y_t \mid x)}{\partial x}\right\| \le K for all x∈Xx \in \mathcal{X} and t∈{0,1}t \in \{0, 1\}.
    2. The loss function L(y1,y2)L(y_1, y_2) has bounded derivative ∣∂L(y1,y2)∂yi∣≤KL\left|\frac{\partial L(y_1, y_2)}{\partial y_i}\right| \le K_L for i∈{1,2}i \in \{1, 2\}, and ∫YL(y1,y2) dy1≤M\int_\mathcal{Y} L(y_1, y_2)\, dy_1 \le M for all y2∈Yy_2 \in \mathcal{Y}.
    3. The hypothesis h:R×{0,1}→Yh: \mathcal{R} \times \{0, 1\} \to \mathcal{Y} has gradient norm bounded by bKbK.

    When G\mathcal{G} is the family of 1-Lipschitz functions, the Integral Probability Metric is the 1-Wasserstein distance Wass(pΦt=1,pΦt=0)\text{Wass}(p_\Phi^{t=1}, p_\Phi^{t=0}). Under the squared loss for factual loss ϵF\epsilon_F, the PEHE risk is bounded by:

    ϵPEHE(h,Φ)≤2ϵFt=0(h,Φ)+2ϵFt=1(h,Φ)−4σY2+2(Mρ(Φ)+b)K⋅KL⋅Wass(pΦt=1,pΦt=0).\epsilon_{\text{PEHE}}(h, \Phi) \le 2 \epsilon_F^{t=0}(h, \Phi) + 2 \epsilon_F^{t=1}(h, \Phi) - 4\sigma_Y^2 + 2\left(M \rho(\Phi) + b\right) K \cdot K_L \cdot \text{Wass}\left(p_\Phi^{t=1}, p_\Phi^{t=0}\right).

    The condition number ρ(Φ)\rho(\Phi) is minimized when Φ\Phi is a linear orthogonal transformation, precluding degenerate representations where Φ\Phi collapses to a constant.

  4. Knowl 4 — PEHE Generalization Error Bound with Maximum Mean Discrepancy

    theoretical result

    Let Hx\mathcal{H}_x and Hr\mathcal{H}_r be Reproducing Kernel Hilbert Spaces (RKHS) over X\mathcal{X} and R\mathcal{R} with kernels kx(⋅,⋅)k_x(\cdot, \cdot) and kr(⋅,⋅)k_r(\cdot, \cdot). Let mt(x)=E[Yt∣x]m_t(x) = \mathbb{E}[Y_t \mid x] and ηYt(x)=∫Y(Yt−mt(x))2p(Yt∣x) dYt\eta_{Y_t}(x) = \sqrt{\int_\mathcal{Y} (Y_t - m_t(x))^2 p(Y_t \mid x)\, dY_t} be the conditional mean and standard deviation functions. Assume:

    1. mt∈Hxm_t \in \mathcal{H}_x with ∥mt∥Hx≤K\|m_t\|_{\mathcal{H}_x} \le K for t∈{0,1}t \in \{0, 1\}.
    2. ηYt∈Hx\eta_{Y_t} \in \mathcal{H}_x with ∥ηYt∥Hx≤M\|\eta_{Y_t}\|_{\mathcal{H}_x} \le M for t∈{0,1}t \in \{0, 1\}.
    3. Φ:X→R\Phi: \mathcal{X} \to \mathcal{R} induces a bounded linear operator ΓΦ:Hr→Hx\Gamma_\Phi: \mathcal{H}_r \to \mathcal{H}_x such that ⟨f,kx(Ψ(r),⋅)⟩Hx=⟨f,ΓΦkr(r,⋅)⟩Hx\langle f, k_x(\Psi(r), \cdot)\rangle_{\mathcal{H}_x} = \langle f, \Gamma_\Phi k_r(r, \cdot)\rangle_{\mathcal{H}_x} with Hilbert-Schmidt norm ∥ΓΦ∥HS≤KΦ\|\Gamma_\Phi\|_{\text{HS}} \le K_\Phi.
    4. Hypothesis h(⋅,t)∈Hrh(\cdot, t) \in \mathcal{H}_r with ∥h(⋅,t)∥Hr≤b\|h(\cdot, t)\|_{\mathcal{H}_r} \le b for t∈{0,1}t \in \{0, 1\}.

    When G={g∈Hr⊗Hr:∥g∥Hr⊗Hr≤1}\mathcal{G} = \{g \in \mathcal{H}_r \otimes \mathcal{H}_r : \|g\|_{\mathcal{H}_r \otimes \mathcal{H}_r} \le 1\}, the Integral Probability Metric corresponds to the Maximum Mean Discrepancy MMD(pΦt=1,pΦt=0)\text{MMD}(p_\Phi^{t=1}, p_\Phi^{t=0}). Under the squared loss, the PEHE risk is bounded by:

    ϵPEHE(h,Φ)≤2ϵFt=0(h,Φ)+2ϵFt=1(h,Φ)−4σY2+4(KΦ2(K2+M2)+b2)MMD(pΦt=1,pΦt=0).\epsilon_{\text{PEHE}}(h, \Phi) \le 2 \epsilon_F^{t=0}(h, \Phi) + 2 \epsilon_F^{t=1}(h, \Phi) - 4\sigma_Y^2 + 4\left(K_\Phi^2(K^2 + M^2) + b^2\right) \text{MMD}\left(p_\Phi^{t=1}, p_\Phi^{t=0}\right).
  5. Knowl 5 — Definitions of PEHE, Factual Loss, and Counterfactual Loss

    definition

    In the Rubin-Neyman potential outcomes framework with covariates x∈Xx \in \mathcal{X}, binary treatment t∈{0,1}t \in \{0, 1\}, and potential outcomes Y0,Y1∈YY_0, Y_1 \in \mathcal{Y}, the true individual treatment effect is τ(x)=E[Y1−Y0∣x]=m1(x)−m0(x)\tau(x) = \mathbb{E}[Y_1 - Y_0 \mid x] = m_1(x) - m_0(x). For a predictive model f(x,t)=h(Φ(x),t)f(x, t) = h(\Phi(x), t), the estimated treatment effect is τ^f(x)=f(x,1)−f(x,0)\hat{\tau}_f(x) = f(x, 1) - f(x, 0).

    1. The Precision in Estimation of Heterogeneous Effect (PEHE) loss is:
    ϵPEHE(f)=∫X(τ^f(x)−τ(x))2p(x) dx.\epsilon_{\text{PEHE}}(f) = \int_{\mathcal{X}} (\hat{\tau}_f(x) - \tau(x))^2 p(x)\, dx.
    1. For a per-unit loss ℓh,Φ(x,t)=∫YL(Yt,h(Φ(x),t))p(Yt∣x) dYt\ell_{h, \Phi}(x, t) = \int_\mathcal{Y} L(Y_t, h(\Phi(x), t)) p(Y_t \mid x)\, dY_t, the expected factual loss ϵF\epsilon_F and counterfactual loss ϵCF\epsilon_{CF} are defined as:
    ϵF(h,Φ)=∫X×{0,1}ℓh,Φ(x,t)p(x,t) dxdt,\epsilon_F(h, \Phi) = \int_{\mathcal{X} \times \{0, 1\}} \ell_{h, \Phi}(x, t) p(x, t)\, dx dt, ϵCF(h,Φ)=∫X×{0,1}ℓh,Φ(x,t)p(x,1−t) dxdt.\epsilon_{CF}(h, \Phi) = \int_{\mathcal{X} \times \{0, 1\}} \ell_{h, \Phi}(x, t) p(x, 1 - t)\, dx dt.

    While ϵF\epsilon_F measures error on the observed treatment assignments, ϵCF\epsilon_{CF} measures expected error under the flipped, counterfactual treatment assignments.

  6. Knowl 6 — Nearest-Neighbor PEHE Approximation for Causal Model Validation

    model/method

    Because counterfactual outcomes are never observed simultaneously with factual outcomes in real-world observational data, standard cross-validation cannot directly evaluate the PEHE loss ϵPEHE(f)\epsilon_{\text{PEHE}}(f). To select hyperparameters without ground-truth counterfactuals, the unobserved potential outcome for unit ii is approximated by the observed outcome yj(i)y_{j(i)} of its nearest Euclidean neighbor j(i)j(i) in the opposite treatment group (tj(i)=1−tit_{j(i)} = 1 - t_i).

    The 1-nearest-neighbor surrogate loss ϵPEHEnn(f)\epsilon_{\text{PEHEnn}}(f) is defined as:

    ϵPEHEnn(f)=1n∑i=1n((1−2ti)(yj(i)−yi)−(f(xi,1)−f(xi,0)))2.\epsilon_{\text{PEHEnn}}(f) = \frac{1}{n} \sum_{i=1}^n \left((1 - 2t_i)(y_{j(i)} - y_i) - (f(x_i, 1) - f(x_i, 0))\right)^2.

    This criterion is computed on a held-out validation set to tune hyperparameters (such as network dimensions, layer counts, batch sizes, and the balance penalty α\alpha) for individual treatment effect estimators.

  7. Knowl 7 — Policy Risk Metric for Treatment Selection Evaluation

    definition

    When individual potential outcomes are unobserved but data includes a randomized trial component, the quality of an ITE estimator f(x,t)f(x, t) can be evaluated through its induced treatment assignment policy πf(x)∈{0,1}\pi_f(x) \in \{0, 1\}. Given a decision threshold λ\lambda, the policy assigns treatment if the predicted effect exceeds λ\lambda:

    πf(x)={1if f(x,1)−f(x,0)>λ,0otherwise.\pi_f(x) = \begin{cases} 1 & \text{if } f(x, 1) - f(x, 0) > \lambda, \\ 0 & \text{otherwise.} \end{cases}

    The policy risk RPol(πf)R_{\text{Pol}}(\pi_f) measures the average loss in value resulting from treating according to πf\pi_f:

    RPol(πf)=1−(E[Y1∣πf(x)=1]⋅p(πf(x)=1)+E[Y0∣πf(x)=0]⋅p(πf(x)=0)).R_{\text{Pol}}(\pi_f) = 1 - \left(\mathbb{E}[Y_1 \mid \pi_f(x) = 1] \cdot p(\pi_f(x) = 1) + \mathbb{E}[Y_0 \mid \pi_f(x) = 0] \cdot p(\pi_f(x) = 0)\right).

    Over a randomized trial subgroup EE where treatment tt is independent of xx, the empirical policy risk is estimated as:

    R^Pol(πf)=1−(E[Y1∣πf(x)=1,t=1]⋅p(πf(x)=1)+E[Y0∣πf(x)=0,t=0]⋅p(πf(x)=0)).\hat{R}_{\text{Pol}}(\pi_f) = 1 - \left(\mathbb{E}[Y_1 \mid \pi_f(x) = 1, t = 1] \cdot p(\pi_f(x) = 1) + \mathbb{E}[Y_0 \mid \pi_f(x) = 0, t = 0] \cdot p(\pi_f(x) = 0)\right).
  8. Knowl 8 — Empirical Performance on IHDP and Jobs Benchmark Datasets

    data/table

    Performance of Counterfactual Regression with Wasserstein distance (CFR WASS) and linear Maximum Mean Discrepancy (CFR MMD), compared to TARNet (α=0\alpha=0) and baseline estimators on the semi-synthetic IHDP dataset (evaluating ϵPEHE\sqrt{\epsilon_{\text{PEHE}}} and ϵATE\epsilon_{\text{ATE}} across 1000 outcome realizations with 63/27/10 splits) and the real-world Jobs dataset (evaluating policy risk RPolR_{\text{Pol}} with λ=0\lambda=0 and ϵATT\epsilon_{\text{ATT}} across 10 splits with 56/24/20 ratios). Lower numbers indicate better performance.

    Within-sample Out-of-sample
    IHDP JOBS IHDP JOBS
    Method ϵPEHE\sqrt{\epsilon_{\text{PEHE}}} ϵATE\epsilon_{\text{ATE}} RPolR_{\text{Pol}} ϵATT\epsilon_{\text{ATT}} ϵPEHE\sqrt{\epsilon_{\text{PEHE}}} ϵATE\epsilon_{\text{ATE}} RPolR_{\text{Pol}} ϵATT\epsilon_{\text{ATT}}
    OLS/LR-1 5.8±.35.8 \pm .3 .73±.04.73 \pm .04 .22±.0.22 \pm .0 .01±.00.01 \pm .00 5.8±.35.8 \pm .3 .94±.06.94 \pm .06 .23±.0.23 \pm .0 .08±.04.08 \pm .04
    OLS/LR-2 2.4±.12.4 \pm .1 .14±.01.14 \pm .01 .21±.0.21 \pm .0 .01±.01.01 \pm .01 2.5±.12.5 \pm .1 .31±.02.31 \pm .02 .24±.0.24 \pm .0 .08±.03.08 \pm .03
    BLR 5.8±.35.8 \pm .3 .72±.04.72 \pm .04 .22±.0.22 \pm .0 .01±.01.01 \pm .01 5.8±.35.8 \pm .3 .93±.05.93 \pm .05 .25±.0.25 \pm .0 .08±.03.08 \pm .03
    k-NN 2.1±.12.1 \pm .1 .14±.01.14 \pm .01 .02±.0.02 \pm .0 .21±.01.21 \pm .01 4.1±.24.1 \pm .2 .79±.05.79 \pm .05 .26±.0.26 \pm .0 .13±.05.13 \pm .05
    TMLE 5.0±.25.0 \pm .2 .30±.01.30 \pm .01 .22±.0.22 \pm .0 .02±.01.02 \pm .01 — — — —
    BART 2.1±.12.1 \pm .1 .23±.01.23 \pm .01 .23±.0.23 \pm .0 .02±.00.02 \pm .00 2.3±.12.3 \pm .1 .34±.02.34 \pm .02 .25±.0.25 \pm .0 .08±.03.08 \pm .03
    RAND. FOR. 4.2±.24.2 \pm .2 .73±.05.73 \pm .05 .23±.0.23 \pm .0 .03±.01.03 \pm .01 6.6±.36.6 \pm .3 .96±.06.96 \pm .06 .28±.0.28 \pm .0 .09±.04.09 \pm .04
    CAUS. FOR. 3.8±.23.8 \pm .2 .18±.01.18 \pm .01 .19±.0.19 \pm .0 .03±.01.03 \pm .01 3.8±.23.8 \pm .2 .40±.03.40 \pm .03 .20±.0.20 \pm .0 .07±.03.07 \pm .03
    BNN 2.2±.12.2 \pm .1 .37±.03.37 \pm .03 .20±.0.20 \pm .0 .04±.01.04 \pm .01 2.1±.12.1 \pm .1 .42±.03.42 \pm .03 .24±.0.24 \pm .0 .09±.04.09 \pm .04
    TARNET .88±.0.88 \pm .0 .26±.01.26 \pm .01 .17±.0.17 \pm .0 .05±.02.05 \pm .02 .95±.0.95 \pm .0 .28±.01.28 \pm .01 .21±.0.21 \pm .0 .11±.04.11 \pm .04
    CFR MMD .73±.0.73 \pm .0 .30±.01.30 \pm .01 .18±.0.18 \pm .0 .04±.01.04 \pm .01 .78±.0.78 \pm .0 .31±.01.31 \pm .01 .21±.0.21 \pm .0 .08±.03.08 \pm .03
    CFR WASS .71±.0.71 \pm .0 .25±.01.25 \pm .01 .17±.0.17 \pm .0 .04±.01.04 \pm .01 .76±.0.76 \pm .0 .27±.01.27 \pm .01 .21±.0.21 \pm .0 .09±.03.09 \pm .03

    The table demonstrates that CFR WASS and CFR MMD achieve significantly lower ITE estimation error (ϵPEHE\sqrt{\epsilon_{\text{PEHE}}}) than both linear models and non-linear causal baselines (BART, Causal Forests, BNN). Balancing representations with α>0\alpha > 0 yields a distinct improvement over TARNet (α=0\alpha = 0), reducing out-of-sample ϵPEHE\sqrt{\epsilon_{\text{PEHE}}} from 0.950.95 to 0.760.76 on IHDP.

  9. Knowl 9 — Impact of Imbalance and Balancing Regularization on Counterfactual Generalization

    empirical result

    To evaluate how distribution shift between treated and control units affects ITE estimation, controlled imbalance was introduced into the IHDP dataset by fitting a propensity score model p^(t=1∣x)\hat{p}(t=1 \mid x) and iteratively removing control observations: with probability q∈{0.0,0.5,1.0}q \in \{0.0, 0.5, 1.0\}, the control unit with propensity score closest to 1 was dropped, and with probability 1−q1 - q, a randomly chosen control unit was dropped (removing 347 units total to leave 400).

    Across 500 realizations of IHDP:

    1. Higher imbalance qq strictly increases the out-of-sample ITE error ϵPEHE\epsilon_{\text{PEHE}} for models without balance regularization (α=0\alpha = 0).
    2. Introducing the Wasserstein IPM regularization penalty α>0\alpha > 0 reduces out-of-sample ϵPEHE\epsilon_{\text{PEHE}} by up to 20–30%20\text{--}30\% relative to α=0\alpha = 0, with optimal error achieved in the range α∈[10−3,10−1]\alpha \in [10^{-3}, 10^{-1}].
    3. The relative gain from IPM balancing persists or grows as the level of confounding imbalance qq increases, confirming that penalizing representation divergence directly mitigates counterfactual prediction errors under selection bias.
  10. Knowl 10 — Stochastic Gradient Computation of Wasserstein Distance via Sinkhorn Iteration

    algorithm

    Computing the exact Wasserstein distance requires linear programming, which is computationally expensive for mini-batch stochastic optimization. An entropic regularized approximation (Sinkhorn distance) enables efficient GPU-compatible gradient computation during representation training.

    Input: Mini-batch with mm control units and m′m' treated units (xi1,0,yi1),…,(xim,0,yim),(xim+1,1,yim+1),…,(xim+m′,1,yim+m′)(x_{i_1}, 0, y_{i_1}), \dots, (x_{i_m}, 0, y_{i_m}), (x_{i_{m+1}}, 1, y_{i_{m+1}}), \dots, (x_{i_{m+m'}}, 1, y_{i_{m+m'}}), representation network ΦW\Phi_W with weights WW.
    Output: Stochastic gradient g1=∇WWass({ΦW(xij)}t=0,{ΦW(xik)}t=1)g_1 = \nabla_W \text{Wass}(\{\Phi_W(x_{i_j})\}_{t=0}, \{\Phi_W(x_{i_k})\}_{t=1}).
    Calculate pairwise Euclidean distance matrix M(ΦW)∈Rm×m′M(\Phi_W) \in \mathbb{R}^{m \times m'}:
        Mkl(ΦW)=∥ΦW(xik)−ΦW(xim+l)∥M_{kl}(\Phi_W) = \|\Phi_W(x_{i_k}) - \Phi_W(x_{i_{m+l}})\|
    Compute regularized kernel matrix Kkl=exp⁡(−λMkl(ΦW)2)K_{kl} = \exp(-\lambda M_{kl}(\Phi_W)^2)
    Compute optimal transport matrix T∗T^* via fixed-point Sinkhorn-Knopp iterations:
        ut+1=nt⊘(ncK(1⊘(utTK)T))u_{t+1} = n_t \oslash (n_c K (1 \oslash (u_t^T K)^T))
    Calculate the gradient with respect to representation weights WW:
        g1=∇W⟨T∗,M(ΦW)⟩g_1 = \nabla_W \langle T^*, M(\Phi_W) \rangle
    return g1g_1

Coverage note — None was omitted; all major theoretical bounds (Theorems 1, 2, 3), algorithms (CFR, TARNet, Sinkhorn gradient), validation metrics (1-NN PEHE, Policy Risk), and empirical findings from the IHDP and Jobs benchmarks are fully covered.

References

  1. 1.MathOverflow: functions with orthogonal Jacobian. https://mathoverflow.net/questions/228964/functions-with-orthogonal-jacobian. Accessed: 2016-05-05.
  2. 2.Athey, Susan and Imbens, Guido. Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27):7353–7360, 2016.
  3. 3.Athey, Susan, Imbens, Guido W, and Wager, Stefan. Efficient inference of average treatment effects in high dimensions via approximate residual balancing. arXiv preprint arXiv:1604.07125, 2016.
  4. 4.Aude, Genevay, Cuturi, Marco, Peyre, Gabriel, and Bach, Francis. Stochastic optimization for large-scale optimal transport. arXiv preprint arXiv:1605.08527, 2016.
  5. 5.Austin, Peter C. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research, 46(3):399–424, 2011.
  6. 6.Balke, Alexander and Pearl, Judea. Bounds on treatment effects from studies with imperfect compliance. Journal of the American Statistical Association, 92(439):1171–1176, 1997.
  7. 7.Bareinboim, Elias and Pearl, Judea. Controlling selection bias in causal inference. In AISTATS, pp. 100–108, 2012.
  8. 8.Bareinboim, Elias and Pearl, Judea. Causal inference and the data-fusion problem. Proceedings of the National Academy of Sciences, 113(27):7345–7352, 2016.
  9. 9.Beck, Nathaniel, King, Gary, and Zeng, Langche. Improving quantitative studies of international conflict: A conjecture. American Political Science Review, 94(01):21–35, 2000.
  10. 10.Belloni, Alexandre, Chernozhukov, Victor, and Hansen, Christian. Inference on treatment effects after selection among high-dimensional controls. The Review of Economic Studies, 81(2):608–650, 2014.
  11. 11.Ben-David, Shai, Blitzer, John, Crammer, Koby, Pereira, Fernando, et al. Analysis of representations for domain adaptation. Advances in neural information processing systems, 19:137, 2007.
  12. 12.Ben-David, Shai, Blitzer, John, Crammer, Koby, Kulesza, Alex, Pereira, Fernando, and Vaughan, Jennifer Wortman. A theory of learning from different domains. Machine learning, 79(1-2):151–175, 2010.
  13. 13.Ben-Israel, Adi. The change-of-variables formula using matrix volume. SIAM Journal on Matrix Analysis and Applications, 21(1):300–312, 1999.
  14. 14.Bengio, Yoshua, Courville, Aaron, and Vincent, Pierre. Representation learning: A review and new perspectives. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 35(8):1798–1828, 2013.
  15. 15.Breiman, Leo. Random forests. Machine learning, 45(1):5–32, 2001.
  16. 16.Cai, Zhihong, Kuroki, Manabu, Pearl, Judea, and Tian, Jin. Bounds on direct effects in the presence of confounded intermediate variables. Biometrics, 64(3):695–701, 2008.
  17. 17.Chernozhukov, Victor, Chetverikov, Denis, Demirer, Mert, Duflo, Esther, Hansen, Christian, et al. Double machine learning for treatment and causal parameters. arXiv preprint arXiv:1608.00060, 2016.
  18. 18.Chipman, Hugh and McCulloch, Robert. BayesTree: Bayesian Additive Regression Trees. https://cran.r-project.org/web/packages/BayesTree, 2016.
  19. 19.Chipman, Hugh A, George, Edward I, and McCulloch, Robert E. BART: Bayesian additive regression trees. The Annals of Applied Statistics, pp. 266–298, 2010.
  20. 20.Cortes, Corinna and Mohri, Mehryar. Domain adaptation and sample bias correction theory and algorithm for regression. Theoretical Computer Science, 519:103–126, 2014.
  21. 21.Cuturi, Marco. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems, pp. 2292–2300, 2013.
  22. 22.Cuturi, Marco and Doucet, Arnaud. Fast computation of Wasserstein barycenters. In Proceedings of The 31st International Conference on Machine Learning, pp. 685–693, 2014.
  23. 23.Daume III, Hal. Frustratingly easy domain adaptation. arXiv preprint arXiv:0907.1815, 2009.
  24. 24.Dehejia, Rajeev H and Wahba, Sadek. Propensity score-matching methods for nonexperimental causal studies. Review of Economics and statistics, 84(1):151–161, 2002.
  25. 25.Dorie, Vincent. NPCI: Non-parametrics for Causal Inference. https://github.com/vdorie/npci, 2016.
  26. 26.Funk, Michele Jonsson, Westreich, Daniel, Wiesen, Chris, Sturmer, Til, Brookhart, M Alan, and Davidian, Marie. Doubly robust estimation of causal effects. American journal of epidemiology, 173(7):761–767, 2011.
  27. 27.Ganin, Yaroslav, Ustinova, Evgeniya, Ajakan, Hana, Germain, Pascal, Larochelle, Hugo, Laviolette, Franc¸ois, Marchand, Mario, and Lempitsky, Victor. Domain-adversarial training of neural networks. Journal of Machine Learning Research, 17(59):1–35, 2016. URL http://jmlr.org/papers/v17/15-239.html.
  28. 28.Gretton, Arthur, Smola, Alex, Huang, Jiayuan, Schmittfull, Marcel, Borgwardt, Karsten, and Scholkopf, Bernhard. Covariate shift by kernel mean matching. Dataset shift in machine learning, 3(4):5, 2009.
  29. 29.Gretton, Arthur, Borgwardt, Karsten M., Rasch, Malte J., Scholkopf, Bernhard, and Smola, Alexander. A kernel two-sample test. J. Mach. Learn. Res., 13:723–773, March 2012. ISSN 1532-4435.
  30. 30.Gruber, Susan and van der Laan, Mark J. tmle: An r package for targeted maximum likelihood estimation. 2011.
  31. 31.Grunewalder, Steffen, Arthur, Gretton, and Shawe-Taylor, John. Smooth operators. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), pp. 1184–1192, 2013.
  32. 32.Hartford, Jason, Lewis, Greg, Leyton-Brown, Kevin, and Taddy, Matt. Counterfactual prediction with deep instrumental variables networks. arXiv preprint arXiv:1612.09596, 2016.
  33. 33.Hill, Jennifer L. Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics, 20(1), 2011.
  34. 34.Hoyer, Patrik O, Janzing, Dominik, Mooij, Joris M, Peters, Jonas, and Scholkopf, Bernhard. Nonlinear causal discovery with additive noise models. In Advances in neural information processing systems, pp. 689–696, 2009.
  35. 35.Imbens, Guido W and Wooldridge, Jeffrey M. Recent developments in the econometrics of program evaluation. Journal of economic literature, 47(1):5–86, 2009.
  36. 36.Johansson, Fredrik D., Shalit, Uri, and Sontag, David. Learning representations for counterfactual inference. In Proceedings of the 33rd International Conference on Machine Learning (ICML), 2016.
  37. 37.Kingma, Diederik and Ba, Jimmy. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  38. 38.Kuang, Max and Tabak, Esteban. Preconditioning of optimal transport. Preprint, 2016.
  39. 39.LaLonde, Robert J. Evaluating the econometric evaluations of training programs with experimental data. The American economic review, pp. 604–620, 1986.
  40. 40.Maathuis, Marloes H, Colombo, Diego, Kalisch, Markus, and Buhlmann, Peter. Predicting causal effects in large-scale systems from observational data. Nature Methods, 7(4):247–248, 2010.
  41. 41.Mansour, Yishay, Mohri, Mehryar, and Rostamizadeh, Afshin. Domain adaptation: Learning bounds and algorithms. 2009.
  42. 42.Mooij, Joris M, Peters, Jonas, Janzing, Dominik, Zscheischler, Jakob, and Scholkopf, Bernhard. Distinguishing cause from effect using observational data: methods and benchmarks. Journal of Machine Learning Research, 17(32):1–102, 2016.
  43. 43.Muller, Alfred. Integral probability metrics and their generating classes of functions. Advances in Applied Probability, pp. 429–443, 1997.
  44. 44.Pan, Sinno Jialin, Tsang, Ivor W, Kwok, James T, and Yang, Qiang. Domain adaptation via transfer component analysis. Neural Networks, IEEE Transactions on, 22(2):199–210, 2011.
  45. 45.Pearl, Judea. Causality. Cambridge university press, 2009.
  46. 46.Pearl, Judea. Detecting latent heterogeneity. Sociological Methods & Research, pp. 0049124115600597, 2015.
  47. 47.Peysakhovich, Alexander and Lada, Akos. Combining observational and experimental data to find heterogeneous treatment effects. arXiv preprint arXiv:1611.02385, 2016.
  48. 48.Rolling, Craig Anthony. Estimation of Conditional Average Treatment Effects. PhD thesis, University of Minnesota, 2014.
  49. 49.Rubin, Donald B. Causal inference using potential outcomes. Journal of the American Statistical Association, 2011.
  50. 50.Shalev-Shwartz, Shai and Ben-David, Shai. Understanding machine learning: From theory to algorithms. Cambridge University Press, 2014.
  51. 51.Shpitser, Ilya and Pearl, Judea. Identification of conditional interventional distributions. In Proceedings of the Twenty-second Conference on Uncertainty in Artificial Intelligence, pp. 437–444. UAI Press, 2006.
  52. 52.Smith, Jeffrey A and Todd, Petra E. Does matching overcome LaLonde’s critique of nonexperimental estimators? Journal of econometrics, 125(1):305–353, 2005.
  53. 53.Sriperumbudur, Bharath K, Fukumizu, Kenji, Gretton, Arthur, Scholkopf, Bernhard, Lanckriet, Gert RG, et al. On the empirical estimation of integral probability metrics. Electronic Journal of Statistics, 6:1550–1599, 2012.
  54. 54.Steinwart, Ingo and Christmann, Andreas. Support vector machines. Springer Science & Business Media, 2008.
  55. 55.Strehl, Alex, Langford, John, Li, Lihong, and Kakade, Sham M. Learning from logged implicit exploration data. In Advances in Neural Information Processing Systems, pp. 2217–2225, 2010.
  56. 56.Sun, Baochen, Feng, Jiashi, and Saenko, Kate. Return of frustratingly easy domain adaptation. In Thirtieth AAAI Conference on Artificial Intelligence, 2016.
  57. 57.Swaminathan, Adith and Joachims, Thorsten. Batch learning from logged bandit feedback through counterfactual risk minimization. Journal of Machine Learning Research, 16:1731–1755, 2015.
  58. 58.Taddy, Matt, Gardner, Matt, Chen, Liyun, and Draper, David. A nonparametric bayesian analysis of heterogenous treatment effects in digital experimentation. Journal of Business & Economic Statistics, 34(4):661–672, 2016.
  59. 59.Triantafillou, Sofia and Tsamardinos, Ioannis. Constraint-based causal discovery from multiple interventions over overlapping variable sets. Journal of Machine Learning Research, 16:2147–2205, 2015.
  60. 60.Villani, Cedric. Optimal transport: old and new, volume 338. Springer Science & Business Media, 2008.
  61. 61.Wager, Stefan and Athey, Susan. Estimation and inference of heterogeneous treatment effects using random forests. arXiv preprint arXiv:1510.04342. https://github.com/susanathey/causalTree, 2015.

Citation

MLA
Shalit, U., et al. “Estimating Individual Treatment Effect: Generalization Bounds and Algorithms”. arXiv, 2016, http://arxiv.org/abs/1606.03976v5.
APA
Shalit, U., Johansson, F. D., & Sontag, D. (2016). Estimating individual treatment effect: generalization bounds and algorithms. arXiv. http://arxiv.org/abs/1606.03976v5
Chicago
Shalit, U., F. D. Johansson, and D. Sontag. 2016. “Estimating Individual Treatment Effect: Generalization Bounds and Algorithms”. arXiv. http://arxiv.org/abs/1606.03976v5.
Harvard
Shalit, U., Johansson, F.D. and Sontag, D. (2016) “Estimating individual treatment effect: generalization bounds and algorithms”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1606.03976v5.
Vancouver
1. Shalit U, Johansson FD, Sontag D (2016) Estimating individual treatment effect: generalization bounds and algorithms. arXiv

BibTeX

@article{shalit2016estimating,
  title = {Estimating individual treatment effect: generalization bounds and algorithms},
  author = {Shalit, Uri and Johansson, Fredrik D. and Sontag, David},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1606.03976v5},
  eprint = {1606.03976}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors