Recursive partitioning for heterogeneous causal effects

Susan AtheyGuido Imbens

article2015Proceedings of the National Academy of Sciences of the United States of America1,739 citations

Develops a decision-tree framework tailored for causal inference that enables researchers to discover heterogeneous treatment effects across subpopulations while preserving valid statistical inference.

Listen

In experimental and observational studies, decision-makers frequently need to understand how treatment effects vary across different subpopulations rather than relying solely on overall population averages. Traditional machine learning techniques, such as standard decision trees, can identify complex patterns among large numbers of baseline characteristics. However, applying these algorithms directly to causal inference creates significant statistical challenges: individual-level causal impacts are never directly observable because no single unit is seen under both treatment and control conditions, and reusing the same data to search for subgroups and estimate effects produces severe selection bias that invalidates standard confidence intervals and hypothesis tests.

To resolve these challenges, the article develops a recursive partitioning framework called Causal Trees alongside an honest estimation strategy. The primary objective is to accurately discover distinct subgroups with heterogeneous treatment effects and construct valid statistical confidence intervals without requiring restrictive assumptions about data sparsity or model complexity.

The proposed approach introduces an honest estimation workflow that splits the available data into distinct samples: one sample builds and prunes the tree structure, while an independent estimation sample calculates the treatment effects and standard errors within the resulting subgroups. In addition, the article redesigns standard splitting and cross-validation criteria to target the mean-squared error of treatment effects rather than overall outcome predictions, explicitly accounting for the trade-off between subgroup granularity and leaf-level estimation variance. The methodology was evaluated through formal statistical analysis and extensive numerical simulations spanning varied sample sizes, covariate dimensions, and treatment effect structures, comparing Causal Trees against several alternative subgroup-discovery approaches.

The simulation results demonstrate several key findings. First, honest estimation successfully achieves nominal 90% confidence interval coverage across all test designs, whereas standard adaptive estimation methods fail, exhibiting coverage rates falling between 74% and 89% due to unchecked selection bias. Second, the honest Causal Tree method consistently outperforms or matches alternative subgrouping algorithms—such as fit-based trees and squared t-statistic trees—by properly balancing treatment effect heterogeneity against within-group variance reduction. Third, although dividing the sample incurs a sample size penalty, the reduction in bias from honest estimation compensates for this loss, leading to superior or highly competitive overall mean-squared error performance.

These findings have immediate practical implications for clinical trials, policy evaluations, and commercial personalization systems where transparent, rule-based decision guidelines are essential. Decision-makers can deploy these methods to uncover meaningful subgroup differences that were not prespecified in initial analysis plans without invalidating downstream statistical testing or falling prey to false discoveries from multiple hypothesis testing. This provides high confidence in identified subgroup impacts, protecting organizations from costly misallocations of treatments or interventions.

Organizations analyzing experimental or observational data should adopt honest recursive partitioning when exploring treatment heterogeneity and deriving simple subpopulation rules. When applying the method to observational datasets, practitioners must ensure that all confounding variables influencing treatment assignment are observed and appropriately adjusted via propensity score weighting. Users should remain cautious in small-sample environments, as sufficient observations per subgroup are required to reliably estimate both treatment and control outcomes.

Cover for Recursive partitioning for heterogeneous causal effects

Abstract

In this paper we study the problems of estimating heterogeneity in causal effects in experimental or observational studies and conducting inference about the magnitude of the differences in treatment effects across subsets of the population. In applications, our method provides a data-driven approach to determine which subpopulations have large or small treatment effects and to test hypotheses about the differences in these effects. For experiments, our method allows researchers to identify heterogeneity in treatment effects that was not specified in a pre-analysis plan, without concern about invalidating inference due to multiple testing. In most of the literature on supervised machine learning (e.g. regression trees, random forests, LASSO, etc.), the goal is to build a model of the relationship between a unit's attributes and an observed outcome. A prominent role in these methods is played by cross-validation which compares predictions to actual outcomes in test samples, in order to select the level of complexity of the model that provides the best predictive power. Our method is closely related, but it differs in that it is tailored for predicting causal effects of a treatment rather than a unit's outcome. The challenge is that the "ground truth" for a causal effect is not observed for any individual unit: we observe the unit with the treatment, or without the treatment, but not both at the same time. Thus, it is not obvious how to use cross-validation to determine whether a causal effect has been accurately predicted. We propose several novel cross-validation criteria for this problem and demonstrate through simulations the conditions under which they perform better than standard methods for the problem of causal effects. We then apply the method to a large-scale field experiment re-ranking results on a search engine.

Table of Contents

  • 1 The Problem
  • 1.1 The Set Up
  • 1.2 Unconfoundedness
  • 1.3 Conditional Average Treatment Effects and Partitioning
  • 2 Honest Inference for Population Averages
  • 2.1 Set Up
  • 2.2 The Honest Target
  • 2.3 The Adaptive Target
  • 2.4 The Implementation of CART
  • 2.5 Honest Splitting
  • 2.6 Honest Crossvalidation
  • 3 Honest Inference for Treatment Effects
  • 3.1 Modifying Conventional CART for Treatment Effects
  • 3.2 Modifying the Honest Approach
  • 4 Four Partitioning Estimators for Causal Effects
  • 4.1 Causal Trees (CT)
  • 4.2 Transformed Outcome Trees (TOT)
  • 4.3 Fit-based Trees (F)
  • 4.4 Squared T-statistic Trees (TS)
  • 4.5 Comparison of the Causal Trees, the Fit Criterion, and the Squared t-statistic Criterion
  • 5 Inference
  • 6 A Simulation Study
  • 7 Observational Studies with Unconfoundedness
  • 8 The Literature
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — Honest Estimation Principle for Causal Partitioning

    model/method

    The honest estimation framework separates the process of selecting the partition structure (model selection) from the process of estimating treatment effects within the partition (estimation). Let (Yiobs,Wi,Xi)i=1N(Y_i^{obs}, W_i, X_i)_{i=1}^N be an independent and identically distributed sample of NN units, where Wi∈{0,1}W_i \in \{0, 1\} is the binary treatment indicator, Xi∈XX_i \in \mathcal{X} is a KK-dimensional vector of pretreatment covariates, and Yiobs=Yi(Wi)Y_i^{obs} = Y_i(W_i) is the realized potential outcome under the Rubin Causal Model with potential outcomes (Yi(0),Yi(1))(Y_i(0), Y_i(1)). The target of inference is the conditional average treatment effect (CATE), τ(x)≡E[Yi(1)−Yi(0)∣Xi=x]\tau(x) \equiv \mathbb{E}[Y_i(1) - Y_i(0) \mid X_i = x].

    Under conventional adaptive recursive partitioning, the same training sample is used both to choose tree splits and to compute leaf estimates. This introduces selection bias (data snooping), because units with extreme idiosyncratic error terms are grouped into the same leaves, resulting in biased leaf estimates and invalid, anti-conservative confidence intervals.

    In honest estimation, the data is split into two independent subsamples:

    1. A training sample Str\mathcal{S}^{tr} of size NtrN^{tr}, used exclusively to construct the tree partition Π={ℓ1,…,ℓ#(Π)}\Pi = \{\ell_1, \dots, \ell_{\#(\Pi)}\} and perform cross-validation/pruning.
    2. An estimation sample Sest\mathcal{S}^{est} of size NestN^{est}, used exclusively to estimate the conditional average treatment effect within each leaf ℓ∈Π\ell \in \Pi: τ^(x;Sest,Π)≡μ^(1,x;Sest,Π)−μ^(0,x;Sest,Π)\hat{\tau}(x; \mathcal{S}^{est}, \Pi) \equiv \hat{\mu}(1, x; \mathcal{S}^{est}, \Pi) - \hat{\mu}(0, x; \mathcal{S}^{est}, \Pi) where for w∈{0,1}w \in \{0, 1\} and Swest={i∈Sest:Wi=w}\mathcal{S}_w^{est} = \{i \in \mathcal{S}^{est} : W_i = w\}, μ^(w,x;Sest,Π)≡1#({i∈Swest:Xi∈ℓ(x;Π)})∑i∈Swest:Xi∈ℓ(x;Π)Yiobs\hat{\mu}(w, x; \mathcal{S}^{est}, \Pi) \equiv \frac{1}{\#(\{i \in \mathcal{S}_w^{est} : X_i \in \ell(x; \Pi)\})} \sum_{i \in \mathcal{S}_w^{est} : X_i \in \ell(x; \Pi)} Y_i^{obs} where ℓ(x;Π)\ell(x; \Pi) denotes the leaf containing covariate vector xx.

    Conditioning on the partition Π\Pi chosen by Str\mathcal{S}^{tr}, the estimates τ^(x;Sest,Π)\hat{\tau}(x; \mathcal{S}^{est}, \Pi) are unbiased for the leaf population average treatment effects τ(x;Π)≡E[Yi(1)−Yi(0)∣Xi∈ℓ(x;Π)]\tau(x; \Pi) \equiv \mathbb{E}[Y_i(1) - Y_i(0) \mid X_i \in \ell(x; \Pi)]. Consequently, standard statistical inference and confidence interval constructions hold conditionally on the tree structure without requiring sparsity assumptions or penalties on tree complexity.

  2. Knowl 2 — Honest Causal Tree Splitting Criterion

    equation

    For a training dataset Str\mathcal{S}^{tr} of size NtrN^{tr} with marginal treatment probability p=Ntreattr/Ntrp = N^{tr}_{treat} / N^{tr} and an anticipated independent estimation sample size NestN^{est}, the splitting objective function −EMSE^τ(Str,Π)-\widehat{\text{EMSE}}_\tau(\mathcal{S}^{tr}, \Pi) for an honest causal tree is given by:

    −EMSE^τ(Str,Π)≡1Ntr∑i∈Strτ^2(Xi;Str,Π)−(1Ntr+1Nest)∑ℓ∈Π(SStreattr2(ℓ)p+SScontroltr2(ℓ)1−p)-\widehat{\text{EMSE}}_\tau(\mathcal{S}^{tr}, \Pi) \equiv \frac{1}{N^{tr}} \sum_{i \in \mathcal{S}^{tr}} \hat{\tau}^2(X_i; \mathcal{S}^{tr}, \Pi) - \left( \frac{1}{N^{tr}} + \frac{1}{N^{est}} \right) \sum_{\ell \in \Pi} \left( \frac{S^2_{\mathcal{S}^{tr}_{treat}}(\ell)}{p} + \frac{S^2_{\mathcal{S}^{tr}_{control}}(\ell)}{1 - p} \right)

    When the training and estimation samples are equal in size (Ntr=NestN^{tr} = N^{est}), this specializes to:

    −EMSE^τ(Str,Π)=1Ntr∑i∈Strτ^2(Xi;Str,Π)−2Ntr∑ℓ∈Π(SStreattr2(ℓ)p+SScontroltr2(ℓ)1−p)-\widehat{\text{EMSE}}_\tau(\mathcal{S}^{tr}, \Pi) = \frac{1}{N^{tr}} \sum_{i \in \mathcal{S}^{tr}} \hat{\tau}^2(X_i; \mathcal{S}^{tr}, \Pi) - \frac{2}{N^{tr}} \sum_{\ell \in \Pi} \left( \frac{S^2_{\mathcal{S}^{tr}_{treat}}(\ell)}{p} + \frac{S^2_{\mathcal{S}^{tr}_{control}}(\ell)}{1 - p} \right)

    where:

    • Π\Pi is a partition of the covariate space into disjoint leaves ℓ\ell.
    • τ^(Xi;Str,Π)=μ^(1,Xi;Str,Π)−μ^(0,Xi;Str,Π)\hat{\tau}(X_i; \mathcal{S}^{tr}, \Pi) = \hat{\mu}(1, X_i; \mathcal{S}^{tr}, \Pi) - \hat{\mu}(0, X_i; \mathcal{S}^{tr}, \Pi) is the estimated leaf-level treatment effect evaluated on Str\mathcal{S}^{tr}.
    • Streattr(ℓ)\mathcal{S}^{tr}_{treat}(\ell) and Scontroltr(ℓ)\mathcal{S}^{tr}_{control}(\ell) denote the treated and control subsamples of Str\mathcal{S}^{tr} falling into leaf ℓ\ell.
    • SStreattr2(ℓ)S^2_{\mathcal{S}^{tr}_{treat}}(\ell) and SScontroltr2(ℓ)S^2_{\mathcal{S}^{tr}_{control}}(\ell) denote the within-leaf sample variances of the observed outcome YiobsY_i^{obs} for treated and control units, respectively.

    The criterion balances two objectives: the first term rewards partitions that capture high heterogeneity in treatment effects across leaves, while the second term penalizes partitions that inflate the sampling variance of treatment effect estimates constructed during the second-stage honest estimation.

  3. Knowl 3 — Expected Mean-Squared Error Objective for Conditional Average Treatment Effects

    theoretical result

    Let τi≡Yi(1)−Yi(0)\tau_i \equiv Y_i(1) - Y_i(0) denote the unit-level causal effect, and let τ(x;Π)≡E[τi∣Xi∈ℓ(x;Π)]\tau(x; \Pi) \equiv \mathbb{E}[\tau_i \mid X_i \in \ell(x; \Pi)] denote the conditional mean treatment effect in leaf ℓ(x;Π)\ell(x; \Pi) of partition Π\Pi. For an independent test sample Ste\mathcal{S}^{te} and estimation sample Sest\mathcal{S}^{est}, the infeasible adjusted mean-squared error for treatment effect prediction is defined as:

    MSEτ(Ste,Sest,Π)≡1#(Ste)∑i∈Ste((τi−τ^(Xi;Sest,Π))2−τi2)\text{MSE}_\tau(\mathcal{S}^{te}, \mathcal{S}^{est}, \Pi) \equiv \frac{1}{\#(\mathcal{S}^{te})} \sum_{i \in \mathcal{S}^{te}} \left( \left(\tau_i - \hat{\tau}(X_i; \mathcal{S}^{est}, \Pi)\right)^2 - \tau_i^2 \right)

    Taking expectations over independent draws of the test sample Ste\mathcal{S}^{te} and estimation sample Sest\mathcal{S}^{est}, the expected mean-squared error objective −EMSEτ(Π)≡−ESte,Sest[MSEτ(Ste,Sest,Π)]-\text{EMSE}_\tau(\Pi) \equiv -\mathbb{E}_{\mathcal{S}^{te}, \mathcal{S}^{est}}[\text{MSE}_\tau(\mathcal{S}^{te}, \mathcal{S}^{est}, \Pi)] decomposes as:

    −EMSEτ(Π)=EXi[τ2(Xi;Π)]−ESest,Xi[V(τ^(Xi;Sest,Π))]-\text{EMSE}_\tau(\Pi) = \mathbb{E}_{X_i}[\tau^2(X_i; \Pi)] - \mathbb{E}_{\mathcal{S}^{est}, X_i}\left[\mathbb{V}\left(\hat{\tau}(X_i; \mathcal{S}^{est}, \Pi)\right)\right]

    where V(τ^(Xi;Sest,Π))\mathbb{V}(\hat{\tau}(X_i; \mathcal{S}^{est}, \Pi)) is the sampling variance of the estimator on sample Sest\mathcal{S}^{est} conditional on Xi∈ℓ(Xi;Π)X_i \in \ell(X_i; \Pi).

    This decomposition enables recursive partitioning without observing the unobserved individual counterfactual τi\tau_i, because both EXi[τ2(Xi;Π)]\mathbb{E}_{X_i}[\tau^2(X_i; \Pi)] and ESest,Xi[V(τ^(Xi;Sest,Π))]\mathbb{E}_{\mathcal{S}^{est}, X_i}[\mathbb{V}(\hat{\tau}(X_i; \mathcal{S}^{est}, \Pi))] can be estimated without bias using the observed training sample Str\mathcal{S}^{tr} and within-leaf treatment and control sample variances.

  4. Knowl 4 — Causal Tree Splitting and Cross-Validation Framework

    model/method

    Causal Tree (CT) estimators adapt classification and regression trees to directly target conditional average treatment effect heterogeneity. Two variants are defined:

    1. Adaptive Causal Trees (CT-A):
    • Splitting criterion: Maximizes the in-sample treatment effect goodness of fit: −MSE^τ(Str,Str,Π)=1Ntr∑i∈Strτ^2(Xi;Str,Π)-\widehat{\text{MSE}}_\tau(\mathcal{S}^{tr}, \mathcal{S}^{tr}, \Pi) = \frac{1}{N^{tr}} \sum_{i \in \mathcal{S}^{tr}} \hat{\tau}^2(X_i; \mathcal{S}^{tr}, \Pi)
    • Cross-validation criterion: Pruning complexity is selected on cross-validation folds (where Str,tr\mathcal{S}^{tr,tr} builds the tree and Str,cv\mathcal{S}^{tr,cv} evaluates it) using the unbiased cross-validation target: −MSE^τ(Str,cv,Str,tr,Π)=−2Ntr,cv∑i∈Str,cvτ^(Xi;Str,cv,Π)⋅τ^(Xi;Str,tr,Π)+1Ntr,cv∑i∈Str,cvτ^2(Xi;Str,tr,Π)-\widehat{\text{MSE}}_\tau(\mathcal{S}^{tr,cv}, \mathcal{S}^{tr,tr}, \Pi) = -\frac{2}{N^{tr,cv}} \sum_{i \in \mathcal{S}^{tr,cv}} \hat{\tau}(X_i; \mathcal{S}^{tr,cv}, \Pi) \cdot \hat{\tau}(X_i; \mathcal{S}^{tr,tr}, \Pi) + \frac{1}{N^{tr,cv}} \sum_{i \in \mathcal{S}^{tr,cv}} \hat{\tau}^2(X_i; \mathcal{S}^{tr,tr}, \Pi)
    1. Honest Causal Trees (CT-H):
    • Splitting criterion: Maximizes the honest expected mean-squared error estimator −EMSE^τ(Str,Π)-\widehat{\text{EMSE}}_\tau(\mathcal{S}^{tr}, \Pi), which explicitly subtracts the expected estimation variance of the second-stage sample Sest\mathcal{S}^{est}.
    • Cross-validation criterion: Evaluated solely on the cross-validation sample Str,cv\mathcal{S}^{tr,cv}: −EMSE^τ(Str,cv,Π)-\widehat{\text{EMSE}}_\tau(\mathcal{S}^{tr,cv}, \Pi) Unlike CT-A, CT-H does not compute cross-products between Str,tr\mathcal{S}^{tr,tr} and Str,cv\mathcal{S}^{tr,cv} because honest estimation anticipates that final leaf treatment effects will be re-estimated on an independent estimation sample Sest\mathcal{S}^{est} rather than Str,tr\mathcal{S}^{tr,tr}.
  5. Knowl 5 — Decomposition of Honest Causal Tree Splitting Criterion

    theoretical result

    Consider a parent leaf being evaluated for a single binary split based on covariate Xi∈{L,R}X_i \in \{L, R\}, yielding unsplit partition ΠN\Pi_N and split partition ΠS={L,R}\Pi_S = \{L, R\}. Let NN be the sample size in the parent leaf, and let p=Ntreat/Np = N_{treat} / N. Let S~2\tilde{S}^2 denote the outcome sample variance in the parent leaf without splitting, and S2S^2 the pooled within-leaf sample variance given the split.

    Define the squared tt-statistics for testing differences in mean outcomes across child leaves LL and RR separately for control (w=0w=0) and treated (w=1w=1) units: Tw2≡(YˉLw−YˉRw)2S2/NLw+S2/NRw,w∈{0,1}T_w^2 \equiv \frac{(\bar{Y}_{Lw} - \bar{Y}_{Rw})^2}{S^2 / N_{Lw} + S^2 / N_{Rw}}, \quad w \in \{0, 1\} where YˉLw,YˉRw\bar{Y}_{Lw}, \bar{Y}_{Rw} are the subgroup mean outcomes and NLw,NRwN_{Lw}, N_{Rw} are the subgroup sample sizes.

    Define the overall outcome goodness-of-fit improvement FF from splitting the leaf as: F≡S~2⋅2(T02+T12)1+2(T02+T12)/NF \equiv \tilde{S}^2 \cdot \frac{2(T_0^2 + T_1^2)}{1 + 2(T_0^2 + T_1^2) / N}

    Define the squared tt-statistic T2T^2 for testing the null hypothesis of equal treatment effects between child leaves (where Yˉℓ≡Yˉℓ1−Yˉℓ0\bar{Y}_\ell \equiv \bar{Y}_{\ell 1} - \bar{Y}_{\ell 0} for ℓ∈{L,R}\ell \in \{L, R\}): T2≡N⋅(YˉL−YˉR)2S2/NL+S2/NRT^2 \equiv N \cdot \frac{(\bar{Y}_L - \bar{Y}_R)^2}{S^2 / N_L + S^2 / N_R}

    Ignoring degrees-of-freedom corrections, the improvement in the honest causal tree criterion from splitting is:

    EMSE^τ(S,ΠN)−EMSE^τ(S,ΠS)=(T2−4)(S~2−F/N)+2S~2p(1−p)\widehat{\text{EMSE}}_\tau(\mathcal{S}, \Pi_N) - \widehat{\text{EMSE}}_\tau(\mathcal{S}, \Pi_S) = \frac{(T^2 - 4)(\tilde{S}^2 - F / N) + 2\tilde{S}^2}{p(1 - p)}

    This demonstrates that the CT-H criterion primary scales with T2T^2 (treatment effect heterogeneity), but also rewards splits that produce higher outcome fit FF, since improved fit reduces the residual variance S~2−F/N\tilde{S}^2 - F/N and thereby reduces the variance of second-stage treatment effect estimates.

  6. Knowl 6 — Alternative Partitioning Methods for Heterogeneous Treatment Effects

    model/method

    Three alternative tree-based methods for estimating heterogeneous treatment effects serve as benchmarks against Causal Trees:

    1. Transformed Outcome Trees (TOT): Defines a transformed outcome variable Yi∗≡Yiobs−Wip(1−p)Y_i^* \equiv \frac{Y_i^{obs} - W_i}{p(1 - p)}, where p=P(Wi=1)p = \mathbb{P}(W_i = 1). Because E[Yi∗∣Xi=x]=τ(x)\mathbb{E}[Y_i^* \mid X_i = x] = \tau(x), standard CART can be applied directly to Yi∗Y_i^* without algorithmic modifications. TOT-A uses CART directly; TOT-H uses the tree structure from TOT-A but re-estimates leaf averages on an independent estimation sample Sest\mathcal{S}^{est}. TOT is inefficient because it discards the treatment indicator during tree construction, leading to high estimation variance unless treated proportions are exactly pp in every leaf.

    2. Fit-based Trees (F): Constructs trees by fitting a linear model Yiobs=αℓ+τℓWi+ϵiY_i^{obs} = \alpha_\ell + \tau_\ell W_i + \epsilon_i within each leaf ℓ\ell, optimizing outcome goodness of fit MSEμ,W(Ste,Sest,Π)≡∑i∈Ste((Yiobs−μ^W(Wi,Xi;Sest,Π))2−(Yiobs)2)\text{MSE}_{\mu, W}(\mathcal{S}^{te}, \mathcal{S}^{est}, \Pi) \equiv \sum_{i \in \mathcal{S}^{te}} ((Y_i^{obs} - \hat{\mu}_W(W_i, X_i; \mathcal{S}^{est}, \Pi))^2 - (Y_i^{obs})^2). F-A applies adaptive CART to this fit criterion, while F-H uses variance-adjusted splitting and independent estimation. A major drawback of the F approach is that it splits on covariates that affect baseline outcome levels even when those covariates do not affect treatment effect heterogeneity.

    3. Squared T-statistic Trees (TS): Splits nodes recursively to maximize the squared two-sample tt-statistic T2=N(YˉL−YˉR)2S2/NL+S2/NRT^2 = N \frac{(\bar{Y}_L - \bar{Y}_R)^2}{S^2/N_L + S^2/N_R} testing the hypothesis that treatment effects are identical across child leaves LL and RR. Cross-validation for TS-A and TS-H is conducted using CT-A and CT-H criteria, respectively. A limitation of TS is that it assigns zero value to splits that improve outcome prediction fit without altering treatment effects, missing variance reduction benefits.

  7. Knowl 7 — Discretized Bucket Splitting Algorithm for Balanced Causal Trees

    algorithm

    Standard recursive partitioning evaluates all unique covariate values as candidate split points, which causes severe sampling instability in treatment effect estimation when single units with extreme baseline outcomes shift across split boundaries. To stabilize splitting and prevent spurious splits, the search space over split points is discretized using balanced buckets of treated and control units.

    Input: Node dataset S\mathcal{S} containing (Yiobs,Wi,Xi)(Y_i^{obs}, W_i, X_i), covariate index k∈{1,…,K}k \in \{1, \dots, K\}, target bucket size bb (default b=4b = 4), minimum leaf size per treatment group nmn_m (default nm=25n_m = 25).
    Output: Best split threshold c∗c^* for covariate XkX_k.
    Separate the node sample into treated units Streat\mathcal{S}_{treat} and control units Scontrol\mathcal{S}_{control}
    Sort Streat\mathcal{S}_{treat} in ascending order of Xi,kX_{i, k}
    Sort Scontrol\mathcal{S}_{control} in ascending order of Xi,kX_{i, k}
    Determine the number of buckets B=min⁡(⌊∣Streat∣/b⌋,⌊∣Scontrol∣/b⌋)B = \min(\lfloor |\mathcal{S}_{treat}| / b \rfloor, \lfloor |\mathcal{S}_{control}| / b \rfloor)
    if B<nmB < n_m then
        Adjust bucket size bb downward such that at least nmn_m buckets are formed
        Recalculate BB
    Partition Streat\mathcal{S}_{treat} into BB ordered buckets {Btreat,1,…,Btreat,B}\{B_{treat, 1}, \dots, B_{treat, B}\}, each containing approximately bb units
    Partition Scontrol\mathcal{S}_{control} into BB ordered buckets {Bcontrol,1,…,Bcontrol,B}\{B_{control, 1}, \dots, B_{control, B}\}, each containing approximately bb units
    for bucket index j=1j = 1 to B−1B - 1 do
        Define candidate left child leaf LL containing buckets 11 through jj of both treatment and control groups
        Define candidate right child leaf RR containing buckets j+1j+1 through BB of both treatment and control groups
        if ∣L∩Streat∣≥nm|L \cap \mathcal{S}_{treat}| \ge n_m and ∣L∩Scontrol∣≥nm|L \cap \mathcal{S}_{control}| \ge n_m and ∣R∩Streat∣≥nm|R \cap \mathcal{S}_{treat}| \ge n_m and ∣R∩Scontrol∣≥nm|R \cap \mathcal{S}_{control}| \ge n_m then
            Evaluate split criterion Δ(j)\Delta(j) (e.g., −EMSE^τ-\widehat{\text{EMSE}}_\tau for CT-H)
    Find bucket index j∗=arg⁡max⁡jΔ(j)j^* = \arg\max_j \Delta(j)
    Set vtreat=max⁡i∈Btreat,j∗Xi,kv_{treat} = \max_{i \in B_{treat, j^*}} X_{i, k}
    Set vcontrol=max⁡i∈Bcontrol,j∗Xi,kv_{control} = \max_{i \in B_{control, j^*}} X_{i, k}
    Set split point c∗=(vtreat+vcontrol)/2c^* = (v_{treat} + v_{control}) / 2
    return c∗c^*
  8. Knowl 8 — Empirical Performance and Confidence Interval Coverage in Simulation Experiments

    data/table

    A simulation study with 6,000 test observations evaluates the four estimators across three data-generating designs with Yi(w)=η(Xi)+12(2w−1)κ(Xi)+ϵiY_i(w) = \eta(X_i) + \frac{1}{2}(2w - 1)\kappa(X_i) + \epsilon_i, where ϵi∼N(0,0.01)\epsilon_i \sim \mathcal{N}(0, 0.01) and Xi,k∼N(0,1)X_{i, k} \sim \mathcal{N}(0, 1):

    • Design 1 (K=2K = 2): η(x)=12x1+x2\eta(x) = \frac{1}{2}x_1 + x_2, κ(x)=12x1\kappa(x) = \frac{1}{2}x_1.
    • Design 2 (K=10K = 10): η(x)=12(x1+x2)+∑k=36xk\eta(x) = \frac{1}{2}(x_1 + x_2) + \sum_{k=3}^6 x_k, κ(x)=∑k=121{xk>0}xk\kappa(x) = \sum_{k=1}^2 \mathbf{1}\{x_k > 0\} x_k.
    • Design 3 (K=20K = 20): η(x)=12∑k=14xk+∑k=58xk\eta(x) = \frac{1}{2}\sum_{k=1}^4 x_k + \sum_{k=5}^8 x_k, κ(x)=∑k=141{xk>0}xk\kappa(x) = \sum_{k=1}^4 \mathbf{1}\{x_k > 0\} x_k.
    Design 1 2 3
    Ntr=NestN^{tr} = N^{est} 500 1000 500 1000 500 1000
    Number of Leaves
    TOT 2.8 3.4 2.1 2.7 4.7 6.1
    F-A 6.1 13.2 6.3 13.1 6.1 13.2
    TS-A 4.0 5.6 2.5 3.3 4.4 8.9
    CT-A 4.0 5.7 2.3 2.5 4.5 6.2
    F-H 6.1 13.2 6.4 13.3 6.3 13.4
    TS-H 4.4 7.7 5.3 11.0 6.0 12.3
    CT-H 4.2 7.5 5.3 11.2 6.2 12.3
    Infeasible MSEτ\text{MSE}_\tau Divided by MSEτ\text{MSE}_\tau for CT-H
    TOT-H 1.77 2.12 1.03 1.04 1.03 1.05
    F-A 1.93 1.54 1.69 2.07 1.63 2.08
    TS-H 1.01 1.02 1.06 0.99 1.24 1.38
    CT-H 1.00 1.00 1.00 1.00 1.00 1.00
    Ratio of Infeasible MSE: Honest to Adaptive
    TOT-H / TOT-A 0.99 0.86 0.76
    F-H / F-A 0.50 0.98 0.91
    TS-H / TS-A 0.92 0.90 0.85
    CT-H / CT-A 0.91 0.93 0.76
    Coverage of 90% Confidence Intervals - Adaptive
    TOT-A 0.83 0.86 0.83 0.83 0.74 0.79
    F-A 0.89 0.89 0.86 0.86 0.82 0.82
    TS-A 0.85 0.85 0.80 0.83 0.77 0.80
    CT-A 0.85 0.85 0.81 0.83 0.80 0.81
    Coverage of 90% Confidence Intervals - Honest
    TOT-H 0.90 0.89 0.90 0.92 0.89 0.89
    F-H 0.91 0.90 0.90 0.90 0.90 0.89
    TS-H 0.89 0.90 0.90 0.90 0.90 0.90
    CT-H 0.90 0.90 0.90 0.89 0.90 0.90

    The table demonstrates two key conclusions:

    1. Nominal Coverage: All honest estimators (CT-H, TOT-H, F-H, TS-H) achieve empirical 90% confidence interval coverage of 89%–92% across all sample sizes and dimensions. In contrast, adaptive estimators suffer severe undercoverage (74%–89%) due to post-selection data reuse.
    2. Estimation Quality: CT-H achieves lower or comparable infeasible MSE compared to alternative honest estimators. Furthermore, despite splitting the effective sample size in half (Ntr=Nest=500N^{tr} = N^{est} = 500 vs. adaptive Ntr=1000N^{tr} = 1000), CT-H achieves a lower MSE than CT-A (ratios 0.76–0.93) because honesty eliminates estimation bias.
  9. Knowl 9 — Adaptation of Honest Partitioning to Observational Studies Under Unconfoundedness

    model/method

    When assignment to treatment is not randomized, honest recursive partitioning extends to observational studies under the unconfoundedness assumption:

    Wi⊥(Yi(0),Yi(1))∣XiW_i \perp (Y_i(0), Y_i(1)) \mid X_i

    Let e(x)≡P(Wi=1∣Xi=x)e(x) \equiv \mathbb{P}(W_i = 1 \mid X_i = x) denote the propensity score. To correct for confounding within each leaf ℓ(x;Π)\ell(x; \Pi) of partition Π\Pi, the honest leaf treatment effect estimate τ^(x;Sest,Π)=μ^(1,x;Sest,Π)−μ^(0,x;Sest,Π)\hat{\tau}(x; \mathcal{S}^{est}, \Pi) = \hat{\mu}(1, x; \mathcal{S}^{est}, \Pi) - \hat{\mu}(0, x; \mathcal{S}^{est}, \Pi) is computed using normalized inverse propensity weighting within the estimation sample Sest\mathcal{S}^{est}:

    μ^(1,x;Sest,Π)≡∑i∈S1est:Xi∈ℓ(x;Π)Yiobs/e(Xi)∑i∈S1est:Xi∈ℓ(x;Π)1/e(Xi)\hat{\mu}(1, x; \mathcal{S}^{est}, \Pi) \equiv \frac{\sum_{i \in \mathcal{S}^{est}_1 : X_i \in \ell(x; \Pi)} Y_i^{obs} / e(X_i)}{\sum_{i \in \mathcal{S}^{est}_1 : X_i \in \ell(x; \Pi)} 1 / e(X_i)}

    μ^(0,x;Sest,Π)≡∑i∈S0est:Xi∈ℓ(x;Π)Yiobs/(1−e(Xi))∑i∈S0est:Xi∈ℓ(x;Π)1/(1−e(Xi))\hat{\mu}(0, x; \mathcal{S}^{est}, \Pi) \equiv \frac{\sum_{i \in \mathcal{S}^{est}_0 : X_i \in \ell(x; \Pi)} Y_i^{obs} / (1 - e(X_i))}{\sum_{i \in \mathcal{S}^{est}_0 : X_i \in \ell(x; \Pi)} 1 / (1 - e(X_i))}

    where S1est\mathcal{S}^{est}_1 and S0est\mathcal{S}^{est}_0 are the treated and control subsamples of Sest\mathcal{S}^{est}. Extreme propensity scores may be trimmed to improve stability.

    Because the honest framework strictly separates partition selection (on Str\mathcal{S}^{tr}) from estimation (on Sest\mathcal{S}^{est}), standard asymptotic normality results for propensity-score weighted estimators apply directly to the leaf estimates τ^(x;Sest,Π)\hat{\tau}(x; \mathcal{S}^{est}, \Pi) without post-selection adjustments.

Coverage note — None was omitted; all primary methodological contributions, theoretical decompositions, algorithm modifications, and empirical simulation tables and findings are fully represented.

References

  1. 1.A. Abadie and G. Imbens, Large Sample Properties of Matching Estimators for Average Treatment Effects, Econometrica, 74(1), 235-267.
  2. 2.A. Beygelzimer and J. Langford, The Offset Tree for Learning with Partial Labels, http://arxiv.org/pdf/0812.4044v2.pdf, (2009).
  3. 3.L. Breiman, Random forests, Machine Learning, 45, (2001), 5-32.
  4. 4.L. Breiman, J. Friedman, R. Olshen, and C. Stone, Classification and Regression Trees, (1984), Wadsworth.
  5. 5.R. Crump, R., J. Hotz, G. Imbens, and O. Mitnik, Nonparametric Tests for Treatment Effect Heterogeneity, Review of Economics and Statistics, 90(3), (2008), 389-405.
  6. 6.M. Dudik, J. Langford and L. Li, Doubly Robust Policy Evaluation and Learning , Proceedings of the 28th International Conference on Machine Learning (ICML-11), (2011).
  7. 7.J. Foster, J. Taylor and S. Ruberg, Subgroup Identification from Randomized Clinical Data, Statistics in Medicine, 30, (2010), 2867-2880.
  8. 8.Green, D., and H. Kern, (2010), Detecting Heterogeneous Treatment Effects in Large-Scale Experiments Using Bayesian Additive Regression Trees, Unpublished Manuscript, Yale University.
  9. 9.T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, Second Edition, (2011), Springer.
  10. 10.K. Hirano, G. Imbens and G. Ridder, Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score, Econometrica, 71 (4), (2003), 1161-1189.
  11. 11.P. Holland, Statistics and Causal Inference (with discussion), Journal of the American Statistical Association, 81, (1986), 945-970.
  12. 12.D. Horvitz, and D. Thompson, A generalization of sampling without replacement from a finite universe, Journal of the American Statistical Association, Vol. 47, (1952), 663685.
  13. 13.K. Imai and M. Ratkovic, Estimating Treatment Effect Heterogeneity in Randomized Program Evaluation, Annals of Applied Statistics, 7(1), (2013), 443-470.
  14. 14.G. Imbens and D. Rubin, Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction, Cambridge University Press, (2015).
  15. 15.J. Pearl, Causality: Models, Reasoning and Inference, Cambridge University Press, (2000).
  16. 16.P. Rosenbaum, Observational Studies, (2002), Springer.
  17. 17.P. Rosenbaum and D. Rubin, The Central Role of the Propensity Score in Observational Studies for Causal Effects, Biometrika, 70, (1983), 41-55.
  18. 18.M. Rosenblum and M. Van Der Laan., Optimizing Randomized Trial Designs to Distinguish which Subpopulations Benefit from Treatment , Biometrika, 98(4), (2011), 845-860.
  19. 19.D. Rubin, Estimating Causal Effects of Treatments in Randomized and Non-randomized Studies Journal of Educational Psychology, 66, (1974), 688-701.
  20. 20.D. Rubin, Bayesian inference for causaleffects: The Role of Randomization, Annals of Statistics, 6, (1978), 34-58.
  21. 21.X. Su, C. Tsai, H. Wang, D. Nickerson, and B. Li, Subgroup Analysis via Recursive Partitioning, Journal of Machine Learning Research, 10, (2009), 141-158.
  22. 22.J. Signovitch, J., Identifying informative biological markers in high-dimensional genomic data and clinical trials, PhD Thesis, Department of Biostatistics, Harvard University, (2007).
  23. 23.L. Tian, A. Alizadeh, A. Gentles, and R. Tibshirani, A Simple Method for Estimating Interactions Between a Treatment and a Large Number of Covariates, Journal of the American Statistical Association, 109(508), (2014) 1517-1532.
  24. 24.R. Tibshirani, Regression shrinkage and selection via the lasso, Journal of the Royal Statistical Society. Series B (Methodological), Volume 58, Issue 1. (1996), 267-288.
  25. 25.M. Taddy, M. Gardner, L. Chen, and D. Draper,, Heterogeneous Treatment Effects in Digital Experimentation, Unpublished Manuscript, (2015), arXiv:1412.8563.
  26. 26.V. Vapnik, Statistical Learning Theory, Wiley, (1998).
  27. 27.M. Van Der Laan, and S. Rose, Targeted Learning: Causal Inference for Observational and Experimental Data, Springer, (2011).
  28. 28.S. Wager, and S. Athey, Estimation and Inference of Heterogeneous Treatment Effects using Random Forests, http://arxiv.org/pdf/1510.04342v2.pdf, (2015).
  29. 29.H. Weisburg, H. and V. Pontes, Post hoc subgroups in Clinical Trials: Anathema or Analytics? Clinical Trials, June, 2015.
  30. 30.A. Zeileis, T. Hothorn, and K. Hornik, Model-based recursive partitioning. Journal of Computational and Graphical Statistics, 17(2), (2008), 492-514.

Citation

MLA
Athey, S., and G. Imbens. “Recursive Partitioning for Heterogeneous Causal Effects”. Proceedings of the National Academy of Sciences, vol. 113, no. 27, 2016, pp. 7353–60, https://doi.org/10.1073/pnas.1510489113.
APA
Athey, S., & Imbens, G. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27), 7353–7360. https://doi.org/10.1073/pnas.1510489113
Chicago
Athey, S., and G. Imbens. 2016. “Recursive Partitioning for Heterogeneous Causal Effects”. Proceedings of the National Academy of Sciences 113 (27): 7353–60. https://doi.org/10.1073/pnas.1510489113.
Harvard
Athey, S. and Imbens, G. (2016) “Recursive partitioning for heterogeneous causal effects”, Proceedings of the National Academy of Sciences, 113(27), pp. 7353–7360. Available at: https://doi.org/10.1073/pnas.1510489113.
Vancouver
1. Athey S, Imbens G (2016) Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113:7353–7360

BibTeX

@article{Athey_2016, title={Recursive partitioning for heterogeneous causal effects}, volume={113}, ISSN={1091-6490}, url={http://dx.doi.org/10.1073/pnas.1510489113}, DOI={10.1073/pnas.1510489113}, number={27}, journal={Proceedings of the National Academy of Sciences}, publisher={National Academy of Sciences}, author={Athey, Susan and Imbens, Guido}, year={2016}, month=July, pages={7353–7360} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF