4 Potential Outcomes and Adjustment
4.1 Motivation: A Third Language for Causality
The first three chapters developed causal inference primarily in the languages of structural equations and directed acyclic graphs. Those two languages are especially effective for expressing interventions syntactically: the structural equation model shows how an intervention replaces an assignment mechanism, and the DAG shows how the same intervention deletes incoming arrows into the treatment node. In this chapter we introduce a third language: the potential outcomes framework of Neyman (1923) and Rubin (1974).
Potential outcomes provide a direct language for defining causal estimands. Quantities such as the average treatment effect, the average treatment effect on the treated, and related causal contrasts are most naturally written in terms of counterfactual outcomes \(Y(1)\), \(Y(0)\), and more generally \(Y(t)\). For this reason, the potential outcomes framework has become standard in statistics, biostatistics, epidemiology, and much of econometrics.
Assumptions such as ignorability and exclusion can certainly be stated in this language, but DAGs and the do-operator often make their structural content more transparent by showing which paths are blocked, which variables are confounders, and which interventions are being considered. Positivity, by contrast, is a support condition on the observed-data distribution and is not encoded by any graph. Thus the two frameworks should be viewed as complementary rather than competing: potential outcomes define the causal quantities of interest, while DAGs and the do-calculus clarify identification.
The purpose of this chapter is therefore not to replace the graphical viewpoint developed earlier, but to connect it to the potential-outcomes notation that dominates much of the applied literature. We first define potential outcomes and the standard causal estimands built from them. We then show how consistency, exchangeability, and positivity together identify the average treatment effect, and how the back-door criterion provides the graphical justification for that identification.
Section 4.2 defines the potential outcome framework rigorously. Section 4.3 develops the two principal causal estimands: ATE and ATT. Section 4.4 develops the identification chain from consistency, exchangeability, and positivity to the adjustment formula, and explains how the back-door criterion provides the graphical justification. Section 4.5 situates the potential outcomes framework relative to the do-calculus and the SEM. For readers who want a formal graphical representation of counterfactual variables, Appendix B introduces single-world intervention graphs (SWIGs).
4.2 The Neyman–Rubin Potential Outcomes Framework
4.2.1 The Potential Outcome
Interpretation. \(Y_i(1)\) is the outcome unit \(i\) would achieve under treatment; \(Y_i(0)\) is the outcome under control. Only one of these is ever observed for any given unit. The unobserved potential outcome is called the counterfactual.
The structure of the problem is visible in a small table — often called the science table — listing the full potential-outcome schedule next to what is actually observed:
| Unit | \(Y_i(0)\) | \(Y_i(1)\) | \(T_i\) | \(Y_i = Y_i(T_i)\) |
|---|---|---|---|---|
| 1 | observed | ? | 0 | \(Y_1(0)\) |
| 2 | ? | observed | 1 | \(Y_2(1)\) |
| 3 | observed | ? | 0 | \(Y_3(0)\) |
Each row is missing exactly one entry, and which entry is missing is determined by \(T_i\): this is the fundamental problem in tabular form. The final column is the consistency equation at work, and the impossibility of filling in the “?” entries within a row is why causal inference must rely on cross-unit comparisons — which is precisely what the exchangeability assumptions of Section 4.4 will license.
4.2.2 SUTVA, Treatment Versions, and Consistency
The potential outcome \(Y_i(t)\) is only well-defined if it does not depend on the treatments assigned to other units, and if there is a single, unambiguous version of each treatment level. These requirements are formalized as SUTVA.
Consistency is the bridge between the potential-outcome world and the observed-data world, and it is conceptually distinct from the exchangeability and positivity assumptions of Section 4.4. It is best regarded as an assumption in its own right rather than as an automatic consequence of SUTVA, although SUTVA is what makes the scalar notation \(Y_i(t)\) in Equation 4.1 adequate.
Two clarifications about SUTVA itself are worth recording. First, interference does not make potential outcomes meaningless; it makes the scalar notation inadequate. Under interference one writes \(Y_i(t_1, \ldots, t_n)\), or reduces the assignment vector through an exposure mapping, and causal analysis proceeds with the richer index set. Second, “no hidden versions” does not require the physical treatment received by any two treated units to be literally identical; it requires that any outcome-relevant versions be either encoded in the treatment variable or causally equivalent for the estimand at hand. When these conditions fail — vaccination that provides herd immunity, or a label \(T{=}1\) that covers interventions with different effects — it is the scalar consistency equation Equation 4.1 that breaks down.
4.2.3 Connection to the Do-Operator
The notation \(Y(t)\) and the expression \(P(Y \mid \doop(T{=}t))\) refer to the same causal idea viewed from two complementary angles. The potential outcome \(Y(t)\) denotes the outcome that would be realized for a unit if the treatment were set to \(t\). The do-operator, by contrast, describes an intervention at the level of the data-generating mechanism: it replaces the original treatment assignment rule by the constant value \(t\).
When \(Y(t)\) is constructed by applying the same structural intervention that defines \(\doop(T{=}t)\), equality of the resulting laws holds by construction: intervening to set \(T = t\) in the underlying structural equation model produces an outcome variable with the same distribution as \(Y(t)\). No interference makes the scalar unit-level notation \(Y_i(t)\) adequate. The separate consistency assumption links these counterfactual variables to observed outcomes when the realized treatment equals \(t\). Thus the potential-outcomes notation and the interventional distribution are two ways of encoding the same manipulated world.
Proposition 4.1 (Potential Outcomes and the Do-Operator) Consider the structural causal model of Chapters 1–3, and construct \(Y(t)\) by replacing the structural equation for \(T\) with the constant assignment \(T := t\) — the same intervention that defines \(\doop(T{=}t)\). Under no interference (SUTVA), the potential outcome has the interventional distribution: \[P\{Y(t) \le y\} \;=\; P\{Y \le y \mid \doop(T{=}t)\} \qquad \text{for every } y; \tag{4.2}\] in particular, \(\E[Y(t)] = \E[Y \mid \doop(T{=}t)]\).
Proof sketch. In the intervened SCM, whose graph is \(\Gcal_{\overline{T}}\), the structural equation for \(T\) is replaced by \(T := t\) for every unit. The structural model then determines the value of \(Y\) for unit \(i\) from \(t\) and unit \(i\)’s own background variables — which is precisely what the potential outcomes framework defines as \(Y_i(t)\). SUTVA’s no-interference condition ensures \(Y_i(t)\) does not depend on other units’ treatment values, so the marginal distribution of \(Y\) in the intervened model equals the marginal distribution of \(Y(t)\). By definition of the do-operator, the former is \(f(y \mid \doop(T{=}t))\). \(\square\)
This proposition should be interpreted carefully. Consistency alone does not create the link between \(Y(t)\) and the do-operator; the link comes from the underlying causal model that assigns meaning to the intervention \(\doop(T{=}t)\). Consistency and no interference ensure that the observed outcome agrees with the appropriate potential outcome at the realized treatment, while the structural model explains how the intervention generates the counterfactual world.
This equivalence matters because it lets us move freely between two notational traditions. When defining estimands such as \(\E[Y(1) - Y(0)]\), potential-outcome notation is often most natural. When proving identification results from a graph, the do-operator is often more convenient because it integrates directly with graph surgery and the do-calculus. In general \(f(y \mid \doop(t))\) differs from \(f(y \mid T{=}t)\) when \(T\) is endogenous (Chapter 1), although equality can occur through special cancellations. Potential-outcome notation marks the same distinction — \(\E[Y(t)]\) and \(\E[Y \mid T{=}t]\) are different expressions — but it does not by itself supply the graphical calculus for deciding when the two coincide. In a well-specified causal model these are not competing definitions, but compatible representations of the same intervention.
4.3 Causal Estimands
4.3.1 The Average Treatment Effect and Its Relatives
Let \(\tau = Y(1) - Y(0)\) denote the unit-level effect and let \(p = P(T{=}1) \in (0,1)\). Alongside the ATT one may define the average treatment effect on the controls, \(\tau_{\mathrm{ATC}} = \E[\tau \mid T{=}0]\). The law of total expectation gives the exact decomposition \[\tau_{\mathrm{ATE}} \;=\; p\,\tau_{\mathrm{ATT}} + (1-p)\,\tau_{\mathrm{ATC}}, \tag{4.3}\] so the ATE and ATT coincide if and only if \(\tau_{\mathrm{ATT}} = \tau_{\mathrm{ATC}}\): the mean effect among the treated must equal the mean effect among the controls. Equivalently, \(\tau_{\mathrm{ATT}} - \tau_{\mathrm{ATE}} = \mathrm{Cov}(T, \tau)/p\). Treatment effects may therefore be heterogeneous across units while the ATE and ATT still coincide; what matters is whether selection into treatment is related to the effect in the mean. Unit-level constancy of the effect is sufficient but not necessary.
None of these quantities is directly observable. The naive estimator \(\hat\tau_{\mathrm{naive}} = \bar{Y}_{T=1} - \bar{Y}_{T=0}\) estimates the observed treatment-group contrast \(D_{\mathrm{obs}} = \E[Y \mid T{=}1] - \E[Y \mid T{=}0]\), which in general differs from the ATE. Under consistency, \(D_{\mathrm{obs}} = \E[Y(1) \mid T{=}1] - \E[Y(0) \mid T{=}0]\), and adding and subtracting \(\E[Y(0) \mid T{=}1]\) yields the exact decomposition \[D_{\mathrm{obs}} - \tau_{\mathrm{ATE}} \;=\; \underbrace{\E[Y(0) \mid T{=}1] - \E[Y(0) \mid T{=}0]}_{\text{baseline-selection bias}} \;+\; \underbrace{\bigl(\tau_{\mathrm{ATT}} - \tau_{\mathrm{ATE}}\bigr)}_{\text{effect-selection difference}}. \tag{4.4}\] The first term reflects baseline differences between the two groups — units who select into treatment may have fared differently even untreated — and the second reflects selection on the size of the effect. In a randomized experiment both terms vanish. Problem 2 works through a numerical case.
4.3.2 Causal Estimands as Functionals of the Structural Model
The potential outcome \(Y_i(t)\) was introduced as a hypothetical: the value unit \(i\) would have exhibited under treatment \(t\). The SEM framework of Chapter 1 gives this hypothetical a precise generative meaning, and thereby expresses the potential-outcome estimands as functionals of the structural equation for \(Y\).
From structural equation to potential outcome. In the SEM, the outcome is determined by a structural equation \[Y \;=\; g(T,\, \mathbf{X},\, U_Y), \tag{4.5}\] where \(U_Y\) collects all sources of variation in \(Y\) not already accounted for by \((T, \mathbf{X})\). The potential outcome under \(\doop(T{=}t)\) is obtained by substituting \(t\) for \(T\) while holding everything else fixed: \[Y_i(t) \;=\; g(t,\, \mathbf{X}_i,\, U_{Y,i}). \tag{4.6}\] Here we assume, as throughout this chapter, that \(\mathbf{X}\) is pretreatment: no component of \(\mathbf{X}\) is a descendant of \(T\), so setting \(T = t\) leaves \(\mathbf{X}_i\) unchanged. If some component of \(\mathbf{X}\) were itself affected by treatment, recursive substitution would replace it by its own counterfactual value \(\mathbf{X}_i(t)\) in Equation 4.6. This is precisely what graph surgery does: the mutilated graph \(\Gcal_{\overline{T}}\) replaces the equation for \(T\) with the constant \(t\), leaving the equation for \(Y\) unchanged. Equation Equation 4.6 makes explicit that the potential outcome is a unit-level quantity determined by \(g\), \(\mathbf{X}_i\), and the unit’s own error \(U_{Y,i}\) — never by the treatments of other units, which is the no-interference component of SUTVA.
ATE and ATT as structural parameters. Substituting Equation 4.6 into the definitions gives: \[\tau_{\mathrm{ATE}} = \E\!\left[g(1, \mathbf{X}, U_Y) - g(0, \mathbf{X}, U_Y)\right], \tag{4.7}\] \[\tau_{\mathrm{ATT}} = \E\!\left[g(1, \mathbf{X}, U_Y) - g(0, \mathbf{X}, U_Y) \mid T{=}1\right]. \tag{4.8}\] Both quantities are averages of the unit-level causal effect \(g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) over different reference populations.
The linear SEM as a special case. In the Gaussian linear SEM of Chapter 1, \[Y \;=\; \beta T + \boldsymbol{\gamma}^{\top}\mathbf{X} + \varepsilon, \tag{4.9}\] the structural equation is additive and separable in \(T\), so \(Y_i(t) = \beta t + \boldsymbol{\gamma}^{\top}\mathbf{X}_i + \varepsilon_i\) and the unit-level effect is \(Y_i(1) - Y_i(0) = \beta\) for every unit. Consequently, \[\tau_{\mathrm{ATE}} \;=\; \tau_{\mathrm{ATT}} \;=\; \beta.\] The structural coefficient \(\beta\) is the average treatment effect: no averaging over heterogeneity is needed because there is none. This homogeneity is a special property of the linear additive model, not a general feature.
Heterogeneous effects. In the nonparametric SEM Equation 4.5, the unit-level effect \(g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) varies across units. The ATE averages this over the full population; the ATT averages it over the treated subpopulation. The two differ whenever treatment selection correlates with the individual effect size — i.e., whenever units who tend to benefit more also tend to self-select into treatment.
4.4 Ignorability, Positivity, and Adjustment
4.4.1 Exchangeability: Mean, Weak, and Strong Forms
The central identifying assumption in observational studies is that treatment assignment is as good as random after conditioning on observed covariates \(X\).
Definition 4.1 (Strong Ignorability (Rosenbaum and Rubin 1983)) The treatment assignment \(T\) is strongly ignorable given \(X\) if:
- Unconfoundedness: \(\bigl(Y(0),\, Y(1)\bigr) \;\indep\; T \mid X\).
- Overlap (positivity): \(0 < P(T{=}1 \mid X) < 1\), \(P_X\)-almost surely (see the positivity taxonomy remark in Section 4.4.4 for why the almost-sure formulation is the natural one).
Definition 4.1 states the joint, or strong, form of the exchangeability condition. For identifying the ATE, strictly weaker conditions suffice, and the hierarchy is worth recording explicitly:
- Mean exchangeability: \(\E[Y(t) \mid T, X] = \E[Y(t) \mid X]\). Together with consistency, positivity, and the relevant moment condition, this suffices to identify the mean \(\E[Y(t)]\).
- Weak (single-world) exchangeability: \(Y(t) \indep T \mid X\) for each \(t\) separately. Together with consistency and positivity, this identifies the entire marginal distribution of each \(Y(t)\).
- Strong (joint) exchangeability: \(\bigl(Y(0), Y(1)\bigr) \indep T \mid X\). This constrains the joint counterfactual pair and is the condition stated in Definition 4.1.
Each condition in the list implies those above it, and none of the implications reverses in general. The adjustment argument below uses only weak exchangeability, one treatment level at a time; the joint form becomes relevant for functionals of the joint law of \((Y(0), Y(1))\), such as the variance of the unit-level effect. Even the joint form, however, is not generally sufficient to identify such functionals: adjustment identifies the two marginal laws, but the observed data do not determine the cross-world dependence between \(Y(0)\) and \(Y(1)\); the covariance term in \(\mathrm{Var}\{Y(1) - Y(0)\}\) requires additional cross-world restrictions, such as rank invariance, for point identification; in the absence of such restrictions one may instead derive partial-identification bounds for the joint counterfactual functional (Appendix B).
4.4.2 Adjustment Formula under Ignorability
Proposition 4.2 (Adjustment under Conditional Exchangeability) Fix \(t \in \{0,1\}\) and assume:
- Consistency: \(T = t\) implies \(Y = Y(t)\);
- Weak conditional exchangeability: \(Y(t) \indep T \mid X\);
- Positivity: \(P(T{=}t \mid X) > 0\), \(P_X\)-almost surely;
- Integrability: \(\E|Y(t)| < \infty\).
Then \(\E[Y(t)] = \E_X\!\left[\E(Y \mid T{=}t,\, X)\right]\). If the assumptions hold for both \(t = 0\) and \(t = 1\), the ATE is identified by the standardization formula \[\tau_{\mathrm{ATE}} \;=\; \E_X\!\left[\E[Y \mid T{=}1, X] - \E[Y \mid T{=}0, X]\right]. \tag{4.10}\]
Strong ignorability (Definition 4.1) implies assumptions 2–3 for both treatment levels, so it is sufficient but not necessary; note also that consistency, which the argument uses explicitly, is not part of Definition 4.1 and must be assumed alongside it. Equation Equation 4.10 is the back-door adjustment formula of Chapter 3 written in potential-outcome notation.
Proof. For the fixed treatment level \(t\), \[\begin{aligned} \E[Y(t)] &= \E_X\!\bigl[\E\{Y(t) \mid X\}\bigr] && \text{(iterated expectations)} \\ &= \E_X\!\bigl[\E\{Y(t) \mid T{=}t,\, X\}\bigr] && \text{(weak exchangeability)} \\ &= \E_X\!\bigl[\E\{Y \mid T{=}t,\, X\}\bigr] && \text{(consistency)}. \end{aligned}\] Positivity is not an algebraic step but a support condition: it guarantees that the conditional mean \(\E(Y \mid T{=}t, X)\) is well-defined for \(P_X\)-almost every covariate value, so the outer expectation is meaningful; integrability licenses the iterated expectations. Subtracting the \(t = 0\) display from the \(t = 1\) display yields Equation 4.10. \(\square\)
4.4.3 Back-Door Interpretation
For a fixed treatment level \(t\), the weak conditional-exchangeability condition \[Y(t) \;\indep\; T \mid X\] states that, after conditioning on \(X\), treatment assignment carries no residual information about the outcome that would be observed under the intervention \(T{=}t\). The joint condition \(\bigl(Y(0), Y(1)\bigr) \indep T \mid X\) of Definition 4.1 is a stronger cross-world package; the adjustment formula requires only the single-world condition for each treatment level separately. In the language of DAGs, the closely related idea is that \(X\) blocks all back-door paths from \(T\) to \(Y\).
The graphical criterion is especially useful because it makes the source of ignorability visible. If \(X\) satisfies the back-door criterion relative to \((T, Y)\), then, under the structural causal model semantics adopted in these notes, treatment assignment is conditionally exchangeable given \(X\), which justifies standardization, regression adjustment, and related methods. In that case the observed conditional distribution within levels of \(X\) can be used to recover the interventional distribution, leading to the adjustment formula Equation 4.10 developed above.
Proposition 4.3 (Back-Door Criterion Implies Ignorability) Under the NPSEM/SWIG semantics adopted in these notes, if \(X\) satisfies the back-door criterion for the effect of \(T\) on \(Y\) — that is,
- no node in \(X\) is a descendant of \(T\), and
- \(X\) blocks every back-door path from \(T\) to \(Y\) —
then weak (single-world) unconfoundedness \(Y(t) \indep T \mid X\) holds for all \(t\).
Proof sketch. The target statement \(Y(t) \indep T \mid X\) is a single-world counterfactual independence: it is not an ordinary d-separation statement in \(\Gcal\) itself, but it is represented as a d-separation statement in the single-world intervention graph \(\Gcal(t)\) (Appendix B). The rigorous translation can be carried out directly in the SWIG \(\Gcal(t)\). A stronger sufficient construction is the NPSEM-IE representation \[Y(t) \;=\; g\!\bigl(t,\, \Pa(Y)\!\setminus\!\{T\},\, U_Y\bigr), \qquad T \;=\; f_T\!\bigl(\Pa(T),\, U_T\bigr),\] with mutually independent exogenous errors (writing, for simplicity, the case in which the non-treatment parents of \(Y\) are non-descendants of \(T\); in general \(Y(t)\) is defined by recursive substitution, with any mediating parents replaced by their own counterfactual values). Under this semantics, \(Y(t)\) and \(T\) depend on disjoint exogenous errors together with their own ancestors, and the back-door criterion implies the d-separation condition on \(\Gcal(t)\) that makes these two sets of inputs conditionally independent given \(X\). Functional independence of inputs then transfers to the outputs, yielding \(Y(t) \indep T \mid X\) for every \(t\). See Appendix B for the explicit SWIG derivation, and Wang et al. (2026) for a discussion of the extra cross-world constraints that NPSEM-IE imposes beyond what is strictly required to identify the ATE. \(\square\)
This is one of the central translations in causal inference: the potential-outcomes notation states the assumption in terms of counterfactual independence, while the DAG states it in terms of blocked paths. Together, consistency, weak exchangeability, and positivity identify the ATE via the adjustment formula Equation 4.10, and the back-door criterion provides the graphical justification for why that formula recovers the causal effect. Appendix B provides a formal graphical representation of this connection through single-world intervention graphs (SWIGs).
4.4.4 Overlap and Positivity
4.5 Where the Frameworks Agree and Diverge
4.5.1 A Systematic Comparison
| Task | Potential Outcomes | Do-Calculus / DAG | SEM |
|---|---|---|---|
| Define causal estimands (ATE, ATT) | \(\checkmark\) Especially natural notation | ATE via \(\E[Y \mid \doop(t)]\); ATT requires counterfactual augmentation (Appendix B) | Via structural equations |
| Encode causal assumptions | Formal counterfactual restrictions; a causal graph may be added but is not implicit in the notation | \(\checkmark\) Explicit directed edges | \(\checkmark\) Structural equations |
| Read conditional independence | Requires auxiliary graph or model to read off | \(\checkmark\) d-separation | Read from the induced graph under dependence restrictions on the exogenous inputs |
| Identification from observational data | Via assignment and counterfactual assumptions plus probability algebra | \(\checkmark\) Back-door, front-door, do-calculus | Via functional, exclusion, rank, and background-input restrictions |
| Likelihood construction | Typically paired with estimating equations or semiparametric models | Identifies observed-data functionals; a likelihood requires an additional statistical model | Induced once the structural functions and error laws are sufficiently specified |
| Cross-world assumptions (e.g. monotonicity) | \(\checkmark\) Natural to state | Not expressible in single-world graphs; see Appendix B | Joint counterfactuals are defined, but restrictions such as monotonicity must be imposed separately |
| Standard in statistics/epidemiology | \(\checkmark\) Dominant | Growing rapidly | Econometrics |
4.6 Summary
The potential outcome \(Y(t)\) is the outcome that would be observed under the intervention \(T = t\). Under the structural causal semantics adopted in these notes (with no interference), \(Y(t)\) has the same distribution as the outcome under the intervention \(\doop(T{=}t)\); the separate consistency assumption links counterfactuals to data through \(Y_i = Y_i(T_i)\).
The ATE and ATT are averages of the unit-level causal effect \(Y_i(1) - Y_i(0) = g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) over the full population and the treated subpopulation respectively. In the linear SEM, both equal the structural coefficient \(\beta\). They diverge under heterogeneous effects when treatment selection correlates with individual effect size.
Consistency, weak conditional exchangeability \(Y(t) \indep T \mid X\), and positivity identify the ATE via Equation 4.10 (Proposition 4.2); strong ignorability (Definition 4.1) is a sufficient package. The back-door criterion provides a sufficient graphical condition for the conditional exchangeability assumption used in adjustment formulas.
Single-world intervention graphs (SWIGs) (Richardson and Robins 2014) provide a formal graphical representation of counterfactual variables, making the single-world independence \(Y(t) \indep T \mid X\), for one fixed \(t\), a d-separation statement in a single diagram that hosts both the potential outcome and the natural treatment. The joint strong-ignorability condition involves counterfactuals from two different worlds and is not represented by any single SWIG. This material is developed in Appendix B and is not required for a first-pass understanding of adjustment.
The frameworks are complementary: potential outcomes define estimands; the do-calculus identifies them. The back-door criterion connects the two by providing the graphical condition under which adjustment recovers the causal effect.
From Ignorability to Randomization
In observational studies, ignorability must be justified by substantive knowledge encoded in a causal graph. Because unobserved confounding can never be ruled out empirically, this justification is often debated.
Randomized experiments provide a fundamentally different solution. When treatment is assigned randomly — with \(T\) denoting the randomized assignment, which coincides with treatment received in the simple full-compliance trial — \((Y(0),Y(1)) \indep T\) holds by design, guaranteeing ignorability without any covariate adjustment. Chapter 5 studies randomized experiments as the canonical design where causal effects are identifiable directly from the data-generating mechanism.
Causal inference \(=\) counterfactual questions \(+\) graphical assumptions \(+\) statistical estimation.
Identification vs. Estimation
The first four chapters have focused on identification: whether a causal quantity \(\tau = \E[Y(1) - Y(0)]\) can be written as \(\tau = \Phi(P(Y,T,X))\) for some functional \(\Phi\) of the observed distribution. Basic estimators — regression adjustment, standardization, stratification, and inverse probability weighting — are introduced in Chapters 5–6; their systematic asymptotic, influence-function, and semiparametric analysis, together with doubly robust estimators and instrumental variables, is developed in Part III.
4.7 Problems
1. SUTVA and consistency. Suppose \(n = 3\) units receive binary treatments \((T_1, T_2, T_3)\) and unit \(i\)’s outcome may depend on all three treatments: \(Y_i = Y_i(T_1, T_2, T_3)\).
- How many potential outcomes does unit 1 have? Write them out explicitly.
- SUTVA imposes no interference: \(Y_i(T_1,T_2,T_3) = Y_i(T_i)\). How many distinct potential outcomes remain?
- Give a real-world example where no-interference plausibly holds and one where it plausibly fails. In the latter case, redefine the potential outcome using one of: the complete assignment vector, an exposure mapping, partial interference, or cluster-level treatment, and explain what additional design or structural assumption would then be needed.
2. ATE, ATT, and selection bias. Let \((Y(0), Y(1), T) \sim P\) with \(P(T{=}1) = 0.5\). Suppose \(\E[Y(1)] = 3\), \(\E[Y(0)] = 1\), \(\E[Y(1) \mid T{=}1] = 4\), \(\E[Y(0) \mid T{=}1] = 2\), \(\E[Y(1) \mid T{=}0] = 2\), \(\E[Y(0) \mid T{=}0] = 0\).
- Compute the ATE and ATT. Are they equal?
- Compute the naive contrast \(D_{\mathrm{obs}} = \E[Y \mid T{=}1] - \E[Y \mid T{=}0]\) (the population quantity targeted by the naive estimator). Decompose the gap between this and the ATE into a selection bias term and an ATT–ATE difference.
- Give a sufficient graphical condition on the DAG under which the naive contrast \(D_{\mathrm{obs}}\) equals the ATE: the empty set satisfies the back-door criterion, i.e. there is no open back-door path from \(T\) to \(Y\). State the corresponding condition in potential-outcome language (marginal ignorability), and explain why the mere existence of some nonempty adjustment set is not enough.
3. Causal estimands from the structural equation. Consider the nonparametric SEM \(Y = g(T, X, U_Y)\) with binary treatment \(T \in \{0,1\}\), observed covariate \(X\), and unobserved error \(U_Y\) with \(\E[U_Y] = 0\); any further distributional assumptions are stated in the individual parts.
- Write \(\tau_{\mathrm{ATE}}\) and \(\tau_{\mathrm{ATT}}\) as expectations of \(g(1, X, U_Y) - g(0, X, U_Y)\) over the appropriate distribution. Under what condition do they coincide?
- Specialize to the linear SEM \(g(t, x, u) = \alpha + \beta t + \gamma x + u\). Show that the unit-level causal effect \(Y_i(1) - Y_i(0)\) is constant across all units, and hence \(\tau_{\mathrm{ATE}} = \tau_{\mathrm{ATT}} = \beta\).
- Now consider the heterogeneous SEM \(g(t, x, u) = (\alpha + u)\,t + \gamma x\), where \(\E[U_Y] = 0\), \(\mathrm{Var}(U_Y) = \sigma^2 < \infty\), \(\mathrm{Cov}(U_Y, T) = \rho\), \(P(T{=}1) = p \in (0,1)\), and \(X\) is independent of \((T, U_Y)\); assume a joint law compatible with these moments (the Cauchy–Schwarz inequality gives the necessary condition \(|\rho| \le \sigma\sqrt{p(1-p)}\)). Compute \(\tau_{\mathrm{ATE}}\) and \(\tau_{\mathrm{ATT}}\). Show that they differ when \(\rho \neq 0\), and interpret this difference in words.
- In part (c), show that the population OLS slope on \(T\) in the regression of \(Y\) on \((1, T, X)\) equals \(\alpha + \rho/p = \tau_{\mathrm{ATT}}\), not \(\tau_{\mathrm{ATE}} = \alpha\). Explain why: the stratum-specific difference \(\E[Y \mid T{=}1, X] - \E[Y \mid T{=}0, X]\) equals \(\alpha + \rho/p\) at every \(X\) (the heterogeneity carried by \(U_Y\) is not observable through \(X\)), so adjusted OLS recovers the ATT, shifted away from \(\tau_{\mathrm{ATE}} = \alpha\) by the selection term \(\E[U_Y \mid T{=}1] = \rho/p\). What additional structure would identify \(\tau_{\mathrm{ATE}}\)? (Hint: randomization of \(T\) would imply \(\E[U_Y \mid T] = 0\). Alternatively, if a richer set of pretreatment covariates \(W\) rendered \(Y(t) \indep T \mid W\), then \(\tau_{\mathrm{ATE}}\) would be identified by standardization over \(W\), even though \(\E[U_Y \mid T]\) need not vanish. A valid instrument does not generally identify the ATE under treatment-effect heterogeneity; in the binary-instrument, binary-treatment setting, relevance, exogeneity, exclusion, and monotonicity together identify the LATE for compliers — see Chapter 7.)
4. SWIGs and ignorability. (Requires Appendix B.) Consider the DAG: \(X \to T\), \(X \to Y\), \(T \to Y\), with \(X\) fully observed.
- Construct the SWIG \(\Gcal(t)\) by splitting \(T\) into its random and fixed halves. Draw the result, labeling the random half, fixed half, and the potential outcome \(Y(t)\).
- In \(\Gcal(t)\), identify all paths between the random half \(T\) and \(Y(t)\). Apply the SWIG d-separation rules to determine which paths are blocked and which are open before conditioning.
- First verify, in the original DAG, that \(X\) satisfies the back-door criterion for \((T, Y)\). Then use d-separation in \(\Gcal(t)\) to verify that \((Y(t) \indep T \mid X)_{\Gcal(t)}\) holds, illustrating the direction of Proposition 4.3.
- Now add a hidden common cause \(U \to T\), \(U \to Y\). Draw the revised SWIG. Does the ignorability argument still hold? State the correct conclusion and identify what additional structure (if any) would be needed for identification.
5. Positivity and target populations. Let \(X \in \{0,1\}\) with \(P(X{=}1) > 0\), and suppose \(P(T{=}1 \mid X{=}1) = 1\).
- Is the ATE identified by nonparametric adjustment over the original target population? Explain.
- State conditions under which the ATT may remain identified.
- Contrast this population positivity failure with a finite sample that happens to contain no control units with \(X = 1\) even though \(P(T{=}0 \mid X{=}1) > 0\) in the population.
- Explain what additional assumption is being relied upon if a parametric outcome model is extrapolated into the unsupported stratum.