4  Potential Outcomes and Adjustment

NoteLearning Objectives

By the end of this chapter, students should be able to:

  1. Define the potential outcome \(Y(t)\), state SUTVA and consistency precisely, and explain how the consistency equation \(Y = Y(T)\) links observed and potential outcomes.
  2. Define the ATE and ATT, explain how they relate to each other, and re-express them as functionals of the structural equation for \(Y\).
  3. Show that the linear SEM structural coefficient \(\beta\) equals the ATE under homogeneous effects, and explain how the ATE and ATT can diverge under treatment-effect heterogeneity when treatment selection is associated with the individual treatment effect.
  4. State the ignorability assumption, distinguish its weak (single-world) and strong (joint) forms, interpret it graphically as a back-door condition, and explain why positivity is required for nonparametric adjustment over the target population.
  5. Derive the adjustment formula for the ATE under consistency, weak conditional exchangeability, and positivity, and explain how the back-door criterion provides the graphical justification for conditional exchangeability.

4.1 Motivation: A Third Language for Causality

The first three chapters developed causal inference primarily in the languages of structural equations and directed acyclic graphs. Those two languages are especially effective for expressing interventions syntactically: the structural equation model shows how an intervention replaces an assignment mechanism, and the DAG shows how the same intervention deletes incoming arrows into the treatment node. In this chapter we introduce a third language: the potential outcomes framework of Neyman (1923) and Rubin (1974).

Potential outcomes provide a direct language for defining causal estimands. Quantities such as the average treatment effect, the average treatment effect on the treated, and related causal contrasts are most naturally written in terms of counterfactual outcomes \(Y(1)\), \(Y(0)\), and more generally \(Y(t)\). For this reason, the potential outcomes framework has become standard in statistics, biostatistics, epidemiology, and much of econometrics.

Assumptions such as ignorability and exclusion can certainly be stated in this language, but DAGs and the do-operator often make their structural content more transparent by showing which paths are blocked, which variables are confounders, and which interventions are being considered. Positivity, by contrast, is a support condition on the observed-data distribution and is not encoded by any graph. Thus the two frameworks should be viewed as complementary rather than competing: potential outcomes define the causal quantities of interest, while DAGs and the do-calculus clarify identification.

The purpose of this chapter is therefore not to replace the graphical viewpoint developed earlier, but to connect it to the potential-outcomes notation that dominates much of the applied literature. We first define potential outcomes and the standard causal estimands built from them. We then show how consistency, exchangeability, and positivity together identify the average treatment effect, and how the back-door criterion provides the graphical justification for that identification.

Foundations (Chs. 1–3) DAGs, do-calculus, identification Potential Outcomes (Ch. 4) ← here Define estimands; bridge frameworks Designs (Chs. 5–9) Randomization, PS, IV, mediation, sensitivity Estimation (Part III) Regression, IPW, doubly robust
Where this chapter sits in the course. Potential outcomes supply the language for defining estimands and bridge the graphical machinery of Part I to the designs and estimation theory that follow.

Section 4.2 defines the potential outcome framework rigorously. Section 4.3 develops the two principal causal estimands: ATE and ATT. Section 4.4 develops the identification chain from consistency, exchangeability, and positivity to the adjustment formula, and explains how the back-door criterion provides the graphical justification. Section 4.5 situates the potential outcomes framework relative to the do-calculus and the SEM. For readers who want a formal graphical representation of counterfactual variables, Appendix B introduces single-world intervention graphs (SWIGs).

4.2 The Neyman–Rubin Potential Outcomes Framework

4.2.1 The Potential Outcome

NoteDefinition: Potential Outcome (Neyman 1923; Rubin 1974)

For each unit \(i\) and each possible treatment value \(t \in \mathcal{T}\), the potential outcome \(Y_i(t)\) is the value of the outcome that unit \(i\) would have exhibited had its treatment been set to \(t\), possibly contrary to fact.

The collection \(\{Y_i(t) : t \in \mathcal{T}\}\) is called the schedule of potential outcomes for unit \(i\). In the binary case \(\mathcal{T} = \{0,1\}\), the schedule is \((Y_i(0),\, Y_i(1))\).

Interpretation. \(Y_i(1)\) is the outcome unit \(i\) would achieve under treatment; \(Y_i(0)\) is the outcome under control. Only one of these is ever observed for any given unit. The unobserved potential outcome is called the counterfactual.

WarningThe Fundamental Problem of Causal Inference

(Holland 1986) For a single treatment occasion, only \(Y_i(T_i)\) is observed; the pair \(\bigl(Y_i(0), Y_i(1)\bigr)\) is never jointly observed for the same unit, so the individual-level causal effect \(Y_i(1) - Y_i(0)\) is never directly observable. Without much stronger structural assumptions, individual treatment effects are not identified. Causal inference therefore typically targets population summaries — such as the ATE and ATT — that can be identified from cross-unit comparisons under explicit assumptions.

The structure of the problem is visible in a small table — often called the science table — listing the full potential-outcome schedule next to what is actually observed:

Unit \(Y_i(0)\) \(Y_i(1)\) \(T_i\) \(Y_i = Y_i(T_i)\)
1 observed ? 0 \(Y_1(0)\)
2 ? observed 1 \(Y_2(1)\)
3 observed ? 0 \(Y_3(0)\)

Each row is missing exactly one entry, and which entry is missing is determined by \(T_i\): this is the fundamental problem in tabular form. The final column is the consistency equation at work, and the impossibility of filling in the “?” entries within a row is why causal inference must rely on cross-unit comparisons — which is precisely what the exchangeability assumptions of Section 4.4 will license.

4.2.2 SUTVA, Treatment Versions, and Consistency

The potential outcome \(Y_i(t)\) is only well-defined if it does not depend on the treatments assigned to other units, and if there is a single, unambiguous version of each treatment level. These requirements are formalized as SUTVA.

NoteDefinition: SUTVA (Rubin 1980)

The stable unit treatment value assumption has two components:

  1. No interference. The potential outcome \(Y_i(t)\) depends only on unit \(i\)’s own treatment, not on the treatments of other units: \(Y_i(t_1, \ldots, t_n) = Y_i(t_i)\).
  2. No hidden versions. For each labeled treatment level \(t\), all outcome-relevant treatment versions are either explicitly encoded in \(t\) or are causally equivalent for the estimand under study, so that \(Y_i(t)\) refers to a unique, well-defined intervention.
NoteDefinition: Consistency

For every unit \(i\) and every treatment level \(t\): if \(T_i = t\), then \(Y_i = Y_i(t)\). Compactly, the observed outcome equals the potential outcome at the treatment actually received: \[Y_i \;=\; Y_i(T_i). \tag{4.1}\] For binary treatment this can be written \(Y_i = T_i\, Y_i(1) + (1 - T_i)\, Y_i(0)\).

Consistency is the bridge between the potential-outcome world and the observed-data world, and it is conceptually distinct from the exchangeability and positivity assumptions of Section 4.4. It is best regarded as an assumption in its own right rather than as an automatic consequence of SUTVA, although SUTVA is what makes the scalar notation \(Y_i(t)\) in Equation 4.1 adequate.

Two clarifications about SUTVA itself are worth recording. First, interference does not make potential outcomes meaningless; it makes the scalar notation inadequate. Under interference one writes \(Y_i(t_1, \ldots, t_n)\), or reduces the assignment vector through an exposure mapping, and causal analysis proceeds with the richer index set. Second, “no hidden versions” does not require the physical treatment received by any two treated units to be literally identical; it requires that any outcome-relevant versions be either encoded in the treatment variable or causally equivalent for the estimand at hand. When these conditions fail — vaccination that provides herd immunity, or a label \(T{=}1\) that covers interventions with different effects — it is the scalar consistency equation Equation 4.1 that breaks down.

4.2.3 Connection to the Do-Operator

The notation \(Y(t)\) and the expression \(P(Y \mid \doop(T{=}t))\) refer to the same causal idea viewed from two complementary angles. The potential outcome \(Y(t)\) denotes the outcome that would be realized for a unit if the treatment were set to \(t\). The do-operator, by contrast, describes an intervention at the level of the data-generating mechanism: it replaces the original treatment assignment rule by the constant value \(t\).

When \(Y(t)\) is constructed by applying the same structural intervention that defines \(\doop(T{=}t)\), equality of the resulting laws holds by construction: intervening to set \(T = t\) in the underlying structural equation model produces an outcome variable with the same distribution as \(Y(t)\). No interference makes the scalar unit-level notation \(Y_i(t)\) adequate. The separate consistency assumption links these counterfactual variables to observed outcomes when the realized treatment equals \(t\). Thus the potential-outcomes notation and the interventional distribution are two ways of encoding the same manipulated world.

Proposition 4.1 (Potential Outcomes and the Do-Operator) Consider the structural causal model of Chapters 1–3, and construct \(Y(t)\) by replacing the structural equation for \(T\) with the constant assignment \(T := t\) — the same intervention that defines \(\doop(T{=}t)\). Under no interference (SUTVA), the potential outcome has the interventional distribution: \[P\{Y(t) \le y\} \;=\; P\{Y \le y \mid \doop(T{=}t)\} \qquad \text{for every } y; \tag{4.2}\] in particular, \(\E[Y(t)] = \E[Y \mid \doop(T{=}t)]\).

Proof sketch. In the intervened SCM, whose graph is \(\Gcal_{\overline{T}}\), the structural equation for \(T\) is replaced by \(T := t\) for every unit. The structural model then determines the value of \(Y\) for unit \(i\) from \(t\) and unit \(i\)’s own background variables — which is precisely what the potential outcomes framework defines as \(Y_i(t)\). SUTVA’s no-interference condition ensures \(Y_i(t)\) does not depend on other units’ treatment values, so the marginal distribution of \(Y\) in the intervened model equals the marginal distribution of \(Y(t)\). By definition of the do-operator, the former is \(f(y \mid \doop(T{=}t))\). \(\square\)

This proposition should be interpreted carefully. Consistency alone does not create the link between \(Y(t)\) and the do-operator; the link comes from the underlying causal model that assigns meaning to the intervention \(\doop(T{=}t)\). Consistency and no interference ensure that the observed outcome agrees with the appropriate potential outcome at the realized treatment, while the structural model explains how the intervention generates the counterfactual world.

This equivalence matters because it lets us move freely between two notational traditions. When defining estimands such as \(\E[Y(1) - Y(0)]\), potential-outcome notation is often most natural. When proving identification results from a graph, the do-operator is often more convenient because it integrates directly with graph surgery and the do-calculus. In general \(f(y \mid \doop(t))\) differs from \(f(y \mid T{=}t)\) when \(T\) is endogenous (Chapter 1), although equality can occur through special cancellations. Potential-outcome notation marks the same distinction — \(\E[Y(t)]\) and \(\E[Y \mid T{=}t]\) are different expressions — but it does not by itself supply the graphical calculus for deciding when the two coincide. In a well-specified causal model these are not competing definitions, but compatible representations of the same intervention.

4.3 Causal Estimands

4.3.1 The Average Treatment Effect and Its Relatives

NoteDefinition: ATE and ATT

For a binary treatment \(T \in \{0,1\}\):

  • The average treatment effect (ATE) is \(\tau_{\mathrm{ATE}} = \E[Y(1) - Y(0)]\).
  • The average treatment effect on the treated (ATT) is \(\tau_{\mathrm{ATT}} = \E[Y(1) - Y(0) \mid T{=}1]\).

Let \(\tau = Y(1) - Y(0)\) denote the unit-level effect and let \(p = P(T{=}1) \in (0,1)\). Alongside the ATT one may define the average treatment effect on the controls, \(\tau_{\mathrm{ATC}} = \E[\tau \mid T{=}0]\). The law of total expectation gives the exact decomposition \[\tau_{\mathrm{ATE}} \;=\; p\,\tau_{\mathrm{ATT}} + (1-p)\,\tau_{\mathrm{ATC}}, \tag{4.3}\] so the ATE and ATT coincide if and only if \(\tau_{\mathrm{ATT}} = \tau_{\mathrm{ATC}}\): the mean effect among the treated must equal the mean effect among the controls. Equivalently, \(\tau_{\mathrm{ATT}} - \tau_{\mathrm{ATE}} = \mathrm{Cov}(T, \tau)/p\). Treatment effects may therefore be heterogeneous across units while the ATE and ATT still coincide; what matters is whether selection into treatment is related to the effect in the mean. Unit-level constancy of the effect is sufficient but not necessary.

None of these quantities is directly observable. The naive estimator \(\hat\tau_{\mathrm{naive}} = \bar{Y}_{T=1} - \bar{Y}_{T=0}\) estimates the observed treatment-group contrast \(D_{\mathrm{obs}} = \E[Y \mid T{=}1] - \E[Y \mid T{=}0]\), which in general differs from the ATE. Under consistency, \(D_{\mathrm{obs}} = \E[Y(1) \mid T{=}1] - \E[Y(0) \mid T{=}0]\), and adding and subtracting \(\E[Y(0) \mid T{=}1]\) yields the exact decomposition \[D_{\mathrm{obs}} - \tau_{\mathrm{ATE}} \;=\; \underbrace{\E[Y(0) \mid T{=}1] - \E[Y(0) \mid T{=}0]}_{\text{baseline-selection bias}} \;+\; \underbrace{\bigl(\tau_{\mathrm{ATT}} - \tau_{\mathrm{ATE}}\bigr)}_{\text{effect-selection difference}}. \tag{4.4}\] The first term reflects baseline differences between the two groups — units who select into treatment may have fared differently even untreated — and the second reflects selection on the size of the effect. In a randomized experiment both terms vanish. Problem 2 works through a numerical case.

NoteRemark: When ATE and ATT Coincide

ATE and ATT coincide when \(\E[Y(t) \mid T] = \E[Y(t)]\) for \(t \in \{0,1\}\), i.e. when treatment assignment is mean-independent of potential outcomes. In a randomized experiment this holds by design, and both terms of the decomposition Equation 4.4 vanish.

4.3.2 Causal Estimands as Functionals of the Structural Model

The potential outcome \(Y_i(t)\) was introduced as a hypothetical: the value unit \(i\) would have exhibited under treatment \(t\). The SEM framework of Chapter 1 gives this hypothetical a precise generative meaning, and thereby expresses the potential-outcome estimands as functionals of the structural equation for \(Y\).

From structural equation to potential outcome. In the SEM, the outcome is determined by a structural equation \[Y \;=\; g(T,\, \mathbf{X},\, U_Y), \tag{4.5}\] where \(U_Y\) collects all sources of variation in \(Y\) not already accounted for by \((T, \mathbf{X})\). The potential outcome under \(\doop(T{=}t)\) is obtained by substituting \(t\) for \(T\) while holding everything else fixed: \[Y_i(t) \;=\; g(t,\, \mathbf{X}_i,\, U_{Y,i}). \tag{4.6}\] Here we assume, as throughout this chapter, that \(\mathbf{X}\) is pretreatment: no component of \(\mathbf{X}\) is a descendant of \(T\), so setting \(T = t\) leaves \(\mathbf{X}_i\) unchanged. If some component of \(\mathbf{X}\) were itself affected by treatment, recursive substitution would replace it by its own counterfactual value \(\mathbf{X}_i(t)\) in Equation 4.6. This is precisely what graph surgery does: the mutilated graph \(\Gcal_{\overline{T}}\) replaces the equation for \(T\) with the constant \(t\), leaving the equation for \(Y\) unchanged. Equation Equation 4.6 makes explicit that the potential outcome is a unit-level quantity determined by \(g\), \(\mathbf{X}_i\), and the unit’s own error \(U_{Y,i}\) — never by the treatments of other units, which is the no-interference component of SUTVA.

ATE and ATT as structural parameters. Substituting Equation 4.6 into the definitions gives: \[\tau_{\mathrm{ATE}} = \E\!\left[g(1, \mathbf{X}, U_Y) - g(0, \mathbf{X}, U_Y)\right], \tag{4.7}\] \[\tau_{\mathrm{ATT}} = \E\!\left[g(1, \mathbf{X}, U_Y) - g(0, \mathbf{X}, U_Y) \mid T{=}1\right]. \tag{4.8}\] Both quantities are averages of the unit-level causal effect \(g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) over different reference populations.

The linear SEM as a special case. In the Gaussian linear SEM of Chapter 1, \[Y \;=\; \beta T + \boldsymbol{\gamma}^{\top}\mathbf{X} + \varepsilon, \tag{4.9}\] the structural equation is additive and separable in \(T\), so \(Y_i(t) = \beta t + \boldsymbol{\gamma}^{\top}\mathbf{X}_i + \varepsilon_i\) and the unit-level effect is \(Y_i(1) - Y_i(0) = \beta\) for every unit. Consequently, \[\tau_{\mathrm{ATE}} \;=\; \tau_{\mathrm{ATT}} \;=\; \beta.\] The structural coefficient \(\beta\) is the average treatment effect: no averaging over heterogeneity is needed because there is none. This homogeneity is a special property of the linear additive model, not a general feature.

Heterogeneous effects. In the nonparametric SEM Equation 4.5, the unit-level effect \(g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) varies across units. The ATE averages this over the full population; the ATT averages it over the treated subpopulation. The two differ whenever treatment selection correlates with the individual effect size — i.e., whenever units who tend to benefit more also tend to self-select into treatment.

NoteRemark: Equivalent in Content, Different in Emphasis

The SEM representation Equation 4.6 clarifies why the potential outcomes framework and the do-calculus are equivalent in content but different in emphasis. The do-calculus works with the interventional distribution \(f(y \mid \doop(T{=}t))\) as a population-level object and asks when it can be recovered from observational data. The potential outcomes framework works with unit-level quantities \(Y_i(t)\) and asks what population summaries (ATE, ATT) are scientifically meaningful. The SEM ties the two together: \(Y_i(t) = g(t, \mathbf{X}_i, U_{Y,i})\) is the unit-level object, and integrating over the distribution of \((\mathbf{X}, U_Y)\) recovers the interventional distribution.

4.4 Ignorability, Positivity, and Adjustment

4.4.1 Exchangeability: Mean, Weak, and Strong Forms

The central identifying assumption in observational studies is that treatment assignment is as good as random after conditioning on observed covariates \(X\).

Definition 4.1 (Strong Ignorability (Rosenbaum and Rubin 1983)) The treatment assignment \(T\) is strongly ignorable given \(X\) if:

  1. Unconfoundedness: \(\bigl(Y(0),\, Y(1)\bigr) \;\indep\; T \mid X\).
  2. Overlap (positivity): \(0 < P(T{=}1 \mid X) < 1\), \(P_X\)-almost surely (see the positivity taxonomy remark in Section 4.4.4 for why the almost-sure formulation is the natural one).

Definition 4.1 states the joint, or strong, form of the exchangeability condition. For identifying the ATE, strictly weaker conditions suffice, and the hierarchy is worth recording explicitly:

  • Mean exchangeability: \(\E[Y(t) \mid T, X] = \E[Y(t) \mid X]\). Together with consistency, positivity, and the relevant moment condition, this suffices to identify the mean \(\E[Y(t)]\).
  • Weak (single-world) exchangeability: \(Y(t) \indep T \mid X\) for each \(t\) separately. Together with consistency and positivity, this identifies the entire marginal distribution of each \(Y(t)\).
  • Strong (joint) exchangeability: \(\bigl(Y(0), Y(1)\bigr) \indep T \mid X\). This constrains the joint counterfactual pair and is the condition stated in Definition 4.1.

Each condition in the list implies those above it, and none of the implications reverses in general. The adjustment argument below uses only weak exchangeability, one treatment level at a time; the joint form becomes relevant for functionals of the joint law of \((Y(0), Y(1))\), such as the variance of the unit-level effect. Even the joint form, however, is not generally sufficient to identify such functionals: adjustment identifies the two marginal laws, but the observed data do not determine the cross-world dependence between \(Y(0)\) and \(Y(1)\); the covariance term in \(\mathrm{Var}\{Y(1) - Y(0)\}\) requires additional cross-world restrictions, such as rank invariance, for point identification; in the absence of such restrictions one may instead derive partial-identification bounds for the joint counterfactual functional (Appendix B).

4.4.2 Adjustment Formula under Ignorability

Proposition 4.2 (Adjustment under Conditional Exchangeability) Fix \(t \in \{0,1\}\) and assume:

  1. Consistency: \(T = t\) implies \(Y = Y(t)\);
  2. Weak conditional exchangeability: \(Y(t) \indep T \mid X\);
  3. Positivity: \(P(T{=}t \mid X) > 0\), \(P_X\)-almost surely;
  4. Integrability: \(\E|Y(t)| < \infty\).

Then \(\E[Y(t)] = \E_X\!\left[\E(Y \mid T{=}t,\, X)\right]\). If the assumptions hold for both \(t = 0\) and \(t = 1\), the ATE is identified by the standardization formula \[\tau_{\mathrm{ATE}} \;=\; \E_X\!\left[\E[Y \mid T{=}1, X] - \E[Y \mid T{=}0, X]\right]. \tag{4.10}\]

Strong ignorability (Definition 4.1) implies assumptions 2–3 for both treatment levels, so it is sufficient but not necessary; note also that consistency, which the argument uses explicitly, is not part of Definition 4.1 and must be assumed alongside it. Equation Equation 4.10 is the back-door adjustment formula of Chapter 3 written in potential-outcome notation.

Proof. For the fixed treatment level \(t\), \[\begin{aligned} \E[Y(t)] &= \E_X\!\bigl[\E\{Y(t) \mid X\}\bigr] && \text{(iterated expectations)} \\ &= \E_X\!\bigl[\E\{Y(t) \mid T{=}t,\, X\}\bigr] && \text{(weak exchangeability)} \\ &= \E_X\!\bigl[\E\{Y \mid T{=}t,\, X\}\bigr] && \text{(consistency)}. \end{aligned}\] Positivity is not an algebraic step but a support condition: it guarantees that the conditional mean \(\E(Y \mid T{=}t, X)\) is well-defined for \(P_X\)-almost every covariate value, so the outer expectation is meaningful; integrability licenses the iterated expectations. Subtracting the \(t = 0\) display from the \(t = 1\) display yields Equation 4.10. \(\square\)

NoteExample: A Two-Stratum Adjustment Calculation

Suppose \(X \in \{0,1\}\) with \(P(X{=}1) = 0.4\) and \(P(X{=}0) = 0.6\), and the observed conditional means are \[\E[Y \mid T{=}1, X{=}1] = 8, \quad \E[Y \mid T{=}0, X{=}1] = 6,\] \[\E[Y \mid T{=}1, X{=}0] = 5, \quad \E[Y \mid T{=}0, X{=}0] = 4.\] Under consistency, weak exchangeability, and positivity, Equation 4.10 gives \[\begin{aligned} \E[Y(1)] &= P(X{=}1)\cdot\E[Y\mid T{=}1,X{=}1] + P(X{=}0)\cdot\E[Y\mid T{=}1,X{=}0] \\ &= 0.4\times 8 + 0.6\times 5 \;=\; 6.2, \\[4pt] \E[Y(0)] &= P(X{=}1)\cdot\E[Y\mid T{=}0,X{=}1] + P(X{=}0)\cdot\E[Y\mid T{=}0,X{=}0] \\ &= 0.4\times 6 + 0.6\times 4 \;=\; 4.8. \end{aligned}\] Hence \(\tau_{\mathrm{ATE}} = \E[Y(1) - Y(0)] = 6.2 - 4.8 = 1.4\). The adjustment formula works by comparing treated and control outcomes within each covariate stratum and then averaging those within-stratum comparisons over the marginal distribution of \(X\).

4.4.3 Back-Door Interpretation

For a fixed treatment level \(t\), the weak conditional-exchangeability condition \[Y(t) \;\indep\; T \mid X\] states that, after conditioning on \(X\), treatment assignment carries no residual information about the outcome that would be observed under the intervention \(T{=}t\). The joint condition \(\bigl(Y(0), Y(1)\bigr) \indep T \mid X\) of Definition 4.1 is a stronger cross-world package; the adjustment formula requires only the single-world condition for each treatment level separately. In the language of DAGs, the closely related idea is that \(X\) blocks all back-door paths from \(T\) to \(Y\).

The graphical criterion is especially useful because it makes the source of ignorability visible. If \(X\) satisfies the back-door criterion relative to \((T, Y)\), then, under the structural causal model semantics adopted in these notes, treatment assignment is conditionally exchangeable given \(X\), which justifies standardization, regression adjustment, and related methods. In that case the observed conditional distribution within levels of \(X\) can be used to recover the interventional distribution, leading to the adjustment formula Equation 4.10 developed above.

Proposition 4.3 (Back-Door Criterion Implies Ignorability) Under the NPSEM/SWIG semantics adopted in these notes, if \(X\) satisfies the back-door criterion for the effect of \(T\) on \(Y\) — that is,

  1. no node in \(X\) is a descendant of \(T\), and
  2. \(X\) blocks every back-door path from \(T\) to \(Y\)

then weak (single-world) unconfoundedness \(Y(t) \indep T \mid X\) holds for all \(t\).

Proof sketch. The target statement \(Y(t) \indep T \mid X\) is a single-world counterfactual independence: it is not an ordinary d-separation statement in \(\Gcal\) itself, but it is represented as a d-separation statement in the single-world intervention graph \(\Gcal(t)\) (Appendix B). The rigorous translation can be carried out directly in the SWIG \(\Gcal(t)\). A stronger sufficient construction is the NPSEM-IE representation \[Y(t) \;=\; g\!\bigl(t,\, \Pa(Y)\!\setminus\!\{T\},\, U_Y\bigr), \qquad T \;=\; f_T\!\bigl(\Pa(T),\, U_T\bigr),\] with mutually independent exogenous errors (writing, for simplicity, the case in which the non-treatment parents of \(Y\) are non-descendants of \(T\); in general \(Y(t)\) is defined by recursive substitution, with any mediating parents replaced by their own counterfactual values). Under this semantics, \(Y(t)\) and \(T\) depend on disjoint exogenous errors together with their own ancestors, and the back-door criterion implies the d-separation condition on \(\Gcal(t)\) that makes these two sets of inputs conditionally independent given \(X\). Functional independence of inputs then transfers to the outputs, yielding \(Y(t) \indep T \mid X\) for every \(t\). See Appendix B for the explicit SWIG derivation, and Wang et al. (2026) for a discussion of the extra cross-world constraints that NPSEM-IE imposes beyond what is strictly required to identify the ATE. \(\square\)

NoteRemark: Which Direction Matters in Practice

This is the direction most important for practice: a graphical adjustment set justifies the counterfactual independence needed for standardization, regression adjustment, and propensity-score methods. The converse — whether every ignorability statement corresponds to a back-door condition — is more delicate and depends on the precise graphical-counterfactual representation used; see Appendix B.

NoteRemark: Independent Errors — Mild or Strong? (Advanced/Optional)

The additional independent-error assumption used in the NPSEM-IE construction — mentioned as one sufficient route in Proposition 4.3, whose conclusion can also be established directly in the SWIG — deserves careful emphasis, because its strength is easy to underestimate.

Consider the causally sufficient triangle \(X \to T\), \(X \to Y\), \(T \to Y\) written as an NPSEM, \[X = U_X, \qquad T = f_T(X, U_T), \qquad Y = f_Y(X, T, U_Y),\] with \(U_X \indep U_T \indep U_Y\) mutually independent. Two lines suffice to show that error independence implies the unconfoundedness half of strong ignorability (Definition 4.1). Given \(X\), the treatment \(T = f_T(X, U_T)\) is a function of \(U_T\) alone, while the entire collection \(\{Y(t) = f_Y(X, t, U_Y) : t \in \{0,1\}\}\) is a function of \(U_Y\) alone; since \(U_T \indep U_Y\) conditional on \(X\) (itself a function of \(U_X\)), independence of the inputs transfers to the outputs: \[\bigl(Y(0),\, Y(1)\bigr) \;\indep\; T \mid X.\] In this causally sufficient model — where \(X\) contains every common cause of \(T\) and \(Y\) — an analyst who writes down a system of regression equations with mutually independent errors and reads it structurally has therefore already assumed conditional ignorability given \(X\): the assumption commonly regarded as heroic follows from the one commonly regarded as routine (Wang et al. 2026). Independent equation-specific errors alone do not remove confounding by a shared background cause omitted from \(X\): the model of Chapter 1, with \(T = f_T(U, \delta)\), \(Y = f_Y(T, U, \varepsilon)\), and \((U, \delta, \varepsilon)\) mutually independent, remains confounded by the shared \(U\).

The implication is strict, because error independence says far more. A single error term \(U_Y\) generates the entire array of counterfactuals \(\{Y(x,t) : \text{all } x, t\}\) simultaneously, and \(U_T\) generates the entire vector \(\{T(x) : \text{all } x\}\), where \(T(x) = f_T(x, U_T)\) denotes the potential treatment under \(X\) set to \(x\). Mutual independence of the errors therefore implies the cross-world statement \[X \;\indep\; \bigl(T(x),\ \forall x\bigr) \;\indep\; \bigl(Y(x,t),\ \forall x, t\bigr)\] (and under the canonical representation, in which the errors are the counterfactual arrays themselves, the two statements are equivalent). This ties together counterfactuals from mutually exclusive interventions — for example, \(Y(x{=}0, t{=}1)\) and \(Y(x{=}1, t{=}0)\) jointly — yet no unit ever inhabits two worlds, so no experiment, not even one randomizing both \(X\) and \(T\), can test these joint constraints. Weak, single-world ignorability, by contrast, restricts one world at a time, and it can be guaranteed by design: randomizing \(T\) within levels of \(X\) enforces it. Even then it is not empirically verified from the observed outcomes of a single experiment; its credibility comes from the known assignment mechanism. Indeed, there exist counterfactual models — the FFRCISTGs of Robins (1986); see Richardson and Robins (2014) — in which every single-world independence holds, so that ignorability is satisfied for each \(t\), while cross-world dependence remains; such models are observationally and experimentally indistinguishable from an NPSEM-IE.

Why, then, does error independence feel mild? Because in ordinary regression practice it is a statement about the observational joint distribution only, where it is close to a bookkeeping convention. What converts it into a strong assumption is autonomy: the structural reading insists that the same equations and the same error distributions persist across interventional worlds, so that \(U_Y\) is no longer “whatever residual makes the regression fit” but the unit’s latent response type, held fixed across all hypothetical treatments. The extra cross-world strength is not vacuous: it supplies one ingredient in the standard identification of the natural direct and indirect effects — quantities that the weaker FFRCISTG model leaves unidentified. The mediation formula additionally requires consistency, appropriate positivity, and the relevant treatment–mediator and mediator–outcome exchangeability conditions; these are developed in Chapter 8.

NoteExample: Labor Training Program

Following LaLonde (1986) and Dehejia and Wahba (1999), let \(T\) = receipt of job training, \(Y\) = earnings two years later, \(X\) = (age, education, prior earnings, race, marital status), all measured before training (pretreatment).

X T Y U

The hidden confounder \(U\) (dashed) is drawn to represent the threat discussed below; the solid graph is the working model.

The back-door path \(T \leftarrow X \to Y\) is blocked by conditioning on \(X\), leaving only the causal path \(T \to Y\) (green). Ignorability holds if this DAG is correctly specified. If there is an unobserved variable \(U\) (motivation, ability) that affects both training participation and earnings, \(X\) no longer satisfies the back-door criterion, so the graph no longer justifies ignorability; weak exchangeability then generally fails (absent special cancellations).

This is one of the central translations in causal inference: the potential-outcomes notation states the assumption in terms of counterfactual independence, while the DAG states it in terms of blocked paths. Together, consistency, weak exchangeability, and positivity identify the ATE via the adjustment formula Equation 4.10, and the back-door criterion provides the graphical justification for why that formula recovers the causal effect. Appendix B provides a formal graphical representation of this connection through single-world intervention graphs (SWIGs).

4.4.4 Overlap and Positivity

WarningWhy Overlap Is Non-Negotiable for Nonparametric Adjustment

Without overlap, some covariate strata contain only treated or only control units. In those strata, \(\E[Y \mid T{=}0, X{=}x]\) or \(\E[Y \mid T{=}1, X{=}x]\) is unobservable, and Equation 4.10 cannot be evaluated at those values of \(x\). Identification of the ATE fails not because the causal structure is wrong, but because the data do not span the support needed.

The ATT can remain identified under a weaker, one-sided overlap condition: together with consistency and the single control-arm exchangeability condition \(Y(0) \indep T \mid X\), it suffices that \(P(T{=}0 \mid X{=}x) > 0\) almost surely on the support of \(X\) in the treated subpopulation, so that the control-arm regression \(\E[Y \mid T{=}0, X{=}x]\) is well-defined at every covariate value where a treated unit appears. Full two-sided overlap is not required: strata containing only control units pose no obstacle for ATT, because ATT averages only over the treated population and such strata never enter that average. Strata containing only treated units, by contrast, are precisely where the one-sided condition fails.

NoteRemark: Three Grades of Positivity Failure

Positivity is a data-support condition, not a causal one: unconfoundedness is a property of the assignment mechanism, but positivity is a property of the joint distribution of \((T, X)\). Three cases should be distinguished.

First, when positivity fails only on a set of \(X\)-values of measure zero, the adjustment formula Equation 4.10 still identifies the estimand, because the offending strata contribute zero mass to the outer expectation — this is why the condition is naturally stated almost-everywhere, \(P(T{=}t \mid X) > 0\) for \(P_X\)-almost every \(x\).

Second, a stratum of positive probability in which treatment level \(t\) is never received — a population zero rather than a sample zero — is a genuine identification failure for the ATE: \(\E[Y \mid T{=}t, X{=}x]\) is undefined on a set of positive mass (though a different estimand, such as the ATT under the one-sided condition above, may remain identified). The failure is one of nonparametric identification by adjustment over the original target population: a parametric outcome model can extrapolate into unsupported strata, but such values are identified by the model restriction, not by the observed data.

Third, small but positive treatment probabilities, or empty cells in a finite sample, are practical estimation problems — sparse strata and inflated variance — and do not by themselves establish population nonpositivity. Part III returns to the estimation consequences of near-positivity violations.

NoteRemark: Continuous Treatments (Technical; Second Reading)

For continuous \(T\), the standardization formula is understood for \(P_T\)-almost every treatment value \(t\) such that the conditional density satisfies \(f_{T \mid X}(t \mid x) > 0\) for \(P_X\)-almost every relevant \(x\): regular conditional distributions are unique only almost everywhere, so a positive density at a single point does not by itself select a unique version of \(\E[Y \mid T{=}t, X{=}x]\) at that point. Pointwise identification at a prespecified treatment value \(t\) additionally requires a regularity condition that selects a unique version of the conditional law, such as continuity in \(t\). The same qualification applies to other identification formulas evaluated at fixed treatment or mediator values, such as the front-door formula of Chapter 8.

4.5 Where the Frameworks Agree and Diverge

4.5.1 A Systematic Comparison

Task Potential Outcomes Do-Calculus / DAG SEM
Define causal estimands (ATE, ATT) \(\checkmark\) Especially natural notation ATE via \(\E[Y \mid \doop(t)]\); ATT requires counterfactual augmentation (Appendix B) Via structural equations
Encode causal assumptions Formal counterfactual restrictions; a causal graph may be added but is not implicit in the notation \(\checkmark\) Explicit directed edges \(\checkmark\) Structural equations
Read conditional independence Requires auxiliary graph or model to read off \(\checkmark\) d-separation Read from the induced graph under dependence restrictions on the exogenous inputs
Identification from observational data Via assignment and counterfactual assumptions plus probability algebra \(\checkmark\) Back-door, front-door, do-calculus Via functional, exclusion, rank, and background-input restrictions
Likelihood construction Typically paired with estimating equations or semiparametric models Identifies observed-data functionals; a likelihood requires an additional statistical model Induced once the structural functions and error laws are sufficiently specified
Cross-world assumptions (e.g. monotonicity) \(\checkmark\) Natural to state Not expressible in single-world graphs; see Appendix B Joint counterfactuals are defined, but restrictions such as monotonicity must be imposed separately
Standard in statistics/epidemiology \(\checkmark\) Dominant Growing rapidly Econometrics
NoteOur Position in This Course

All three frameworks are indispensable. We use potential outcomes to define causal estimands. We use the do-calculus and DAGs for identification. The back-door criterion is the graphical condition that justifies conditional ignorability and hence the adjustment formula through which causal effects are identified. For a formal graphical representation that hosts both frameworks simultaneously, see Appendix B.

Both notations mark the interventional/observational distinction syntactically: \(\E[Y(t)]\) and \(\E[Y \mid T{=}t]\) are different expressions, just as \(f(y \mid \doop(t))\) and \(f(y \mid T{=}t)\) are. The do-operator is preferred in this course because it integrates directly with graph surgery and the do-calculus — the tools by which identification is actually decided. After an interventional estimand has been identified as an observed-data functional, Part III derives estimating equations, influence functions, and efficiency bounds for that functional.

4.6 Summary

  1. The potential outcome \(Y(t)\) is the outcome that would be observed under the intervention \(T = t\). Under the structural causal semantics adopted in these notes (with no interference), \(Y(t)\) has the same distribution as the outcome under the intervention \(\doop(T{=}t)\); the separate consistency assumption links counterfactuals to data through \(Y_i = Y_i(T_i)\).

  2. The ATE and ATT are averages of the unit-level causal effect \(Y_i(1) - Y_i(0) = g(1, \mathbf{X}_i, U_{Y,i}) - g(0, \mathbf{X}_i, U_{Y,i})\) over the full population and the treated subpopulation respectively. In the linear SEM, both equal the structural coefficient \(\beta\). They diverge under heterogeneous effects when treatment selection correlates with individual effect size.

  3. Consistency, weak conditional exchangeability \(Y(t) \indep T \mid X\), and positivity identify the ATE via Equation 4.10 (Proposition 4.2); strong ignorability (Definition 4.1) is a sufficient package. The back-door criterion provides a sufficient graphical condition for the conditional exchangeability assumption used in adjustment formulas.

  4. Single-world intervention graphs (SWIGs) (Richardson and Robins 2014) provide a formal graphical representation of counterfactual variables, making the single-world independence \(Y(t) \indep T \mid X\), for one fixed \(t\), a d-separation statement in a single diagram that hosts both the potential outcome and the natural treatment. The joint strong-ignorability condition involves counterfactuals from two different worlds and is not represented by any single SWIG. This material is developed in Appendix B and is not required for a first-pass understanding of adjustment.

  5. The frameworks are complementary: potential outcomes define estimands; the do-calculus identifies them. The back-door criterion connects the two by providing the graphical condition under which adjustment recovers the causal effect.

NoteRemark: How Much Extra Structure Does NPSEM-IE Impose?

A binary example, following Wang et al. (2026), makes the comparison quantitative. Take \(X\), \(T\), and \(Y\) all binary, so the full joint distribution of \((X,\, T(x),\, Y(t,x))\) for \(t, x \in \{0,1\}\) has \(2^{1 + 2 + 4} - 1 = 127\) free parameters. Three nested assumptions progressively shrink this model:

  • Weak (single-world) ignorability \(T \indep Y(t) \mid X\) for each \(t\) separately: 123 parameters.
  • Strong cross-world ignorability \((Y(0), Y(1)) \indep T \mid X\): 121 parameters.
  • Full NPSEM-IE (all exogenous errors mutually independent): 19 parameters.

NPSEM-IE therefore imposes \(123 - 19 = 104\) additional constraints beyond what is needed for ATE identification. Most are cross-world statements such as \(T(X{=}0) \indep Y(t, X{=}1)\) that remain untestable even under randomized treatment. Working under NPSEM-IE is a substantive modeling choice, not a free consequence of using structural equations.

From Ignorability to Randomization

In observational studies, ignorability must be justified by substantive knowledge encoded in a causal graph. Because unobserved confounding can never be ruled out empirically, this justification is often debated.

Randomized experiments provide a fundamentally different solution. When treatment is assigned randomly — with \(T\) denoting the randomized assignment, which coincides with treatment received in the simple full-compliance trial — \((Y(0),Y(1)) \indep T\) holds by design, guaranteeing ignorability without any covariate adjustment. Chapter 5 studies randomized experiments as the canonical design where causal effects are identifiable directly from the data-generating mechanism.

Causal inference \(=\) counterfactual questions \(+\) graphical assumptions \(+\) statistical estimation.

Identification vs. Estimation

The first four chapters have focused on identification: whether a causal quantity \(\tau = \E[Y(1) - Y(0)]\) can be written as \(\tau = \Phi(P(Y,T,X))\) for some functional \(\Phi\) of the observed distribution. Basic estimators — regression adjustment, standardization, stratification, and inverse probability weighting — are introduced in Chapters 5–6; their systematic asymptotic, influence-function, and semiparametric analysis, together with doubly robust estimators and instrumental variables, is developed in Part III.

4.7 Problems

1. SUTVA and consistency. Suppose \(n = 3\) units receive binary treatments \((T_1, T_2, T_3)\) and unit \(i\)’s outcome may depend on all three treatments: \(Y_i = Y_i(T_1, T_2, T_3)\).

  1. How many potential outcomes does unit 1 have? Write them out explicitly.
  2. SUTVA imposes no interference: \(Y_i(T_1,T_2,T_3) = Y_i(T_i)\). How many distinct potential outcomes remain?
  3. Give a real-world example where no-interference plausibly holds and one where it plausibly fails. In the latter case, redefine the potential outcome using one of: the complete assignment vector, an exposure mapping, partial interference, or cluster-level treatment, and explain what additional design or structural assumption would then be needed.

2. ATE, ATT, and selection bias. Let \((Y(0), Y(1), T) \sim P\) with \(P(T{=}1) = 0.5\). Suppose \(\E[Y(1)] = 3\), \(\E[Y(0)] = 1\), \(\E[Y(1) \mid T{=}1] = 4\), \(\E[Y(0) \mid T{=}1] = 2\), \(\E[Y(1) \mid T{=}0] = 2\), \(\E[Y(0) \mid T{=}0] = 0\).

  1. Compute the ATE and ATT. Are they equal?
  2. Compute the naive contrast \(D_{\mathrm{obs}} = \E[Y \mid T{=}1] - \E[Y \mid T{=}0]\) (the population quantity targeted by the naive estimator). Decompose the gap between this and the ATE into a selection bias term and an ATT–ATE difference.
  3. Give a sufficient graphical condition on the DAG under which the naive contrast \(D_{\mathrm{obs}}\) equals the ATE: the empty set satisfies the back-door criterion, i.e. there is no open back-door path from \(T\) to \(Y\). State the corresponding condition in potential-outcome language (marginal ignorability), and explain why the mere existence of some nonempty adjustment set is not enough.

3. Causal estimands from the structural equation. Consider the nonparametric SEM \(Y = g(T, X, U_Y)\) with binary treatment \(T \in \{0,1\}\), observed covariate \(X\), and unobserved error \(U_Y\) with \(\E[U_Y] = 0\); any further distributional assumptions are stated in the individual parts.

  1. Write \(\tau_{\mathrm{ATE}}\) and \(\tau_{\mathrm{ATT}}\) as expectations of \(g(1, X, U_Y) - g(0, X, U_Y)\) over the appropriate distribution. Under what condition do they coincide?
  2. Specialize to the linear SEM \(g(t, x, u) = \alpha + \beta t + \gamma x + u\). Show that the unit-level causal effect \(Y_i(1) - Y_i(0)\) is constant across all units, and hence \(\tau_{\mathrm{ATE}} = \tau_{\mathrm{ATT}} = \beta\).
  3. Now consider the heterogeneous SEM \(g(t, x, u) = (\alpha + u)\,t + \gamma x\), where \(\E[U_Y] = 0\), \(\mathrm{Var}(U_Y) = \sigma^2 < \infty\), \(\mathrm{Cov}(U_Y, T) = \rho\), \(P(T{=}1) = p \in (0,1)\), and \(X\) is independent of \((T, U_Y)\); assume a joint law compatible with these moments (the Cauchy–Schwarz inequality gives the necessary condition \(|\rho| \le \sigma\sqrt{p(1-p)}\)). Compute \(\tau_{\mathrm{ATE}}\) and \(\tau_{\mathrm{ATT}}\). Show that they differ when \(\rho \neq 0\), and interpret this difference in words.
  4. In part (c), show that the population OLS slope on \(T\) in the regression of \(Y\) on \((1, T, X)\) equals \(\alpha + \rho/p = \tau_{\mathrm{ATT}}\), not \(\tau_{\mathrm{ATE}} = \alpha\). Explain why: the stratum-specific difference \(\E[Y \mid T{=}1, X] - \E[Y \mid T{=}0, X]\) equals \(\alpha + \rho/p\) at every \(X\) (the heterogeneity carried by \(U_Y\) is not observable through \(X\)), so adjusted OLS recovers the ATT, shifted away from \(\tau_{\mathrm{ATE}} = \alpha\) by the selection term \(\E[U_Y \mid T{=}1] = \rho/p\). What additional structure would identify \(\tau_{\mathrm{ATE}}\)? (Hint: randomization of \(T\) would imply \(\E[U_Y \mid T] = 0\). Alternatively, if a richer set of pretreatment covariates \(W\) rendered \(Y(t) \indep T \mid W\), then \(\tau_{\mathrm{ATE}}\) would be identified by standardization over \(W\), even though \(\E[U_Y \mid T]\) need not vanish. A valid instrument does not generally identify the ATE under treatment-effect heterogeneity; in the binary-instrument, binary-treatment setting, relevance, exogeneity, exclusion, and monotonicity together identify the LATE for compliers — see Chapter 7.)

4. SWIGs and ignorability. (Requires Appendix B.) Consider the DAG: \(X \to T\), \(X \to Y\), \(T \to Y\), with \(X\) fully observed.

  1. Construct the SWIG \(\Gcal(t)\) by splitting \(T\) into its random and fixed halves. Draw the result, labeling the random half, fixed half, and the potential outcome \(Y(t)\).
  2. In \(\Gcal(t)\), identify all paths between the random half \(T\) and \(Y(t)\). Apply the SWIG d-separation rules to determine which paths are blocked and which are open before conditioning.
  3. First verify, in the original DAG, that \(X\) satisfies the back-door criterion for \((T, Y)\). Then use d-separation in \(\Gcal(t)\) to verify that \((Y(t) \indep T \mid X)_{\Gcal(t)}\) holds, illustrating the direction of Proposition 4.3.
  4. Now add a hidden common cause \(U \to T\), \(U \to Y\). Draw the revised SWIG. Does the ignorability argument still hold? State the correct conclusion and identify what additional structure (if any) would be needed for identification.

5. Positivity and target populations. Let \(X \in \{0,1\}\) with \(P(X{=}1) > 0\), and suppose \(P(T{=}1 \mid X{=}1) = 1\).

  1. Is the ATE identified by nonparametric adjustment over the original target population? Explain.
  2. State conditions under which the ATT may remain identified.
  3. Contrast this population positivity failure with a finite sample that happens to contain no control units with \(X = 1\) even though \(P(T{=}0 \mid X{=}1) > 0\) in the population.
  4. Explain what additional assumption is being relied upon if a parametric outcome model is extrapolated into the unsupported stratum.
Dehejia, Rajeev H., and Sadek Wahba. 1999. “Causal Effects in Nonexperimental Studies: Reevaluating the Evaluation of Training Programmes.” Journal of the American Statistical Association 94 (448): 1053–62.
Holland, Paul W. 1986. “Statistics and Causal Inference.” Journal of the American Statistical Association 81 (396): 945–60.
LaLonde, Robert J. 1986. “Evaluating the Econometric Evaluations of Training Programmes with Experimental Data.” American Economic Review 76 (4): 604–20.
Neyman, Jerzy Splawa. 1923. “On the Application of Probability Theory to Agricultural Experiments: Essay on Principles, Section 9.” Statistical Science 5 (4): 465–72.
Richardson, Thomas S., and James M. Robins. 2014. ACE Bounds; Single World Intervention Graphs (SWIGs) and Identification of Causal Effects. University of Washington.
Robins, James M. 1986. “A New Approach to Causal Inference in Mortality Studies with a Sustained Exposure Period—Application to Control of the Healthy Worker Survivor Effect.” Mathematical Modelling 7 (9–12): 1393–512.
Rosenbaum, Paul R., and Donald B. Rubin. 1983. “The Central Role of the Propensity Score in Observational Studies for Causal Effects.” Biometrika 70 (1): 41–55.
Rubin, Donald B. 1974. “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies.” Journal of Educational Psychology 66 (5): 688–701.
Rubin, Donald B. 1980. “Randomization Analysis of Experimental Data: The Fisher Randomization Test.” Journal of the American Statistical Association 75 (371): 591–93.
Wang, Linbo, Thomas S. Richardson, and James M. Robins. 2026. “Causal Inference: A Tale of Three Frameworks.” Journal of Data Science 24 (1): 53–85. https://doi.org/10.6339/25-JDS1211.