5 Randomization and Back-Door Adjustment
5.1 Randomized Experiments
5.1.1 The Do-Calculus of Randomization
The defining feature of a randomized experiment is that the analyst assigns treatment by a known randomized mechanism, independent of the potential outcomes and of all pretreatment characteristics. (Under complete randomization the individual assignments are not mutually independent — their total is fixed — but the assignment vector as a whole is independent of everything substantive.) In do-calculus terms, the experimental joint law is a randomized mixture of treatment-specific interventional laws, \[p_R(t, y) \;=\; p_R(t)\, f\!\bigl(y \mid \doop(T{=}t)\bigr),\] where \(p_R(t)\) is the known assignment probability. The experimental law is not itself the degenerate law generated by fixing \(T\) at one value — treatment remains random — but within it, conditioning on the assigned treatment recovers the interventional distribution: \(p_R(y \mid T{=}t) = f(y \mid \doop(T{=}t))\). This equality, which fails in general for observational data, is what the following box establishes.
What the equality assumes beyond the graph. A precision is worth recording about where the extra assumptions enter. The fundamental equality Equation 5.1, read as a statement about randomized assignment, holds by the graph-surgery argument alone: randomization identifies the causal effect of assignment without any compliance assumption. The further conditions below are needed to interpret that assignment effect as the point-treatment effect developed in this chapter.
First, SUTVA in the sense of Chapter 4 — no interference (one unit’s outcome does not depend on another unit’s treatment) and no hidden treatment versions — which makes the scalar notation \(Y_i(t)\) adequate, together with the separate consistency assumption \(Y_i = Y_i(T_i)\). These are assumptions about the science — the treatment and the outcome-generating mechanism — and no randomization protocol, however well executed, can guarantee them.
Second, full compliance is what equates the effect of assignment with the effect of treatment received. It is worth separating the two variables that the notation \(T\) conflates: write \(Z_i\) for the randomized assignment and \(D_i\) for the treatment received. The design randomizes \(Z\), never \(D\). Throughout this chapter we assume full compliance, \(D_i = Z_i\), and write \(T_i\) for their common value. When compliance fails, \(Z\) remains randomized and identifies the intention-to-treat effect of assignment, but \(D\) is no longer randomized: the right-hand side \(f(y \mid D{=}t)\) describes outcomes among those who actually received \(t\) rather than those who were assigned \(t\), and no longer matches the interventional distribution even though the design has severed all arrows into \(Z\). The intent-to-treat versus per-protocol distinction then becomes substantive and is the subject of Chapters 7 and 13; interference and spillover settings require separate frameworks beyond the scope of these notes.
Graphical argument. In the observational DAG, there are arrows into \(T\) from observed covariates \(X\) and unobserved confounders \(U\), creating back-door paths from \(T\) to \(Y\). Randomization physically severs these arrows: assignment is determined by an exogenous randomizer, not by \(X\) or \(U\). Call the resulting graph the randomized-assignment graph: \(T\) remains a random variable, but its only parent is the randomizer. This graph shares the incoming-edge structure of the mutilated graph \(\Gcal_{\overline{T}}\) — no arrows into \(T\) from \(X\) or \(U\) — but the two objects should not be identified: in the randomized-assignment graph \(T\) is stochastic, whereas under \(\doop(T{=}t)\) the mutilated graph fixes \(T\) at the single value \(t\). What matters for identification is the shared feature: there are no back-door paths to block, because none exist.
Potential outcomes statement. In the potential outcomes language, complete or Bernoulli randomization with a common assignment probability implies: \[\bigl(Y(0),\, Y(1)\bigr) \;\indep\; T.\] The potential outcomes are independent of assignment unconditionally — no covariate adjustment is needed. This is the unconditional counterpart of strong ignorability (which conditions on covariates \(X\)). The scoping matters: in stratified, blocked, or covariate-adaptive designs, where the allocation probability depends on design variables, the natural statement is conditional exchangeability given those variables. Throughout this chapter, “randomized experiment” means complete randomization unless stated otherwise.
5.1.2 Estimation of the ATE in a Randomized Experiment
Lemma 5.1 (Identification under Complete Randomization) Under complete randomization, \[\E\bigl[Y(t)\bigr] \;=\; \E\bigl[Y \mid T{=}t\bigr], \qquad t \in \{0,1\}. \tag{5.2}\]
Proof. By the potential-outcomes/do-operator correspondence of Chapter 4, \(Y(t)\) has the same distribution as \(Y\) under the interventional distribution \(f(y \mid \doop(T{=}t))\), so \(\E[Y(t)] = \E[Y \mid \doop(T{=}t)]\). In a completely randomized experiment, the treatment node has no parents in the DAG, so Equation 5.1 gives \(p_R(y \mid T{=}t) = f(y \mid \doop(T{=}t))\). Therefore, within the experimental law, \(\E[Y(t)] = \E[Y \mid T{=}t]\). \(\square\)
A word on what probability law the expectations refer to. Complete randomization by itself does not conjure an external superpopulation: in this lemma, \(\E\) is taken either with respect to a randomly selected unit from the trial population (the empirical finite-population distribution supplying the law) or under the two-phase superpopulation model discussed below, in which units are sampled before being randomized. Randomization supports the exact finite-population statements of Theorem 5.1; it does not make a nonrepresentative trial sample representative of any wider population.
Under complete randomization, the ATE is estimated by the difference-in-means estimator: \[\hat\tau_{\mathrm{DIM}} = \bar Y_1 - \bar Y_0 = \frac{1}{n_1}\sum_{i:\, T_i=1} Y_i - \frac{1}{n_0}\sum_{i:\, T_i=0} Y_i,\] where \(n_1\) and \(n_0\) are the numbers of treated and control units.
Two targets, two frameworks. Before stating the exact properties of \(\hat\tau_{\mathrm{DIM}}\), it is worth separating two estimands that the generic symbol \(\tau_{\mathrm{ATE}}\) tends to blur. The identity Equation 5.2 may be read either with respect to a uniformly selected unit from the trial population or under the two-phase superpopulation model. The theorem below instead conditions on the realized finite population: it treats the \(n\) enrolled units’ potential outcomes as fixed numbers, targets the finite-population ATE \[\bar\tau_n \;=\; \frac{1}{n}\sum_{i=1}^n \bigl\{Y_i(1) - Y_i(0)\bigr\},\] and derives all randomness from the assignment mechanism alone. Randomization by itself licenses design-based inference for \(\bar\tau_n\) — the effect among the units actually randomized (internal validity). Extending conclusions to a wider population requires an additional random sampling or transportability argument (external validity).
Theorem 5.1 (Neyman’s Theorem for the Difference-in-Means Estimator) Under complete randomization, let \(\mathcal{F}_n = \{(Y_i(0), Y_i(1)) : i = 1, \ldots, n\}\) denote the set of potential outcomes, treated as fixed, and let the expectations and variances below be taken with respect to the randomization distribution of the assignment vector. Set \(f_1 = n_1/n\) and let \[\bar Y(t) = \frac{1}{n}\sum_{i=1}^n Y_i(t), \qquad S_t^2 = \frac{1}{n-1}\sum_{i=1}^n \bigl(Y_i(t) - \bar Y(t)\bigr)^2,\] \[S_{01} = \frac{1}{n-1}\sum_{i=1}^n \bigl(Y_i(0) - \bar Y(0)\bigr)\bigl(Y_i(1) - \bar Y(1)\bigr).\] Then the following hold.
Unbiasedness. \(\E\bigl[\hat\tau_{\mathrm{DIM}} \mid \mathcal{F}_n\bigr] = \bar Y(1) - \bar Y(0) = \bar\tau_n\).
Design variance. \[\mathrm{Var}\bigl(\hat\tau_{\mathrm{DIM}} \mid \mathcal{F}_n\bigr) \;=\; \frac{1}{n}\!\left(\frac{n_0}{n_1}\, S_1^2 + \frac{n_1}{n_0}\, S_0^2 + 2\, S_{01}\right) \;=\; \frac{S_1^2}{n_1} + \frac{S_0^2}{n_0} - \frac{S_\tau^2}{n}, \tag{5.3}\] where \(S_\tau^2 = (n-1)^{-1}\sum_{i=1}^n (\tau_i - \bar\tau_n)^2\) is the finite-population variance of the unit-level effects \(\tau_i = Y_i(1) - Y_i(0)\).
Conservative variance estimator. Let \(\hat S_1^2\) and \(\hat S_0^2\) be the within-arm sample variances. Then \[\hat V \;=\; \frac{\hat S_1^2}{n_1} + \frac{\hat S_0^2}{n_0} \tag{5.4}\] satisfies \(\E[\hat V \mid \mathcal{F}_n] - \mathrm{Var}(\hat\tau_{\mathrm{DIM}} \mid \mathcal{F}_n) = S_\tau^2 / n \geq 0\), so \(\hat V\) is conservative. It is unbiased in the special case of a constant unit-level effect, in which case \(S_\tau^2 = 0\).
5.1.3 Fisher’s Randomization Tests vs. Neyman’s Repeated-Randomization Estimation
Both Fisher and Neyman take the randomization mechanism seriously, but they use it for different inferential purposes.
Two inferential targets. Fisher’s randomization framework is built around the sharp null hypothesis of no treatment effect for any unit, \(Y_i(1) = Y_i(0)\) for all \(i\). Under this null, every missing potential outcome is known once the observed outcomes are given. This makes it possible to compute the exact randomization distribution of a chosen test statistic under repeated reassignments of treatment. The exactness rests on four ingredients: a fully known assignment mechanism, completely observed outcomes, a well-defined sharp null, and SUTVA’s no-interference clause; remove any one and the enumeration argument breaks.
Neyman’s framework instead focuses on average causal effects. Its design-based target is the finite-population ATE \(\bar\tau_n\), with the superpopulation ATE arising under the two-phase sampling framework. Neyman’s approach studies the repeated-sampling behavior of estimators such as the difference in sample means.
Practical distinction. Fisher’s framework is naturally aligned with exact hypothesis testing under a sharp null, whereas Neyman’s is naturally aligned with point estimation and uncertainty quantification. The two approaches are complementary rather than contradictory: both rely on the treatment assignment mechanism, but they answer different questions.
5.2 Ignorability
5.2.1 Two Routes to the Same Estimand
Section 5.1 established that a randomized experiment achieves the fundamental equality and that the difference-in-means estimator is exactly unbiased for the finite-population ATE. The mechanism is graphical: randomization physically severs every arrow into \(T\), so no back-door path from \(T\) to \(Y\) exists.
In observational studies, no such physical severance occurs. The treatment was not assigned by the researcher; it was selected — by patients, firms, governments, or individuals — based on their characteristics. Those characteristics may include unmeasured variables \(U\) that also affect the outcome, opening back-door paths that observational data alone cannot close.
The key insight of this chapter is that, when defined for the same target population, both settings can target the same estimand — the ATE — but reach it by fundamentally different routes:
| Randomized experiment | Observational study | |
|---|---|---|
| How ignorability arises | By design: the researcher severs all arrows into \(T\) | By assumption: the analyst asserts that \(X\) is a sufficient pretreatment adjustment set |
| Status of \((Y(0),Y(1)) \indep T \mid X\) | Guaranteed — holds unconditionally (\(X\) not needed) | Assumed — requires \(X\) to block every back-door path |
| Credibility | As strong as the randomization protocol | As strong as substantive knowledge of the data-generating process |
| Testability | Can audit the assignment mechanism | Cannot be verified from the data alone |
| Failure mode | Protocol violations, non-compliance | An unmeasured common cause opening a back-door path that no observed set blocks |
The central pedagogical point. When statisticians speak of “controlling for confounders” or “adjusting for \(X\),” they are invoking the observational route. Statistical adjustment can implement ignorability once it is assumed, but it cannot create it. No amount of covariate adjustment — however flexible the model — can close a back-door path that no observed pretreatment variable intercepts. (When an observed variable does lie on the path, adjustment can block it even though the common cause itself is unmeasured — the graph \(T \leftarrow U \to W \to Y\) of Chapter 3, with \(W\) observed, is the canonical example, and the case study of Section 5.4.4 is a real one.) This is why a well-conducted randomized experiment is considered more credible than even the most carefully adjusted observational study.
In observational studies, the three methods developed in Section 5.3–Section 5.4 all implement the observational route — they assume ignorability and then estimate the identified quantity. The same regression machinery reappears in Section 5.3.5 for a different purpose: precision improvement under randomization, where identification is supplied by the design and no ignorability assumption is needed. Chapter 6 extends the observational route with the propensity score. Chapter 7 abandons the ignorability route entirely when unobserved confounders are present.
5.2.2 Terminology
The formal definition of strong ignorability was given in Chapter 4: joint unconfoundedness \((Y(0), Y(1)) \indep T \mid X\), together with overlap \(0 < P(T{=}1 \mid X) < 1\) almost surely (Rosenbaum and Rubin 1983). The pointwise (weak) form \(Y(t) \indep T \mid X\) for each \(t\) separately is implied by the joint statement and is what is actually used to identify \(\E[Y(t)]\) for a single \(t\), as in Lemma 5.2 below. Functionals of the joint law of \((Y(0), Y(1))\) require additional cross-world restrictions; strong joint exchangeability constrains the assignment mechanism for the pair but is not by itself generally sufficient to identify that dependence.
Several closely related names are used in the literature: unconfoundedness, selection on observables, no unmeasured confounders, and conditional exchangeability. Conventions vary across authors — over weak versus joint exchangeability, mean versus distributional forms, and whether positivity or consistency is bundled into the package — but all express the same core idea. In a randomized experiment, the joint independence holds unconditionally — \(X\) is not needed, and the condition reduces to \((Y(0), Y(1)) \indep T\), the strongest possible form.
5.2.3 Three Languages for Ignorability
Chapter 4 established that under the structural causal model semantics adopted in these notes, a set \(X\) satisfying the back-door criterion implies ignorability. The condition has parallel sufficient formulations in each of the three languages of the causal trinity.
| Language | Statement of ignorability |
|---|---|
| SEM | In the point-treatment SEM \(Y = f_Y(T, X, U_Y)\) with \(T = f_T(X, U_T)\), let \(U_Y\) denote the entire collection of exogenous inputs determining \(Y(t)\). A sufficient condition is \(U_Y \indep T \mid X\) — which follows, for example, from \(U_T \indep U_Y \mid X\). This presumes no shared background cause remains outside \(X\): independence of an equation-specific noise term alone is not sufficient if an additional latent variable enters both mechanisms, as in the confounded SEM of Chapter 1, where \(\varepsilon \indep T\) holds yet \(U\) confounds. |
| DAG / do-calculus | \(X\) satisfies the back-door criterion for the effect of \(T\) on \(Y\): \(X\) blocks all back-door paths from \(T\) to \(Y\) and contains no descendant of \(T\) (Chapter 3). |
| Potential outcomes | \(Y(t) \indep T \mid X\) for \(t \in \{0,1\}\) (weak ignorability; see the exchangeability hierarchy in Chapter 4). |
These three rows are corresponding sufficient formulations under the structural causal model semantics adopted in these notes, together with consistency and no interference — not literally equivalent statements. In particular, the SEM row is sufficient in the displayed point-treatment model, but a general \(Y(t)\) can depend on multiple exogenous inputs through mediators and other ancestors, and strict logical equivalence across the rows requires further bridge assumptions (Appendix B). In practice, which language is most useful depends on the task: the SEM formulation is natural when writing down a structural model and asking whether any omitted variable enters both equations; the DAG formulation is natural for reading off the criterion from a graph; the potential outcomes formulation is natural for writing down estimators and proofs.
5.2.4 What Ignorability Requires
Ignorability is a substantive, not a statistical, assumption. It does not require observing every common cause itself: it requires an observed pretreatment set \(X\) that blocks every back-door path from \(T\) to \(Y\). An unmeasured common cause defeats adjustment only when no observed set intercepts all of the paths it creates; a variable lying downstream of the unobserved cause can block the relevant path even though the cause itself is never seen. The shorthand “no unmeasured confounders” should be read in this path-blocking sense.
5.2.5 From Ignorability to Identification
Chapter 4 established the back-door adjustment formula from consistency, weak conditional exchangeability, positivity, and integrability. We restate the result here as a lemma, with each step of the proof traced back to the assumption it invokes. The explicit tracing makes the failure mode of each assumption transparent: a positivity failure may be visible empirically and may motivate re-targeting to a supported subpopulation, but it prevents nonparametric identification of the original target; weak ignorability and consistency can fail without an equally direct empirical signature.
Lemma 5.2 (Identification of \(\E\{Y(t)\}\)) Suppose: (i) weak ignorability: \(Y(t) \indep T \mid X\); (ii) overlap: \(P(T{=}t \mid X{=}x) > 0\) for \(P_X\)-almost every \(x\); (iii) consistency: \(Y = Y(T)\); (iv) integrability: \(\E|Y(t)| < \infty\). Then: \[\E\bigl[Y(t)\bigr] = \E\Bigl[\E\bigl[Y \mid T{=}t,\, X\bigr]\Bigr] = \int \E\bigl[Y \mid T{=}t,\, X{=}x\bigr]\, p(x)\, dx. \tag{5.7}\]
Proof. Apply the law of iterated expectations, then use each assumption in turn: \[\begin{aligned} \E\bigl[Y(t)\bigr] &= \E\Bigl[\E\bigl[Y(t) \mid X\bigr]\Bigr] && \text{(LIE)} \\ &= \E\Bigl[\E\bigl[Y(t) \mid T{=}t,\, X\bigr]\Bigr] && \text{(weak ignorability)} \\ &= \E\Bigl[\E\bigl[Y \mid T{=}t,\, X\bigr]\Bigr] && \text{(consistency)}. \end{aligned}\] The step labeled “weak ignorability” uses \(Y(t) \indep T \mid X\), which gives \(\E[Y(t) \mid X] = \E[Y(t) \mid T{=}t, X]\) for any value of \(T\) that has positive probability given \(X\) — and overlap guarantees this almost surely, so \(\E[Y \mid T{=}t, X{=}x]\) is well-defined everywhere in the support of \(X\). The step labeled “consistency” uses \(Y = Y(t)\) on the event \(\{T = t\}\). Integrability licenses the law of iterated expectations. \(\square\)
The ATE and ATT follow immediately as corollaries.
Theorem 5.2 (Identification of the ATE and ATT) Under the conditions of Lemma 5.2, imposed for both treatment levels \(t \in \{0,1\}\): \[\tau_{\mathrm{ATE}} = \int \bigl[\E(Y \mid T{=}1, X{=}x) - \E(Y \mid T{=}0, X{=}x)\bigr]\, p(x)\, dx, \tag{5.8}\] \[\tau_{\mathrm{ATT}} = \int \bigl[\E(Y \mid T{=}1, X{=}x) - \E(Y \mid T{=}0, X{=}x)\bigr]\, p(x \mid T{=}1)\, dx. \tag{5.9}\]
Proof. For the ATE, apply Lemma 5.2 separately to \(t = 1\) and \(t = 0\) and subtract.
For the ATT, the lemma cannot simply be re-weighted — it is an unconditional statement, whereas the ATT conditions on \(T{=}1\) — so we argue conditionally. By the law of iterated expectations within the treated subpopulation, \[\tau_{\mathrm{ATT}} = \E_{X \mid T=1}\Bigl[\E\bigl\{Y(1) - Y(0) \mid T{=}1, X\bigr\}\Bigr].\] For the treated potential outcome, consistency alone gives \(\E\{Y(1) \mid T{=}1, X\} = \E(Y \mid T{=}1, X)\). For the counterfactual mean, \[\begin{aligned} \E\bigl\{Y(0) \mid T{=}1, X\bigr\} &= \E\bigl\{Y(0) \mid X\bigr\} && \text{by } Y(0) \indep T \mid X \\ &= \E\bigl\{Y(0) \mid T{=}0, X\bigr\} && \text{by } Y(0) \indep T \mid X \text{ again} \\ &= \E(Y \mid T{=}0, X) && \text{by consistency,} \end{aligned}\] where control-arm positivity ensures the conditional expectation in the second line is well-defined where it is needed. Substituting and integrating against \(p(x \mid T{=}1)\) yields Equation 5.9. \(\square\)
What the proof uses and where it can fail. The table below traces each step back to the assumption it invokes.
| Proof step | Assumption invoked | Failure mode |
|---|---|---|
| \(\E[Y(t) \mid X] = \E[Y(t) \mid T{=}t, X]\) | Weak ignorability | Unmeasured confounder \(U\): \(Y(t)\) depends on \(T\) even within \(X\)-strata |
| \(\E[Y \mid T{=}t, X{=}x]\) is well-defined | Overlap | Structural zero: \(P(T{=}t \mid X{=}x) = 0\) on a positive-probability set; the estimand is not identified for the original population |
| \(\E[Y(t) \mid T{=}t, X] = \E[Y \mid T{=}t, X]\) | Consistency | Interference: \(Y_i\) depends on \(T_j\) for \(j \ne i\); or hidden versions of treatment |
Equations Equation 1.6–Equation 5.9 are the identification results. The next sections present methods for estimating the identified quantities from data.
5.2.6 Which Variables to Condition On: The Pre-Treatment Requirement
The back-door formula conditions on \(X\). Before estimating anything, there is a prior question that is purely graphical: which variables should be in \(X\)? The answer is not “all available variables.” Conditioning on the wrong variable can introduce bias rather than remove it.
The practical rule. Before running any regression or computing any cell means, classify every candidate covariate as pre-treatment or post-treatment using the causal graph. Only pre-treatment variables that satisfy the back-door criterion belong in \(X\). Post-treatment variables should be left out unless the research question is specifically about a direct or mediated effect — in which case Chapter 8 provides the appropriate framework.
5.3 Regression Adjustment and Standardization
5.3.1 The Common Three-Step Logic
Both regression adjustment and standardization are implementations of the same back-door formula Equation 5.8. They differ only in how they estimate the conditional mean \(\mu(t, x) = \E[Y \mid T{=}t, X{=}x]\), not in the underlying procedure. Throughout this section \(X\) denotes a set of pre-treatment covariates satisfying the back-door criterion.
This estimator is called the outcome regression (OR) estimator or the g-computation estimator (Robins 1986). The g-formula name emphasizes that step 3 averages out the covariates \(X\) by plugging in the empirical distribution of \(X\), exactly as the back-door formula requires.
Why the three steps correspond to the identifying formula. The back-door formula says \(\tau_{\mathrm{ATE}} = \int [\mu(1,x) - \mu(0,x)]\, p(x)\, dx\). Step 1 estimates \(\mu\); step 2 evaluates the contrast at the observed covariate value of each unit; step 3 approximates the integral by the sample average, using the empirical distribution of \(X\) as a nonparametric estimate of \(p(x)\). The estimator is plug-in: replace every population quantity in the identification formula with its sample counterpart.
5.3.2 A Worked Example
Consider a binary covariate \(X \in \{0,1\}\) (e.g. sex) and a continuous outcome \(Y\) (e.g. earnings in thousands of dollars). An OLS model fit separately within each treatment arm gives:
| Treated (\(T=1\)): \(n\) | \(\bar{Y}\) | Control (\(T=0\)): \(n\) | \(\bar{Y}\) | |
|---|---|---|---|---|
| \(X = 0\) | 30 | 42 | 70 | 35 |
| \(X = 1\) | 70 | 58 | 30 | 50 |
| All | 100 | 53.2 | 100 | 39.5 |
The unadjusted difference in means is \(53.2 - 39.5 = 13.7\). But \(X{=}1\) units (with higher baseline earnings) are overrepresented in the treated arm: 70% treated vs. 30% control. \(X\) confounds the comparison.
Step 1. The estimated conditional means are \(\hat\mu(1, 0) = 42\), \(\hat\mu(1, 1) = 58\), \(\hat\mu(0, 0) = 35\), \(\hat\mu(0, 1) = 50\).
Step 2. The imputed within-stratum effects are \(42 - 35 = 7\) (for \(X{=}0\) units) and \(58 - 50 = 8\) (for \(X{=}1\) units).
Step 3. Averaging over the marginal distribution of \(X\): 50% of the full sample has \(X{=}0\), 50% has \(X{=}1\) (equal mix in the combined sample of 200), giving \[\hat\tau_{\mathrm{ATE}} = 7 \times 0.5 + 8 \times 0.5 = 7.5.\] After adjusting for \(X\), the estimated treatment effect is 7.5, not 13.7. The difference arises because the unadjusted comparison conflates the treatment effect with the higher baseline earnings of \(X{=}1\) units who happen to be treated more often.
5.3.3 Standardization as a Special Case
Standardization is the name traditionally used in epidemiology and biostatistics for the special case of regression adjustment in which \(X\) is categorical with finite support \(\mathcal{X}\). When \(X\) is categorical, no parametric model for \(\mu(t,x)\) is needed: the conditional mean is simply the sample mean of \(Y\) within each cell \((T{=}t, X{=}x)\). This is equivalent to fitting a fully saturated outcome model.
The resulting estimator is \[\hat\tau_{\mathrm{STD}} = \sum_{x \in \mathcal{X}} \bigl[\hat\mu(1,x) - \hat\mu(0,x)\bigr]\,\hat{p}(x), \tag{5.11}\] where \(\hat\mu(t,x) = \bar{Y}_{t,x}\) and \(\hat{p}(x) = n_x/n\). Equation Equation 5.11 is algebraically identical to the plug-in form of Equation 5.10. Standardization is therefore regression adjustment with a saturated outcome model, and the two estimators coincide whenever \(X\) is categorical.
The worked example above is a standardization: the cell means were computed directly from the data without any parametric model. The formula is also known as the g-formula (Robins 1986) and as direct standardization in epidemiology.
| Regression adjustment | Standardization | |
|---|---|---|
| How \(\hat\mu(t,x)\) is estimated | Regression model: parametric (OLS, logistic) or flexible nonparametric learner | Cell means: \(\bar Y\) within \((T{=}t, X{=}x)\) |
| Works when \(X\) is | Continuous or moderately high-dimensional (through modeling) | Discrete and low-dimensional |
| Main vulnerability | Model misspecification | Sparse or empty cells: undefined cell means, instability |
| Steps 2–3 | Identical: predict both conditional means, average the contrasts | Identical |
When \(X\) is continuous, cell means are unavailable. \(\hat\mu(t,x)\) is then estimated by a regression model and the three-step recipe applies with no other change.
5.3.4 Model Specification and What Can Go Wrong
To see exactly what consistency of the plug-in estimator requires, write \(P_n f = n^{-1}\sum_{i=1}^n f(X_i)\) and \(Pf = \E\{f(X)\}\), with \(\hat\mu_t(x) = \hat\mu(t,x)\). Then \[\hat\tau_{\mathrm{OR}} - \tau_{\mathrm{ATE}} = P_n(\hat\mu_1 - \mu_1) \;-\; P_n(\hat\mu_0 - \mu_0) \;+\; (P_n - P)(\mu_1 - \mu_0).\] Suppose the observations are i.i.d., \(\E|\mu_1(X) - \mu_0(X)| < \infty\), and the estimated regressions satisfy the empirical-\(L_1\) condition \[P_n\bigl|\hat\mu_t - \mu_t\bigr| = \frac{1}{n}\sum_{i=1}^n \bigl|\hat\mu(t, X_i) - \mu(t, X_i)\bigr| \;\overset{p}{\to}\; 0, \qquad t \in \{0, 1\}.\] The first two terms are then \(o_p(1)\), while the law of large numbers handles the third. Consequently \(\hat\tau_{\mathrm{OR}} \overset{p}{\to} \tau_{\mathrm{ATE}}\).
The empirical form of the first condition matters: because \(\hat\mu\) is fitted and evaluated on the same observations, convergence in the population sense is not by itself enough — a learner can have negligible population error while behaving anomalously at its own training points. The empirical-\(L_1\) condition follows, for example, from uniform convergence; from population-\(L_1(P_X)\) convergence together with suitable Glivenko–Cantelli or stability conditions; or from independent (cross-fitted) evaluation under integrability conditions. A correctly specified finite-dimensional parametric model satisfies it under standard regularity conditions.
The condition is sufficient, not necessary: if \(\hat\mu(t,\cdot)\) converges to some limit \(\mu_t^*\), the OR estimator targets \(\E[\mu_1^*(X) - \mu_0^*(X)]\), and it is consistent whenever this integrated contrast equals \(\tau_{\mathrm{ATE}}\) even though the two functions are pointwise wrong everywhere. Section 5.3.5 and the lab of Section 5.6 are built around exactly this cancellation under randomization. In an observational study, however, no design guarantees the cancellation, and consistent estimation of the outcome regressions is the practical requirement. This is the main vulnerability.
Misspecification bias. If the true \(\mu(t,x)\) is nonlinear but a linear model is used, the estimated treatment effect absorbs part of the functional form error. The bias can be in either direction. Flexible learners can reduce approximation bias, though they may introduce greater estimation variability and overfitting; in practice they are paired with sample splitting (cross-fitting), which separates fitting from evaluation. For plain outcome regression that is cross-fitting’s whole role: it does not repair an inconsistent outcome model, and it does not confer first-order robustness to nuisance-estimation error. That stronger role emerges only when cross-fitting is paired with an orthogonal score, as in the AIPW/DML estimators of Chapter 12.
Extrapolation. Step 2 predicts \(\hat{Y}_i(0)\) for treated units with covariate values that may lie outside the support of the control group. In regions with no control data, \(\hat\mu(0, X_i)\) is extrapolated from the model rather than estimated from comparable units. The propensity score overlap condition (Chapter 6) formalizes when extrapolation is unavoidable.
5.3.5 Regression Adjustment in Randomized Experiments: Lin (2013)
In a randomized experiment, outcome regression can be used to reduce variance even though covariate adjustment is not needed for unbiasedness. The regression-adjusted estimator of Lin (2013) fits the fully interacted OLS model \[Y_i = \alpha + \beta T_i + \gamma^\top \tilde{X}_i + \delta^\top (T_i \cdot \tilde{X}_i) + \varepsilon_i, \tag{5.12}\] where \(\tilde{X}_i = X_i - \bar{X}\) are mean-centered covariates, and takes \(\hat\beta\) as the point estimate of the finite-population ATE \(\bar\tau_n\). Before discussing its properties we explain, in two short steps, why \(\hat\beta\) coincides with the three-step outcome-regression estimator fit arm-by-arm with a linear model.
Step A: Single-arm regression with a centered covariate. Consider the single-arm OLS fit minimizing \(\sum_i T_i\{Y_i - \beta_0 - \beta_1^\top (X_i - \bar X)\}^2\). The first-order condition for \(\beta_0\) gives \(\hat\beta_0 = \bar Y_1 - \hat\beta_1^\top(\bar X_1 - \bar X)\). On the other hand, the three-step OR estimator with a linear within-arm model \(\hat\mu(1, x) = \hat a_1 + \hat b_1^\top x\), averaged over all units, equals \(\hat\mu_{1,\mathrm{reg}} = \hat a_1 + \hat b_1^\top \bar X\). The within-treated slope \(\hat b_1\) coincides with \(\hat\beta_1\) (centering shifts the intercept but not the slope), and at the within-treated sample mean \(\bar Y_1 = \hat a_1 + \hat b_1^\top \bar X_1\), so \[\hat a_1 + \hat b_1^\top \bar X = \bar Y_1 - \hat b_1^\top(\bar X_1 - \bar X) = \hat\beta_0.\] Hence \(\hat\beta_0 = \hat\mu_{1,\mathrm{reg}}\): when the covariate is centered at the full-sample mean \(\bar X\), the OLS intercept is the regression-adjusted estimator of the treated-arm mean. The symmetric argument in the control arm gives \(\hat\mu_{0,\mathrm{reg}}\).
Step B: Interacted regression over both arms. Write \(\widehat{Y_i(t)} = \hat\mu_{t,\mathrm{reg}} + \hat B_1(t)^\top (X_i - \bar X)\), where \(\hat B_1(t)\) is the within-arm slope. The arm-specific predictions combine as \[\hat Y_i = \hat\mu_{0,\mathrm{reg}} + T_i\bigl(\hat\mu_{1,\mathrm{reg}} - \hat\mu_{0,\mathrm{reg}}\bigr) + \hat B_1(0)^\top (X_i - \bar X) + T_i (X_i - \bar X)^\top \bigl(\hat B_1(1) - \hat B_1(0)\bigr).\] The right-hand side is exactly the fit of the interacted regression, and reading off coefficients identifies \(\hat\beta\) with \(\hat\mu_{1,\mathrm{reg}} - \hat\mu_{0,\mathrm{reg}}\). Setting \(\hat\tau_{\mathrm{reg}} := \hat\mu_{1,\mathrm{reg}} - \hat\mu_{0,\mathrm{reg}}\), we have the algebraic identity \[\hat\beta \;=\; \hat\tau_{\mathrm{reg}}. \tag{5.13}\]
The Lin estimator is therefore not a separate method: it is the three-step outcome-regression estimator with linear within-arm models, recast as a single regression. The interacted representation makes computation of the estimator and of its Huber–White sandwich standard error immediate in standard software; the asymptotic validity of that variance estimator under the randomization distribution is not an algebraic consequence of the recasting but part of Lin’s theoretical result.
Theorem 5.3 (Design-Consistency of the Regression-Adjusted Estimator (Lin 2013)) Consider a sequence of finite populations with fixed covariate dimension in which \(n_1/n \to \rho \in (0,1)\), the population Gram matrices converge to a nonsingular limit, and the covariates and potential outcomes satisfy moment and leverage conditions sufficient for the arm-specific OLS coefficients to be root-\(n\) consistent for their finite-population projection coefficients. Under complete randomization, \[\hat\beta \;-\; \bar\tau_n \;\overset{p}{\to}\; 0\] regardless of whether the linear working models are correctly specified. The result is asymptotic: unlike \(\hat\tau_{\mathrm{DIM}}\), the estimator \(\hat\beta\) need not be exactly unbiased in finite samples.
Under the two-phase superpopulation framework, the finite-population target itself converges, \(\bar\tau_n \to_p \tau_{\mathrm{ATE}}\), so \(\hat\beta\) is consistent for the superpopulation ATE as well; the design-based statement is the more primitive one.
Byproduct: a linearization of the regression estimator. Equation Equation 5.15, applied to both arms, gives \[\hat\tau_{\mathrm{reg}} \;\simeq\; \bar\tau_n + \frac{1}{n_1}\sum_{i=1}^n T_i\, e_i(1) - \frac{1}{n_0}\sum_{i=1}^n (1 - T_i)\, e_i(0), \tag{5.16}\] the linearization at the heart of the design-based analysis. Replacing \(Y_i\) by the smaller residual \(e_i(t)\) is the mechanism behind the potential variance reduction. Marginal residual-variance reduction alone does not settle the comparison: the design variance also contains the cross-potential-outcome term, exactly as in Theorem 5.1, and it too changes under residualization. Lin’s theorem establishes that, for the fully interacted specification under its regularity conditions, the combined asymptotic variance cannot exceed that of the unadjusted difference in means. The full calculation is in Lin (2013); see also Freedman (2008).
5.4 Stratification
5.4.1 The Idea
Stratification (subclassification) is a nonparametric implementation of the back-door formula. Rather than modeling \(\E[Y \mid T, X]\), the analyst divides the sample into strata — subgroups of units with the same or similar values of \(X\) — and estimates the treatment effect within each stratum.
Within a stratum \(\mathcal{S}_k = \{i : X_i \in B_k\}\), the treated and control units have similar covariate distributions. Narrowness alone is not enough: coarsening \(X\) into bins does not automatically inherit exchangeability, and residual confounding can persist within every finite-width stratum. What justifies the method is a combination of conditions: exchangeability given the exact \(X\), treatment propensity and outcome regressions that vary smoothly within strata, stratum widths that shrink with the sample size, and adequate within-stratum overlap. Under these conditions the within-stratum difference in means \[\hat\tau_k = \bar Y_{1,k} - \bar Y_{0,k}\] is an approximately unconfounded comparison: its residual coarsening bias vanishes as the strata shrink. Finite-width strata are not literally unbiased.
5.4.2 The Stratified Estimator
The overall ATE is estimated as a weighted average of within-stratum estimates: \[\hat\tau_{\mathrm{STRAT}} = \sum_{k=1}^K \hat\tau_k \cdot \frac{n_k}{n}, \tag{5.17}\] where \(K\) is the number of strata and \(n_k\) the number of units in stratum \(k\). Cochran (1968) showed — in his setting of a single continuous confounder under particular distributional models, not as a universal fact — that with \(K = 5\) strata of equal size, roughly 90% of the bias is removed; with \(K = 10\) strata, over 95%.
5.4.3 Limitations
Stratification faces the curse of dimensionality: with \(p\) binary covariates there are \(2^p\) cells, many of which will be empty in practice. This motivates two solutions studied in this course:
- Regression modeling (Section 5.3) permits interpolation across continuous or sparse cells; exact cell standardization remains nonparametric but becomes infeasible as the dimension grows.
- Propensity score methods (Chapter 6) collapse the multivariate \(X\) to a scalar \(\pi(X) = P(T{=}1 \mid X)\) and adjust on the scalar. The Propensity Score Theorem of Chapter 6 guarantees that conditioning on the true score \(\pi(X)\) balances the full multivariate \(X\); an estimated score delivers only approximate balance, which must be checked with diagnostics.
5.4.4 Case Study: The Payoff to Attending a Selective College
The identification assumptions of Section 5.2, and the stratified estimator just developed, are best appreciated through a real study in which ignorability is initially not credible, and a clever design partially restores it — by constructing the strata. Dale and Krueger (2002) ask whether attending a more selective college causally raises later earnings. Let \(T\) denote the selectivity of the college a student attends (measured by the school’s average SAT score) and \(Y\) the student’s log earnings observed in 1995, roughly two decades after the 1976 entering cohort they study began college. Earlier studies had found that a 100-point increase in school-average SAT is associated with earnings 3–7 percent higher; the basic model of Dale and Krueger (2002), which adjusts for the student’s own SAT score, high school rank, race, sex, athletic participation, and parental income, gives 7.6 percent in the College and Beyond survey of 30 selective institutions.
Is this a causal effect? The threat is an unmeasured confounder — a variable that influences both where a student enrolls and how much the student later earns, but that no data set records. Model the admissions decision of college \(j\) on applicant \(i\) as a threshold rule, \[\text{admit $i$ to college $j$} \quad\Longleftrightarrow\quad Z_{ij} = \gamma_1 X_{1i} + \gamma_2 X_{2i} + e_{ij} > C_j, \tag{5.18}\] where \(X_1\) collects characteristics observed by both the admissions committee and the analyst (SAT score, high school rank), \(X_2\) collects characteristics observed by the committee but not by the analyst (motivation, ambition, and maturity as judged from essays, interviews, and recommendation letters), \(e_{ij}\) is idiosyncratic committee taste, and \(C_j\) is the cutoff of college \(j\). If the labor market also rewards \(X_2\), then students at selective colleges have higher earnings capacity regardless of where they enroll, and adjustment for \(X_1\) alone leaves the back-door path through \(X_2\) open: ignorability given \(X_1\) fails, and the 7.6 percent figure is biased upward.
In the potential-outcomes language of Chapter 4 the problem is stated in one line. Simplify to a binary treatment: \(T = 1\) if the student enrolls in the more selective tier. The naive contrast then decomposes as \[\underbrace{\E[Y \mid T{=}1] - \E[Y \mid T{=}0]}_{\text{observed gap}} = \underbrace{\E[Y(1) - Y(0) \mid T{=}1]}_{\tau_{\mathrm{ATT}}} + \underbrace{\E[Y(0) \mid T{=}1] - \E[Y(0) \mid T{=}0]}_{\text{selection bias}}, \tag{5.19}\] obtained by adding and subtracting \(\E[Y(0) \mid T{=}1]\) and applying consistency. The selection-bias term compares the earnings the selective-college students would have had anyway with the earnings of everyone else. Under the admissions model Equation 5.18 this term is plausibly positive, but a step of care is needed: the model governs admission, while \(T\) records the selectivity attended, and admission does not determine how students choose among the schools that admit them. Attending a selective school does require clearing its high cutoff, so attendees have high values of the latent index — and hence of \(X_2\) — on average; if, in addition, the labor market rewards \(X_2\) and matriculation choice does not systematically reverse this ranking, then \(\E[Y(0) \mid T{=}1] > \E[Y(0) \mid T{=}0]\), and the observed gap overstates \(\tau_{\mathrm{ATT}}\). Adjusting for \(X_1\) shrinks the selection-bias term but cannot drive it to zero, because part of it flows through \(X_2\).
The design. The insight of Dale and Krueger (2002) is that the admissions process itself measures \(X_2\), imperfectly but repeatedly. Each accept or reject decision reveals, through Equation 5.18, whether the latent index cleared the threshold \(C_j\). A student admitted by a college with cutoff \(C_1\) but rejected by one with cutoff \(C_2 > C_1\) has, absent noise, an index bracketed between \(C_1\) and \(C_2\). As accept/reject decisions accumulate over a fine range of cutoffs, the profile \(A\) of admissions outcomes pins down the index ever more tightly. Conditioning on \(A\) therefore approximates conditioning on \(\gamma_1 X_1 + \gamma_2 X_2\) — a function of the unobservable — even though \(X_2\) itself is never seen. Operationally, Dale and Krueger (2002) group students who applied to, and were accepted and rejected by, the same set of colleges (“matched applicants”) and include a dummy variable for each group in the wage regression, comparing only students within a group who nonetheless enrolled at schools of different selectivity. The authors describe this as extending selection on the observables to “selection on the observables and unobservables.”
The table below shows how the grouping works for seven hypothetical applicants. Schools whose average SAT scores fall in the same 25-point interval are treated as equally selective.
| Student | App 1 | App 2 | App 3 | Matched group |
|---|---|---|---|---|
| A | 1300 / R | 1240 / A* | 1190 / A | 1 |
| B | 1305 / R | 1245 / A | 1185 / A* | 1 |
| C | 1355 / A* | 1310 / A | — | 2 |
| D | 1350 / A | 1305 / A* | — | 2 |
| E | 1370 / A* | 1315 / A | — | 2 |
| F | 1200 / A* | — | — | excluded |
| G | 1400 / R | 1150 / A* | — | no match |
Students A and B applied to equivalent — not literally identical — schools and received identical decisions, so they form group 1; within the group they enrolled at schools 55 points apart, which is exactly the within-group variation the estimator uses. Students C, D, and E form group 2; note that D chose the less selective option from the same admitted set — informative variation, but also precisely the kind of choice the matriculation caveat below is about. In the actual data, 6,335 of the 14,238 workers could be matched this way, forming 1,232 groups. The exclusions are not innocuous for the estimand: because only matchable applicants contribute, the design naturally targets an effect for the matched (overlap) population rather than the full entering cohort.
The estimator itself should look familiar, though one distinction matters: with one dummy variable per group, the wage regression compares earnings only among students who share a group label — the matched groups play exactly the role of strata, so the design principle is that of the stratified estimator Equation 5.17. The two are not, however, generally algebraically identical. For a binary treatment with no additional covariates, the common-slope fixed-effects regression yields \[\hat\beta_{\mathrm{FE}} = \frac{\sum_{k=1}^K (n_{1k} n_{0k} / n_k)\, \hat\tau_k}{\sum_{k=1}^K n_{1k} n_{0k} / n_k},\] which weights each within-group contrast by its within-group treatment variation, whereas Equation 5.17 uses the target-population weights \(n_k/n\). In the actual application, moreover, treatment is continuous school selectivity, so the regression estimates a common within-group linear slope. The matched groups supply a stratified design; fixed-effects regression supplies a closely related within-stratum estimator.
If students choose from their admitted set for reasons unrelated to \(X_2\), then there is no \(X_2 \to T\) arrow, the set \((X_1, A)\) satisfies the back-door criterion, and ignorability \(Y(t) \indep T \mid X_1, A\) holds even though a confounder remains unmeasured. Two qualifications keep the claim honest. First, a choice driver may be safely left outside the adjustment set only if it does not also affect later earnings: geography or cost that influences both college choice and earnings is simply another confounder. Second, identification also requires positivity within the target admissions profiles: matched groups in which every student attended a school of the same selectivity exhibit no within-profile treatment variation. The lesson is that an unobserved common cause need not itself be measured: what identification requires is an observed pretreatment quantity that blocks every back-door path the unobservable opens, and a design can sometimes locate one.
Findings. Within matched-applicant groups, the large positive association between school selectivity and earnings disappears: most selection-adjusted specifications are near zero, and one is significantly negative. All coefficients are from Table III of Dale and Krueger (2002): the coefficient on school-average SAT (per 100 points) falls from \(0.076\) (s.e. \(0.016\)) in the basic model to \(-0.016\) (\(0.022\)) when matching on schools of similar selectivity, and to \(-0.106\) (\(0.036\)) when matching students who applied to, and were accepted or rejected by, exactly the same schools; matching on Barron’s selectivity categories gives \(0.004\) (\(0.016\)). A nested variant, the “self-revelation” model — which controls for the average SAT score of the schools the student applied to — gives \(-0.001\) (\(0.018\)).
Two further results sharpen the interpretation. First, the average SAT score of the schools that rejected the student predicts earnings strongly (coefficient \(0.072\), s.e. \(0.012\)). Under the maintained assumption that rejected-school selectivity has no causal path to later earnings (itself a substantive exclusion, not a logical fact), its predictive power is a pure signature of confounding — the application-and-admissions record is informative about \(X_2\) — and it falsifies the unadjusted comparison without any appeal to the unobservable itself. Anticipating the language of Chapter 9, this is a negative control or falsification argument. (Translating its magnitude into the bias of the treatment-effect estimate requires additional calibration structure; detection and quantification are different tasks.) Second, the near-zero average masks heterogeneity: interacting selectivity with parental income shows a gain of about 8 percent (per 200 SAT points) for students from bottom-decile-income families, versus essentially nil at mean income — a reminder that “no average effect” is not “no effect for anyone.”
5.5 Simpson’s Paradox
Simpson’s paradox — the reversal of an association when conditioning on a third variable — was introduced in Chapter 1. There, a flu-vaccination example showed that the vaccine appeared harmful in the aggregate yet beneficial within every age stratum, because elderly individuals were disproportionately vaccinated. Conditioning on age via the back-door formula recovered the correct causal effect.
This chapter’s development gives that example a second layer of interpretation. The back-door formula used in Chapter 1 is precisely the standardization estimator Equation 5.11: within-stratum conditional means averaged over the marginal distribution of the confounder rather than its treatment-conditional distribution. The weighted average corrects the confounding; the pooled association does not, because it weights strata by \(P(X \mid T)\) instead of \(P(X)\). One qualifier keeps the resolution honest: standardizing over age settles the example only under the maintained causal assumptions — that age is a sufficient adjustment set, that positivity holds, and that consistency is credible. The reversal by itself is a fact about conditional proportions; it does not reveal which comparison, marginal or conditional, is the causal one. That verdict comes from the graph.
5.5.1 When Should You Not Condition on \(X\)?
Simpson’s paradox also has a mirror image: conditioning on a variable can create a spurious association or make a true causal effect disappear. This occurs when \(X\) is a mediator or a collider. The back-door criterion gives the correct prescription: include \(X\) in the adjustment set only if it blocks back-door paths without opening collider paths.
We revisit the mediator case in detail in Chapter 8; the collider case was covered in Chapter 2.
5.6 Lab: Simulation Study of the Outcome Regression Estimator
The identification results of Section 5.2 determine the target functional before any outcome model is chosen. Estimating that functional by the outcome-regression plug-in then creates a bias–variance tradeoff. In an observational study, consistent estimation of the treatment-specific outcome regressions is the standard sufficient condition for plug-in consistency. Under a completely randomized design (CRD), by contrast, Theorem 5.3 gives the fully interacted linear OR estimator a separate design-consistency guarantee even when its working outcome models are misspecified; the guarantee disappears as soon as treatment is confounded.
Estimators. Two outcome regression estimators of the form Equation 5.10 are compared. Both follow the three-step recipe. Linear OR: within each arm, fit the simple linear regression \(Y \sim X\) by OLS. This model is misspecified: the true conditional mean \(\mu(t, x) = t + 3x^2 + 3tx^2\) contains a quadratic-in-\(X\) term and a \(tx^2\) interaction. Local-linear OR: within each arm, estimate \(\mu(t,x_0)\) by local linear regression using a Gaussian kernel with bandwidth \(h = 0.5 \cdot n^{-1/5}\). Local linear regression removes the leading-order boundary bias, which matters here because the observational propensity score places treated and control units in different regions of \([0,1]\).
Results (\(n = 1{,}000\), \(B = 2{,}000\) replications, seed 2024; true \(\tau_{\mathrm{ATE}} = 2.000\)):
| Design | Estimator | Mean | Bias | SD | RMSE |
|---|---|---|---|---|---|
| CRD | Linear OR | 1.9964 | −0.0036 | 0.0709 | 0.0710 |
| CRD | Local-linear OR | 2.0260 | +0.0260 | 0.0685 | 0.0732 |
| Observational | Linear OR | 1.9674 | −0.0326 | 0.0895 | 0.0952 |
| Observational | Local-linear OR | 2.0273 | +0.0273 | 0.0951 | 0.0989 |
1. Under CRD, the linear OR estimator is consistent despite model misspecification. Conditional on each simulated finite population, Theorem 5.3 gives \(\hat\tau_{\mathrm{OR}} - \bar\tau_n \overset{p}{\to} 0\) under the randomization distribution, even though the linear outcome model is wrong. Because the simulation first draws the units i.i.d., \(\bar\tau_n \overset{p}{\to} \tau_{\mathrm{ATE}} = 2\); combining the two phases gives consistency. Randomization, rather than correctness of the working model, supplies the guarantee.
2. Under an observational design, the misspecified linear OR is biased. The bias grows from \(-0.0036\) under CRD to \(-0.0326\). The omitted \(X^2\) and \(TX^2\) terms cause \(\hat\mu(t,x)\) to misrepresent the outcome surface, and because high-\(X\) units are predominantly treated (the propensity score ranges from 0.12 to 0.95), the misfit is systematically amplified in the direction of the confounding. In a CRD the same misfit is present, but the balanced assignment averages it out.
3. The local-linear OR reduces bias at the cost of higher variance. Under the observational design the local-linear OR reduces the absolute bias from \(0.0326\) to \(0.0273\), but its standard deviation is 0.0951 compared to 0.0895 for the misspecified linear model. The RMSE comparison (0.0952 vs. 0.0989) therefore favors the linear estimator, even though it is biased. This is the classic bias–variance tradeoff.
5.7 Where the Propensity Score Fits
The three methods developed above — regression adjustment, standardization, and stratification — all implement the same identifying formula Equation 5.8, and all three therefore rest on the same causal identification conditions: consistency, conditional exchangeability, and positivity. What distinguishes them is only how they handle the conditioning on the high-dimensional covariate vector \(X\), and what each additionally requires at the estimation stage:
| Method | How it handles \(X\) | Method-specific estimation requirement |
|---|---|---|
| Outcome regression / standardization | Estimates \(\mu(t,x)\): parametrically, nonparametrically, or by saturated cells | Sufficient convergence of the estimated conditional means — correct parametric specification is one route, not the only one |
| Stratification | Coarsens \(X\) into cells; estimates effect in each | Vanishing coarsening bias (narrow strata), adequate within-cell overlap, growing cell sizes |
| Propensity score (Ch. 6) | Reduces \(X\) to the scalar balancing score \(\pi(X)\) | Consistent score estimation, or estimator-specific balance conditions, together with overlap (a correct propensity model does not repair unmeasured confounding) |
When \(X\) is low-dimensional, stratification and regression work well. When \(X\) has many components, direct stratification suffers from the curse of dimensionality. The propensity score provides an identification-preserving dimension reduction: Chapter 6 shows that adjusting for the scalar \(\pi(X)\) suffices for identification even when \(X\) is high-dimensional. (The theorem reduces the dimension of the balancing variable, not every statistical difficulty — estimating \(\pi\) with high-dimensional \(X\) can itself be hard.)
5.8 Summary
Randomization as graph surgery. A randomized experiment removes all arrows into \(T\) from variables in the scientific causal system, so the fundamental equality \(p_R(y \mid T{=}t) = f(y \mid \doop(T{=}t))\) holds within the experimental law and the difference-in-means estimator is exactly unbiased for the finite-population ATE \(\bar\tau_n\), without covariate adjustment. Reaching the superpopulation ATE additionally requires a sampling or transportability argument.
Ignorability and support. Under consistency, weak conditional exchangeability, positivity for both treatment levels, and integrability, the back-door adjustment formula identifies the ATE. Strong ignorability is a sufficient package for the exchangeability and positivity parts; consistency remains a separate assumption. In an observational study, exchangeability and consistency are substantive conditions that cannot be verified from the observed law; population positivity is likewise not fully testable from a finite sample, although empirical overlap can and should be diagnosed.
Three estimation strategies. Regression adjustment (g-formula), standardization, and stratification all implement the same identifying formula via different modeling choices. In observational studies, a standard sufficient condition for regression and standardization is consistent estimation of the relevant outcome regressions, and stratification requires strata fine enough to achieve approximate balance; under a randomized experiment, regression adjustment remains design-consistent for the finite-population ATE even when used as a misspecified assisting model (Theorem 5.3).
Simpson’s paradox. Pooled associations can reverse within subgroups when a confounder determines treatment selection. The correct causal estimate requires standardization over the marginal covariate distribution, guided by the back-door criterion.
The propensity score dimension reduction. When \(X\) is high-dimensional, direct stratification or cell-by-cell standardization breaks down. The propensity score \(\pi(X)\) provides a scalar balancing score that preserves the adjustment identification result. Chapter 6 develops the theory and estimation methods.
5.9 Problems
1. Randomization and the do-calculus. Let the observational DAG be \(\{U \to T,\; U \to Y,\; X \to T,\; X \to Y,\; T \to Y\}\) with \(U\) unobserved.
- Identify all back-door paths from \(T\) to \(Y\).
- Does \(X\) satisfy the back-door criterion? Does \(U\)? Explain.
- Now suppose treatment is randomized. Draw the modified DAG and identify all back-door paths. Show that \(p_R(y \mid T{=}t) = f(y \mid \doop(T{=}t))\) holds within the randomized experimental law using Rule 2 of the do-calculus. (Hint: identify the appropriate \((X, Z, W)\) instantiation, construct \(\Gcal_{\underline{T}}\), and verify the d-separation condition.)
- Under randomization, is covariate adjustment on \(X\) necessary for unbiasedness? Is it ever beneficial? Explain.
2. Ignorability and the back-door criterion. Consider the DAG \(\{X \to T,\; X \to Y,\; T \to Y,\; U \to Y\}\) with \(U\) unobserved.
- Does \(\{X\}\) satisfy the back-door criterion? Write out the identifying formula for \(\tau_{\mathrm{ATE}}\).
- Now add the edge \(U \to T\). Does \(\{X\}\) still satisfy the back-door criterion? What does the criterion require when both \(X\) and \(U\) affect \(T\)?
- Explain in words what “unconfoundedness given \(X\)” means about the role of \(U\) in the data-generating process.
3. Standardization. A study of job training (\(T\)) and earnings (\(Y\), in thousands) produces the following cell means, with \(P(X{=}0) = 0.4\) and \(P(X{=}1) = 0.6\):
| \(X\) | \(\hat\mu(1,x)\) | \(\hat\mu(0,x)\) | \(n_x/n\) |
|---|---|---|---|
| 0 | 28 | 22 | 0.4 |
| 1 | 35 | 31 | 0.6 |
- Compute \(\hat\tau_{\mathrm{ATE}}\) using the standardization formula.
- Compute \(\hat\tau_{\mathrm{ATT}}\), given that \(P(T{=}1) = 0.5\) and \(P(X{=}1 \mid T{=}1) = 0.8\) — so that \(P(X{=}1 \mid T{=}0) = 0.4\), consistent with the marginal \(P(X{=}1) = 0.6\). The ATT-weighting density in Equation 5.9 places mass \(0.2\) on \(X{=}0\) and \(0.8\) on \(X{=}1\).
- The unadjusted difference in means is \(33.6 - 25.6 = 8\). Compare to your answers in (a) and (b) and explain the discrepancy.
4. Stratified standardization. You have a binary confounder \(X \in \{0,1\}\) and form two strata. Within stratum \(X{=}0\): \(n_{X=0} = 600\), \(\bar Y_1 = 10\), \(\bar Y_0 = 8\). Within stratum \(X{=}1\): \(n_{X=1} = 400\), \(\bar Y_1 = 15\), \(\bar Y_0 = 12\).
- Compute the stratified ATE estimator \(\hat\tau_{\mathrm{STRAT}}\).
- Suppose an unadjusted difference-in-means gives \(\hat\tau_{\mathrm{DIM}} = 5.5\). Explain why the two estimates differ and, under the assumption that \(X\) is a sufficient adjustment variable, which one is the appropriate causal estimate.
- Describe one limitation that would arise if \(X\) were a continuous variable with 15 dimensions.
5. Simpson’s paradox. A hospital reports that patients who receive intensive care (ICU admission, \(T{=}1\)) have a higher mortality rate than those who do not: \(P(Y{=}1 \mid T{=}1) = 0.30\) vs. \(P(Y{=}1 \mid T{=}0) = 0.10\). For concreteness, assume 1,000 patients in each treatment arm.
- Construct a numerical example (a \(2 \times 2 \times 2\) table with disease severity \(X\) as the confounder) consistent with the pooled numbers and yet showing ICU admission reduces mortality within both severity strata.
- Compute the standardized ATE using the marginal distribution of \(X\).
- Draw the DAG. Identify the back-door path that the pooled comparison fails to block.
- A colleague suggests that since the hospital treats patients with both mild and severe illness, the pooled statistic is the right answer. Explain, using potential outcomes notation, why this argument is incorrect.
6. Finite-population enumeration. Consider a finite population of \(n = 4\) units with potential-outcome schedule
| Unit \(i\) | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| \(Y_i(0)\) | 0 | 1 | 1 | 2 |
| \(Y_i(1)\) | 1 | 3 | 2 | 4 |
and complete randomization with \(n_1 = n_0 = 2\).
- Compute the finite-population ATE \(\bar\tau_n\) and the finite-population variances \(S_1^2\), \(S_0^2\), \(S_\tau^2\), and the covariance \(S_{01}\).
- Enumerate all \(\binom{4}{2} = 6\) equally likely assignment vectors and compute \(\hat\tau_{\mathrm{DIM}}\) for each. Verify by direct averaging that \(\E[\hat\tau_{\mathrm{DIM}} \mid \mathcal{F}_n] = \bar\tau_n\) (Theorem 5.1).
- Compute the exact design variance from your enumeration and verify that it equals Equation 5.3.
7. Conservative Neyman variance. Continue with the schedule of the previous problem.
- For each of the 6 assignments, compute the Neyman variance estimator \(\hat V\) of Equation 5.4, and compute \(\E[\hat V \mid \mathcal{F}_n]\) by averaging. (Hint: under simple random sampling without replacement, the within-arm sample variance is unbiased for the finite-population variance.)
- Verify that \(\E[\hat V \mid \mathcal{F}_n] - \mathrm{Var}(\hat\tau_{\mathrm{DIM}} \mid \mathcal{F}_n) = S_\tau^2/n\), as Theorem 5.1 asserts.
- For what modification of the schedule — keeping the same \(Y_i(0)\) column — would the gap be exactly zero? Explain in terms of treatment-effect heterogeneity.
8. Lin’s estimator as arm-specific standardization. Consider the fully interacted regression of \(Y_i\) on \((1,\, T_i,\, X_i - \bar X,\, T_i(X_i - \bar X))\), where \(\bar X\) is the full-sample covariate mean.
- Show that this regression is equivalent to fitting separate OLS regressions of \(Y\) on \((1, X)\) within each treatment arm.
- Let \((\hat a_t, \hat b_t)\) denote the arm-\(t\) intercept and slope. Show that the OLS coefficient on \(T_i\) equals \([\hat a_1 + \hat b_1 \bar X] - [\hat a_0 + \hat b_0 \bar X] = n^{-1}\sum_i \{\hat\mu(1, X_i) - \hat\mu(0, X_i)\}\), i.e. the OR estimator Equation 5.10 with linear working models.
- Explain why the centering at \(\bar X\) is essential: what does the coefficient on \(T_i\) estimate if the covariate is not centered?
9. Exact unbiasedness versus design consistency. Under complete randomization, \(\hat\tau_{\mathrm{DIM}}\) is exactly unbiased for \(\bar\tau_n\) in every finite sample (Theorem 5.1), while the regression-adjusted estimator \(\hat\beta\) is only design-consistent (Theorem 5.3).
- Identify the step in the proof of Theorem 5.1 that delivers exact unbiasedness, and explain why no analogous exact argument is available for \(\hat\beta\). (Hint: which quantity in the linearization of \(\hat\beta\) is random but appears inside a nonlinear function of the assignment vector?)
- Locate where asymptotics enter the proof of Theorem 5.3: which remainder term is \(O_p(n^{-1})\), and why is asymptotic negligibility, rather than exact unbiasedness, the relevant conclusion for it?
- A colleague argues that since the regression-adjusted estimator can be biased in finite samples, the unadjusted difference in means should always be preferred. Evaluate this argument, referring to the order of the bias and the variance comparison in the remark on the role of the regression model in Section 5.3.5.