7 Instrumental Variables
7.1 Why Instrumental Variables?
Chapter 6 showed how causal effects can be identified when all confounders are observed and can be blocked by conditioning on \(X\). When some confounders are unobserved, back-door adjustment fails. This chapter develops instrumental variables (IV) as an alternative identification strategy: rather than blocking the confounding path \(T \leftarrow U \rightarrow Y\), IV exploits an external variable \(Z\) whose effect on \(T\) is free of confounding by \(U\). This chapter asks what IV identifies and under what assumptions; Chapter 13 asks how that estimand is computed and tested in practice.
7.1.1 The Endogeneity Problem
Chapter 6 established that the propensity score provides a powerful surrogate for randomization, but only under unconfoundedness: \((Y(0), Y(1)) \indep T \mid X\). This assumption requires that every variable affecting both treatment and outcome is observed and included in \(X\). In many empirical settings this is implausible:
- In labor economics, unobserved ability or motivation affects both schooling decisions and wages.
- In epidemiology, unobserved health behaviors affect both treatment uptake and outcomes.
- In survey sampling, voluntary participation is driven by factors that also affect the variables of interest.
Whenever an unobserved confounder \(U\) creates a back-door path \(T \leftarrow U \rightarrow Y\), the adjustment formula fails: \[\int f(y \mid t, x)\, p(x)\, dx \;\neq\; f\!\bigl(y \mid \doop(T{=}t)\bigr).\] The gap is the endogeneity bias. We need a different identification strategy.
7.1.2 The IV Idea
An instrumental variable \(Z\) is an observed variable that:
- Moves treatment \(T\) (relevance).
- Does so exogenously — \(Z\) is unrelated to the unobserved confounder \(U\) (exogeneity).
- Affects \(Y\) only through \(T\) — \(Z\) has no direct path to \(Y\) (exclusion).
Under these three conditions, the variation in \(T\) induced by \(Z\) is free of confounding by \(U\), so the ratio of the \(Z\)-induced variation in \(Y\) to the \(Z\)-induced variation in \(T\) becomes a meaningful target for identification. Which causal estimand it identifies — the ATE, the ATT, or the LATE for compliers — depends on a fourth structural restriction on treatment-effect heterogeneity, taken up in Section 7.3 and Section 7.7.
Instrumental variables do not eliminate the confounding path \(T \leftarrow U \rightarrow Y\), nor do they make treatment as-if randomly assigned for the full population. Instead, IV identifies causal effects by isolating a component of treatment variation that is induced by \(Z\) and is therefore exogenous under the IV assumptions. In that sense, IV avoids confounding rather than controlling for it — a distinction that matters both for interpreting what is identified (the LATE for compliers, not the ATE) and for understanding which assumption does the heaviest lifting (exclusion, not unconfoundedness).
7.2 Graphical Setup and Core Assumptions
This section is the conceptual foundation of the IV chapter. We fix the causal graph for the IV design and state the three core assumptions in the three causal languages used throughout these notes: DAGs and do-calculus, structural equations, and potential outcomes. The purpose is not merely stylistic. Each language highlights a different aspect of the design: the DAG makes the causal pathways visible, the do-calculus formalizes what those pathways imply about interventional distributions, the structural model supplies the moment conditions used in the Wald derivation, and the potential-outcomes formulation prepares the ground for the LATE framework in Section 7.7.
Before proceeding, it is important to separate two questions that are often conflated. The first is: when is an instrument valid? Validity requires the three core IV assumptions — relevance, exogeneity, and exclusion — which are the subject of this section. The second is: what does a valid instrument identify? That answer depends on a fourth structural assumption. Under linearity and homogeneous treatment effects (Section 7.3), the Wald ratio identifies a single parameter \(\beta\) equal to the ATE. Under heterogeneous effects with monotonicity (Section 7.7), the same Wald ratio identifies the LATE for compliers. The three IV assumptions determine whether the design is valid; the fourth assumption determines what causal quantity it recovers.
7.2.1 The IV DAG
The causal structure for a basic IV model with covariates \(X\) is:
The three IV assumptions correspond to three distinct features of this DAG:
- Relevance: there is a directed path \(Z \to T\) (edge labeled (1) above), so the instrument shifts treatment.
- Exogeneity: conditional on observed covariates \(X\), the instrument is free of open back-door paths to the latent determinants of the outcome — the DAG implies \(Z \indep U \mid X\) by d-separation. Intuitively, \(Z\) is as-if randomized with respect to the unobserved factors that confound \(T\) and \(Y\).
- Exclusion: every directed path from \(Z\) to \(Y\) in \(\Gcal\) passes through \(T\). In the simple IV DAG above this rules out the direct edge \(Z \to Y\); in richer graphs it also rules out indirect channels \(Z \to M \to Y\) that bypass \(T\) through any unmodeled mediator \(M\).
The distinction between exogeneity and exclusion is fundamental. Exogeneity says that the instrument is not confounded with the latent causes of the outcome. Exclusion says that the instrument has no causal channel to the outcome except through treatment. An instrument can satisfy one and fail the other. For example, a randomized encouragement may be exogenous by design, yet still violate exclusion if the encouragement changes outcomes through information, motivation, stigma, or administrative treatment apart from the treatment itself.
What the DAG rules out. The absence of a direct \(Z \to Y\) arrow encodes the exclusion restriction. Any such arrow — no matter how small the coefficient — would violate exclusion and invalidate the Wald estimand. The DAG forces the researcher to commit to this assumption explicitly; the potential outcomes framework expresses it as the equality \(Y_i(t,z) = Y_i(t,z')\).
7.2.2 The Three Assumptions in Three Languages
The same three assumptions can be expressed in three causal languages. These are parallel formulations, not literally identical statements: each language highlights a different aspect of the underlying design. The graphical formulation is most useful for causal design, because it makes confounding paths and forbidden arrows explicit. The structural formulation is most useful for deriving moment restrictions such as \(\E[Z\varepsilon \mid X] = 0\). The potential-outcomes formulation is most useful when we later introduce compliance types and the LATE theorem. The assumptions are ordered relevance → exogeneity → exclusion, reflecting the natural sequence in which a researcher assesses them.
| Assumption | Potential outcomes | Structural / econometric | Do-calculus / DAG |
|---|---|---|---|
| Relevance | \(\Pr\bigl(T_i(1) \neq T_i(0)\bigr) > 0\) | \(\pi \neq 0\) in \(T = \pi Z + \delta^\top X + \eta\) | \(Z \to T\) in \(\Gcal\) (no d-separation) |
| Exogeneity | \(Z \indep (Y(0), Y(1), T(0), T(1)) \mid X\) | \(\E[\varepsilon \mid Z, X] = 0\) | \(Z \indep U \mid X\) |
| Exclusion | \(Y_i(t, z) = Y_i(t, z')\) for all \(z, z'\) | \(Z\) absent from structural equation for \(Y\) | \(f(y \mid x, \doop(T{=}t), z) = f(y \mid x, \doop(T{=}t))\) |
Why the do-calculus column is preferred. The exclusion restriction in the do-calculus column reads \[f\!\bigl(y \mid x,\, \doop(T{=}t),\, z\bigr) \;=\; f\!\bigl(y \mid x,\, \doop(T{=}t)\bigr).\] This is a statement about the interventional density — the distribution of \(Y\) after we have set \(T = t\) by do-surgery. It says that once \(T\) is fixed by intervention, \(Z\) carries no additional information about \(Y\). The do-operator makes it impossible to confuse the exclusion restriction with the observational statement \(f(y \mid x, T{=}t, z) = f(y \mid x, T{=}t)\), which is a much weaker condition.
7.2.3 Relevance
The graphical analogue of relevance is that \(Z\) and \(T\) are not d-separated in \(\Gcal\). Under faithfulness, a directed path \(Z \to T\) implies an observable association between \(Z\) and \(T\); without faithfulness, non-d-separation does not by itself guarantee a non-zero conditional covariance, and an exact cancellation between alternative paths is mathematically possible. Conversely, an observable association between \(Z\) and \(T\) can arise without a causal \(Z \to T\) effect when \(Z\) and \(T\) share a common cause. Graphical non-d-separation, observable association, and a non-zero causal effect are therefore three related but not identical conditions. The conditional covariance \(\mathrm{Cov}(Z, T \mid X)\) and the first-stage \(F\)-statistic are empirical diagnostics for the observable association; the causal claim that \(Z\) shifts \(T\) is design-based.
7.2.4 Exogeneity
Exogeneity means that the instrument carries no information about latent determinants of the outcome other than through treatment. In the structural linear model, the same requirement appears as the moment orthogonality condition \(\E[Z\varepsilon \mid X] = 0\), because the structural residual \(\varepsilon\) is a function of \(U\). In potential-outcomes language, \(Z\) must be independent of the counterfactual quantities that the IV argument requires — outcomes \(Y(0), Y(1)\) and treatment indicators \(T(0), T(1)\) — given \(X\). These are related manifestations of the same identification requirement, but they live in different formal languages and are not literally the same statement.
Econometric implication. The moment condition \(\E[\varepsilon Z \mid X] = 0\) is a consequence of the graphical assumption, not a definition of exogeneity. This distinction matters: a researcher who adopts the moment condition as a primitive has no guarantee that \(Z\) is free of back-door paths to the outcome, only that a particular linear projection is zero.
Potential-outcomes analogue. In the potential-outcomes language, exogeneity is expressed as the full independence of the instrument from all counterfactual outcomes and counterfactual treatments: \[Z \indep \bigl(Y(0),\, Y(1),\, T(0),\, T(1)\bigr) \;\big|\; X.\] This is a stronger statement than the moment condition alone: it also covers the joint independence of \(Z\) with the treatment potential outcomes \(T(0), T(1)\), which the LATE proof in Section 7.7 requires in addition to \(\E[\varepsilon Z \mid X] = 0\).
Key point. All three versions of exogeneity are untestable: the unobserved confounder \(U\) is, by definition, unobserved, and counterfactual potential outcomes are never jointly observed. The assumption must be justified by institutional knowledge or the design of the study. See Section 7.6 for the limited consistency check that overidentification provides.
7.2.5 Exclusion
Exclusion is a causal restriction, not a conditional-correlation restriction. It is emphatically not the observational statement \(Y \indep Z \mid T, X\), which can hold in the data even when \(Z\) has a direct structural effect on \(Y\). What exclusion requires is that changing \(Z\) while holding treatment fixed — that is, intervening to set \(T = t\) — would leave \(Y\) unchanged. In graphical terms, every directed path from \(Z\) to \(Y\) must pass through \(T\). In the structural linear model, this means no \(Z\) term appears in the outcome equation for \(Y\).
The do-calculus makes the violation visible. Adding a direct arrow \(Z \to Y\) to the DAG is exactly the formal counterpart of exclusion failing: the arrow represents a structural effect of \(Z\) on \(Y\) that bypasses \(T\). In the mutilated graph \(\Gcal_{\overline{T}}\) (where we have removed all arrows into \(T\)), the path \(Z \to Y\) remains open. Rule 1 of the do-calculus, which would allow us to remove \(Z\) from the density of \(Y\) given \(\doop(T{=}t)\), no longer applies. The exclusion restriction is therefore visibly a statement about the structure of \(\Gcal_{\overline{T}}\), not of \(\Gcal\) itself.
7.3 Identification in the Linear Homogeneous-Effect Model
7.3.1 The Linear Structural Model
Consider the linear structural model: \[Y = \alpha + \beta T + \gamma^\top X + \varepsilon, \tag{7.1}\] \[T = \pi Z + \delta^\top X + \eta, \tag{7.2}\] where \(\varepsilon\) and \(\eta\) are structural errors with standard deviations \(\sigma_\varepsilon\) and \(\sigma_\eta\) and \(\mathrm{Cov}(\varepsilon, \eta) = \rho\sigma_\varepsilon\sigma_\eta \neq 0\). The non-zero covariance is the source of endogeneity: OLS applied to Equation 7.1 gives a biased estimator of \(\beta\).
The parameter of interest is \(\beta\), the causal effect of \(T\) on \(Y\): \[\beta = \E\!\bigl[Y \mid \doop(T{=}t+1)\bigr] - \E\!\bigl[Y \mid \doop(T{=}t)\bigr].\] Equation Equation 7.1 is the outcome equation (or second stage); Equation 7.2 is the first-stage equation. The reduced form is obtained by substituting Equation 7.2 into Equation 7.1: \[Y = \alpha + \beta\pi Z + (\beta\delta + \gamma)^\top X + (\beta\eta + \varepsilon).\] The reduced-form coefficient on \(Z\) is \(\beta\pi\): the total effect of the instrument on the outcome. Dividing by the first-stage coefficient \(\pi\) recovers \(\beta\), provided \(\pi \neq 0\).
7.3.2 The OLS Bias
To see why OLS fails, apply OLS to Equation 7.1. The probability limit of the OLS estimator is \[\mathrm{plim}\; \hat\beta_{\mathrm{OLS}} = \beta + \frac{\mathrm{Cov}(T, \varepsilon)}{\mathrm{Var}(T)} = \beta + \frac{\rho\sigma_\varepsilon\sigma_\eta}{\pi^2 \mathrm{Var}(Z) + \sigma_\eta^2}. \tag{7.3}\] The second term is the endogeneity bias. It is zero only if \(\rho = 0\) (no unobserved confounding) or if the instrument perfectly determines \(T\) (infeasible in practice). The direction of the bias depends on the sign of \(\rho\).
The point of IV is not that OLS is always wrong, but that when \(T\) is endogenous, IV can recover a causal parameter from the exogenous variation in \(T\) that \(Z\) induces — variation that OLS mixes indiscriminately with the confounded component.
7.3.3 Derivation of the Wald Estimand
The derivation has three ingredients: the first stage, which measures how much the instrument moves treatment; the reduced form, which measures how much the instrument moves the outcome; and the exclusion restriction, which implies that any effect of \(Z\) on \(Y\) must operate through \(T\). Under the homogeneous-effect linear model, these three ingredients imply that the reduced-form effect of \(Z\) on \(Y\) is exactly \(\beta\) times the first-stage effect of \(Z\) on \(T\).
We derive the identification of \(\beta\) step by step, using the structural model as the vehicle and the do-calculus to interpret the exclusion step.
Exogeneity (\(Z \indep U \mid X\) in \(\Gcal\)) implies, for the structural residual \(\varepsilon\), that \(\E[\varepsilon \mid Z, X] = \E[\varepsilon \mid X]\). With the normalization \(\E[\varepsilon \mid X] = 0\) (absorbed into \(\alpha\)), we obtain \(\E[\varepsilon \mid Z, X] = 0\).
Exclusion (\(Z\) absent from the structural equation for \(Y\)) is what permits writing Equation 7.1 without a \(Z\) term in the first place: under exclusion, the structural residual \(\varepsilon\) depends on \(U\) but does not absorb a direct effect of \(Z\) on \(Y\). Combined with exogeneity, \(\E[\varepsilon \mid Z, X] = 0\) immediately yields the moment condition \(\E[\varepsilon \cdot Z \mid X] = 0\). The graphical counterpart is in \(\Gcal_{\overline{T}}\): exclusion means there is no direct edge \(Z \to Y\), so \(Y \indep Z \mid T, X\) holds in the mutilated graph and Rule 1 of the do-calculus removes \(Z\) from \(f(y \mid x, \doop(T{=}t), z)\).
Multiplying Equation 7.1 through by \((Z - \E[Z \mid X])\) and taking expectations, the residual term \(\E[\varepsilon (Z - \E[Z\mid X])]\) vanishes by step 2, leaving \[\E\!\bigl[(Y - \alpha - \gamma^\top X)(Z - \E[Z\mid X])\bigr] = \beta\, \E\!\bigl[(T - \E[T\mid X])(Z - \E[Z\mid X])\bigr].\] In the simpler case with no covariates \(X\), this collapses to the reduced-form decomposition \(\mathrm{Cov}(Y, Z) = \beta\, \mathrm{Cov}(T, Z)\), which shows that the \(Z\)–\(Y\) covariance is entirely attributable to the causal path \(Z \to T \to Y\) (by exclusion), and equals \(\beta\) times the first-stage covariance.
Relevance (\(\mathrm{Cov}(Z, T \mid X) \neq 0\)) ensures the denominator is non-zero, so we can solve: \[\beta = \frac{\mathrm{Cov}(Y,\, Z \mid X)}{\mathrm{Cov}(T,\, Z \mid X)}. \tag{7.4}\]
The key lesson is that the Wald ratio identifies \(\beta\) in this section only because treatment effects are assumed to be constant across units. The observed formula Equation 7.4 is an identifying expression, not an estimand by itself. Its causal meaning depends on the structural assumptions maintained around it. In Section 7.7, the same ratio will survive, but its interpretation will change from a common causal effect to a complier-specific average effect.
7.4 Why the IV Assumptions Matter
The three IV assumptions are not merely technical conditions: each one is load-bearing, and each failure mode produces a distinct, quantifiable distortion of the Wald estimand. Violating exogeneity or exclusion contaminates the reduced form, and dividing by a weak first stage magnifies that contamination. A weak instrument is therefore not merely inefficient; it makes any small violation of the identifying assumptions more consequential.
When relevance fails. If \(\pi = 0\), the Wald estimand is undefined: its denominator is zero and there is no exogenous variation to exploit. When \(\pi\) is small but nonzero, the instrument is weak. The estimator’s variance diverges as \(\pi \to 0\) because the denominator amplifies noise, and finite-sample bias pulls the IV estimate toward the OLS estimate at a rate proportional to \(1/F\), where \(F\) is the first-stage \(F\)-statistic.
When exclusion fails. If the outcome equation contains a direct effect \(Y = \alpha + \beta T + \delta Z + \gamma^\top X + \varepsilon\) with \(\delta \neq 0\), the reduced form picks up both the indirect path \(\beta\pi\) and the direct path \(\delta\), so the Wald estimand converges to \(\beta + \delta/\pi\). The bias \(\delta/\pi\) is amplified by a weak first stage: a small direct effect combined with a weak instrument can produce large bias. This is why a weak instrument with a plausible exclusion violation is not “nearly valid” — it may be severely misleading.
When exogeneity fails. If \(\mathrm{Cov}(Z, \varepsilon \mid X) \neq 0\), the IV moment condition fails and the Wald estimand converges to \(\beta + \mathrm{Cov}(\varepsilon, Z)/\mathrm{Cov}(T, Z)\). This is the IV analogue of OLS omitted-variable bias: IV fails when the instrument itself is endogenous, just as OLS fails when the treatment is endogenous. Again the bias is amplified when the first stage is weak.
| Assumption | Directly testable? | Basis for assessment |
|---|---|---|
| Relevance | Association testable; causal claim design-based | First-stage \(F\)-statistic and \(\mathrm{Cov}(Z,T \mid X)\) test the observable association; the causal \(Z \to T\) link rests on the design |
| Exogeneity | No | Institutional knowledge; randomization (if available); placebo regressions on pre-determined outcomes |
| Exclusion | No in just-identified case; partially in overidentified case | Institutional argument; overidentification test (\(J\)-test, Chapter 13) checks mutual consistency of instruments but cannot confirm all are valid |
The key asymmetry is that the hardest work in defending an IV design — the institutional arguments for exogeneity and exclusion — is precisely the work that the data cannot do for the researcher.
7.5 Lab: OLS vs. IV Across Instrument Strengths
The analytic results of Section 7.3 and Section 7.4 established two facts: OLS is biased by a precise, computable amount whenever \(\rho \neq 0\); and IV is consistent but its variance diverges as \(\pi \to 0\). What the analytics do not immediately reveal is the finite-sample consequence: for weak instruments, the IV variance penalty is so severe that the biased OLS estimator can have lower mean squared error. This simulation studies estimator behavior conditional on the IV assumptions being true; it does not address the separate and fundamentally substantive question of whether a proposed instrument is valid in an applied study.
Estimators. OLS regresses \(Y\) on \(T\); from Equation 7.3 it converges to \(\beta + \rho/(\pi^2 + 1) = 1 + 0.8/(\pi^2+1)\). The bias is always positive (since \(\rho > 0\)), decreasing in \(|\pi|\), but never zero for finite \(\pi\). IV (Wald) uses \(Z\) as the instrument; it is consistent for \(\beta\) for any \(\pi \neq 0\), with finite-sample variance approximately \(\sigma^2/(n\pi^2)\), diverging as \(\pi \to 0\).
Results (\(n = 500\), \(B = 2{,}000\) replications, seed 2024; true \(\beta = 1.000\)):
| \(\pi\) | \(F\) | Theory bias | OLS mean | OLS bias | OLS SD | OLS RMSE | IV mean | IV bias | IV SD | IV RMSE |
|---|---|---|---|---|---|---|---|---|---|---|
| 0.00 | 1 | +0.800 | 1.801 | +0.801 | 0.026 | 0.802 | — | — | — | — |
| 0.10 | 6 | +0.792 | 1.791 | +0.791 | 0.027 | 0.792 | 0.872 | −0.128 | 4.391 | 4.393 |
| 0.15 | 13 | +0.782 | 1.782 | +0.782 | 0.027 | 0.783 | 0.764 | −0.236 | 6.487 | 6.491 |
| 0.20 | 21 | +0.769 | 1.769 | +0.769 | 0.027 | 0.770 | 0.940 | −0.060 | 0.380 | 0.385 |
| 0.30 | 46 | +0.734 | 1.735 | +0.735 | 0.028 | 0.735 | 0.975 | −0.025 | 0.163 | 0.165 |
| 0.50 | 127 | +0.640 | 1.639 | +0.639 | 0.028 | 0.640 | 0.996 | −0.004 | 0.091 | 0.091 |
| 1.00 | 500 | +0.400 | 1.401 | +0.401 | 0.027 | 0.402 | 0.999 | −0.001 | 0.045 | 0.045 |
| 2.00 | 2006 | +0.160 | 1.160 | +0.160 | 0.019 | 0.161 | 1.000 | 0.000 | 0.022 | 0.022 |
1. The OLS bias formula is exact. The theory-bias column matches the simulated OLS bias to within Monte Carlo error across all eight values of \(\pi\). OLS is biased in the direction of \(\rho\) at every value of \(\pi\), including \(\pi = 0\) where there is no instrument at all. Crucially, the OLS bias decreases as \(\pi\) grows, because a stronger first stage means more of \(T\)’s variance comes from the exogenous source \(Z\). At \(\pi = 2\) the bias is only \(0.160\), yet this still produces RMSE \(= 0.161\), which is \(7\times\) larger than the IV RMSE of \(0.022\) at the same \(\pi\).
2. IV is consistent but has catastrophically heavy tails when the instrument is weak. At \(\pi = 0.10\) (\(F \approx 6\)) and \(\pi = 0.15\) (\(F \approx 13\)) the IV mean is far from the true value, suggesting bias. This is misleading: the IV median is \(\approx 1.01\) at both values, confirming that the estimator is consistent and the typical replication recovers \(\beta = 1\). The mean is dragged off by a small fraction of replications in which the first stage happens to be near zero, making the Wald ratio explode. Indeed, the just-identified IV estimator possesses no finite moments, so the mean and SD entries at \(\pi = 0.10\) and \(0.15\) have no population counterparts: they are determined by the few most extreme replications and vary substantially across seeds. The take-away is that the weak instrument problem is not inconsistency; it is uncontrolled tail behavior.
3. The RMSE crossover occurs near \(F \approx 20\). OLS RMSE ranges from \(0.802\) down to \(0.161\), while IV RMSE starts infinite, is catastrophic at \(\pi = 0.10\)–\(0.15\), then drops sharply to \(0.385\) at \(\pi = 0.20\) (\(F \approx 21\)) and to \(0.165\) at \(\pi = 0.30\). IV first beats OLS on RMSE at \(\pi = 0.20\): this is the crossover. Below this threshold, the variance penalty of IV exceeds the bias penalty of OLS and the biased estimator is preferable in mean squared error terms. The Staiger–Stock rule of thumb (\(F \geq 10\)) is therefore slightly too lenient: at \(F \approx 13\) the IV RMSE in this run is still an order of magnitude above OLS. A more conservative threshold of \(F \geq 20\)–\(25\) is needed before IV dominates OLS in this DGP.
4. A strong instrument eliminates both problems. At \(\pi = 1.00\) (\(F \approx 500\)), IV is virtually unbiased with RMSE \(= 0.045\), while OLS still carries a bias of \(0.401\), giving RMSE \(= 0.402\) — a 9-fold improvement from IV. With a strong instrument, IV is simultaneously unbiased and more efficient than OLS in MSE terms, because OLS efficiency is illusory: its small variance is offset by a large, persistent bias.
7.6 Multiple Instruments and Overidentification
This section gives only the identification-level intuition for using more than one instrument. The computational details of 2SLS with multiple instruments, GMM weighting, and the Sargan–Hansen \(J\)-test are deferred to Chapter 13.
When there are exactly as many instruments as endogenous variables (\(q = p\)), the model is just-identified: the IV assumptions pin down the causal parameter exactly, leaving no surplus identifying variation. When \(q > p\), the model is overidentified: the extra instruments impose additional moment restrictions beyond those needed for identification. Under the homogeneous-effect linear model, every valid instrument must imply the same structural coefficient \(\beta\), so those extra restrictions are testable. Under heterogeneous treatment effects, valid instruments can legitimately identify different LATEs because they shift treatment for different complier populations; the overidentification test then targets equality of probability limits across instruments, and a rejection admits more interpretations than instrument invalidity. In both cases, the Sargan–Hansen \(J\)-test (Sargan 1958; Hansen 1982) formalizes the mutual-consistency check; rejection indicates that at least one of the instruments either violates exogeneity or exclusion or identifies a different LATE, but does not localize which.
The terminology distinguishes the order condition — a counting requirement that \(q \geq p\) — from the rank condition, which requires the instruments to be linearly independent in the first stage. Both are necessary for identification. The order condition is met by inspection; the rank condition is testable from the first-stage coefficient matrix.
7.7 Heterogeneous Treatment Effects and the LATE Framework
7.7.1 Compliance Types
In the binary-instrument, binary-treatment setting, each unit’s response to the instrument is fully described by the pair of potential treatment decisions \((T_i(0), T_i(1))\): the treatment the unit would take under each value of \(Z\). This pair defines the unit’s compliance type — a latent causal classification, because \((T_i(0), T_i(1))\) is never jointly observed in data.
The instrument \(Z\) only shifts treatment for compliers: always-takers and never-takers have the same treatment status regardless of \(Z\), so they contribute nothing to the denominator \(\E[T \mid Z{=}1] - \E[T \mid Z{=}0]\). This is the first indication that the Wald ratio will be driven by the complier subgroup rather than by the full population.
7.7.2 The Monotonicity Assumption
The three core IV assumptions alone do not yield a clean causal interpretation for the Wald ratio when effects are heterogeneous. The difficulty is that the instrument may push some units toward treatment and others away from it: in the presence of defiers, the numerator and denominator of the Wald ratio conflate effects in opposite directions. A fourth assumption rules out this ambiguity.
Interpretation. Monotonicity is not a generic law of causal inference. It is a design-specific claim about how this particular instrument changes treatment behavior. Switching the instrument from \(0\) to \(1\) may induce some units to take treatment (compliers) and leave others unaffected (always-takers or never-takers), but it should not reverse anyone’s treatment decision. This is most plausible when the instrument is a randomized encouragement, access rule, or administrative assignment mechanism. In many observational IV settings the no-defiers assumption is substantively harder to defend and requires explicit justification.
7.7.3 The LATE Theorem
Under the three core IV assumptions, heterogeneous treatment effects, and monotonicity, the Wald ratio no longer identifies the ATE. Instead it recovers the average treatment effect for the units whose treatment status is actually changed by the instrument — the compliers. The effect is local not because it is estimated with nearby observations, but because it pertains to a local margin of behavioral response defined by the instrument itself.
Theorem 7.1 (LATE Theorem (Angrist and Imbens 1994)) Suppose the following hold:
- Exogeneity (PO form): \(Z \indep \bigl(Y(0),\, Y(1),\, T(0),\, T(1)\bigr)\);
- Exclusion: \(Y_i(t, z) = Y_i(t)\) for all \(z\);
- Relevance: \(\E[T(1) - T(0)] \neq 0\);
- Monotonicity: \(T_i(1) \geq T_i(0)\) for all \(i\).
Then the Wald estimand identifies the Local Average Treatment Effect (LATE): \[\frac{\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0]}{\E[T \mid Z{=}1] - \E[T \mid Z{=}0]} \;=\; \E\!\bigl[Y(1) - Y(0) \;\big|\; T_i(1) > T_i(0)\bigr] \;\equiv\; \tau_{\mathrm{LATE}}. \tag{7.7}\] That is, the Wald estimand identifies the average causal effect for compliers only. When covariates are present, the same result holds with all assumptions and the Wald formula stated conditionally on \(X\).
Proof. We decompose the numerator and denominator by compliance type. Throughout, consistency gives \(T_i = T_i(Z_i)\) and \(Y_i = Y_i(T_i, Z_i)\), and exclusion reduces the latter to \(Y_i = Y_i(T_i(Z_i))\).
Denominator. By consistency for \(T\) and exogeneity (\(Z \indep (T(0), T(1))\)): \[\begin{aligned} \E[T \mid Z{=}1] - \E[T \mid Z{=}0] &= \E[T(1) \mid Z{=}1] - \E[T(0) \mid Z{=}0] \\ &= \E[T(1)] - \E[T(0)] = \E[T(1) - T(0)], \end{aligned}\] and monotonicity gives \(\E[T(1) - T(0)] = P(\text{complier})\), since always-takers contribute \(1 - 1 = 0\) and never-takers contribute \(0 - 0 = 0\).
Numerator. By consistency, exclusion, and exogeneity: \[\begin{aligned} \E[Y \mid Z{=}1] - \E[Y \mid Z{=}0] &= \E[Y(T(1))] - \E[Y(T(0))] \\ &= \sum_{c} P(c)\, \E\!\bigl[Y(T_c(1)) - Y(T_c(0)) \;\big|\; \text{type}=c\bigr]. \end{aligned}\] For always-takers, \(T(1) = T(0) = 1\), so the contribution is \(0\). For never-takers, \(T(1) = T(0) = 0\), contribution \(0\). For compliers, \(T(1) = 1\) and \(T(0) = 0\), so \(Y(T(1)) - Y(T(0)) = Y(1) - Y(0)\). No defiers exist by monotonicity. Therefore \[\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0] = P(\text{complier})\,\E\!\bigl[Y(1) - Y(0) \mid \text{complier}\bigr].\]
Ratio. Dividing numerator by denominator gives Equation 7.7. \(\square\)
7.8 Interpreting IV Estimands
The two frameworks — linear homogeneous effects and the LATE framework — give the Wald estimand different interpretations. This section synthesizes those interpretations and clarifies when they agree, when they disagree, and what the difference means for applied work.
7.8.1 What the Two Frameworks Say
| Framework 1 (linear, homogeneous) | Framework 2 (heterogeneous effects) | |
|---|---|---|
| Key assumption | \(\tau_i = \beta\) for all \(i\) | Monotonicity; no defiers |
| What IV identifies | \(\beta = \mathrm{ATE} = \mathrm{ATT} = \mathrm{LATE}\) | \(\tau_{\mathrm{LATE}} = \E[\tau_i \mid \text{complier}]\) |
| Estimand depends on instrument? | No (same \(\beta\) regardless of \(Z\)) | Yes (different \(Z\) \(\Longrightarrow\) different compliers \(\Longrightarrow\) different LATE) |
| Identifies the ATE? | Yes, automatically | Only if all units are compliers or effects homogeneous |
Framework 1 is a special case of Framework 2: when \(\tau_i = \beta\) for all \(i\), the LATE equals the ATE equals \(\beta\), and the instrument does not affect the estimand — only identification. Framework 2 is the more general and realistic setting. In applied work, the default interpretation of the Wald estimand is the LATE; the ATE interpretation requires the additional homogeneity argument of Framework 1.
7.8.2 When Does LATE Equal ATE?
LATE equals ATE only under additional structure, most notably treatment-effect homogeneity or special forms of heterogeneity that make complier effects representative of the full population. Neither condition should be assumed without argument.
In general, \(\tau_{\mathrm{LATE}} \neq \mathrm{ATE}\) when treatment effects are heterogeneous and some units are not compliers. To see this, decompose the ATE by compliance type: \[\mathrm{ATE} = P(\mathrm{co})\,\E[\tau_i \mid \mathrm{co}] + P(\mathrm{at})\,\E[\tau_i \mid \mathrm{at}] + P(\mathrm{nt})\,\E[\tau_i \mid \mathrm{nt}],\] where co, at, and nt abbreviate complier, always-taker, and never-taker. The LATE equals only the first term divided by its probability weight. The ATE and LATE coincide if and only if:
- Mean treatment effects are equal across compliance types — a strict weakening of the constant-effect assumption of Section 7.3, but still a non-trivial restriction on heterogeneity; or
- Everyone is a complier (\(P(\mathrm{at}) = P(\mathrm{nt}) = 0\)), which would require \(Z\) to perfectly determine \(T\); or
- The average effects for always-takers and never-takers happen to equal the LATE — an untestable coincidence.
In practice, none of these conditions is likely to hold exactly. The LATE is a well-defined, identifiable parameter, but it is not the ATE.
7.8.3 Different Instruments, Different Estimands
Because the LATE is specific to the complier population, and different instruments select different complier populations, two valid instruments for the same treatment can legitimately identify different LATEs. This is a precise, substantive fact — not a contradiction. Different instruments can legitimately identify different causal effects because they shift treatment for different margins of the population; the diversity of LATEs is informative about treatment effect heterogeneity, not a symptom of model failure.
The dependence of the LATE on the instrument is sometimes described as a limitation of IV. It is better understood as a precise statement about what question is being answered. A researcher using compulsory schooling laws is estimating the return to schooling for students at the compulsory leaving margin; a researcher using distance to college is estimating the return for students deterred by geography. These are different causal questions, and it is informative — not troubling — that they can yield different answers.
7.8.4 The Policy Relevance of LATE
For many policy questions, the LATE is exactly the right estimand. If a policy is designed to encourage a subset of the population to take treatment — for example, an outreach program that reaches only some potential participants — then the effect on compliers is precisely what the policy-maker wants to know. LATE is most policy-relevant when the contemplated intervention resembles the instrument, because then the complier population under the study design is close to the policy-relevant margin.
When the ATE over the full population is required for policy analysis, IV alone is insufficient under heterogeneous effects. Additional assumptions — such as an explicit model of treatment effect heterogeneity, or a second instrument that identifies effects for a different subpopulation — are needed to extrapolate from the LATE to the ATE.
7.9 IV versus Back-Door Adjustment
Having developed both strategies in full, it is instructive to compare them directly. The two strategies differ in what they require, what they identify, and how they fail.
7.9.1 A Direct Comparison
| Dimension | Back-door / propensity score | Instrumental variables |
|---|---|---|
| Core assumption | All confounders observed: \((Y(0),Y(1)) \indep T \mid X\) | Valid instrument: relevance, exogeneity, exclusion |
| Unobserved confounders | Fatal: back-door adjustment fails | Permitted: IV routes around \(U\) |
| Estimand | ATE, ATT, or ATC depending on design and overlap; all coincide under homogeneity | LATE (compliers only); reduces to a common \(\beta\) under homogeneous effects |
| Testability | Unconfoundedness untestable; overlap testable | Relevance testable; exogeneity and exclusion untestable (just-identified case) |
| Main threat | Unmeasured confounder | Exclusion restriction violation |
| Identifies the ATE? | Yes, under strong ignorability | Only under homogeneous effects |
| Typical setting | Rich administrative or survey data; RCT with imperfect compliance | Natural experiment; RCT with non-compliance; policy change |
7.9.2 Complementary Failure Modes
Back-door adjustment fails when \(X\) does not capture all confounders — that is, when an unobserved variable \(U\) opens a back-door path. The estimator is then inconsistent even with infinite data, because the path \(T \leftarrow U \rightarrow Y\) transmits spurious association that conditioning on \(X\) cannot close.
IV fails when the exclusion restriction is violated — that is, when \(Z\) has a direct effect on \(Y\) beyond its effect through \(T\). As derived in Section 7.4, the bias in the Wald estimand is \(\delta/\pi\), amplified by weak instruments.
The two failures are in a sense orthogonal: back-door adjustment requires many observed covariates but tolerates no unobserved ones, while IV tolerates unobserved confounders but requires an instrument with no direct effect on the outcome. Back-door adjustment fails when confounding remains after conditioning; IV fails when the proposed source of exogenous variation is not truly exogenous or does not act solely through treatment. When designing a study, the choice of identification strategy should be guided by which assumption is more plausible in the specific empirical context.
7.9.3 When Both Strategies Are Available
When a valid instrument and a sufficient adjustment set \(X\) are both available, comparing the two estimates can be informative, but disagreement between them does not by itself tell us which method is wrong. The two strategies typically target different estimands — back-door adjustment identifies the ATE or ATT over the full population or the treated, while IV identifies the LATE for compliers — and they rely on different identifying assumptions. Agreement is therefore reassuring under treatment-effect homogeneity but is neither required nor sufficient otherwise.
The Hausman (1978) endogeneity test operationalizes this comparison: under the null hypothesis that \(T\) is exogenous given \(X\), both the OLS and IV estimators are consistent, and a large discrepancy is evidence of endogeneity. Under the null and standard regularity conditions, an appropriately scaled quadratic form in the discrepancy \(\hat\beta_{\mathrm{IV}} - \hat\beta_{\mathrm{OLS}}\) is asymptotically \(\chi^2_p\)-distributed, where \(p\) is the number of components of \(T\) tested for endogeneity; the formal construction is given in Chapter 13. Rejection means at least one strategy is inconsistent; it does not identify which.
7.10 Practical Guidance on Defending an IV Design
There is no algorithm for finding a valid instrument; credible instruments arise from institutional knowledge and careful reasoning about the data-generating process. A strong IV design is defended primarily by institutional knowledge, design logic, and causal structure; statistical diagnostics are supportive but secondary.
A researcher proposing an instrument should be able to answer five questions explicitly:
What exactly is the instrument? Specify \(Z\) precisely: its source of variation, the level at which it varies, and the population to which it applies.
Why does it shift treatment? Articulate the causal mechanism by which \(Z\) moves \(T\). Relevance can be verified empirically with the first-stage \(F\)-statistic, but the \(F\)-statistic is a diagnostic for instrument strength, not a substitute for a causal account of the \(Z \to T\) link.
Why is it as-if random relative to latent outcome determinants? Argue why \(Z\) is unrelated to the unobserved causes of \(Y\). The most credible sources are designed randomization (lotteries, randomized encouragement), natural experiments with institutional quasi-randomness (policy discontinuities, geographic boundaries, biological quirks), and shift-share designs (Bartik 1991; Goldsmith-Pinkham et al. 2020). Placebo regressions on pre-determined outcomes provide partial — but not definitive — evidence.
Why can it affect the outcome only through treatment? The exclusion restriction is untestable in just-identified models, so it must rest on a structural argument that no direct path \(Z \to Y\) exists. A useful diagnostic is to ask: how large would the direct effect \(\delta\) have to be, relative to the first-stage coefficient \(\pi\), to overturn the estimated causal effect? When the first stage is weak, the answer is: not very large at all.
What population margin does it shift? Identify the complier population — the units whose treatment status changes with \(Z\). This determines the LATE that is being identified and governs the external validity of the estimates for other populations or policy margins.
7.11 Applied Example: Charter School Lotteries and the KIPP Lynn Study
This example illustrates a canonical randomized-encouragement IV design. The key distinction is that the lottery randomizes offer status, not actual treatment. Winning the lottery does not mechanically force a student to attend KIPP, and losing the lottery does not make later attendance impossible in all cases. Thus the lottery offer is the instrument \(Z\), while actual years of KIPP attendance is the endogenous treatment \(S\). Here \(S\) plays the role of the generic treatment variable \(T\) used throughout the rest of this chapter; the paper’s notation is retained to keep the empirical discussion close to the source.
This distinction is exactly what makes the design an IV design rather than a randomized controlled trial on treatment itself. The lottery generates exogenous variation in access to KIPP, and the IV analysis uses only the portion of attendance variation induced by that randomized offer to identify a causal effect.
7.11.1 Setting and Instrument
KIPP (Knowledge Is Power Program) schools follow a “No Excuses” model: extended school days, a longer academic year, selective teacher hiring, and strict behavioral norms. KIPP Academy Lynn was substantially oversubscribed beginning in 2005. Massachusetts law requires oversubscribed charter schools to select students by lottery, so the school conducted randomized admissions lotteries from 2005 through 2008. The treatment is years of KIPP attendance \(S\); in the panel structure of the data this is denoted \(s_{igt}\), the number of calendar years student \(i\) has spent at KIPP by the time of test \((g, t)\) — a continuous, endogenous variable, because families self-select into applying and attending. The outcome \(Y_{igt}\) is the student’s standardized score on the Massachusetts Comprehensive Assessment System (MCAS), normalized to mean zero and standard deviation one within each subject–grade–year cell statewide.
7.11.2 Mapping the Three Assumptions to the KIPP Context
The KIPP lottery design makes the logic of the three IV assumptions unusually transparent. One point deserves emphasis before proceeding. Relevance and exogeneity are especially natural in a lottery-based design: because lottery sequence numbers are randomly assigned within cohort, offer status is independent of the latent determinants of achievement by construction. Exclusion, however, still requires institutional argument. Random assignment of the instrument does not, by itself, imply that the instrument has no direct effect on the outcome — it only guarantees that the instrument is exogenous.
Relevance. The lottery offer must shift years of KIPP attendance. Lottery winners were offered a seat and about 80 percent accepted; losers rarely enrolled elsewhere at KIPP. The first-stage regression of \(s_{igt}\) on \(Z_i\) (with year and grade controls) yields a coefficient of approximately \(1.2\): at the time of each MCAS exam, lottery winners had spent about 1.2 more years at KIPP than lottery losers. The first-stage \(F\)-statistic is far above conventional thresholds. The first stage is less than the theoretical maximum because compliance is partial (some winners do not attend; some losers find entry through later cohorts or attrition slots) and the follow-up window captures students at different points in their KIPP tenure.
Exogeneity. The lottery offer must be independent of all potential outcomes and potential treatments — formally, \(Z \indep (Y(0), Y(1), T(0), T(1))\) within each application cohort. Because offer status was determined by randomly drawn lottery-sequence numbers, this independence holds by design — an especially strong basis for exogeneity compared to most observational IV applications. The authors verify it empirically: a joint test of covariate balance across lottery winners and losers yields a \(p\)-value of \(0.615\) for demographic characteristics and baseline test scores, consistent with the null of no pre-lottery differences. Pre-lottery variables are used for the balance check precisely because post-lottery variables such as LEP or SPED classification may themselves be affected by school attended.
Exclusion. The lottery offer must affect test scores only through KIPP attendance, not through any direct channel. Because the offer merely provides access to a school — it does not itself deliver instruction — this restriction is institutionally plausible. One potential violation is a discouragement effect: losing the lottery might demoralize students, suppressing their achievement regardless of where they enroll. The authors address this by noting that the scores of lottery losers are typical of demographically comparable students in Lynn, which is inconsistent with large discouragement effects.
7.11.3 First Stage, Reduced Form, and the 2SLS Estimand
The equations below illustrate the components of the Wald/2SLS logic — first stage, reduced form, and their ratio. Formal estimation theory for 2SLS, including asymptotic variance and cluster-robust inference, is deferred to Chapter 13.
The structural equation for test scores is \[y_{igt} = \alpha_t + \beta_g + \sum_j \delta_j d_{ij} + \gamma' X_i + \theta\, s_{igt} + \varepsilon_{igt}, \tag{7.8}\] where \(\alpha_t\) and \(\beta_g\) are year-of-test and grade-of-test fixed effects, \(d_{ij}\) are application-cohort dummies (cohort membership determines the probability of winning, so cohort is an essential control), \(X_i\) is a vector of baseline demographics, and \(\theta\) is the causal effect of interest per year at KIPP. We write \(\theta\) rather than \(\rho\) (as in the original paper) to avoid clashing with the endogeneity correlation parameter of the same name used earlier in this chapter. The first-stage equation is \[s_{igt} = \lambda_t + \kappa_g + \sum_j \mu_j d_{ij} + \Gamma' X_i + \pi Z_i + \eta_{igt}. \tag{7.9}\] The model is just-identified (one excluded instrument per endogenous variable), so the 2SLS estimator of \(\theta\) equals the ratio of the reduced-form coefficient on \(Z_i\) to the first-stage coefficient \(\pi\).
| Subject | First stage | Reduced form | 2SLS |
|---|---|---|---|
| Math | \(1.221\) \((0.068)\) | \(0.430\) \((0.067)\) | \(0.352\) \((0.053)\) |
| ELA | \(1.228\) \((0.068)\) | \(0.164\) \((0.073)\) | \(0.133\) \((0.059)\) |
Each year at KIPP raises math scores by approximately \(0.35\) standard deviations and ELA scores by approximately \(0.13\) standard deviations. The reduced-form estimate for math (\(0.43\sigma\)) is larger than the 2SLS estimate (\(0.35\sigma\)) because the first stage exceeds \(1\): lottery winners accumulated somewhat more than one additional year at KIPP per unit of follow-up time.
7.11.4 LATE Interpretation
The lottery design identifies the causal effect of KIPP attendance only for the students whose attendance behavior is changed by the lottery offer. These are the lottery compliers: students who attend KIPP if offered a seat and do not attend KIPP if not offered one. Compliance is partial in both directions: some lottery winners do not enroll and some lottery losers eventually find entry through later cohorts or attrition slots. Always-takers are largely ruled out at the extensive margin by the lottery’s control over seat access; never-takers remain — winners who decline the offer — and both margins reappear in the intensive dimension (years of enrollment).
This is the key interpretive point of the example. The 2SLS estimand \(\theta\) is not the effect of KIPP on all students in Lynn, nor on all applicants — it is the average per-year treatment effect for the lottery compliers, the specific margin of students whose enrollment decision is altered by randomized access to the school.
An important feature of this LATE is that it may differ from the effect the school would have on a randomly selected Lynn student who was not part of the applicant pool. KIPP applicants already had parents motivated enough to apply. The authors find that KIPP applicants have baseline test scores slightly lower than the district average, so the applicant pool is not positively selected on prior achievement, but it may still be selected on parental engagement in ways that affect how students respond to the KIPP environment.
The design is therefore strongest not because it identifies a universally generalizable effect, but because it identifies a clearly interpretable causal effect for a well-defined subpopulation under highly credible exogeneity.
7.11.5 Treatment Effect Heterogeneity
Subgroup analysis reveals that the LATE varies substantially across observable student characteristics. Reading gains (\(\approx 0.13\sigma\) overall) are driven almost entirely by students classified as having limited English proficiency (LEP, \(\approx 0.43\sigma\)) and special education needs (SPED, \(\approx 0.27\sigma\)); non-LEP, non-SPED students show negligible ELA gains. Math effects are large and positive across all subgroups but are largest for LEP and lower-achieving students. An interaction model that adds the product of baseline score with years at KIPP — identified by including \(Z_i \times \text{baseline score}\) as a second instrument — yields a significantly negative interaction term in both subjects: each additional standard deviation of baseline disadvantage is associated with an additional \(0.08\)–\(0.17\sigma\) gain per year at KIPP.
These findings connect directly to Section 7.8.3: were we to define separate subgroup-specific instruments (LEP-lottery and non-LEP-lottery), each would identify a distinct subpopulation LATE. The overall 2SLS estimate is a weighted average of these subgroup LATEs, with weights proportional to each subgroup’s share of the complier population. This heterogeneity is not a nuisance detail; it is exactly why IV estimates must be interpreted together with the margin of compliance they capture.
The KIPP example ties together the main lessons of this chapter. All three IV assumptions are visible in the design: relevance is confirmed by a strong, precisely estimated first stage; exogeneity within the applicant sample follows from the randomization protocol; and exclusion rests on an institutional argument that the lottery offer affects achievement only through attendance. Because compliance is partial, the resulting estimate is a LATE for lottery compliers, not an ATE for all applicants or all students in Lynn. A lottery-based IV design is strongest not because it answers every causal question, but because it answers a clearly defined causal question for a clearly defined subpopulation.
7.12 Summary
IV identifies effects from exogenous treatment variation. When back-door adjustment fails because an unobserved \(U\) creates a path \(T \leftarrow U \rightarrow Y\), a valid instrument \(Z\) identifies the causal effect by exploiting only the component of treatment variation that \(Z\) induces. IV does not block the confounding path — it avoids it by isolating exogenous variation and ignoring the rest.
Three assumptions, ordered by testability. Relevance can be assessed with the first-stage \(F\)-statistic, though the \(F\)-statistic is a sample diagnostic, not the assumption itself. Exogeneity and exclusion are primarily substantive and must be defended by institutional knowledge, design logic, and causal structure. Each violated assumption produces a distinct, quantifiable bias in the Wald estimand, amplified by weak instruments (Section 7.4).
Framework 1: homogeneous-effect SEM \(\Longrightarrow\) Wald identifies \(\beta\). Under constant treatment effects and the linear structural model, the Wald estimand identifies the single causal parameter \(\beta\), which equals the ATE, ATT, and LATE simultaneously. The identification follows from a three-step argument using the reduced form and first stage (Section 7.3).
Framework 2: heterogeneity + monotonicity \(\Longrightarrow\) Wald identifies LATE. Under heterogeneous treatment effects and monotonicity, the Wald estimand identifies the average treatment effect for compliers only — those whose treatment status changes with the instrument, a latent subgroup defined by potential treatment statuses \((T(0), T(1))\), not by observable characteristics (Section 7.7).
Different instruments identify different effects. The LATE depends on the instrument through the complier population it selects. Different instruments can legitimately identify different causal effects because they shift treatment for different margins of the population. LATE equals ATE only under additional structure, most notably treatment-effect homogeneity (Section 7.8).
IV versus back-door adjustment. The two strategies have complementary failure modes: back-door adjustment fails when confounders are unobserved; IV fails when the exclusion restriction is violated or the instrument is not truly exogenous. IV identifies the LATE (not the ATE) under heterogeneous effects; back-door adjustment identifies the ATE or ATT. They are tools for different identification problems, not competitors (Section 7.9).
Applied example: the KIPP Lynn lottery. Angrist et al. (2012) use a randomized admissions lottery as a canonical randomized-encouragement instrument for years of charter school attendance. Within the applicant sample, conditional on application cohort, random assignment of the offer supports exogeneity; exclusion is defended institutionally; relevance is confirmed by a first-stage coefficient of \(1.2\). The estimates identify the LATE for lottery compliers — \(0.35\sigma\) per year in math, \(0.13\sigma\) in ELA — with the largest gains for LEP, SPED, and low-baseline-score students (Section 7.11).
Estimation deferred to Chapter 13. This chapter establishes what IV identifies and under what assumptions. How the Wald ratio is estimated from finite data — the reduced form regression, two-stage least squares, asymptotic inference, and overidentification tests — is the subject of Chapter 13.
7.13 Problems
1. The three IV assumptions in three languages. Consider the DAG \(\{Z \to T,\; T \to Y,\; U \to T,\; U \to Y,\; X \to T,\; X \to Y,\; X \to Z\}\) with \(U\) unobserved.
- List all back-door paths from \(T\) to \(Y\). Does \(X\) alone satisfy the back-door criterion? Explain.
- Verify the three IV assumptions using d-separation: (i) relevance: show \(Z\) and \(T\) are not d-separated in \(\Gcal\); (ii) exogeneity: show \(Z \indep U \mid X\) in \(\Gcal\); (iii) exclusion: show \(Y \indep Z \mid T, X\) in \(\Gcal_{\overline{T}}\).
- Now add the arrow \(Z \to Y\) to the DAG. Which IV assumption is violated? Show explicitly which step of the Wald derivation in Section 7.3 breaks down.
- Translate each of the three IV assumptions into the structural language: write the equations for \(T\) and \(Y\) and identify which coefficient restriction corresponds to each assumption.
2. Bias under assumption violations. Let \(Y = \beta T + \varepsilon\) and \(T = \pi Z + \eta\) with \(\E[\varepsilon \mid Z] = 0\) and \(\pi \neq 0\).
- Starting from \(\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0]\), substitute the structural equation for \(Y\) and simplify. What role does exogeneity play?
- Show that \(\E[T \mid Z{=}1] - \E[T \mid Z{=}0] = \pi\) in the linear first-stage model. What role does relevance play?
- Derive the Wald estimand and confirm it equals \(\beta\).
- Now suppose the exclusion restriction fails and \(Y = \beta T + \delta Z + \varepsilon\) with \(\delta \neq 0\). Derive the probability limit of the Wald estimator. Confirm the bias formula from Section 7.4.
- Suppose instead that exogeneity fails: \(\E[\varepsilon \mid Z] = cZ\) for some constant \(c \neq 0\). Derive the probability limit of the Wald estimator and express the bias in terms of \(c\) and \(\pi\). Compare the structure of this bias with the exclusion violation bias.
3. Order, rank, and the limits of overidentification. Consider a model with one endogenous variable \(T\) and two instruments \(Z_1\) and \(Z_2\), both satisfying exogeneity and exclusion.
- State the order condition and verify it is satisfied.
- State the rank condition. What would it mean geometrically if the rank condition failed — i.e., if \(Z_1\) and \(Z_2\) were perfectly collinear in the first-stage regression?
- Explain intuitively why having two valid instruments rather than one should improve estimation precision.
- Now suppose \(Z_1\) is valid but \(Z_2\) violates the exclusion restriction. The Sargan–Hansen \(J\)-test is applied. Under what conditions does the \(J\)-test have power to detect \(Z_2\)’s invalidity? Under what conditions does the test fail to detect it?
- Why does passing the \(J\)-test not confirm that both \(Z_1\) and \(Z_2\) are valid? Give a concrete example of a situation in which both instruments are invalid and the \(J\)-test has no power.
4. Compliance types and the LATE. In a binary instrument, binary treatment study, suppose the population has the following composition: 30% compliers with average treatment effect \(\tau_c = 6\); 25% always-takers with \(\tau_a = 3\); 45% never-takers with \(\tau_n = 1\); no defiers.
- Compute \(P(\text{complier}) = \E[T \mid Z{=}1] - \E[T \mid Z{=}0]\).
- Compute the ATE as a weighted average of \(\tau_c\), \(\tau_a\), \(\tau_n\) with appropriate weights.
- The Wald estimand equals \(\tau_c = 6\). By how much does this overstate the ATE, and why?
- A second study uses a different binary instrument \(Z'\) with a complier population of 50% and a LATE of 2. Is this contradictory? What can you infer about the relative treatment effect in the two complier populations?
- Explain, using compliance type language, why the denominator of the Wald estimand equals \(P(\text{complier})\).
5. The exclusion restriction: plausibility and violations. Evaluate the exclusion restriction for each of the following proposed instruments. For each, (i) state whether the restriction is plausible and why; (ii) describe a specific mechanism by which it could be violated; and (iii) assess whether the violation would bias the IV estimate upward or downward.
- Instrument: rainfall in the home region of a politician, used as an instrument for government infrastructure spending. Outcome: local economic growth.
- Instrument: distance to the nearest hospital, used as an instrument for hospital admission. Outcome: 30-day mortality.
- Instrument: a randomly assigned financial incentive to enroll in a health screening program, used as an instrument for screening uptake. Outcome: health status two years later.
- Instrument: lottery number in the Vietnam-era draft lottery, used as an instrument for military service. Outcome: lifetime earnings. (This is the Angrist (1990) study; discuss why this instrument is widely regarded as satisfying the exclusion restriction.)
6. IV versus back-door adjustment. A researcher studies the effect of job training (\(T\)) on earnings (\(Y\)). Two strategies are available: (A) a rich set of pre-treatment covariates \(X\) and a propensity-score estimator; (B) a lottery that randomly selected units to be offered training (not required to attend), used as an instrument \(Z\).
- Under what assumption does strategy (A) identify the ATE? What specific unobserved variable would most plausibly violate this assumption?
- Strategy (B) identifies a LATE. Describe the complier population in words. Is the LATE likely to be larger or smaller than the ATE in this setting? Explain.
- Both strategies are implemented and yield estimates of $1,800 and $2,400 per year, respectively. Describe a Hausman-type test that uses both estimates. Under what null hypothesis does the test have an approximate \(\chi^2\) distribution?
- If the two estimates differ significantly, which strategy would you trust more and why? What additional evidence would help distinguish the two explanations (endogeneity bias in (A) versus LATE \(\neq\) ATE in (B))?
7. Nonparametric nonidentification of the ATE under valid IV assumptions (advanced / second pass). Consider the IV DAG \(Z \to T\), \(T \to Y\), \(U \to T\), \(U \to Y\), with \(U\) unobserved, \(Z \indep U\), and no direct \(Z \to Y\) edge, so that relevance, exogeneity, and exclusion all hold. This problem shows constructively that these assumptions do not nonparametrically identify the ATE, even with a genuinely relevant instrument, as the nonidentification remark opening Section 7.7 asserts.
- Construct two SEMs \(\mathcal{M}_1\) and \(\mathcal{M}_2\), each compatible with the IV graph, that are observationally indistinguishable but causally distinct: \[P_{\mathcal{M}_1}(Z, T, Y) = P_{\mathcal{M}_2}(Z, T, Y), \qquad P_{\mathcal{M}_1}\!\left(Y{=}1 \mid \doop(T{=}1)\right) \ne P_{\mathcal{M}_2}\!\left(Y{=}1 \mid \doop(T{=}1)\right).\] (Hint: let \(Z \sim \mathrm{Bern}(1/2)\) independently of a latent type \(U \in \{\mathrm{C}, \mathrm{A}, \mathrm{N}\}\) (complier, always-taker, never-taker) with \(P(\mathrm{C}) = 1/2\) and \(P(\mathrm{A}) = P(\mathrm{N}) = 1/4\), and in both models set \(T = Z\) for \(U = \mathrm{C}\), \(T = 1\) for \(U = \mathrm{A}\), and \(T = 0\) for \(U = \mathrm{N}\), so that \(P(T{=}1 \mid Z{=}1) - P(T{=}1 \mid Z{=}0) = 1/2\): the instrument is relevant. In both models set \(Y = T\) for \(U = \mathrm{C}\) and \(Y = 1\) for \(U = \mathrm{A}\); let the models differ only for \(U = \mathrm{N}\), with \(Y = 0\) in \(\mathcal{M}_1\) and \(Y = T\) in \(\mathcal{M}_2\).)
- Verify that both models generate the same joint \(P(Z, T, Y)\), then compute \(P(Y{=}1 \mid \doop(T{=}1))\) and \(P(Y{=}1 \mid \doop(T{=}0))\) in each model and show that the two ATEs differ.
- Verify that both models satisfy monotonicity (there are no defiers), compute the LATE \(\E[Y(1) - Y(0) \mid U = \mathrm{C}]\) in each model, and confirm that it coincides across the two models and equals the Wald ratio computed from the common observed distribution. What does this comparison show about which causal quantity the instrument identifies without further assumptions?