7  Instrumental Variables

NoteLearning Objectives

By the end of this chapter, students should be able to:

  1. Diagnose the endogeneity problem: explain why an unobserved confounder makes back-door adjustment fail, and describe how an instrumental variable provides an alternative identification route.
  2. State the three IV assumptions — relevance, exogeneity, and exclusion — in the DAG/do-calculus, structural-equation, and potential-outcomes languages, and explain which are testable and which require institutional justification.
  3. Follow how the IV assumptions identify the Wald estimand, and interpret the reduced form and first stage as its two observable components.
  4. Explain what the Wald estimand identifies when treatment effects are heterogeneous — the Local Average Treatment Effect for compliers — why that estimand depends on the choice of instrument, and when it coincides with the ATE.
  5. Compare IV and back-door adjustment on the assumptions each requires, the estimand each identifies, and the ways each can fail.

7.1 Why Instrumental Variables?

Chapter 6 showed how causal effects can be identified when all confounders are observed and can be blocked by conditioning on \(X\). When some confounders are unobserved, back-door adjustment fails. This chapter develops instrumental variables (IV) as an alternative identification strategy: rather than blocking the confounding path \(T \leftarrow U \rightarrow Y\), IV exploits an external variable \(Z\) whose effect on \(T\) is free of confounding by \(U\). This chapter asks what IV identifies and under what assumptions; Chapter 13 asks how that estimand is computed and tested in practice.

7.1.1 The Endogeneity Problem

Chapter 6 established that the propensity score provides a powerful surrogate for randomization, but only under unconfoundedness: \((Y(0), Y(1)) \indep T \mid X\). This assumption requires that every variable affecting both treatment and outcome is observed and included in \(X\). In many empirical settings this is implausible:

  • In labor economics, unobserved ability or motivation affects both schooling decisions and wages.
  • In epidemiology, unobserved health behaviors affect both treatment uptake and outcomes.
  • In survey sampling, voluntary participation is driven by factors that also affect the variables of interest.

Whenever an unobserved confounder \(U\) creates a back-door path \(T \leftarrow U \rightarrow Y\), the adjustment formula fails: \[\int f(y \mid t, x)\, p(x)\, dx \;\neq\; f\!\bigl(y \mid \doop(T{=}t)\bigr).\] The gap is the endogeneity bias. We need a different identification strategy.

7.1.2 The IV Idea

An instrumental variable \(Z\) is an observed variable that:

  1. Moves treatment \(T\) (relevance).
  2. Does so exogenously\(Z\) is unrelated to the unobserved confounder \(U\) (exogeneity).
  3. Affects \(Y\) only through \(T\)\(Z\) has no direct path to \(Y\) (exclusion).

Under these three conditions, the variation in \(T\) induced by \(Z\) is free of confounding by \(U\), so the ratio of the \(Z\)-induced variation in \(Y\) to the \(Z\)-induced variation in \(T\) becomes a meaningful target for identification. Which causal estimand it identifies — the ATE, the ATT, or the LATE for compliers — depends on a fourth structural restriction on treatment-effect heterogeneity, taken up in Section 7.3 and Section 7.7.

Instrumental variables do not eliminate the confounding path \(T \leftarrow U \rightarrow Y\), nor do they make treatment as-if randomly assigned for the full population. Instead, IV identifies causal effects by isolating a component of treatment variation that is induced by \(Z\) and is therefore exogenous under the IV assumptions. In that sense, IV avoids confounding rather than controlling for it — a distinction that matters both for interpreting what is identified (the LATE for compliers, not the ATE) and for understanding which assumption does the heaviest lifting (exclusion, not unconfoundedness).

NoteExample: Returns to Schooling as Motivation

A researcher wants to estimate the causal effect of years of schooling \(T\) on log wages \(Y\). Unobserved ability \(U\) raises both schooling and wages, so \(\mathrm{Cov}(T, \varepsilon) > 0\) and OLS is biased upward. No set of observed covariates \(X\) fully captures ability. An instrumental variable \(Z\) that shifts schooling for reasons unrelated to ability — such as proximity to a school or a policy change in compulsory attendance laws — provides a way to isolate the causal effect of schooling on wages.

T Y U X Back-door approach condition on X to block T ← U → Y Z T Y U IV IV approach use only Z-induced variation in T
Two identification strategies for a confounded treatment. Dashed red arrows denote paths through the unobserved confounder $U$. In the back-door approach, conditioning on the observed $X$ (green) blocks the confounding path. In the IV approach, $Z$ is used to select only the part of $T$'s variation that is free of $U$.

7.2 Graphical Setup and Core Assumptions

This section is the conceptual foundation of the IV chapter. We fix the causal graph for the IV design and state the three core assumptions in the three causal languages used throughout these notes: DAGs and do-calculus, structural equations, and potential outcomes. The purpose is not merely stylistic. Each language highlights a different aspect of the design: the DAG makes the causal pathways visible, the do-calculus formalizes what those pathways imply about interventional distributions, the structural model supplies the moment conditions used in the Wald derivation, and the potential-outcomes formulation prepares the ground for the LATE framework in Section 7.7.

Before proceeding, it is important to separate two questions that are often conflated. The first is: when is an instrument valid? Validity requires the three core IV assumptions — relevance, exogeneity, and exclusion — which are the subject of this section. The second is: what does a valid instrument identify? That answer depends on a fourth structural assumption. Under linearity and homogeneous treatment effects (Section 7.3), the Wald ratio identifies a single parameter \(\beta\) equal to the ATE. Under heterogeneous effects with monotonicity (Section 7.7), the same Wald ratio identifies the LATE for compliers. The three IV assumptions determine whether the design is valid; the fourth assumption determines what causal quantity it recovers.

7.2.1 The IV DAG

The causal structure for a basic IV model with covariates \(X\) is:

Z T Y X U (1)

The three IV assumptions correspond to three distinct features of this DAG:

  1. Relevance: there is a directed path \(Z \to T\) (edge labeled (1) above), so the instrument shifts treatment.
  2. Exogeneity: conditional on observed covariates \(X\), the instrument is free of open back-door paths to the latent determinants of the outcome — the DAG implies \(Z \indep U \mid X\) by d-separation. Intuitively, \(Z\) is as-if randomized with respect to the unobserved factors that confound \(T\) and \(Y\).
  3. Exclusion: every directed path from \(Z\) to \(Y\) in \(\Gcal\) passes through \(T\). In the simple IV DAG above this rules out the direct edge \(Z \to Y\); in richer graphs it also rules out indirect channels \(Z \to M \to Y\) that bypass \(T\) through any unmodeled mediator \(M\).

The distinction between exogeneity and exclusion is fundamental. Exogeneity says that the instrument is not confounded with the latent causes of the outcome. Exclusion says that the instrument has no causal channel to the outcome except through treatment. An instrument can satisfy one and fail the other. For example, a randomized encouragement may be exogenous by design, yet still violate exclusion if the encouragement changes outcomes through information, motivation, stigma, or administrative treatment apart from the treatment itself.

Z T Y U violation Exogeneity failure Z ← U → Y: shared common cause Z T Y U violation Exclusion failure Z → Y directly, bypassing T
The two main IV assumption failures. Blue arrows represent a valid IV design structure; the thick red arrow in each panel is the additional path that constitutes the violation. The two failures are logically independent: a randomized instrument can fail exclusion, and a non-randomized instrument can satisfy it.

What the DAG rules out. The absence of a direct \(Z \to Y\) arrow encodes the exclusion restriction. Any such arrow — no matter how small the coefficient — would violate exclusion and invalidate the Wald estimand. The DAG forces the researcher to commit to this assumption explicitly; the potential outcomes framework expresses it as the equality \(Y_i(t,z) = Y_i(t,z')\).

7.2.2 The Three Assumptions in Three Languages

The same three assumptions can be expressed in three causal languages. These are parallel formulations, not literally identical statements: each language highlights a different aspect of the underlying design. The graphical formulation is most useful for causal design, because it makes confounding paths and forbidden arrows explicit. The structural formulation is most useful for deriving moment restrictions such as \(\E[Z\varepsilon \mid X] = 0\). The potential-outcomes formulation is most useful when we later introduce compliance types and the LATE theorem. The assumptions are ordered relevance → exogeneity → exclusion, reflecting the natural sequence in which a researcher assesses them.

Assumption Potential outcomes Structural / econometric Do-calculus / DAG
Relevance \(\Pr\bigl(T_i(1) \neq T_i(0)\bigr) > 0\) \(\pi \neq 0\) in \(T = \pi Z + \delta^\top X + \eta\) \(Z \to T\) in \(\Gcal\) (no d-separation)
Exogeneity \(Z \indep (Y(0), Y(1), T(0), T(1)) \mid X\) \(\E[\varepsilon \mid Z, X] = 0\) \(Z \indep U \mid X\)
Exclusion \(Y_i(t, z) = Y_i(t, z')\) for all \(z, z'\) \(Z\) absent from structural equation for \(Y\) \(f(y \mid x, \doop(T{=}t), z) = f(y \mid x, \doop(T{=}t))\)
NoteRemark: Three Languages, One Identification Argument

These three formulations are aligned but not literally identical. The graphical statement encodes causal structure — which arrows are present or absent in \(\Gcal\). The potential-outcomes statement encodes counterfactual independence — which potential outcomes are independent of \(Z\). The structural formulation encodes moment orthogonality — which inner product with the residual is zero. In a well-specified IV model they support the same identification argument, but students should not treat them as interchangeable symbols.

Substantively: relevance is about the \(Z \to T\) link — \(Z\) must genuinely move \(T\), not merely be correlated with it. Exogeneity is about the absence of omitted common causes linking \(Z\) to \(Y\)\(Z\) must be as-if randomly assigned relative to the latent determinants of the outcome. Exclusion is about the absence of any direct causal path \(Z \to Y\) that bypasses \(T\) — it rules out both a direct arrow in the DAG and any parallel channel through unmodeled variables.

Why the do-calculus column is preferred. The exclusion restriction in the do-calculus column reads \[f\!\bigl(y \mid x,\, \doop(T{=}t),\, z\bigr) \;=\; f\!\bigl(y \mid x,\, \doop(T{=}t)\bigr).\] This is a statement about the interventional density — the distribution of \(Y\) after we have set \(T = t\) by do-surgery. It says that once \(T\) is fixed by intervention, \(Z\) carries no additional information about \(Y\). The do-operator makes it impossible to confuse the exclusion restriction with the observational statement \(f(y \mid x, T{=}t, z) = f(y \mid x, T{=}t)\), which is a much weaker condition.

7.2.3 Relevance

NoteDefinition: Relevance

The instrument \(Z\) is relevant if it has a non-zero causal effect on the treatment \(T\) within at least some stratum of covariates \(X\). Formally, \(P\bigl(T \mid \doop(Z{=}z),\, X{=}x\bigr)\) varies with \(z\) for some \(x\) in the support; in the linear first-stage model \(T = \pi Z + \delta^\top X + \eta\), this reduces to \(\pi \neq 0\).

The graphical analogue of relevance is that \(Z\) and \(T\) are not d-separated in \(\Gcal\). Under faithfulness, a directed path \(Z \to T\) implies an observable association between \(Z\) and \(T\); without faithfulness, non-d-separation does not by itself guarantee a non-zero conditional covariance, and an exact cancellation between alternative paths is mathematically possible. Conversely, an observable association between \(Z\) and \(T\) can arise without a causal \(Z \to T\) effect when \(Z\) and \(T\) share a common cause. Graphical non-d-separation, observable association, and a non-zero causal effect are therefore three related but not identical conditions. The conditional covariance \(\mathrm{Cov}(Z, T \mid X)\) and the first-stage \(F\)-statistic are empirical diagnostics for the observable association; the causal claim that \(Z\) shifts \(T\) is design-based.

7.2.4 Exogeneity

NoteDefinition: Exogeneity (Graphical Assumption)

The instrument \(Z\) is exogenous given covariates \(X\) if there is no open back-door path from \(Z\) to \(Y\) through unobserved causes. In DAG terms: \(Z \indep U \mid X\) in \(\Gcal\) by d-separation.

Exogeneity means that the instrument carries no information about latent determinants of the outcome other than through treatment. In the structural linear model, the same requirement appears as the moment orthogonality condition \(\E[Z\varepsilon \mid X] = 0\), because the structural residual \(\varepsilon\) is a function of \(U\). In potential-outcomes language, \(Z\) must be independent of the counterfactual quantities that the IV argument requires — outcomes \(Y(0), Y(1)\) and treatment indicators \(T(0), T(1)\) — given \(X\). These are related manifestations of the same identification requirement, but they live in different formal languages and are not literally the same statement.

Econometric implication. The moment condition \(\E[\varepsilon Z \mid X] = 0\) is a consequence of the graphical assumption, not a definition of exogeneity. This distinction matters: a researcher who adopts the moment condition as a primitive has no guarantee that \(Z\) is free of back-door paths to the outcome, only that a particular linear projection is zero.

Potential-outcomes analogue. In the potential-outcomes language, exogeneity is expressed as the full independence of the instrument from all counterfactual outcomes and counterfactual treatments: \[Z \indep \bigl(Y(0),\, Y(1),\, T(0),\, T(1)\bigr) \;\big|\; X.\] This is a stronger statement than the moment condition alone: it also covers the joint independence of \(Z\) with the treatment potential outcomes \(T(0), T(1)\), which the LATE proof in Section 7.7 requires in addition to \(\E[\varepsilon Z \mid X] = 0\).

NoteExample: Quarter of Birth as an Instrument

Angrist and Krueger (1991) used quarter of birth as an instrument for schooling to estimate the return to education. In graphical terms, exogeneity requires no open back-door path from quarter of birth to wages through unobserved variables such as ability or family background; in potential-outcomes terms, it requires that birth timing is independent of what wages a student would earn under any schooling level. Both are plausible because birth timing is largely outside parental control. Bound et al. (1995) later raised concerns about weak first stages in some specifications, illustrating that even a well-motivated instrument can fail the relevance requirement in certain samples.

Key point. All three versions of exogeneity are untestable: the unobserved confounder \(U\) is, by definition, unobserved, and counterfactual potential outcomes are never jointly observed. The assumption must be justified by institutional knowledge or the design of the study. See Section 7.6 for the limited consistency check that overidentification provides.

7.2.5 Exclusion

NoteDefinition: Exclusion Restriction

The instrument \(Z\) satisfies the exclusion restriction if it affects the outcome \(Y\) only through its effect on the treatment \(T\). In do-calculus terms: \[f\!\bigl(y \mid x,\, \doop(T{=}t),\, z\bigr) \;=\; f\!\bigl(y \mid x,\, \doop(T{=}t)\bigr) \quad \text{for all } z.\] In potential outcomes terms: \(Y_i(t, z) = Y_i(t, z')\) for all \(z \neq z'\).

Exclusion is a causal restriction, not a conditional-correlation restriction. It is emphatically not the observational statement \(Y \indep Z \mid T, X\), which can hold in the data even when \(Z\) has a direct structural effect on \(Y\). What exclusion requires is that changing \(Z\) while holding treatment fixed — that is, intervening to set \(T = t\) — would leave \(Y\) unchanged. In graphical terms, every directed path from \(Z\) to \(Y\) must pass through \(T\). In the structural linear model, this means no \(Z\) term appears in the outcome equation for \(Y\).

WarningWhy Exclusion Is the Hardest Assumption

Relevance can be tested. Exogeneity can sometimes be defended by design (e.g., randomized encouragement). But the exclusion restriction — that \(Z\) has no direct effect on \(Y\) whatsoever — is:

  • Untestable in just-identified models (one instrument, one treatment).
  • Easy to violate in practice. An instrument that “moves treatment” often also moves other inputs. For example: if \(Z\) is distance to a hospital (instrument for treatment uptake), it may also directly affect health outcomes through travel time in emergencies.
  • Consequential. Even a small direct effect of \(Z\) on \(Y\) — if correlated with the instrument — can cause substantial bias in the Wald estimand, especially when the first stage is weak. This bias is derived explicitly in Section 7.4.

The do-calculus makes the violation visible. Adding a direct arrow \(Z \to Y\) to the DAG is exactly the formal counterpart of exclusion failing: the arrow represents a structural effect of \(Z\) on \(Y\) that bypasses \(T\). In the mutilated graph \(\Gcal_{\overline{T}}\) (where we have removed all arrows into \(T\)), the path \(Z \to Y\) remains open. Rule 1 of the do-calculus, which would allow us to remove \(Z\) from the density of \(Y\) given \(\doop(T{=}t)\), no longer applies. The exclusion restriction is therefore visibly a statement about the structure of \(\Gcal_{\overline{T}}\), not of \(\Gcal\) itself.

NoteRemark: Three Assumptions Are Necessary but Not Sufficient for Point Identification

The three IV assumptions are necessary but not sufficient for point identification of any causal estimand. Under these three assumptions alone, the observed data restrict the distribution of potential outcomes but do not pin it down to a single value; the Wald ratio is a well-defined function of observable quantities, but it does not equal any specific causal parameter without additional structure (Hernán and Robins 2006; Levis et al. 2024). A fourth structural assumption on the counterfactual distribution of treatment response or treatment effect heterogeneity is required.

The choice of fourth assumption determines which causal quantity the Wald estimand identifies. Two leading choices appear in this chapter: constant treatment effects (Framework 1, Section 7.3), under which the Wald estimand identifies the ATE; and monotonicity (Framework 2, Section 7.7), under which it identifies the LATE for compliers. A third important family relies on structural mean model (SMM) restrictions. Under an additive SMM that rules out direct effect modification of \(Y\) by \(Z\) among the treated, the Wald ratio identifies the ATT — the average effect of treatment among those who actually received it — without requiring either constant effects or monotonicity (Hernán and Robins 2006; Levis et al. 2024). Moving further to identify the population ATE from the ATT requires an additional step: that the average effect among the treated equals the average effect among the untreated.

Strikingly, across these different identification schemes the same conditional Wald formula \(\mathrm{Cov}(Z, Y \mid X)/\mathrm{Cov}(Z, T \mid X)\) appears as the identifying expression in many cases. The qualification “many” is precise: when identification relies on a multiplicative SMM instead of an additive one, the identifying formula for the ATE takes a different form entirely and the Wald ratio is no longer consistent for the target estimand (Hernán and Robins 2006). What changes in the additive cases is not the observed-data formula but its causal target: the same estimator can consistently estimate the ATE, the LATE, or the ATT depending on which structural assumption is maintained.

A further consequence deserves emphasis: the three core IV assumptions are all expressible within the graphical language of this course. Relevance is the directed edge \(Z \to T\); exogeneity is the absence of an open back-door path from \(Z\) to \(Y\) via \(U\); exclusion is the absence of a direct edge \(Z \to Y\) in the mutilated graph \(\Gcal_{\overline{T}}\). The fourth assumption, by contrast, lies outside the graphical language in every case. Monotonicity (\(T_i(1) \geq T_i(0)\) for all \(i\)) is a restriction on the joint distribution of potential treatments across two intervention values simultaneously — a cross-world statement that cannot be encoded by the presence or absence of any arrow in the DAG. Constant treatment effects similarly constrains the covariance structure of the potential outcome pair, which the DAG does not represent. This is a genuine scope limitation of the graphical approach: a DAG specifies the causal structure of a single world, but the fourth assumption always reaches across worlds.

NoteExample: Same Observables, Different ATEs

Two distinct data-generating processes satisfy all three IV assumptions and produce identical observable statistics, yet imply different ATEs. No amount of data can distinguish them without a fourth structural assumption.

Observable statistics. Let \(Z, T \in \{0,1\}\) and let \(Y\) be a continuous outcome. Suppose \[\E[T \mid Z{=}1] - \E[T \mid Z{=}0] = 0.2 \quad \text{(first stage)},\] \[\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0] = 0.1 \quad \text{(reduced form)},\] giving a Wald ratio of \(0.1/0.2 = 0.5\). Both DGPs share the same compliance structure — under monotonicity, \(P(\mathrm{co}) = 0.2\), \(P(\mathrm{at}) = 0.4\), \(P(\mathrm{nt}) = 0.4\) — though under DGP A the compliance types are simply a convenient labeling; what drives identification there is the constant-effect restriction.

DGP A: constant treatment effects. Every unit has the same effect \(\tau_i = 0.5\). Then the reduced form is \(P(\mathrm{co}) \cdot \tau_{\mathrm{co}} = 0.2 \times 0.5 = 0.1\) \(\checkmark\), and \(\mathrm{ATE} = 0.5\).

DGP B: heterogeneous effects with monotonicity. Effects differ by type: \(\tau_{\mathrm{co}} = 0.5\), \(\tau_{\mathrm{at}} = 0\), \(\tau_{\mathrm{nt}} = 0\). The reduced form is again \(0.2 \times 0.5 = 0.1\) \(\checkmark\), but \[\mathrm{ATE} = 0.2 \times 0.5 + 0.4 \times 0 + 0.4 \times 0 = 0.1.\]

What this shows. Both DGPs produce the same first stage, the same reduced form, and therefore the same Wald ratio of \(0.5\). Yet the ATE is \(0.5\) under DGP A and \(0.1\) under DGP B — a fivefold difference. The fourth assumption is what resolves the ambiguity: under DGP A, constant effects makes the Wald ratio equal the ATE; under DGP B, monotonicity makes it equal the LATE (\(=0.5\)) for compliers, a different quantity from the ATE (\(=0.1\)). Committing to one fourth assumption or the other changes not the estimator but what it estimates.

7.3 Identification in the Linear Homogeneous-Effect Model

NoteFramework 1: Constant Treatment Effect and Linear Structure

This section works entirely within a linear structural model in which the treatment effect is the same for every unit and every treatment contrast: \[Y_i(t) - Y_i(t') = \beta(t - t') \quad \text{for all } i,\, t,\, t'.\] Under this assumption, IV identifies the single parameter \(\beta\), which equals the ATE, the ATT, and the LATE simultaneously. This section should therefore be read as a special-case identification result. In Section 7.7 we drop homogeneity, allow treatment effects to vary across units, and reinterpret the same Wald ratio as the LATE for compliers. The estimator does not change; what changes is the estimand.

Note on assumption strength. Constant treatment effects is the strongest form of the homogeneity family of fourth assumptions: it requires not only that unobserved confounders \(U\) do not modify the effect of \(T\) on \(Y\), but that the effect is exactly the same for every unit within each stratum of \(X\). Weaker homogeneity conditions — such as requiring only that \(U\) not act as an additive effect modifier on \(Y\) — can also identify the ATE under IV (Wang and Tchetgen Tchetgen 2018).

7.3.1 The Linear Structural Model

Consider the linear structural model: \[Y = \alpha + \beta T + \gamma^\top X + \varepsilon, \tag{7.1}\] \[T = \pi Z + \delta^\top X + \eta, \tag{7.2}\] where \(\varepsilon\) and \(\eta\) are structural errors with standard deviations \(\sigma_\varepsilon\) and \(\sigma_\eta\) and \(\mathrm{Cov}(\varepsilon, \eta) = \rho\sigma_\varepsilon\sigma_\eta \neq 0\). The non-zero covariance is the source of endogeneity: OLS applied to Equation 7.1 gives a biased estimator of \(\beta\).

The parameter of interest is \(\beta\), the causal effect of \(T\) on \(Y\): \[\beta = \E\!\bigl[Y \mid \doop(T{=}t+1)\bigr] - \E\!\bigl[Y \mid \doop(T{=}t)\bigr].\] Equation Equation 7.1 is the outcome equation (or second stage); Equation 7.2 is the first-stage equation. The reduced form is obtained by substituting Equation 7.2 into Equation 7.1: \[Y = \alpha + \beta\pi Z + (\beta\delta + \gamma)^\top X + (\beta\eta + \varepsilon).\] The reduced-form coefficient on \(Z\) is \(\beta\pi\): the total effect of the instrument on the outcome. Dividing by the first-stage coefficient \(\pi\) recovers \(\beta\), provided \(\pi \neq 0\).

7.3.2 The OLS Bias

To see why OLS fails, apply OLS to Equation 7.1. The probability limit of the OLS estimator is \[\mathrm{plim}\; \hat\beta_{\mathrm{OLS}} = \beta + \frac{\mathrm{Cov}(T, \varepsilon)}{\mathrm{Var}(T)} = \beta + \frac{\rho\sigma_\varepsilon\sigma_\eta}{\pi^2 \mathrm{Var}(Z) + \sigma_\eta^2}. \tag{7.3}\] The second term is the endogeneity bias. It is zero only if \(\rho = 0\) (no unobserved confounding) or if the instrument perfectly determines \(T\) (infeasible in practice). The direction of the bias depends on the sign of \(\rho\).

The point of IV is not that OLS is always wrong, but that when \(T\) is endogenous, IV can recover a causal parameter from the exogenous variation in \(T\) that \(Z\) induces — variation that OLS mixes indiscriminately with the confounded component.

7.3.3 Derivation of the Wald Estimand

The derivation has three ingredients: the first stage, which measures how much the instrument moves treatment; the reduced form, which measures how much the instrument moves the outcome; and the exclusion restriction, which implies that any effect of \(Z\) on \(Y\) must operate through \(T\). Under the homogeneous-effect linear model, these three ingredients imply that the reduced-form effect of \(Z\) on \(Y\) is exactly \(\beta\) times the first-stage effect of \(Z\) on \(T\).

WarningAssumptions Active in This Derivation
  1. Linear homogeneous-effect model: \(Y_i(t) - Y_i(t') = \beta(t - t')\) for all \(i\), \(t\), \(t'\), and the structural equations are linear.
  2. Relevance: \(\mathrm{Cov}(Z, T \mid X) \neq 0\).
  3. Exogeneity: \(\mathrm{Cov}(Z, \varepsilon \mid X) = 0\).
  4. Exclusion: \(Z\) has no direct effect on \(Y\); it does not appear in the structural equation for \(Y\).

Assumptions 2–4 are the three core IV assumptions; Assumption 1 is the homogeneity restriction specific to this section. Under heterogeneous effects the same Wald ratio identifies a different quantity; see Section 7.7.

We derive the identification of \(\beta\) step by step, using the structural model as the vehicle and the do-calculus to interpret the exclusion step.

  1. Exogeneity (\(Z \indep U \mid X\) in \(\Gcal\)) implies, for the structural residual \(\varepsilon\), that \(\E[\varepsilon \mid Z, X] = \E[\varepsilon \mid X]\). With the normalization \(\E[\varepsilon \mid X] = 0\) (absorbed into \(\alpha\)), we obtain \(\E[\varepsilon \mid Z, X] = 0\).

  2. Exclusion (\(Z\) absent from the structural equation for \(Y\)) is what permits writing Equation 7.1 without a \(Z\) term in the first place: under exclusion, the structural residual \(\varepsilon\) depends on \(U\) but does not absorb a direct effect of \(Z\) on \(Y\). Combined with exogeneity, \(\E[\varepsilon \mid Z, X] = 0\) immediately yields the moment condition \(\E[\varepsilon \cdot Z \mid X] = 0\). The graphical counterpart is in \(\Gcal_{\overline{T}}\): exclusion means there is no direct edge \(Z \to Y\), so \(Y \indep Z \mid T, X\) holds in the mutilated graph and Rule 1 of the do-calculus removes \(Z\) from \(f(y \mid x, \doop(T{=}t), z)\).

  3. Multiplying Equation 7.1 through by \((Z - \E[Z \mid X])\) and taking expectations, the residual term \(\E[\varepsilon (Z - \E[Z\mid X])]\) vanishes by step 2, leaving \[\E\!\bigl[(Y - \alpha - \gamma^\top X)(Z - \E[Z\mid X])\bigr] = \beta\, \E\!\bigl[(T - \E[T\mid X])(Z - \E[Z\mid X])\bigr].\] In the simpler case with no covariates \(X\), this collapses to the reduced-form decomposition \(\mathrm{Cov}(Y, Z) = \beta\, \mathrm{Cov}(T, Z)\), which shows that the \(Z\)\(Y\) covariance is entirely attributable to the causal path \(Z \to T \to Y\) (by exclusion), and equals \(\beta\) times the first-stage covariance.

  4. Relevance (\(\mathrm{Cov}(Z, T \mid X) \neq 0\)) ensures the denominator is non-zero, so we can solve: \[\beta = \frac{\mathrm{Cov}(Y,\, Z \mid X)}{\mathrm{Cov}(T,\, Z \mid X)}. \tag{7.4}\]

NoteThe Wald Estimand

In the binary instrument case (\(Z \in \{0,1\}\)) without additional covariates, Equation 7.4 simplifies to \[\beta \;=\; \frac{\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0]}{\E[T \mid Z{=}1] - \E[T \mid Z{=}0]}. \tag{7.5}\] The numerator is the reduced form: the total effect of \(Z\) on \(Y\). The denominator is the first stage: the effect of \(Z\) on \(T\). The exclusion restriction is what guarantees that the entire reduced form operates through \(T\), with no parallel channel; dividing by the first stage then strips out the \(Z \to T\) piece and leaves \(\beta\).

The key lesson is that the Wald ratio identifies \(\beta\) in this section only because treatment effects are assumed to be constant across units. The observed formula Equation 7.4 is an identifying expression, not an estimand by itself. Its causal meaning depends on the structural assumptions maintained around it. In Section 7.7, the same ratio will survive, but its interpretation will change from a common causal effect to a complier-specific average effect.

NoteRemark: Estimation of the Wald Estimand

The Wald estimand above is an identification result: it expresses the causal parameter \(\beta\) as a ratio of observable quantities. Estimation — that is, how to consistently estimate this ratio from a finite sample, including the correct treatment of standard errors — is taken up in Chapter 13.

NoteExample: Returns to Schooling

In the Mincer earnings equation \(\log W = \alpha + \beta S + \gamma^\top X + \varepsilon\), where \(S\) is years of schooling, the OLS estimator of \(\beta\) is biased upward because unobserved ability \(U\) raises both \(S\) and \(W\) (\(\rho > 0\)). Angrist and Krueger (1991) use quarter of birth as \(Z\): proximity to mandatory school-leaving age at different birth quarters generates exogenous variation in completed schooling. The IV estimate of the return to schooling is approximately \(0.08\)\(0.10\) per year, similar to the OLS estimate in this case, suggesting the ability bias is modest.

7.4 Why the IV Assumptions Matter

The three IV assumptions are not merely technical conditions: each one is load-bearing, and each failure mode produces a distinct, quantifiable distortion of the Wald estimand. Violating exogeneity or exclusion contaminates the reduced form, and dividing by a weak first stage magnifies that contamination. A weak instrument is therefore not merely inefficient; it makes any small violation of the identifying assumptions more consequential.

When relevance fails. If \(\pi = 0\), the Wald estimand is undefined: its denominator is zero and there is no exogenous variation to exploit. When \(\pi\) is small but nonzero, the instrument is weak. The estimator’s variance diverges as \(\pi \to 0\) because the denominator amplifies noise, and finite-sample bias pulls the IV estimate toward the OLS estimate at a rate proportional to \(1/F\), where \(F\) is the first-stage \(F\)-statistic.

When exclusion fails. If the outcome equation contains a direct effect \(Y = \alpha + \beta T + \delta Z + \gamma^\top X + \varepsilon\) with \(\delta \neq 0\), the reduced form picks up both the indirect path \(\beta\pi\) and the direct path \(\delta\), so the Wald estimand converges to \(\beta + \delta/\pi\). The bias \(\delta/\pi\) is amplified by a weak first stage: a small direct effect combined with a weak instrument can produce large bias. This is why a weak instrument with a plausible exclusion violation is not “nearly valid” — it may be severely misleading.

When exogeneity fails. If \(\mathrm{Cov}(Z, \varepsilon \mid X) \neq 0\), the IV moment condition fails and the Wald estimand converges to \(\beta + \mathrm{Cov}(\varepsilon, Z)/\mathrm{Cov}(T, Z)\). This is the IV analogue of OLS omitted-variable bias: IV fails when the instrument itself is endogenous, just as OLS fails when the treatment is endogenous. Again the bias is amplified when the first stage is weak.

Testability of the three IV assumptions. Although exogeneity and exclusion are not directly testable, they are not entirely beyond empirical scrutiny. The IV model imposes inequality restrictions on the observable data distribution; when these are violated, falsification tests can detect the violation (Kitagawa 2015; Hernán and Robins 2006). Importantly, such tests are not consistent against all alternatives: for most data distributions under which the assumptions fail, the tests will not reject even in large samples.
Assumption Directly testable? Basis for assessment
Relevance Association testable; causal claim design-based First-stage \(F\)-statistic and \(\mathrm{Cov}(Z,T \mid X)\) test the observable association; the causal \(Z \to T\) link rests on the design
Exogeneity No Institutional knowledge; randomization (if available); placebo regressions on pre-determined outcomes
Exclusion No in just-identified case; partially in overidentified case Institutional argument; overidentification test (\(J\)-test, Chapter 13) checks mutual consistency of instruments but cannot confirm all are valid

The key asymmetry is that the hardest work in defending an IV design — the institutional arguments for exogeneity and exclusion — is precisely the work that the data cannot do for the researcher.

7.5 Lab: OLS vs. IV Across Instrument Strengths

The analytic results of Section 7.3 and Section 7.4 established two facts: OLS is biased by a precise, computable amount whenever \(\rho \neq 0\); and IV is consistent but its variance diverges as \(\pi \to 0\). What the analytics do not immediately reveal is the finite-sample consequence: for weak instruments, the IV variance penalty is so severe that the biased OLS estimator can have lower mean squared error. This simulation studies estimator behavior conditional on the IV assumptions being true; it does not address the separate and fundamentally substantive question of whether a proposed instrument is valid in an applied study.

NoteDGP for Lab 7

The simulation implements the linear structural model of Section 7.3 with no covariates. Each replication draws \(n = 500\) observations from \[Y_i = \beta T_i + \varepsilon_i, \qquad T_i = \pi Z_i + \eta_i,\] with \(Z_i \overset{\mathrm{i.i.d.}}{\sim} \mathrm{N}(0,1)\) and the structural errors generated as \[\eta_i = U_i, \qquad \varepsilon_i = \rho U_i + \sqrt{1 - \rho^2}\,\xi_i, \qquad U_i,\, \xi_i \overset{\mathrm{i.i.d.}}{\sim} \mathrm{N}(0,1).\] Here \(U_i\) is the unobserved confounder shared by both equations, and \(\xi_i\) is the idiosyncratic component of \(\varepsilon_i\). This gives \(\mathrm{Var}(\varepsilon_i) = \mathrm{Var}(\eta_i) = 1\) and \(\mathrm{Cov}(\varepsilon_i, \eta_i) = \rho\). The fixed parameter values are \(\beta = 1\) (true causal effect) and \(\rho = 0.8\) (strong positive endogeneity). The first-stage coefficient \(\pi\) is varied across eight values: \[\pi \;\in\; \{0,\; 0.10,\; 0.15,\; 0.20,\; 0.30,\; 0.50,\; 1.00,\; 2.00\}.\] The expected first-stage \(F\)-statistic is approximately \(n\pi^2 = 500\pi^2\). The Wald estimand is undefined at \(\pi = 0\); the Staiger–Stock rule of thumb \(F \geq 10\) corresponds to \(\pi \geq 0.141\) in this design.

NotePotential Outcomes Interpretation of the DGP

The structural DGP has an exact translation into the potential outcomes language of Chapter 4, which makes explicit why OLS fails and why IV succeeds.

Potential outcomes. The latent background factor \(\varepsilon_i\) is fixed for each unit across all counterfactual worlds. The potential outcome under treatment value \(t\) is therefore \[Y_i(t) \;=\; \beta t + \varepsilon_i. \tag{7.6}\] The observed outcome satisfies \(Y_i = Y_i(T_i)\), which is consistency. Because \(\varepsilon_i\) does not depend on \(t\), the individual treatment effect is constant: \(Y_i(t) - Y_i(t') = \beta(t - t')\) for all \(i\). This is exactly Framework 1: \(\beta\) is simultaneously the ATE, the ATT, and the LATE.

The unobserved confounder. The shared factor \(U_i\) enters both the treatment assignment and the background outcome. Because \(\varepsilon_i\) depends on \(U_i\) and \(\eta_i = U_i\), units with large \(T_i\) (driven by large \(U_i\)) also tend to have large \(\varepsilon_i\), and hence large \(Y_i\), even if \(\beta = 0\). Formally, \[\mathrm{Cov}\!\bigl(Y_i(t),\, T_i\bigr) = \mathrm{Cov}(\varepsilon_i,\, \eta_i) = \rho \;\neq\; 0,\] so unconfoundedness fails. The parameter \(\rho\) is precisely the degree of violation.

Why IV restores identification. The instrument \(Z_i\) is drawn independently of \(U_i\) and \(\xi_i\), so \(Z_i \indep U_i\) implies \(Z_i \indep \varepsilon_i\) — exogeneity, satisfied by construction. The instrument does not appear in \(Y_i(t)\) (Equation 7.6), which is the exclusion restriction: \(Y_i(t, z) = Y_i(t)\) for all \(z\). And \(Z_i\) moves \(T_i\) through \(\pi\), which is relevance. All three IV assumptions hold exactly by construction.

What varies across the grid. The potential outcomes \(Y_i(t)\) are the same for every value of \(\pi\) — the parameter \(\pi\) governs only how strongly the instrument shifts treatment. All that changes for OLS as \(\pi\) grows is that more of \(T_i\)’s variance comes from the clean source \(Z_i\) rather than from the confounded source \(U_i\), which partially dilutes the endogeneity bias. But because the potential outcomes themselves are unchanged, \(\beta = 1\) remains the target at every row of the table.

Estimators. OLS regresses \(Y\) on \(T\); from Equation 7.3 it converges to \(\beta + \rho/(\pi^2 + 1) = 1 + 0.8/(\pi^2+1)\). The bias is always positive (since \(\rho > 0\)), decreasing in \(|\pi|\), but never zero for finite \(\pi\). IV (Wald) uses \(Z\) as the instrument; it is consistent for \(\beta\) for any \(\pi \neq 0\), with finite-sample variance approximately \(\sigma^2/(n\pi^2)\), diverging as \(\pi \to 0\).

Results (\(n = 500\), \(B = 2{,}000\) replications, seed 2024; true \(\beta = 1.000\)):

Monte Carlo performance of OLS and IV across instrument strengths. Theoretical OLS bias \(= 0.8/(\pi^2+1)\). “—” denotes \(\pi = 0\), where IV is not defined. IV mean and SD at \(\pi = 0.10\) and \(0.15\) are dominated by extreme outliers; the IV median is \(\approx 1.01\) at both values.
\(\pi\) \(F\) Theory bias OLS mean OLS bias OLS SD OLS RMSE IV mean IV bias IV SD IV RMSE
0.00 1 +0.800 1.801 +0.801 0.026 0.802
0.10 6 +0.792 1.791 +0.791 0.027 0.792 0.872 −0.128 4.391 4.393
0.15 13 +0.782 1.782 +0.782 0.027 0.783 0.764 −0.236 6.487 6.491
0.20 21 +0.769 1.769 +0.769 0.027 0.770 0.940 −0.060 0.380 0.385
0.30 46 +0.734 1.735 +0.735 0.028 0.735 0.975 −0.025 0.163 0.165
0.50 127 +0.640 1.639 +0.639 0.028 0.640 0.996 −0.004 0.091 0.091
1.00 500 +0.400 1.401 +0.401 0.027 0.402 0.999 −0.001 0.045 0.045
2.00 2006 +0.160 1.160 +0.160 0.019 0.161 1.000 0.000 0.022 0.022

1. The OLS bias formula is exact. The theory-bias column matches the simulated OLS bias to within Monte Carlo error across all eight values of \(\pi\). OLS is biased in the direction of \(\rho\) at every value of \(\pi\), including \(\pi = 0\) where there is no instrument at all. Crucially, the OLS bias decreases as \(\pi\) grows, because a stronger first stage means more of \(T\)’s variance comes from the exogenous source \(Z\). At \(\pi = 2\) the bias is only \(0.160\), yet this still produces RMSE \(= 0.161\), which is \(7\times\) larger than the IV RMSE of \(0.022\) at the same \(\pi\).

2. IV is consistent but has catastrophically heavy tails when the instrument is weak. At \(\pi = 0.10\) (\(F \approx 6\)) and \(\pi = 0.15\) (\(F \approx 13\)) the IV mean is far from the true value, suggesting bias. This is misleading: the IV median is \(\approx 1.01\) at both values, confirming that the estimator is consistent and the typical replication recovers \(\beta = 1\). The mean is dragged off by a small fraction of replications in which the first stage happens to be near zero, making the Wald ratio explode. Indeed, the just-identified IV estimator possesses no finite moments, so the mean and SD entries at \(\pi = 0.10\) and \(0.15\) have no population counterparts: they are determined by the few most extreme replications and vary substantially across seeds. The take-away is that the weak instrument problem is not inconsistency; it is uncontrolled tail behavior.

3. The RMSE crossover occurs near \(F \approx 20\). OLS RMSE ranges from \(0.802\) down to \(0.161\), while IV RMSE starts infinite, is catastrophic at \(\pi = 0.10\)\(0.15\), then drops sharply to \(0.385\) at \(\pi = 0.20\) (\(F \approx 21\)) and to \(0.165\) at \(\pi = 0.30\). IV first beats OLS on RMSE at \(\pi = 0.20\): this is the crossover. Below this threshold, the variance penalty of IV exceeds the bias penalty of OLS and the biased estimator is preferable in mean squared error terms. The Staiger–Stock rule of thumb (\(F \geq 10\)) is therefore slightly too lenient: at \(F \approx 13\) the IV RMSE in this run is still an order of magnitude above OLS. A more conservative threshold of \(F \geq 20\)\(25\) is needed before IV dominates OLS in this DGP.

4. A strong instrument eliminates both problems. At \(\pi = 1.00\) (\(F \approx 500\)), IV is virtually unbiased with RMSE \(= 0.045\), while OLS still carries a bias of \(0.401\), giving RMSE \(= 0.402\) — a 9-fold improvement from IV. With a strong instrument, IV is simultaneously unbiased and more efficient than OLS in MSE terms, because OLS efficiency is illusory: its small variance is offset by a large, persistent bias.

WarningRMSE Cannot Be the Only Criterion

The table shows that OLS has lower RMSE than IV for \(\pi < 0.20\) in this DGP. This does not vindicate OLS. An estimator with RMSE \(= 0.78\) because it is biased by \(0.80\) is useless for causal inference: the bias is systematic and does not shrink with sample size. As \(n \to \infty\), IV RMSE shrinks to zero while OLS RMSE stays at \(0.80\). The RMSE comparison is only meaningful in finite samples. For a fixed \(n\), it quantifies when a weak instrument is so unreliable that IV is not yet practically useful — but that is an argument for finding a stronger instrument, not for using OLS.

NoteRemark: The First-Stage \(F\)-Statistic as a Diagnostic

The approximate formula \(F \approx n\pi^2\) gives first-stage \(F\)-statistics of \(6\), \(13\), \(21\), \(46\), \(127\), \(500\), and \(2006\) for the seven non-zero \(\pi\) values. The Staiger–Stock threshold of \(F \geq 10\) is widely used in practice and corresponds to IV RMSE within a factor of roughly two of the strong-instrument limit; the weak-instrument problem it addresses was documented forcefully by Bound et al. (1995). The simulation here suggests this threshold may understate the problem when endogeneity is strong (\(\rho = 0.8\)): at \(F = 13\) the IV RMSE in this run is still far above any useful threshold. A researcher should report the first-stage \(F\) alongside any IV estimate, and treat values below \(20\)\(25\) with particular caution.

7.6 Multiple Instruments and Overidentification

This section gives only the identification-level intuition for using more than one instrument. The computational details of 2SLS with multiple instruments, GMM weighting, and the Sargan–Hansen \(J\)-test are deferred to Chapter 13.

When there are exactly as many instruments as endogenous variables (\(q = p\)), the model is just-identified: the IV assumptions pin down the causal parameter exactly, leaving no surplus identifying variation. When \(q > p\), the model is overidentified: the extra instruments impose additional moment restrictions beyond those needed for identification. Under the homogeneous-effect linear model, every valid instrument must imply the same structural coefficient \(\beta\), so those extra restrictions are testable. Under heterogeneous treatment effects, valid instruments can legitimately identify different LATEs because they shift treatment for different complier populations; the overidentification test then targets equality of probability limits across instruments, and a rejection admits more interpretations than instrument invalidity. In both cases, the Sargan–Hansen \(J\)-test (Sargan 1958; Hansen 1982) formalizes the mutual-consistency check; rejection indicates that at least one of the instruments either violates exogeneity or exclusion or identifies a different LATE, but does not localize which.

The terminology distinguishes the order condition — a counting requirement that \(q \geq p\) — from the rank condition, which requires the instruments to be linearly independent in the first stage. Both are necessary for identification. The order condition is met by inspection; the rank condition is testable from the first-stage coefficient matrix.

WarningWhat Overidentification Does Not Do

Passing the \(J\)-test does not confirm that any instrument is valid. It only shows that the sample moment conditions are mutually compatible — a weaker conclusion than validity. In the just-identified case the test has no power at all. Multiple invalid instruments can agree with one another if they share the same violation, so a consistent set of instruments is not necessarily a valid one. Heterogeneous treatment effects or model misspecification can also affect the test’s behavior. Overidentification is an opportunity for a consistency check, not a substitute for the institutional argument that makes a design credible.

NoteRemark: GMM and Overidentification

In the overidentified case, the surplus moment conditions can also be exploited for efficient estimation: rather than discarding the extra instruments, one can combine all \(q\) moment conditions optimally. The generalized method of moments (GMM) framework provides the natural setting for this, treating the just-identified IV estimator as a special case. This connection is developed in Chapter 13.

7.7 Heterogeneous Treatment Effects and the LATE Framework

NoteFramework 2: Heterogeneous Treatment Effects

Section 7.3 imposed a common treatment effect \(\beta\) for every unit. That assumption is often too strong for applied work. We now drop homogeneity and allow \(\tau_i = Y_i(1) - Y_i(0)\) to vary across units. Throughout this section we work in the binary instrument, binary treatment setting (\(Z, T \in \{0,1\}\)).

In this more realistic setting, the Wald ratio generally no longer identifies the ATE. Instead, under an additional assumption of monotonicity, it identifies the Local Average Treatment Effect (LATE): the average treatment effect for the subpopulation whose treatment status is changed by the instrument. The contrast with Framework 1 should be kept explicit:

  • In Framework 1, the Wald ratio identifies a single common effect \(\beta = \mathrm{ATE} = \mathrm{ATT} = \mathrm{LATE}\).
  • In Framework 2, the same ratio identifies a subgroup-specific average effect for compliers: \(\tau_{\mathrm{LATE}} \neq \mathrm{ATE}\) in general.

The observed-data formula remains the same; the causal target changes.

NoteRemark: Nonidentification of the ATE Is Structural, Not Estimator-Specific

The failure is not a defect of the Wald ratio: under relevance, exogeneity, exclusion, and monotonicity alone, no functional of the observed joint distribution \(P(Z, T, Y)\) identifies the ATE. Problem 7 constructs two models that agree on \(P(Z, T, Y)\) — and hence on the Wald ratio and, as the problem shows, on the LATE — yet disagree on the ATE, in the two-model style of the completeness discussion of Chapter 3. Point identification of the ATE therefore requires something beyond the core IV assumptions: effect homogeneity as in Framework 1, an extrapolation assumption such as those discussed in Section 7.8.2, or a different target (such as the LATE itself, or partial-identification bounds, Chapter 9).

7.7.1 Compliance Types

In the binary-instrument, binary-treatment setting, each unit’s response to the instrument is fully described by the pair of potential treatment decisions \((T_i(0), T_i(1))\): the treatment the unit would take under each value of \(Z\). This pair defines the unit’s compliance type — a latent causal classification, because \((T_i(0), T_i(1))\) is never jointly observed in data.

NoteDefinition: Compliance Types (Angrist et al. 1996)

For binary \(Z, T \in \{0,1\}\), define the potential treatment \(T_i(z)\) as the treatment unit \(i\) would take under instrument value \(z\). The four compliance types are:

Compliance type \(T_i(0)\) \(T_i(1)\)
Complier 0 1
Always-taker 1 1
Never-taker 0 0
Defier 1 0

A complier takes treatment if and only if the instrument is switched on. An always-taker takes treatment regardless of \(Z\). A never-taker never takes treatment regardless of \(Z\). A defier does the opposite of what the instrument suggests.

The instrument \(Z\) only shifts treatment for compliers: always-takers and never-takers have the same treatment status regardless of \(Z\), so they contribute nothing to the denominator \(\E[T \mid Z{=}1] - \E[T \mid Z{=}0]\). This is the first indication that the Wald ratio will be driven by the complier subgroup rather than by the full population.

7.7.2 The Monotonicity Assumption

The three core IV assumptions alone do not yield a clean causal interpretation for the Wald ratio when effects are heterogeneous. The difficulty is that the instrument may push some units toward treatment and others away from it: in the presence of defiers, the numerator and denominator of the Wald ratio conflate effects in opposite directions. A fourth assumption rules out this ambiguity.

NoteDefinition: Monotonicity (Angrist and Imbens 1994)

The treatment assignment is monotone in \(Z\) if \(T_i(1) \geq T_i(0)\) for all \(i\). Equivalently: there are no defiers in the population.

Interpretation. Monotonicity is not a generic law of causal inference. It is a design-specific claim about how this particular instrument changes treatment behavior. Switching the instrument from \(0\) to \(1\) may induce some units to take treatment (compliers) and leave others unaffected (always-takers or never-takers), but it should not reverse anyone’s treatment decision. This is most plausible when the instrument is a randomized encouragement, access rule, or administrative assignment mechanism. In many observational IV settings the no-defiers assumption is substantively harder to defend and requires explicit justification.

7.7.3 The LATE Theorem

Under the three core IV assumptions, heterogeneous treatment effects, and monotonicity, the Wald ratio no longer identifies the ATE. Instead it recovers the average treatment effect for the units whose treatment status is actually changed by the instrument — the compliers. The effect is local not because it is estimated with nearby observations, but because it pertains to a local margin of behavioral response defined by the instrument itself.

Theorem 7.1 (LATE Theorem (Angrist and Imbens 1994)) Suppose the following hold:

  1. Exogeneity (PO form): \(Z \indep \bigl(Y(0),\, Y(1),\, T(0),\, T(1)\bigr)\);
  2. Exclusion: \(Y_i(t, z) = Y_i(t)\) for all \(z\);
  3. Relevance: \(\E[T(1) - T(0)] \neq 0\);
  4. Monotonicity: \(T_i(1) \geq T_i(0)\) for all \(i\).

Then the Wald estimand identifies the Local Average Treatment Effect (LATE): \[\frac{\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0]}{\E[T \mid Z{=}1] - \E[T \mid Z{=}0]} \;=\; \E\!\bigl[Y(1) - Y(0) \;\big|\; T_i(1) > T_i(0)\bigr] \;\equiv\; \tau_{\mathrm{LATE}}. \tag{7.7}\] That is, the Wald estimand identifies the average causal effect for compliers only. When covariates are present, the same result holds with all assumptions and the Wald formula stated conditionally on \(X\).

Proof. We decompose the numerator and denominator by compliance type. Throughout, consistency gives \(T_i = T_i(Z_i)\) and \(Y_i = Y_i(T_i, Z_i)\), and exclusion reduces the latter to \(Y_i = Y_i(T_i(Z_i))\).

Denominator. By consistency for \(T\) and exogeneity (\(Z \indep (T(0), T(1))\)): \[\begin{aligned} \E[T \mid Z{=}1] - \E[T \mid Z{=}0] &= \E[T(1) \mid Z{=}1] - \E[T(0) \mid Z{=}0] \\ &= \E[T(1)] - \E[T(0)] = \E[T(1) - T(0)], \end{aligned}\] and monotonicity gives \(\E[T(1) - T(0)] = P(\text{complier})\), since always-takers contribute \(1 - 1 = 0\) and never-takers contribute \(0 - 0 = 0\).

Numerator. By consistency, exclusion, and exogeneity: \[\begin{aligned} \E[Y \mid Z{=}1] - \E[Y \mid Z{=}0] &= \E[Y(T(1))] - \E[Y(T(0))] \\ &= \sum_{c} P(c)\, \E\!\bigl[Y(T_c(1)) - Y(T_c(0)) \;\big|\; \text{type}=c\bigr]. \end{aligned}\] For always-takers, \(T(1) = T(0) = 1\), so the contribution is \(0\). For never-takers, \(T(1) = T(0) = 0\), contribution \(0\). For compliers, \(T(1) = 1\) and \(T(0) = 0\), so \(Y(T(1)) - Y(T(0)) = Y(1) - Y(0)\). No defiers exist by monotonicity. Therefore \[\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0] = P(\text{complier})\,\E\!\bigl[Y(1) - Y(0) \mid \text{complier}\bigr].\]

Ratio. Dividing numerator by denominator gives Equation 7.7. \(\square\)

NoteRemark: The Two Senses of “Local”

The LATE is local in two senses: it is local to the complier group defined by this instrument, and local to the particular instrument that defines that group. A different instrument, even for the same treatment, will generally select a different complier population and identify a different LATE. This is the subject of Section 7.8.3.

7.8 Interpreting IV Estimands

The two frameworks — linear homogeneous effects and the LATE framework — give the Wald estimand different interpretations. This section synthesizes those interpretations and clarifies when they agree, when they disagree, and what the difference means for applied work.

7.8.1 What the Two Frameworks Say

Framework 1 (linear, homogeneous) Framework 2 (heterogeneous effects)
Key assumption \(\tau_i = \beta\) for all \(i\) Monotonicity; no defiers
What IV identifies \(\beta = \mathrm{ATE} = \mathrm{ATT} = \mathrm{LATE}\) \(\tau_{\mathrm{LATE}} = \E[\tau_i \mid \text{complier}]\)
Estimand depends on instrument? No (same \(\beta\) regardless of \(Z\)) Yes (different \(Z\) \(\Longrightarrow\) different compliers \(\Longrightarrow\) different LATE)
Identifies the ATE? Yes, automatically Only if all units are compliers or effects homogeneous

Framework 1 is a special case of Framework 2: when \(\tau_i = \beta\) for all \(i\), the LATE equals the ATE equals \(\beta\), and the instrument does not affect the estimand — only identification. Framework 2 is the more general and realistic setting. In applied work, the default interpretation of the Wald estimand is the LATE; the ATE interpretation requires the additional homogeneity argument of Framework 1.

NoteRemark: Same Formula, Different Estimand

The two frameworks illustrate a broader principle in IV analysis. Across a range of fourth assumptions — constant effects, monotonicity, structural mean model restrictions, and others — the conditional Wald formula serves as the identifying expression in many (though not all) cases (Levis et al. 2024). Accordingly, two researchers who use the same instrument and compute the same Wald ratio may be consistently estimating different causal quantities if they maintain different structural assumptions: one estimates the ATE (under constant effects), another estimates the LATE (under monotonicity), and a third estimates the ATT (under a structural mean model restriction). The point estimates and confidence intervals can be numerically identical; what differs is the causal quantity they represent.

This is not a deficiency of IV but a precise illustration of why causal modeling is not merely a formal preliminary to estimation. The data alone cannot resolve which estimand the Wald ratio identifies; that determination requires the researcher to commit to a structural assumption about treatment response. Stating that commitment explicitly — and defending it on substantive grounds — is as important as any statistical calculation.

7.8.2 When Does LATE Equal ATE?

LATE equals ATE only under additional structure, most notably treatment-effect homogeneity or special forms of heterogeneity that make complier effects representative of the full population. Neither condition should be assumed without argument.

In general, \(\tau_{\mathrm{LATE}} \neq \mathrm{ATE}\) when treatment effects are heterogeneous and some units are not compliers. To see this, decompose the ATE by compliance type: \[\mathrm{ATE} = P(\mathrm{co})\,\E[\tau_i \mid \mathrm{co}] + P(\mathrm{at})\,\E[\tau_i \mid \mathrm{at}] + P(\mathrm{nt})\,\E[\tau_i \mid \mathrm{nt}],\] where co, at, and nt abbreviate complier, always-taker, and never-taker. The LATE equals only the first term divided by its probability weight. The ATE and LATE coincide if and only if:

  1. Mean treatment effects are equal across compliance types — a strict weakening of the constant-effect assumption of Section 7.3, but still a non-trivial restriction on heterogeneity; or
  2. Everyone is a complier (\(P(\mathrm{at}) = P(\mathrm{nt}) = 0\)), which would require \(Z\) to perfectly determine \(T\); or
  3. The average effects for always-takers and never-takers happen to equal the LATE — an untestable coincidence.

In practice, none of these conditions is likely to hold exactly. The LATE is a well-defined, identifiable parameter, but it is not the ATE.

7.8.3 Different Instruments, Different Estimands

Because the LATE is specific to the complier population, and different instruments select different complier populations, two valid instruments for the same treatment can legitimately identify different LATEs. This is a precise, substantive fact — not a contradiction. Different instruments can legitimately identify different causal effects because they shift treatment for different margins of the population; the diversity of LATEs is informative about treatment effect heterogeneity, not a symptom of model failure.

NoteExample: Compulsory Schooling Laws vs. Distance to College

Two classic instruments for years of schooling are (i) compulsory schooling laws (Angrist and Krueger 1991) and (ii) distance to the nearest college (Card 1995). The compliers for instrument (i) are individuals at the margin of dropping out before the compulsory leaving age — typically lower-income students. The compliers for instrument (ii) are students deterred from college by geographic distance — again, disproportionately lower-income. Both instruments identify the return to schooling for their respective complier populations, and the LATEs can legitimately differ even if both instruments are valid.

The dependence of the LATE on the instrument is sometimes described as a limitation of IV. It is better understood as a precise statement about what question is being answered. A researcher using compulsory schooling laws is estimating the return to schooling for students at the compulsory leaving margin; a researcher using distance to college is estimating the return for students deterred by geography. These are different causal questions, and it is informative — not troubling — that they can yield different answers.

7.8.4 The Policy Relevance of LATE

For many policy questions, the LATE is exactly the right estimand. If a policy is designed to encourage a subset of the population to take treatment — for example, an outreach program that reaches only some potential participants — then the effect on compliers is precisely what the policy-maker wants to know. LATE is most policy-relevant when the contemplated intervention resembles the instrument, because then the complier population under the study design is close to the policy-relevant margin.

When the ATE over the full population is required for policy analysis, IV alone is insufficient under heterogeneous effects. Additional assumptions — such as an explicit model of treatment effect heterogeneity, or a second instrument that identifies effects for a different subpopulation — are needed to extrapolate from the LATE to the ATE.

7.9 IV versus Back-Door Adjustment

Having developed both strategies in full, it is instructive to compare them directly. The two strategies differ in what they require, what they identify, and how they fail.

7.9.1 A Direct Comparison

Dimension Back-door / propensity score Instrumental variables
Core assumption All confounders observed: \((Y(0),Y(1)) \indep T \mid X\) Valid instrument: relevance, exogeneity, exclusion
Unobserved confounders Fatal: back-door adjustment fails Permitted: IV routes around \(U\)
Estimand ATE, ATT, or ATC depending on design and overlap; all coincide under homogeneity LATE (compliers only); reduces to a common \(\beta\) under homogeneous effects
Testability Unconfoundedness untestable; overlap testable Relevance testable; exogeneity and exclusion untestable (just-identified case)
Main threat Unmeasured confounder Exclusion restriction violation
Identifies the ATE? Yes, under strong ignorability Only under homogeneous effects
Typical setting Rich administrative or survey data; RCT with imperfect compliance Natural experiment; RCT with non-compliance; policy change

7.9.2 Complementary Failure Modes

Back-door adjustment fails when \(X\) does not capture all confounders — that is, when an unobserved variable \(U\) opens a back-door path. The estimator is then inconsistent even with infinite data, because the path \(T \leftarrow U \rightarrow Y\) transmits spurious association that conditioning on \(X\) cannot close.

IV fails when the exclusion restriction is violated — that is, when \(Z\) has a direct effect on \(Y\) beyond its effect through \(T\). As derived in Section 7.4, the bias in the Wald estimand is \(\delta/\pi\), amplified by weak instruments.

The two failures are in a sense orthogonal: back-door adjustment requires many observed covariates but tolerates no unobserved ones, while IV tolerates unobserved confounders but requires an instrument with no direct effect on the outcome. Back-door adjustment fails when confounding remains after conditioning; IV fails when the proposed source of exogenous variation is not truly exogenous or does not act solely through treatment. When designing a study, the choice of identification strategy should be guided by which assumption is more plausible in the specific empirical context.

7.9.3 When Both Strategies Are Available

When a valid instrument and a sufficient adjustment set \(X\) are both available, comparing the two estimates can be informative, but disagreement between them does not by itself tell us which method is wrong. The two strategies typically target different estimands — back-door adjustment identifies the ATE or ATT over the full population or the treated, while IV identifies the LATE for compliers — and they rely on different identifying assumptions. Agreement is therefore reassuring under treatment-effect homogeneity but is neither required nor sufficient otherwise.

The Hausman (1978) endogeneity test operationalizes this comparison: under the null hypothesis that \(T\) is exogenous given \(X\), both the OLS and IV estimators are consistent, and a large discrepancy is evidence of endogeneity. Under the null and standard regularity conditions, an appropriately scaled quadratic form in the discrepancy \(\hat\beta_{\mathrm{IV}} - \hat\beta_{\mathrm{OLS}}\) is asymptotically \(\chi^2_p\)-distributed, where \(p\) is the number of components of \(T\) tested for endogeneity; the formal construction is given in Chapter 13. Rejection means at least one strategy is inconsistent; it does not identify which.

NoteRemark: Interpreting a Hausman Rejection

A significant Hausman test can arise from two distinct sources: (a) back-door adjustment is biased because \(T\) is endogenous, or (b) IV is biased because the exclusion restriction is violated or the instrument is weak. Even when the test is non-significant, the two estimates may still differ for the legitimate reason that they identify different estimands — LATE versus ATE — rather than because either strategy has failed. External evidence — instrument strength diagnostics and institutional arguments for exclusion — is needed to interpret any discrepancy.

7.10 Practical Guidance on Defending an IV Design

There is no algorithm for finding a valid instrument; credible instruments arise from institutional knowledge and careful reasoning about the data-generating process. A strong IV design is defended primarily by institutional knowledge, design logic, and causal structure; statistical diagnostics are supportive but secondary.

A researcher proposing an instrument should be able to answer five questions explicitly:

  1. What exactly is the instrument? Specify \(Z\) precisely: its source of variation, the level at which it varies, and the population to which it applies.

  2. Why does it shift treatment? Articulate the causal mechanism by which \(Z\) moves \(T\). Relevance can be verified empirically with the first-stage \(F\)-statistic, but the \(F\)-statistic is a diagnostic for instrument strength, not a substitute for a causal account of the \(Z \to T\) link.

  3. Why is it as-if random relative to latent outcome determinants? Argue why \(Z\) is unrelated to the unobserved causes of \(Y\). The most credible sources are designed randomization (lotteries, randomized encouragement), natural experiments with institutional quasi-randomness (policy discontinuities, geographic boundaries, biological quirks), and shift-share designs (Bartik 1991; Goldsmith-Pinkham et al. 2020). Placebo regressions on pre-determined outcomes provide partial — but not definitive — evidence.

  4. Why can it affect the outcome only through treatment? The exclusion restriction is untestable in just-identified models, so it must rest on a structural argument that no direct path \(Z \to Y\) exists. A useful diagnostic is to ask: how large would the direct effect \(\delta\) have to be, relative to the first-stage coefficient \(\pi\), to overturn the estimated causal effect? When the first stage is weak, the answer is: not very large at all.

  5. What population margin does it shift? Identify the complier population — the units whose treatment status changes with \(Z\). This determines the LATE that is being identified and governs the external validity of the estimates for other populations or policy margins.

7.11 Applied Example: Charter School Lotteries and the KIPP Lynn Study

This example illustrates a canonical randomized-encouragement IV design. The key distinction is that the lottery randomizes offer status, not actual treatment. Winning the lottery does not mechanically force a student to attend KIPP, and losing the lottery does not make later attendance impossible in all cases. Thus the lottery offer is the instrument \(Z\), while actual years of KIPP attendance is the endogenous treatment \(S\). Here \(S\) plays the role of the generic treatment variable \(T\) used throughout the rest of this chapter; the paper’s notation is retained to keep the empirical discussion close to the source.

This distinction is exactly what makes the design an IV design rather than a randomized controlled trial on treatment itself. The lottery generates exogenous variation in access to KIPP, and the IV analysis uses only the portion of attendance variation induced by that randomized offer to identify a causal effect.

7.11.1 Setting and Instrument

KIPP (Knowledge Is Power Program) schools follow a “No Excuses” model: extended school days, a longer academic year, selective teacher hiring, and strict behavioral norms. KIPP Academy Lynn was substantially oversubscribed beginning in 2005. Massachusetts law requires oversubscribed charter schools to select students by lottery, so the school conducted randomized admissions lotteries from 2005 through 2008. The treatment is years of KIPP attendance \(S\); in the panel structure of the data this is denoted \(s_{igt}\), the number of calendar years student \(i\) has spent at KIPP by the time of test \((g, t)\) — a continuous, endogenous variable, because families self-select into applying and attending. The outcome \(Y_{igt}\) is the student’s standardized score on the Massachusetts Comprehensive Assessment System (MCAS), normalized to mean zero and standard deviation one within each subject–grade–year cell statewide.

Z S Y U X relevance θ no U → Z (exogeneity) no Z → Y (exclusion)
Causal DAG for the KIPP Lynn lottery IV design. $Z$ = lottery offer (instrument); $S$ = years enrolled at KIPP (endogenous treatment); $Y$ = MCAS test score; $U$ = unobserved determinants of both enrollment and achievement; $X$ = observed baseline covariates. Two absent arrows encode the IV assumptions: no $U \to Z$ edge (exogeneity) and no direct $Z \to Y$ arrow (exclusion).

7.11.2 Mapping the Three Assumptions to the KIPP Context

The KIPP lottery design makes the logic of the three IV assumptions unusually transparent. One point deserves emphasis before proceeding. Relevance and exogeneity are especially natural in a lottery-based design: because lottery sequence numbers are randomly assigned within cohort, offer status is independent of the latent determinants of achievement by construction. Exclusion, however, still requires institutional argument. Random assignment of the instrument does not, by itself, imply that the instrument has no direct effect on the outcome — it only guarantees that the instrument is exogenous.

Relevance. The lottery offer must shift years of KIPP attendance. Lottery winners were offered a seat and about 80 percent accepted; losers rarely enrolled elsewhere at KIPP. The first-stage regression of \(s_{igt}\) on \(Z_i\) (with year and grade controls) yields a coefficient of approximately \(1.2\): at the time of each MCAS exam, lottery winners had spent about 1.2 more years at KIPP than lottery losers. The first-stage \(F\)-statistic is far above conventional thresholds. The first stage is less than the theoretical maximum because compliance is partial (some winners do not attend; some losers find entry through later cohorts or attrition slots) and the follow-up window captures students at different points in their KIPP tenure.

Exogeneity. The lottery offer must be independent of all potential outcomes and potential treatments — formally, \(Z \indep (Y(0), Y(1), T(0), T(1))\) within each application cohort. Because offer status was determined by randomly drawn lottery-sequence numbers, this independence holds by design — an especially strong basis for exogeneity compared to most observational IV applications. The authors verify it empirically: a joint test of covariate balance across lottery winners and losers yields a \(p\)-value of \(0.615\) for demographic characteristics and baseline test scores, consistent with the null of no pre-lottery differences. Pre-lottery variables are used for the balance check precisely because post-lottery variables such as LEP or SPED classification may themselves be affected by school attended.

Exclusion. The lottery offer must affect test scores only through KIPP attendance, not through any direct channel. Because the offer merely provides access to a school — it does not itself deliver instruction — this restriction is institutionally plausible. One potential violation is a discouragement effect: losing the lottery might demoralize students, suppressing their achievement regardless of where they enroll. The authors address this by noting that the scores of lottery losers are typical of demographically comparable students in Lynn, which is inconsistent with large discouragement effects.

7.11.3 First Stage, Reduced Form, and the 2SLS Estimand

The equations below illustrate the components of the Wald/2SLS logic — first stage, reduced form, and their ratio. Formal estimation theory for 2SLS, including asymptotic variance and cluster-robust inference, is deferred to Chapter 13.

The structural equation for test scores is \[y_{igt} = \alpha_t + \beta_g + \sum_j \delta_j d_{ij} + \gamma' X_i + \theta\, s_{igt} + \varepsilon_{igt}, \tag{7.8}\] where \(\alpha_t\) and \(\beta_g\) are year-of-test and grade-of-test fixed effects, \(d_{ij}\) are application-cohort dummies (cohort membership determines the probability of winning, so cohort is an essential control), \(X_i\) is a vector of baseline demographics, and \(\theta\) is the causal effect of interest per year at KIPP. We write \(\theta\) rather than \(\rho\) (as in the original paper) to avoid clashing with the endogeneity correlation parameter of the same name used earlier in this chapter. The first-stage equation is \[s_{igt} = \lambda_t + \kappa_g + \sum_j \mu_j d_{ij} + \Gamma' X_i + \pi Z_i + \eta_{igt}. \tag{7.9}\] The model is just-identified (one excluded instrument per endogenous variable), so the 2SLS estimator of \(\theta\) equals the ratio of the reduced-form coefficient on \(Z_i\) to the first-stage coefficient \(\pi\).

IV estimates of KIPP Lynn attendance on MCAS scores (per year at KIPP). Source: Angrist et al. (2012), Table 4. Standard errors clustered at the student level in parentheses. \(N = 833\) student-by-test observations for both subjects. All regressions include year-of-test, grade-of-test, and application-cohort fixed effects, plus demographic and baseline-score controls.
Subject First stage Reduced form 2SLS
Math \(1.221\) \((0.068)\) \(0.430\) \((0.067)\) \(0.352\) \((0.053)\)
ELA \(1.228\) \((0.068)\) \(0.164\) \((0.073)\) \(0.133\) \((0.059)\)

Each year at KIPP raises math scores by approximately \(0.35\) standard deviations and ELA scores by approximately \(0.13\) standard deviations. The reduced-form estimate for math (\(0.43\sigma\)) is larger than the 2SLS estimate (\(0.35\sigma\)) because the first stage exceeds \(1\): lottery winners accumulated somewhat more than one additional year at KIPP per unit of follow-up time.

NoteRemark: Treatment Heterogeneity and the Weighted-Average Interpretation

When the treatment is continuous, Angrist and Imbens (1995) show that the IV estimand is a weighted average of the marginal causal effects of each additional unit of treatment, with weights proportional to the first-stage effect of the instrument at each dose level. In the KIPP setting, \(\theta\) is therefore a weighted average of the per-year effect across the different years a student might spend at KIPP, rather than the effect of a single fixed dose. The authors note this interpretation explicitly and adopt the conservative coding convention that partial years count as full years, which slightly attenuates the 2SLS estimates.

7.11.4 LATE Interpretation

The lottery design identifies the causal effect of KIPP attendance only for the students whose attendance behavior is changed by the lottery offer. These are the lottery compliers: students who attend KIPP if offered a seat and do not attend KIPP if not offered one. Compliance is partial in both directions: some lottery winners do not enroll and some lottery losers eventually find entry through later cohorts or attrition slots. Always-takers are largely ruled out at the extensive margin by the lottery’s control over seat access; never-takers remain — winners who decline the offer — and both margins reappear in the intensive dimension (years of enrollment).

This is the key interpretive point of the example. The 2SLS estimand \(\theta\) is not the effect of KIPP on all students in Lynn, nor on all applicants — it is the average per-year treatment effect for the lottery compliers, the specific margin of students whose enrollment decision is altered by randomized access to the school.

An important feature of this LATE is that it may differ from the effect the school would have on a randomly selected Lynn student who was not part of the applicant pool. KIPP applicants already had parents motivated enough to apply. The authors find that KIPP applicants have baseline test scores slightly lower than the district average, so the applicant pool is not positively selected on prior achievement, but it may still be selected on parental engagement in ways that affect how students respond to the KIPP environment.

The design is therefore strongest not because it identifies a universally generalizable effect, but because it identifies a clearly interpretable causal effect for a well-defined subpopulation under highly credible exogeneity.

7.11.5 Treatment Effect Heterogeneity

Subgroup analysis reveals that the LATE varies substantially across observable student characteristics. Reading gains (\(\approx 0.13\sigma\) overall) are driven almost entirely by students classified as having limited English proficiency (LEP, \(\approx 0.43\sigma\)) and special education needs (SPED, \(\approx 0.27\sigma\)); non-LEP, non-SPED students show negligible ELA gains. Math effects are large and positive across all subgroups but are largest for LEP and lower-achieving students. An interaction model that adds the product of baseline score with years at KIPP — identified by including \(Z_i \times \text{baseline score}\) as a second instrument — yields a significantly negative interaction term in both subjects: each additional standard deviation of baseline disadvantage is associated with an additional \(0.08\)\(0.17\sigma\) gain per year at KIPP.

These findings connect directly to Section 7.8.3: were we to define separate subgroup-specific instruments (LEP-lottery and non-LEP-lottery), each would identify a distinct subpopulation LATE. The overall 2SLS estimate is a weighted average of these subgroup LATEs, with weights proportional to each subgroup’s share of the complier population. This heterogeneity is not a nuisance detail; it is exactly why IV estimates must be interpreted together with the margin of compliance they capture.

NoteRemark: Lottery Design and Assumption Credibility

The KIPP study illustrates a general principle: instruments derived from designed randomization — lotteries, random assignment, randomized encouragement — provide the most transparent basis for the exogeneity assumption, because the distribution of \(Z\) is known by construction. The exclusion restriction still requires institutional argument (here, that the lottery offer affects only schooling, not morale or parental behavior in other dimensions), but exogeneity within the applicant sample, conditional on application cohort, is not merely plausible — it follows from the randomization protocol. This is why lottery-based IV designs occupy a privileged position in the program evaluation literature (Angrist and Pischke 2009).

The KIPP example ties together the main lessons of this chapter. All three IV assumptions are visible in the design: relevance is confirmed by a strong, precisely estimated first stage; exogeneity within the applicant sample follows from the randomization protocol; and exclusion rests on an institutional argument that the lottery offer affects achievement only through attendance. Because compliance is partial, the resulting estimate is a LATE for lottery compliers, not an ATE for all applicants or all students in Lynn. A lottery-based IV design is strongest not because it answers every causal question, but because it answers a clearly defined causal question for a clearly defined subpopulation.

7.12 Summary

  1. IV identifies effects from exogenous treatment variation. When back-door adjustment fails because an unobserved \(U\) creates a path \(T \leftarrow U \rightarrow Y\), a valid instrument \(Z\) identifies the causal effect by exploiting only the component of treatment variation that \(Z\) induces. IV does not block the confounding path — it avoids it by isolating exogenous variation and ignoring the rest.

  2. Three assumptions, ordered by testability. Relevance can be assessed with the first-stage \(F\)-statistic, though the \(F\)-statistic is a sample diagnostic, not the assumption itself. Exogeneity and exclusion are primarily substantive and must be defended by institutional knowledge, design logic, and causal structure. Each violated assumption produces a distinct, quantifiable bias in the Wald estimand, amplified by weak instruments (Section 7.4).

  3. Framework 1: homogeneous-effect SEM \(\Longrightarrow\) Wald identifies \(\beta\). Under constant treatment effects and the linear structural model, the Wald estimand identifies the single causal parameter \(\beta\), which equals the ATE, ATT, and LATE simultaneously. The identification follows from a three-step argument using the reduced form and first stage (Section 7.3).

  4. Framework 2: heterogeneity + monotonicity \(\Longrightarrow\) Wald identifies LATE. Under heterogeneous treatment effects and monotonicity, the Wald estimand identifies the average treatment effect for compliers only — those whose treatment status changes with the instrument, a latent subgroup defined by potential treatment statuses \((T(0), T(1))\), not by observable characteristics (Section 7.7).

  5. Different instruments identify different effects. The LATE depends on the instrument through the complier population it selects. Different instruments can legitimately identify different causal effects because they shift treatment for different margins of the population. LATE equals ATE only under additional structure, most notably treatment-effect homogeneity (Section 7.8).

  6. IV versus back-door adjustment. The two strategies have complementary failure modes: back-door adjustment fails when confounders are unobserved; IV fails when the exclusion restriction is violated or the instrument is not truly exogenous. IV identifies the LATE (not the ATE) under heterogeneous effects; back-door adjustment identifies the ATE or ATT. They are tools for different identification problems, not competitors (Section 7.9).

  7. Applied example: the KIPP Lynn lottery. Angrist et al. (2012) use a randomized admissions lottery as a canonical randomized-encouragement instrument for years of charter school attendance. Within the applicant sample, conditional on application cohort, random assignment of the offer supports exogeneity; exclusion is defended institutionally; relevance is confirmed by a first-stage coefficient of \(1.2\). The estimates identify the LATE for lottery compliers — \(0.35\sigma\) per year in math, \(0.13\sigma\) in ELA — with the largest gains for LEP, SPED, and low-baseline-score students (Section 7.11).

  8. Estimation deferred to Chapter 13. This chapter establishes what IV identifies and under what assumptions. How the Wald ratio is estimated from finite data — the reduced form regression, two-stage least squares, asymptotic inference, and overidentification tests — is the subject of Chapter 13.

7.13 Problems

1. The three IV assumptions in three languages. Consider the DAG \(\{Z \to T,\; T \to Y,\; U \to T,\; U \to Y,\; X \to T,\; X \to Y,\; X \to Z\}\) with \(U\) unobserved.

  1. List all back-door paths from \(T\) to \(Y\). Does \(X\) alone satisfy the back-door criterion? Explain.
  2. Verify the three IV assumptions using d-separation: (i) relevance: show \(Z\) and \(T\) are not d-separated in \(\Gcal\); (ii) exogeneity: show \(Z \indep U \mid X\) in \(\Gcal\); (iii) exclusion: show \(Y \indep Z \mid T, X\) in \(\Gcal_{\overline{T}}\).
  3. Now add the arrow \(Z \to Y\) to the DAG. Which IV assumption is violated? Show explicitly which step of the Wald derivation in Section 7.3 breaks down.
  4. Translate each of the three IV assumptions into the structural language: write the equations for \(T\) and \(Y\) and identify which coefficient restriction corresponds to each assumption.

2. Bias under assumption violations. Let \(Y = \beta T + \varepsilon\) and \(T = \pi Z + \eta\) with \(\E[\varepsilon \mid Z] = 0\) and \(\pi \neq 0\).

  1. Starting from \(\E[Y \mid Z{=}1] - \E[Y \mid Z{=}0]\), substitute the structural equation for \(Y\) and simplify. What role does exogeneity play?
  2. Show that \(\E[T \mid Z{=}1] - \E[T \mid Z{=}0] = \pi\) in the linear first-stage model. What role does relevance play?
  3. Derive the Wald estimand and confirm it equals \(\beta\).
  4. Now suppose the exclusion restriction fails and \(Y = \beta T + \delta Z + \varepsilon\) with \(\delta \neq 0\). Derive the probability limit of the Wald estimator. Confirm the bias formula from Section 7.4.
  5. Suppose instead that exogeneity fails: \(\E[\varepsilon \mid Z] = cZ\) for some constant \(c \neq 0\). Derive the probability limit of the Wald estimator and express the bias in terms of \(c\) and \(\pi\). Compare the structure of this bias with the exclusion violation bias.

3. Order, rank, and the limits of overidentification. Consider a model with one endogenous variable \(T\) and two instruments \(Z_1\) and \(Z_2\), both satisfying exogeneity and exclusion.

  1. State the order condition and verify it is satisfied.
  2. State the rank condition. What would it mean geometrically if the rank condition failed — i.e., if \(Z_1\) and \(Z_2\) were perfectly collinear in the first-stage regression?
  3. Explain intuitively why having two valid instruments rather than one should improve estimation precision.
  4. Now suppose \(Z_1\) is valid but \(Z_2\) violates the exclusion restriction. The Sargan–Hansen \(J\)-test is applied. Under what conditions does the \(J\)-test have power to detect \(Z_2\)’s invalidity? Under what conditions does the test fail to detect it?
  5. Why does passing the \(J\)-test not confirm that both \(Z_1\) and \(Z_2\) are valid? Give a concrete example of a situation in which both instruments are invalid and the \(J\)-test has no power.

4. Compliance types and the LATE. In a binary instrument, binary treatment study, suppose the population has the following composition: 30% compliers with average treatment effect \(\tau_c = 6\); 25% always-takers with \(\tau_a = 3\); 45% never-takers with \(\tau_n = 1\); no defiers.

  1. Compute \(P(\text{complier}) = \E[T \mid Z{=}1] - \E[T \mid Z{=}0]\).
  2. Compute the ATE as a weighted average of \(\tau_c\), \(\tau_a\), \(\tau_n\) with appropriate weights.
  3. The Wald estimand equals \(\tau_c = 6\). By how much does this overstate the ATE, and why?
  4. A second study uses a different binary instrument \(Z'\) with a complier population of 50% and a LATE of 2. Is this contradictory? What can you infer about the relative treatment effect in the two complier populations?
  5. Explain, using compliance type language, why the denominator of the Wald estimand equals \(P(\text{complier})\).

5. The exclusion restriction: plausibility and violations. Evaluate the exclusion restriction for each of the following proposed instruments. For each, (i) state whether the restriction is plausible and why; (ii) describe a specific mechanism by which it could be violated; and (iii) assess whether the violation would bias the IV estimate upward or downward.

  1. Instrument: rainfall in the home region of a politician, used as an instrument for government infrastructure spending. Outcome: local economic growth.
  2. Instrument: distance to the nearest hospital, used as an instrument for hospital admission. Outcome: 30-day mortality.
  3. Instrument: a randomly assigned financial incentive to enroll in a health screening program, used as an instrument for screening uptake. Outcome: health status two years later.
  4. Instrument: lottery number in the Vietnam-era draft lottery, used as an instrument for military service. Outcome: lifetime earnings. (This is the Angrist (1990) study; discuss why this instrument is widely regarded as satisfying the exclusion restriction.)

6. IV versus back-door adjustment. A researcher studies the effect of job training (\(T\)) on earnings (\(Y\)). Two strategies are available: (A) a rich set of pre-treatment covariates \(X\) and a propensity-score estimator; (B) a lottery that randomly selected units to be offered training (not required to attend), used as an instrument \(Z\).

  1. Under what assumption does strategy (A) identify the ATE? What specific unobserved variable would most plausibly violate this assumption?
  2. Strategy (B) identifies a LATE. Describe the complier population in words. Is the LATE likely to be larger or smaller than the ATE in this setting? Explain.
  3. Both strategies are implemented and yield estimates of $1,800 and $2,400 per year, respectively. Describe a Hausman-type test that uses both estimates. Under what null hypothesis does the test have an approximate \(\chi^2\) distribution?
  4. If the two estimates differ significantly, which strategy would you trust more and why? What additional evidence would help distinguish the two explanations (endogeneity bias in (A) versus LATE \(\neq\) ATE in (B))?

7. Nonparametric nonidentification of the ATE under valid IV assumptions (advanced / second pass). Consider the IV DAG \(Z \to T\), \(T \to Y\), \(U \to T\), \(U \to Y\), with \(U\) unobserved, \(Z \indep U\), and no direct \(Z \to Y\) edge, so that relevance, exogeneity, and exclusion all hold. This problem shows constructively that these assumptions do not nonparametrically identify the ATE, even with a genuinely relevant instrument, as the nonidentification remark opening Section 7.7 asserts.

  1. Construct two SEMs \(\mathcal{M}_1\) and \(\mathcal{M}_2\), each compatible with the IV graph, that are observationally indistinguishable but causally distinct: \[P_{\mathcal{M}_1}(Z, T, Y) = P_{\mathcal{M}_2}(Z, T, Y), \qquad P_{\mathcal{M}_1}\!\left(Y{=}1 \mid \doop(T{=}1)\right) \ne P_{\mathcal{M}_2}\!\left(Y{=}1 \mid \doop(T{=}1)\right).\] (Hint: let \(Z \sim \mathrm{Bern}(1/2)\) independently of a latent type \(U \in \{\mathrm{C}, \mathrm{A}, \mathrm{N}\}\) (complier, always-taker, never-taker) with \(P(\mathrm{C}) = 1/2\) and \(P(\mathrm{A}) = P(\mathrm{N}) = 1/4\), and in both models set \(T = Z\) for \(U = \mathrm{C}\), \(T = 1\) for \(U = \mathrm{A}\), and \(T = 0\) for \(U = \mathrm{N}\), so that \(P(T{=}1 \mid Z{=}1) - P(T{=}1 \mid Z{=}0) = 1/2\): the instrument is relevant. In both models set \(Y = T\) for \(U = \mathrm{C}\) and \(Y = 1\) for \(U = \mathrm{A}\); let the models differ only for \(U = \mathrm{N}\), with \(Y = 0\) in \(\mathcal{M}_1\) and \(Y = T\) in \(\mathcal{M}_2\).)
  2. Verify that both models generate the same joint \(P(Z, T, Y)\), then compute \(P(Y{=}1 \mid \doop(T{=}1))\) and \(P(Y{=}1 \mid \doop(T{=}0))\) in each model and show that the two ATEs differ.
  3. Verify that both models satisfy monotonicity (there are no defiers), compute the LATE \(\E[Y(1) - Y(0) \mid U = \mathrm{C}]\) in each model, and confirm that it coincides across the two models and equals the Wald ratio computed from the common observed distribution. What does this comparison show about which causal quantity the instrument identifies without further assumptions?
Angrist, Joshua D. 1990. “Lifetime Earnings and the Vietnam Era Draft Lottery: Evidence from Social Security Administrative Records.” American Economic Review 80 (3): 313–36.
Angrist, Joshua D., Susan M. Dynarski, Thomas J. Kane, Parag A. Pathak, and Christopher R. Walters. 2012. “Who Benefits from KIPP?” Journal of Policy Analysis and Management 31 (4): 837–60.
Angrist, Joshua D., and Guido W. Imbens. 1994. “Identification and Estimation of Local Average Treatment Effects.” Econometrica 62 (2): 467–75.
Angrist, Joshua D., and Guido W. Imbens. 1995. “Two-Stage Least Squares Estimation of Average Causal Effects in Models with Variable Treatment Intensity.” Journal of the American Statistical Association 90 (430): 431–42.
Angrist, Joshua D., Guido W. Imbens, and Donald B. Rubin. 1996. “Identification of Causal Effects Using Instrumental Variables.” Journal of the American Statistical Association 91 (434): 444–55.
Angrist, Joshua D., and Alan B. Krueger. 1991. “Does Compulsory School Attendance Affect Schooling and Earnings?” Quarterly Journal of Economics 106 (4): 979–1014.
Angrist, Joshua D., and Jörn-Steffen Pischke. 2009. Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press.
Bartik, Timothy J. 1991. Who Benefits from State and Local Economic Development Policies? W. E. Upjohn Institute for Employment Research.
Bound, John, David A. Jaeger, and Regina M. Baker. 1995. “Problems with Instrumental Variables Estimation When the Correlation Between the Instruments and the Endogenous Explanatory Variable Is Weak.” Journal of the American Statistical Association 90 (430): 443–50.
Card, David. 1995. “Using Geographic Variation in College Proximity to Estimate the Return to Schooling.” In Aspects of Labour Market Behaviour: Essays in Honour of John Vanderkamp, edited by Louis N. Christofides, E. Kenneth Grant, and Robert Swidinsky. University of Toronto Press.
Goldsmith-Pinkham, Paul, Isaac Sorkin, and Henry Swift. 2020. “Bartik Instruments: What, When, Why, and How.” American Economic Review 110 (8): 2586–624.
Hansen, Lars Peter. 1982. “Large Sample Properties of Generalized Method of Moments Estimators.” Econometrica 50 (4): 1029–54. https://doi.org/10.2307/1912775.
Hausman, Jerry A. 1978. “Specification Tests in Econometrics.” Econometrica 46 (6): 1251–71.
Hernán, Miguel A., and James M. Robins. 2006. “Instruments for Causal Inference: An Epidemiologist’s Dream?” Epidemiology 17 (4): 360–72.
Kitagawa, Toru. 2015. “A Test for Instrument Validity.” Econometrica 83 (5): 2043–63. https://doi.org/10.3982/ECTA11974.
Levis, Alexander W., Edward H. Kennedy, and Luke Keele. 2024. “Nonparametric Identification and Efficient Estimation of Causal Effects with Instrumental Variables.” arXiv Preprint arXiv:2402.09332.
Sargan, John D. 1958. “The Estimation of Economic Relationships Using Instrumental Variables.” Econometrica 26 (3): 393–415. https://doi.org/10.2307/1907619.
Wang, Linbo, and Eric Tchetgen Tchetgen. 2018. “Bounded, Efficient and Multiply Robust Estimation of Average Treatment Effects Using Instrumental Variables.” Biometrika 105 (2): 387–97.