9 Sensitivity Analysis and Partial Identification
Proofs, derivations, extended lab material, and additional problems marked here as optional are collected in the separate chapter supplement; nothing needed for a first reading depends on it.
How to read this chapter. The core first reading comprises all of Section 9.1 (uncertainty and the general framework), the bias decomposition of Section 9.2, the E-value (Section 9.3.2), benchmarking and reporting (Section 9.4), partial identification and overlap (Section 9.5), and the boundary with modern estimation (Section 9.7). The Rosenbaum \(\Gamma\)-model (Section 9.3.1) and the marginal sensitivity model (Section 9.3.3) are specialized method capsules; the IV and mediation material (Section 9.6) gives design-specific applications; and the lab (Section 9.8) is optional computational material.
9.1 Why Sensitivity Analysis?
Chapters 5–8 developed a sequence of identification strategies: back-door adjustment, propensity-score re-weighting, instrumental variables, and front-door and mediation analysis. Each strategy identifies a causal parameter as a functional of the observed-data distribution under a corresponding set of assumptions. Each set of assumptions is, however, stated in terms of unobserved quantities — potential outcomes, hypothetical interventions, or unmeasured variables — and cannot be verified from the data alone.
The chapter ahead (Chapter 10) opens by noting that identification does not by itself provide a statistically reliable estimator: once a parameter is identified, a separate theory of estimation and inference is still needed. The present chapter makes the complementary point, which closes out Part II:
Statistical reliability does not by itself validate identification.
A confidence interval reports sampling variation conditional on a maintained identification assumption. It is silent about whether that assumption is correct. A causal analysis is credible only if both components — identification and estimation — are credible. Sensitivity analysis is the tool for reporting how much the conclusion depends on the first.
9.1.1 Sampling Uncertainty vs. Identification Uncertainty
Consider an analyst who reports \(\hat\tau = 2.0\) with 95% CI \([1.2,\ 2.8]\). This interval answers one question and only one:
If the identifying assumptions are correct, what range of values of \(\tau\) is consistent with the observed sampling variation?
As \(n \to \infty\) the interval will shrink around \(\tau\), provided the identifying assumptions hold. If the assumptions do not hold, the interval will instead shrink around a biased target. More data drawn from the same observational regime do not remove this bias: the identification gap can only be closed by adding new assumptions, new design information, or new measurements relevant to the source of nonidentification — validation data, randomized encouragement, negative controls, instrumental variables, repeated measurements, or richer confounder measurements. A real-data illustration of the negative-control idea appears in the college-selectivity case study of Chapter 5: the average SAT score of the colleges that rejected a student — a variable with no causal path to the student’s earnings — strongly predicts earnings, directly exposing the confounding in the unadjusted comparison (Dale and Krueger 2002).
Formally, we may distinguish:
- Sampling uncertainty: \(\hat P_n \neq P\). The empirical distribution differs from the population. This uncertainty vanishes as \(n \to \infty\).
- Model misspecification: the working parametric or semiparametric model for a nuisance function (say the outcome regression \(\mu_t(x)\) or the propensity score \(\pi(x)\)) is not correct. This uncertainty is partly addressable by flexible modeling, doubly robust construction, and cross-fitting (Chapters 10–12).
- Identification uncertainty: the causal parameter \(\psi\) is not in fact identified by the functional \(\Psi(P)\) assumed by the analysis. This uncertainty does not vanish merely by increasing the sample size from the same observational regime.
The three sources of error compose. The usual confidence interval addresses only the first; the estimation chapters to come address primarily the second; this chapter addresses the third.
9.1.2 The Role of Sensitivity Analysis
Sensitivity analysis is a structured way to vary the strength of violations of identifying assumptions and to report how the causal conclusion responds. It is not a test of the assumptions. Identifying assumptions are generally not fully verifiable from the observed outcome data, although some of their observed-data implications can be assessed — a proposed DAG may imply testable conditional independences, overidentified models imply moment restrictions, negative controls can falsify aspects of a design, known randomization protocols supply design information, and positivity can be partly diagnosed empirically. It is a statement of conditional robustness: how far can the assumptions be violated before the substantive conclusion changes?
Three interpretations of the same analysis are useful, and this chapter develops each:
- Sensitivity curve. A plot of the estimand against a scalar sensitivity parameter \(\lambda\), with the neutral value \(\lambda = \lambda_0\) recovering the baseline identifying assumption.
- Sensitivity bounds. The range of estimand values consistent with any \(\lambda\) in a specified plausibility set \(\mathcal{L}\).
- Tipping point. The smallest violation magnitude sufficient to change the sign, magnitude, or statistical significance of the conclusion.
9.1.3 The General Framework
The three views of sensitivity analysis — curves, bounds, tipping points — all arise from a single construction. Let \(\psi\) be the causal estimand of interest, and let \(A_0\) be a baseline identifying assumption such that, under \(A_0\), \(\psi = \Psi(P)\) for a known functional \(\Psi\). A sensitivity model is an indexed relaxation \(\{A_\lambda : \lambda \in \mathcal{L}\}\) of \(A_0\), with \(A_0\) recovered at a neutral value \(\lambda_0 \in \mathcal{L}\). In the simplest, point-valued case, the estimand under \(A_\lambda\) is a function of both the observed-data distribution and the sensitivity parameter: \[\psi(\lambda) = \Psi(P;\, \lambda), \qquad \Psi(P;\, \lambda_0) = \Psi(P); \tag{9.1}\] the general case, in which a fixed \(\lambda\) pins \(\psi\) down only to a set, is developed in Section 9.5. The key reporting object is therefore not a single number but the attainable set \[\{\psi(\lambda) : \lambda \in \mathcal{L}\}. \tag{9.2}\]
Across this chapter we will encounter several sensitivity parameters, each quantifying the magnitude of a specific violation:
- \(\lambda\) = strength of unmeasured confounding, parameterized either as an odds-ratio bound, as a pair of association strengths in the VanderWeele–Ding E-value framework, or as an odds-ratio bound in the marginal sensitivity model.
- \(\delta\) = magnitude of a structural outcome-effect coefficient that the baseline identifying assumption treats as inert. Used in two contexts: the \(U\)-outcome effect in the linear and binary-confounder sensitivity models (Section 9.2.2–Section 9.2.3); and the direct effect of an instrument on the outcome violating the IV exclusion restriction (Section 9.6), where \(\delta = 0\) is the baseline assumption. The two roles never appear together.
- \(\alpha\) = propensity-score trimming threshold, addressing near-violations of positivity (Section 9.5.6).
- \(\rho\) = residual correlation between mediator and outcome disturbances, violating sequential ignorability (Section 9.6.4).
In each case the neutral value returns the baseline identifying assumption — \(\Gamma = 1\) and \(\Lambda = 1\) for the odds-ratio confounding bounds, and \(\delta = 0\), \(\alpha = 0\), \(\rho = 0\) for the remaining parameters — and the distance from the neutral value quantifies the strength of the violation.
One entry in this list is of a different kind. The trimming threshold \(\alpha\) does not relax an identifying assumption while holding the causal estimand fixed; it indexes the retargeted estimands \(\tau_\alpha = \E\{Y(1) - Y(0) \mid \alpha \leq \pi(X) \leq 1-\alpha\}\). We therefore treat trimming as target sensitivity — sensitivity of the conclusion to the choice of population — rather than as an instance of the fixed-estimand sensitivity model.
A related notational point concerns the two capital lambdas. We write \(\mathcal{L}\) for the plausibility set in which the sensitivity parameter is assumed to lie, and reserve the upright \(\Lambda\) for the scalar level of the marginal sensitivity model (Section 9.3.3), where it is standard in the literature (Tan 2006; Zhao et al. 2019). The two are different objects: \(\mathcal{L}\) is a set of admissible parameter values, while \(\Lambda\) is one particular sensitivity parameter, whose own plausibility set is an interval \([1, \Lambda_{\max}] \subseteq \mathcal{L}\).
The symbol \(\alpha\) appears elsewhere in this chapter in two unrelated roles: as a regression intercept in the linear outcome and IV models, and, in the lab, as the subscripted parameter \(\alpha_U\) denoting the strength of \(U\) in the treatment mechanism. Similarly, in Section 9.6.3 the symbols \(\pi_c\) and \(\pi_d\) denote complier and defier population shares; they are unrelated to the propensity score \(\pi(x)\).
Definition 9.1 (Sensitivity Model) A sensitivity model for the causal estimand \(\psi\) and baseline identifying assumption \(A_0\) is a triple \((\{A_\lambda\}_{\lambda \in \mathcal{L}},\, \lambda_0,\, \Psi(\cdot;\cdot))\) such that (i) \(\lambda_0 \in \mathcal{L}\) is a neutral value at which \(A_{\lambda_0} = A_0\), recovering the original identifying assumption; and (ii) under \(A_\lambda\) and the observable restrictions imposed by the data, \(\psi = \Psi(P;\lambda)\). The parameter \(\lambda\) may be scalar- or vector-valued.
Several of the models in this chapter pin \(\psi\) down only to a set at each \(\lambda\) — the MSM at fixed \(\Lambda\) delivers an interval for the estimand, a Rosenbaum analysis at fixed \(\Gamma\) bounds a randomization \(p\)-value, and the bias factor of Section 9.3.2 is two-dimensional before its scalar summary. The set-valued extension of this framework is developed in Section 9.5, where it merges naturally with partial identification.
Two further comments on the definition. First, the neutral value is not always zero. It is \(\lambda_0 = 0\) for the additive parameters used in this chapter (\(\delta\), \(\rho\), \(\alpha_U\)) but \(\lambda_0 = 1\) for the multiplicative ones (\(\Gamma\), \(\Lambda\)), since those are odds-ratio bounds for which no violation means a ratio of one. Writing \(\lambda_0\) rather than \(0\) keeps the framework honest about this.
Second, continuity of \(\Psi(\cdot;\lambda)\) in \(\lambda\) is deliberately not part of the definition. It is a convenient regularity condition — it makes sensitivity curves drawable and tipping points well defined — and it holds for most models in this chapter. But bound endpoints obtained from an optimization can be nonsmooth, and a sensitivity region can change discretely as constraints become active, so continuity should be imposed where it is needed rather than assumed throughout.
9.1.3.1 Sensitivity Curves
When \(\mathcal{L} \subseteq \mathbb{R}\) is one-dimensional, the natural reporting object is the sensitivity curve \(\lambda \mapsto \hat\psi(\lambda)\), where \(\hat\psi(\lambda)\) is a consistent estimator of \(\psi(\lambda)\) under \(A_\lambda\). A sensitivity curve is useful when the reader wishes to see the effect smoothly deform as the assumption is relaxed. Many sensitivity curves are monotone — the bias moves in one direction as \(\lambda\) increases — though this is not guaranteed.
9.1.3.2 Sensitivity Bounds
When multiple parameters index the violation, or when the sensitivity parameter is high-dimensional, reporting a full curve is infeasible. The natural alternative is to report the extremes over \(\mathcal{L}\). One caveat should be recorded first: because continuity is not part of Definition 9.1, the image \(\{\psi(\lambda) : \lambda \in \mathcal{L}\}\) need not be an interval. The bounds below are therefore an enclosure of that set; intermediate values inside \([\psi_L, \psi_U]\) need not be attainable by any admissible \(\lambda\). With that caveat recorded, define \[\psi_L = \inf_{\lambda \in \mathcal{L}} \psi(\lambda), \qquad \psi_U = \sup_{\lambda \in \mathcal{L}} \psi(\lambda). \tag{9.3}\] The interval \([\psi_L,\psi_U]\) is the sensitivity interval. In practice, one estimates the lower and upper bounds and reports them with sampling-based uncertainty intervals; we caution that inference for the endpoints of a partially identified set is more delicate than inference for an ordinary scalar parameter, and the appropriate construction depends on whether the target is the identified set or the parameter itself.
9.1.3.3 Tipping-Point Analysis
A third summary is the smallest \(\lambda\) at which the substantive conclusion changes. If the conclusion of interest is “the effect is positive,” the natural tipping point is \[\lambda^\star = \inf\{\lambda \in \mathcal{L} : \psi(\lambda) \leq 0\}, \tag{9.4}\] or, if inference is at issue rather than the point estimate, \[\lambda^\star_{\mathrm{CI}} = \inf\{\lambda \in \mathcal{L} : 0 \in \hat{\mathrm{CI}}(\lambda)\}, \tag{9.5}\] where \(\hat{\mathrm{CI}}(\lambda)\) is the confidence interval for \(\psi(\lambda)\). Note that \(\psi(\lambda)\) in Equation 9.4 is the pointwise value at \(\lambda\), distinct from the global bound \(\psi_L\) of Equation 9.3.
Three cautions attach to Equation 9.4. First, the definition is deliberately directional: it is stated for the conclusion “the effect is positive,” and the mirror image applies for a negative one. Second, writing \(\psi(\lambda) \leq 0\) rather than \(\psi(\lambda) = 0\) avoids requiring an exact crossing, which need not exist once continuity is not assumed. Third, and most important, the infimum presupposes an ordered, nested scalar scale: \(\lambda' > \lambda\) must mean that \(A_{\lambda'}\) is a strictly weaker assumption, so that larger \(\lambda\) is unambiguously “more violation.” For a vector-valued \(\lambda\), or a scalar one whose sensitivity sets are not nested, there is no unique smallest violation until one fixes a norm, a path through the parameter space, or a partial order — and that choice is a substantive modeling decision, not a technicality. If \(\psi\) is not monotone, \(\lambda^\star\) is only the first crossing and should be reported as such.
Example 9.1 (Tipping-Point Reading) Suppose \(\hat\psi(0) = 1.8\) with 95% CI \([1.1, 2.5]\), and the sensitivity model produces \(\hat\psi(\lambda) = 1.8 - 2.0\lambda\). Then:
- The sensitivity curve is linear in \(\lambda\).
- The tipping point for the sign is \(\lambda^\star = 0.9\).
- The tipping point for significance is \(\lambda^\star_{\mathrm{CI}}\), the value at which the lower endpoint of the CI — which in this simple linear model equals \(1.1 - 2.0\lambda\) — hits zero, so \(\lambda^\star_{\mathrm{CI}} = 0.55\).
The significance tipping point is always at least as strict as the sign tipping point. An analysis is robust “at level \(\lambda^\star\)” for a conclusion that holds throughout \([0, \lambda^\star)\).
9.2 Sensitivity to Unmeasured Confounding
The most common application of sensitivity analysis is to unmeasured confounding under the back-door framework of Chapter 5. The baseline identifying assumption is conditional exchangeability, \(Y(t) \indep T \mid X\), together with consistency and positivity. Under these, \[\tau = \E\{\mu_1(X) - \mu_0(X)\}, \qquad \mu_t(x) = \E(Y \mid T=t, X=x). \tag{9.6}\] Sensitivity analysis here contemplates a world in which there exists an unobserved \(U\) such that \[Y(t) \indep T \mid X, U, \qquad\text{but}\qquad Y(t) \nindep T \mid X. \tag{9.7}\] In words: ignorability would hold if \(U\) were observed, but \(U\) is not observed, so the working identification formula Equation 9.6 is biased. The structure is the familiar back-door path \(T \leftarrow U \rightarrow Y\) that adjustment for \(X\) alone cannot block.
9.2.1 The Bias Decomposition
Let \[\Delta_{\mathrm{obs}} = \E\{\E(Y \mid T=1, X) - \E(Y \mid T=0, X)\} \tag{9.8}\] denote the observed adjusted contrast — the quantity that a back-door analysis actually computes when only \(X\) is adjusted for. Under ignorability conditional on \(X\) alone, \(\Delta_{\mathrm{obs}} = \tau\). Under the relaxed condition Equation 9.7, the observed adjusted contrast departs from \(\tau\) by a quantity that depends on how \(U\) is distributed conditional on \((T, X)\).
Theorem 9.1 (Bias Decomposition under Unmeasured Confounding) Suppose Equation 9.7 holds along with consistency and positivity given \((X, U)\), that is \(P(T=t \mid X, U) > 0\) almost surely for \(t \in \{0,1\}\) — the stronger requirement needed for \(m_t\) below to be well defined, and not implied by positivity given \(X\) alone. Let \(f(u \mid \cdot)\) denote the conditional density (or mass function) of \(U\), define \[m_t(x, u) = \E(Y \mid T=t, X=x, U=u),\] and the distributional difference \[g_t(x, u) = f(u \mid T=t, X=x) - f(u \mid X=x).\] Then \(\Delta_{\mathrm{obs}} = \tau + B\), where the confounding bias is \[B = \E\!\left\{\sum_t (-1)^{1-t} \int m_t(X, u)\, g_t(X, u)\, du\right\}, \tag{9.9}\] with the integral replaced by a sum if \(U\) is discrete.
Proof sketch. Condition on \((X, U)\); conditional exchangeability given \((X, U)\) identifies \(m_t(X, u)\) from the corresponding treatment arm, and the observed contrast differs from the causal one exactly through the distributional differences \(g_t\), which yields Equation 9.9 upon rearranging. The full derivation is in the optional supplement. \(\square\)
Equation Equation 9.9 is the master formula underlying most sensitivity models for unmeasured confounding. Each model amounts to a parameterization of the two quantities \(m_t(x, u)\) and \(g_t(x, u)\): the effect of \(U\) on \(Y\), and the imbalance of \(U\) across treatment arms given \(X\).
9.2.2 A Simple Linear Sensitivity Model
A closed-form version of Equation 9.9 is available in a fully linear model, which is the form most often encountered in econometrics and the one easiest to benchmark. Suppose the outcome follows \[Y = \alpha + \tau T + \gamma^\top X + \delta U + \varepsilon, \qquad \E(\varepsilon \mid T, X, U) = 0, \tag{9.10}\] and the treatment assignment depends on \((X, U)\) through \(T = h(X, U, \eta)\) with \(\eta \indep (U, \varepsilon)\).
Lemma 9.1 (Linear Sensitivity Bias) Under the model above and conditional exchangeability given \((X, U)\), \[B = \delta \cdot \E\!\left\{\E(U \mid T=1, X) - \E(U \mid T=0, X)\right\}. \tag{9.11}\]
Proof. Specializing Equation 9.9 to the linear outcome model, \(m_t(x, u) = \alpha + \tau t + \gamma^\top x + \delta u\), so the inner integrals reduce to \(\delta \int u\, g_t(X, u)\, du\), because the components of \(m_t\) that do not depend on \(u\) multiply integrals of \(g_t(X, u)\), which is \(0\) by construction. Combining, \[B = \delta\,\E\!\left\{\int u\, [g_1(X, u) - g_0(X, u)]\, du\right\} = \delta\,\E\{\E(U \mid T=1, X) - \E(U \mid T=0, X)\}. \qquad\square\]
Formula Equation 9.11 is the canonical teaching decomposition. Under the linear outcome model Equation 9.10 it is an exact identity, not an approximation; approximate language is appropriate only when Equation 9.10 is being used as a local working approximation to a nonlinear outcome model: \[\text{bias} \;=\; \underbrace{\delta}_{U\text{-outcome effect}} \;\times\; \underbrace{\E\{\E(U \mid T=1, X) - \E(U \mid T=0, X)\}}_{U\text{-treatment imbalance}}. \tag{9.12}\] Both factors are sensitivity parameters because neither is identified: \(U\) is unobserved, so its effect on \(Y\) and its imbalance across treatment arms can only be posited, not estimated.
9.2.3 Binary Unmeasured Confounder
A useful concrete case arises when \(U \in \{0, 1\}\). Define \(p_1(x) = P(U=1 \mid T=1, X=x)\), \(p_0(x) = P(U=1 \mid T=0, X=x)\), and \(\delta(t, x) = m_t(x, 1) - m_t(x, 0)\), the \(U\)-outcome contrast within the cell defined by \((T, X)\).
Lemma 9.2 (Binary-Confounder Bias) If \(U \in \{0,1\}\) and the \(U\)-outcome contrast is constant across \((t, x)\) at value \(\delta\), then \[B = \delta \cdot \E\{p_1(X) - p_0(X)\}. \tag{9.13}\]
Proof. With \(U \in \{0,1\}\), \(m_t(x, u) = m_t(x, 0) + u\,\delta(t, x)\). The constant-contrast assumption gives \(\delta(t,x) = \delta\). Applying Equation 9.9 and using that \(g_t(x, u)\) integrates to \(0\) in \(u\), \[\int m_t(X, u)\, g_t(X, u)\, du = \delta\,[P(U{=}1 \mid T{=}t, X) - P(U{=}1 \mid X)],\] so \(B = \delta\,\E\{p_1(X) - p_0(X)\}\). \(\square\)
Formula Equation 9.13 is the workhorse of applied sensitivity analysis under the back-door design. It has two sensitivity parameters — the \(U\)-outcome effect \(\delta\) and the mean \(U\)-treatment imbalance — each of which has a concrete interpretation. A sensitivity analysis then consists of computing \(\hat\tau_{\mathrm{sens}} = \hat\Delta_{\mathrm{obs}} - \hat B\) over a grid of values and reporting the resulting surface, contours, tipping points, or bounds.
9.2.4 A Sensitivity Table
For reporting, a simple table tabulates the adjusted effect over a grid of sensitivity parameters, in the style of Rosenbaum (2002) and VanderWeele and Ding (2017), for an observed adjusted effect \(\hat\Delta_{\mathrm{obs}} = 2.0\).
| \(U\)-outcome effect \(\delta\) | \(p_1 - p_0 = 0.10\) | \(0.30\) | \(0.50\) |
|---|---|---|---|
| \(\delta = 2\) | 1.80 | 1.40 | 1.00 |
| \(\delta = 4\) | 1.60 | 0.80 | 0.00 |
| \(\delta = 8\) | 1.20 | \(-0.40\) | \(-2.00\) |
The table reads straightforwardly: the effect remains clearly positive when the posited \(U\)-outcome effect is small relative to the outcome scale, is substantially attenuated once \(\delta\) and the imbalance are both moderate on that scale, and reverses when both are large. (Verbal labels such as “weak” or “strong” have no meaning without specifying the scale of \(Y\), which is why the rows are indexed by \(\delta\) itself.) The tipping point for the sign is not a single cell but the contour \(\delta \cdot (p_1 - p_0) = 2.0\), a hyperbola in the \((\delta,\, p_1-p_0)\) plane. The grid lands exactly on that contour at \((\delta = 4,\, p_1 - p_0 = 0.5)\); the cell \((\delta = 8,\, p_1 - p_0 = 0.3)\) lies just beyond it. The analyst then asks whether an unmeasured \(U\) anywhere near this strong is scientifically plausible — a question addressed through benchmarking in Section 9.4.
9.3 Three Canonical Sensitivity Models
The bias decomposition of Section 9.2.1 is a general identity. Three canonical models specialize it by imposing a structural restriction on either \(g_t\) (the \(U\)-treatment mechanism) or on the induced observed-data relationships. Each model is a named instance of the framework of Section 9.1.3 with a specific choice of sensitivity parameter.
We present the three models in the sequence of matched-design sensitivity analysis, risk-ratio-scale sensitivity summaries, and weighting-based sensitivity intervals. This is a sequence of settings, not an ordering by generality — the E-value is not tied to an estimator in the way the other two are.
9.3.1 Rosenbaum’s \(\Gamma\)-Sensitivity Model
Rosenbaum (2002) formulated sensitivity analysis for matched observational studies in terms of an odds-ratio bound on the conditional treatment probabilities. Define the conditional treatment odds given \((X, U)\) and the sensitivity parameter \[\Gamma = \sup_{x, u, u'} \frac{\mathrm{odds}(T=1 \mid X=x, U=u)}{\mathrm{odds}(T=1 \mid X=x, U=u')}. \tag{9.14}\] In words, \(\Gamma\) bounds the factor by which any two units with the same observed covariates can differ in their odds of treatment because of unobserved \(U\).
The parameter \(\Gamma\) is the sensitivity parameter in the sense of Definition 9.1, with \(\lambda = \Gamma\), plausibility set \(\mathcal{L} = [1, \infty)\), and neutral value \(\lambda_0 = 1\). A value \(\Gamma = 2\) states that units with the same \(X\) may differ by up to a factor of 2 in their treatment odds because of unmeasured \(U\).
Under the \(\Gamma\)-model — testing Fisher’s sharp null of no effect (or a specified constant additive effect), with assignments independent across matched pairs — worst-case null distributions for matched-pair test statistics are obtained by pushing every pair-specific treatment probability to the endpoints \(1/(1+\Gamma)\) and \(\Gamma/(1+\Gamma)\) permitted by Equation 9.15. For the sign test the resulting bounds are binomial; for Wilcoxon’s signed-rank and other signed-score statistics they are distributions of rank-weighted sums of independent Bernoulli variables. The formal statement and its proof are in the optional supplement.
The practical outcome of a \(\Gamma\)-analysis is a tipping \(\Gamma^\star\): the smallest \(\Gamma\) at which the test of no effect fails to reject. This value is not a property of the data alone. It is defined only relative to a stated test statistic, a stated worst-case direction, a one- or two-sided alternative, and a significance level; changing any of these changes \(\Gamma^\star\), and all four should be reported alongside it. Small \(\Gamma^\star\) (near 1) indicates a fragile conclusion; large \(\Gamma^\star\) indicates a robust one.
9.3.2 The E-Value and the VanderWeele–Ding Bound
Ding and VanderWeele (2016) and VanderWeele and Ding (2017) proposed a sensitivity summary that requires no matched design, takes the form of a single closed-form number, and can be read off a reported risk ratio, subject to the scope restrictions recorded at the end of this subsection. It has become the most widely reported sensitivity summary in the applied literature.
The setup uses two risk ratios as sensitivity parameters. Within a covariate stratum \(x\), define \[\mathrm{RR}_{TU}(x) = \sup_u \frac{dP(u \mid T=1, X=x)}{dP(u \mid T=0, X=x)}, \tag{9.16}\] \[\mathrm{RR}_{UY}(x) = \max_{t \in \{0,1\}} \frac{\sup_u P(Y=1 \mid T=t, X=x, U=u)}{\inf_u P(Y=1 \mid T=t, X=x, U=u)}. \tag{9.17}\] Write \(\mathrm{RR}_{TU}\) and \(\mathrm{RR}_{UY}\) for the maxima of these over \(x\). These are, respectively, the maximum relative-risk association of \(U\) with \(T\) and the maximum relative-risk association of \(U\) with \(Y\) given \((T, X)\); both are at least \(1\).
The force of the Ding–VanderWeele result is that with the parameters defined this generally, the bias-factor bound below requires no further structure: \(U\) need not be binary, need not be a single variable, and may interact with \(T\) in its effect on \(Y\). That generality is why the E-value is usable in practice, where nothing is known about \(U\). A transparent special-case derivation (binary \(U\), no \(T \times U\) interaction) appears in the optional supplement.
Let \(\mathrm{RR}_{TY \mid X}^{\mathrm{obs}}(x)\) denote the observed risk ratio within stratum \(X = x\), and \(\mathrm{RR}_{TY \mid X}^{\mathrm{true}}(x)\) the corresponding causal risk ratio. Define the bias factor \[B(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}) = \frac{\mathrm{RR}_{TU}\, \mathrm{RR}_{UY}}{\mathrm{RR}_{TU} + \mathrm{RR}_{UY} - 1}. \tag{9.18}\]
Theorem 9.2 (VanderWeele–Ding Bound) Assume consistency; conditional exchangeability given the full confounder set, \(\{Y(0), Y(1)\} \indep T \mid X, U\); positivity of \(P(T=t \mid X, U)\); and the binary-outcome risk-ratio setup above, with \(P(U \mid T{=}1, X{=}x)\) absolutely continuous with respect to \(P(U \mid T{=}0, X{=}x)\) so that Equation 9.16 is finite. Then for every \(x\), \[\mathrm{RR}_{TY \mid X}^{\mathrm{true}}(x) \;\geq\; \frac{\mathrm{RR}_{TY \mid X}^{\mathrm{obs}}(x)}{B(\mathrm{RR}_{TU}, \mathrm{RR}_{UY})}. \tag{9.19}\] Write \(R = \mathrm{RR}_{TY \mid X}^{\mathrm{obs}}(x) \geq 1\) and define the E-value \[\mathrm{EV} = R + \sqrt{R\,(R - 1)}. \tag{9.20}\] Then \(\mathrm{EV}\) is the smallest common value \(e\) such that the symmetric configuration \(\mathrm{RR}_{TU} = \mathrm{RR}_{UY} = e\) drives the bound Equation 9.19 to the null. Equivalently, \[\mathrm{EV} = \min\big\{\max(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}) : B(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}) \geq R\big\}. \tag{9.21}\] Consequently, for the observed risk ratio to be explained away completely, both associations must be at least \(R\) and at least one must be at least \(\mathrm{EV}\): \[\min(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}) \geq R, \qquad \max(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}) \geq \mathrm{EV}. \tag{9.22}\]
A transparent closed-form derivation of Equation 9.19, the E-value formula, the min–max characterization, and the Cornfield pair — for the special case of a binary \(U\) with no \(T \times U\) interaction — is given in the optional supplement.
Example 9.2 (An E-Value Calculation) An observational study reports an adjusted risk ratio of \(2.5\) with 95% CI \([1.7, 3.7]\). The E-value for the point estimate is \[\mathrm{EV} = 2.5 + \sqrt{2.5 \cdot 1.5} = 2.5 + \sqrt{3.75} \approx 4.44.\] The E-value for the lower confidence limit is \(1.7 + \sqrt{1.7 \cdot 0.7} \approx 2.79\).
Interpretation: to reduce the point estimate to the null, an unmeasured confounder must have both risk-ratio associations — with treatment, and with the outcome given treatment — at least \(2.5\), and at least one of them at least \(4.44\). The symmetric configuration \(\mathrm{RR}_{TU} = \mathrm{RR}_{UY} = 4.44\) is sufficient but not necessary: the asymmetric pair \((3.00,\, 10.00)\) also explains the point estimate away. For the lower confidence limit the corresponding thresholds are \(1.7\) and \(2.79\).
Scope of the E-value formula. Equation Equation 9.20 is stated for a risk ratio \(R \geq 1\). For a protective effect, \(R < 1\), the E-value is computed from \(1/R\) and interpreted on that scale. For odds ratios and hazard ratios the formula is not automatically exact: an odds ratio approximates the corresponding risk ratio only under a rare-outcome condition, and a hazard ratio carries additional causal interpretive difficulties (selection over time among survivors) that the bias factor of Theorem 9.2 does not address. Applying the E-value to these measures requires the conversions and caveats given by VanderWeele and Ding (2017), not a direct substitution into Equation 9.20.
9.3.3 The Marginal Sensitivity Model
The \(\Gamma\)-model is formulated in terms of the unobservable conditional distribution \(P(T \mid X, U)\); the E-value collapses the \(U\)-treatment and \(U\)-outcome associations to a symmetric scalar. Tan (2006) proposed an alternative formulation — the marginal sensitivity model (MSM) — that bounds the violation directly by an odds ratio comparing the nominal propensity \(P(T=1 \mid X)\) with the latent propensity \(P(T=1 \mid X, U)\), making it natural to combine with weighting estimators in Chapters 10–11.
Define the nominal and true propensity scores \(\pi(x) = P(T=1 \mid X=x)\) and \(\pi^{\mathrm{true}}(x, u) = P(T=1 \mid X=x, U=u)\).
The parameter \(\Lambda\) bounds, on the odds scale, how far the covariate-conditional propensity can depart from the true propensity once \(U\) is marginalized. Read literally, \(\Lambda = 1\) says only that the represented \(U\) does not alter the treatment odds within levels of \(X\); under the maintained latent exchangeability of Theorem 9.3 below, this eliminates the additional treatment-selection dependence on \(U\) and thereby recovers conditional exchangeability given \(X\) alone.
Theorem 9.3 (MSM Sensitivity Interval for the ATE) Assume consistency; latent conditional exchangeability \(\{Y(0), Y(1)\} \indep T \mid X, U\); positivity of the true propensity; suitable integrability conditions; and the marginal sensitivity model Equation 9.23 at level \(\Lambda\). Then the average treatment effect satisfies \(\tau_L(\Lambda) \leq \tau \leq \tau_U(\Lambda)\), where \((\tau_L, \tau_U)\) are functionals of the observed-data distribution obtained by optimizing the weighting estimand over a valid outer envelope of weight perturbations implied by Equation 9.23. The interval is valid in the sense of Definition 9.2: it contains \(\tau\) whenever the model holds. It is not claimed to be sharp.
The interval is computed empirically by optimizing the self-normalized (Hájek) weighting estimator over a valid outer envelope of weight perturbations — a linear-fractional program whose solution is a threshold rule in the ordered outcomes, which is why it takes a percentile form — with percentile-bootstrap inference (Zhao et al. 2019). The derivation of the perturbation envelope and the optimization details are in the optional supplement.
9.3.4 Comparing the Three Models
The three models share the general logic of indexed assumption relaxation. The MSM and the E-value bias-factor model fit the causal identified-set formulation directly; Rosenbaum’s usual \(p\)-value analysis is its inferential analogue and fits the estimand-set formulation after test inversion (Section 9.5). They differ in (i) what object the sensitivity parameter bounds, (ii) what estimator they naturally pair with, and (iii) how the sensitivity parameter is benchmarked.
| Rosenbaum \(\Gamma\) | E-value | MSM (\(\Lambda\)) | |
|---|---|---|---|
| Parameter bounds | Odds ratio of treatment given \(X, U\) | Pair of risk-ratio associations \(\mathrm{RR}_{TU}, \mathrm{RR}_{UY}\) | Odds ratio of nominal to true propensity |
| Natural design or estimator setting | Matched-pair and signed-score analyses | Estimator-agnostic risk-ratio-scale effect estimate | IPW (augmented extensions cited separately) |
| Typical sensitivity output | Tipping \(\Gamma^\star\) | \(\mathrm{EV}\) (symmetric diagonal) | Sensitivity interval \([\tau_L, \tau_U]\) or tipping \(\Lambda^\star\) |
| Primary reference | Rosenbaum (2002) | VanderWeele and Ding (2017) | Tan (2006); Zhao et al. (2019) |
9.4 Benchmarking and Reporting
Every sensitivity model in the previous section has one or more parameters that are not identified by the data. A sensitivity curve of the form “at \(\lambda = 0.2\) the effect is halved” is mathematically precise but scientifically empty until \(\lambda = 0.2\) has been anchored to something concrete. Benchmarking is the discipline of giving sensitivity parameters an empirical referent by comparing them to the analogous quantities computed for the observed covariates (Cinelli and Hazlett 2020).
9.4.1 Why Sensitivity Parameters Need Interpretation
Consider a statement that would otherwise close a sensitivity analysis: “the conclusion is robust up to \(\Gamma = 1.8\).” To evaluate this statement scientifically, one must know whether an unmeasured confounder inducing a treatment-odds ratio of \(1.8\) is a plausible hypothesis in the particular study, or an implausible one. The same \(\Gamma\)-value can be a comfortable margin in one application and an alarming fragility in another.
Similarly, “the E-value equals 3.5” means that, to explain the effect away, an unmeasured confounder must have risk-ratio associations with \(T\) and with \(Y\) that are both at least the observed risk ratio \(R\), with at least one of them at least \(3.5\). Whether a confounder of this magnitude is plausible depends on the context, and the benchmark must name its scale: in a study where the strongest observed covariate has an equivalent E-value of \(2.0\), an \(\mathrm{EV} = 3.5\) is reassuring; in a study where measured covariates routinely have equivalent E-values of \(4\) or more, it is not. Quoting a covariate’s raw bias factor, or one of its two associations alone, in place of its equivalent E-value would compare different objects.
9.4.2 Benchmarking Against Observed Covariates
The core idea of benchmarking is to compute the sensitivity-parameter values implied by the observed covariates and to compare them to the proposed range for an unmeasured confounder.
For the binary-confounder setup, the sensitivity parameters are the \(U\)-outcome effect \(\delta\) and the \(U\)-treatment imbalance \(p_1(X) - p_0(X)\). For each observed covariate \(X_j\), one can compute the analogous quantities. The treatment-side benchmark must match the estimator. If the primary estimator is OLS adjustment, the exact analogue is the Frisch–Waugh–Lovell auxiliary coefficient \[d_j^{\mathrm{OLS}} = \text{coefficient on } T \text{ in the regression of } X_j \text{ on } (1, T, X_{-j}),\] so the covariate’s implied bias is \(\hat\gamma_j\,d_j^{\mathrm{OLS}}\). If the primary estimator is a standardized conditional-mean contrast, the matching benchmark is instead \(\E_{X_{-j}}\{\E(X_j \mid T{=}1, X_{-j}) - \E(X_j \mid T{=}0, X_{-j})\}\). The two need not coincide — the lab in Section 9.8 makes exactly this distinction for the unmeasured confounder itself, and the observed-covariate benchmark must respect it for the same reason.
For the E-value, benchmarking is particularly clean. For each observed covariate \(X_j\), compute its risk-ratio association with \(T\) (call it \(\mathrm{RR}_{TX_j}\)) and with \(Y\) (adjusted for \(T\) and the other covariates, call it \(\mathrm{RR}_{X_j Y}\)). The corresponding bias factor is \[B_j = \frac{\mathrm{RR}_{TX_j}\,\mathrm{RR}_{X_j Y}}{\mathrm{RR}_{TX_j} + \mathrm{RR}_{X_j Y} - 1}.\] By Theorem 9.2, a confounder with these associations can explain the observed effect away only if \(B_j \geq \mathrm{RR}_{TY \mid X}^{\mathrm{obs}}\). Hence if \(B_j < \mathrm{RR}_{TY \mid X}^{\mathrm{obs}}\) for every observed covariate, no unmeasured confounder as strong as any observed covariate can account for the effect. One precision matters here: the comparison is one joint pair at a time. “As strong as covariate \(X_j\)” means having the same pair as that single covariate. It does not rule out a hypothetical confounder combining the strongest treatment association seen for one covariate with the strongest outcome association seen for another: because \(B(\cdot,\cdot)\) is increasing in each argument, such mixed maxima can yield a larger bias factor than any individual observed covariate exhibits.
Two scales are in play here, and mixing them is a common error. The bias factor \(B_j\) lives on the scale of the observed risk ratio, so the comparison just made is \(B_j\) against \(\mathrm{RR}_{TY \mid X}^{\mathrm{obs}}\) — not \(B_j\) against \(\mathrm{EV}\), which is on the scale of an individual association. To compare on the E-value scale instead, convert each covariate to the symmetric pair of associations that would produce the same bias factor: \[e_j = B_j + \sqrt{B_j\,(B_j - 1)}, \tag{9.24}\] which solves \(B(e_j, e_j) = B_j\) by the same algebra that produced Equation 9.20. Call \(e_j\) the equivalent E-value of covariate \(X_j\). Because \(B\) is increasing in each argument, \(e_j\) and \(\mathrm{EV}\) are directly comparable, and the conclusion survives a confounder as strong as \(X_j\) precisely when \(\mathrm{EV} > e_j\).
Cinelli and Hazlett (2020) develop the idea systematically for the linear regression setting, introducing partial-\(R^2\) benchmarks that quantify how much of the variation in \(T\) and \(Y\) an unmeasured confounder would need to explain — expressed as multiples of the partial-\(R^2\) of the strongest observed covariate.
E-values are defined for ratio measures. When the outcome is continuous, an E-value can still be reported, but only after the effect has been placed on a ratio scale; the usual device is to convert a standardized mean difference \(d\) to an approximate risk ratio \(\exp(0.91\,d)\) and compute the E-value from that (VanderWeele and Ding 2017). Quoting an E-value directly for a raw difference — a dollar amount, a test-score gap — is a category error, because the bias factor of Theorem 9.2 multiplies a ratio and has no interpretation as an additive displacement.
9.4.3 Reporting a Benchmarking Analysis
A minimal benchmarked sensitivity report contains three components: a point estimate with sampling-based confidence interval; a sensitivity summary (a tipping point, a bound, or the E-value); and a benchmarking comparison to observed covariates. The third component is essential. Without it, the sensitivity summary is a mathematical object without a scientific referent, and the report risks conveying a false precision about the robustness of the conclusion.
9.4.4 Reporting Guidelines
At minimum, an applied report should include the point estimate \(\hat\psi\) and its sampling 95% confidence interval; a sensitivity summary — one of (a) a sensitivity curve, (b) a sensitivity bound over a specified \(\mathcal{L}\), or (c) an E-value or tipping point; a benchmarking statement; and a plain-language interpretation.
9.4.4.1 Recommended Template
A sample report paragraph, in the style recommended throughout this chapter:
Under the back-door adjustment for \(X\), the estimated risk ratio is \(1.85\) with 95% CI \([1.25, 2.70]\). The E-value is \(3.10\) for the point estimate and \(1.81\) for the lower confidence limit. Among observed covariates, the bias factors \(B_j\) range from \(1.15\) (age) to \(1.60\) (baseline health status); converting each to its equivalent E-value by Equation 9.24 gives \(1.57\) and \(2.58\) respectively. Comparing on that common scale, \(3.10 > 2.58\): the point estimate is robust to an unmeasured confounder as strong as any single observed covariate, with roughly 20% to spare. The lower confidence limit is not, since \(1.81 < 2.58\). Explaining away the point estimate would therefore require unmeasured confounding stronger, on the equivalent-E-value scale, than any single measured covariate. Whether confounding of that strength is scientifically implausible depends on whether baseline health is a credible comparator for the important pathways that may remain unmeasured; the numbers alone do not settle that question.
Two features of this template are deliberate. The effect is reported as a risk ratio rather than as an additive ATE, because the E-value formula is defined on the ratio scale; attaching an E-value to a mean difference would report two incompatible objects. And the comparison is \(e_j\) against \(\mathrm{EV}\) throughout, never \(B_j\) against \(\mathrm{EV}\), which would mix scales.
9.5 Partial Identification and Overlap
9.5.1 From Sensitivity Models to Identified Sets
Definition 9.1 is stated for a point-valued relaxation, in which each \(\lambda\) pins \(\psi\) down to a single number. Several of the models in this chapter are not of that kind. The general object is therefore set-valued: \[\mathcal{I}_\lambda(P) = \{\psi(Q) : Q \text{ compatible with } P \text{ and } A_\lambda\}, \tag{9.25}\] the collection of estimand values consistent with the observed-data distribution \(P\) once the assumption has been relaxed to \(A_\lambda\). The neutral value is characterized by \(\mathcal{I}_{\lambda_0}(P) = \{\Psi(P)\}\). Write \(\underline\psi(\lambda) = \inf \mathcal{I}_\lambda(P)\) and \(\overline\psi(\lambda) = \sup \mathcal{I}_\lambda(P)\) for the pointwise lower and upper endpoints at a given \(\lambda\). When \(\mathcal{I}_\lambda(P)\) is a singleton these coincide and we recover the point-valued case; a sensitivity curve is the appropriate report exactly then. Otherwise the report is a band between \(\underline\psi\) and \(\overline\psi\).
One distinction within this formulation deserves emphasis. The MSM and the E-value bias-factor model fit the estimand-valued object directly: each restricts the possible values of, or discrepancy affecting, the causal estimand itself. Rosenbaum’s \(\Gamma\)-analysis is ordinarily an inferential sensitivity analysis: at fixed \(\Gamma\) it bounds a randomization \(p\)-value under a sharp null, not the causal estimand. It enters the estimand-set framework only after test inversion — for example, by inverting tests over a constant additive treatment effect to obtain a \(\Gamma\)-sensitivity interval. The \(\Gamma\)-model thus shares the indexed-relaxation logic of Definition 9.1, but its usual \(p\)-value output is not literally an element of \(\mathcal{I}_\Gamma(P)\) for \(\psi\) — a distinction worth marking in a chapter whose central theme is the separation of identification from inference.
For set-valued models, the tipping-point definition Equation 9.4 applies with \(\psi\) replaced by the pointwise lower endpoint \(\underline\psi\).
The sensitivity models above each restrict the magnitude of a violation through a finite-dimensional sensitivity parameter or plausibility region. A more radical approach is to impose no parametric restriction on the violation at all — or, equivalently, to ask what the observed data can say about the causal parameter without any untestable identifying assumption. The answer is a set of values rather than a single point. This is the theory of partial identification (Manski 1990, 2003).
9.5.2 Point Identification vs. Partial Identification
We have so far worked under point identification: given the baseline assumption \(A_0\), the estimand equals a specific functional of the observed-data distribution, \(\psi = \Psi(P)\). Under partial identification, the assumptions pin down not a single value of \(\psi\) but a set. The identified set \(\Psi_A(P)\) is the sharp answer to the question “given the assumption set \(A\) and the population distribution \(P\), what values of \(\psi\) are possible?” Point identification is the limiting case \(\Psi_A(P) = \{\psi\}\). Sensitivity bounds as defined in Equation 9.3 are a parametric restriction of the partial-identification idea.
The identified set is sharp in the sense that every value in it is consistent with some admissible \(Q\), and every value outside is not. Refining \(A\) (adding assumptions) shrinks \(\Psi_A(P)\); weakening \(A\) expands it.
Definition 9.2 (Sharp Endpoint Bounds and Conservative Bounds) The bounds \((\psi_L, \psi_U)\) are sharp endpoint bounds under assumption set \(A\) if \(\psi_L = \inf \Psi_A(P)\) and \(\psi_U = \sup \Psi_A(P)\). If \(\Psi_A(P)\) is an interval, then \(\Psi_A(P) = [\psi_L, \psi_U]\) and this interval is the sharp identified set; if \(\Psi_A(P)\) is nonconvex, \([\psi_L, \psi_U]\) is only its convex hull, which contains values attainable under no admissible \(Q\), and full sharpness requires reporting \(\Psi_A(P)\) itself. A bound is conservative (or merely valid) if \(\psi_L \leq \inf \Psi_A(P)\) and \(\psi_U \geq \sup \Psi_A(P)\): it contains the identified set but may be wider than \(A\) and \(P\) require.
Sharp endpoint bounds cannot be improved without adding assumptions; conservative bounds can sometimes be tightened by a more careful derivation under the same assumptions. Whenever bounds are reported, one should state which kind they are, because a sharpness claim has real content: it asserts that no further information can be extracted from \((A, P)\). For the average treatment effect and the related scalar estimands of this chapter, \(\Psi_A(P)\) is typically an interval, and sharp endpoint bounds coincide with the sharp identified set.
9.5.3 Manski’s No-Assumption Bound
The widest identified set is obtained by imposing no identifying assumption at all beyond what the observed data mechanically reveal. The label “no-assumption” is conventional shorthand and slightly overstates the case: the result below still requires consistency, a binary treatment, and known bounds on the support of both potential outcomes — not merely of the observed outcome, since the proof bounds the unobserved counterfactual means. What it dispenses with is any assumption about the treatment-assignment mechanism.
Theorem 9.4 (Manski’s No-Assumption Bound on the ATE) Assume consistency, \(Y = T\,Y(1) + (1-T)\,Y(0)\), and suppose both potential outcomes are bounded: \(Y(t) \in [y_L, y_U]\) almost surely for \(t \in \{0, 1\}\). Let \(p = P(T=1)\). Then \(\tau = \E\{Y(1) - Y(0)\}\) satisfies \(\tau \in [\tau_L, \tau_U]\) with \[\tau_L = \big[\E(Y \mid T=1)\,p + y_L\,(1 - p)\big] - \big[\E(Y \mid T=0)\,(1 - p) + y_U\,p\big], \tag{9.26}\] \[\tau_U = \big[\E(Y \mid T=1)\,p + y_U\,(1 - p)\big] - \big[\E(Y \mid T=0)\,(1 - p) + y_L\,p\big]. \tag{9.27}\] These bounds are sharp, and the width of the identified set is \[\tau_U - \tau_L = y_U - y_L. \tag{9.28}\]
Proof. By the law of total probability, \[\E\{Y(1)\} = \E\{Y(1) \mid T=1\}\, p + \E\{Y(1) \mid T=0\}\,(1 - p),\] and symmetrically for \(\E\{Y(0)\}\). Consistency identifies two of the four terms from the data: \(\E\{Y(1) \mid T=1\} = \E(Y \mid T=1)\) and \(\E\{Y(0) \mid T=0\} = \E(Y \mid T=0)\). The other two are unobserved counterfactual means, restricted only by the bounded-outcome condition to lie in \([y_L, y_U]\). The extremes of \(\E\{Y(1)\}\) are therefore attained at \(\E\{Y(1) \mid T=0\} \in \{y_L, y_U\}\), giving \[\E\{Y(1)\} \in \big[\E(Y \mid T=1)\,p + y_L\,(1 - p),\ \E(Y \mid T=1)\,p + y_U\,(1 - p)\big],\] and similarly for \(\E\{Y(0)\}\). The identified set for \(\tau\) is the Minkowski difference of the two intervals, giving Equation 9.26–Equation 9.27. The width is \([y_U - y_L](1-p) + [y_U - y_L]p = y_U - y_L\). Sharpness follows from the existence of potential-outcome distributions attaining each extreme, constructed explicitly as degenerate distributions concentrated at \(y_L\) or \(y_U\) in the unobserved arm. \(\square\)
The width \(y_U - y_L\) is the immediate and striking consequence: the no-assumption bound is as wide as the support of the outcome. If \(Y\) is a \([0,1]\)-valued indicator, the width is \(1\); if \(Y\) is a continuous health score on \([0, 100]\), the width is \(100\). In either case the bound is typically too wide to resolve the sign of \(\tau\).
Example 9.3 (Manski Bound in a Binary Outcome) Suppose \(Y \in \{0,1\}\), \(P(T=1) = 0.4\), \(\E(Y \mid T=1) = 0.6\), and \(\E(Y \mid T=0) = 0.3\). The observed association is \(0.3\). Applying Theorem 9.4 with \(y_L = 0\), \(y_U = 1\): \[\tau_L = (0.6)(0.4) + (0)(0.6) - [(0.3)(0.6) + (1)(0.4)] = 0.24 - 0.58 = -0.34,\] \[\tau_U = (0.6)(0.4) + (1)(0.6) - [(0.3)(0.6) + (0)(0.4)] = 0.84 - 0.18 = 0.66.\] The identified set is \(\tau \in [-0.34, 0.66]\), which includes zero. With only the bounded-outcome restriction, the observed positive association of \(0.3\) is consistent with a causal effect anywhere from a meaningful harm to a substantial benefit.
This example illustrates both the discipline and the limitation of pure partial identification. The discipline: the width of the set transparently conveys how much the conclusion depends on untestable assumptions. The limitation: without assumptions, the set is usually too wide to be informative. Productive use of partial identification therefore typically proceeds by adding shape restrictions that are weaker than point identification but strong enough to narrow the set.
9.5.4 Shape Restrictions
Monotone treatment response. The assumption \(Y(1) \geq Y(0)\) almost surely restricts the counterfactual outcomes to a half-plane: no unit is harmed by treatment. Under monotone treatment response, \(\E\{Y(1) \mid T=0\} \geq \E(Y \mid T=0)\) and \(\E\{Y(0) \mid T=1\} \leq \E(Y \mid T=1)\). Inserting these restrictions into the derivation of Theorem 9.4 tightens the bounds.
Monotone treatment selection. The assumption \(\E\{Y(t) \mid T=1\} \geq \E\{Y(t) \mid T=0\}\) for each \(t\) states that treated units, on average, would have had better outcomes than untreated units regardless of treatment level. This is plausible in many observational studies where selection into treatment is favorable.
Combining monotone treatment response with monotone treatment selection can produce informative bounds even when neither assumption alone suffices. Manski (2003) develops these and other shape-restriction hierarchies.
9.5.5 Bounds as Sensitivity Analysis
Partial identification and sensitivity analysis are two views of the same underlying object. A sensitivity model together with its bounds is a parametric partial-identification analysis; Manski’s no-assumption bound is the limit when \(\mathcal{L}\) is taken to permit arbitrary violations. Intermediate models — the \(\Gamma\)-model, the MSM — lie between these extremes, trading informativeness for robustness.
9.5.6 Positivity, Overlap, and Retargeting
Chapter 6 introduced positivity — \(0 < \pi(X) < 1\) almost surely — as a structural requirement for the identification formulas underlying propensity-score methods. When positivity fails, some counterfactual comparisons are not supported by the observed data, and the identification argument breaks down locally. When positivity holds formally but weakly, the estimators behave poorly even though the identification formula remains technically valid. Sensitivity analysis with respect to positivity addresses both failure modes.
9.5.6.1 Positivity as an Identification Assumption
The back-door identification formula involves the counterfactual outcome at both treatment levels for every value of \(X\). If positivity fails on a set \(S_0\) of positive \(P_X\)-probability — that is, \(\pi(x) = 0\) throughout \(S_0\) — there are no treated units with \(X \in S_0\) to inform \(\mu_1\) there, and the counterfactual mean \(\E\{Y(1) \mid X \in S_0\}\) is not identified by the data. The qualification matters: an isolated zero at a \(P_X\)-null point does not obstruct identification of the population ATE. On such a region, \(\mu_1(x_0)\) is not uniquely determined by the observed data, and any numerical value assigned to the missing arm-specific regression is supplied by a model-based extrapolation assumption rather than by observed support.
The statistical consequence is that the ATE itself may not be point-identified over the full covariate support even though it is identified over the subset where positivity holds. This motivates redefining the estimand to match the support where the data are informative.
9.5.6.2 Practical Positivity Violations
Even when \(0 < \pi(x) < 1\) in a strict sense, estimation becomes unstable when \(\pi(x)\) is near \(0\) or \(1\) over regions with non-negligible covariate mass. IPW weights are of order \(1/\pi(X)\), so near-violations generate extreme weights that inflate variance and destabilize estimates. The instability is compounded by the fact that errors in \(\hat\pi\) near the boundary translate into much larger errors in \(1/\hat\pi\), so that low-overlap regions are also the regions where propensity misspecification is most damaging. Crump et al. (2009) address this directly by asking which subpopulation yields the most precisely estimable treatment effect; they characterize the variance-optimal overlap subsample and show that it takes the form of a threshold rule on \(\pi(x)\), which in practice motivates discarding units with \(\pi(x)\) outside roughly \([0.1, 0.9]\). That interval is a widely used rule of thumb motivated by the Crump criterion, not a variance-optimal threshold for every design and estimand.
9.5.6.3 Trimming as a Sensitivity Analysis
Define the trimmed average treatment effect \[\tau_\alpha = \E\{Y(1) - Y(0) \mid \alpha \leq \pi(X) \leq 1 - \alpha\} \tag{9.29}\] for \(\alpha \in [0, 0.5)\). The choice \(\alpha = 0\) recovers the full-population ATE; \(\alpha > 0\) restricts attention to a subpopulation with adequate overlap.
Lemma 9.3 (Trimmed Estimand Identity) Let \(\alpha \in (0, 1/2)\), let \(S_\alpha = \{X : \alpha \leq \pi(X) \leq 1 - \alpha\}\), and suppose \(q_\alpha = P(X \in S_\alpha) > 0\). Under consistency and conditional exchangeability, \[\tau_\alpha = \frac{1}{q_\alpha}\,\E\!\left\{\big[\mu_1(X) - \mu_0(X)\big]\,\mathbf{1}\{X \in S_\alpha\}\right\}. \tag{9.30}\] Moreover, the IPW weights are bounded on the trimmed population: \(w(T, X) \leq 1/\alpha\) on \(S_\alpha\).
The restriction to \(\alpha \in (0, 1/2)\) is essential and not merely technical. At \(\alpha = 0\) the set \(S_0\) is the entire support, which in the setting motivating this section contains points with \(\pi(x) \in \{0, 1\}\); at such points \(\mu_1(x)\) or \(\mu_0(x)\) is not nonparametrically identified, so Equation 9.30 does not hold as an observed-data identity. It holds at \(\alpha = 0\) only if full-population positivity is separately assumed — which is precisely the assumption in question here. The causal estimand \(\tau_0\) still exists without positivity; what fails is its nonparametric identification by the displayed functional.
Proof. Conditional on \(X \in S_\alpha\), exchangeability and positivity continue to hold, the latter with slack \(\alpha > 0\), so the back-door formula applies pointwise. Integrating over the conditional distribution of \(X\) given \(X \in S_\alpha\) gives Equation 9.30. The IPW weight bound follows directly from \(\alpha \leq \pi(X) \leq 1 - \alpha\). \(\square\)
The key conceptual point is that \(\tau_\alpha\) and \(\tau_0\) are different estimands: trimming is not a variance-reduction trick that leaves the target unchanged, but a restatement of the scientific question. A trimmed analysis answers “what is the causal effect on the subpopulation for which the treatment decision is not already essentially forced by covariates?” For some scientific questions this is the better target; for others, it is not.
Trimming should therefore be interpreted as changing the estimand to a region of covariate overlap. It is a robustness-and-retargeting strategy, not a sensitivity correction for the original population ATE: the bias produced by a positivity violation in the full population is not removed by reporting a trimmed estimand, it is sidestepped by redefining what is being estimated.
A further distinction: the estimand Equation 9.29 is defined through the true propensity \(\pi\), but analysts trim on the estimated \(\hat\pi\), which makes the retained population data-adaptive. A careful report separates the conceptual overlap population, the empirical rule, and the additional uncertainty introduced by estimating which units are retained.
9.6 Sensitivity for Instrumental Variables and Mediation
Chapter 7 identified three assumptions that underwrite instrumental variables: relevance, exogeneity, and exclusion. Of the three, relevance is the most empirically diagnosable — weak-instrument \(F\)-tests and related diagnostics evaluate it directly from the first-stage regression. Exogeneity and exclusion, by contrast, involve unobserved relationships between the instrument and the outcome, and are not testable without additional structure. A sensitivity analysis for IV therefore concentrates on these two untestable assumptions.
9.6.1 What Weak-Instrument Diagnostics Do Not Address
A useful starting point is to be precise about what first-stage strength does and does not deliver. A strong first stage is not sufficient for instrument validity: it says nothing about exogeneity or exclusion, and a strong instrument may still directly affect the outcome. Conversely, an instrument may satisfy exogeneity and exclusion while having a weak first stage, in which case it is valid but provides practically uninformative — and possibly nonregular — inference. A nonzero first stage is of course necessary for identification by the Wald ratio; the point concerns strength, not relevance. The two concerns are otherwise orthogonal:
- Relevance is a statement about \((Z, T)\) given \(X\). It is observable: the first-stage coefficient is an identifiable functional of the observed distribution.
- Exogeneity restricts the relationship between \(Z\) and the latent determinants of treatment and outcome. Exclusion restricts the causal pathways from \(Z\) to \(Y\). Neither condition is generally verifiable from the observed distribution alone.
A complete robustness argument therefore has two parts: (a) an evaluation of first-stage strength (Chapter 7) and (b) a sensitivity analysis for exogeneity and exclusion (this section). The two parts are complementary, not substitutes.
9.6.2 Direct-Effect Violation of Exclusion
Consider the canonical scalar-IV model of Chapter 7 with a direct effect \(\delta\) of \(Z\) on \(Y\): \[Y = \alpha + \beta\, T + \delta\, Z + \varepsilon, \qquad \E(\varepsilon \mid Z) = 0. \tag{9.31}\] The condition \(\E(\varepsilon \mid Z) = 0\) maintains instrument exogeneity; the sensitivity parameter \(\delta\) relaxes only the exclusion restriction, which is the condition \(\delta = 0\). What follows is therefore an exclusion-only sensitivity analysis conditional on maintained exogeneity. A two-parameter analysis covering both assumptions would add a second term \(\mathrm{Cov}(Z, \varepsilon)/\mathrm{Cov}(T, Z)\) to the bias formula; we do not pursue it here.
Theorem 9.5 (IV Bias under Exclusion Violation) Under Equation 9.31 and the standard IV regularity conditions (relevance \(\mathrm{Cov}(T, Z) \neq 0\) and orthogonality of \(Z\) and \(\varepsilon\)), \[\frac{\mathrm{Cov}(Y, Z)}{\mathrm{Cov}(T, Z)} = \beta + \delta\,\frac{\mathrm{Var}(Z)}{\mathrm{Cov}(T, Z)}. \tag{9.32}\] Equivalently, the Wald estimand at posited \(\delta\) is \[\beta(\delta) = \frac{\mathrm{Cov}(Y, Z) - \delta\,\mathrm{Var}(Z)}{\mathrm{Cov}(T, Z)}. \tag{9.33}\]
Proof. Take the covariance of Equation 9.31 with \(Z\): \[\mathrm{Cov}(Y, Z) = \beta\,\mathrm{Cov}(T, Z) + \delta\,\mathrm{Var}(Z) + \mathrm{Cov}(\varepsilon, Z),\] and \(\mathrm{Cov}(\varepsilon, Z) = 0\) by orthogonality. Dividing by \(\mathrm{Cov}(T, Z)\) gives Equation 9.32. \(\square\)
Formula Equation 9.32 has a striking pedagogical consequence: the bias term is proportional to \(\mathrm{Var}(Z)/\mathrm{Cov}(T, Z)\), which is the inverse of the first-stage slope. A weak first stage therefore amplifies the bias from any exclusion violation: the same \(\delta\) causes more damage in a weak-IV analysis than in a strong-IV analysis. This is the opposite of the intuition that weak instruments can at least be “rescued” by being exogenous.
Example 9.4 (IV Sensitivity Curve) A study has \(\mathrm{Cov}(Y, Z) = 3\), \(\mathrm{Cov}(T, Z) = 1.5\), and \(\mathrm{Var}(Z) = 1\). The Wald point estimate is \(\hat\beta(0) = 3/1.5 = 2.0\). Applying Equation 9.33, \[\hat\beta(\delta) = \frac{3 - \delta}{1.5}, \qquad \hat\beta(0.5) = 1.67,\ \hat\beta(1.0) = 1.33,\ \hat\beta(2.0) = 0.67,\ \hat\beta(3.0) = 0.\] The tipping point for the sign is \(\delta^\star = \mathrm{Cov}(Y,Z)/\mathrm{Var}(Z) = 3\). It coincides numerically with \(\mathrm{Cov}(Y, Z)\) here only because \(\mathrm{Var}(Z) = 1\); in general the two differ. In a weaker-instrument setting with \(\mathrm{Cov}(T, Z) = 0.5\) but the same numerator, the tipping point would still be \(\delta^\star = 3\) but the rate of bias accumulation would be tripled — each unit of \(\delta\) moves the estimate by \(2\) rather than by \(2/3\).
9.6.3 Connection to LATE and Monotonicity
Under heterogeneous treatment effects, the Wald estimand identifies a local average treatment effect (LATE, Chapter 7) rather than an ATE, and the identification relies on a monotonicity assumption in addition to the three core IV assumptions. Sensitivity analysis can also address monotonicity violations, at the cost of additional modeling. Angrist et al. (1996) discuss a bound that includes “defiers” as an unobserved compliance type, and a natural first sensitivity parameter is the prevalence of defiers. Prevalence alone will not do, however. Writing \(\pi_c, \pi_d\) for the complier and defier proportions and \(\tau_c, \tau_d\) for their average treatment effects, the Wald estimand becomes \[\frac{\Delta_Y}{\Delta_T} = \frac{\pi_c \tau_c - \pi_d \tau_d}{\pi_c - \pi_d},\] so the distortion depends jointly on how many defiers there are, on how their average effect compares with the compliers’, and on the net first stage \(\pi_c - \pi_d\) in the denominator, which amplifies both. A sensitivity analysis must therefore posit at least two components, not one.
9.6.4 Sensitivity in Mediation Analysis
Sensitivity analysis is especially important in mediation analysis because the identifying assumptions are strictly stronger than those required for total effects. Chapter 8 derived the mediation formula under sequential ignorability (Imai et al. 2010): no unmeasured \(T\)–\(M\) confounding, no unmeasured \(M\)–\(Y\) confounding given \(T\), and a cross-world independence condition. Even in a randomized trial, where randomization of \(T\) eliminates the first, the second and third cannot generally be guaranteed.
9.6.4.1 Why Mediation Needs Sensitivity Analysis
The controlled direct effect (CDE) is the effect of setting \(T\) and \(M\) jointly to specified values; the back-door formula for the joint intervention identifies it under the stated conditions, which require that all relevant \(M\)–\(Y\) confounders be observed and appropriately adjusted for. Natural direct and indirect effects (NDE, NIE) additionally require the relevant cross-world independence condition; the standard graphical sufficient condition is that no \(M\)–\(Y\) confounder — observed or unobserved — be affected by treatment. This is a strictly stronger condition, and observing such a confounder does not repair it: adjusting for the realized value of a treatment-induced confounder does not in general identify the natural effects.
Panel (a) already defeats \(M\)–\(Y\) ignorability. In panel (b), \(U\) is in addition affected by \(T\); the cross-world independence condition then fails as well, and the NDE/NIE identification formula is biased even in a randomized trial. Strictly, the formal requirement for the natural effects is the cross-world independence; “no treatment-induced mediator–outcome confounder” is the standard graphical sufficient condition for it, which is why panel (b) is fatal. The residual-correlation parameter introduced below is aimed primarily at panel (a); panel (b) is the harder case, in which no scalar residual correlation fully captures the violation.
9.6.4.2 Residual-Correlation Sensitivity Parameter
A clean parameterization, following Imai et al. (2010), uses the correlation between the structural disturbances of the mediator and outcome equations. Consider the linear mediation system of Chapter 8: \[M = a_0 + a\,T + a_X^\top X + \varepsilon_M, \tag{9.34}\] \[Y = b_0 + \tau'\,T + b\,M + b_X^\top X + \varepsilon_Y, \tag{9.35}\] and define the sensitivity parameter as the correlation between the structural disturbances: \[\rho = \mathrm{Corr}(\varepsilon_M, \varepsilon_Y), \tag{9.36}\] where \((\varepsilon_M, \varepsilon_Y)\) are assumed independent of \((T, X)\) with a constant correlation. Within this maintained linear model, sequential ignorability implies \(\rho = 0\): the disturbances are orthogonal because any shared source of variation has been conditioned out, and a nonzero \(\rho\) represents exactly the unobserved \(M\)–\(Y\) confounding that the assumption rules out. The converse direction is model-bound, and the equivalence should not be read as a general nonparametric statement: the entire Imai et al. (2010) sensitivity calculation is explicitly model-based. Under misspecification of the joint mediator–outcome system, \(\rho = 0\) and sequential ignorability are not the same condition.
Under this parameterization, Imai et al. (2010) derive the bias in the estimated indirect effect as a smooth function of \(\rho\) and of the observed residual variances, giving a sensitivity curve \(\rho \mapsto \hat{\mathrm{NIE}}(\rho)\). The tipping point \(\rho^\star\) at which the indirect effect reaches zero is the reportable summary.
9.6.4.3 Reporting Mediation Sensitivity
A well-reported mediation sensitivity analysis states three things: the point estimate of the indirect effect (with its sampling CI), the tipping value \(\rho^\star\), and a benchmark for what magnitude of \(\rho\) is plausible in the given scientific setting. An example:
Under sequential ignorability (\(\rho = 0\)), the estimated indirect effect of the behavioral intervention on depression through sleep quality is \(0.38\) with 95% CI \([0.18, 0.58]\). The sensitivity curve crosses zero at \(\rho^\star = 0.31\), and the lower confidence limit crosses zero at \(\rho^\star_{\mathrm{CI}} = 0.15\). A residual correlation of \(\rho = 0.3\) corresponds to a squared residual correlation of \(0.09\).
It is tempting to describe \(\rho^2 = 0.09\) as “9% of the joint variation explained,” and this is often done. It should be resisted. A squared correlation between two disturbances acquires a variance-explained reading only under an additional latent-factor model specifying how a common cause loads on \(M\) and on \(Y\); without that model, \(\rho^2\) is simply a squared correlation.
9.7 Sensitivity Analysis and Modern Estimators
Chapters 10–13 develop estimating-equation theory, doubly robust estimation, orthogonal scores, cross-fitting, and semiparametrically efficient IV estimation. These tools make estimators more robust — but robust to a specific class of perturbations, namely nuisance-model misspecification. They are silent about identification.
9.7.1 What Robust Estimation Does Not Do
Orthogonality and cross-fitting do not solve any of the following:
- Unmeasured confounding. The target functional \(\E\{\mu_1(X) - \mu_0(X)\}\) does not equal \(\tau\) if \(Y(t) \nindep T \mid X\). No amount of accurate nuisance estimation corrects this; the nuisances are conditional on \(X\), and \(X\) is the wrong conditioning set.
- Positivity failure. If \(\pi(x) = 0\) on a region with positive covariate mass, the full-population target is not nonparametrically identified there and the influence-function theory underlying AIPW and DML breaks down. Note that an outcome-regression-based estimator may still return a finite number, because a fitted \(\hat\mu_1(x)\) extrapolates into the region where no treated units were observed. Numerical finiteness is not identification: the reported value is then a property of the extrapolating model, not of the data.
- IV invalidity. If the exclusion restriction fails, 2SLS, efficient GMM (Chapter 13), and any other IV estimator converge to a biased target. The bias formula Equation 9.32 applies to any estimator that is consistent for the scalar Wald ratio in model Equation 9.31. Other IV estimands — overidentified GMM, nonlinear models, and heterogeneous-effect targets such as the LATE — require their own sensitivity derivations.
- Mediation assumption failure. The NIE and NDE are well-defined causal estimands whatever the data can support — they are fixed by the causal model, not by an assumption. What sequential ignorability governs is whether the mediation formula identifies them. When it fails, the estimand still exists, but the mediation functional no longer identifies it from the same observed data without additional assumptions.
A compact summary:
Neyman orthogonality and cross-fitting protect the estimator against nuisance-estimation error. They do not protect against violations of causal identification assumptions.
This is the reason sensitivity analysis belongs in Part II (identification) rather than Part III (estimation). Every method in Chapters 10–13 assumes identification and refines the estimation. This chapter sits at the boundary: it refines the assessment of how much the identification can bend before the conclusion breaks.
Steps 1–4 are the customary content of a causal analysis. Step 5 is the contribution of this chapter, and step 6 brings the two outputs together. For observational analyses — or any analysis resting on substantively contestable, nonverifiable assumptions — omitting step 5 generally leaves the robustness argument incomplete.
9.8 Lab (Optional): A Tipping-Point Analysis for an Observational ATE
This lab implements a minimal sensitivity analysis end-to-end: a deliberately confounded observational ATE, a naive back-door adjustment that ignores the unmeasured confounder, and a tipping-point analysis that locates the strength of unmeasured confounding required to reverse the conclusion.
By construction, the true average treatment effect is \(\tau = 1.5\). Ignorability given \((X, U)\) holds, but ignorability given \(X\) alone is violated by the \(U\) arrows. The observed \(X\) enters the structural outcome equation with coefficient \(0.8\) — an oracle DGP coefficient, not the analyst-facing outcome-side benchmark, which is derived below as \(\hat\gamma_X \approx 0.605\) from the analyst’s observed regression.
9.8.1 Naive Adjusted Estimator
The analyst’s first estimator adjusts only for \(X\): \(\hat\tau_X\) is the OLS coefficient on \(T\) in the regression of \(Y\) on \((1, T, X)\). Running \(500\) Monte Carlo replicates of sample size \(n = 2{,}000\): the mean of \(\hat\tau_X\) across replicates is \(3.172\); the Monte Carlo standard deviation is \(0.099\); the true ATE is \(1.500\).
The Monte Carlo standard deviation describes spread across simulation runs; the standard error an analyst computes from one realized data set is a different object, though here the usual OLS standard error tracks it closely, giving a typical single-replicate 95% interval of about \([2.98,\, 3.37]\) — narrow, and not containing the truth. A consumer of this single report, relying on the CI for uncertainty, would conclude that the ATE is approximately \(3.2\) and confidently positive. This is the identification-uncertainty failure mode of Section 9.1.1 in concrete form.
9.8.2 Sensitivity Adjustment
Because \(U\) is Gaussian rather than binary, the relevant sensitivity model here is the continuous-\(U\) linear model of Section 9.2.2, not the binary-confounder formula. The bias quantity must also be matched to the estimator actually in use. Since \(\hat\tau_X\) is the OLS coefficient on \(T\) in a regression of \(Y\) on \((1, T, X)\), the Frisch–Waugh–Lovell theorem gives the exact omitted-variable-bias identity \[\operatorname{plim} \hat\tau_X = \tau + \delta\,\theta_U(\alpha_U), \qquad \theta_U(\alpha_U) = \frac{\E(\tilde T U)}{\E(\tilde T^2)}, \tag{9.37}\] where \(\delta\) is the outcome coefficient on \(U\) and \(\theta_U\) is the coefficient on \(T\) in the population regression of \(U\) on \((1, T, X)\), with \(\tilde T\) the residual from projecting \(T\) on \((1, X)\).
Three quantities are easily confused here, and only the last is correct for this estimator: the marginal difference \(\E(U \mid T{=}1) - \E(U \mid T{=}0)\), the standardized conditional contrast \(\E\{\E(U \mid T{=}1, X) - \E(U \mid T{=}0, X)\}\), and the regression coefficient \(\theta_U\). In this particular numerical design the last two happen to agree to three decimals; that is a coincidence of the coefficients used here, not a general identity, and it is not implied by the independence or normality of \(X\) and \(U\). Under other coefficient values the two differ appreciably — by more than 15% in some settings — because OLS weights strata by the conditional variance of \(T\) rather than by covariate mass. The marginal difference is smaller than either, because it ignores confounding by \(X\). Note also that \(\theta_U\) is built from conditional means of a standard normal variate, not from probabilities.
The analyst-facing analysis posits the pair \((\theta_U, \delta)\) directly. This is the estimator-matched parameterization: by Equation 9.37 these two numbers are exactly what determine the bias of \(\hat\tau_X\), and positing \(\theta_U\) itself avoids routing the analysis through a treatment-model coefficient \(\alpha_U\) that the analyst can neither observe nor estimate.
| Posited \(\delta\) | \(\theta_U = 0.25\) | \(0.50\) | \(0.75\) | \(1.00\) |
|---|---|---|---|---|
| \(1.0\) | \(2.922\) | \(2.672\) | \(2.422\) | \(2.172\) |
| \(2.0\) | \(2.672\) | \(2.172\) | \(1.672\) | \(1.172\) |
| \(3.0\) | \(2.422\) | \(1.672\) | \(0.922\) | \(0.172\) |
9.8.3 Tipping Contour and Benchmarking
The zero contour of the surface is the hyperbola \[\delta \cdot \theta_U = \hat\tau_X = 3.172: \tag{9.38}\] any posited pair whose product reaches \(3.172\) explains the entire adjusted estimate away, and no pair below the contour does.
The benchmark for the observed covariate must be estimator-matched, and both of its components must come from the analyst’s observed regressions. The treatment-side component is the auxiliary coefficient \(d_X^{\mathrm{OLS}} = \mathrm{coef}_T\{X \sim 1 + T\} = 0.472\). The outcome-side component is the coefficient on \(X\) in the regression the analyst actually ran: \[\hat\gamma_X = \mathrm{coef}_X\{Y \sim 1 + T + X\} \approx 0.605,\] not the structural coefficient \(0.8\) of the data-generating process. The two differ because omitting \(U\) biases the coefficient on \(X\) as well: conditioning on the collider \(T\) induces a negative partial association between \(X\) and \(U\) (\(\mathrm{coef}_X\{U \sim 1 + T + X\} \approx -0.097\)), and \(0.8 + 2 \times (-0.097) \approx 0.605\). The observed covariate’s implied bias is therefore \[\hat\gamma_X \, d_X^{\mathrm{OLS}} \approx 0.605 \times 0.472 \approx 0.286,\] an order of magnitude below the contour value \(3.172\): no unmeasured confounder with the same joint pair of associations as \(X\) comes close to explaining the estimate away.
Oracle verification: the true configuration \((\alpha_U = 1.0,\, \delta = 2.0)\) induces \(\theta_U = 0.832\), and \(3.172 - 2.0 \times 0.832 = 1.508 \approx \tau = 1.5\), confirming Equation 9.37.
9.8.4 Interpretation
The lab reinforces four points:
- The naive single-replicate CI \([2.98, 3.37]\) is narrow but does not contain the true ATE \(\tau = 1.5\). Sampling uncertainty is not identification uncertainty.
- Sensitivity adjustment at the correct unmeasured-confounding magnitude recovers the truth to within Monte Carlo error. The method works when the sensitivity parameters are correctly posited.
- The tipping point depends jointly on both sensitivity parameters, and two slices through the contour must not be conflated. At the observed imbalance of \(X\), \(d_X^{\mathrm{OLS}} = 0.472\), the tipping value is \(\delta^\star = 3.172/0.472 \approx 6.72\), about \(11\) times the observed partial outcome coefficient \(\hat\gamma_X \approx 0.605\) — this is the observed-covariate benchmark. Along the oracle slice \(\theta_U = 0.832\), \(\delta^\star = 3.172/0.832 \approx 3.81\) — a hypothetical slice through the surface, not a “confounder-as-strong-as-\(X\)” comparison.
- Benchmarking against the observed \(X\) anchors the analysis: \(X\)’s implied bias \(\approx 0.286\) sits an order of magnitude below the tipping contour \(3.172\).
9.9 Chapter Summary
This chapter closed Part II by addressing the complement of what Chapters 5–8 provided. Those chapters gave conditions under which a causal parameter is identified as a functional of the observed-data distribution; this chapter provided tools for reporting how much the identification can be violated before the causal conclusion changes.
- Sampling uncertainty, model misspecification, and identification uncertainty are distinct sources of error. A confidence interval addresses only the first, and identification uncertainty does not vanish merely by collecting more data from the same observational regime.
- A sensitivity analysis introduces a parameter \(\lambda\) that quantifies the strength of a violation, with a neutral value \(\lambda_0\) recovering the baseline assumption — zero for the additive parameters, one for the odds-ratio bounds \(\Gamma\) and \(\Lambda\). The three reporting objects are the sensitivity curve, the sensitivity bounds, and the tipping point.
- The master bias decomposition \(\Delta_{\mathrm{obs}} = \tau + B\) underwrites sensitivity analysis for unmeasured confounding, and specializes cleanly in the linear and binary-\(U\) cases.
- Three canonical sensitivity models differ by the object that the sensitivity parameter bounds. Rosenbaum’s \(\Gamma\) bounds a treatment-odds ratio. The VanderWeele–Ding E-value \(\mathrm{EV} = R + \sqrt{R(R-1)}\) is the smallest achievable value of the larger of \(\mathrm{RR}_{TU}\) and \(\mathrm{RR}_{UY}\) among confounders able to explain the effect away, so at least one of the two associations must reach it, while asymmetric pairs with one component below it can still suffice — though by the Cornfield inequality the smaller association must in every case be at least the observed risk ratio \(R\). The marginal sensitivity model bounds an odds ratio of nominal to true propensity, yielding valid — not sharp — sensitivity intervals for weighting estimators.
- Sensitivity parameters need benchmarking against observed covariates to be interpretable, and the benchmark is informative only when the comparator is scientifically comparable to the hypothesized unmeasured confounder.
- Partial identification reports an identified set rather than a single value. Manski’s no-assumption ATE bound has width \(y_U - y_L\), the maintained common support range of the two potential outcomes; shape restrictions narrow it, and a careful report distinguishes sharp endpoint bounds from conservative ones.
- Exact positivity failure on a positive-mass region is an identification failure; near-positivity violations are primarily a source of estimation and inferential instability. Trimming changes the estimand to a region of covariate overlap; it is a retargeting strategy, not a correction for the original population ATE.
- An exclusion violation of size \(\delta\) shifts the scalar Wald estimand by \(\delta \cdot \mathrm{Var}(Z)/\mathrm{Cov}(T, Z)\) — an exclusion-only analysis under maintained exogeneity — and the bias is amplified by weak first stages. Weak-instrument diagnostics do not address exclusion or exogeneity violations.
- Mediation sensitivity analysis proceeds, within a maintained linear mediation model, by a residual-correlation parameter \(\rho\) that captures ordinary unmeasured \(M\)–\(Y\) confounding. A treatment-induced mediator–outcome confounder is the harder failure mode and is not captured by a scalar \(\rho\).
- Modern estimation methods (AIPW, TMLE, DML, efficient GMM) address nuisance estimation, not identification. Neyman orthogonality protects the estimator; it does not protect the causal claim.
9.10 Problems
1. Sampling vs. identification uncertainty. In at most half a page, explain why identification uncertainty does not shrink as the sample size grows. Sketch on paper (no simulation required) a data-generating process with one observed covariate \(X\) and one unmeasured confounder \(U\) in which a back-door-adjusted estimator converges to a value different from the true ATE; state which identifying assumption is violated and give the limiting bias in terms of the model coefficients. (A full simulation version of this exercise is in the optional supplement.)
2. Binary unmeasured confounder. Suppose the observed adjusted effect is \(\hat\Delta_{\mathrm{obs}} = 2.0\). Assume a binary unmeasured confounder \(U\) with constant outcome contrast \(\delta = 4\), and suppose the \(U\)-treatment imbalance is uniform across \(X\) with value \(p_1 - p_0 = 0.3\).
- Use Lemma 9.2 to compute the confounding bias \(B\) and the sensitivity-adjusted estimate \(\hat\tau_{\mathrm{sens}}\).
- What outcome contrast \(\delta\) would be required to zero out the estimate at the same imbalance level?
- What imbalance \(p_1 - p_0\) would be required to zero out the estimate at \(\delta = 4\)?
3. Tipping point in a linear bias model. For \(\hat\Delta_{\mathrm{obs}} = 1.5\) and bias model \(B(\lambda) = 0.4\,\lambda\):
- Find the tipping point \(\lambda^\star\) for the sign of the adjusted estimate.
- Suppose the sampling standard error of the observed estimate is \(0.3\), independent of \(\lambda\). Find the tipping point \(\lambda^\star_{\mathrm{CI}}\) for the lower confidence limit to reach zero (95% level). Compare to \(\lambda^\star\).
4. Benchmarking. A researcher reports a sensitivity analysis with \(\hat\tau_X = 2.4\) under the linear model Equation 9.10, in which the tipping point for the \(U\)-outcome effect is \(\delta^\star = 3.0\) at a posited treatment imbalance of \(d = 0.8\). The observed covariates have estimated outcome coefficients and treatment imbalances:
| Covariate | \(\hat\gamma_j\) | \(\hat d_j\) |
|---|---|---|
| age | 0.3 | 0.20 |
| income | 1.8 | 0.35 |
| education | 2.4 | 0.10 |
- Explain why the outcome coefficient \(\hat\gamma_j\) alone is not a complete benchmark under the linear sensitivity model, and state what the second component contributes.
- Compute each covariate’s implied bias \(\hat\gamma_j \hat d_j\) and identify which covariate is the strongest benchmark on that scale. Note that the ranking differs from the ranking by \(\hat\gamma_j\) alone.
- Compare \(\delta^\star = 3.0\) with the largest observed outcome coefficient, while holding the imbalance fixed at \(d = 0.8\). Is the result robust to an unmeasured confounder with the same outcome association as the strongest observed covariate?
- Write two sentences of interpretation suitable for an applied report, making clear which quantity is being held fixed.
5. E-value calculation. An observational study reports an adjusted risk ratio of \(1.9\) with 95% CI \([1.3, 2.8]\).
- Compute the E-value for the point estimate using Equation 9.20.
- Compute the E-value for the lower confidence limit.
- Write one sentence of interpretation for each, stating the symmetric reading correctly.
- Exhibit a pair \((\mathrm{RR}_{TU}, \mathrm{RR}_{UY})\) with \(\mathrm{RR}_{TU}\) strictly below the E-value of part (a) that nonetheless explains the point estimate away, and verify it using Equation 9.18. Explain why this does not contradict Equation 9.21, and why one may not conclude that both associations must individually exceed the E-value.
6. Manski bound calculation. Suppose consistency holds and both potential outcomes satisfy \(Y(t) \in [0, 10]\) almost surely. Also suppose \(P(T = 1) = 0.3\), \(\E(Y \mid T=1) = 6.5\), and \(\E(Y \mid T=0) = 4.0\).
- Compute Manski’s no-assumption bound for the ATE using Equation 9.26–Equation 9.27.
- Compute the width of the bound and verify Equation 9.28.
- The observed association is \(2.5\). Does the no-assumption identified set include zero? Discuss what this implies about the informational content of the data alone.
7. Modern estimators and identification. A colleague argues that because DML and AIPW are “doubly robust and orthogonalized,” they “automatically correct for unmeasured confounding provided the machine-learning models are good enough.” Explain why this is incorrect, referring to the bias decomposition of Theorem 9.1. Be precise about what is and is not estimable: machine learning can consistently estimate the observed-data functional \(\Delta_{\mathrm{obs}}\) and the observed conditional means \(\mu_t(X)\), but it cannot recover the split of \(\Delta_{\mathrm{obs}}\) into \(\tau\) and \(B\), nor the \(U\)-indexed quantities \(m_t(X, U)\) and \(g_t(X, U)\), without information beyond the observed data. Explain why no amount of flexibility in the nuisance estimators changes this.