Appendix C — Pathwise Differentiability, Tangent Spaces, and the Efficient Influence Function
Chapter 10 introduced the efficient influence function of the average treatment effect functional and stated, without proof, that this object is the canonical gradient of \(\Psi\) relative to the nonparametric tangent space. This appendix supplies the geometric machinery behind that statement. The treatment follows Bickel et al. (1993) and Tsiatis (2006), with notation aligned to Chapter 10.
The reader should think of this appendix as the formal counterpart of the asymptotic-linearity and efficiency sections of Chapter 10: those sections showed how an influence function determines the asymptotic distribution of an estimator; here we develop the parallel notion of an influence function as a derivative of a functional, and characterize the unique influence function that achieves the semiparametric efficiency bound. The notation is standard in semiparametric theory and is introduced as needed.
The applied consequences — the AIPW estimator, double robustness, Neyman orthogonality, cross-fitting, and rate conditions — are developed in Chapters 11 and 12. This appendix supplies only the foundational geometry on which those chapters rest. Section C.1 collects the Hilbert-space background used throughout, for readers who have not previously encountered the projection theorem and Riesz representation. Section C.7 closes the appendix by carrying out the EIF derivation explicitly for the average treatment effect, recovering the AIPW influence function from the projection construction of Theorem C.1.
How to read this appendix. The material divides into a core path and a set of refinements, and it is worth separating them on a first reading. The core is Section C.1 through the projection theorem and Riesz representation; Definition C.1 and the meaning of a score; Definition C.2 with Example C.1; Definition C.3 and Definition C.4; Theorem C.1; the ATE derivation of Section C.7 through Equation C.13 and the bound Equation C.14; and the synthesis table of Section C.8. That path is self-contained and suffices for reading Chapters 11 and 12.
The remaining material is second-pass: the density counterexample (Example C.2); the attainability discussion after Definition C.4; the normal-location example (Example C.3); the convolution discussion following Theorem C.2; the weak-positivity subtleties in the definition of \(\mathcal{H}_T\); and the IPW-versus-AIPW remark at the end of Section C.7. Exercises marked [Advanced] belong to the same tier. None of it is optional in the long run — each item exists because a natural first guess about the theory is wrong — but none of it is needed to follow the main line.
C.1 Hilbert-Space Background
The geometry of pathwise differentiability lives in the Hilbert space \(L_2(P)\) of square-integrable real-valued functions on \(\mathcal{O}\), equipped with the inner product \[\langle f, g \rangle_P \;=\; \E_P[f(O)\, g(O)], \qquad \|f\|_P \;=\; \langle f, f \rangle_P^{1/2}.\] Two functions are orthogonal, written \(f \perp g\), when \(\langle f, g \rangle_P = 0\). All scores and influence functions introduced below live in the closed subspace of mean-zero functions, \[L_2^0(P) \;=\; \{f \in L_2(P) : \E_P[f(O)] = 0\}.\] This section records the four Hilbert-space facts used repeatedly in the rest of the appendix. Readers familiar with these may skip ahead; for a full treatment, see Vaart (1998) (§25.7) or the appendices of Tsiatis (2006).
Closed linear span. For a subset \(A \subset L_2^0(P)\), the closed linear span \(\overline{\mathrm{span}}(A)\) is the smallest closed subspace of \(L_2^0(P)\) containing \(A\); concretely, it consists of \(L_2(P)\)-limits of finite linear combinations of elements of \(A\). Tangent spaces below are defined as closed linear spans, rather than as raw sets of scores, because the set of scores of regular parametric submodels need not itself be closed.
Projection theorem. If \(V \subset L_2^0(P)\) is a closed subspace, every \(f \in L_2^0(P)\) admits a unique decomposition \[f \;=\; f_V \,+\, f_{V^\perp}, \qquad f_V \in V, \quad f_{V^\perp} \in V^\perp,\] where \(V^\perp = \{g \in L_2^0(P) : \langle g, h \rangle_P = 0 \text{ for all } h \in V\}\) is the orthogonal complement. The map \(\Pi[\,\cdot \mid V\,]: f \mapsto f_V\) is the orthogonal projection onto \(V\) and is characterized equivalently by either of the following:
- \(f - \Pi[f \mid V] \perp V\) (orthogonality of the residual);
- \(\Pi[f \mid V]\) is the unique element of \(V\) minimizing \(\|f - g\|_P\) over \(g \in V\) (best approximation).
The decomposition \(L_2^0(P) = V \oplus V^\perp\) is the engine behind the canonical-gradient construction in Section C.6.
Riesz representation. Every continuous linear functional \(\Lambda: L_2^0(P) \to \mathbb{R}\) admits a unique representer \(r_\Lambda \in L_2^0(P)\) such that \[\Lambda(f) \;=\; \langle r_\Lambda, f \rangle_P \qquad \text{for all } f \in L_2^0(P),\] and conversely every \(r \in L_2^0(P)\) defines such a continuous functional via \(f \mapsto \langle r, f \rangle_P\). The same statement holds verbatim with \(L_2^0(P)\) replaced by any closed subspace, in particular by the tangent space \(\mathcal{T}\).
Pathwise differentiability of a functional \(\Psi\) (Definition C.2 below) is precisely the statement that the score-to-derivative map \(S \mapsto \partial_\varepsilon \Psi(P_\varepsilon)|_0\) extends to a continuous linear functional \[\dot\Psi_P : \mathcal{T} \to \mathbb{R}\] on the tangent space. Riesz representation then supplies a unique \(\varphi^* \in \mathcal{T}\) with \[\dot\Psi_P(g) \;=\; \E_P[\varphi^*(O)\, g(O)], \qquad g \in \mathcal{T},\] and this unique representer in \(\mathcal{T}\) is the canonical gradient. It is important to keep two statements apart. Uniqueness holds within \(\mathcal{T}\); a function \(\varphi \in L_2^0(P)\) represents the same derivative on every score if and only if \[\varphi \;=\; \varphi^* + h \qquad \text{for some } h \in \mathcal{T}^\perp,\] so representers in the ambient space \(L_2^0(P)\) are generally non-unique whenever \(\mathcal{T} \neq L_2^0(P)\). Their projections onto \(\mathcal{T}\) all coincide. Section C.4 takes up this non-uniqueness, and Theorem C.1 states the resulting dichotomy formally.
Codimension. A non-zero continuous linear functional \(\Lambda\) on a Hilbert space \(H\) has closed kernel \(\ker(\Lambda) := \{f \in H : \Lambda(f) = 0\}\) of codimension one, and the orthogonal complement of that kernel within \(H\) is the line spanned by the Riesz representer: \[H \cap \ker(\Lambda)^\perp \;=\; \mathrm{span}\{r_\Lambda\}.\] Applied with \(H = \mathcal{T}\) and \(\Lambda = \dot\Psi_P\), this describes the geometry of a scalar target: the tangent space splits into the nuisance directions, along which \(\psi\) does not move to first order, and the single remaining direction along which it does. The codimension-one statement is a geometric consequence, not the source of uniqueness — uniqueness of \(\varphi^*\) comes from Riesz representation on \(\mathcal{T}\) and requires no rank condition. In the degenerate case \(\dot\Psi_P = 0\) one has \(\varphi^* = 0\) and \(\ker(\dot\Psi_P) = \mathcal{T}\), which has codimension zero.
C.2 Regular Parametric Submodels and Scores
Let \(\mathcal{P}\) be a statistical model for the distribution of the observed data \(O\) on a sample space \(\mathcal{O}\), with true distribution \(P \in \mathcal{P}\). For simplicity this appendix works throughout in a dominated model: all densities below are taken with respect to one common \(\sigma\)-finite reference measure \(\mu\). This is an assumption on \(\mathcal{P}\), not a triviality. The general theory of differentiable paths does not require a single measure dominating the whole model — domination local to each path is enough — but the global convention keeps the notation light and costs nothing in the examples treated here. The parameter of interest is a smooth real-valued functional \[\Psi:\mathcal{P} \to \mathbb{R}, \qquad \psi = \Psi(P).\]
For vector-valued \(\Psi: \mathcal{P} \to \mathbb{R}^p\), each component is pathwise differentiable in the sense of Definition C.2, and we write \(\varphi_j^* \in \mathcal{T}\) for the canonical gradient of the \(j\)th component (Definition C.5). Stacking the componentwise canonical gradients gives \(\varphi^* = (\varphi_1^*, \ldots, \varphi_p^*)^\top\), and the semiparametric efficiency bound is the covariance matrix \[V^* \;=\; \E_P[\varphi^*(O)\, \varphi^*(O)^\top] \;\in\; \mathbb{R}^{p \times p}.\] For any regular asymptotically linear estimator with influence function \(\varphi\), the matrix difference \(\E_P[\varphi(O)\, \varphi(O)^\top] - V^*\) is positive semidefinite. The word canonical is doing real work here: stacking arbitrary componentwise gradients would admit terms in \(\mathcal{T}^\perp\) and inflate the covariance matrix, so the resulting object would not be the bound.
Asymptotic-variance comparisons are then made in the Loewner partial order, or equivalently through all scalar linear contrasts \(c^\top \Psi\), \(c \in \mathbb{R}^p\). No new Hilbert-space principle is needed. Two caveats are worth recording. The bound matrix may be singular: there is then a non-zero \(c\) with \(c^\top V^* c = 0\), i.e. \(c^\top \varphi^* = 0\) almost surely, so the contrast \(c^\top \Psi\) has zero canonical gradient and zero first-order efficiency bound. This usually signals a redundant target vector or a locally flat contrast, and it should not be read as an impossibility theorem ruling out every non-degenerate estimator of that contrast. Second, individual contrasts may differ in their regularity properties, so pathwise differentiability of every component is a genuine requirement rather than a formality. For clarity we restrict attention to \(p = 1\) throughout.
To define what it means for \(\Psi\) to be “smooth,” we need a notion of a path through \(\mathcal{P}\) that begins at \(P\). The standard device is the regular parametric submodel.
Definition C.1 (Regular Parametric Submodel) A regular parametric submodel of \(\mathcal{P}\) passing through \(P\) is a one-parameter family \[\{P_\varepsilon : \varepsilon \in (-\delta, \delta)\} \subset \mathcal{P}, \qquad P_0 = P,\] that is differentiable in quadratic mean (DQM) at \(\varepsilon = 0\): there exists \(S \in L_2(P)\) such that \[\int \left\{\sqrt{p_\varepsilon(o)} \,-\, \sqrt{p(o)} \,-\, \tfrac{1}{2}\,\varepsilon\, S(o)\, \sqrt{p(o)}\right\}^2 d\mu(o) \;=\; o(\varepsilon^2) \qquad (\varepsilon \to 0). \tag{C.1}\] The function \(S\) is the score of the submodel at \(P\).
The definition differentiates the square-root density \(\varepsilon \mapsto p_\varepsilon^{1/2}\) in \(L_2(\mu)\) rather than the log density in \(L_2(P)\). This is the standard construction, and the choice is not cosmetic: \(L_2(P)\)-differentiability of \(\log p_\varepsilon\) is neither the usual definition nor, without further integrability conditions, sufficient to license the likelihood and integration manipulations used below. Differentiability in quadratic mean supplies a well-defined square-integrable score and the likelihood expansion underlying local asymptotic normality. It does not by itself license differentiation of an unbounded functional such as \(\E_{P_\varepsilon}[Y]\); that step may additionally require local moment or uniform-integrability conditions, as Example C.1 illustrates. In particular Equation C.1 implies \[\E_P[S(O)] = 0, \qquad \E_P[S(O)^2] < \infty, \tag{C.2}\] so every score lies in \(L_2^0(P)\), and the score \(S\) is determined by the path up to \(P\)-null sets. See Vaart (1998) (§7.2) for the proof and for the local asymptotic normality that DQM buys.
The familiar log-derivative formula survives as a consequence rather than a definition. When \(\varepsilon \mapsto p_\varepsilon(o)\) is differentiable pointwise and the usual domination conditions permit interchanging differentiation and integration, the DQM score satisfies \[S(o) \;=\; \left.\frac{\partial}{\partial \varepsilon}\log p_\varepsilon(o)\right|_{\varepsilon = 0} \qquad \text{for } P\text{-almost every } o,\] and this is the representation used in every concrete calculation below. A regular submodel is, intuitively, a smooth one-dimensional curve through \(\mathcal{P}\) that can be probed by ordinary parametric methods, and its score is the directional derivative of the log-likelihood at \(P\) along that curve.
C.3 Pathwise Differentiability
The functional \(\Psi\) is differentiable along a submodel if the map \(\varepsilon \mapsto \Psi(P_\varepsilon)\) is differentiable at \(\varepsilon = 0\) in the ordinary sense. Pathwise differentiability asserts that this derivative can be represented as an inner product between a fixed function and the score, with the same representer for every regular submodel.
Definition C.2 (Pathwise Differentiability and Influence Function) The functional \(\Psi: \mathcal{P} \to \mathbb{R}\) is pathwise differentiable at \(P\) if there exists a function \(\varphi \in L_2^0(P)\) such that, for every regular parametric submodel \(\{P_\varepsilon\}\) with score \(S\), \[\left.\frac{\partial}{\partial \varepsilon}\Psi(P_\varepsilon)\right|_{\varepsilon=0} \;=\; \E_P[\varphi(O)\, S(O)]. \tag{C.3}\] Any such \(\varphi\) is called an influence function (or gradient) of \(\Psi\) at \(P\). The set of influence functions is denoted \(\mathrm{IF}(\Psi, P)\).
Equation Equation C.3 is the defining identity of semiparametric theory. It says the perturbation of \(\Psi\) along a submodel is fully encoded by the \(L_2(P)\) inner product of a fixed function \(\varphi\) with the score \(S\). Geometrically, \(\varphi\) is a representer for the linear functional \(S \mapsto \partial_\varepsilon \Psi(P_\varepsilon)|_0\) restricted to the space of scores. Because the right-hand side of Equation C.3 is continuous in \(S\) with respect to the \(L_2(P)\) norm, this map extends uniquely by linearity and continuity to the whole tangent space; we write \[\dot\Psi_P : \mathcal{T} \to \mathbb{R}, \qquad \dot\Psi_P(g) \;=\; \E_P[\varphi(O)\, g(O)],\] and call \(\dot\Psi_P\) the pathwise derivative of \(\Psi\) at \(P\). Note that \(\dot\Psi_P\) does not depend on which gradient \(\varphi\) is used to define it, since any two agree on \(\mathcal{T}\).
Example C.1 (Mean Functional) Let \(O = Y\) with \(\E_P[Y^2] < \infty\) and \(\Psi(P) = \E_P[Y]\). Let \(\{P_\varepsilon\}\) be a regular submodel with score \(S\), and assume the local integrability condition that permits differentiating \(\varepsilon \mapsto \E_{P_\varepsilon}[Y]\) under the integral sign at \(\varepsilon = 0\). Then \[\frac{\partial}{\partial \varepsilon}\int y\, p_\varepsilon(y)\, d\mu(y)\bigg|_{0} = \int y\, S(y)\, p(y)\, d\mu(y) = \E_P[Y \cdot S(O)].\] Subtracting \(\E_P[Y]\cdot \E_P[S] = 0\) gives \(\partial_\varepsilon \Psi(P_\varepsilon)|_0 = \E_P[(Y - \E_P Y)\, S(O)]\), so \[\varphi(O) = Y - \Psi(P)\] is an influence function of the mean functional. Under the nonparametric model (Section C.5), this is in fact the unique influence function, and hence the efficient influence function.
The example illustrates two features common to all gradient calculations. First, \(\varphi\) is identified only up to functions orthogonal to the model’s tangent space (defined in Section C.5); the constant \(\Psi(P)\) above was subtracted to make \(\varphi\) mean-zero. Second, the calculation is an exchange of differentiation and integration. It is worth being precise about what licenses that exchange. Differentiability in quadratic mean supplies the score structure — that \(\partial_\varepsilon p_\varepsilon\) behaves like \(S p\) in the appropriate \(L_2\) sense — but it does not by itself justify differentiating the integral of an unbounded integrand such as \(y\) against \(p_\varepsilon\). For bounded \(Y\) the interchange is immediate. For unbounded \(Y\) it follows from a local second-moment condition, for instance \(\sup_{|\varepsilon| < \delta} \E_{P_\varepsilon}[Y^2] < \infty\), which gives uniform integrability along the path. Conditions of this kind are assumed silently in most treatments; they are recorded here because the appendix is meant to be the formal counterpart of Chapter 10 rather than a summary of it.
Not every functional of interest is pathwise differentiable, and the failure is not an edge case. The next example is the reason the hypothesis in Definition C.2 has content.
Example C.2 (A Functional with No Gradient) Second reading.
Let \(O = Y\) and let \(\mathcal{P}\) consist of the distributions on \(\mathbb{R}\) whose density admits a continuous version in a neighborhood of a fixed point \(y_0\), with \(p(y_0) > 0\). Evaluating that continuous version at \(y_0\) defines \(\Psi(P) = p(y_0)\).
Fixing a version matters: an element of \(L_2(P)\) is an equivalence class, so pointwise evaluation is not otherwise well defined. For the same reason the perturbations must be restricted to bounded continuous \(h \in L_2^0(P)\), so that \(p_\varepsilon = (1 + \varepsilon h)\, p\) remains a continuous density near \(y_0\) and stays inside \(\mathcal{P}\). Each such path is regular with score \(h\), by the computation in Section C.5, and \[\left.\frac{\partial}{\partial\varepsilon}\Psi(P_\varepsilon)\right|_0 \;=\; p(y_0)\, h(y_0).\]
Now let \(h_k\) be continuous mean-zero bumps supported on \([y_0 - 1/k,\, y_0 + 1/k]\) with peak height \(\sqrt{k}\). Then \(\|h_k\|_P = O(1)\), because the squared peak \(k\) is offset by the support width \(2/k\), while \(\partial_\varepsilon \Psi(P_{\varepsilon,k})|_0 \asymp \sqrt{k}\, p(y_0) \to \infty\). The score-to-derivative map is therefore unbounded on the unit ball of \(\mathcal{T}\): it is not continuous, and it has no Riesz representer in \(L_2(P)\). Informally, the only object that would represent point evaluation is a Dirac delta, which is not a function in \(L_2(P)\).
Two conclusions follow, and they are not the same conclusion. The functional is not pathwise differentiable, so by the remark on the connection to estimation below, no regular asymptotically linear estimator of \(p(y_0)\) exists in this model. Ruling out every regular \(\sqrt{n}\)-consistent estimator is a strictly stronger statement and needs the local asymptotic minimax or convolution machinery of Vaart (1998) (Chapter 25); under that framework it does indeed follow. This is the analytic content of the familiar fact that density estimation converges more slowly than \(\sqrt{n}\).
The same phenomenon affects causal targets. A regression function \(\E_P[Y \mid X = x_0]\) at a fixed value of a continuously distributed covariate, or a dose-response curve \(\E_P[Y(t_0)]\) at a single level of a continuous treatment, is not pathwise differentiable in the unrestricted model. Students sometimes assume that any causal estimand that is identified must possess an efficient influence function; identification and pathwise differentiability are independent properties. It is equally tempting, and equally wrong, to suppose that adding smoothness fixes the problem. Restrictions do change the tangent space and can improve rates, but ordinary finite-order smoothness or shape constraints — Hölder classes, monotonicity — generally leave pointwise targets slower than \(\sqrt{n}\). What typically restores pathwise differentiability is smoothing or averaging the target itself, as the ATE averages a regression contrast over \(P_X\) in Section C.7, or imposing a sufficiently strong structural, often finite-dimensional, restriction on the model.
C.4 Non-Uniqueness of Influence Functions
Definition C.2 does not produce a unique influence function. If \(\varphi\) satisfies Equation C.3 and \(h \in L_2^0(P)\) is orthogonal to every score of the model in \(L_2(P)\), then \(\varphi + h\) also satisfies Equation C.3, since \(\E_P[(\varphi+h)S] = \E_P[\varphi S] + \E_P[h S] = \E_P[\varphi S]\). Whether \(\varphi\) is unique depends on how rich the collection of scores is — a point made precise in Section C.5 through the notion of the tangent space.
The terminological distinction Chapter 10 alluded to can now be made precise:
- An estimating function is any function \(U(O;\theta)\) whose root defines an estimator.
- An asymptotic influence function of an estimator is the function \(\varphi\) in its asymptotic expansion.
- A (pathwise) influence function of a functional is a \(\varphi \in L_2^0(P)\) satisfying Equation C.3.
For a regular asymptotically linear estimator of a pathwise-differentiable functional, the estimator’s asymptotic influence function is also a pathwise influence function of the functional. The relationship to estimating functions is looser: for an M-estimator solving \(n^{-1}\sum_{i=1}^n U(O_i;\theta) = 0\), the standard expansion gives \[\varphi(O) \;=\; -A^{-1}\, U(O;\theta_0), \qquad A \;=\; \E_P\!\left[\partial U(O;\theta_0)/\partial \theta^\top\right],\] so a generic estimating function \(U\) equals the influence function \(\varphi\) only after this normalization (or when \(U\) is written directly in influence-function form). The first two notions can be defined without choosing a statistical model class \(\mathcal{P}\) — an estimator influence function is still defined at a particular distribution \(P\) — while the third depends crucially on \(\mathcal{P}\) through the collection of admissible submodels. This dependence is what makes \(\mathrm{IF}(\Psi, P)\) generally non-unique in semiparametric models.
C.5 The Tangent Space
The space of admissible perturbations of \(P\) within the model \(\mathcal{P}\) is called the tangent space.
Definition C.3 (Tangent Space) The tangent space of \(\mathcal{P}\) at \(P\), denoted \(\mathcal{T}_P(\mathcal{P})\) or simply \(\mathcal{T}\), is the closed linear span in \(L_2^0(P)\) of all scores of regular parametric submodels of \(\mathcal{P}\) passing through \(P\).
The tangent space is a closed subspace of the Hilbert space \(L_2^0(P)\). Two extreme cases are illustrative.
Unrestricted nonparametric model. Suppose \(\mathcal{P}\) is the unrestricted dominated model, containing every distribution on \(\mathcal{O}\) that is absolutely continuous with respect to \(\mu\). For bounded \(h \in L_2^0(P)\) the tilted family \[p_\varepsilon(o) \;=\; \bigl(1 + \varepsilon\, h(o)\bigr) p(o)\] is a genuine density for \(|\varepsilon| < 1/\|h\|_\infty\): it is non-negative there, and it integrates to \(1 + \varepsilon\, \E_P[h] = 1\) exactly, so no normalizing constant is required. The path is regular with score \(S = h\). Since bounded mean-zero functions are dense in \(L_2^0(P)\), taking the closed linear span gives \(\mathcal{T} = L_2^0(P)\).
It is worth resisting the habit of treating “nonparametric” and “maximal tangent space” as synonyms. The conclusion \(\mathcal{T} = L_2^0(P)\) holds because the model admits every bounded tilt. A model can be infinite-dimensional and still impose support, shape, monotonicity, smoothness, or moment restrictions, any of which excludes some tilts and yields a proper closed subspace \(\mathcal{T} \subsetneq L_2^0(P)\).
Fully parametric model. If \(\mathcal{P} = \{P_\theta : \theta \in \Theta \subset \mathbb{R}^k\}\) is parametric with score components \(S_{\theta_0, 1}, \ldots, S_{\theta_0, k}\) at \(\theta_0\), then \(\mathcal{T} = \mathrm{span}\{S_{\theta_0,1}, \ldots, S_{\theta_0,k}\}\) is finite-dimensional, of dimension \(\mathrm{rank}\{I(\theta_0)\} \leq k\), where \(I(\theta_0)\) is the Fisher information matrix. The dimension is exactly \(k\) precisely when the score components are linearly independent in \(L_2(P)\), that is, when \(I(\theta_0)\) is non-singular. A singular information matrix means that some parameter direction produces no first-order perturbation of \(P\). That may reflect local non-identification or a redundant parameterization, but it may also arise from an irregular yet perfectly injective parameterization: the model \(P_\theta = N(\theta^3, 1)\) has zero score at \(\theta_0 = 0\) while \(\theta \mapsto P_\theta\) remains one-to-one. Singular information is therefore best described as first-order degeneracy rather than as failure of identification.
Many semiparametric causal problems have a finite-dimensional target \(\psi = \Psi(P)\) and infinite-dimensional nuisance. In fully nonparametric observed-data models, the tangent space is often the whole space \(L_2^0(P)\); in restricted semiparametric models — for instance, those imposing a known treatment mechanism, a known missingness mechanism, or a structural restriction on \(P\) — the tangent space \(\mathcal{T}\) is a proper infinite-dimensional subspace of \(L_2^0(P)\).
Definition C.4 (First-Order Nuisance Tangent Space) Suppose \(\Psi\) is pathwise differentiable at \(P\) with pathwise derivative \(\dot\Psi_P : \mathcal{T} \to \mathbb{R}\). The nuisance tangent space of \(\Psi\) at \(P\) is the kernel of that derivative, \[\mathcal{T}_\eta \;:=\; \ker(\dot\Psi_P) \;=\; \{\, g \in \mathcal{T} \;:\; \dot\Psi_P(g) = 0 \,\}.\] Because \(\dot\Psi_P\) is continuous, \(\mathcal{T}_\eta\) is a closed linear subspace of \(\mathcal{T}\).
A regular submodel whose score lies in \(\mathcal{T}_\eta\) is called a nuisance submodel: it moves \(P\) without moving \(\psi\) to first order. Defining \(\mathcal{T}_\eta\) as a kernel rather than as the closed span of attainable nuisance scores is a deliberate choice, and the two are not automatically the same object. Every score of a nuisance submodel does lie in the kernel, so the closed span of such scores is contained in \(\mathcal{T}_\eta\); the reverse inclusion requires that every tangent direction annihilated by \(\dot\Psi_P\) be attainable — or at least approximable in \(L_2(P)\) — by scores of actual nuisance submodels. In sufficiently rich models, including the unrestricted nonparametric models used throughout Chapters 10–12, that local richness holds and the two descriptions agree. It is not automatic in a general model, and taking the kernel as the definition avoids leaving an attainability assumption unstated.
The nuisance tangent space is a closed subspace of \(\mathcal{T}\), and hence of \(L_2^0(P)\). Throughout the appendix, its orthogonal complement is taken in \(L_2^0(P)\): \[\mathcal{T}_\eta^\perp \;:=\; \{f \in L_2^0(P) : \E_P[f g] = 0 \text{ for all } g \in \mathcal{T}_\eta\}. \tag{C.4}\] With this convention, the Hilbert-space decomposition \[L_2^0(P) \;=\; \mathcal{T}_\eta \;\oplus\; \mathcal{T}_\eta^\perp \tag{C.5}\] holds automatically by the projection theorem, and provides the geometric structure underlying semiparametric efficiency.
Proposition C.1 (Every Influence Function Lies in \(\mathcal{T}_\eta^\perp\)) Let \(\Psi\) be pathwise differentiable at \(P\). Then every influence function \(\varphi \in \mathrm{IF}(\Psi, P)\) satisfies \(\varphi \in \mathcal{T}_\eta^\perp\).
Proof. Let \(\varphi \in \mathrm{IF}(\Psi, P)\). By Definition C.2 and the continuous extension described in Section C.3, \(\E_P[\varphi\, g] = \dot\Psi_P(g)\) for every \(g \in \mathcal{T}\). If in addition \(g \in \mathcal{T}_\eta = \ker(\dot\Psi_P)\), the right-hand side vanishes, so \(\E_P[\varphi\, g] = 0\). Hence \(\varphi \perp \mathcal{T}_\eta\), i.e. \(\varphi \in \mathcal{T}_\eta^\perp\). \(\square\)
Proposition C.1 is the key structural fact: being orthogonal to \(\mathcal{T}_\eta\) is automatic for influence functions, not a restriction. The distinguishing property of the efficient influence function, introduced in the next section, is instead that it is the unique influence function lying in the full tangent space \(\mathcal{T}\).
The abstract claim that gradients may be non-unique deserves a model in which one can watch it happen. The nonparametric examples above are the wrong place to look: there \(\mathcal{T} = L_2^0(P)\), so \(\mathcal{T}^\perp = \{0\}\) and the gradient is unique. Non-uniqueness requires a restricted model, and the smallest interesting one suffices.
Example C.3 (Non-Uniqueness in a Normal Location Model) Second reading.
Let \(Y \sim N(\mu, 1)\) with \(\mathcal{P} = \{N(\mu,1) : \mu \in \mathbb{R}\}\) and target \(\Psi(P_\mu) = \mu\). The model is one-dimensional, with score \(S(Y) = Y - \mu\) and \[\mathcal{T} \;=\; \mathrm{span}\{Y - \mu\} \;\subsetneq\; L_2^0(P).\] Write \(Z = Y - \mu \sim N(0,1)\). Since \(\partial_\mu \Psi(P_\mu) = 1\) and \(\E_P[Z^2] = 1\), the function \(\varphi_1(Y) = Y - \mu\) is a gradient. So, however, is \[\varphi_2(Y) \;=\; Y - \mu + c\left\{(Y-\mu)^2 - 1\right\} \qquad\text{for every } c \in \mathbb{R},\] because \(\E_P[(Z^2-1)Z] = \E_P[Z^3] - \E_P[Z] = 0\), so the added term is orthogonal to the only score in the model and contributes nothing to Equation C.3. Indeed \(\E_P[\varphi_2 S] = 1\) for all \(c\). The added term lies in \(\mathcal{T}^\perp\), exactly as Theorem C.1 (i) will require.
Projecting onto \(\mathcal{T}\) recovers the canonical gradient: \[\Pi[\varphi_2 \mid \mathcal{T}] \;=\; \frac{\langle \varphi_2,\, Z\rangle_P}{\|Z\|_P^2}\, Z \;=\; Z \;=\; \varphi_1,\] independently of \(c\). Comparing variances, \[\E_P[\varphi_2^2] \;=\; 1 + 2c^2 \;\geq\; 1 \;=\; \E_P[\varphi_1^2],\] with equality only at \(c = 0\) — the Pythagorean inequality of Theorem C.1 (iii) in the simplest possible setting. The efficiency bound is \(V^* = 1\), attained by the sample mean, and the extra term \(c\{(Y-\mu)^2-1\}\) is precisely the kind of variance inflation that projection removes.
The non-canonical gradients are not merely formal. Each is realized by an actual estimator: for \(c\) fixed, put \[\hat\mu_c \;=\; \bar Y \;+\; c\left\{\frac{1}{n}\sum_{i=1}^n (Y_i - \bar Y)^2 - 1\right\}.\] Because \(n^{-1}\sum_i (Y_i - \bar Y)^2 = n^{-1}\sum_i (Y_i - \mu)^2 - (\bar Y - \mu)^2\) and \(\sqrt{n}(\bar Y - \mu)^2 = O_P(n^{-1/2})\), \[\sqrt{n}(\hat\mu_c - \mu) \;=\; \frac{1}{\sqrt{n}}\sum_{i=1}^n \Bigl[(Y_i - \mu) + c\{(Y_i-\mu)^2 - 1\}\Bigr] \;+\; o_P(1),\] so \(\hat\mu_c\) is regular and asymptotically linear with influence function \(\varphi_2\), and its asymptotic variance is \(1 + 2c^2\). Only \(c = 0\), the sample mean, attains the bound. This matters because the connection between asymptotic linearity and gradients runs in one direction only — exhibiting a gradient does not by itself produce an estimator — and here the estimator can be written down.
Two lessons generalize. First, non-uniqueness is a statement about the model, not the functional: enlarging \(\mathcal{P}\) to the unrestricted dominated finite-variance model on \(\mathbb{R}\) shrinks \(\mathcal{T}^\perp\) to \(\{0\}\) and makes the gradient unique. Second, the non-canonical gradients are not wrong — each is a legitimate influence function of a regular asymptotically linear estimator — they are merely inefficient.
C.6 The Canonical Gradient and the Efficiency Bound
Proposition C.1 showed that every influence function lies in \(\mathcal{T}_\eta^\perp\). What distinguishes a single influence function within this set? The answer is orthogonality to \(\mathcal{T}^\perp\), equivalently membership in the full tangent space \(\mathcal{T}\): among all influence functions, there is a unique one lying in \(\mathcal{T}\), and it is the one with smallest variance.
Theorem C.1 (Canonical Gradient) Let \(\Psi\) be pathwise differentiable at \(P\), so that the set \(\mathrm{IF}(\Psi, P)\) of gradients is non-empty. Then:
- Any two influence functions differ by an element of \(\mathcal{T}^\perp\): for \(\varphi_1, \varphi_2 \in \mathrm{IF}(\Psi, P)\), \(\varphi_1 - \varphi_2 \in \mathcal{T}^\perp\).
- There exists a unique \(\varphi^* \in \mathrm{IF}(\Psi, P)\) with \(\varphi^* \in \mathcal{T}\), namely \[\varphi^* \;=\; \Pi[\varphi \mid \mathcal{T}] \qquad\text{for any }\varphi \in \mathrm{IF}(\Psi, P),\] where \(\Pi[\,\cdot \mid \mathcal{T}\,]\) denotes orthogonal projection in \(L_2^0(P)\) onto \(\mathcal{T}\).
- For every \(\varphi \in \mathrm{IF}(\Psi, P)\), \(\E_P[\varphi(O)^2] \geq \E_P[\varphi^*(O)^2]\), with equality iff \(\varphi = \varphi^*\) in \(L_2(P)\).
Proof. (i) For \(\varphi_1, \varphi_2 \in \mathrm{IF}(\Psi, P)\) and every score \(S\) of a regular submodel, \(\E_P[(\varphi_1 - \varphi_2) S] = \partial_\varepsilon\Psi|_0 - \partial_\varepsilon\Psi|_0 = 0\). By linearity and \(L_2\)-continuity, \(\E_P[(\varphi_1 - \varphi_2) g] = 0\) for every \(g \in \mathcal{T}\), i.e., \(\varphi_1 - \varphi_2 \in \mathcal{T}^\perp\).
Fix any \(\varphi \in \mathrm{IF}(\Psi, P)\) and set \(\varphi^* := \Pi[\varphi \mid \mathcal{T}]\), so \(\varphi - \varphi^* \in \mathcal{T}^\perp\). For any score \(S \in \mathcal{T}\), \[\E_P[\varphi^* S] \;=\; \E_P[\varphi S] - \E_P[(\varphi - \varphi^*) S] \;=\; \E_P[\varphi S] \;=\; \left.\frac{\partial}{\partial\varepsilon}\Psi(P_\varepsilon)\right|_0,\] so \(\varphi^* \in \mathrm{IF}(\Psi, P)\). Uniqueness: if \(\tilde\varphi^* \in \mathrm{IF}(\Psi, P) \cap \mathcal{T}\), then by (i) \(\varphi^* - \tilde\varphi^* \in \mathcal{T}^\perp \cap \mathcal{T} = \{0\}\).
Write \(\varphi = \varphi^* + (\varphi - \varphi^*)\) with the two summands orthogonal in \(L_2(P)\) (since \(\varphi^* \in \mathcal{T}\) and \(\varphi - \varphi^* \in \mathcal{T}^\perp\)). The Pythagorean identity gives \[\E_P[\varphi^2] \;=\; \E_P[(\varphi^*)^2] + \E_P[(\varphi - \varphi^*)^2] \;\geq\; \E_P[(\varphi^*)^2],\] with equality iff \(\varphi = \varphi^*\). \(\square\)
The figure below summarizes the geometry of the theorem in a single picture, and is worth consulting alongside Example C.3.
Definition C.5 (Efficient Influence Function and Efficiency Bound) The unique element \(\varphi^* \in \mathrm{IF}(\Psi, P) \cap \mathcal{T}\) is called the efficient influence function (EIF), or canonical gradient, of \(\Psi\) at \(P\). Its variance, \(V^*(\Psi, P) = \E_P[\varphi^*(O)^2]\), is the semiparametric efficiency bound for estimating \(\Psi(P)\) in the model \(\mathcal{P}\).
As stated, \(V^*(\Psi, P)\) is finite whenever it is defined at all, since \(\varphi^* \in L_2^0(P)\) by construction. If the pathwise derivative has no continuous \(L_2(P)\) representer, then \(\Psi\) is not pathwise differentiable at \(P\) and no bound exists in the sense of Definition C.5. It is convenient to record that situation through the extended-value convention \[V^*(\Psi, P) \;=\; +\infty \qquad\text{when no } L_2(P) \text{ canonical gradient exists}, \tag{C.6}\] and we use Equation C.6 below when discussing what weak overlap does to the ATE bound.
The efficiency bound has the following interpretation, which is the formal version of the heuristic statement in Chapter 10.
Theorem C.2 (Asymptotic Efficiency Bound) Let \(O_1, \ldots, O_n\) be i.i.d. from \(P\), and let \(\hat\psi_n\) be a regular asymptotically linear estimator of \(\Psi(P)\) in the model \(\mathcal{P}\). Then its asymptotic variance satisfies \[\mathrm{AVar}(\sqrt{n}(\hat\psi_n - \psi)) \;\geq\; V^*(\Psi, P),\] and the bound is achieved if and only if the influence function of \(\hat\psi_n\) is \(\varphi^*\) in \(L_2(P)\).
For regular asymptotically linear estimators, the variance bound in Theorem C.2 follows directly from Theorem C.1 (iii) applied to the estimator’s influence function, which is itself a gradient of \(\Psi\). This is the statement the appendix actually needs, and it is elementary given the projection theorem.
The Hájek–Le Cam convolution theorem extends the lower-bound interpretation beyond the asymptotically linear class, but its content should be stated carefully. Under the local asymptotic normality and regularity conditions the theorem requires, every weak limit law of a regular estimator has the convolution form \(N(0, V^*) * M\) for some probability measure \(M\); equivalently \(\sqrt{n}(\hat\psi_n - \psi) \rightsquigarrow Z + W\) with \(Z \sim N(0, V^*)\) independent of \(W \sim M\). If that limit law has a finite second moment, its variance is at least \(V^*\), and one recovers a variance comparison. Without a finite second moment the comparison by “taking variances” is unavailable — a regular estimator need not have a limit distribution with any moments at all — and the correct conclusion is the convolution statement itself, or the equivalent local asymptotic minimax bound. The limits are also in general subsequential. A precise treatment requires the local asymptotic normality framework of Vaart (1998) (Chapter 25). For our purposes the content is that no regular estimator improves on \(V^*\) in the convolution ordering, and that a regular asymptotically linear estimator achieves the bound exactly when its influence function equals the EIF.
C.7 Worked Example: The ATE Functional
This section makes the geometric machinery of the appendix concrete by carrying out the EIF derivation for the average treatment effect under the nonparametric observed-data model. The end product is the AIPW influence function of Chapter 10. The value of the derivation lies in showing how it arises directly from Theorem C.1 as the Riesz representer in \(L_2^0(P)\) of the pathwise derivative of \(\tau\) — recovering the AIPW formula from first principles rather than guessing it.
Setup. The observed data are \(O = (X, T, Y)\) with \(X\) pre-treatment covariates, \(T \in \{0, 1\}\) a treatment indicator, and \(Y\) the outcome. Under the identification assumptions of Chapters 4–5 — consistency, conditional exchangeability, and positivity — the ATE is identified with the observed-data functional \[\tau(P) \;=\; \E_P\!\left[\mu_1(X) - \mu_0(X)\right], \qquad \mu_t(X) \;=\; \E_P[Y \mid T = t,\, X].\] Write \[\pi_t(x) \;=\; P(T = t \mid X = x), \qquad \pi(x) \;:=\; \pi_1(x), \quad \pi_0(x) = 1 - \pi(x),\] for the propensity score, and \(\sigma_t^2(X) = \mathrm{Var}_P(Y \mid T = t, X)\) for the conditional outcome variances. The model \(\mathcal{P}\) is the unrestricted dominated model for the joint law of \((X, T, Y)\) subject to positivity, \(0 < \pi(X) < 1\) almost surely, so by the nonparametric example in Section C.5 the tangent space is \(\mathcal{T} = L_2^0(P)\).
Score factorization. The joint density factors as \[p(x, t, y) \;=\; p_X(x) \cdot p_{T \mid X}(t \mid x) \cdot p_{Y \mid T, X}(y \mid t, x).\] Differentiating \(\log p_\varepsilon\) along any regular submodel \(\{P_\varepsilon\}\) yields the additive score decomposition \[S(O) \;=\; S_X(X) \,+\, S_T(T \mid X) \,+\, S_Y(Y \mid T, X),\] where \(S_X = \partial_\varepsilon \log p_{X,\varepsilon}|_0\) and analogously for \(S_T\) and \(S_Y\). These satisfy the conditional mean-zero relations \[\E_P[S_X(X)] = 0, \qquad \E_P[S_T(T \mid X) \mid X] = 0, \qquad \E_P[S_Y(Y \mid T, X) \mid T, X] = 0.\] Define the closed subspaces of \(L_2^0(P)\): \[\begin{aligned} \mathcal{H}_X &= \{a(X) : a \in L_2(P_X),\; \E_P[a(X)] = 0\}, \\ \mathcal{H}_T &= \{h(T, X) \in L_2^0(P) \;:\; \E_P[h(T,X) \mid X] = 0\}, \\ \mathcal{H}_Y &= \{c(O) \in L_2^0(P) \;:\; \E_P[c(O) \mid T, X] = 0\}. \end{aligned}\] The treatment space \(\mathcal{H}_T\) is specified by a conditional mean-zero condition rather than by an explicit parametrization, because the two are not equivalent under weak positivity. For binary \(T\), every \(h \in \mathcal{H}_T\) can indeed be written as \[h(T, X) \;=\; b(X)\,\{T - \pi(X)\}, \qquad b(X) \;=\; \frac{h(1, X)}{1 - \pi(X)}, \tag{C.8}\] but the square-integrability requirement on \(b\) is weighted by the conditional variance of \(T\): \[\E_P[h(T,X)^2] \;=\; \E_P\!\left[b(X)^2\, \pi(X)\{1 - \pi(X)\}\right] \;<\; \infty.\] A treatment score can therefore be square-integrable while \(b \notin L_2(P_X)\), precisely in the regions where \(\pi(X)\) approaches \(0\) or \(1\). Defining \(\mathcal{H}_T\) as \(\{b(X)(T - \pi(X)) : b \in L_2(P_X)\}\) would exclude such directions and give a space strictly smaller than the treatment tangent space. Under strong overlap the weighted condition and \(b \in L_2(P_X)\) are equivalent, and the distinction is immaterial.
A short conditioning calculation shows that the three subspaces are pairwise orthogonal in \(L_2^0(P)\). For \(a(X) \in \mathcal{H}_X\) and \(h(T,X) \in \mathcal{H}_T\), \[\E_P\!\left[a(X)\, h(T,X)\right] \;=\; \E_P\!\left[a(X)\, \E_P\{h(T,X) \mid X\}\right] \;=\; 0,\] and the other two pairs follow analogously by conditioning on \((T, X)\). Joint spanning is explicit rather than merely assertible. Given any \(g \in L_2^0(P)\), set \[\begin{aligned} g_X &\;=\; \E_P[g \mid X], \\ g_T &\;=\; \E_P[g \mid T, X] \,-\, \E_P[g \mid X], \\ g_Y &\;=\; g \,-\, \E_P[g \mid T, X]. \end{aligned} \tag{C.9}\] These three terms sum to \(g\) by cancellation, and each lies in the advertised subspace: \(\E_P[g_X] = \E_P[g] = 0\) puts \(g_X\) in \(\mathcal{H}_X\); \(\E_P[g_T \mid X] = 0\) by the tower property puts \(g_T\) in \(\mathcal{H}_T\); and \(\E_P[g_Y \mid T, X] = 0\) puts \(g_Y\) in \(\mathcal{H}_Y\). Since Equation C.9 is a nested sequence of conditional expectations, it is precisely the orthogonal projection of \(g\) onto the three subspaces. Hence \[L_2^0(P) \;=\; \mathcal{H}_X \,\oplus\, \mathcal{H}_T \,\oplus\, \mathcal{H}_Y. \tag{C.10}\] The score decomposition \(S = S_X + S_T + S_Y\) above is precisely the projection decomposition of \(S \in L_2^0(P)\) relative to Equation C.10.
The pathwise derivative. The calculation below differentiates the conditional outcome means \(\mu_{t,\varepsilon}(x) = \E_{P_\varepsilon}[Y \mid T = t, X = x]\) under the conditional integral, which as in Example C.1 requires more than differentiability in quadratic mean. We therefore restrict attention to regular submodels satisfying the local conditional-moment conditions needed for that interchange; equivalently, one may establish the identity first for bounded-score, locally square-integrable paths and then extend it to the tangent-space closure by continuity. With that understood, differentiating \(\tau(P_\varepsilon) = \int (\mu_{1,\varepsilon}(x) - \mu_{0,\varepsilon}(x))\, p_{X,\varepsilon}(x)\, d\mu(x)\) at \(\varepsilon = 0\) and applying the product rule gives \[\begin{aligned} \left.\frac{\partial}{\partial\varepsilon}\tau(P_\varepsilon)\right|_0 &\;=\; \underbrace{\int \{\mu_1(x) - \mu_0(x)\}\, S_X(x)\, p_X(x)\, d\mu(x)}_{(I)} \\ &\qquad\;+\; \underbrace{\int \biggl(\partial_\varepsilon \mu_{1,\varepsilon}(x) - \partial_\varepsilon \mu_{0,\varepsilon}(x)\biggr)\bigg|_{\varepsilon = 0} p_X(x)\, d\mu(x)}_{(II)}. \end{aligned} \tag{C.11}\] No \(S_T\) term appears: \(\tau(P)\) is a functional of \(p_X\) and \(p_{Y \mid T, X}\) only, so perturbations of the treatment law do not affect \(\tau\) to first order. This already shows that \(\mathcal{H}_T \subset \mathcal{T}_\eta\).
For Term (I), since \(\E_P[S_X] = 0\) we may subtract any constant from \(\mu_1(X) - \mu_0(X)\) without changing the inner product; the choice \(\tau = \E_P[\mu_1(X) - \mu_0(X)]\) is the unique one that places the result in \(\mathcal{H}_X\), giving \[(I) \;=\; \E_P\!\left[\{\mu_1(X) - \mu_0(X) - \tau\}\, S_X(X)\right] \;=\; \langle \varphi_X,\, S_X \rangle_P,\] with \(\varphi_X(X) := \mu_1(X) - \mu_0(X) - \tau \in \mathcal{H}_X\).
For Term (II), fix \(t \in \{0, 1\}\) and compute \[\left.\partial_\varepsilon \mu_{t,\varepsilon}(x)\right|_0 \;=\; \int y\, S_Y(y \mid t, x)\, p(y \mid t, x)\, d\mu(y) \;=\; \E_P\!\left[(Y - \mu_t(x))\, S_Y(Y \mid t, X) \,\big|\, T = t,\, X = x\right],\] where the second equality uses \(\E_P[S_Y(Y \mid t, x) \mid T = t, X = x] = 0\) to subtract \(\mu_t(x)\). Multiplying by \(p_X(x)\) and integrating gives a conditional expectation that can be re-expressed as an unconditional one through the identity \[p_X(x)\, p(y \mid t, x) \;=\; \frac{p(x, T = t, y)}{\pi_t(x)},\] which follows from the factorization \(p(x, T = t, y) = p_X(x)\, \pi_t(x)\, p(y \mid t, x)\), with \(\pi_t\) as defined in the setup above. Applying this identity for \(t = 1\) and \(t = 0\) and combining, \[(II) \;=\; \E_P\!\left[\left\{\frac{T}{\pi(X)} (Y - \mu_1(X)) \,-\, \frac{1-T}{1-\pi(X)} (Y - \mu_0(X))\right\} S_Y(Y \mid T, X)\right] \;=\; \langle \varphi_Y,\, S_Y \rangle_P,\] with \[\varphi_Y(O) \;:=\; \frac{T\,(Y - \mu_1(X))}{\pi(X)} \;-\; \frac{(1-T)\,(Y - \mu_0(X))}{1-\pi(X)}.\] A short conditioning check confirms \(\E_P[\varphi_Y \mid T, X] = 0\), so \(\varphi_Y \in \mathcal{H}_Y\).
The canonical gradient. Combining the two terms and using the orthogonality of the three subspaces in Equation C.10, \[\left.\frac{\partial}{\partial\varepsilon}\tau(P_\varepsilon)\right|_0 \;=\; \langle \varphi^*,\, S \rangle_P, \qquad \varphi^*(O) \;:=\; \varphi_X(X) + \varphi_Y(O), \tag{C.12}\] for every score \(S = S_X + S_T + S_Y \in \mathcal{T} = L_2^0(P)\). By Definition C.2, \(\varphi^*\) is an influence function of \(\tau\). By construction \(\varphi^* \in \mathcal{H}_X \oplus \mathcal{H}_Y \subset L_2^0(P) = \mathcal{T}\), so \(\varphi^*\) is the canonical gradient of Theorem C.1 (ii). Writing it out explicitly, \[\varphi^*(O) \;=\; \mu_1(X) - \mu_0(X) - \tau \;+\; \frac{T\,(Y - \mu_1(X))}{\pi(X)} \;-\; \frac{(1-T)\,(Y - \mu_0(X))}{1-\pi(X)}, \tag{C.13}\] which is the AIPW influence function of Chapter 10. The figure below records the bookkeeping: each of the three score directions contributes its own piece of the gradient, and the treatment direction contributes nothing.
Dimension of \(\mathcal{T}_\eta^\perp\). The pathwise derivative \(\dot\tau_P : \mathcal{T} \to \mathbb{R}\), \(\dot\tau_P(S) = \partial_\varepsilon \tau(P_\varepsilon)|_0\), is by Equation C.12 a continuous linear functional on \(\mathcal{T} = L_2^0(P)\) with Riesz representer \(\varphi^*\), and its kernel is \(\mathcal{T}_\eta\) by Definition C.4. Suppose \(\dot\tau_P \neq 0\), equivalently \(\varphi^* \neq 0\) in \(L_2(P)\); by the efficiency-bound formula below this holds unless \(Y\) is almost surely a deterministic function of \((T,X)\) and the conditional effect \(\mu_1(X) - \mu_0(X)\) is almost surely constant. Then the codimension fact of Section C.1 gives \(\ker(\dot\tau_P)\) codimension one in \(L_2^0(P)\), whence \[\mathcal{T}_\eta^\perp \;=\; \mathcal{T} \cap \mathcal{T}_\eta^\perp \;=\; \mathrm{span}\{\varphi^*\},\] the one-dimensional subspace promised above. Two features of the unrestricted model are doing work here and should be named: the tangent space is all of \(L_2^0(P)\), so \(\mathcal{T}^\perp = \{0\}\) and the warning following Proposition C.1 is vacuous; and the derivative is non-zero. In the degenerate case \(\dot\tau_P = 0\) one has \(\varphi^* = 0\), \(\mathcal{T}_\eta = L_2^0(P)\), and \(\mathcal{T}_\eta^\perp = \{0\}\) rather than a line.
Reading off the efficiency bound. By Definition C.5, the semiparametric efficiency bound for estimating \(\tau(P)\) in the nonparametric model is \(V^*(\tau, P) = \E_P[\varphi^*(O)^2]\). Since \(\varphi_X \in \mathcal{H}_X\) and \(\varphi_Y \in \mathcal{H}_Y\) are orthogonal, the Pythagorean identity splits the bound into an outcome-noise term and a treatment-effect-heterogeneity term: \[V^*(\tau, P) \;=\; \E_P[\varphi_Y(O)^2] \;+\; \E_P[\varphi_X(X)^2].\] The second term is \(\E_P[\{\mu_1(X)-\mu_0(X)-\tau\}^2]\) by definition. For the first, note that \(T(1-T) = 0\) kills the cross product, so \[\varphi_Y(O)^2 \;=\; \frac{T\{Y - \mu_1(X)\}^2}{\pi(X)^2} \;+\; \frac{(1-T)\{Y - \mu_0(X)\}^2}{\{1-\pi(X)\}^2},\] and conditioning on \((T, X)\) and then on \(X\) gives, for \(t = 1\), \[\begin{aligned} \E_P\!\left[\frac{T\{Y-\mu_1(X)\}^2}{\pi(X)^2}\right] &\;=\; \E_P\!\left[\frac{\E_P[\,T\{Y-\mu_1(X)\}^2 \mid X\,]}{\pi(X)^2}\right] \\ &\;=\; \E_P\!\left[\frac{\pi(X)\,\sigma_1^2(X)}{\pi(X)^2}\right] \;=\; \E_P\!\left[\frac{\sigma_1^2(X)}{\pi(X)}\right], \end{aligned}\] with the analogous computation for \(t = 0\). Collecting terms, \[V^*(\tau, P) \;=\; \E_P\!\left[\frac{\sigma_1^2(X)}{\pi(X)} + \frac{\sigma_0^2(X)}{1-\pi(X)} + \{\mu_1(X) - \mu_0(X) - \tau\}^2\right], \tag{C.14}\] the classical semiparametric variance bound for the ATE (Hahn 1998; Robins et al. 1994). Comparing Equation C.14 with Equation C.7 confirms the claim made at the outset: the assumed moment condition is exactly the statement \(V^*(\tau, P) < \infty\), equivalently \(\varphi^* \in L_2^0(P)\). The \(1/\pi\) and \(1/(1-\pi)\) weights show directly how weak overlap inflates the bound and, in the limit, destroys pathwise differentiability altogether.
It remains to say what this does and does not establish. Equation Equation C.14 is a parameter-side result: it is a property of \(\tau\) and \(P\), computed before any estimator is proposed, and by Theorem C.2 no regular asymptotically linear estimator of \(\tau\) can have smaller asymptotic variance. It does not follow that every procedure labeled AIPW attains it. A feasible AIPW estimator attains the bound when the estimator-side conditions developed in Chapters 11 and 12 hold: consistent nuisance estimation, a product-rate condition controlling the second-order remainder, empirical-process control or cross-fitting to license the linearization, and a consistent variance estimator for inference. The EIF calculation supplies the target; those chapters supply the argument that a particular estimator hits it.
C.7.1 IPW versus AIPW: The Same Function in Two Models
Second reading.
Now that \(\mathcal{H}_X\), \(\mathcal{H}_T\), \(\mathcal{H}_Y\) and \(\varphi^*\) are all in hand, the status of the IPW estimating function \[U_{\mathrm{IPW}}(O) \;=\; \frac{T\, Y}{\pi(X)} \;-\; \frac{(1-T)\, Y}{1-\pi(X)} \;-\; \tau\] can be settled exactly. A first point to clear away: evaluated at the true propensity score, and provided the displayed terms are integrable, \(U_{\mathrm{IPW}}\) has mean zero whether or not the analyst happens to know \(\pi\). Unbiasedness is a property of \(P\), not of the analyst’s information. What knowledge of \(\pi\) changes is the model, and therefore the tangent space.
Throughout this remark we assume in addition that \[U_{\mathrm{IPW}} \in L_2^0(P), \qquad\text{equivalently}\qquad \E_P\!\left[\frac{\E_P[Y^2 \mid T = 1, X]}{\pi(X)} + \frac{\E_P[Y^2 \mid T = 0, X]}{1-\pi(X)}\right] < \infty. \tag{C.15}\] This is a genuinely stronger requirement than Equation C.7: the canonical gradient involves only the residual variances \(\sigma_t^2\) and the contrast \(\mu_1 - \mu_0\), whereas \(U_{\mathrm{IPW}}\) also carries the common level of \(\mu_1\) and \(\mu_0\), inverse-weighted. A function must lie in \(L_2^0(P)\) before it can be a gradient in any model, so Equation C.15 is needed before the question below is even well posed. Strong overlap together with \(\E_P[Y^2] < \infty\) suffices.
Direct algebra gives the exact decomposition \[U_{\mathrm{IPW}}(O) \;=\; \varphi^*(O) \;+\; \{T - \pi(X)\}\left\{\frac{\mu_1(X)}{\pi(X)} \,+\, \frac{\mu_0(X)}{1-\pi(X)}\right\}, \tag{C.16}\] in which the second term, written \[q(O) \;=\; \{T - \pi(X)\}\left\{\frac{\mu_1(X)}{\pi(X)} \,+\, \frac{\mu_0(X)}{1-\pi(X)}\right\},\] has the form \(b(X)\{T - \pi(X)\}\) and hence lies in the treatment-score space \(\mathcal{H}_T\) (see Equation C.8). Everything follows from where \(q\) sits.
Unrestricted treatment mechanism. Here \(\mathcal{H}_T \subset \mathcal{T}\), while \(\dot\tau_P\) vanishes on \(\mathcal{H}_T\), so \(\mathcal{H}_T \subset \mathcal{T}_\eta\). By Proposition C.1 every gradient of \(\tau\) is orthogonal to \(\mathcal{H}_T\). Since \(\varphi^* \perp \mathcal{H}_T\) and \(q \in \mathcal{H}_T\), \[\E_P[U_{\mathrm{IPW}}\, q] \;=\; \E_P[\varphi^* q] + \E_P[q^2] \;=\; \E_P[q^2].\] So \(U_{\mathrm{IPW}}\) is orthogonal to \(\mathcal{H}_T\) if and only if \(q = 0\) in \(L_2(P)\). Generically \(q \neq 0\), and then \(U_{\mathrm{IPW}}\) is not a gradient of \(\tau\) at all — efficient or otherwise. The exceptional case is worth stating because it is a genuine equivalence rather than a one-way implication: if \(q = 0\) almost surely, then \(U_{\mathrm{IPW}} = \varphi^*\) and the IPW function is the canonical gradient, even in the unrestricted model.
Known treatment mechanism. If \(\pi\) is known, no submodel may perturb \(p_{T \mid X}\), so treatment-score directions drop out of the tangent space: \(\mathcal{T} = \mathcal{H}_X \oplus \mathcal{H}_Y\) and \(\mathcal{H}_T \subset \mathcal{T}^\perp\). The second term of Equation C.16 now lies in \(\mathcal{T}^\perp\), so by Theorem C.1 (i) \(U_{\mathrm{IPW}}\) is a perfectly valid gradient of \(\tau\) — just not the canonical one. Projecting it onto \(\mathcal{T}\) deletes the \(\mathcal{H}_T\) component and returns \(\varphi^*\).
The efficiency bound is the same in both models, since \(\varphi^*\) is unchanged; knowing \(\pi\) does not help one estimate \(\tau\) more precisely. What changes is which functions count as gradients, and hence which estimators are regular. This is the cleanest illustration in the appendix of the fact that “efficient” is always efficient relative to a model, and it explains the otherwise puzzling result of Chapter 10 that estimating \(\pi\) can reduce the variance of the unaugmented Horvitz–Thompson IPW estimator relative to using the true \(\pi\).
C.8 Synthesis
The appendix has introduced several objects that are easy to conflate, partly because the literature calls more than one of them an “influence function.” The table below collects them.
| Object | Lives in | Defining property | Unique? |
|---|---|---|---|
| Score \(S\) | \(\mathcal{T} \subset L_2^0(P)\) | DQM derivative along a regular submodel (Definition C.1) | No: one per submodel |
| Pathwise gradient \(\varphi\) | \(L_2^0(P)\) | \(\dot\Psi_P(S) = \E_P[\varphi S]\) for every score (Definition C.2) | No, unless \(\mathcal{T} = L_2^0(P)\) |
| Canonical gradient \(\varphi^*\) | \(\mathcal{T} \cap \mathcal{T}_\eta^\perp\) | Riesz representer of \(\dot\Psi_P\) on \(\mathcal{T}\); equivalently \(\Pi[\varphi \mid \mathcal{T}]\) (Theorem C.1) | Yes |
| Estimator influence function | \(L_2^0(P)\) | \(\sqrt{n}(\hat\psi_n - \psi) = n^{-1/2}\sum_i \varphi(O_i) + o_P(1)\) | Yes, for a fixed estimator |
| Efficiency bound \(V^*\) | \([0, \infty)\), or \(+\infty\) by convention | \(\E_P[(\varphi^*)^2]\) (Definition C.5); finite whenever \(\Psi\) is pathwise differentiable, and \(+\infty\) under Equation C.6 when no \(L_2\) representer exists | Yes |
Three points are worth carrying away. First, an estimating function is not a gradient until it is normalized and shown to satisfy Equation C.3; unbiasedness is not enough. Second, non-uniqueness of gradients is a symptom of a restricted model, and projection onto \(\mathcal{T}\) is the cure. Third, every efficiency claim is relative to a model: the same function can be a valid gradient in one model and no gradient at all in a larger one, as Section C.7.1 shows for IPW.
C.9 Exercises
1. A regular tilted submodel. [(b)–(d) advanced] Let \(h \in L_2^0(P)\) be bounded and set \(p_\varepsilon(o) = \{1 + \varepsilon h(o)\}\, p(o)\) for \(|\varepsilon| < 1/\|h\|_\infty\).
- Verify that \(p_\varepsilon\) is non-negative and integrates to one, with no normalizing constant.
- Verify differentiability in quadratic mean Equation C.1 directly, and show that the score is \(h\). (Hint: expand \(\sqrt{1 + \varepsilon h}\) and use boundedness of \(h\) to control the remainder.)
- Conclude that the tangent space of the unrestricted dominated model is \(L_2^0(P)\), explaining where the density of bounded mean-zero functions in \(L_2^0(P)\) is used.
- Where does this linear-tilt construction fail when \(h\) is unbounded? Explain why this does not show that an unbounded \(h \in L_2^0(P)\) can never be the score of some other regular submodel.
2. Non-uniqueness in a normal location model. Let \(Y \sim N(\mu, 1)\) with \(\Psi(P_\mu) = \mu\), as in Example C.3.
- Identify \(\mathcal{T}\) and \(\mathcal{T}^\perp\).
- Show that \(\varphi_1(Y) = Y - \mu\) is a gradient.
- Show that \(\varphi_1(Y) + c\{(Y-\mu)^2 - 1\}\) is a gradient for every \(c \in \mathbb{R}\), and explain which property of the normal distribution makes this work.
- Project this gradient onto \(\mathcal{T}\) and compare variances, identifying the EIF.
- Now enlarge \(\mathcal{P}\) to the unrestricted dominated model of distributions on \(\mathbb{R}\) with finite variance, keeping \(\Psi(P) = \E_P[Y]\). What happens to the family of gradients, and why?
- Verify that \(\hat\mu_c\) of Example C.3 is asymptotically linear with influence function \(\varphi_2\), and compute its asymptotic variance directly.
3. Kernel and nuisance tangent space. Let \(\dot\Psi_P : \mathcal{T} \to \mathbb{R}\) be a non-zero continuous linear functional with representer \(\varphi^* \in \mathcal{T}\). Prove that \[\ker(\dot\Psi_P) \;=\; \{g \in \mathcal{T} : \E_P[\varphi^* g] = 0\}, \qquad \mathcal{T} \;=\; \ker(\dot\Psi_P) \,\oplus\, \mathrm{span}\{\varphi^*\}.\] Then state what each assertion becomes when \(\dot\Psi_P = 0\), and reconcile your answer with the codimension discussion of Section C.1.
4. Orthogonal decomposition for \((X, T, Y)\). For arbitrary \(g \in L_2^0(P)\) define \(g_X\), \(g_T\), \(g_Y\) as in Equation C.9.
- Show \(g = g_X + g_T + g_Y\) and that the three terms are pairwise orthogonal.
- For binary \(T\), show that any \(g_T\) takes the form \(b(X)\{T - \pi(X)\}\) and identify \(b\).
- [Advanced] Show that the correct integrability condition on \(b\) is \(\E_P[b(X)^2 \pi(X)\{1 - \pi(X)\}] < \infty\), and construct a \(P\) satisfying positivity for which some \(g_T \in \mathcal{H}_T\) has \(b \notin L_2(P_X)\).
- Show that under strong overlap the two conditions coincide.
5. IPW versus AIPW geometry. With \(U_{\mathrm{IPW}}\), \(q\), and \(\varphi^*\) as in Section C.7.1, and assuming \(U_{\mathrm{IPW}} \in L_2^0(P)\) as in Equation C.15:
- Verify Equation C.16 by direct algebra.
- Explain why \(q\) is a treatment-score direction.
- Show that \(U_{\mathrm{IPW}}\) is a gradient of \(\tau\) in the unrestricted model if and only if \(q = 0\) almost surely, and that otherwise it is not a gradient. Then explain why it is a valid non-canonical gradient in the known-propensity model.
- Verify that projecting \(U_{\mathrm{IPW}}\) onto the known-propensity tangent space returns \(\varphi^*\).
- Show that \(\E_P[U_{\mathrm{IPW}}^2] - V^*(\tau, P) = \E_P[q^2]\), and interpret the excess as a variance penalty.
- Exhibit a \(P\) satisfying Equation C.7 but violating Equation C.15, and explain what goes wrong.
6. Deriving the ATE efficiency bound. Starting from Equation C.13, prove Equation C.14. Then:
- Give sufficient conditions on \(\pi\) and the conditional distribution of \(Y\) for the bound to be finite.
- Construct a \(P\) satisfying \(0 < \pi(X) < 1\) almost surely for which the bound is infinite. What does Definition C.2 say about \(\tau\) at such a \(P\)? What qualitative instability should one expect from IPW and AIPW estimators there, and why is this an asymptotic warning rather than a universal finite-sample theorem?
- Explain why the bound does not depend on whether \(\pi\) is known.
7. Pointwise versus integrated functionals. [Advanced] Let \(X\) have a density that is continuous and positive near \(x_0\), and suppose the regression \(m_P(x) = \E_P[Y \mid X = x]\) has a specified continuous version near \(x_0\). Consider the pointwise functional \(\Psi_{\mathrm{pt}}(P) = m_P(x_0)\).
- Using bounded continuous bump-function tilts as in Example C.2, show that the score-to-derivative map for \(\Psi_{\mathrm{pt}}\) is not continuous in the \(L_2(P)\) norm. Explain why the tilts must be continuous and why a version of \(m_P\) must be fixed in advance.
- Conclude that no regular asymptotically linear estimator of \(\Psi_{\mathrm{pt}}\) exists in the unrestricted model, and state the additional local asymptotic minimax conclusion needed to rule out regular root-\(n\) estimation more generally.
- Contrast this with the integrated functional \(\tau(P) = \E_P[\mu_1(X) - \mu_0(X)]\). Explain why averaging over \(P_X\) smooths the target, and state the positivity and moment condition Equation C.7 under which its canonical gradient lies in \(L_2^0(P)\).
C.10 Bibliographic Notes
The modern formulation of pathwise differentiability and tangent spaces is developed in the monographs of Bickel et al. (1993) and Vaart (1998) (Chapter 25); the latter is the standard reference for the convolution theorem and local asymptotic normality underlying Theorem C.2. Tsiatis (2006) gives a treatment oriented specifically toward missing data and causal inference, and is a natural companion to the material developed here.
For the ATE functional specifically, Hahn (1998) derives the efficiency bound Equation C.14 and analyzes the role of the propensity score, including the point developed in Section C.7.1 that knowledge of \(\pi\) does not lower the bound; Robins et al. (1994) obtain the same influence function from the augmented estimating-equation perspective that Chapter 11 follows. The failure of pathwise differentiability for pointwise functionals (Example C.2) and the resulting slower-than-\(\sqrt{n}\) rates are treated in Vaart (1998) (§25.3).