2 DAGs and d-Separation
Readers who want a gentler introduction to conditional independence, the three basic path motifs, and the Markov-property interpretation may consult Appendix A before or alongside this chapter.
2.1 Directed Acyclic Graphs
DAGs allow us to translate qualitative causal assumptions into quantitative statistical restrictions. This single idea underlies everything in this chapter: once we draw a graph encoding our causal assumptions, d-separation tells us which conditional independences are implied by the graph and helps identify candidate adjustment sets for removing confounding.
2.1.1 Basic Definition
In a causal DAG each node represents a variable and each directed edge \(A \to B\) encodes that \(A\) is a direct cause of \(B\) relative to the variables included in the graph: the structural mechanism generating \(B\) depends nontrivially on \(A\) when the other parents of \(B\) are held fixed. A direct edge does not preclude additional mediated influence — in the mediation example below, \(T\) affects \(Y\) both directly and through \(M\). Strictly speaking, a drawn arrow represents genuine dependence on the parent only under a causal minimality convention, which we adopt informally throughout: edges are not drawn gratuitously.
The acyclicity condition is essential for the standard DAG semantics: it guarantees that the nodes admit a topological ordering, in which every parent precedes its children. Acyclicity alone, however, imposes no restriction on a joint distribution — by the chain rule, any density factorizes as \(p(v_1,\dots,v_k) = \prod_{i} p(v_i \mid v_1,\dots,v_{i-1})\) along any ordering. We say that a distribution \(P\) is Markov with respect to \(\Gcal\) if, along a topological ordering, the predecessors can be replaced by the parents: \[p(v_1,\dots,v_k) = \prod_{i=1}^{k} p(v_i \mid \Pa(v_i)).\] This factorization is a property of the pair \((\Gcal, P)\), not a consequence of acyclicity alone — it is precisely the substantive probabilistic restriction imposed by the graphical model, developed formally in Section 2.4. If directed feedback were present, the graph would not be a DAG, and the recursive semantics developed here would not apply without additional equilibrium or dynamic structure; cyclic systems are outside the scope of these notes.
2.1.2 Structural Relationships
Example 2.1 (The IV DAG) The canonical instrumental variable model involves four nodes: \(Z\) (instrument), \(T\) (treatment), \(Y\) (outcome), and \(U\) (unobserved confounder).
\(\Pa(T) = \{Z,U\}\), \(\Pa(Y) = \{T,U\}\), \(\Pa(Z) = \Pa(U) = \varnothing\). The descendants of \(Z\) are \(\{T, Y\}\); \(U\) and \(Z\) are non-descendants of each other.
2.2 The Three Motifs
A DAG encodes direct causal relationships through its edges, but variables can also be statistically related through longer routes that mix directed and reversed edges. The concept of a path formalizes these routes, and every intermediate node on a path plays exactly one of three structural roles: a link in a chain, the apex of a fork, or a collider. This section introduces the three motifs through concrete examples and then develops the blocking rules that determine whether each motif transmits or blocks statistical dependence, depending on which variables are conditioned upon. Section 2.3 then assembles these single-path rules into the d-separation criterion. The key insight is that conditioning can both close channels that were open (by blocking a chain or fork) and open channels that were closed (by activating a collider).
2.2.1 Paths and the Three Motifs
Every intermediate node \(V_m\) on a path is classified by the orientation of the two edges meeting it. The node \(V_m\) is a collider on the path if both arrows point into it: \(V_{m-1} \to V_m \leftarrow V_{m+1}\). Otherwise \(V_m\) is a non-collider. The non-collider cases have familiar names: in a chain (\(V_{m-1} \to V_m \to V_{m+1}\), or its mirror image \(V_{m-1} \leftarrow V_m \leftarrow V_{m+1}\)), \(V_m\) is a causal intermediary through which influence is transmitted; in a fork (\(V_{m-1} \leftarrow V_m \to V_{m+1}\)), \(V_m\) is a common cause of its two neighbors on the path. Chain and fork are useful motif names, but the formal d-separation rule of Section 2.3 distinguishes only colliders from non-colliders. The collider is the crucial asymmetric case whose behavior under conditioning is the opposite of the other two.
The three motifs correspond to three familiar causal patterns, one per configuration. Each story below should be read as an idealized three-node model: the stated conclusions hold for the displayed graph, and need not hold in the corresponding real-world setting, where omitted common causes, additional paths, or measurement error may intervene.
Chain: Smoking \(\to\) Tar \(\to\) Cancer. Let \(A\) = smoking, \(M\) = tar deposits in the lungs, \(B\) = lung cancer. The path \(A \to M \to B\) is a chain: smoking causes tar accumulation, which in turn causes cancer, so \(M\) is a causal intermediary. Marginally, the chain transmits association: heavy smokers have elevated cancer rates.
Fork: Poverty \(\to\) Poor Diet; Poverty \(\to\) Lack of Exercise. Let \(M\) = poverty, \(A\) = poor diet, \(B\) = lack of exercise. The path \(A \leftarrow M \to B\) is a fork: poverty is a common cause of both outcomes. Unconditionally, the fork also transmits association: diet quality and exercise are correlated — not because one causes the other, but because both are driven by poverty.
Collider: Accident \(\to\) Hospitalization \(\leftarrow\) Cancer. Let \(A\) = traffic accident, \(B\) = cancer, \(M\) = hospitalization. Both accidents and cancer independently cause hospitalization, so \(M\) is a collider: a common effect of its two neighbors. The collider is the odd one out already at the marginal level: in the general population, having a traffic accident tells you nothing about your cancer risk, \(A \indep B\). The path is closed by default.
What happens to each motif when the middle node is conditioned upon is the subject of the next subsection.
2.2.2 Blocking Rules
Conditioning on a node affects the paths it lies on differently depending on its structural role:
| Structure | Pattern | Effect of conditioning | Key intuition |
|---|---|---|---|
| Chain | \(A \to M \to B\) | Blocks the path | Causal transmission |
| Fork | \(A \leftarrow M \to B\) | Blocks the path | Common cause |
| Collider | \(A \to M \leftarrow B\) | Opens the path (conditioning on \(M\) or a descendant of \(M\)) | Common effect |
The critical asymmetry — colliders are closed by default and opened by conditioning, while chains and forks are open by default and closed by conditioning — is what makes collider bias so easy to overlook in practice. Returning to the three stories of Section 2.2.1 anchors each rule to a concrete setting.
Chain. If we condition on tar level (e.g., restrict to patients with the same measured tar burden), smoking no longer predicts cancer within that stratum: once the tar level is fixed, the causal channel is fully accounted for. Conditioning on the intermediary \(M\) blocks the path.
Fork. Conditioning on poverty status removes the spurious association: among people at the same level of poverty, the diet–exercise correlation disappears. Conditioning on the common cause \(M\) blocks the path.
Collider. Among hospitalized patients — conditioning on \(M\) — accident and cancer typically become negatively associated. If we learn that a hospitalized patient was not in a traffic accident, cancer or another serious illness becomes more likely as an explanation for the admission. Conditioning on the common effect \(M\) opens the path. (Opening the path licenses dependence; the negative sign reflects the monotone selection mechanism — either cause suffices to produce hospitalization — not the arrow pattern alone.) This is Berkson’s bias (Berkson 1946).
The blocking rules need not be taken on faith. The next example derives the chain rule directly from the Markov factorization stated in Section 2.1; the remark that follows explains how the remaining rules are obtained by the same kind of computation.
Example 2.2 (A Direct Probability Proof for the Chain) Consider the chain \(X_1 \to X_2 \to X_3\). By the Markov factorization, \[p(x_1, x_2, x_3) = p(x_1)\,p(x_2 \mid x_1)\,p(x_3 \mid x_2).\] Conditioning on \(X_2 = x_2\), for any \(x_2\) with \(p(x_2) > 0\): \[p(x_1, x_3 \mid x_2) = \frac{p(x_1, x_2, x_3)}{p(x_2)} = \frac{p(x_1)\,p(x_2 \mid x_1)\,p(x_3 \mid x_2)}{p(x_2)}.\] Rearranging, \[p(x_1, x_3 \mid x_2) = \frac{p(x_1)\,p(x_2 \mid x_1)}{p(x_2)}\,p(x_3 \mid x_2) = p(x_1 \mid x_2)\,p(x_3 \mid x_2).\] Hence \(X_1 \indep X_3 \mid X_2\). This derivation shows concretely how a graphical blocking statement becomes an ordinary conditional-independence identity in the distribution. The corresponding fork case \(X_1 \leftarrow X_2 \to X_3\) can be verified similarly and is left as an exercise.
2.3 d-Separation
The blocking rules of Section 2.2 settle the status of a single path. In any realistic DAG, however, two variables are typically joined by several paths at once, and association can flow along any one of them: a statistical relationship is ruled out by the graph only when every route is shut. d-Separation — the “d” stands for directional, because the rules depend on the orientation of the arrows along each path — is the criterion that performs exactly this aggregation. Given two nodes and a conditioning set \(\mathbf{S}\), it examines each path in turn and declares the nodes separated precisely when \(\mathbf{S}\) blocks all of the paths. It is a purely graphical computation: it can be carried out from the arrows alone, without reference to any probability distribution.
Why does this bookkeeping deserve a section of its own? Because d-separation is the bridge over which the graph speaks about data. The soundness theorem below (Theorem 2.1) shows that, for any distribution that is Markov with respect to the DAG, every d-separation statement translates into a genuine conditional independence of that distribution. The criterion thereby converts qualitative causal assumptions — arrows — into concrete probabilistic restrictions, and it is the engine behind the identification criteria of later chapters: the back-door criterion of Chapter 3 is, at bottom, a d-separation requirement in a modified graph.
2.3.1 The d-Separation Criterion
Before stating the formal criterion, it is useful to summarize the intuition: a path stays active unless we block a non-collider on it, or unless it contains a collider that has not been conditioned on (and has no conditioned-on descendant). The definition below simply turns this intuition into a rule that applies to any path in any DAG.
Definition 2.1 (d-Blocking and d-Separation) A path \(\pi\) is d-blocked by a set \(\mathbf{S}\) if:
- \(\pi\) contains a non-collider \(M\) (a chain or fork node) with \(M \in \mathbf{S}\); or
- \(\pi\) contains a collider \(A \to C \leftarrow B\) with \(C \notin \mathbf{S}\) and no descendant of \(C\) is in \(\mathbf{S}\).
Nodes \(X\) and \(Y\) are d-separated by \(\mathbf{S}\), written \((X \indep Y \mid \mathbf{S})_{\Gcal}\), if every path between \(X\) and \(Y\) in \(\Gcal\) is d-blocked by \(\mathbf{S}\). If some path is not d-blocked, \(X\) and \(Y\) are d-connected given \(\mathbf{S}\), written \((X \nindep Y \mid \mathbf{S})_{\Gcal}\). The subscript \(\Gcal\) always marks a statement about the graph; unsubscripted statements such as \(X \indep Y \mid \mathbf{S}\) refer to the distribution \(P\).
Example 2.3 (Applying the Criterion to the Three Toy Settings) We revisit the three motif stories of Section 2.2.1 and verify each blocking claim of Section 2.2.2 directly using the definition.
(i) Chain — Smoking \(\to\) Tar \(\to\) Cancer. Let \(A\) = smoking, \(M\) = tar, \(B\) = cancer. The only path between \(A\) and \(B\) is the chain \(A \to M \to B\).
- \(\mathbf{S} = \varnothing\): the path contains no conditioned-on non-collider, so clause (1) does not apply. The collider clause (2) is also irrelevant (there is no collider). The path is not d-blocked. \(\Rightarrow (A \nindep B)_{\Gcal}\) (generically, smokers have higher cancer rates).
- \(\mathbf{S} = \{M\}\): the chain node \(M \in \mathbf{S}\), so clause (1) applies and the path is d-blocked. \(\Rightarrow (A \indep B \mid M)_{\Gcal}\) (at fixed tar level, smoking no longer predicts cancer).
(ii) Fork — Poverty \(\to\) Poor Diet; Poverty \(\to\) Lack of Exercise. Let \(M\) = poverty, \(A\) = poor diet, \(B\) = lack of exercise. The only path is the fork \(A \leftarrow M \to B\).
- \(\mathbf{S} = \varnothing\): no non-collider on the path is in \(\mathbf{S}\), so the path is not d-blocked. \(\Rightarrow (A \nindep B)_{\Gcal}\) (generically, poor diet and lack of exercise are correlated through shared poverty).
- \(\mathbf{S} = \{M\}\): the fork node \(M \in \mathbf{S}\), so clause (1) applies and the path is d-blocked. \(\Rightarrow (A \indep B \mid M)_{\Gcal}\) (at fixed poverty status, the diet–exercise association disappears).
(iii) Collider — Accident \(\to\) Hospitalization \(\leftarrow\) Cancer. Let \(A\) = accident, \(M\) = hospitalization, \(B\) = cancer. The only path is the collider \(A \to M \leftarrow B\).
- \(\mathbf{S} = \varnothing\): the path contains collider \(M\) with \(M \notin \mathbf{S}\) and no descendant of \(M\) in \(\mathbf{S}\), so clause (2) applies and the path is d-blocked. \(\Rightarrow (A \indep B)_{\Gcal}\) (accidents and cancer are independent in the general population).
- \(\mathbf{S} = \{M\}\): now \(M \in \mathbf{S}\), so clause (2) no longer blocks the path. No other clause applies, so the path is open. \(\Rightarrow (A \nindep B \mid M)_{\Gcal}\) (generically, among hospitalized patients, ruling out an accident raises the probability of cancer — Berkson’s bias).
The following theorem is what gives d-separation its statistical force: it guarantees that every d-separation statement in the graph corresponds to a genuine conditional independence in the joint distribution.
Theorem 2.1 (Soundness of d-Separation (Pearl 2009, Theorem 1.2.4)) If \((X \indep Y \mid \mathbf{S})_{\Gcal}\), then \(X \indep Y \mid \mathbf{S}\) in every distribution that is Markov with respect to \(\Gcal\).
2.3.2 Practical Ways to Check d-Separation
The formal definition of d-separation can appear abstract when applied to larger graphs. Two complementary algorithmic viewpoints are useful in practice, both equivalent to the definition. On a first reading, path-by-path application of the definition — systematized by the tracing rules below — is the essential skill; the moral-graph method may be treated as optional second-pass material.
1. Active-Path Tracing. Imagine releasing a ball from \(X\) along each path and asking whether it can reach \(Y\), given that nodes in \(\mathbf{S}\) are observed (shaded). On any given path, the ball follows four local rules:
- At a non-collider (chain or fork node) not in \(\mathbf{S}\): ball passes through.
- At a non-collider in \(\mathbf{S}\): ball stops.
- At a collider not in \(\mathbf{S}\) (and no descendant in \(\mathbf{S}\)): ball stops.
- At a collider in \(\mathbf{S}\) (or with a descendant in \(\mathbf{S}\)): ball passes through.
If the ball can reach \(Y\) from \(X\) along some path, then \(X\) and \(Y\) are d-connected given \(\mathbf{S}\): the graph does not imply \(X \indep Y \mid \mathbf{S}\). If every path is stopped, then \((X \indep Y \mid \mathbf{S})_{\Gcal}\). These rules restate the d-blocking criterion path by path; a related, direction-sensitive search algorithm — Bayes-Ball, which visits each node at most a bounded number of times rather than enumerating paths — is developed by Shachter (1998).
2. Moral Graph Transformation. There is also a graph-transformation approach, which connects DAG models to undirected graphical models (Markov random fields).
- Restrict the DAG to the ancestral set of \(\{X, Y\} \cup \mathbf{S}\): keep the nodes \(X\), \(Y\), and \(\mathbf{S}\) themselves together with all of their ancestors, and discard everything else.
- Moralize the graph: for every collider \(A \to C \leftarrow B\) in the ancestral subgraph, add an undirected edge \(A - B\) (“marry the parents”).
- Remove all edge directions to obtain an undirected graph.
- Delete all nodes in \(\mathbf{S}\) (and their incident edges).
\(X\) and \(Y\) are d-separated by \(\mathbf{S}\) in \(\Gcal\) if and only if they are disconnected in the resulting undirected graph. The moralization step is what makes the collider rule visible in the undirected representation: without adding the \(A\)–\(B\) edge, the path through \(C\) would appear disconnected even when \(C\) is conditioned upon, which would give the wrong answer. A full treatment is in Lauritzen (1996, Ch. 3).
These two viewpoints — active-path tracing and moral-graph separation — are equivalent algorithmic ways of evaluating the d-separation criterion. Students are encouraged to use whichever they find most natural, and to cross-check with the other when in doubt.
Example 2.4 (Checking d-Separation in a Four-Node DAG) Consider the DAG with edges \(Z \to T\), \(U \to T\), \(T \to Y\), \(U \to Y\).
Query 1. Is \((Z \indep U)_{\Gcal}\), i.e. \(\mathbf{S} = \varnothing\)?
Path tracing. The only path from \(Z\) to \(U\) is \(Z \to T \leftarrow U\). Node \(T\) is a collider with \(T \notin \mathbf{S}\) and no descendant of \(T\) in \(\mathbf{S}\), so the ball is stopped at \(T\). No ball can reach \(U\), hence \((Z \indep U)_{\Gcal}\).
Moral graph. (1) Ancestral set of \(\{Z, U\} \cup \mathbf{S} = \{Z, U\}\): neither \(Z\) nor \(U\) has any parents, so the ancestral set is \(\{Z, U\}\) itself; nodes \(T\) and \(Y\) are discarded. (2) Moralize the subgraph \(\{Z, U\}\): it has no edges and no colliders, so nothing is added. (3) Undirected graph: isolated nodes \(Z\) and \(U\). (4) Delete \(\mathbf{S} = \varnothing\): nothing changes. \(Z\) and \(U\) are disconnected \(\Rightarrow (Z \indep U)_{\Gcal}\). Note that the key step is restricting to the ancestral set: this removes \(T\) from the graph before moralization, so the collider \(T\) never has a chance to create a moral edge between \(Z\) and \(U\).
Query 2. Is \((Z \indep U \mid T)_{\Gcal}\), i.e. \(\mathbf{S} = \{T\}\)?
Path tracing. Same path \(Z \to T \leftarrow U\). Now \(T \in \mathbf{S}\): the collider is observed, so the ball passes through \(T\). The path is open, hence \((Z \nindep U \mid T)_{\Gcal}\).
Moral graph. (1) Ancestral set of \(\{Z, U\} \cup \mathbf{S} = \{Z, U, T\}\): the parents of \(T\) are \(Z\) and \(U\), both already in the set; node \(Y\) is discarded. (2) Moralize: the subgraph contains the collider \(Z \to T \leftarrow U\), so add the moral edge \(Z - U\). (3) Undirected graph: edges \(Z - T\), \(U - T\), \(Z - U\). (4) Delete \(\mathbf{S} = \{T\}\) and its incident edges. Remaining graph: single edge \(Z - U\). \(Z\) and \(U\) are connected \(\Rightarrow (Z \nindep U \mid T)_{\Gcal}\). The moralization step is decisive: it adds the edge \(Z - U\) before \(T\) is deleted, ensuring the dependence induced by conditioning on the collider is visible in the undirected graph.
Comparing the two queries: \(Z\) and \(U\) are marginally independent (the instrument is exogenous) but become dependent once we condition on \(T\). This is collider bias — of which Berkson’s bias is the classic instance — treated in full in Section 2.5.
Example 2.5 (Practice DAG — Eight-Node Graph) Consider the DAG below with nodes \(W\), \(X_1\), \(T\), \(M_1\), \(M_2\), \(Y\), \(C\), and \(D\).
For each query below, state whether the independence holds and identify every relevant path and its blocking status.
(1) Is \(X_1 \indep Y\)? No. There are four paths from \(X_1\) to \(Y\). Two are open: \(X_1 \leftarrow W \to Y\) (fork at \(W\), unblocked) and \(X_1 \leftarrow W \to T \to M_1 \to Y\) (fork at \(W\), then chain; also open). Two pass through the collider at \(C\): \(X_1 \to C \leftarrow M_2 \leftarrow M_1 \to Y\) and \(X_1 \to C \leftarrow M_2 \leftarrow M_1 \leftarrow T \leftarrow W \to Y\). With \(\mathbf{S} = \varnothing\), neither \(C\) nor any descendant of \(C\) is in \(\mathbf{S}\), so both collider paths are blocked. One open path suffices: \((X_1 \nindep Y)_{\Gcal}\).
(2) Is \(X_1 \indep Y \mid W\)? Yes. Both open paths above pass through the fork \(W\); conditioning on \(W\) blocks them. The two remaining paths each contain the collider at \(C\) with \(C \notin \{W\}\) (and no descendant of \(C\) in \(\{W\}\)), so both remain blocked. All paths blocked \(\Rightarrow (X_1 \indep Y \mid W)_{\Gcal}\).
(3) Is \(T \indep M_2 \mid M_1\)? Yes. There are four paths between \(T\) and \(M_2\). The direct path \(T \to M_1 \to M_2\) is a chain with \(M_1 \in \{M_1\}\) — blocked. The path \(T \leftarrow W \to X_1 \to C \leftarrow M_2\) has a collider at \(C\) with \(C \notin \{M_1\}\) — blocked. The path \(T \leftarrow W \to Y \leftarrow M_1 \to M_2\) has a collider at \(Y\) (\(W \to Y \leftarrow M_1\)) with \(Y \notin \{M_1\}\) — blocked (and also blocked at the fork node \(M_1 \in \{M_1\}\)). Finally, the path \(T \to M_1 \to Y \leftarrow W \to X_1 \to C \leftarrow M_2\) is blocked at the chain node \(M_1 \in \{M_1\}\) (and also at the colliders \(Y\) and \(C\)). All paths blocked \(\Rightarrow (T \indep M_2 \mid M_1)_{\Gcal}\).
(4) Does conditioning on \(C\) open a collider path between \(X_1\) and \(M_2\)? Yes. \(C\) is a collider on the path \(X_1 \to C \leftarrow M_2\), which is blocked marginally and is opened by conditioning on \(C\). Note, however, that \(X_1\) and \(M_2\) are not marginally d-separated: the path \(X_1 \leftarrow W \to T \to M_1 \to M_2\) is open even without any conditioning. A cleaner comparison conditions on \(W\) throughout. First, \((X_1 \indep M_2 \mid W)_{\Gcal}\): conditioning on \(W\) blocks the fork path above, and the two collider paths remain blocked. Second, \((X_1 \nindep M_2 \mid \{W, C\})_{\Gcal}\): adding \(C\) to the conditioning set opens \(X_1 \to C \leftarrow M_2\). Conditioning on the collider thus introduces a new, non-causal source of association — it does not create the only connection between \(X_1\) and \(M_2\), but it destroys an independence that held given \(W\) alone.
(5) Does conditioning on \(D\) open the path \(X_1 \to C \leftarrow M_2\)? Yes. \(D\) is a descendant of \(C\). By clause (2) of the d-separation definition, a collider path is unblocked whenever the collider or any of its descendants is in the conditioning set. Since \(D \in \{D\}\), the collider at \(C\) is activated, opening the path \(X_1 \to C \leftarrow M_2\) even though \(C\) itself is not conditioned on.
2.4 The Markov Property and Factorization
Up to this point, we have used DAGs qualitatively, to decide which paths are open or blocked. We now connect the graph to probability algebra: the same parent structure that governs d-separation also determines how the joint distribution factorizes. Appendix A provides a gentler preview of the same ideas.
Without structural assumptions, any joint distribution can always be written by repeated conditioning, but that generic representation is often too high-dimensional to reveal much structure. A DAG becomes statistically meaningful because, together with the Markov property, it replaces the generic factorization by a sparse one involving only the parents of each node. This is the sense in which a graphical model is not merely a picture: it imposes probabilistic structure on the joint distribution. Whether that structure is empirically testable depends on which variables are observed; implications involving only observed variables yield conditional independences that can be checked against data, while those involving latent variables generally cannot.
Under Markov compatibility — the assumption that \(P\) satisfies Equation 2.1 for the graph at hand — the parent sets determine the sparse factorization of the joint law; the factorization, not the picture alone, is the bridge between causal structure and probability.
What does the equivalence buy in the smallest possible case? For the three-node chain, the global property contains exactly one substantive statement, \(X_1 \indep X_3 \mid X_2\) — and the factorization delivers it in the two lines already carried out above. We record the result for reference: it is the miniature version of Theorem 2.1 that the reader has proved by hand.
Proposition 2.1 (Conditional Independence in the Three-Node Chain) Let \(X_1 \to X_2 \to X_3\) be a chain with Markov factorization \(p(x_1, x_2, x_3) = p(x_1)\,p(x_2 \mid x_1)\,p(x_3 \mid x_2)\). Then \(X_1 \indep X_3 \mid X_2\). The analogous result for the fork \(X_1 \leftarrow X_2 \to X_3\) is left as an exercise.
2.5 Collider Bias and the IV DAG
Throughout the d-separation analyses of Section 2.3, the collider was the asymmetric case: closed by default, and opened — often unintentionally — by conditioning. This section studies the resulting phenomenon — collider bias — systematically: first by defining collider bias and illustrating it through Berkson-type selection, and then through a full d-separation analysis of the instrumental-variables DAG, where a collider is simultaneously the central danger for the naive analyst and, in later chapters, part of the identification strategy itself.
2.5.1 Berkson’s Bias
2.5.2 Full d-Separation Analysis of the IV DAG
We work through the IV DAG with edges \(Z \to T\), \(T \to Y\), \(U \to T\), \(U \to Y\). This is the same graph as in Example 2.4, but now \(U\) is treated as unobserved. That assumption is precisely what earns the name “IV DAG”: because \(U\) cannot be conditioned on, the back-door path \(T \leftarrow U \to Y\) cannot be blocked by adjustment, and the instrument \(Z\) motivates an alternative identification strategy. The graph alone does not nonparametrically identify \(P(y \mid \doop(T{=}t))\); Chapter 7 introduces the additional assumptions — such as linearity or monotonicity — under which particular IV estimands are identified.
Before listing the paths, it is worth classifying them by type, since students often conflate causal and non-causal sources of dependence.
| Endpoints | Path | Type | Why |
|---|---|---|---|
| \(Z \leftrightarrow Y\) | \(Z \to T \to Y\) | Directed causal | Directed from \(Z\) to \(Y\); carries the IV signal |
| \(Z \leftrightarrow Y\) | \(Z \to T \leftarrow U \to Y\) | Non-causal, collider-blocked | Blocked by collider at \(T\) unless \(T\) (or a descendant) is conditioned on |
| \(Z \leftrightarrow U\) | \(Z \to T \leftarrow U\) | Non-causal, collider-blocked | Collider at \(T\); dormant unless \(T\) conditioned on |
| \(Z \leftrightarrow U\) | \(Z \to T \to Y \leftarrow U\) | Non-causal, collider-blocked | Collider at \(Y\); dormant unless \(Y\) (or a descendant) conditioned on |
The key distinction: a causal path from \(Z\) to \(Y\) is one on which every arrow points from \(Z\) toward \(Y\); every other path is non-causal. A non-causal path does not represent a mechanism by which \(Z\) produces \(Y\), but it can transmit statistical association. Non-causal paths have no single default blocking status: a back-door fork such as \(T \leftarrow U \to Y\) is open without any conditioning, whereas the three non-causal paths in the present IV DAG each contain a collider (\(T\) on the first two, \(Y\) on the third) and are blocked unless the corresponding collider, or one of its descendants, is conditioned upon. With this in mind, the four questions become much easier to diagnose.
Question 1. Is \(Z \indep Y\)? There are two paths between \(Z\) and \(Y\). The directed path \(Z \to T \to Y\) is open, since it is a chain and we are not conditioning on \(T\). The path \(Z \to T \leftarrow U \to Y\) contains a collider at \(T\); since we are not conditioning on \(T\) or any descendant of \(T\), that path is blocked. Because the causal path remains open, \((Z \nindep Y)_{\Gcal}\).
Question 2. Is \(Z \indep Y \mid T\)? The causal path \(Z \to T \to Y\) is blocked because conditioning on \(T\) blocks the chain. The path \(Z \to T \leftarrow U \to Y\) contains a collider at \(T\), and because we are conditioning on \(T\), that collider is opened. \(\Rightarrow (Z \nindep Y \mid T)_{\Gcal}\).
Question 3. Is \(Z \indep Y \mid \{T, U\}\)? The causal path is blocked at \(T\). The path \(Z \to T \leftarrow U \to Y\) is opened at the collider \(T\), but the same path then passes through \(U\), and because we are also conditioning on \(U\), the path is blocked there. Hence both paths are blocked, so \((Z \indep Y \mid \{T, U\})_{\Gcal}\).
Question 4. Is \(Z \indep U\)? There are two paths between \(Z\) and \(U\): \(Z \to T \leftarrow U\) and \(Z \to T \to Y \leftarrow U\). The first contains a collider at \(T\); the second contains a collider at \(Y\). Since we are not conditioning on \(T\), on \(Y\), or on any descendant of either, both colliders remain closed, so both paths are blocked. Hence \((Z \indep U)_{\Gcal}\).
Taken together, the four questions give a graphical account of the structure behind the three IV assumptions, together with one critical warning.
Relevance. Question 1 shows that the graph places \(Z\) on an open causal path to \(Y\) through \(T\). Relevance, however, concerns the first stage: it requires that changing \(Z\) actually shifts the distribution of \(T\) — the edge \(Z \to T\) together with a nonzero first-stage effect. A drawn arrow is compatible with an arbitrarily weak (or, absent a minimality convention, even null) effect, so relevance is a substantive condition beyond the displayed topology.
Exogeneity. Question 4 expresses exogeneity: every path from \(Z\) to the confounder \(U\) is a dormant collider path, so the graph implies \(Z \indep U\).
Exclusion. Question 3 is the d-separation footprint of the exclusion restriction. The restriction itself is structural: every directed path from \(Z\) to \(Y\) must pass through \(T\). In this four-node graph, that is equivalent to \(Z\) being absent from the mechanism generating \(Y\); in a richer graph, absence from \(Y\)’s own structural equation would rule out only a direct \(Z \to Y\) edge, and the path formulation is the general one. The conditional independence \((Z \indep Y \mid \{T, U\})_{\Gcal}\) is an implication of that structure, not its definition. Note also that verifying the corresponding probabilistic independence directly would require observing \(U\), which is unavailable by assumption. The IV strategy instead uses exclusion and exogeneity jointly: in a centered linear outcome model without additional covariates, the two assumptions yield the orthogonality restriction \(\E[Z \varepsilon] = 0\), where \(\varepsilon\) is the structural disturbance in the outcome equation — exclusion keeps \(Z\) out of that equation, and exogeneity makes \(Z\) orthogonal to the latent determinants collected in \(\varepsilon\). Chapter 7 introduces observed covariates and states the covariate-adjusted version, \(\E[\varepsilon Z \mid X] = 0\), precisely.
Collider warning. Question 2 is the warning: conditioning on \(T\) alone — as a naive analyst might do — blocks the directed path \(Z \to T \to Y\) and opens the non-causal path \(Z \to T \leftarrow U \to Y\). The remaining association between \(Z\) and \(Y\) within levels of \(T\) is therefore collider-induced and non-causal; it contains no open causal component from \(Z\) to \(Y\).
It is important to recognize that the IV DAG encodes the substantive assumptions behind instrument validity graphically — relevance, exogeneity, and exclusion — but does not make them testable in any strong sense. Some observed-data implications of the IV DAG may be checked for compatibility with data, but the core IV assumptions are not generally fully testable from observational data alone.
2.6 Worked Example: The Education–Earnings DAG
This section is the template for how we will use DAGs throughout the course. Starting from a single causal graph, we move step by step through the full pipeline: first specify the DAG, then read off the factorization, then use d-separation to identify the implied conditional independences, and finally interpret the result in causal terms. Later chapters will follow the same pattern, but with richer identification questions.
Practice: d-separation workflow. Before working through the example below, return to Example 2.5 and rework each of the five queries from scratch, without looking at the answers. For each query, follow the same three steps: (1) list every path between the two nodes of interest; (2) classify every intermediate node on each path as a chain, fork, or collider; and (3) determine which paths are blocked or open after conditioning on the given set. This is exactly the procedure that d-separation always requires, in graphs of any size.
The causal story. Education (\(E\)) affects earnings (\(Y\)); family background (\(B\)) is a common cause of both education and earnings; neighborhood (\(N\)) affects education but has no direct effect on earnings. All four displayed variables are observed. In addition, we assume causal sufficiency relative to these variables: no common cause of any pair among \(N\), \(B\), \(E\), and \(Y\) has been omitted. This is a substantive simplifying assumption, not a consequence of observing the displayed variables — in reality, neighborhood and family background are often associated through broader socioeconomic factors, which would constitute exactly such an omitted common cause; we exclude it to keep the d-separation analysis tractable.
Step 1 — Markov factorization. The parents are \(\Pa(N) = \Pa(B) = \varnothing\), \(\Pa(E) = \{N, B\}\), and \(\Pa(Y) = \{E, B\}\). The Markov factorization is therefore \[p(n, b, e, y) \;=\; p(n)\,p(b)\,p(e \mid n, b)\,p(y \mid e, b).\] Every factor on the right-hand side involves only observed variables, so in principle each term is estimable from data. This is the starting point for identification: the full joint distribution is expressed in terms of estimable quantities.
Step 2 — d-Separation. There are two paths between \(N\) and \(Y\): Path 1 is \(N \to E \to Y\) (a chain through \(E\)); Path 2 is \(N \to E \leftarrow B \to Y\) (a collider at \(E\), followed by the fork leg \(B \to Y\)). We examine four queries.
Query (i): Is \((N \indep Y)_{\Gcal}\)? Path 1 is a chain with nothing conditioned on, so it is open. Path 2 has a collider at \(E\), and since \(E\) is not conditioned on, that path is blocked. Because Path 1 remains open, \((N \nindep Y)_{\Gcal}\).
Query (ii): Is \((N \indep Y \mid E)_{\Gcal}\)? Path 1 is a chain, and conditioning on \(E\) blocks it. Path 2 has a collider at \(E\), and conditioning on \(E\) opens it. The activated path \(N \to E \leftarrow B \to Y\) continues through \(B\), which is not conditioned on, so the path is open. Hence \((N \nindep Y \mid E)_{\Gcal}\). This is precisely the collider bias of Section 2.5: conditioning on \(E\) opens the path \(N \to E \leftarrow B \to Y\), thereby inducing a spurious association between \(N\) and \(Y\) through \(B\).
Query (iii): Is \((N \indep Y \mid \{E, B\})_{\Gcal}\)? Path 1 is blocked by conditioning on \(E\). Path 2 is opened at the collider \(E\), but then blocked at \(B\) because \(B\) is also conditioned on. Both paths are blocked: \((N \indep Y \mid E, B)_{\Gcal}\).
Query (iv): Is \((N \indep B)_{\Gcal}\)? The only path is \(N \to E \leftarrow B\), which has a collider at \(E\). Since \(E\) is not conditioned on, the path is blocked. Therefore \((N \indep B)_{\Gcal}\).
Step 3 — Conditional Independence. We now move from the graphical analysis to the probabilistic implications, being careful about the direction of the inference. By Theorem 2.1, the graph entails the two independence statements \[N \indep B, \qquad N \indep Y \mid E, B.\] The two remaining graphical conclusions are d-connections: \(N\) and \(Y\) are d-connected marginally, and also given \(E\). These mean only that the graph does not imply the corresponding independences; they become dependence predictions, \[N \nindep Y, \qquad N \nindep Y \mid E,\] under the additional faithfulness (or generic-parameter) reading discussed in the soundness remark of Section 2.3.1.
Because all four variables are observed, all four empirical relationships can in principle be investigated. Under the Markov assumption alone, however, only violations of the two entailed conditional independences contradict the graphical model — and agreement with those restrictions does not by itself verify the causal graph. The d-connection given \(E\) is especially important as a warning to practitioners. If the target is the total causal effect of neighborhood \(N\) on earnings \(Y\), conditioning on education \(E\) blocks the mediated causal path \(N \to E \to Y\) and opens the collider path \(N \to E \leftarrow B \to Y\): the adjustment distorts the very quantity being estimated rather than removing bias.
Step 4 — Identification. We wish to identify \(P(y \mid \doop(E{=}e))\), the distribution of earnings under an intervention that sets education to \(e\). The back-door criterion (Chapter 3) requires an adjustment set \(\mathbf{S}\) that (i) blocks every back-door path from \(E\) to \(Y\) and (ii) contains no descendant of \(E\).
The only back-door path is \(E \leftarrow B \to Y\), a fork at \(B\). Consider two candidate adjustment sets:
- \(\mathbf{S} = \{B\}\): this blocks the path \(E \leftarrow B \to Y\), and \(B\) is not a descendant of \(E\). Hence \(\{B\}\) is valid.
- \(\mathbf{S} = \{N\}\): this does not block the path \(E \leftarrow B \to Y\), so \(\{N\}\) is invalid.
This illustrates why every back-door path must be blocked: although \(N\) is upstream of \(E\), conditioning on \(N\) does nothing to close the confounding fork \(E \leftarrow B \to Y\).
As a preview of the back-door adjustment formula derived rigorously in Chapter 3, with \(\mathbf{S} = \{B\}\) the interventional distribution takes the form \[P(y \mid \doop(E{=}e)) \;=\; \sum_{b} P(y \mid e, b)\,P(b).\] Every term on the right-hand side involves only the observed distribution, so the causal effect is identified entirely from observational data — the payoff of the graphical analysis. Beyond the path criterion, the formula also relies on the causal interpretation of the DAG with its intervention semantics, and on a positivity condition: for discrete \(E\), \(P(E{=}e \mid B{=}b) > 0\) for every relevant \(b\); for continuous \(E\), the intervention level \(e\) must lie in the conditional support of \(E\) given \(B = b\). (For continuous \(B\), the sum over \(b\) is likewise replaced by an integral with respect to the distribution of \(B\).) The path analysis identifies \(B\) as the appropriate adjustment variable; Chapter 3 shows, using intervention graphs, why this graphical condition yields the standardization formula above.
2.7 The Big Picture
The concepts developed in this chapter form the graphical foundation for the identification and estimation theory developed later in the course. The central logic runs from a qualitative causal graph to an identification formula for an interventional quantity:
Each arrow represents a distinct step in causal reasoning. First, we encode substantive assumptions in a DAG. Second, we use d-separation to read off the conditional independence structure implied by that graph. Third, under the Markov property, we translate those graphical statements into probabilistic restrictions. Fourth, adding the intervention semantics developed in Chapter 3 — the causal reading of arrows, modularity, and d-separation in suitably modified graphs — we use those restrictions to derive identification formulas for interventional quantities such as \(P(y \mid \doop(T{=}t))\). Observational conditional independences alone do not yield causal identification; the intervention step is what carries the causal content. Section 2.6 illustrates the first three steps in full and previews the fourth through the back-door formula.
In more abstract form, the logic of the course runs \[\begin{aligned} \text{substantive causal assumptions} &\;\Longrightarrow\; \text{causal DAG, with intervention semantics (Ch. 3)}\\ &\;\Longrightarrow\; \text{Markov factorization and d-separation relations}\\ &\;\Longrightarrow\; \text{conditional independences under the Markov property}\\ &\;\Longrightarrow\; \text{identification formulas in later chapters}. \end{aligned}\] This chapter establishes the graph syntax, the Markov semantics, and the d-separation machinery in this chain; the intervention step is the subject of Chapter 3. Chapters 3 and beyond use the same graphical machinery to derive specific identification results, including back-door adjustment, front-door adjustment, and do-calculus formulas.
2.8 Summary
DAGs as causal structure. A DAG is a directed acyclic graph whose nodes represent variables and whose directed edges represent direct causal relationships. Once a DAG and a treatment–outcome target are specified, the graph shows which paths are causal, which are noncausal, and which variables may block or open those paths.
Three-node motifs. Every path is built from three local structures: chains, forks, and colliders. Chains and forks are open by default and are blocked by conditioning on the middle node. Colliders are blocked by default and are opened by conditioning on the collider or on one of its descendants.
d-Separation and conditional independence. The d-separation criterion is a graphical rule: \(X\) and \(Y\) are d-separated by a set \(\mathbf{S}\) exactly when every path between them is blocked by \(\mathbf{S}\). Under the Markov property, a d-separation statement implies the corresponding conditional independence in the distribution; a d-connection implies dependence only under the additional faithfulness assumption. Some such implications may be assessed empirically, but agreement with the data does not by itself verify the graph, and assumptions involving unobserved variables are generally not testable.
Markov factorization. The local Markov property yields the factorization \(p(v_1,\dots,v_k) = \prod_{i=1}^{k} p(v_i \mid \Pa(v_i))\), which expresses the joint distribution in terms of local conditional distributions. This factorization is the bridge from graphical structure to probability calculus and underlies later identification arguments.
Collider bias. Conditioning on a collider — or on a descendant of a collider — can induce a new association between otherwise independent variables or distort an existing one (Berkson 1946). This is the opposite of confounding adjustment: conditioning on a common cause removes spurious association, whereas conditioning on a common effect creates it.
A practical workflow. To analyze a DAG, proceed in four steps: identify the graph structure, determine which paths are open or blocked, translate d-separation statements into conditional independences under the Markov property, and then interpret those independences in light of the causal question. This workflow will be reused throughout the rest of the course.
2.9 Problems
1. Warm-up: a single collider. Consider the DAG \(X \to Y \leftarrow Z\), where \(X\) and \(Z\) have no other connections.
- Identify the structural role of \(Y\) on the path \(X \to Y \leftarrow Z\).
- Are \(X\) and \(Z\) d-separated marginally? Apply the d-separation criterion to the only path between \(X\) and \(Z\), and state which blocking rule applies.
- Are \(X\) and \(Z\) d-separated given \(Y\)? Explain what happens to the path when \(Y\) is conditioned on, state what the faithfulness assumption would add to the graphical conclusion, and describe in one sentence the real-world phenomenon this illustrates.
2. d-Separation practice. Consider the DAG: \(A \to B \to D\), \(A \to C \to D\), \(B \to E\), \(C \to E\).
- List all paths between \(A\) and \(E\). (Hint: there are four paths in total; two pass through \(D\).)
- For each path, identify the role (chain, fork, collider) of each intermediate node.
- Does \(\{B, C\}\) d-separate \(A\) and \(E\)?
- Does \(\{D\}\) d-separate \(B\) and \(C\)? What type of node is \(D\) on the path \(B \to D \leftarrow C\)?
3. Berkson’s bias. Suppose \(X\) and \(Y\) are independent standard normal variables, and let \(S = \mathbf{1}[X + Y > 0]\) (selected into a sample).
- Verify analytically that \(\mathrm{Cov}(X, Y \mid S{=}1) < 0\). (Hint: rotate coordinates. \(R = (X+Y)/\sqrt{2}\) and \(Q = (X-Y)/\sqrt{2}\) are independent standard normals, and selection is simply \(R > 0\).)
- Draw the DAG for \((X, Y, S)\) and identify \(S\) as a collider.
- Explain in one sentence why restricting the analysis to the subsample with \(S=1\) biases estimates of any association between \(X\) and \(Y\), and name the type of bias this illustrates.
4. Markov factorization and collider activation. Consider the DAG with edges \(A \to E\), \(A \to W\), \(F \to E\), \(E \to W\), where \(A\) = ability, \(F\) = family income, \(E\) = education, \(W\) = wages.
- Write down the Markov factorization \(p(a, f, e, w)\).
- Is \((F \indep W)_{\Gcal}\)? List all paths between \(F\) and \(W\) and determine which are open.
- Is \((F \indep W \mid E)_{\Gcal}\)? Identify the role of \(E\) on each path and explain whether conditioning on \(E\) opens or closes each one.
- A researcher regresses \(W\) on \(E\) and \(F\), omitting \(A\). Is the coefficient on \(E\) a causal effect of education on wages? Explain using the graph.
5. Soundness of d-separation for the collider. Consider the collider \(A \to M \leftarrow B\) with Markov factorization \(p(a, m, b) = p(a)\,p(b)\,p(m \mid a, b)\).
- By marginalizing over \(M\), show that \(A \indep B\) in the joint distribution (i.e., \(p(a,b) = p(a)\,p(b)\)). This verifies Theorem 2.1 for the case \(\mathbf{S} = \varnothing\), where Example 2.3 established d-separation graphically.
- Show that conditioning on \(M\) need not preserve the marginal independence: write out \(p(a, b \mid m)\) and explain why it generally does not factorize into \(p(a \mid m)\,p(b \mid m)\). Under what special parameterizations could conditional independence nevertheless occur? What does this imply, generically, about \(A\) and \(B\) among hospitalized patients in the accident–hospitalization–cancer example?
- Explain in one sentence why parts (a) and (b) together are consistent with Theorem 2.1. (Hint: Theorem 2.1 is a one-directional statement.)
6. Terminology check. Consider the DAG with edges \(U \to X\), \(U \to Z\), \(X \to W\), \(Z \to W\), \(W \to Y\).
- Identify the parents, children, ancestors, descendants, and non-descendants of node \(W\).
- List all pairs of adjacent nodes (connected by a single edge).
- Which pairs of nodes are connected by a directed path? List every such pair and all corresponding directed paths (some pairs are connected by more than one).
- Write the Markov factorization \(p(u, x, z, w, y)\) implied by this DAG.
7. Markov factorization and local Markov property. Consider the DAG with edges \(X_1 \to X_2\), \(X_1 \to X_3\), \(X_2 \to X_4\), \(X_3 \to X_4\).
- Write the joint density \(p(x_1, x_2, x_3, x_4)\) implied by the Markov factorization.
- State the local Markov property for each of the four nodes. For each node, identify the conditioning set \(\Pa(X_i)\) and the independence set \(\Nd(X_i) \setminus \Pa(X_i)\).
- Is \((X_2 \indep X_3)_{\Gcal}\)? Is \((X_2 \indep X_3 \mid X_1)_{\Gcal}\)? Justify each answer by listing all paths between \(X_2\) and \(X_3\) and checking whether each is blocked.
8. Toy proof: conditional independence in the fork. Consider the fork \(X_1 \leftarrow X_2 \to X_3\) with Markov factorization \(p(x_1, x_2, x_3) = p(x_2)\,p(x_1 \mid x_2)\,p(x_3 \mid x_2)\).
- Show directly, by conditioning on \(X_2 = x_2\), that \(X_1 \indep X_3 \mid X_2\). Follow the same three steps used in Example 2.2 for the chain.
- Does the DAG imply \(X_1 \indep X_3\) marginally? Justify your answer graphically (using d-separation) and algebraically: show that \(p(x_1, x_3) = \int p(x_2)\,p(x_1 \mid x_2)\,p(x_3 \mid x_2)\,\mathrm{d}x_2\) generally does not factorize into \(p(x_1)\,p(x_3)\), although special parameterizations can make it do so.
- Explain in one sentence what the fork represents substantively and why conditioning on the common cause \(X_2\) removes the association between \(X_1\) and \(X_3\).
9. d-Separation in the practice DAG. Refer to the eight-node DAG in Example 2.5 (nodes \(W\), \(X_1\), \(T\), \(M_1\), \(M_2\), \(Y\), \(C\), \(D\); edges \(W \to X_1\), \(W \to T\), \(W \to Y\), \(T \to M_1\), \(M_1 \to Y\), \(M_1 \to M_2\), \(X_1 \to C\), \(M_2 \to C\), \(C \to D\)). For each query below, state whether it is true or false and justify your answer by listing all relevant paths and determining whether each is blocked or open. This problem is intended as retrieval practice: do not refer to the worked solutions until you have written your own answer.
- \(X_1 \indep Y\)
- \(X_1 \indep Y \mid W\)
- \(T \indep M_2 \mid M_1\)
- Are \(X_1\) and \(M_2\) d-separated given \(W\)? Given \(\{W, C\}\)? Identify the specific path that conditioning on \(C\) opens, and explain why the association it transmits is non-causal.
- Does conditioning on \(D\) (a descendant of \(C\)) open the path \(X_1 \to C \leftarrow M_2\)? State which clause of the d-separation definition applies.