What is Causal DAG?
Causal DAG is a Directed Acyclic Graph whose nodes represent causally defined variables and whose arrows encode direct causal assumptions, enabling graphical reasoning about paths, conditional independence, interventions, and adjustment.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Chains, forks, and colliders respond differently to conditioning
Conditioning on a non-collider blocks a chain A->M->Y or fork A<-U->Y. A collider A->C<-Y blocks its path by default, but conditioning on the collider or its descendant can open association. This is why adjusting every predictive feature can block mediation or create selection bias. The AHRQ causal-DAG guide connects these path rules to counterfactual effects and covariate selection.
The backdoor criterion converts a graph into an adjustment claim
For the total effect of treatment A on outcome Y, an adjustment set must contain no descendants of A and must block every path entering A through a back door. Multiple valid sets may exist, and the smallest or most predictive set is not automatically best under measurement error or limited overlap. Unmeasured common causes should appear as latent nodes rather than being silently omitted.
Observed data can refute implications but rarely certify the graph
Conditional-independence tests can contradict a DAG under faithfulness and measurement assumptions, but many different graphs are Markov equivalent and imply the same observed independences. Causal discovery therefore needs background constraints, interventions, or stronger functional assumptions. Version the graph, variable definitions, time indices, target estimand, and adjustment set; use sensitivity analysis for uncertain edges.
Key Characteristics
- Represents causally defined variables as nodes and assumptions as arrows
- Requires directed paths without same-time causal cycles
- Uses d-separation to reason about open and blocked paths
- Distinguishes confounders, mediators, colliders, and descendants
- Supports adjustment-set and identification analysis before estimation
- Encodes assumptions that data alone usually cannot prove
Common Use Cases
- Selecting pre-treatment covariates for effect estimation
- Finding collider or mediator adjustment mistakes in a feature pipeline
- Documenting assumptions behind an Instrumental Variables design
- Separating causal identification from predictive feature selection
- Reviewing how time order, missing causes, or selection affect an analysis
Example
Loading code...Frequently Asked Questions
Does drawing an arrow in a Causal DAG prove causation?
No. An arrow records a causal assumption supported by domain knowledge, design, timing, or prior evidence. The graph helps derive consequences of that model. Observational fit can reveal contradictions under additional assumptions, but it does not turn asserted arrows into proven mechanisms.
Why must a Causal DAG be acyclic?
A DAG cannot contain a directed path that returns to its starting node. Real feedback is represented by distinct time-indexed nodes, such as state at `t` causing state at `t+1`. Collapsing time can hide feedback and make adjustment or intervention semantics ambiguous.
How does a DAG determine which variables to adjust for?
For a total effect, choose a set that blocks every backdoor path from treatment to outcome without including treatment descendants. Do not adjust for mediators when targeting the total effect or for colliders that would open paths. The graph must include relevant latent causes for this reasoning to be credible.
Can causal discovery recover the true DAG from data?
Usually not uniquely from observational data. Markov-equivalent DAGs can encode the same conditional independences, and results depend on faithfulness, causal sufficiency, measurement, and model-class assumptions. Time order, experiments, background constraints, and sensitivity analysis are needed.
What is the difference between a Causal DAG and a Bayesian Network?
Both can use a DAG to factorize a distribution, but a Bayesian Network's arrows need not support intervention semantics. A Causal DAG asserts a structural causal interpretation: changing a parent through intervention changes descendants according to the model while other mechanisms remain stable.