What is Causal Discovery?

Causal Discovery is the task of learning identifiable features of causal structure from data under explicit assumptions about graph form, sampling, independence, latent variables, interventions, and the data-generating mechanisms.

Quick Facts

SpecificationOfficial Specification

How It Works

Every algorithm family trades on a different assumption set

PC commonly assumes acyclicity, causal Markov, faithfulness, correct conditional-independence tests, and causal sufficiency; FCI relaxes causal sufficiency and can return a Partial Ancestral Graph. Score-based methods require a suitable score and search assumptions, while functional models gain orientation from restrictions such as linear non-Gaussian or additive noise. Glymour, Zhang, and Spirtes review constraint, score, and functional-model approaches and their scope.

The honest output is often an equivalence class

A CPDAG represents DAGs with the same skeleton and unshielded colliders under standard assumptions; a PAG represents ancestral features when latent confounding or selection may exist. Circles and undirected endpoints encode unresolved orientation, not implementation defects. Returning one fully directed graph without uncertainty can invent causal direction that the data do not identify.

Graph recovery needs stability and external validation

Conditional-independence errors, weak signals, mixed data types, measurement error, selection, cycles, distribution shift, and hidden causes can change the graph. Use background constraints, temporal order, bootstrap or subsample stability, multiple defensible algorithms, synthetic benchmarks with known truth, and interventions where feasible. A discovered edge is a hypothesis for review and testing, not automatic authorization for policy or adjustment.

Key Characteristics

  • Learns structural features rather than estimating one prespecified effect
  • Includes constraint-based, score-based, functional, hybrid, and continuous methods
  • Depends on causal Markov, faithfulness, graph, and sampling assumptions
  • May return a CPDAG or PAG instead of a uniquely oriented DAG
  • Must distinguish latent confounding from observed adjacency and direction
  • Requires stability, domain, temporal, and interventional validation

Common Use Cases

  1. Generating candidate causal graphs for scientific review
  2. Finding conditional-independence contradictions in a proposed DAG
  3. Prioritizing interventions that can distinguish competing graph orientations
  4. Comparing graph stability across environments or time windows
  5. Benchmarking structure-learning algorithms on simulated systems with known truth

Example

loading...
Loading code...

Frequently Asked Questions

Can Causal Discovery recover the true DAG from observational data?

Usually not uniquely without strong assumptions. Markov-equivalent DAGs can imply the same observational conditional independences. Functional restrictions, time order, background knowledge, multiple environments, or interventions can orient more edges, but every added orientation inherits those assumptions.

What is the difference between PC, FCI, and GES?

PC is constraint-based and typically assumes no latent confounding among measured variables. FCI is constraint-based but allows latent confounding and can output a PAG. GES is score-based and searches equivalence classes using a penalized fit score. Their assumptions and output semantics differ.

Why does Causal Discovery return a CPDAG or PAG?

The data may determine only features shared by many causal graphs. A CPDAG encodes a Markov equivalence class of DAGs; a PAG also represents uncertainty compatible with latent confounding or selection. Ambiguous endpoints should remain ambiguous until additional evidence resolves them.

Does NOTEARS remove the assumptions of Causal Discovery?

No. NOTEARS turns acyclic structure learning into continuous optimization, changing the search procedure rather than eliminating causal assumptions. Its result still depends on the selected model class, loss, regularization, acyclicity, sampling, latent-variable, and identifiability conditions.

How should a discovered causal graph be validated?

Check data definitions and timing, test implied independences, assess bootstrap and environment stability, compare algorithms with compatible assumptions, enforce defensible background constraints, evaluate on known simulations, and use targeted interventions where possible. Predictive fit alone is insufficient.

Related Terms