What is Difference-in-Differences?
Difference-in-Differences is a quasi-experimental design that estimates a treatment effect by subtracting the comparison group's outcome change from the treated group's change under a parallel-trends assumption.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Parallel trends concerns an unobserved counterfactual path
DiD commonly targets an Average Treatment Effect on the Treated. Parallel trends equates the untreated outcome change the treated group would have experienced with the observed change in the comparison group, possibly conditional on pre-treatment covariates. Similar outcome levels are unnecessary, but a convenient comparison group is not enough. Treatment-correlated shocks, changing measurement, selective attrition, and spillovers can invalidate the contrast.
Pre-treatment plots diagnose but cannot prove the assumption
Event-study leads, placebo dates, alternate comparison groups, covariate balance, and outcome-definition checks can reveal contradictions. Failure to reject pre-trend differences may simply reflect low power, while conditioning on a noisy pre-test can distort inference. State the plausible deviation from parallel trends and perform sensitivity analysis rather than declaring the assumption verified.
Staggered adoption requires cohort-time estimands
With different adoption dates and heterogeneous or dynamic effects, a conventional two-way fixed-effects coefficient can mix valid and already-treated comparisons with difficult or negative weights. Callaway and Sant'Anna identify group-time effects under no anticipation and parallel-trends conditions, then aggregate them for declared questions. Report the comparison group, event-time support, aggregation weights, clustered uncertainty, and simultaneous bands.
Key Characteristics
- Compares changes rather than post-treatment outcome levels
- Commonly targets an effect for treated units or treatment cohorts
- Relies on a parallel untreated-outcome trend assumption
- Requires treatment timing and anticipation to be defined explicitly
- Needs cohort-aware estimators under staggered adoption
- Supports placebo, pre-trend, and sensitivity diagnostics without proving validity
Common Use Cases
- Evaluating a policy introduced in one region before another
- Measuring a product rollout when randomized assignment was unavailable
- Estimating a program effect from repeated panel or cross-sectional data
- Comparing adoption cohorts with never-treated or not-yet-treated units
- Testing whether an event-study conclusion survives alternate trend assumptions
Example
Loading code...Frequently Asked Questions
What is the parallel-trends assumption in Difference-in-Differences?
It says the treated group's average untreated Potential Outcome would have changed like the comparison group's outcome over the relevant period, possibly after conditioning on declared pre-treatment covariates. It concerns an unobserved counterfactual trend, not necessarily equal observed outcome levels.
Do flat pre-trends prove a DiD design is valid?
No. Pre-period estimates can detect some violations but have limited power and do not observe the post-period untreated counterfactual. Event-study leads should be paired with design knowledge, placebo outcomes or dates, alternate controls, and sensitivity analysis for plausible trend deviations.
Why can two-way fixed effects fail with staggered treatment?
When cohorts adopt at different times and effects vary, TWFE can use already-treated units as controls and combine underlying effects with difficult or negative weights. Cohort-time estimators instead define valid comparisons first and aggregate explicit group-time effects.
Does Difference-in-Differences require panel data?
No. It can use unit-level panels or repeated cross-sections when the corresponding sampling and composition assumptions hold. Panels support within-unit changes; repeated cross-sections compare changing samples and therefore require stable population composition or appropriate reweighting.
How should uncertainty be estimated in DiD?
Account for the level at which treatment is assigned and errors are dependent, often through cluster-robust or design-specific inference. With few treated clusters, conventional asymptotics can be poor. Event studies also need simultaneous rather than isolated pointwise interpretation.