什么是 混杂(Confounding)?
混杂(Confounding)是 Treatment Assignment 与 Outcome 的共同原因存在分布差异,导致观测到的 Treatment-Outcome Association 不能代表目标 Causal Effect 的系统性偏差。
快速了解
| 规范文档 | 官方规范 |
|---|
工作原理
Confounder 是因果角色,不是表字段类型
同一个变量对一组 Treatment-Outcome 可能是 Confounder,对另一组可能是 Mediator,换一个时间边界后也可能无关。角色必须由领域知识与时序决定。AHRQ 的 Causal DAG 指南解释 Open Backdoor Path 如何表示 Confounding,以及为什么调整 Collider 会打开原本关闭的路径。
Adjustment 追求可比性,而不是预测准确率
Restriction、Randomization、Stratification、Matching、Standardization、Weighting 与 Outcome Regression 都可以在各自假设下处理 Measured Confounding。应根据 Causal Structure 与 Estimand 选择 Treatment 前协变量,再检查 Overlap 与 Balance。区分 Treatment 很准的模型可能恶化 Overlap,高预测率 Outcome Model 也可能遗漏因果识别所需变量。
Residual 与 Unmeasured Confounding 必须如实限定
Cochrane ROBINS-I 指南区分 Confounding、Selection Bias 与 Information Bias,并要求预先声明 Confounding Domain。Measurement Error、粗糙 Proxy 与缺失 Common Cause 都会留下偏差。Sensitivity Analysis、Negative Control 与替代 Set 可量化脆弱性,却不能证明隐藏混杂不存在。
主要特点
- 使 Observed Association 偏离目标 Causal Effect
- 取决于 Treatment、Outcome、时间顺序与 Causal Structure
- 通常涉及同时影响 Assignment 与 Outcome 的共同原因
- 不能只通过相关性、显著性或 Feature Importance 判定
- 会因遗漏或测量误差在调整后继续残留
- 不同于 Selection Bias、Mediation、Effect Modification 与 Collider Bias
常见用途
- 解释自选择使用功能的用户为何留存更高
- 为观察性 Treatment Study 选择 Baseline Covariate
- 检测不同风险 Stratum 间的 Simpson's Paradox
- 模型拟合前审计 Treatment 后变量
- 为未测量 Common Cause 设计 Sensitivity Analysis
示例
Loading code...常见问题
同时与 Treatment 和 Outcome 相关的变量都是 Confounder 吗?
不是。统计相关只能作为筛选线索,Confounder 由因果角色与时序定义。Mediator、Collider、Instrument 与 Proxy 也可能同时相关,但调整它们可能改变 Estimand 或引入偏差。
Machine Learning 能消除 Confounding 吗?
Machine Learning 可以为已测量协变量拟合灵活 Propensity 或 Outcome Function,却不能保证所有 Common Cause 都已测量,也不能修复 Treatment 后调整、创造 Overlap,或从预测准确率确定 Causal Graph。
Residual Confounding 是什么?
Residual Confounding 是调整后仍存在的偏差,原因可能是 Confounding Domain 被遗漏、测量有误、表示过于粗糙或建模不充分。更大数据集会减少随机误差,却不会自动消除这种系统偏差。
为什么调整更多变量反而可能增加偏差?
对 Collider 条件化会打开非因果路径;调整 Mediator 则可能阻断 Total Effect 的一部分。Treatment 后变量还可能同时编码多种机制。Adjustment Set 应服从因果问题和假设图,而不是使用全部特征。
应该如何处理 Unmeasured Confounding?
应把它写成 Identification Risk,利用设计知识限制可能的 Common Cause,并尽量执行 Quantitative Sensitivity Analysis、Negative Control 或替代设计。普通 Balance Metric 或 Regression Fit 都不能证明未测量混杂不存在。