What is HSIC?
HSIC (Hilbert-Schmidt Independence Criterion) is a kernel dependence measure defined as the squared Hilbert-Schmidt norm of the cross-covariance operator between RKHS feature maps of two paired random variables.
Quick Facts
| Full Name | Hilbert-Schmidt Independence Criterion |
|---|---|
| Created | Introduced as an RKHS dependence criterion by Gretton and collaborators in 2005 |
| Specification | Official Specification |
How It Works
Measure RKHS cross-covariance
HSIC is the squared Hilbert-Schmidt norm of the cross-covariance operator between feature maps of X and Y. For n paired observations, a common biased estimator is proportional to tr(KHLH), where K and L are Gram matrices and H = I - 11^T/n centers both.
The kernel independence test provides significance procedures around this statistic. With suitable characteristic kernels, population HSIC is zero exactly under independence; arbitrary kernels can miss forms of dependence.
Separate estimators from test calibration
Biased V-statistic and unbiased U-statistic formulas have different diagonals, denominators, finite-sample ranges, and null behavior. The unbiased estimate can be negative; changing formulas between experiments invalidates direct score comparison.
Permutation tests shuffle one variable relative to the other under exchangeability. Paired time series, repeated entities, spatial samples, family groups, or matched experiments need block, cluster, or design-preserving permutations. Record the unit being shuffled, kernel parameters, estimator, number of permutations, and random seed.
Interpret dependence without claiming causality
Kernel bandwidth controls sensitivity: narrow kernels emphasize local matches, while wide kernels can miss fine structure. High-dimensional distances, outliers, unequal preprocessing, and small samples can reduce power or inflate variance. Quadratic Gram matrices also limit scale.
A positive HSIC result shows statistical dependence under the test design, not direction, mechanism, or causality. A hidden common cause can create dependence, and selection or conditioning can distort it. Use causal assumptions and interventions separately; use random features or incomplete statistics only after validating test size and power.
Key Characteristics
- Measures the RKHS cross-covariance norm between paired variables
- Can detect nonlinear and multivariate dependence with suitable kernels
- Uses centered Gram matrices rather than explicit feature coordinates
- Has biased and unbiased estimators with different finite-sample behavior
- Requires exchangeability-aware calibration for independence testing
- Does not identify causal direction or remove confounding
Common Use Cases
- Testing nonlinear dependence between two multivariate variables
- Screening candidate features for predictive dependence
- Checking residual dependence in model diagnostics
- Comparing representation dependence across layers or training runs
- Supporting causal discovery procedures under additional assumptions
Example
Loading code...Frequently Asked Questions
What does HSIC measure?
HSIC measures the squared norm of the cross-covariance operator between RKHS feature maps of two variables. With suitable kernels it captures nonlinear, multivariate dependence, but the numerical value remains kernel- and scale-dependent.
How is HSIC different from Pearson correlation?
Pearson correlation measures linear association after standardization. HSIC compares centered similarity structures and can detect broader nonlinear dependence. It needs kernel and bandwidth choices, costs more to compute, and still requires calibrated inference.
Does HSIC equal zero only when variables are independent?
At the population level, that equivalence requires kernels with the relevant characteristic or universality properties and regularity assumptions. A finite empirical estimate can be near zero because of low power, poor bandwidths, or limited sampling.
Can HSIC prove that one variable causes another?
No. Dependence is compatible with either direction, common causes, selection effects, or feedback. HSIC can be a component inside a causal-discovery method, but causal conclusions require additional structural assumptions and validation.
How can HSIC be scaled beyond quadratic Gram matrices?
Incomplete U-statistics, block estimators, Nyström approximations, and random features can reduce cost. Validate type-I error and power after approximation; preserving an average score does not guarantee preservation of the test decision.