What is Centered Kernel Alignment?

Centered Kernel Alignment (CKA) is a normalized similarity measure between two centered Gram matrices, commonly used to compare how paired examples are represented by different neural network layers or models.

Quick Facts

CreatedApplied to neural representation comparison by Kornblith and collaborators in 2019
SpecificationOfficial Specification

How It Works

Compare paired sample geometry

For centered activations X and Y, linear CKA can be written as ||Y^T X||_F^2 / (||X^T X||_F ||Y^T Y||_F). The equivalent Gram form normalizes HSIC(K,L) by the geometric mean of HSIC(K,K) and HSIC(L,L).

The original representation-similarity study uses this structure to compare layers with different widths. Rows must denote the same examples in the same order; a high score after an accidental permutation, duplicated records, or mismatched preprocessing has no valid paired interpretation.

State invariances precisely

Linear CKA is invariant to orthogonal transformations and isotropic scaling of a representation. It is not invariant to arbitrary invertible linear transformations, and that narrower invariance is intentional: complete invariance becomes uninformative when feature dimension exceeds sample count.

Kernel CKA can capture nonlinear similarity but adds kernel and bandwidth dependence. Centering convention, activation normalization, token or spatial aggregation, and whether examples or features occupy rows must be fixed before comparing scores.

Treat the score as one diagnostic

CKA can reveal block structure across layers, correspondence across random initializations, or changes during fine-tuning. It cannot by itself identify causal features, guarantee equal predictions, or determine that one representation is better.

Research on CKA reliability demonstrates sensitivity to outliers and transformations that can preserve linear separability. Report uncertainty across sample subsets, inspect influential examples, pair CKA with task probes and behavioral tests, and use an estimator designed for minibatches rather than averaging arbitrary batch scores.

Key Characteristics

  • Compares centered sample-similarity matrices from paired representations
  • Normalizes HSIC to remove isotropic scale differences
  • Supports linear and kernel variants
  • Allows different feature widths while requiring the same ordered examples
  • Is invariant to orthogonal transforms but not arbitrary invertible transforms
  • Can be sensitive to sample composition, outliers, kernels, and minibatch estimation

Common Use Cases

  1. Comparing corresponding layers across independently trained networks
  2. Tracking representation changes during fine-tuning or continual learning
  3. Diagnosing redundancy or stage structure inside a deep model
  4. Comparing teacher and student representations during distillation
  5. Selecting layers for probes while retaining task-level validation

Example

loading...
Loading code...

Frequently Asked Questions

What does a high CKA score mean?

It means the two representations induce similar centered pairwise geometry on the evaluated examples under the chosen kernel and estimator. It does not prove identical neurons, predictions, concepts, causal mechanisms, or generalization.

Can CKA compare layers with different widths?

Yes. CKA compares sample-by-sample Gram matrices, so feature counts may differ. The number, identity, order, preprocessing, and weighting of examples must still match, and layer aggregation must be defined consistently.

What transformations leave linear CKA unchanged?

Linear CKA is invariant to orthogonal changes of basis and isotropic scaling. It is not generally invariant to arbitrary invertible linear transforms, nonlinear transforms, anisotropic rescaling, or changes in the evaluated sample distribution.

How is CKA related to HSIC?

CKA normalizes HSIC between two centered Gram matrices by their self-HSIC values. HSIC was designed as a dependence measure and test statistic; CKA uses the normalized quantity as a representation-similarity index.

Can CKA scores be averaged across minibatches?

A simple mean of batch CKA values generally does not equal full-dataset CKA because centering and normalization are nonlinear. Use a validated streaming or unbiased minibatch estimator, keep sample composition stable, and report uncertainty.

Related Terms