What is Concept Whitening?

Concept Whitening is an intrinsic interpretability technique that replaces a normalization layer with whitening and a learned orthogonal rotation so selected latent axes align with predefined human-interpretable concepts during model training.

Quick Facts

SpecificationOfficial Specification

How It Works

Whiten the latent representation

At a selected layer, estimate activation means and covariance over the training batch or maintained statistics, center the features, and apply an inverse square-root covariance transform. The output has approximately zero mean and identity covariance under the estimator. The Concept Whitening paper uses this normalization freedom as the basis for orienting latent axes toward user-specified concepts.

Rotate axes with concept supervision

Because any orthogonal rotation preserves identity covariance, optimize a rotation matrix so designated coordinates respond strongly to examples of their assigned concepts. Training can alternate task updates, whitening-statistic updates, and rotation updates on concept batches. The concept data, axis assignment, update schedule, and whether concepts overlap are part of the model specification, not post-hoc display choices.

Test axis purity and downstream use

Evaluate concept discrimination and calibration on held-out entities, cross-activation on non-target concepts, stability across seeds, task performance, and robustness under concept prevalence shift. Compare with ordinary normalization and post-hoc probes under equal budgets. Intervene on aligned coordinates to test downstream effects, while checking that other axes or later layers do not reconstruct the same concept or carry label leakage.

Key Characteristics

  • Changes model training instead of analyzing only a frozen checkpoint
  • Centers and whitens selected latent features using covariance statistics
  • Learns an orthogonal rotation that assigns concepts to chosen axes
  • Requires labeled concept examples and an explicit axis assignment
  • Preserves decorrelation under rotation but does not guarantee independence
  • Needs held-out purity, leakage, stability, and intervention evaluation

Common Use Cases

  1. Training visual classifiers with inspectable concept-aligned latent coordinates
  2. Monitoring whether selected semantic attributes emerge during optimization
  3. Comparing intrinsic concept alignment with post-hoc concept probes
  4. Auditing cross-talk between correlated concepts assigned to different axes
  5. Creating candidate concept controls for downstream intervention experiments

Example

loading...
Loading code...

Frequently Asked Questions

How is Concept Whitening different from ordinary whitening?

Ordinary whitening centers features and transforms their covariance toward the identity, leaving the orientation arbitrary. Concept Whitening additionally learns an orthogonal rotation using labeled concept examples so selected coordinates align with declared semantics. The concept objective and supervision are therefore essential parts of the method.

How is Concept Whitening different from TCAV?

TCAV analyzes a frozen model by fitting a post-hoc concept direction and testing target sensitivity along it. Concept Whitening changes training so selected representation axes become concept aligned. Both depend on concept examples, but one diagnoses an existing space while the other deliberately constructs an aligned space.

Does whitening make concepts statistically independent?

No. Identity covariance removes estimated linear second-order correlations under the whitening sample; it does not eliminate nonlinear dependence, confounding, label leakage, or semantic overlap. Concept axes can still co-activate, and later layers can recombine them. Independence requires stronger assumptions and dedicated tests.

What happens when concepts are correlated or incomplete?

Correlated concepts can compete for orthogonal axes or leak into one another, while an incomplete vocabulary leaves task-relevant variation in unassigned coordinates. Report the concept co-occurrence structure, cross-axis activation, task-concept tradeoff, and results under alternate concept sets rather than interpreting every axis as isolated ground truth.

How should a Concept Whitening model be validated?

Use held-out entity-level splits to test concept discrimination, calibration, cross-talk, seed stability, task quality, and distribution-shift robustness. Compare against the same architecture with ordinary normalization and a post-hoc probe. Coordinate interventions should then test downstream effects and reveal bypass or reconstruction paths.

Related Terms