What is TCAV?

TCAV (Testing with Concept Activation Vectors) is a post-hoc interpretability method that learns a direction for a human-defined concept in a model's activation space and measures how often a target output is locally sensitive in that direction.

Quick Facts

Full NameTesting with Concept Activation Vectors
SpecificationOfficial Specification

How It Works

Construct a concept experiment

Choose a model checkpoint, layer, target output, concept definition, representative positive examples, and multiple independently sampled random counterexample sets. Run all examples through identical preprocessing and extract the same activation representation. The original TCAV paper uses linear classifiers to separate concept activations from random activations; held-out separator accuracy is a quality check, not the final TCAV result.

Compute directional sensitivity and repeat

Use the fitted separator normal as the CAV, then compute the gradient of the target logit or score with respect to the selected activation for each target example. Its dot product with the CAV is the local directional derivative. Count the fraction whose sign matches the declared direction, repeat the experiment with several random sets and seeds, and compare against random concepts with confidence intervals or a documented significance test.

Audit concept validity and conclusion scope

Results can change with the layer, concept images, random distribution, correlated concepts, class subset, gradient saturation, and CAV orientation. Check concept-set diversity, train-test separation, random controls, multiple-testing policy, effect stability, and alternate layers. TCAV shows local sensitivity to a learned direction; it does not prove that the concept is uniquely represented, globally important, or causally used by the model.

Key Characteristics

  • Represents a user-defined concept as a direction in one layer's activation space
  • Fits the direction from concept examples and explicit random counterexamples
  • Measures target-score directional derivatives rather than raw activation magnitude
  • Summarizes the fraction of target examples with a declared derivative sign
  • Requires repeated random sets, seed variation, and statistical controls
  • Establishes local concept sensitivity, not contribution percentage or causality

Common Use Cases

  1. Testing whether an image classifier is sensitive to texture, color, or shape concepts
  2. Comparing concept sensitivity across layers, classes, and model checkpoints
  3. Auditing whether a medical model responds to clinically meaningful morphology
  4. Screening suspected shortcuts before targeted ablation or counterfactual testing
  5. Tracking whether fine-tuning changes sensitivity to a predefined concept

Example

loading...
Loading code...

Frequently Asked Questions

What does a TCAV score mean?

For the declared target set, layer, score function, and CAV orientation, it is the fraction of examples whose target score has a positive directional derivative along the learned concept direction. A score of 0.8 means 80% met that sign condition; it does not mean the concept caused or contributed 80% of each prediction.

How are Concept Activation Vectors learned?

Activations are collected for concept examples and for a random counterexample set at the same layer. A linear separator is fitted, and its normal supplies the concept direction, subject to a documented sign convention. The original protocol repeats this process with multiple random sets because one convenient baseline can create an unstable result.

Is high CAV classifier accuracy enough to validate TCAV?

No. Accuracy only shows that the selected activations separate the supplied concept and random examples. It does not show that the concept set is representative, that the model's target is sensitive to the direction, or that results survive alternate random sets, layers, seeds, correlated concepts, and held-out target examples.

How is TCAV different from Linear Probing?

Both fit linear boundaries on frozen activations, but their estimands differ. A Linear Probe reports whether labels are decodable from representations. TCAV uses the learned normal as a concept direction and evaluates target-score directional derivatives. Neither method alone proves that the model causally relies on the decoded concept.

Can TCAV establish that a concept causes a prediction?

No. TCAV is a local sensitivity test in a learned activation direction. Correlated concepts, off-manifold directions, gradient saturation, and downstream redundancy can all complicate interpretation. A causal claim needs controlled concept or activation interventions, behavior measurements, negative controls, and tests of alternative pathways.

Related Terms