What is TCAV?
TCAV (Testing with Concept Activation Vectors) is a post-hoc interpretability method that learns a direction for a human-defined concept in a model's activation space and measures how often a target output is locally sensitive in that direction.
Quick Facts
| Full Name | Testing with Concept Activation Vectors |
|---|---|
| Specification | Official Specification |
How It Works
Construct a concept experiment
Choose a model checkpoint, layer, target output, concept definition, representative positive examples, and multiple independently sampled random counterexample sets. Run all examples through identical preprocessing and extract the same activation representation. The original TCAV paper uses linear classifiers to separate concept activations from random activations; held-out separator accuracy is a quality check, not the final TCAV result.
Compute directional sensitivity and repeat
Use the fitted separator normal as the CAV, then compute the gradient of the target logit or score with respect to the selected activation for each target example. Its dot product with the CAV is the local directional derivative. Count the fraction whose sign matches the declared direction, repeat the experiment with several random sets and seeds, and compare against random concepts with confidence intervals or a documented significance test.
Audit concept validity and conclusion scope
Results can change with the layer, concept images, random distribution, correlated concepts, class subset, gradient saturation, and CAV orientation. Check concept-set diversity, train-test separation, random controls, multiple-testing policy, effect stability, and alternate layers. TCAV shows local sensitivity to a learned direction; it does not prove that the concept is uniquely represented, globally important, or causally used by the model.
Key Characteristics
- Represents a user-defined concept as a direction in one layer's activation space
- Fits the direction from concept examples and explicit random counterexamples
- Measures target-score directional derivatives rather than raw activation magnitude
- Summarizes the fraction of target examples with a declared derivative sign
- Requires repeated random sets, seed variation, and statistical controls
- Establishes local concept sensitivity, not contribution percentage or causality
Common Use Cases
- Testing whether an image classifier is sensitive to texture, color, or shape concepts
- Comparing concept sensitivity across layers, classes, and model checkpoints
- Auditing whether a medical model responds to clinically meaningful morphology
- Screening suspected shortcuts before targeted ablation or counterfactual testing
- Tracking whether fine-tuning changes sensitivity to a predefined concept
Example
Loading code...Frequently Asked Questions
What does a TCAV score mean?
For the declared target set, layer, score function, and CAV orientation, it is the fraction of examples whose target score has a positive directional derivative along the learned concept direction. A score of 0.8 means 80% met that sign condition; it does not mean the concept caused or contributed 80% of each prediction.
How are Concept Activation Vectors learned?
Activations are collected for concept examples and for a random counterexample set at the same layer. A linear separator is fitted, and its normal supplies the concept direction, subject to a documented sign convention. The original protocol repeats this process with multiple random sets because one convenient baseline can create an unstable result.
Is high CAV classifier accuracy enough to validate TCAV?
No. Accuracy only shows that the selected activations separate the supplied concept and random examples. It does not show that the concept set is representative, that the model's target is sensitive to the direction, or that results survive alternate random sets, layers, seeds, correlated concepts, and held-out target examples.
How is TCAV different from Linear Probing?
Both fit linear boundaries on frozen activations, but their estimands differ. A Linear Probe reports whether labels are decodable from representations. TCAV uses the learned normal as a concept direction and evaluates target-score directional derivatives. Neither method alone proves that the model causally relies on the decoded concept.
Can TCAV establish that a concept causes a prediction?
No. TCAV is a local sensitivity test in a learned activation direction. Correlated concepts, off-manifold directions, gradient saturation, and downstream redundancy can all complicate interpretation. A causal claim needs controlled concept or activation interventions, behavior measurements, negative controls, and tests of alternative pathways.