What is Concept Bottleneck Model (CBM)?
Concept Bottleneck Model (CBM) is a predictive architecture that maps inputs to a declared set of human-interpretable concept variables and computes the target prediction from those concepts, creating an inspectable and potentially editable intermediate interface.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define the concept contract
Choose concepts that domain users can label consistently and that are available at the intended decision point. Specify whether each variable is binary, categorical, ordinal, continuous, or uncertain, together with annotation guidance and missing-value policy. The original Concept Bottleneck Models paper formalizes an input-to-concept model followed by a concept-to-label model and evaluates interventions on predicted concepts.
Choose independent, sequential, or joint training
Independent training fits both stages with ground-truth concepts; sequential training first fits concept prediction and then trains the task head on predicted concepts; joint training optimizes concept and task losses together. These choices trade task accuracy, concept fidelity, robustness to concept prediction errors, and ease of intervention. Any residual input path or unconstrained latent channel must be disclosed because it weakens the bottleneck claim.
Evaluate concepts and interventions separately
Report concept discrimination, calibration, subgroup error, missing-label behavior, task performance, and robustness under concept distribution shift. Measure intervention gain using realistic correction policies and compare against random or equally costly corrections. Test whether the declared concept set is complete enough for the target and whether intervention effects remain sensible when concepts are correlated or logically constrained.
Key Characteristics
- Places named concept variables between the raw input and target prediction
- Supports inspection and controlled correction of intermediate concept values
- Can be trained independently, sequentially, jointly, or with hybrid objectives
- Requires explicit concept schemas, labels, uncertainty, and missing-value handling
- Must be audited for incomplete concepts, leakage, shortcuts, and bypass paths
- Separates concept quality, task quality, and intervention utility as distinct metrics
Common Use Cases
- Allowing clinicians to review predicted findings before a diagnostic output
- Debugging which semantic attributes drive a visual classification
- Correcting sensor or component-state concepts before a maintenance decision
- Comparing concept annotation schemes for human-model collaboration
- Testing whether a target can be predicted through an auditable semantic interface
Example
Loading code...Frequently Asked Questions
What makes a model a true Concept Bottleneck Model?
Its target prediction must be computed from a declared concept representation rather than merely displaying concept scores beside an unrestricted predictor. If raw features or latent residuals bypass the concept interface, the architecture is a hybrid and should be described as such because corrections may not control the full decision path.
How do independent, sequential, and joint CBM training differ?
Independent training fits the target head on ground-truth concepts, sequential training fits it on concepts predicted by a frozen first stage, and joint training optimizes concept and target losses together. They expose different train-test mismatches and task-concept tradeoffs, so results should not be compared without the training protocol.
What is concept leakage in a CBM?
Leakage occurs when the learned concept representation carries target-relevant information beyond the declared concept meaning, for example through continuous scores that encode within-class detail or through annotation artifacts. Strong task accuracy can then overstate semantic faithfulness. Bottleneck capacity controls, adversarial tests, and intervention behavior help diagnose it.
How should concept interventions be evaluated?
Define who supplies corrections, which concepts they can observe, how uncertainty and dependencies are handled, and what correction budget is realistic. Measure target improvement, harmful changes, subgroup effects, and calibration against random or policy-matched interventions. Oracle correction of every concept is an upper bound, not normal operating performance.
Does a CBM guarantee an accurate or complete explanation?
No. The concept set may omit decisive information, labels may be noisy, predicted concepts may be miscalibrated, and the task head may exploit correlations among concepts. A CBM provides a structured interface whose fidelity and usefulness can be tested; it does not automatically make the full decision process correct, causal, or understandable.