What is Network Dissection?

Network Dissection is a post-hoc interpretability method that quantifies how well an individual hidden unit's spatial activation aligns with human-labeled visual concepts by comparing thresholded activation maps with pixel-level concept masks.

Quick Facts

SpecificationOfficial Specification

How It Works

Align unit activations with concept masks

Select a checkpoint and convolutional layer, then collect each channel's activation map on images that have pixel-level concept annotations. The Network Dissection study uses a high activation quantile to obtain a unit-specific threshold, upsamples the binary activation region, and compares it with masks from a broad visual concept vocabulary. Preprocessing and interpolation must match across all units.

Score detectors with intersection over union

For a unit and concept, aggregate the intersection and union of their binary masks over eligible images and compute Intersection over Union (IoU). Assign the best concept only when the score exceeds a documented threshold, and report unmatched as a valid outcome. Record the layer, unit, quantile, dataset version, concept frequency, resize rule, IoU aggregation, and tie policy so counts are reproducible.

Separate semantic alignment from mechanism

A concept may be distributed across many units, while one unit may be polysemantic. Rotating a representation can preserve network behavior while changing which axes appear interpretable, so detector counts are basis dependent. Dataset imbalance and missing concepts also cap what can be named. Use ablation, activation replacement, or controlled synthesis to test whether an aligned unit is necessary or sufficient for behavior.

Key Characteristics

  • Evaluates individual spatial units against pixel-level semantic annotations
  • Thresholds activation maps and compares them with concept masks using IoU
  • Depends on the layer, activation quantile, upsampling rule, and concept vocabulary
  • Allows unmatched units instead of forcing every channel to receive a label
  • Can miss distributed concepts and assign one label to polysemantic units
  • Measures representational alignment rather than causal contribution

Common Use Cases

  1. Comparing the number and type of concept-aligned units across architectures
  2. Auditing whether scene models learn objects, materials, textures, or colors
  3. Locating channels associated with a suspected foreground or background shortcut
  4. Generating unit-level hypotheses for ablation and activation intervention
  5. Testing how training objectives or supervision change interpretable unit counts

Example

loading...
Loading code...

Frequently Asked Questions

What does Network Dissection measure?

It measures spatial overlap between a thresholded hidden-unit activation map and a human-annotated concept mask, usually with IoU aggregated over a dataset. The result says that one unit axis aligns with one vocabulary concept under a declared protocol; it does not directly measure prediction importance or causal use.

How is the activation threshold chosen?

A common protocol selects a high quantile from that unit's activation distribution, then applies the resulting cutoff across the evaluation images. The quantile, sampling population, interpolation, and resolution must be fixed and reported. Choosing a threshold after viewing labels can inflate detector counts and invalidates comparisons.

Why does the concept vocabulary matter?

A unit can only receive labels that the dataset defines and annotates. Missing concepts become unmatched even if the unit is semantically coherent, while broad or imbalanced labels can dominate scores. Comparisons are credible only when vocabulary, annotation version, concept support, and image distribution are held constant.

Can Network Dissection detect distributed concepts?

Not reliably in its standard unit-by-unit form. A concept encoded across several channels may have no single high-IoU detector, and a polysemantic unit may align with different concepts in different contexts. Subspace methods, multivariate probes, and intervention experiments can complement the single-axis analysis.

Does a high-IoU unit cause the model's decision?

No. High IoU establishes representational alignment with annotated regions. The unit may be redundant, downstream computation may ignore it, or both masks may track a shared correlate. Test causal relevance with controlled ablation, replacement, or stimulation and measure class-specific behavior against negative controls.

Related Terms

Related Articles