What is Label Shift?
Label Shift is a distribution-change assumption under which class prevalence changes from source to target, `P_source(Y) != P_target(Y)`, while the class-conditional input distribution remains stable, `P_source(X|Y) = P_target(X|Y)`.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Identify priors through a stable measurement channel
Black Box Shift Estimation uses a fixed classifier as a noisy measurement of labels. Estimate its confusion matrix P(prediction|Y) on held-out labeled source data, measure prediction frequencies on unlabeled target data, and solve the resulting linear system for target priors. Lipton, Wang, and Smola establish this BBSE approach and its assumptions. If the confusion matrix is singular or nearly singular, the classes are not identifiable through that classifier.
Distinguish prior correction from probability calibration
Estimated class-prior ratios can reweight source examples or adjust posterior odds when the model and assumption support that operation. They do not automatically calibrate scores, repair ranking errors, or prove P(X|Y) stability. Soft estimators use predicted probabilities rather than hard classes and need those scores to satisfy their own calibration or moment conditions. Always record the classifier revision, thresholding rule, label map, source prior, and estimation method.
Validate feasibility and downstream utility
Check that estimated priors are non-negative, sum to one, and remain stable under sampling uncertainty and reasonable slices. Inspect the confusion matrix condition number or bootstrap sensitivity; inversion amplifies noise when class signatures are similar. Compare corrected and uncorrected decisions on representative target labels, including rare classes and cost-sensitive outcomes. A mathematically feasible prior estimate can still make the product worse if the invariance assumption is false.
Key Characteristics
- Changes `P(Y)` while assuming `P(X|Y)` remains invariant
- Uses a stable classifier confusion matrix as a label measurement channel
- Can estimate target priors from unlabeled target predictions under identifiability
- Becomes unstable when classes have similar prediction signatures
- Does not cover new classes, relabeling, or changed within-class inputs
- Requires target-labeled validation before corrected decisions are trusted
Common Use Cases
- Adjusting a classifier when disease or fraud prevalence changes
- Estimating target class mix before enough target labels arrive
- Reweighting evaluation or training under a justified prior-shift model
- Diagnosing prediction-rate changes with a known source confusion matrix
- Testing whether prior correction improves cost-sensitive target decisions
Example
Loading code...Frequently Asked Questions
What is the difference between Label Shift and Covariate Shift?
Label Shift changes class priors `P(Y)` while assuming class-conditional inputs `P(X|Y)` stay stable. Covariate Shift changes `P(X)` while assuming `P(Y|X)` stays stable. The corrections and required evidence differ, so observed prediction-rate drift does not identify either one by itself.
Can Label Shift be estimated without target labels?
Under the stable `P(X|Y)` assumption, a fixed classifier's source confusion matrix and unlabeled target prediction frequencies can identify target priors. The matrix must distinguish classes and remain stable. Target labels are still needed to validate the assumption and downstream decisions.
Why can BBSE return negative class priors?
Sampling noise, an ill-conditioned confusion matrix, changed class-conditional inputs, a new class, or pipeline mismatch can make the unconstrained inverse infeasible. Do not silently clamp the result; diagnose the violation, quantify uncertainty, and use a documented constrained estimator only if justified.
Is Label Shift the same as changing label definitions?
No. Label Shift preserves the label space and the within-class input distribution while prevalence changes. Renaming, merging, splitting, or changing annotation criteria modifies the target semantics and requires dataset and model migration rather than prior-only correction.
How should a Label Shift correction be evaluated?
Use representative target labels to compare corrected and uncorrected loss, calibration, decision cost, and class-specific errors. Report prior-estimation uncertainty, matrix conditioning, classifier revision, and slices. Improvement in estimated priors alone is not a production success criterion.