What is Out-of-Distribution Detection?

Out-of-Distribution Detection is the sample-level task of deciding whether an input is sufficiently incompatible with a declared in-distribution reference that the normal model decision should be rejected, deferred, or handled differently.

Quick Facts

SpecificationOfficial Specification

How It Works

Choose a score whose direction and evidence are explicit

Baselines include Maximum Softmax Probability, energy scores, distances in learned representations, ensembles, and task-specific consistency checks. Hendrycks and Gimpel established a widely used softmax baseline, not a universal detector. Neural classifiers can remain confident on unfamiliar inputs, and a score that separates one far-OOD benchmark may fail on semantically close or adversarial cases. Freeze preprocessing, model revision, layer, score formula, and whether larger values mean more OOD.

Calibrate the threshold on untouched ID data

Use held-out ID data representative of the approved operating domain to select a threshold for a declared ID false-positive budget. A finite-sample corrected empirical quantile can support that contract under exchangeability, but it says nothing about OOD recall. Do not tune the score and report performance on the same examples. Thresholds are model-, representation-, population-, and preprocessing-specific and must be recalibrated after those components change.

Evaluate realistic unknowns and the fallback system

Report AUROC or AUPR for ranking, FPR at a stated true-positive rate, ID false-positive rate at the production threshold, latency, and downstream cost. Include near-OOD, far-OOD, corruption, new classes, subgroup slices, and unknown sets that were not used for selection. OpenOOD shows why consistent datasets and protocols matter across methods. Finally test the reject path: fallback capacity, timeout, manual review, safe default, and recovery.

Key Characteristics

  • Makes a sample-level decision relative to an explicit ID reference
  • Requires a score direction, threshold, and fallback action
  • Can control an ID false-positive budget without guaranteeing OOD recall
  • Must test near-OOD and unseen unknowns, not only easy far-OOD examples
  • Is sensitive to model, representation, preprocessing, and population changes
  • Does not by itself prove model error or aggregate Distribution Shift

Common Use Cases

  1. Deferring unfamiliar medical, industrial, or document inputs for review
  2. Rejecting unsupported image classes in an open-world classifier
  3. Routing unusual requests to a broader model or safer workflow
  4. Monitoring whether new sensors or data sources leave an approved domain
  5. Preventing automated decisions when input support is insufficient

Example

loading...
Loading code...

Frequently Asked Questions

What counts as Out-of-Distribution?

Only what falls outside a declared ID contract: classes, population, sensors, preprocessing, time period, and operating conditions. OOD is not an intrinsic label attached to an object. The same sample can be ID for one system and OOD for another.

Is OOD Detection the same as Distribution Shift detection?

No. OOD Detection decides how to handle an individual input relative to an ID reference. Distribution Shift detection compares populations or windows. A changed class mix can shift a batch without making its samples OOD, and one OOD alert does not establish population shift.

Does a low softmax confidence reliably detect OOD inputs?

No. Maximum Softmax Probability is a useful baseline, but neural networks can assign high confidence to unfamiliar or adversarial inputs. Compare multiple justified scores on realistic near- and far-OOD sets, then validate the chosen operating threshold and fallback.

How should an OOD threshold be selected?

Select it on untouched, representative ID calibration data using a declared false-positive budget and finite-sample rule. Evaluate OOD recall separately on multiple unseen unknown sets. Recalibrate after model, representation, preprocessing, or operating-domain changes.

What should happen after an OOD input is detected?

The contract should specify abstention, human review, a broader model, additional sensing, a safe default, or explicit rejection. Test queue capacity, latency, authorization, timeout, and recovery. A detector without a dependable fallback only relocates the failure.

Related Terms

Related Articles