What is Deep Ensembles?

Deep Ensembles are collections of neural networks trained as distinct members for the same predictive task and combined at inference to form an aggregate predictive distribution.

Quick Facts

SpecificationOfficial Specification

How It Works

Create useful diversity without sacrificing member quality

The Deep Ensembles paper trains proper-scoring-rule models from different initializations and combines their predictions. Random seeds and data order may be sufficient to reach distinct solutions, while bootstraps, architectures, priors, or hyperparameters can broaden diversity when justified. Diversity alone is not useful: every member still needs leakage-free validation, and correlated failures must be tested with domain-relevant slices and perturbations.

Aggregate predictive distributions, not arbitrary scores

For classification, average normalized class probabilities from members and then evaluate the mixture; averaging logits defines a different predictor. For regression members that output a mean and conditional variance, the law of total variance gives total predictive variance as mean within-member variance plus variance of member means. The first term can represent modeled Aleatoric Uncertainty, while the second is only the epistemic proxy represented by this ensemble.

Evaluate value under shift and full serving cost

Measure task quality, Log Loss or Brier Score, calibration, selective risk, and uncertainty behavior on in-distribution and realistic shifted data. Ovadia et al. benchmark uncertainty methods under dataset shift, but no ranking transfers automatically to another domain. Count training, storage, latency, energy, and correlated availability risk; distillation or a smaller ensemble is acceptable only after preserving the decision metrics that justified the ensemble.

Key Characteristics

  • Combines multiple separately trained neural-network predictors
  • Uses probability mixtures or distribution mixtures rather than arbitrary score averaging
  • Can improve predictive accuracy and expose variation across fitted solutions
  • Represents only uncertainty induced by the chosen member-generation process
  • Can remain confidently wrong when all members share data or architectural blind spots
  • Multiplies training, storage, inference, monitoring, and release-management costs

Common Use Cases

  1. Improving predictive distributions for high-value classification or regression
  2. Estimating member disagreement for abstention and review policies
  3. Separating within-member and between-member variance in probabilistic regression
  4. Benchmarking uncertainty behavior under realistic dataset shifts
  5. Creating a stronger teacher distribution for validated model distillation

Example

loading...
Loading code...

Frequently Asked Questions

How do Deep Ensembles work?

Train several competent models for the same task with a declared source of variation, freeze their versions, and combine their predictive distributions at inference. Evaluate both the aggregate prediction and member disagreement on untouched in-domain and shifted data.

Are Deep Ensembles Bayesian?

Not in general. Randomly initialized, independently optimized members do not constitute samples from a known Bayesian posterior unless a specific construction proves that interpretation. Deep Ensembles are useful predictive mixtures whose uncertainty meaning comes from their generation and validation protocol.

How many ensemble members are enough?

There is no universal count. Plot task metrics, proper scores, calibration, disagreement stability, latency, and cost as members are added on a fixed validation protocol. Stop when incremental decision value plateaus or the resource budget is exceeded.

Should a classifier ensemble average logits or probabilities?

A standard predictive mixture averages normalized class probabilities. Averaging logits and then applying Softmax creates a different aggregation with different confidence behavior. Declare the rule and validate its proper scores and calibration rather than treating them as interchangeable.

Can Deep Ensembles detect every out-of-distribution input?

No. Members can agree on unsupported inputs because they share training data, representations, architecture, or objective. Evaluate Near-OOD and Far-OOD cases against a declared reference and pair disagreement with deterministic validation and fallback controls.

Related Terms

Related Articles