What is Active Learning?

Active Learning is a supervised-learning process in which a learner chooses which unlabeled examples or queries should receive labels under a defined annotation budget and evaluation protocol.

Quick Facts

SpecificationOfficial Specification

How It Works

Define the loop and the true annotation budget

Begin with a leakage-free seed set, untouched evaluation set, and versioned unlabeled pool. At each round, train or update the model, score only eligible candidates, select a batch, collect labels with provenance, resolve adjudication, and record cost before retraining. The Cambridge Machine Learning Group overview covers foundational statistical and Bayesian active-learning work. Budget should include expert time, disagreements, failed queries, latency, and model retraining rather than only label count.

Balance informativeness, representation, and batch diversity

Least confidence, margin, entropy, Query-by-Committee, BALD, expected model change, and expected error reduction define different utilities. Uncertainty can select noisy outliers or duplicate boundary cases; density and core-set methods can choose representative but already easy samples. Current scikit-activeml documentation separates informativeness, representativeness, and hybrid strategies and distinguishes top-k from diverse batches. Choose with an ablation against random sampling.

Evaluate the adaptive process without selection leakage

Plot performance and decision cost against cumulative annotation cost on a fixed, representative test set, repeating full acquisition trajectories across seeds. Actively selected labels are not an IID evaluation sample and can distort prevalence and calibration. Track class and subgroup coverage, duplicate rate, oracle disagreement, skipped cases, and pool drift. Stop when marginal value falls below a predefined threshold, not when the selected set happens to score well on itself.

Key Characteristics

  • Alternates model fitting, acquisition scoring, labeling, and dataset updates
  • Operates under a declared pool, stream, or query-synthesis setting
  • Can combine uncertainty, disagreement, expected value, and representation
  • Requires explicit annotation cost, batch, eligibility, and stopping policies
  • Creates adaptive selection bias that must be separated from evaluation data
  • Must outperform random or business-as-usual acquisition on repeated learning curves

Common Use Cases

  1. Prioritizing expensive expert labels for medical or scientific data
  2. Selecting diverse edge cases for document or image classification
  3. Adapting a model to a new domain with a limited labeling budget
  4. Collecting preference judgments for ranking or recommendation systems
  5. Maintaining production classifiers as new data patterns emerge

Example

loading...
Loading code...

Frequently Asked Questions

How does Active Learning reduce labeling cost?

It directs a limited budget toward examples expected to improve the target decision more than passive sampling. Savings are not guaranteed: they must be measured as a learning curve against random or current acquisition using total annotation and retraining cost.

Is uncertainty sampling enough for Active Learning?

Not usually. High uncertainty may identify mislabeled, out-of-domain, or repetitive cases. Batch systems often combine informativeness with diversity, representation, eligibility, and cost, then test each contribution with controlled ablations.

What is the cold-start problem in Active Learning?

With too few initial labels, the model and its uncertainty scores may be unreliable, so early queries reinforce a poor boundary. Use a diverse seed set, prior knowledge where defensible, and random or representation-based baselines before trusting model-driven acquisition.

Does Active Learning create biased training data?

Yes, the labeled set is selected adaptively rather than sampled IID. This can distort class prevalence and probability calibration. Preserve query propensities and provenance when possible, and never use the acquired set itself as the sole performance estimate.

When should an Active Learning loop stop?

Stop according to a predeclared rule based on total budget, marginal improvement on an untouched evaluation set, target-risk attainment, pool exhaustion, or unacceptable oracle delay. Repeatedly checking and stopping on noisy gains should be handled explicitly.

Related Terms

Related Articles