What is Weak Supervision?

Weak Supervision is a family of learning settings in which the available training signal is incomplete, coarse, noisy, uncertain, or produced by imperfect sources instead of being a large set of fully verified instance-level labels.

Quick Facts

SpecificationOfficial Specification

How It Works

Classify the weakness before selecting a method

Zhou's taxonomy separates incomplete, inexact, and inaccurate supervision. Semi-Supervised Learning addresses one incomplete-label setting; Multiple Instance Learning uses coarse bag labels; noisy-label learning handles unreliable targets. Programmatic rules are one important source family, not the definition of all Weak Supervision. Record each source's scope, output space, abstention behavior, provenance, cost, and expected failure modes.

Treat source aggregation as an estimation problem

The Snorkel system paper models unknown Labeling Function accuracies and selected dependencies from agreements and conflicts, producing probabilistic labels for a discriminative model. The result depends on source diversity, overlap, conditional-dependence assumptions, class balance, and identifiability. Compare learned aggregation with majority vote, simple rules, and supervised baselines; duplicated or near-duplicated sources must not masquerade as independent evidence.

Measure both label quality and end-model utility

On an untouched gold set, report per-source coverage, precision, recall, abstention, conflict, class and subgroup behavior, and stability after source changes. Then evaluate aggregated-label calibration and the End Model on a separate test set. Keep source code, prompts, model revisions, knowledge snapshots, and generated labels versioned. Review high-impact conflicts manually and monitor whether source or population drift invalidates earlier accuracy estimates.

Key Characteristics

  • Covers incomplete, inexact, and inaccurate forms of supervision
  • Can use rules, knowledge bases, crowds, sensors, prompts, or existing models
  • Allows sources to vote, abstain, overlap, conflict, and vary by data slice
  • May aggregate weak sources into probabilistic labels before End Model training
  • Is vulnerable to correlated errors, source duplication, bias, and drift
  • Still requires independent gold labels for development and final evaluation

Common Use Cases

  1. Encoding expert rules for domain-specific text classification
  2. Using knowledge-base matches as noisy relation-extraction labels
  3. Learning instance predictions from document-level or bag-level annotations
  4. Combining multiple annotators, models, and heuristics under one label contract
  5. Bootstrapping a training set before targeted manual annotation

Example

loading...
Loading code...

Frequently Asked Questions

Is Weak Supervision the same as Semi-Supervised Learning?

No. Semi-Supervised Learning is an incomplete-supervision setting with labeled and unlabeled examples. Weak Supervision is an umbrella that also includes coarse labels, noisy labels, distant supervision, crowds, and programmatic sources. A workflow can be both weakly supervised and semi-supervised, but the assumptions differ.

Does Weak Supervision require Labeling Functions or Snorkel?

No. Labeling Functions and Snorkel implement Programmatic Weak Supervision. The broader field also covers bag-level labels, partial labels, noisy annotators, distant supervision, and other imperfect signals. Name the exact source and aggregation method rather than using Weak Supervision as a complete protocol.

Why is majority vote often insufficient for weak labels?

Sources can differ in accuracy, coverage, class behavior, and correlation. Five copied rules are not five independent observations, and a high-coverage weak rule can overwhelm a precise specialist. Majority vote remains a useful baseline, while learned aggregation must prove value on independent labels and disclose its assumptions.

Can a Label Model remove all weak-label noise?

No. It estimates latent labels under a model of source behavior. Shared blind spots, unobserved dependencies, missing class support, adversarial sources, and population shift can leave systematic errors. Preserve source lineage, inspect conflicts and slices, and validate both generated labels and the End Model against gold data.

How much manually labeled data is needed for Weak Supervision?

There is no universal count. Build a representative development set for source design and a strictly untouched test set for final evaluation, stratified by important classes, sources, and risks. Plot uncertainty and learning curves as labels are added, then stop when the intended decisions have adequate evidence rather than meeting a fixed record quota.

Related Terms

Related Articles