What is Differential Privacy?

Differential Privacy is a mathematical framework that bounds how much the probability of any output can change when one declared privacy unit is added to, removed from, or replaced in a dataset.

Quick Facts

SpecificationOfficial Specification

How It Works

Define adjacency and the privacy unit first

Decide whether neighboring datasets differ by one event, one row, one user, one household, or another entity. User-level protection must bound all contributions from one user; event-level protection can leave repeated contributors exposed. NIST SP 800-226 treats the privacy unit, privacy parameters, algorithm, utility, bias, trust model, security, and collection practice as parts of one evaluable guarantee.

Calibrate mechanisms to sensitivity

For a numeric query, sensitivity is the largest output change allowed between neighboring datasets. The Laplace mechanism commonly calibrates noise to sensitivity / epsilon; Gaussian mechanisms usually provide approximate (epsilon, delta) DP under their stated accountant. Means, sums, gradients, and histograms need clipping or contribution limits before sensitivity is finite. Post-processing preserves DP, but querying the source again creates another release.

Account for composition and implementation hazards

Every release against overlapping protected units spends privacy loss. Maintain a ledger across dashboards, experiments, model rounds, and retries instead of resetting epsilon per endpoint. Specify the accountant and sampling assumptions for DP-SGD. Test secure randomness, floating-point behavior, side channels, bounds, empty groups, and small slices; measure confidence intervals and disparate utility because noise can burden small populations more heavily.

Key Characteristics

  • Defines privacy through output distributions on neighboring datasets
  • Requires an explicit privacy unit and adjacency relation
  • Uses epsilon and, for approximate DP, delta to quantify privacy loss
  • Calibrates randomness to bounded sensitivity or clipped contributions
  • Supports post-processing and formal composition across releases
  • Trades privacy against task-specific utility, bias, and operational cost

Common Use Cases

  1. Publishing counts or histograms from sensitive population data
  2. Training models with clipped and noised gradients under DP-SGD
  3. Collecting telemetry with central, local, or distributed trust models
  4. Generating synthetic statistics under an explicit privacy budget
  5. Auditing repeated analytics releases against one privacy ledger

Example

loading...
Loading code...

Frequently Asked Questions

What do epsilon and delta mean in Differential Privacy?

Epsilon bounds the multiplicative change between output probabilities on neighboring datasets; smaller values generally mean a stronger bound. Delta permits a limited additive probability of outcomes not covered by that multiplicative bound. They are meaningful only with the same privacy unit, adjacency rule, mechanism, accountant, and release scope.

Is adding random noise enough to claim Differential Privacy?

No. The mechanism must calibrate randomness to a proven sensitivity or contribution bound and use the declared adjacency relation. The implementation also needs secure randomness, correct accounting, controlled side channels, and protection of raw inputs. Arbitrary perturbation may reduce accuracy without providing a valid DP guarantee.

What is a privacy budget?

A privacy budget is the allowed cumulative privacy loss for a defined population and scope. Repeated analyses generally compose, so each release consumes part of the budget. A production ledger should bind expenditures to dataset versions, privacy units, mechanisms, parameters, jobs, and outputs rather than resetting epsilon for every request.

Does Differential Privacy prevent every data breach?

No. DP limits what released outputs reveal about one protected contribution under its assumptions. It does not encrypt raw data, repair broken authorization, stop insiders from copying inputs, guarantee fairness, or make an unsafe data-collection purpose acceptable. Those risks require separate controls.

How should a Differential Privacy deployment be evaluated?

Review the privacy unit, adjacency relation, epsilon, delta, contribution bounds, mechanism, trust model, accountant, number of releases, randomness, and side channels. Then measure utility with uncertainty across relevant slices and verify the implementation against the claimed algorithm. Reporting only epsilon is not an adequate assessment.

Related Terms

Related Articles