What is Data Poisoning?
Data Poisoning is a training-time integrity or availability attack in which an adversary controls part of the data used to fit or update a machine learning model so that the learned behavior serves the adversary's objective.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Model the attack before measuring it
NIST defines data poisoning as a poisoning attack in which an adversary controls part of the training data. Turn that boundary into an experiment: declare whether the attacker can add, edit, relabel, delete, or order records; whether training is centralized, online, or federated; and whether the goal is availability, targeted integrity, or a trigger-based backdoor. Keep the poison budget and defender knowledge fixed when comparing attacks.
Measure the intended failure and collateral damage
Report clean utility, target success, affected slices, poison rate, persistence after retraining, and variance across seeds. For a backdoor, measure attack success on triggered inputs and false activation on clean inputs. For broad degradation, report task metrics and calibration by source or subgroup. Compare against random corruption and natural data errors so an optimized attack is not mistaken for ordinary data quality failure.
Defend the training supply chain in layers
Record source, consent, license, collection time, transformation lineage, labeler, checksum, and split membership for every dataset version. Quarantine new sources, deduplicate before splitting, validate labels with independent evidence, compare source-stratified holdouts, cap each contributor's influence, and retain rollback-ready checkpoints. Outlier detection, robust training, and Robust Aggregation can reduce specific attacks, but no single sanitizer proves that a large or adaptive dataset is clean.
Key Characteristics
- Acts through data consumed during training, fine-tuning, online learning, or federated updates
- Can target broad availability, selected predictions, or trigger-activated behavior
- Includes dirty-label, clean-label, feature, selection, deletion, and feedback-loop variants
- May preserve aggregate validation accuracy while harming a narrow target or slice
- Requires malicious intent and must be separated from noise, bias, drift, and leakage
- Is evaluated under explicit access, knowledge, poison-budget, and persistence assumptions
Common Use Cases
- Threat-modeling web-scale pretraining and third-party dataset ingestion
- Red-teaming fine-tuning pipelines that accept user or vendor examples
- Testing online learners exposed to feedback manipulation
- Auditing malicious clients in federated or collaborative training
- Planning quarantine, rollback, and retraining after poisoned records are discovered
Example
Loading code...Frequently Asked Questions
How is Data Poisoning different from ordinary bad data?
Poisoning is intentional manipulation designed to influence a learning outcome. Bad data can result from sensor faults, ambiguous labels, stale records, or sampling error without an adversary. The symptoms may overlap, so an investigation needs provenance, access evidence, timing, and a reproducible attack objective rather than inferring intent from low accuracy alone.
Is every Backdoor Attack a Data Poisoning attack?
No. Training-data poisoning is a common way to implant a backdoor, but an attacker can also modify model weights, training code, an adapter, or a supplied checkpoint. Conversely, many poisoning attacks only degrade utility or alter selected predictions and do not create a reusable trigger-conditioned backdoor.
Can clean validation accuracy rule out Data Poisoning?
No. A targeted or backdoor attack may preserve average clean accuracy while changing one class, subgroup, rare input, or trigger condition. Validation should include source and slice analysis, target-specific tests, trusted holdouts, trigger or anomaly probes, multiple seeds, and comparison with the previous approved dataset and model.
Which controls reduce Data Poisoning risk?
Use authenticated data sources, immutable lineage, checksums, access separation, review of label and transformation changes, source-level quotas, deduplication, trusted holdouts, staged ingestion, robust training where justified, and rollback-ready artifacts. Monitor both data distributions and model behavior. Each control addresses a different attacker capability and none proves universal immunity.
How should a poisoning defense be evaluated?
Fix the threat model, poison budget, attack knowledge, task, dataset version, and random seeds. Report clean utility, attack objective success, false positives, subgroup effects, compute cost, persistence, and recovery behavior. Test adaptive attacks that know the defense, and compare with accidental corruption so the claimed gain is attributable to the defense.