What is Backdoor Attack?

A Backdoor Attack is an adversarial manipulation that implants hidden behavior in a machine learning model so an attacker-chosen trigger causes a targeted output while ordinary inputs can continue to behave normally.

Quick Facts

SpecificationOfficial Specification

How It Works

Define the trigger, target, and insertion path

The BadNets study demonstrated a model that preserved normal task performance but changed predictions when a chosen pattern appeared. A modern assessment must state whether the attacker controls records, labels, local updates, training code, adapters, or weights; whether the trigger is fixed, dynamic, semantic, or physical; and whether the target is a class, generated behavior, safety bypass, or downstream action.

Test clean behavior and triggered behavior separately

Report clean utility, attack success rate across eligible source classes, false activation on benign inputs, trigger prevalence, confidence, persistence after fine-tuning, and results under transformations such as crop, paraphrase, noise, or context changes. Hold out trigger families and source classes from attack development. Aggregate accuracy alone can hide a reliable backdoor, while one successful trigger does not prove universal exploitability.

Use supply-chain and behavioral controls together

Pin model, dataset, adapter, and training-code digests; verify provenance and signatures; separate build and approval identities; and retain a trusted evaluation set. Add trigger search, activation or representation analysis, targeted fuzzing, output invariants, and canary monitoring where the threat warrants them. Fine-tuning, pruning, filtering, or Machine Unlearning may reduce a tested backdoor but must be re-evaluated against adaptive triggers and clean-utility regressions.

Key Characteristics

  • Creates trigger-conditioned attacker-chosen behavior
  • Can remain dormant while aggregate clean performance looks normal
  • May be inserted through data, updates, code, adapters, or model weights
  • Uses visible, invisible, lexical, semantic, physical, or contextual triggers
  • Requires joint measurement of clean utility, attack success, and false activation
  • Can persist across fine-tuning, transfer, compression, or deployment transformations

Common Use Cases

  1. Auditing pretrained models and adapters from third-party registries
  2. Red-teaming outsourced training and fine-tuning pipelines
  3. Testing federated models for malicious-client trigger behavior
  4. Verifying safety-critical vision, speech, and text classifiers
  5. Investigating a narrow production failure that appears only under a repeatable condition

Example

loading...
Loading code...

Frequently Asked Questions

How is a Backdoor Attack different from Data Poisoning?

Data Poisoning describes malicious control of training data and can cause broad degradation, selected errors, or a backdoor. A Backdoor Attack describes the resulting hidden trigger-conditioned behavior and may instead be implanted through weights, code, adapters, or a supplied checkpoint. The concepts overlap, but neither is a synonym for the other.

Why can a backdoored model pass normal validation?

The attacker optimizes the model to behave normally when the trigger is absent, so a clean validation set may sample only the benign behavior. Detection requires trigger-aware or behavior-oriented tests, provenance review, and analysis across source classes and transformations. High clean accuracy is therefore a required utility measure, not evidence that no backdoor exists.

What metrics should a Backdoor Attack evaluation report?

Report clean utility, attack success rate on eligible triggered inputs, false-trigger activation on benign inputs, target and source classes, trigger visibility and prevalence, confidence, persistence after model changes, and uncertainty across seeds. The test should include unseen trigger variants and adaptive attacks when making a defense claim.

Can fine-tuning or pruning remove a backdoor?

Sometimes, but neither is a general guarantee. A backdoor may share parameters with normal behavior, survive transfer learning, or adapt to the detector. Measure clean utility and attack success after every repair, test transformed and unseen triggers, preserve rollback, and reject the artifact when its provenance or residual risk is unacceptable.

Is Prompt Injection a model backdoor?

Usually not. Prompt Injection supplies adversarial instructions or retrieved content at inference time and can affect an otherwise unmodified model. A model backdoor is implanted before use and waits for a trigger. An attacker can combine them, but their trust boundaries, detection evidence, and mitigations differ.

Related Terms

Related Articles