What is Multiple Object Tracking Accuracy?
Multiple Object Tracking Accuracy is a CLEAR MOT score that subtracts missed detections, false detections, and identity switches from the total number of ground-truth detections.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Count errors only after protocol-defined matching
At each frame, the evaluator creates a one-to-one matching between eligible ground-truth and predicted detections. Eligibility depends on the declared localization similarity and threshold, such as bounding-box IoU of at least 0.5. Unmatched truth detections become FN, unmatched predictions become FP, and a matched truth identity changing from its previous matched prediction creates an ID switch under the evaluator's convention.
The original CLEAR MOT paper defines the metric family. Modern implementations may differ in ignore handling, continuity rules, and aggregation, so the formula alone is not a complete evaluation protocol.
Compute MOTA and inspect the error decomposition
MOTA = 1 - (FN + FP + IDSW) / GT, where GT is the total number of ground-truth detections across the evaluated sequences. Equivalently, MOTA = (TP - FP - IDSW) / (TP + FN). A perfect result is 1, while errors exceeding GT make MOTA negative; zero is not a universal baseline.
The Go example aggregates counts before taking the ratio, which is the correct micro-averaging pattern. Averaging per-frame or per-sequence MOTA values without the benchmark's prescribed weights can produce a different score.
Pair MOTA with association and localization evidence
TrackEval's CLEAR implementation prioritizes continuity from the previous frame, then localization similarity, before computing TP, FN, FP, and IDSW. This makes class filters, ignored regions, similarity thresholds, and identity-switch conventions material to reproducibility.
MOTA gives every FN, FP, and IDSW the same additive cost, even though detections usually outnumber identity switches. Report HOTA or IDF1 for association behavior, MOTP or LocA for localization, and task-specific safety or latency measures instead of treating MOTA as a complete quality guarantee.
Key Characteristics
- Combines false negatives, false positives, and identity switches
- Normalizes total errors by ground-truth detection count
- Ranges from negative infinity to a perfect score of one
- Depends on frame-level matching and switch conventions
- Usually reflects detection quality more strongly than association
- Does not include localization quality among accepted matches
Common Use Cases
- Reproducing CLEAR MOT and MOTChallenge evaluations
- Comparing tracker error counts under one fixed protocol
- Auditing how missed and false detections affect a tracker
- Monitoring regressions in established tracking pipelines
- Reporting a legacy score alongside HOTA and IDF1
Example
Loading code...Frequently Asked Questions
How is MOTA calculated?
Add false negatives, false positives, and identity switches over the evaluation set, divide by the total number of ground-truth detections, and subtract the ratio from one. Aggregate counts before division unless the benchmark explicitly defines another weighting rule.
Can MOTA be negative?
Yes. MOTA has an upper bound of one but no finite lower bound. If the combined FN, FP, and IDSW count exceeds the number of ground-truth detections, the score becomes negative.
Does MOTA measure identity consistency well?
Only partially. It includes identity switches, but each switch has the same unit cost as one missed or false detection. Detection errors are often much more numerous, so MOTA can hide substantial association differences.
What is the difference between MOTA and MOTP?
MOTA counts missed detections, false detections, and identity switches. MOTP averages localization similarity or distance only over accepted matches. One tracker can therefore have high MOTP but poor MOTA, or the reverse.
Can MOTA scores from different benchmarks be compared?
Not reliably unless dataset, class filters, ignored regions, localization representation, matching threshold, switch definition, sequence aggregation, and evaluator version are compatible. Those choices alter the underlying counts.