What is Average Multi-Object Tracking Accuracy?
Average Multi-Object Tracking Accuracy is a confidence-swept multi-object tracking score that averages an accuracy value over predefined recall operating points instead of selecting one detection threshold.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Turn track confidence into recall operating points
The evaluator assigns each predicted trajectory a confidence, sorts or thresholds tracks, and chooses confidence cutoffs corresponding to a configured recall grid. At every cutoff it repeats frame matching and recomputes TP, FN, FP, identity switches, and localization statistics.
The original 3D MOT metrics paper introduced integral scores so evaluation would cover the full confidence-threshold spectrum. The current nuScenes devkit uses track-level confidence, a configured minimum recall and number of thresholds, and explicit worst values for recall targets a tracker cannot reach.
Average MOTAR under the current nuScenes contract
For evaluator-measured recall r, nuScenes computes MOTAR_r = max(0, 1 - ((FN + FP + IDS) - (1-r)*GT) / (r*GT)). Subtracting (1-r)*GT removes misses implied solely by operating at recall r, while division by r*GT normalizes the remaining error burden among available matches.
Current nuScenes then computes AMOTA = mean_r(MOTAR_r). Its published configuration uses 40 hypothetical recall targets from 0.1 to 1.0; an unattained target receives the worst AMOTA contribution of zero. The released code derives r from its MATCH event count divided by GT, with SWITCH events supplied separately, so it may differ from both the requested grid point and a conventional recall implementation.
Do not compare names without comparing protocols
The early paper defined unscaled AMOTA by averaging ordinary MOTA and called the recall-normalized variant sAMOTA. nuScenes later maps the public AMOTA field to averaged MOTAR. A result from another toolkit may therefore share the acronym while using a different per-recall formula, clipping rule, recall grid, or treatment of unreachable recall.
Matching geometry is equally material. nuScenes currently matches class-specific 3D tracks by ground-plane center distance with a 2 m cutoff, then averages class scores. AMOTA still mixes detection and identity errors and does not measure matched localization; pair it with AMOTP, HOTA, IDF1, raw counts, and deployment-specific safety measures.
Key Characteristics
- Sweeps confidence thresholds through predefined recall targets
- Averages recall-normalized MOTAR in the current nuScenes evaluator
- Penalizes recall targets a tracker cannot achieve
- Combines false positives, identity switches, and excess misses
- Depends on track confidence, matching, recall grid, and clipping rules
- Does not measure localization quality among accepted matches
Common Use Cases
- Ranking 3D multi-object trackers on the nuScenes benchmark
- Comparing tracker behavior across confidence operating points
- Auditing whether one selected threshold hides poor recall behavior
- Evaluating autonomous-driving perception with a fixed protocol
- Reproducing AMOTA tables with evaluator-version provenance
Example
Loading code...Frequently Asked Questions
How is AMOTA calculated in nuScenes?
nuScenes evaluates multiple track-confidence cutoffs tied to a recall grid, computes recall-normalized MOTAR at each achieved point, assigns zero to unreachable recall targets, and averages the resulting values. Final benchmark reporting also averages across its declared tracking classes.
What is the difference between AMOTA and MOTAR?
MOTAR is the recall-normalized accuracy calculated at one operating point. AMOTA aggregates MOTAR across the configured recall points. Reporting only AMOTA hides the shape of the MOTAR-versus-recall curve.
Is AMOTA the average of ordinary MOTA?
It depends on the evaluator. The original 3D MOT paper used AMOTA for averaged ordinary MOTA and sAMOTA for the scaled variant. The current nuScenes devkit maps AMOTA to the mean of MOTAR, so the implementation version is part of the score.
Why does AMOTA penalize unreachable recall?
Without a penalty, a tracker could omit difficult objects and average only its easy operating range. nuScenes assigns zero at recall targets above maximum achieved recall, making coverage part of the result.
Can AMOTA values from different datasets be compared?
Not directly unless the class set, confidence aggregation, recall grid, matching representation and cutoff, ignored-data rules, MOTAR formula, clipping, unreachable-recall policy, class aggregation, and evaluator version are compatible.