What is Average Multi-Object Tracking Precision?

Average Multi-Object Tracking Precision is a confidence-swept localization measure that averages the per-match MOTP value over a configured set of recall operating points.

Quick Facts

SpecificationOfficial Specification

How It Works

Reuse the same recall sweep as AMOTA

For each class, the evaluator converts trajectory confidence into cutoffs for a predefined recall grid. At every cutoff it keeps eligible tracks, matches predicted and Ground Truth objects, and computes MOTP over accepted pairs. AMOTP is the arithmetic mean of those per-cutoff MOTP values.

The original 3D MOT metrics paper introduced AMOTP to replace one hand-selected confidence threshold with an integral view. The recall grid, confidence aggregation, duplicate thresholds, and unreachable-recall policy are therefore part of the metric rather than plotting details.

Interpret current nuScenes AMOTP as meters of center error

The current nuScenes devkit computes XY Euclidean center distances, rejects pairs at or above its configured 2 m matching cutoff, and defines per-point MOTP_r = sum(distance of matches) / TP_r. It then averages MOTP over 40 hypothetical recall targets from 0.1 to 1.0.

If a class cannot reach a recall target, the published configuration inserts the worst AMOTP contribution of 2.0 m. This intentionally makes maximum recall affect the aggregate even though per-point MOTP only averages accepted matches.

Keep score direction, geometry, and averaging visible

A low nuScenes AMOTP means accepted centers are close on average; it does not mean the tracker found most objects or preserved their identities. An easy, low-recall subset can have excellent per-point MOTP, while penalties at unreachable recall points increase AMOTP.

Do not compare a meter-valued nuScenes AMOTP with an IoU-valued, higher-is-better AMOTP from another evaluator. Record coordinate frame, distance function, matching cutoff, recall targets, missing-point value, class weighting, and evaluator commit. Pair AMOTP with AMOTA, recall, HOTA or IDF1, and class-specific error slices.

Key Characteristics

  • Averages per-match localization across recall operating points
  • Uses lower-is-better XY center distance in current nuScenes
  • Carries physical units when the underlying MOTP is a distance
  • Penalizes recall targets that a tracker cannot reach
  • Depends on confidence, matching cutoff, coordinate frame, and class set
  • Does not directly count misses, false tracks, or identity switches

Common Use Cases

  1. Reporting nuScenes 3D tracking localization over a recall sweep
  2. Comparing tracker center error under one frozen evaluator
  3. Auditing localization changes separately from AMOTA
  4. Detecting confidence regions with unstable spatial estimates
  5. Reproducing leaderboard results with explicit units and penalties

Example

loading...
Loading code...

Frequently Asked Questions

How is AMOTP calculated in nuScenes?

At every configured recall target, nuScenes averages XY center distance over accepted matches. It inserts `2.0 m` for unreachable targets, averages all per-target values for each class, and then averages the declared class results.

Is lower AMOTP better?

Yes for current nuScenes because AMOTP is measured as center-distance error in meters. Some historical toolkits average IoU similarity instead, where higher is better, so the acronym alone does not reveal direction.

What is the difference between AMOTP and MOTP?

MOTP describes accepted-match localization at one confidence operating point. AMOTP averages MOTP across a recall grid and can include worst-value penalties for recall levels the tracker cannot reach.

Why can AMOTP worsen when matched localization is unchanged?

Maximum recall may have fallen, causing more recall targets to receive the worst-distance fallback. Confidence calibration can also select different track subsets at nominal recall points, changing which localization errors enter the average.

Can AMOTP prove that a 3D tracker is good?

No. It summarizes localization among accepted matches across the configured sweep. It does not independently expose false tracks, missed objects, identity continuity, latency, or safety; report AMOTA, recall, association metrics, and operational slices too.

Related Terms

Related Articles