What is Track Mean Average Precision?
Track Mean Average Precision is a confidence-ranked tracking metric that computes average precision after matching predicted and ground-truth trajectories by a declared track-similarity threshold.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define similarity over an entire trajectory
Track mAP needs a trajectory similarity rather than only a frame-level similarity. A common box or mask definition is spatiotemporal IoU: sum the intersection area over every frame and divide by the sum of union area, treating an absent object as an empty region. Another historical variant first thresholds frame-level matches and computes a Jaccard score over their TP, FN, and FP counts.
The YouTube-VIS paper uses mask-track IoU and averages AP over IoU thresholds from 0.50 through 0.95. TAO uses box-track 3D IoU and a federated annotation protocol, showing why the representation and labeling policy belong in the metric contract.
Rank tracks, match greedily, and integrate precision-recall
Within each class, predicted tracks are sorted by decreasing confidence. A prediction becomes a true-positive track if it exceeds the trajectory-similarity threshold and claims an eligible, still-unmatched Ground Truth; otherwise it is a false positive. Cumulative precision and recall form a curve whose interpolated area is AP.
The reported mAP may average AP across classes, trajectory-IoU thresholds, datasets, or all three. TrackEval defaults to thresholds 0.50:0.05:0.95 and 101 recall samples. A report labeled only Track mAP is incomplete unless it declares these axes.
Use Track mAP for ranked discovery, not failure diagnosis
TrackEval's TrackMAP implementation supports box and mask tracks, confidence ranking, area ranges, duration ranges, crowd ignores, and categories that are not exhaustively labeled. These choices are essential in open-world datasets such as TAO.
A hard trajectory threshold gives no partial credit below the cutoff, and assigning one confidence to a long track is itself a modeling choice. Report AP by threshold, class, size, and duration when possible, then add HOTA, IDF1, MOTA, or track-set distances to expose localization and association failure modes.
Key Characteristics
- Ranks complete predicted trajectories by confidence
- Matches tracks using a declared trajectory similarity
- Builds an interpolated precision-recall curve per class
- Can average over classes and multiple IoU thresholds
- Supports bounding-box and segmentation-mask trajectories
- Uses hard match thresholds and needs calibrated track scores
Common Use Cases
- Evaluating video instance segmentation systems
- Ranking large-vocabulary trackers on TAO-style benchmarks
- Comparing confidence-scored track detection systems
- Breaking down performance by class, size, or duration
- Auditing temporal-spatial coverage of complete trajectories
Example
Loading code...Frequently Asked Questions
How is Track mAP calculated?
For each class, sort predicted tracks by confidence, greedily match each to an eligible unmatched ground-truth trajectory above the similarity threshold, build cumulative precision and recall, and integrate interpolated precision. The final mean may also average classes and thresholds.
What is trajectory IoU?
Trajectory IoU sums spatial intersections over all frames and divides by the summed unions over those frames, with absent regions treated as empty. It jointly reflects localization, temporal support, misses, and extra predictions.
How is Track mAP different from detection mAP?
Detection mAP ranks and matches individual frame detections. Track mAP ranks and matches complete trajectories, so identity fragmentation or incorrect temporal extent reduces trajectory similarity even when many individual detections are accurate.
Why does Track mAP require a track confidence score?
Average precision is defined from a ranking of predictions. If a system emits per-detection confidence only, the benchmark must prescribe how those values become one track score; averaging is common but not universal.
Can Track mAP replace HOTA or IDF1?
No. Track mAP is useful for confidence-ranked track discovery and multi-class benchmarks, but hard thresholds and global matching make failure diagnosis difficult. HOTA and IDF1 answer different detection and identity questions.