What is Higher Order Tracking Accuracy?

Higher Order Tracking Accuracy is a multi-object tracking score that combines detection accuracy and higher-order association accuracy at each localization threshold, then averages across thresholds.

Quick Facts

SpecificationOfficial Specification

How It Works

Build detection and association accuracy from matched detections

At a localization threshold alpha, truth and predicted detections are matched within each frame. DetA_alpha = TP / (TP + FN + FP) measures the Jaccard accuracy of detection matching. For each true-positive match, its Association IoU compares how often that truth identity and predicted identity match against detections assigned to either identity but not both; AssA_alpha averages those values over true positives.

The original HOTA paper defines HOTA_alpha = sqrt(DetA_alpha * AssA_alpha). The Go example computes one threshold from already matched detections. Full HOTA must repeat matching across the declared threshold grid and average the resulting HOTA_alpha values.

Inspect the decomposition instead of optimizing one scalar

DetA separates into detection precision and recall, while AssA separates into association precision and recall. LocA averages localization similarity over true-positive matches. These components reveal whether a score changed because of missing objects, false detections, fragmented identities, merged identities, or box alignment.

The geometric mean gives detection and association symmetric multiplicative influence at each threshold, but it does not encode application-specific harm. A surveillance workflow, robot safety monitor, and offline video index can value missed objects, identity continuity, latency, and localization differently.

Reproduce the benchmark protocol before comparing scores

TrackEval is the reference HOTA implementation used by MOTChallenge, KITTI Tracking, and other benchmarks. It also reports MOTA, MOTP, and Identity metrics, which helps compare HOTA with IDF1 under one preprocessing pipeline.

Scores are comparable only when class filters, ignored regions, visibility rules, localization similarity, threshold grid, duplicate handling, sequence weighting, and evaluator version match. Track mAP serves confidence-ranked track discovery, while ATA and local tracking metrics test strict or horizon-specific identity. OSPA(2) and T-GOSPA remain track-set distances with different error semantics.

Key Characteristics

  • Balances detection and association through a geometric mean
  • Evaluates higher-order identity alignment for every matched detection
  • Averages results across multiple localization thresholds
  • Provides DetA, AssA, LocA, precision, and recall diagnostics
  • Depends on benchmark preprocessing and matching conventions
  • Produces a normalized score rather than a metric-space distance

Common Use Cases

  1. Ranking visual multi-object trackers on standard benchmarks
  2. Separating detection failures from association failures
  3. Evaluating tracking through occlusion, fragmentation, and identity merges
  4. Comparing trackers across bounding-box or segmentation thresholds
  5. Auditing whether a benchmark improvement comes from detection or association

Example

loading...
Loading code...

Frequently Asked Questions

How is HOTA calculated?

At each localization threshold, HOTA takes the square root of Detection Accuracy times Association Accuracy. The reported HOTA score averages those threshold-specific values across the evaluator's localization-threshold grid.

What are DetA, AssA, and LocA?

DetA is the Jaccard accuracy of matched versus missed and false detections. AssA averages identity-pair Association IoU over matched detections. LocA averages localization similarity for true positives and is reported as a diagnostic component.

How is HOTA different from MOTA?

MOTA subtracts false negatives, false positives, and identity switches from a ground-truth detection count and can be dominated by detection errors. HOTA separately measures detection and association, then balances them at multiple localization thresholds.

How is HOTA different from IDF1?

IDF1 uses one sequence-level assignment between truth and predicted identities and emphasizes globally correct identity detections. HOTA gives partial association credit for each matched detection and explicitly balances that association quality with detection accuracy.

Can HOTA scores from different benchmarks be compared?

Only when datasets and evaluation protocols are compatible. Class selection, ignored regions, visibility, matching similarity, threshold grids, sequence aggregation, and evaluator versions can all change the score independently of the tracker.

Related Terms

Related Articles