What is Tracking Every Thing Accuracy?

Tracking Every Thing Accuracy is a large-vocabulary multi-object tracking metric that separately measures localization, association, and classification, then combines the three components with an arithmetic mean.

Quick Facts

SpecificationOfficial Specification

How It Works

Use local clusters to bound false-positive evidence

For each evaluated class, every Ground Truth box anchors a local cluster. Predictions sufficiently near an anchor under the configured IoU margin enter a cluster; those not selected as a match become localization false positives. Predictions outside all clusters are ignored in the non-exhaustive setting because the dataset may contain valid but unannotated objects.

The original TETA paper uses this policy to avoid treating every unmatched prediction as false while still penalizing duplicates near known objects. The cluster margin changes that tradeoff and must be reported.

Compute localization, association, and classification explicitly

LocA = LocTP / (LocTP + LocFP + LocFN) is a Jaccard score for localized detections. For every localized true positive, TETA computes a HOTA-style identity-pair Jaccard from true-positive, false-positive, and false-negative associations, then averages it as AssocA. ClsA is a classwise Jaccard over correctly and incorrectly classified well-localized matches.

The combined score is TETA = (LocA + AssocA + ClsA) / 3. The arithmetic mean prevents a near-zero long-tail classification score from collapsing localization and association to zero, but the three components still receive equal formal weight rather than application-specific cost.

Keep annotation mode and threshold sweep in the contract

The reference implementation evaluates arrays of localization thresholds, performs global identity alignment, and treats classification false positives differently when annotations are declared exhaustive. Class averaging and threshold reduction therefore affect the reported scalar.

TETA can reveal that a tracker localizes and associates an object despite semantic confusion, but it does not prove open-vocabulary recognition, calibration, safety, or real-time behavior. Compare its three components, base versus novel classes, object frequency, threshold curves, and raw errors; use HOTA, IDF1, or Track mAP when their benchmark contracts better match the task.

Key Characteristics

  • Separates localization, association, and classification quality
  • Groups predictions by location rather than predicted class
  • Supports non-exhaustive large-vocabulary annotations
  • Uses HOTA-style higher-order association accounting
  • Combines three components with an arithmetic mean
  • Depends on cluster margin, thresholds, annotation mode, and class averaging

Common Use Cases

  1. Evaluating large-vocabulary multi-object trackers
  2. Benchmarking open-world and open-vocabulary tracking
  3. Separating semantic confusion from trajectory association
  4. Handling datasets with non-exhaustively annotated categories
  5. Comparing base, novel, common, and rare class performance

Example

loading...
Loading code...

Frequently Asked Questions

How is TETA calculated?

TETA computes Localization Accuracy, higher-order Association Accuracy, and Classification Accuracy under its local-cluster protocol, then takes their arithmetic mean. The evaluator may additionally average over localization thresholds and classes.

Why does TETA separate classification from tracking?

Large-vocabulary trackers often localize and follow an object while confusing a rare or similar class. Class-first grouping would discard that tracking evidence. TETA groups by location, scores trajectory association, and reports classification separately.

How does TETA handle incomplete annotations?

In the non-exhaustive setting, it penalizes unmatched predictions inside local clusters around annotated objects but ignores predictions outside every cluster. This limits false punishment of valid unannotated objects, although the cluster margin remains a modeling assumption.

What is the difference between TETA and HOTA?

HOTA balances detection and association, normally after class-based evaluation. TETA further separates localization from classification and adds local-cluster handling for incomplete labels. For exhaustive single-class data, TETA becomes HOTA-like but uses an arithmetic rather than geometric mean.

Can a high TETA score prove open-world safety?

No. TETA is an offline benchmark aggregate. It does not encode class-specific harm, confidence calibration, latency, missed annotation outside clusters, distribution shift, or downstream planning risk; those require separate evaluation.

Related Terms

Related Articles