What is MOTSA and sMOTSA?

MOTSA and sMOTSA are multi-object tracking and segmentation metrics that extend MOTA to pixel masks, with sMOTSA replacing each hard true-positive count by its matched-mask Intersection over Union.

Quick Facts

SpecificationOfficial Specification

How It Works

Match non-overlapping instance masks in each frame

Ground Truth and predicted masks are matched within each frame when their mask IoU exceeds 0.5. In the original MOTS protocol, masks within each result or annotation are non-overlapping, so more than one prediction cannot exceed 0.5 IoU with the same truth mask. Ignore regions and neighboring classes are filtered by the evaluator before unmatched predictions become FP.

The original MOTS paper also defines identity switches from the matched prediction identity associated with a Ground Truth trajectory over time. Changing the mask representation, overlap threshold, ignore policy, or switch convention changes the score.

Separate hard accuracy from soft mask credit

Let M be all Ground Truth masks, and let softTP = sum IoU(h, c(h)) over accepted prediction-to-truth matches. Then MOTSA = (TP - FP - IDS) / |M|, equivalently 1 - (FN + FP + IDS) / |M|. Its accepted TP contributes one regardless of IoU above the threshold.

sMOTSA = (softTP - FP - IDS) / |M| replaces that unit credit with mask overlap. MOTSP = softTP / TP reports average IoU among accepted matches. MOTSA and sMOTSA can be negative when FP and IDS costs exceed credited matches.

Use the decomposition to diagnose mask tracking

The official MOTS evaluation tools aggregate TP, FN, FP, IDS, and summed IoU before computing dataset-level scores. Averaging sequence-level ratios without the benchmark's weighting can produce another result.

A higher sMOTSA can come from better masks, fewer misses, fewer false masks, or fewer identity switches, so it does not isolate which subsystem improved. Report MOTSA, MOTSP, raw counts, HOTA components, and task-specific latency or safety slices. J&F serves prompt- or first-frame-guided video object segmentation and does not impose the same discovery and identity penalties.

Key Characteristics

  • Evaluates temporally identified instance-segmentation masks
  • Uses mask IoU above 0.5 for accepted matches
  • MOTSA gives every accepted match one unit of credit
  • sMOTSA replaces hard TP credit with summed mask IoU
  • Subtracts false positives and identity switches
  • Can be negative and depends on ignore and aggregation rules

Common Use Cases

  1. Reproducing KITTI MOTS and MOTSChallenge evaluations
  2. Comparing joint detection, segmentation, and tracking systems
  3. Measuring whether mask quality changes tracking accuracy
  4. Auditing identity switches in pixel-level object tracks
  5. Reporting legacy MOTS scores alongside HOTA components

Example

loading...
Loading code...

Frequently Asked Questions

How are MOTSA and sMOTSA calculated?

MOTSA divides `TP - FP - IDS` by the number of Ground Truth masks. sMOTSA replaces TP with the sum of mask IoU over accepted matches. Both use counts accumulated under the benchmark's matching and ignore rules.

What is the difference between MOTSA and sMOTSA?

MOTSA gives every accepted mask match full credit, even if its IoU barely passes the threshold. sMOTSA gives that match only its IoU value, so it incorporates segmentation quality into the combined tracking score.

What is the difference between sMOTSA and MOTSP?

sMOTSA divides summed matched IoU minus FP and IDS by all Ground Truth masks. MOTSP divides only summed matched IoU by TP, so it measures localization quality among accepted matches without directly penalizing misses or extra masks.

Can MOTSA or sMOTSA be negative?

Yes. Their upper bound is one, but false positives and identity switches can exceed credited true positives or IoU mass. A negative score indicates that penalties outweigh credited matches under that protocol.

How do MOTSA and sMOTSA differ from J&F?

MOTSA and sMOTSA evaluate discovery, identity, and masks for multi-object tracks. DAVIS J&F averages region overlap and boundary quality for task-defined objects; its semi-supervised and unsupervised protocols use different identity or proposal matching rules.

Related Terms

Related Articles