What is sPTQ?

sPTQ is a class-averaged panoptic tracking metric that subtracts the IoU credit of matched segments at identity-switch frames from frame-level Panoptic Quality statistics.

Quick Facts

Full NameSoft Panoptic Tracking Quality
SpecificationOfficial Specification

How It Works

Start from frame-level Panoptic Quality matches

Within each class and frame, predicted and Ground Truth segments with IoU above the benchmark threshold become true-positive pairs. Accumulate their IoU, plus unmatched predicted segments as FP and unmatched Ground Truth segments as FN. Stuff classes contribute segmentation statistics but have no temporal instance switches.

The MOPT paper introduced PTQ and sPTQ for joint semantic, instance, and temporal evaluation. The input masks, Thing/Stuff taxonomy, IoU threshold, ignored points, minimum segment size, and class averaging are part of the metric contract.

Remove switch-frame IoU instead of a hard unit

For class c, sPTQ_c = (sum_TP IoU - sum_IDS IoU_switch) / (TP + 0.5*FP + 0.5*FN). The final score is the arithmetic mean over valid classes. A switched true-positive segment remains in the denominator but its current-frame IoU credit is removed from the numerator.

Hard PTQ_c subtracts the number of identity switches instead. It can penalize one switch by more than the matched IoU that was available, whereas sPTQ weights the loss by mask quality. This does not make the score threshold-free: the segment must first pass frame-level matching.

Reproduce the switch convention and read companion metrics

The current nuScenes Devkit compares a Ground Truth instance's accepted prediction in two adjacent frames and adds the current IoU when the predicted ID changes. If either frame lacks an accepted match, that transition is not the same switch event. Other evaluators may differ.

Report PTQ, sPTQ, PQ, raw TP/FP/FN/IDS, and evaluator version. Use LSTQ for point-centric full-sequence association, PAT for instance-centric Tracking Quality, and MOTSA/sMOTSA when only thing-mask tracks belong to the task.

Key Characteristics

  • Extends frame-level Panoptic Quality with identity penalties
  • Subtracts switched-segment IoU rather than one hard unit
  • Uses a TP plus half-FP plus half-FN denominator
  • Computes per-class scores before macro averaging
  • Includes stuff segmentation without stuff identity switches
  • Depends on segment matching and adjacent-frame switch rules

Common Use Cases

  1. Reproducing MOPT and Panoptic nuScenes evaluations
  2. Comparing panoptic trackers with different mask quality
  3. Auditing whether switch penalties dominate PTQ
  4. Evaluating joint thing tracking and stuff segmentation
  5. Reporting a legacy frame-centric metric beside LSTQ and PAT

Example

loading...
Loading code...

Frequently Asked Questions

How is sPTQ calculated?

For each class, add IoU over accepted frame-level segment matches, subtract the IoU of matches at identity-switch frames, and divide by `TP + 0.5 FP + 0.5 FN`. Average the valid class scores.

What is the difference between PTQ and sPTQ?

PTQ subtracts one full unit for each identity switch. sPTQ subtracts the switched segment's IoU, so its penalty follows the mask credit that the segment would otherwise contribute.

What is the difference between sPTQ and sMOTSA?

sPTQ uses the PQ denominator, applies IoU-weighted switch penalties, and class-averages thing and stuff scores. sMOTSA subtracts FP and hard IDS from summed mask IoU, normalizes by all Ground Truth thing masks, and excludes stuff.

Does sPTQ measure long-term association?

Only indirectly through identity-switch events defined by the evaluator. The nuScenes implementation compares accepted matches in adjacent frames; LSTQ and PAT provide different sequence- or instance-centric association views.

Can sPTQ be compared across benchmarks?

Only with compatible Thing/Stuff classes, segment-IoU threshold, minimum size, ignored points, identity-switch convention, sequence boundaries, class aggregation, and evaluator version.

Related Terms

Related Articles