What is STQ?

STQ (Segmentation and Tracking Quality) is a video panoptic segmentation metric that combines semantic mean Intersection over Union with full-sequence, class-agnostic pixel association through a geometric mean.

Quick Facts

Full NameSegmentation and Tracking Quality
SpecificationOfficial Specification

How It Works

Measure class-agnostic association over complete sequences

For Ground Truth Thing track g and predicted identity p, TPA = |g intersect p|, FNA = |g| - TPA, and FPA = |p| - TPA, counted in pixels across the sequence. Their Association IoU is TPA / (TPA + FNA + FPA). The score for g is sum_p TPA * AssociationIoU(p,g) / |g|, and AQ averages this over Ground Truth tracks.

The STEP paper makes AQ class-agnostic: a Van mislabeled as Car is penalized by semantic SQ, not again by association. Reusing an earlier identity after a temporary mistake can improve AQ, while fragmentation and merges reduce overlap without a hard segment threshold.

Compute semantic SQ as mIoU, then balance both terms

SQ = mIoU is computed from the semantic confusion matrix over all evaluated pixels and classes, ignoring Instance IDs. This SQ differs from the identically named PQ component, which averages IoU only over segment pairs that already passed a matching threshold. Report the full name of each component to avoid mixing the two.

STQ = sqrt(AQ * SQ) gives both components symmetric multiplicative influence and reaches zero if either is zero. The score does not encode Confidence Ranking, Boundary Quality, latency, calibration, physical risk, or application cost, so AQ, SQ, per-class IoU, and track slices should accompany STQ.

Reproduce dataset-level aggregation and ignore rules

The KITTI-STEP benchmark uses dense semantic labels for every pixel and Track IDs for Car and Pedestrian. The pinned DeepLab2 implementation excludes Ground Truth Crowd pixels from AQ, aggregates semantic confusion matrices globally, averages AQ over Ground Truth Tubes, and also emits per-sequence scores.

STQ differs from VPQ's windowed, thresholded Tube matching. Its algebra resembles LSTQ, but STQ evaluates image pixels in Video Panoptic Segmentation while LSTQ applies a point-centric contract to 4D LiDAR. Dataset scores are not interchangeable without matching domains and protocols.

Key Characteristics

  • Combines semantic mIoU and association with a geometric mean
  • Measures association from full-sequence pixel intersections
  • Avoids threshold-based segment matching for AQ
  • Keeps semantic class errors out of the association component
  • Allows identity recovery to regain association credit
  • Averages association over Ground Truth Thing tracks

Common Use Cases

  1. Ranking KITTI-STEP video panoptic segmentation submissions
  2. Evaluating long-sequence pixel association without Tube thresholds
  3. Separating semantic errors from identity-association errors
  4. Diagnosing fragmentation, merges, and class recovery
  5. Comparing dense video scene-understanding systems

Example

loading...
Loading code...

Frequently Asked Questions

How is STQ calculated?

STQ is `sqrt(AQ * SQ)`. AQ averages each Ground Truth Thing track's pixel-weighted Association IoU contributions over predicted identities, while SQ is semantic mean IoU over evaluated classes.

Does STQ use an IoU threshold to match segments?

No for its association component. Every overlapping Ground Truth and predicted Thing identity contributes according to pixel intersection and union, so below-threshold association evidence is not discarded.

Why is STQ association class-agnostic?

Semantic mistakes already reduce SQ. Ignoring class labels in AQ prevents double punishment and lets a stable identity retain association credit when its semantic label is corrected later.

How is STQ different from VPQ?

STQ combines full-sequence, threshold-free pixel association with semantic mIoU. VPQ constructs tubes within configured windows, matches them only above IoU 0.5, and applies the PQ recognition denominator.

How is STQ different from LSTQ?

They use closely related semantic-and-association decompositions, but STQ is defined for video pixels and STEP-style protocols; LSTQ is point-centric for 4D LiDAR and inherits different taxonomies, ID namespaces, sampling, and ignore rules.

Related Terms

Related Articles