What is PQ?

PQ (Panoptic Quality) is a class-averaged metric for panoptic segmentation that combines the overlap of matched segments with penalties for unmatched predictions and ground-truth segments.

Quick Facts

Full NamePanoptic Quality
SpecificationOfficial Specification

How It Works

Match same-class segments above the strict threshold

For one class, a prediction and Ground Truth segment are a True Positive pair when IoU > 0.5. Because both panoptic maps contain non-overlapping segments, the original proof shows that this strict threshold makes matching unique; no greedy order or Hungarian assignment is needed. A pair at exactly 0.5 is not a match.

Unmatched predictions become FP and unmatched Ground Truth segments become FN. Category equality, Void pixels, Crowd regions, and malformed segment metadata must be resolved according to the evaluator before these counts are trusted.

Decompose PQ into matched-mask SQ and recognition RQ

For class c, PQ_c = sum IoU / (TP + 0.5*FP + 0.5*FN). It can be written as SQ_c * RQ_c, where SQ_c = sum IoU / TP and RQ_c = TP / (TP + 0.5*FP + 0.5*FN). RQ is F1-like, while SQ is conditional on accepted matches; a model can have high SQ after missing many objects, which RQ exposes.

The dataset score is the arithmetic mean of valid class scores. Segment area does not directly weight PQ inside a class, and class frequency does not directly weight the macro average. Report PQ_th, PQ_st, per-class PQ, SQ, and RQ to reveal which population drives a change.

Preserve evaluator handling and metric boundaries

The pinned COCO PanopticAPI evaluator removes Ground Truth Void overlap from a prediction's union, ignores Crowd regions as matches, and suppresses an unmatched prediction when more than half overlaps Void plus same-class Crowd. It also checks that PNG segment IDs and JSON metadata agree.

PQ has no confidence sweep and no temporal identity term. Do not compare it with AP, semantic mIoU, VPQ, STQ, sPTQ, or PAT as if the values shared one error model. Pin the taxonomy, split, encoding, ignored-region rules, and evaluator commit before comparing systems.

Key Characteristics

  • Matches same-class non-overlapping segments at IoU strictly above 0.5
  • Combines matched-mask overlap with FP and FN penalties
  • Decomposes into Segmentation Quality and Recognition Quality
  • Macro-averages valid classes rather than weighting by frequency
  • Reports separate Thing and Stuff summaries in common evaluators
  • Depends on Void, Crowd, category, and segment-metadata handling

Common Use Cases

  1. Ranking image panoptic segmentation systems
  2. Separating mask quality from missed and spurious segments
  3. Comparing Thing and Stuff performance under one protocol
  4. Auditing per-class scene-parsing failures
  5. Validating COCO-style panoptic prediction exports

Example

loading...
Loading code...

Frequently Asked Questions

How is Panoptic Quality calculated?

For each class, sum IoU over same-class segment pairs whose IoU is strictly greater than 0.5, then divide by `TP + 0.5 FP + 0.5 FN`. The reported PQ is the arithmetic mean over evaluated classes.

Why does PQ use an IoU threshold greater than 0.5?

With non-overlapping panoptic segments, a strict threshold above 0.5 guarantees one-to-one matching. Exactly 0.5 does not qualify, and lowering the threshold would require an explicit assignment strategy.

What is the difference between PQ's SQ and semantic mIoU?

PQ's SQ averages IoU only over matched segment pairs. Semantic mIoU merges all pixels by class and ignores instance identities. A model can therefore have strong matched-mask SQ while its RQ reveals many missed segments.

Does PQ weight large objects more than small objects?

Not directly. Within a class, matched segments contribute one IoU each and unmatched segments contribute counts; the final result macro-averages classes. Pixel area still affects each segment's IoU and benchmark ignore rules.

Can PQ scores from different datasets or evaluators be compared?

Not reliably unless the class taxonomy, Thing/Stuff mapping, split, Void and Crowd handling, segment encoding, minimum-area policy, and evaluator implementation are compatible. Report those details with the score.

Related Terms

Related Articles