What is PartPQ?

PartPQ (Part-Aware Panoptic Quality) is a class-averaged metric that evaluates panoptic segment recognition and, for classes with annotated parts, the quality of part labels inside matched parent segments.

Quick Facts

Full NamePart-Aware Panoptic Quality
SpecificationOfficial Specification

How It Works

Match parent segments before scoring parts

The original PPS paper first matches predicted and Ground Truth parent segments of the same scene class using binary mask IoU > 0.5. Unmatched predictions and Ground Truth segments become FP and FN. Part labels cannot rescue a parent pair that fails this scene-level gate.

This ordering preserves Panoptic Quality's recognition semantics. It also means a PartPQ decrease can originate from a missed parent, a spurious parent, poor parent overlap, or wrong internal part labels. Report raw TP/FP/FN and class-level results before attributing the change to the part head.

Replace matched overlap with part-level mean IoU

For classes with parts, each accepted parent pair contributes IoU_p, the mean Intersection over Union over the evaluated part labels within that parent context. For classes without parts, IoU_p is the binary parent-segment IoU. Per class, PartPQ_c = sum IoU_p / (TP + 0.5*FP + 0.5*FN) and can be decomposed into PartSQ and PartRQ.

PartRQ uses the same accepted parent matches and denominator as PQ, while PartSQ averages the credited IoU_p values over TP. A strong PartRQ with weak PartSQ can indicate that objects are found but their masks or parts are poor; PartSQ alone cannot reveal missed objects.

Pin EvalSpec, ignored regions, and implementation

The pinned reference evaluator derives class groups from an EvalSpec, removes Ground Truth segments whose eligible part region is entirely unlabeled, excludes undefined part pixels plus Void/Crowd from part confusion, and suppresses unmatched predictions dominated by ignored regions.

Published protocols report all classes, classes with parts, and classes without parts separately. PartPQ is not interchangeable with PQ, part mIoU, or Part-Whole Quality, and it does not measure boundary safety, hierarchy calibration, latency, or temporal consistency. Compare only runs with the same dataset version, label grouping, EvalSpec, and evaluator commit.

Key Characteristics

  • Matches same-class parent segments at IoU strictly above 0.5
  • Uses part-label mean IoU for matched classes with parts
  • Uses parent binary IoU for classes without parts
  • Retains the PQ denominator for FP and FN penalties
  • Macro-averages valid scene classes and supports split reports
  • Depends on EvalSpec, part grouping, Void, and Crowd handling

Common Use Cases

  1. Ranking models on Cityscapes Panoptic Parts
  2. Evaluating PASCAL Panoptic Parts predictions
  3. Separating parent recognition from matched-part quality
  4. Comparing classes with parts against classes without parts
  5. Validating hierarchical segmentation exports and EvalSpecs

Example

loading...
Loading code...

Frequently Asked Questions

How is PartPQ calculated?

Match same-class parent segments at binary IoU strictly above 0.5. Sum part-label mean IoU for matched classes with parts, or binary mask IoU for classes without parts, then divide by `TP + 0.5 FP + 0.5 FN` and macro-average valid classes.

How is PartPQ different from Panoptic Quality?

Both use the same parent-segment recognition gate and denominator. PQ credits each match with parent mask IoU; PartPQ credits classes with parts using internal part-label mean IoU, while no-part classes retain parent IoU.

Can correct part labels compensate for a missed parent segment?

No. Part labels are scored only after the predicted and Ground Truth parent segments pass the same-class `IoU > 0.5` gate. A failed parent match contributes FP and FN rather than part overlap.

What do PartSQ and PartRQ mean?

PartSQ averages the credited match quality over accepted parent pairs. PartRQ is the F1-like recognition term from TP, FP, and FN. Their product is PartPQ for that class, but both inherit the evaluator's part and ignore rules.

Can PartPQ scores from different EvalSpecs be compared?

Not reliably. Scene classes, part classes, grouped labels, unknown-part treatment, Void/Crowd masks, and eligible segments can differ. Report the dataset version, EvalSpec, evaluator commit, and with-parts/no-parts splits.

Related Terms

Related Articles