What is VPQ?
VPQ (Video Panoptic Quality) is a windowed video panoptic segmentation metric that applies Panoptic Quality to same-class spatiotemporal tubes formed by consistent semantic and instance IDs.
Quick Facts
| Full Name | Video Panoptic Quality |
|---|---|
| Specification | Official Specification |
How It Works
Build tubes inside a declared temporal window
The original VPS paper groups frame segments with the same (class, ID) pair into a Tube. Tube intersection and union are accumulated over the annotated frames inside one sliding window. A same-class pair with Tube IoU > 0.5 is TP; unmatched predicted and Ground Truth Tubes become FP and FN.
Identity fragmentation splits one Ground Truth object across predicted Tubes, reducing overlap and adding unmatched Tube penalties. An ID merge can enlarge a predicted Tube with pixels from another object. Spatial mask, semantic, detection, and identity errors therefore interact before the threshold rather than appearing as independently reported components.
Aggregate per class and across configured spans
For span k and class c, VPQ_c^k = sum TubeIoU / (TP + 0.5*FP + 0.5*FN). Statistics are accumulated across windows and videos before class scores are macro-averaged. In the original Cityscapes-VPS protocol, annotations occur every lambda = 5 frames, windows slide by that stride, k belongs to {0, 5, 10, 15}, and final VPQ averages the four VPQ^k values.
Here k is frame-index span, not the count of annotated masks: k = 0 uses one annotated frame and equals image PQ, while k = 10 uses frames t, t+5, and t+10. Other datasets can use different spans or report a single VPQ window, so the bare label VPQ is incomplete without protocol details.
Read implementation details and companion metrics
The pinned VPSNet reference evaluator stacks annotated masks, aggregates segment area by ID, applies the strict Tube-IoU threshold, handles Void and Crowd similarly to PanopticAPI, and reports All, Thing, and Stuff scores for each configured span.
VPQ emphasizes identities that remain coherent within its windows, but a short-window score does not prove arbitrary-length tracking. Report every VPQ^k, Thing/Stuff splits, and window settings. Use STQ when full-sequence threshold-free pixel association and separate semantic mIoU are required.
Key Characteristics
- Extends Panoptic Quality from frame segments to temporal tubes
- Matches same-class tubes at IoU strictly above 0.5
- Uses the PQ denominator for unmatched tube penalties
- Macro-averages classes after aggregating window statistics
- Can average multiple temporal spans into one reported score
- Depends on annotation stride, window, sequence, Void, and Crowd rules
Common Use Cases
- Evaluating Cityscapes-VPS under its original window protocol
- Measuring short-range mask and identity consistency together
- Comparing Thing and Stuff video-panoptic performance
- Diagnosing degradation as temporal windows grow
- Validating video-panoptic exports with stable segment IDs
Example
Loading code...Frequently Asked Questions
How is Video Panoptic Quality calculated?
Within each temporal window, aggregate same-ID segments into tubes, match same-class tubes at IoU strictly above 0.5, compute the PQ formula per class, and macro-average classes. A protocol may then average several window spans.
Is VPQ with k equals zero the same as image PQ?
Yes under the original definition: `k = 0` contains one annotated frame, so tube matching reduces to frame-segment matching. The evaluator's class, Void, Crowd, and encoding rules must still match.
Why must a VPQ score include its window settings?
Longer windows expose more identity fragmentation and merge errors, while annotation stride changes which frames enter a tube. Different span sets and averaging policies therefore define different measurements even when all are called VPQ.
How is VPQ different from STQ?
VPQ threshold-matches spatiotemporal tubes inside configured windows and couples mask, recognition, and identity errors. STQ uses full-sequence, threshold-free pixel association plus a separate semantic mIoU component.
Does a high VPQ prove long-term tracking quality?
Only for the evaluated window spans and sequences. A short-window average can miss failures after long occlusion or re-entry, and it says nothing about runtime latency, Online/Offline access, calibration, or operational safety.