What is DVPQ?

DVPQ (Depth-Aware Video Panoptic Quality) is a joint metric that masks panoptic predictions at pixels exceeding a relative depth-error threshold and then applies Video Panoptic Quality over a declared temporal window.

Quick Facts

Full NameDepth-Aware Video Panoptic Quality
SpecificationOfficial Specification

How It Works

Gate panoptic predictions with relative depth error

The original DVPS paper compares predicted depth d_hat with valid Ground Truth depth d using relative error abs(d_hat - d) / d. If that error exceeds lambda, the pixel's predicted semantic label is replaced by Void before panoptic evaluation. Pixels within the threshold retain their predicted class and instance ID.

Depth gating is asymmetric with respect to Ground Truth: a rejected prediction does not remove the Ground Truth pixel. It can reduce Tube overlap or create unmatched evidence. Invalid Ground Truth depth must follow the dataset mask and must not be silently treated as an accurate zero.

Apply VPQ after gating inside each temporal window

After depth masking, DVPQ concatenates the gated predictions and Ground Truth labels over each k-frame window, forms same-ID Tubes, matches same-class Tubes at IoU strictly above 0.5, and uses the PQ denominator TP + 0.5*FP + 0.5*FN. Statistics are aggregated before class macro averaging, with Thing and Stuff results reported separately.

A depth error can therefore lower the metric even when the original panoptic label was correct. Conversely, accurate depth does not rescue a wrong class, fragmented identity, missed Tube, or poor mask. DVPQ intentionally couples the errors; companion metrics are needed for diagnosis.

Preserve the threshold-window grid and evaluator behavior

The original protocol evaluates lambda in {0.1, 0.25, 0.5}. It uses k in {1, 2, 3, 4} for six-frame Cityscapes-DVPS clips and {1, 5, 10, 20} for longer SemKITTI-DVPS sequences, then averages the configured combinations. These dataset-specific window sets must not be mixed.

The pinned reference evaluator masks only valid-depth pixels whose relative error is greater than the threshold, uses strict Tube IoU > 0.5, and applies dataset-specific Void handling. Report the code version, depth encoding and scale, valid mask, class map, sequence grouping, every cell in the grid, VPQ without depth gating, and a standalone depth metric.

Key Characteristics

  • Masks predictions using a relative depth-error threshold
  • Applies windowed Video Panoptic Quality after depth gating
  • Couples depth, mask, class, recognition, and identity errors
  • Uses strict same-class Tube matching above 0.5 IoU
  • Supports Thing, Stuff, per-threshold, and per-window reports
  • Depends on depth scale, valid mask, window grid, and evaluator version

Common Use Cases

  1. Benchmarking Cityscapes-DVPS models
  2. Evaluating SemKITTI-DVPS temporal perception
  3. Testing whether panoptic labels remain valid at accurate depths
  4. Measuring degradation across stricter depth thresholds
  5. Diagnosing joint depth, segmentation, and tracking systems

Example

loading...
Loading code...

Frequently Asked Questions

How is DVPQ calculated?

Mask each predicted panoptic pixel whose relative depth error exceeds `lambda`, aggregate the remaining labels into `k`-frame Tubes, apply strict same-class VPQ matching and the PQ denominator, then macro-average classes. A protocol may average several lambda and k combinations.

What do lambda and k mean in DVPQ?

`lambda` is the maximum accepted relative depth error, and `k` is the number of frames in the evaluation clip. Smaller lambda makes depth stricter; larger k tests identity and mask consistency over a longer temporal span.

How is DVPQ different from VPQ?

VPQ evaluates panoptic Tubes without depth. DVPQ first converts pixels with excessive relative depth error to an ignored prediction category, so inaccurate depth can reduce Tube overlap and recognition even when the original panoptic label was correct.

Does a high DVPQ identify which subtask is strong?

No. Depth, mask, class, detection, and identity errors interact before aggregation. Report the complete DVPQ grid alongside ungated VPQ or STQ, standalone depth metrics, per-class scores, and Thing/Stuff splits.

Can Cityscapes-DVPS and SemKITTI-DVPS DVPQ be compared directly?

No. They use different source annotations, depth density, sequence lengths, window sets, taxonomies, and visibility. Even within one dataset, depth scale, valid mask, preprocessing, sequence grouping, and evaluator commit must match.

Related Terms

Related Articles