What is DVPS?

DVPS (Depth-Aware Video Panoptic Segmentation) is a dense video-perception task that jointly predicts per-pixel depth, semantic class, and temporally consistent Thing-instance identity.

Quick Facts

Full NameDepth-Aware Video Panoptic Segmentation
SpecificationOfficial Specification

How It Works

Predict depth, semantics, and identity on aligned pixels

The ViP-DeepLab paper defines DVPS as the joint prediction of monocular depth and video panoptic labels. Every evaluated pixel needs a depth value and semantic class; Thing regions additionally need temporally coherent instance IDs. The three outputs must share image coordinates, crop, resize, timestamp, and validity masks.

A model can use separate heads, task-specific decoders, shared object queries, or a hybrid representation. Joint training may transfer useful geometry and object cues, but it can also create gradient competition. Subtask metrics remain necessary because one aggregate does not identify whether depth, segmentation, or tracking failed.

Keep dataset derivation and visibility explicit

The official release derives Cityscapes-DVPS by combining Cityscapes-VPS panoptic labels with Cityscapes depth, while SemKITTI-DVPS projects annotated SemanticKITTI point clouds into the image plane. These sources differ in image cadence, depth density, occlusion, field of view, taxonomy, and sequence structure.

Projected LiDAR depth is not equivalent to dense camera depth: many pixels can remain invalid, and collisions require a projection policy. Resizing depth like a color image, evaluating across mismatched frames, or treating missing depth as zero error can silently invalidate results. Persist calibration and preprocessing metadata with every prediction artifact.

Evaluate the joint task and each failure surface

DVPQ converts pixels whose relative depth error exceeds a threshold into an ignored prediction category, then applies windowed Video Panoptic Quality. It therefore requires segmentation, identity, and depth to be correct at the same pixels. Report every depth threshold and temporal window rather than an unlabeled average.

Also report VPQ or STQ without depth gating, depth metrics on the same valid mask, Thing/Stuff and per-class scores, temporal degradation, and latency. DVPQ does not measure uncertainty calibration, completed geometry behind occlusions, physical scale safety outside the dataset, or causal response time.

Key Characteristics

  • Predicts depth and panoptic labels in the same image coordinates
  • Preserves Thing-instance identities across video frames
  • Covers Stuff and Thing pixels under one dense output contract
  • Requires synchronized frames, calibration, and valid-depth masks
  • Supports separate, shared, and hybrid multi-task architectures
  • Depends on depth units, dataset derivation, and temporal protocol

Common Use Cases

  1. Estimating scene geometry while tracking road users
  2. Building monocular video perception for autonomous systems
  3. Studying shared representations across depth and segmentation
  4. Comparing Cityscapes-DVPS and SemKITTI-DVPS pipelines
  5. Auditing aligned depth, mask, class, and identity failures

Example

loading...
Loading code...

Frequently Asked Questions

What does a DVPS model predict?

For each evaluated video pixel, it predicts a semantic class and depth. Thing pixels also receive an instance identity that stays coherent across frames, while Stuff pixels remain part of the complete panoptic partition.

How is DVPS different from video panoptic segmentation?

VPS predicts dense semantic and temporal instance labels but no geometry. DVPS adds an aligned depth estimate for each valid pixel, so evaluation can require spatial, semantic, and identity correctness together.

Does DVPS reconstruct a complete 3D scene?

No. Common DVPS protocols estimate depth for visible image pixels. They do not necessarily fill occluded space, fuse a persistent world model, recover absolute scale outside calibration assumptions, or guarantee geometrically complete point clouds.

Why are Cityscapes-DVPS and SemKITTI-DVPS not interchangeable?

One combines camera panoptic annotations with depth, while the other projects labeled LiDAR points into the image. Their depth density, cadence, taxonomy, visibility, image geometry, sequence boundaries, and evaluation settings differ.

How should a production DVPS system be evaluated?

Report DVPQ by threshold and window, ungated VPQ or STQ, depth errors on one valid mask, per-class and Thing/Stuff slices, identity failures, calibration and synchronization checks, causal latency, memory, and robustness under weather and sensor faults.

Related Terms

Related Articles