What is LSTQ?
LSTQ is a point-centric 4D LiDAR panoptic segmentation metric that combines semantic classification quality and sequence-level instance association quality with a geometric mean.
Quick Facts
| Full Name | LiDAR Segmentation and Tracking Quality |
|---|---|
| Specification | Official Specification |
How It Works
Compute semantic quality independently of instances
S_cls is the mean Intersection over Union across evaluated semantic classes. It ignores instance identities, so a car point can receive semantic credit even if its track ID is fragmented. Stuff and thing classes enter according to the benchmark's class list and ignore policy.
The original 4D Panoptic LiDAR Segmentation paper deliberately separates semantic classification from association. This makes a low score diagnosable: inspect S_cls before changing the tracker.
Measure point-to-track association over the sequence
For a Ground Truth thing track g and predicted track p, let TPA = |g intersect p|, FNA = |g| - TPA, and FPA = |p| - TPA, counted in points across the sequence. Their Association IoU is TPA / (TPA + FNA + FPA). The score for g is sum_p TPA * AssociationIoU(g,p) / |g|; S_assoc averages this value over Ground Truth thing tracks.
Splitting one truth track across predictions lowers each pair's overlap, while merging truth tracks creates FPA for the competing pair. SemanticKITTI additionally requires instance IDs to be unique across classes because association is evaluated independently of semantic correctness.
Use the components and protocol, not the scalar alone
LSTQ = sqrt(S_cls * S_assoc) gives both components multiplicative influence and reaches zero when either is zero. It does not encode confidence ranking, real-time latency, geometric risk, calibration, or application-specific costs. Track-equal association can also make a small object influential despite far fewer points than a large region.
Report both components, per-class IoU, track-length and point-count slices, ignored classes, minimum points, sequence boundaries, and evaluator commit. Compare sPTQ when frame-level segment matching and switch-weighted IoU matter, or PAT when an instance-centric tracking term should be balanced with PQ.
Key Characteristics
- Combines semantic classification and temporal association
- Uses class mean IoU for the semantic component
- Uses point-level sequence intersections for association
- Averages association over ground-truth thing tracks
- Avoids thresholded whole-segment matching for association credit
- Depends on ID namespace, class, ignore, and instance-size rules
Common Use Cases
- Evaluating SemanticKITTI 4D panoptic segmentation
- Benchmarking LiDAR semantic and instance tracking jointly
- Diagnosing semantic errors separately from association errors
- Comparing long-sequence point-cloud track consistency
- Auditing fragmentation and merges at point level
Example
Loading code...Frequently Asked Questions
How is LSTQ calculated?
LSTQ is the square root of semantic classification quality times association quality. `S_cls` is class mean IoU; `S_assoc` averages point-weighted Association IoU contributions over ground-truth thing tracks.
What do S_cls and S_assoc measure?
`S_cls` measures whether LiDAR points have correct semantic labels without considering instances. `S_assoc` measures whether points from each ground-truth thing track are assigned consistently to predicted tracks without requiring semantic correctness.
Does LSTQ use a segment IoU matching threshold?
Its association component uses point intersections across complete sequence tracks rather than first accepting frame segments above a fixed IoU threshold. Dataset preprocessing and minimum-instance rules can still remove points or tracks.
How is LSTQ different from sPTQ?
LSTQ uses point-centric sequence association and geometric combination with semantic mIoU. sPTQ starts from frame-level PQ matching and removes the IoU credit of matched segments at identity-switch frames before class averaging.
Can LSTQ scores from different datasets be compared?
Not directly. Class taxonomies, thing and stuff definitions, ignored points, minimum instance size, ID namespace, sequence length, sampling, and evaluator implementation all change the evaluated problem.