What is J&F Score?

J&F Score is the arithmetic mean of DAVIS region similarity J and contour accuracy F, combining mask-area overlap with boundary alignment for video object segmentation.

Quick Facts

SpecificationOfficial Specification

How It Works

Measure region overlap and boundary alignment separately

J = |P intersection G| / |P union G| compares predicted mask P with Ground Truth mask G. Boundary F = 2PR / (P + R) computes precision from predicted contour pixels near a truth contour and recall from truth contour pixels near a predicted contour.

The official DAVIS evaluator uses a boundary tolerance equal to a configured fraction of the image diagonal, 0.008 by default, rounded up to pixels. Resolution and tolerance must therefore remain part of a reproducible result.

Aggregate by frame, object, and metric before combining

DAVIS computes per-frame J and F for each evaluated object, then derives each object's mean. The benchmark averages object-level J means and F means and reports J&F = (J_mean + F_mean) / 2. It also exposes Recall, the fraction of valid frames scoring above 0.5, and Decay, the mean first-quartile score minus the mean last-quartile score.

The official semi-supervised protocol supplies each object mask in its first frame and evaluates subsequent propagation. Averaging every pixel globally or averaging sequence summaries with different weights does not reproduce this contract.

Distinguish semi-supervised and unsupervised matching

In semi-supervised DAVIS, the target identities are given by first-frame masks. In the unsupervised DAVIS 2019 task, the evaluator performs bipartite matching between Ground Truth objects and predicted video-object proposals to maximize their mean J&F; extra proposals beyond matched annotated objects are not penalized under that challenge rule.

This means equal J&F values from different tasks can represent different failure sets. Pair J&F with per-object scores, J/F Recall and Decay, runtime, memory, and failure slices. For systems that must discover, classify, and maintain many objects, use MOTSA, HOTA, Track mAP, or TETA under the appropriate annotation protocol.

Key Characteristics

  • Combines region Jaccard overlap with contour F-measure
  • Uses tolerance-aware boundary matching
  • Averages frame scores through object-level summaries
  • Reports J and F means before their arithmetic mean
  • Supports Recall and temporal Decay diagnostics
  • Changes meaning across semi-supervised and unsupervised protocols

Common Use Cases

  1. Evaluating semi-supervised video object segmentation
  2. Comparing prompt-guided or first-frame mask propagation
  3. Separating mask-area errors from contour errors
  4. Diagnosing temporal degradation with J and F Decay
  5. Reproducing DAVIS challenge evaluation results

Example

loading...
Loading code...

Frequently Asked Questions

How is the J&F Score calculated?

DAVIS computes per-frame region Jaccard J and boundary F for each object, summarizes each object, averages J and F over objects, and reports their arithmetic mean: `J&F = (J_mean + F_mean) / 2`.

What is the difference between J and F?

J measures foreground-region overlap through mask IoU. F measures contour alignment through tolerance-aware boundary precision and recall. A mask can cover most of an object while still receiving a weaker F score for rough or displaced edges.

What do J Recall, F Recall, and Decay mean?

Recall is the fraction of valid frames whose per-frame metric exceeds `0.5`. Decay is the mean score in the first temporal quartile minus the mean in the last quartile, so positive decay indicates deterioration over the sequence.

Does DAVIS use an exact boundary-pixel match?

No. The official implementation dilates contours by a tolerance derived from the image diagonal, `0.008` by default, before computing boundary precision and recall. The tolerance and resolution affect F.

Can J&F be compared between DAVIS tasks?

Only with the same task and evaluator. Semi-supervised evaluation follows provided object identities, while unsupervised DAVIS matches proposals to annotated objects and may not penalize extra proposals. Those protocols expose different errors.

Related Terms

Related Articles