What is Precision-Recall Curve?

Precision-Recall Curve is a threshold-sweep plot that places Recall on the horizontal axis and Precision on the vertical axis to show the positive-class coverage and reliability tradeoff of a scoring classifier.

Quick Facts

SpecificationOfficial Specification

How It Works

Group tied scores before creating operating points

Examples with identical scores cannot be ordered by the model, so they should enter the predicted-positive set together. Arbitrarily ordering ties can invent intermediate points and alter a finite-sample summary. Current scikit-learn documentation returns points at distinct thresholds plus an endpoint that completes the plot; threshold and endpoint conventions should be recorded.

Distinguish Average Precision from trapezoidal PR-AUC

Average Precision (AP) weights Precision at each operating point by the increase in Recall: sum((R_n-R_(n-1))*P_n). Current scikit-learn documentation explicitly distinguishes this non-interpolated summary from trapezoidal area, whose linear interpolation can be optimistic. Never publish PR-AUC without naming the integration method.

Interpret rare-event performance on the deployment population

Davis and Goadrich establish the relationship between ROC and PR spaces and show that ROC dominance corresponds to PR dominance on the same dataset, while area optimization and interpolation differ. Saito and Rehmsmeier demonstrate why PR plots can expose poor positive-prediction reliability hidden by ROC views on heavily imbalanced data. These findings do not make one universal metric sufficient.

Key Characteristics

  • Plots Precision against Recall while sweeping distinct score thresholds
  • Focuses evaluation on the declared positive class and its predictions
  • Has a prevalence-dependent no-skill Precision baseline
  • Requires tied scores to be handled as one indistinguishable group
  • Can be summarized by non-interpolated Average Precision or another declared area rule
  • Does not choose a threshold, encode costs, or measure probability calibration

Common Use Cases

  1. Comparing rare-event detection rankings on the same evaluation population
  2. Selecting thresholds under alert-quality and coverage constraints
  3. Inspecting fraud, safety, retrieval, or diagnosis score tradeoffs
  4. Reporting Average Precision with an explicit integration convention
  5. Monitoring ranking changes alongside workload and calibrated probabilities

Example

loading...
Loading code...

Frequently Asked Questions

How is a Precision-Recall Curve created?

Sort examples by the score for the declared positive class, group identical scores, and lower the threshold across each distinct group. At every operating point compute Precision from predicted positives and Recall from actual positives, retaining the threshold and support.

When is a PR Curve more useful than a ROC curve?

It is often more revealing when positives are rare and the reliability of positive alerts matters. ROC false-positive rate can remain small because its denominator contains many negatives, while Precision exposes how many produced alerts are false. Report both when they answer relevant questions.

Is Average Precision the same as PR-AUC?

Not automatically. Average Precision usually weights each observed Precision by the increase in Recall without linear interpolation. Trapezoidal PR-AUC assumes straight lines between points and can produce a different value. Name the formula and library implementation.

What is the baseline for a Precision-Recall Curve?

For an uninformative random ranking, expected Precision is the positive-class prevalence. Because that baseline changes with the population, PR curves and their summaries should be compared on the same sampling and weighting scheme.

Does the best point on a PR Curve define the production threshold?

No. Choose a threshold on validation data using false-positive and false-negative costs, review capacity, minimum coverage, uncertainty, and policy constraints. Then evaluate that frozen operating point on untouched data and monitor it after deployment.

Related Terms

Related Articles