What is Answer Correctness?
Answer Correctness is an evaluation property that measures whether the factual claims or task results in a generated answer agree with a declared trusted reference, executable oracle, or adjudicated ground truth.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Choose an oracle that matches the task
Use an executable oracle for code, calculations, database state, or structured business rules when one exists. Use a versioned reference answer or claim set for open-ended factual questions, and preserve acceptable alternatives, units, qualifiers, and validity dates. A single reference phrasing is not the truth boundary: semantic equivalents should match, while contradictions and unsupported additions remain separate errors.
Make the denominator inspectable
The current Ragas Factual Correctness protocol derives response and reference claims, then reports precision, recall, or F1. RAGChecker likewise uses claim-level entailment to separate overall accuracy from retrieval and generator failures. Their implementations are not interchangeable: record the claim extractor, matching rubric, judge, aggregation mode, empty-answer policy, and reference revision.
Use errors to diagnose the pipeline
A false-positive claim can originate from model memory, a noisy retrieved chunk, an ambiguous question, or a stale reference. A false negative can come from missing retrieval, incomplete synthesis, an over-strict refusal, or a reference that demands irrelevant detail. Preserve claim-level labels and source IDs so a score change leads to a specific retrieval, generation, data, or evaluator fix.
Key Characteristics
- Compares generated claims or task results with a declared oracle
- Separates added incorrect claims from omitted required claims
- Can report precision, recall, F1, exact match, or task-specific success
- Depends on reference quality, claim granularity, and evaluator version
- Does not prove that the answer used the supplied context
- Supports release decisions only when slice and uncertainty are retained
Common Use Cases
- Evaluating factual answers against a versioned reference claim set
- Checking calculations, code, SQL, or structured outputs with executable oracles
- Distinguishing incorrect additions from incomplete answers
- Comparing RAG releases on paired questions and reference revisions
- Building per-slice release gates for high-impact factual tasks
Example
Loading code...Frequently Asked Questions
Is Answer Correctness the same as Answer Faithfulness?
No. Correctness compares the response with a trusted reference or oracle. Faithfulness checks whether response claims follow from the supplied context. A response can be faithful to an incorrect source or correct despite ignoring the provided evidence.
How is Answer Correctness calculated?
One common protocol extracts atomic claims from the response and reference, matches equivalent claims, and computes precision, recall, or F1 from true positives, false positives, and false negatives. Other tasks may use exact match, execution, or a documented rubric.
Should Answer Correctness use precision, recall, or F1?
Choose from the error cost. Precision emphasizes avoiding incorrect additions, recall emphasizes covering required facts, and F1 balances both. Report the components because equal F1 values can hide very different omission and fabrication risks.
Can an LLM judge measure Answer Correctness reliably?
It can estimate semantic claim agreement, but the result depends on decomposition, rubric, prompt, model, and reference quality. Calibrate against representative human labels, pin the evaluator version, and use deterministic oracles whenever the task permits them.
What if the reference answer is incomplete or outdated?
Mark the case invalid or revise the reference before using it for a release decision. Preserve source revision and validity, allow documented alternative answers, and review disagreements instead of treating every difference from one reference string as an error.