What is Noise Sensitivity?
Noise Sensitivity is a RAG evaluation metric that measures how often retrieved content induces incorrect answer claims, with protocols commonly separating noise inside relevant chunks from noise in irrelevant chunks.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Separate relevant and irrelevant noise
The current Ragas protocol offers relevant and irrelevant modes. Relevant noise is an incorrect response claim attributable to a chunk that also contains useful evidence; irrelevant noise is an incorrect claim attributable to a chunk judged irrelevant to the reference. The labels are not properties of text alone: they depend on the question, reference, evidence granularity, and entailment rubric.
Use paired perturbations for causal evidence
A static attribution score shows association, not that a distractor caused the error. Run paired cases with the same question, reference, model, prompt, and required evidence while adding, removing, or relocating one controlled distractor. Compare claim errors, refusal behavior, citation support, and latency. A change is more actionable when the added error appears only in the noisy condition and can be traced to the injected passage.
Diagnose retrieval and generation separately
RAGChecker separates incorrect claims attributable to relevant chunks, irrelevant chunks, and no retrieved chunk. This distinguishes a generator that blindly trusts context from one that hallucinates beyond it. Test stale, contradictory, near-duplicate, authority-mismatched, multilingual, and permission-filtered distractors instead of relying on random unrelated paragraphs alone.
Key Characteristics
- Counts incorrect response claims linked to retrieved context
- Separates noise embedded in relevant chunks from irrelevant chunks
- Uses a trusted reference and an explicit claim-attribution protocol
- Can be strengthened with paired clean-versus-noisy perturbations
- Differs from retrieval precision because a generator may ignore noise
- Should be reported with correctness, faithfulness, refusal, and citation metrics
Common Use Cases
- Testing whether irrelevant retrieved passages change a correct answer
- Finding misleading statements inside otherwise relevant documents
- Comparing reranking, deduplication, filtering, and context-ordering changes
- Building regression suites from known noisy-context incidents
- Setting release limits for incorrect claims induced by untrusted evidence
Example
Loading code...Frequently Asked Questions
Is Noise Sensitivity the same as Context Precision?
No. Context Precision evaluates whether relevant retrieval results are ranked ahead of irrelevant ones. Noise Sensitivity evaluates whether retrieved content induces incorrect response claims. A generator may ignore a noisy retrieval list, or copy a misleading sentence from a relevant chunk.
Is a lower Noise Sensitivity score better?
Yes under the common claim-error protocol, because the score is the fraction of response claims that are incorrect and attributable to the selected noise class. Always record the implementation and mode, since custom robustness tests may use different directions or denominators.
What is relevant noise in RAG?
Relevant noise is misleading or incorrect content inside a chunk that also contains evidence useful to the question. Labeling an entire chunk as relevant does not make every sentence within it trustworthy.
How can teams test Noise Sensitivity?
Run paired clean and noisy cases while holding the question, required evidence, prompt, model, and decoding settings fixed. Add one controlled distractor, then compare claim correctness, citation support, refusal behavior, and latency.
Does low Noise Sensitivity prove a RAG system is good?
No. The retriever may still miss required evidence, the answer may be irrelevant or incomplete, or the model may refuse everything. Pair the metric with context recall and precision, correctness, faithfulness, refusal correctness, citation quality, and task success.