A side-by-side comparison of Unsupported-claim Rate and Hallucination. Understand how a measurement of unsupported claims relates to the broader failure of false or fabricated AI output.
Veredicto rápido: Use Unsupported-claim Rate when quantifying unsupported assertions; use Hallucination when describing the underlying generative AI failure mode.
Unsupported-claim Rate describes percentage of claims in a model response that are not grounded in supporting evidence.
Contexto: Most relevant for evaluation reports that need a measurable rate of unsupported claims.
Hallucination describes AI-generated output that appears plausible or confident but is false, unsupported, misleading, or fabricated.
Contexto: Most relevant when describing a generative AI reliability or safety failure.
| Aspecto | Unsupported-claim Rate | Hallucination |
|---|---|---|
| What it measures | Unsupported-claim rate measures the percentage of claims in a response that lack supporting evidence. | Hallucination describes a type of output that is false, unsupported, misleading, or fabricated. |
| Best use case | Use UCR to compare systems, prompts, or retrieval settings using a defined evaluation rubric. | Use hallucination to describe the failure mode observed in a model response or incident. |
| Failure mode | A high UCR indicates that many response claims are not grounded in evidence. | A hallucination may be one unsupported claim, an invented citation, or a broader fabricated explanation. |
| Threshold sensitivity | UCR depends on how claims are segmented, what counts as evidence, and which threshold is used for support. | Hallucination assessment depends on factual verification, grounding checks, and the task context. |
| Common mistake | A common mistake is treating a low UCR as proof that every answer is correct. | A common mistake is using hallucination as a vague label without recording which claims were false or unsupported. |
In practice, hallucination is the problem category, while unsupported-claim rate is one way to measure and govern part of that problem.
Reporting hallucination risk without a measurable evaluation method.
Counting unsupported claims without defining what qualifies as a claim or supporting evidence.
Assuming citations are valid without checking whether they support the statement.
Using aggregate UCR to hide severe individual failures.
Use Unsupported-claim Rate when evaluating responses against source material or a grounding standard. It is appropriate for dashboards, regression tests, and evidence packs where unsupported assertions need to be counted.
Use Hallucination when describing the substantive error pattern in generative AI output. It is the right term for user-facing risks, incident descriptions, and reliability controls.
For ISO/IEC 42001 and NIST AI RMF evidence, unsupported-claim metrics can support monitoring and validation of hallucination risk, especially where AI outputs inform decisions or external communications.
No. Unsupported-claim rate measures unsupported claims in a response, while hallucination is a broader failure category that may include false facts, fabricated sources, and misleading reasoning.
Yes. A claim may cite a source but still misread it or rely on unreliable evidence. Grounding and factuality should be evaluated separately where the use case requires it.
It turns an abstract hallucination risk into a repeatable measurement. That makes it easier to track changes across prompts, models, retrieval settings, and releases.
No recently viewed comparisons yet.