Glossary · term

Groundedness

Groundedness is the degree to which the claims in a generated response are supported by a specified source context. In a retrieval-augmented system, that context is usually the retrieved passages supplied for the request. Evaluation may split a response into claims and check whether each is entailed or otherwise supported by those passages. The property is source-relative: a statement can be factually true yet ungrounded if the designated context does not verify it, and a well-grounded statement can repeat an error present in the source.

LLMOps2021-04-30Wave 2 · 2024Maturity: 3/5

Origin and context

Groundedness predates generative AI as a general idea connecting language to evidence or an environment. BEGIN benchmarked source attribution in knowledge-grounded dialogue in 2021. For the narrower RAG-evaluation scope, TruLens later grouped groundedness with context relevance and answer relevance in the RAG Triad. Microsoft documented a production metric that verifies response claims against user-provided context, while a 2024 NAACL Findings study examined support from retrieved documents or a model's pretraining corpus.

Sources: s1, s2, s3, s4

Why it matters

RAG can retrieve useful evidence without ensuring that a generator actually follows it. Measuring groundedness isolates that generation-stage failure from two different questions: whether retrieval found relevant material and whether the final answer addresses the user. Claim-level results also help reviewers locate unsupported passages instead of relying on a single impression of fluency. The metric is therefore useful for evaluation and debugging, but it does not by itself establish truth, relevance, completeness, or safety.

Sources: s1, s2, s3

Example

For a policy assistant, an evaluation set can store each question, the exact policy passages retrieved at that time, and the generated response. Reviewers or an evaluator split the response into material claims, mark which passage supports each claim, and record unsupported claims separately. The team reports both claim-level evidence and an aggregate score, calibrates an automated evaluator against human judgments, and repeats the test after changes to retrieval, prompts, models, or source documents. Correct citations are checked independently from mere source support.

Sources: s1, s2, s3

How it differs

AI hallucination

Hallucination is a broader and inconsistently defined failure family involving fabricated, unsupported, or incorrect output. Groundedness has a narrower test: whether claims are supported by a designated context. Microsoft explicitly notes that a factually correct answer can still be scored ungrounded when the supplied source does not verify it.

Retrieval-Augmented Generation

RAG is an architecture that retrieves context before or during generation. Groundedness is a property or evaluation dimension of the resulting answer. Adding retrieval can improve access to evidence, but it does not guarantee that retrieval is relevant, that the model uses it faithfully, or that the underlying documents are correct.

Maturity and evidence

Maturity is rated 3. TruLens, Microsoft, and academic researchers independently use a recognizable source-support construct, and claim-level groundedness is now a practical RAG evaluation dimension. It is not rated higher because evaluator prompts, score scales, context boundaries, aggregation rules, and relationships to faithfulness or factuality vary across implementations.

Sources: s1, s2, s3, s4

Limits and open questions

A groundedness score inherits the quality and completeness of the chosen context. If a source is false, stale, contradictory, or unauthorized, support from that source does not make the answer trustworthy. Automated judges can miss paraphrases, over-credit weak evidence, or vary with model and prompt; thresholds must be calibrated on representative human-labeled cases. Scores should preserve the evaluated context and evaluator version, and teams should assess citation attribution, factual accuracy, relevance, completeness, and retrieval quality separately. Groundedness is evidence about one relationship, not a certification of an answer or system.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as rag evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai grounding citations skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.