Glossary · term

AI hallucination

An AI hallucination is generated content that is false, erroneous, contradictory, or unsupported by the relevant source or prompt, yet may be presented fluently and confidently. The boundary depends on the task: a statement can be factually true but still unfaithful to a supplied document. NIST uses “confabulation” as a formal risk label and describes “hallucination” and “fabrication” as colloquial alternatives.

Culture2015-05-21Wave 1 · 2023Maturity: 4/5

Origin and context

In May 2015, Andrej Karpathy described a character-level RNN generating a plausible but nonexistent URL and wrote that the model had “hallucinated” it. This is the earliest direct generative-model usage in the reviewed evidence, not a claim that he invented every earlier AI use of the metaphor. In September 2018, Rohrbach and colleagues supplied a peer-reviewed image-captioning milestone by studying captions that mention objects absent from the image under the name object hallucination. A 2022 survey then organized work across summarization, dialogue, question answering, data-to-text, translation, and visual-language generation; ACM published the reviewed survey in March 2023. In July 2024, NIST categorized confidently stated false content as confabulation and connected it to the colloquial term hallucination.

Sources: s5, s4, s1, s2, s3

Why it matters

Fluency can make unsupported output appear more reliable than it is. A fabricated citation, incorrect policy summary, or invented product fact can mislead a user and contaminate downstream decisions or automated actions. The risk rises when people over-rely on a system or when an agent passes generated claims to tools without verification. Managing hallucination therefore requires task-specific evaluation, provenance and grounding where appropriate, review of cited sources, and escalation for high-impact decisions. It cannot be reduced to a single universal benchmark score.

Sources: s1, s2, s3

Example

A research assistant is asked for a paper supporting a claim and returns a plausible title, author list, and DOI that do not exist. The answer is a hallucination because its central evidence is fabricated, even though its format is convincing. A safer workflow searches an authoritative index, opens the cited record, and reports uncertainty when no match is found. Retrieval can reduce unsupported generation by supplying evidence, but it does not guarantee correctness if retrieval fails or the model misreads the source.

Sources: s1, s2, s3

How it differs

Confabulation

In current generative-AI risk guidance, confabulation and hallucination often refer to the same family of failures. NIST prefers confabulation for confidently presented false or erroneous content and calls hallucination colloquial. This glossary retains hallucination as the canonical public-facing entry because it is the established search term, while treating confabulation as an alias or reference rather than a separate technical mechanism.

Maturity and evidence

Maturity is rated 4. Hallucination has a peer-reviewed cross-task survey, extensive measurement and mitigation literature, and explicit treatment in NIST risk guidance. The concept is established across research and operations. It remains below 5 because definitions and metrics vary by task, factuality and source faithfulness are not identical, and no mitigation reliably eliminates unsupported generation across models and deployment contexts.

Sources: s4, s1, s2, s3

Limits and open questions

The term can blur different failure modes: contradiction, unsupported detail, stale knowledge, retrieval error, or an intentionally creative response. Automatic detectors may disagree with human reviewers and can miss domain-specific errors. Teams should define what counts as unsupported for the application, test on representative cases, preserve source evidence, and avoid implying that a single confidence score proves factuality.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as hallucination detection skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai output verification skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai grounding citations skill.