Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) is a pattern in which a generative model receives evidence retrieved from an external collection at query time and uses that evidence while producing an answer. The original formulation combined a pretrained sequence-to-sequence model's parametric memory with a dense vector index of Wikipedia as non-parametric memory. Modern systems vary in how they index, retrieve, rerank, assemble, and cite evidence.
Origin and context
The 2020 paper framed RAG as a way to improve knowledge-intensive language tasks and make factual knowledge easier to update or inspect than knowledge stored only in model parameters. Later research broadened the label beyond one architecture. A 2023 survey distinguishes naive, advanced, and modular RAG and organizes the field around retrieval, generation, augmentation, and evaluation choices.
Why it matters
RAG lets an application draw on private, recent, or domain-specific material without retraining the base model for every document change. Retrieved passages can also provide an evidence trail for users and evaluators. Its practical value depends on the whole pipeline: collection quality, chunking, indexing, query construction, retrieval recall, ranking, context assembly, and answer behavior. RAG is therefore an application architecture, not a guarantee that an answer is current or correct.
Example
An internal support assistant can index approved product manuals and incident runbooks. When an engineer asks about an error code, the system retrieves the most relevant passages, places them in the model context, and asks for an answer with citations. A robust implementation also checks access permissions, records which passages were used, and declines when retrieval returns weak or conflicting evidence.
How it differs
GraphRAG
GraphRAG is a specialized family of retrieval-augmented approaches that derives graph structure and summaries to answer relationship-heavy or corpus-wide questions. Ordinary RAG can use flat text chunks and does not require a knowledge graph.
Long-context language models
Long-context models increase how much material can be supplied in one request. RAG selects a subset before generation. The two can be combined; a larger context window does not by itself decide which evidence is relevant, current, or authorized.
Maturity and evidence
Maturity is rated 4 because the pattern has a peer-reviewed origin, a substantial research literature, multiple architectural variants, and an established evaluation vocabulary. The rating applies to RAG as a broad pattern, not to the quality of any particular retriever or deployment.
Limits and open questions
Retrieval can miss decisive evidence, surface stale or adversarial text, or return passages that look similar but do not answer the question. Generation can ignore, distort, or overgeneralize retrieved material. Chunk boundaries may destroy context, while aggressive retrieval increases latency and context cost. Teams need retrieval and answer-level evaluation, provenance, permission filtering, update processes, and defenses against instructions embedded in untrusted documents.
Sources: s2
Related terms
References
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksFacebook AI Research, UCL, and NYU / NeurIPS · 2020-05-22 · class A
- Retrieval-Augmented Generation for Large Language Models: A SurveyIndependent academic collaboration / arXiv · 2023-12-18 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as retrieval augmented generation skill.
This term is also covered in the Skills Atlas as information retrieval skill.
This term is also covered in the Skills Atlas as rag evaluation skill.