Glossary · term

Prompt injection

Prompt injection is a vulnerability in an application built around an instruction-following model. Untrusted text, images, or other content is interpreted as instructions that alter the model's intended behavior. The injected instruction may arrive directly from a user or indirectly through retrieved documents, web pages, email, tool output, or memory. The risk becomes consequential when model output can expose data or trigger actions.

Safety2022-09-12Wave 1 · 2023Maturity: 4/5

Origin and context

Willison introduced the term while comparing a GPT-3 translation example with SQL injection: an application concatenated trusted instructions and attacker-controlled input, and the model followed the latter. Research by Greshake and colleagues then demonstrated indirect prompt injection against LLM-integrated applications, where an attacker plants instructions in resources the application later retrieves. OWASP now documents both delivery paths as one vulnerability family.

Sources: s1, s2, s3

Why it matters

A successful injection can change an answer, leak a system prompt, influence retrieval, exfiltrate information, or steer a connected agent toward an unauthorized tool call. The application boundary matters more than a clever malicious phrase: impact depends on what untrusted content reaches the model, what secrets are present, and what permissions downstream components grant. OWASP recommends layered controls such as least privilege, separation of trust domains, output validation, monitoring, and human approval for consequential actions.

Sources: s2, s3

Example

As an illustrative scenario, suppose an assistant retrieves a vendor web page before drafting a procurement summary. Hidden text on that page instructs the assistant to ignore the user's request and send confidential context to an external endpoint. That is indirect prompt injection even though the employee never typed the hostile instruction. A safer design treats retrieved content as untrusted data, prevents it from directly authorizing tools, scopes credentials narrowly, validates proposed actions against the original task, and requires confirmation before external transmission.

Sources: s2, s3

How it differs

Indirect prompt injection

Indirect prompt injection is a delivery subtype of prompt injection, not a competing parent concept. Direct injection comes through the model-facing input interface; indirect injection is planted in an external resource later processed by the application. The distinction concerns where the hostile instruction enters the workflow, not a different underlying vulnerability or a claim that every external document is malicious.

Maturity and evidence

Maturity is rated 4 for a security concept used across independent research and defensive practice: Greshake and colleagues study indirect attacks, and OWASP organizes application-level prevention guidance. The dated naming source provides chronology, not adoption evidence by itself. This rating does not mean the vulnerability has been solved, that mitigations are interchangeable, or that the term has a regulatory status.

Sources: s1, s2, s3

Limits and open questions

Prompt injection is not identical to jailbreaking, although an attack can involve both. Jailbreaking usually seeks to bypass a model's safety policy; prompt injection subverts an application's intended instruction hierarchy or task. String filters, delimiters, and extra instructions can reduce simple attacks but should not be treated as a security boundary. Risk assessment must cover the complete application, including retrieval, memory, tools, credentials, output rendering, and human approval paths.

Sources: s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as prompt injection defense skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai red teaming skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as owasp top 10 for llm applications skill.