Glossary · term

Indirect prompt injection

A variant of prompt injection (Greshake et al. 2023) where malicious instructions are injected not by the user but through data the agent retrieves: a web page, an email, a PDF. The agent "reads" the instructions as though they came from the user. Especially dangerous for autonomous agents. It spurred the development of sandboxing, KYA, and Constitutional Classifiers.

Safety2024Wave 1 · 2023Maturity: 3/5

Maturity rationale

Indirect prompt injection — a central problem of agents

References

Author: Simon Willison