Glossary · term
Indirect prompt injection
A variant of prompt injection (Greshake et al. 2023) where malicious instructions are injected not by the user but through data the agent retrieves: a web page, an email, a PDF. The agent "reads" the instructions as though they came from the user. Especially dangerous for autonomous agents. It spurred the development of sandboxing, KYA, and Constitutional Classifiers.
Safety2024Wave 1 · 2023Maturity: 3/5
Maturity rationale
Indirect prompt injection — a central problem of agents
References
Author: Simon Willison