Agent protocol exploits
Agent protocol exploits are attacks whose exploitable path manifests in structured exchanges among an AI agent, tool server, peer agent or user-interface bridge. They abuse metadata, discovery, identity or authorization, message sequencing, context propagation or lifecycle events to cause unauthorized behavior. The label is an umbrella, not one exploit: every finding still needs a narrower mechanism, affected protocol and trust boundary.
Origin and context
Ferrag and colleagues placed `protocol exploits` in a June 2025 paper title, although their four-domain taxonomy calls the relevant category `Protocol Vulnerabilities`. CMU later used `communication protocol exploits`; OWASP formalized insecure inter-agent communication and protocol abuse; and OATF defines executable `agent-protocol attacks`. That convergence supports the concept, but not a single settled label. The paper's 30-plus catalog covers all four threat domains, not 30-plus protocol exploits.
Why it matters
Agent protocols turn messages into discovery, delegation, tool use and other consequential actions while moving data across trust boundaries. Layering is important: content filtering does not correct a token-audience error, TLS does not establish application authorization, and a signed message can still carry unsafe semantics. A useful threat model therefore identifies the sender, receiver, message or lifecycle event, trust transition, granted capability and resulting action rather than treating every failure as prompt injection.
Example
Suppose a travel agent discovers a remote booking agent through A2A and then calls a local MCP payment tool. An attacker supplies a forged discovery descriptor, embeds an instruction in returned content and reuses a token accepted for the wrong audience. The chain should be decomposed into descriptor or communication abuse, a semantic prompt payload, authorization failure and unsafe tool action. `Agent protocol exploit` can summarize the chain, but should not replace those specific findings.
How it differs
Prompt injection
Prompt injection manipulates model instructions or interpreted content. It may travel inside a protocol message, but protocol exploitation also covers discovery, identity, authorization, routing and lifecycle failures that require no injected prompt.
Model Context Protocol
MCP is one protocol whose implementations and deployments expose security-relevant boundaries. It is neither an exploit nor evidence that every MCP vulnerability is a defect in the protocol design.
MCP rug pull
An MCP rug pull is a narrower temporal attack in which previously trusted tool behavior or metadata changes. It can sit under the umbrella, but is not synonymous with all agent protocol exploits.
Maturity and evidence
Maturity is rated 3. A research survey, an independent CMU systematization, OWASP taxonomy, protocol-specific engineering analyses and OATF's operational format now describe a recognizable protocol attack surface. Maturity 4 would overstate the evidence: labels and boundaries still vary, OATF remains version 0.1 with provisional protocol bindings, and the reviewed sources do not measure deployment prevalence or validate a stable control baseline.
Limits and open questions
This entry is not a vulnerability identifier, certification, prevalence estimate or claim that agent protocols are inherently unsafe. Report concrete weaknesses at their narrowest useful level and state the protocol version and deployment assumptions. Keep ordinary software flaws, dependency compromise and network attacks outside the category unless the exploit actually depends on agent-protocol messages or semantics. Mitigations are protocol- and implementation-specific; authentication, encryption, schema validation and content controls address different layers and none is universal protection.
Related terms
References
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents WorkflowsFerrag et al. / arXiv · 2025-06-29 · class A
- Open Agent Threat Format, Specification v0.1Open Agent Threat Format · 2026 · class A
- SoK: Bridging Research and Practice in LLM Agent SecurityCarnegie Mellon University Software Engineering Institute · 2025-11 · class B
- OWASP Top 10 for Agentic Applications 2026OWASP GenAI Security Project · 2025-12-09 · class A
- MCP Tools: Attack and Defense RecommendationsElastic Security Labs · 2025-09-19 · class B
- Authorization — Model Context Protocol Specification 2025-06-18Model Context Protocol · 2025-06-18 · class A
- A Security Engineer's Guide to the A2A ProtocolSemgrep · 2025-12-17 · class B
- RFC 9113: HTTP/2 — Cross-Protocol AttacksRFC Editor / IETF · 2022-06 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as model context protocol skill.