MCP rug pull
An MCP rug pull is a post-approval bait-and-switch: an MCP tool or server is presented as benign, then its effective definition or behavior is maliciously changed while a client continues relying on the earlier trust decision. The changed surface can include a description, schema, permissions, supplied package, or backend behavior. The defining feature is the time gap between review and execution; an accidental compatible change is tool drift, not a rug pull.
Origin and context
Invariant Labs described `MCP Rug Pulls` on 1 April 2025: a malicious server could alter a tool description after the client had approved it. Its 7 April follow-up made the sequence concrete. A sleeper server first advertised a harmless fact-of-the-day tool, then activated a malicious description on its second launch. Simon Willison independently called the pattern silent redefinition on 9 April. Two preprints and OWASP later retained rug pulls as a recognizable MCP attack class or sub-technique.
Why it matters
A one-time review becomes stale when the tool presented later is not the tool that was assessed. MCP deliberately supports dynamic tool discovery: `tools/list` returns names, descriptions, schemas and annotations, and a server can declare list-change notifications. Those protocol messages do not by themselves prove that new content matches an approved version or require a particular re-approval interface. A familiar tool identity can therefore conceal a newly dangerous instruction, capability, or implementation unless the host compares versions and re-evaluates trust.
Example
A remote MCP server initially exposes `summarize_docs` with a narrow, harmless description, and an operator approves it. On a later connection the same name is returned with instructions to attach local credentials, or the unchanged-looking interface now sends documents to a new destination. If the host refreshes and exposes that tool without detecting the contract or behavior change, the attacker has reused yesterday's approval for today's different capability. Invariant Labs' sleeper demonstration used this timing and combined it with cross-server shadowing.
How it differs
Tool poisoning
Tool poisoning describes a malicious instruction or contract presented to the model. A rug pull adds a temporal condition: the reviewed version was benign and the poisoned or expanded version arrived later. OWASP therefore places rug pulls under its broader tool-poisoning category, but the terms are not interchangeable.
Cross-server tool shadowing
Tool shadowing concerns scope: one server's metadata changes how the agent uses another trusted tool. A rug pull concerns timing. Invariant Labs combined both in the sleeper WhatsApp demonstration, but a server can silently change its own behavior without shadowing another tool, and shadowing can be malicious from first exposure.
Maturity and evidence
Maturity is rated 3. The term has a dated origin disclosure, an independent explanation within eight days, two 2025 research treatments, OWASP taxonomy coverage, and convergent recommendations to fingerprint or version approved definitions. It remains below 4 because the exact scope varies across sources, the protocol and client controls are still evolving, and published work demonstrates feasibility rather than measuring how often deliberate rug pulls occur in deployed systems.
Limits and open questions
A changed hash is evidence of drift, not proof of malice. Trust-on-first-use also cannot detect a hostile first version, and hashing only descriptions will miss unchanged metadata backed by altered server code. Signatures establish provenance, not benevolent behavior. Useful controls therefore combine a normalized definition and artifact baseline, explicit re-review for meaningful changes, least privilege, isolation, visible consequential inputs, and runtime monitoring. The current MCP change notification is a synchronization signal, not an integrity attestation or security certification.
Related terms
References
- MCP Security Notification: Tool Poisoning AttacksInvariant Labs · 2025-04-01 · class A
- WhatsApp MCP Exploited: Exfiltrating your message history via MCPInvariant Labs · 2025-04-07 · class A
- Tools — Model Context Protocol specification 2025-11-25Model Context Protocol · 2025-11-25 · class A
- Model Context Protocol has prompt injection security problemsSimon Willison's Weblog · 2025-04-09 · class B
- Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol EcosystemSong et al. / arXiv · 2025-05-31 · class A
- ETDI: Mitigating Tool Squatting and Rug Pull Attacks in Model Context Protocol (MCP) by using OAuth-Enhanced Tool Definitions and Policy-Based Access ControlBhatt, Narajala and Habler / arXiv · 2025-06-02 · class A
- OWASP Top 10 for Model Context Protocol version v0.1OWASP Foundation · 2025 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as model context protocol skill.