What Is MCP Tool Poisoning and Prompt Injection, and How Do You Defend Against It?
MCP tool poisoning is an attack where a malicious MCP server hides instructions in a tool description. The model reads the full description and may follow it, while the user usually sees only a short summary. It is one form of prompt injection, where text a model reads is treated as a command. Defenses are layered: vet and pin servers, give agents only the tools they need, require approval for risky actions, and inspect and log every call. Metorial applies several of these through Protoguard, tool filters, version pinning, and tracing.
An MCP tool description is text written by whoever runs the server, and the model reads all of it. If that text contains hidden instructions, the model may follow them, and the person approving the tool often never sees them.
What is prompt injection?
OWASP (the Open Worldwide Application Security Project) defines a prompt injection vulnerability as one where inputs alter a model's behavior in unintended ways. It splits into direct injection, where the user's own prompt does it, and indirect injection, where the model reads an external source such as a web page or file that carries instructions.
The core problem is that a model reads instructions and data as one stream of text. Without tools, the worst outcome is a bad answer. With tools and real credentials attached, the model can act on the injected text.
What is MCP tool poisoning?
Invariant Labs described it in April 2025. A malicious server puts hidden instructions inside a tool description. The model sees the complete description, while the client interface typically shows the user a simplified version.
Their example was a tool named add that appeared to add two numbers. Its description told the model to read the user's MCP configuration file and an SSH (Secure Shell) private key, and pass the contents along as a parameter. The model, running in Cursor, complied.
The specification already treats this as a risk: it says clients must consider tool annotations untrusted unless they come from a trusted server.
Why does MCP make this easier to attempt?
An assistant can connect to servers from many different authors, and it loads each server's tool descriptions into the model's context automatically so the model can choose among them. Anyone who publishes a server therefore gets to write text that the model reads before every decision. The specification says tools represent arbitrary code execution and must be treated with caution, which is why a description from an unknown author deserves the same suspicion as a script from an unknown author.
The same applies to tool results. A server that reads a ticket or an email returns that content to the model, and any instructions inside it arrive through the same channel.
What are rug pulls and tool shadowing?
Invariant Labs named two escalations. A rug pull is a server that changes its tool descriptions after you approve it, so a tool that was safe when you approved it is not safe later. Shadowing is a malicious server whose description changes how the model uses a trusted tool from a different server, without the user ever calling the malicious tool.
What is the lethal trifecta?
In June 2025, Simon Willison named three capabilities that become dangerous together: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent with all three can be told by a poisoned page to read a private file and send it out. Meta later framed the same idea as the Agents Rule of Two: an agent should satisfy no more than two of the three within a session.
How do you defend against it?
No single control stops it, so stack them.
Vet servers before connecting them and pin the version you reviewed. Read the tool descriptions yourself.
Give each agent only the tools it needs. A read-only agent cannot be talked into deleting anything. This is OWASP's least-privilege guidance.
Require human approval for actions that change or send data. The specification says a human should be able to deny tool calls.
Break the trifecta where you can, for example by keeping agents that read untrusted content away from agents that hold private data.
Inspect and record every call. You cannot investigate what you did not log, and a flagged call can be blocked before it runs. Filtering lowers the risk but does not remove it, so it never replaces the other controls. For the wider picture, see Are MCP servers secure?
What does it look like in Metorial?
Take a team that connects a third-party provider to a shared assistant. In Metorial, an admin can limit which of its tools are exposed, for instance read-only ones, using tool filters, so a poisoned description has no write action to trigger. See restricting agents to read-only tools.
Protoguard reviews messages and tool requests before an agent acts and checks them for prompt injection. It also monitors schema changes: if a provider changes its tools, the change is flagged, and version pinning keeps the integration on the version you reviewed. Tracing records each session with the tool, arguments, and result.
Protoguard reduces prompt injection risk rather than eliminating it.
Frequently asked questions
Is tool poisoning the same as prompt injection?
Tool poisoning is a kind of prompt injection. Prompt injection is the general problem of a model treating text it reads as instructions. Tool poisoning is the case where that text sits in a tool description supplied by an MCP server.
What is a rug pull in MCP?
A rug pull is when a server changes a tool description or behavior after the user has approved it. The tool looked safe at approval time and is not safe later. Pinning a version and alerting on tool changes are the usual defenses.
What is the lethal trifecta?
A term from Simon Willison for an agent that combines access to private data, exposure to untrusted content, and a way to communicate externally. With all three, one poisoned piece of text can lead to data theft. Removing any one of them reduces the risk.
Can a filter or scanner fully stop prompt injection?
No filter is a guarantee. Scanning calls for injection lowers risk and gives you alerts, but it should sit alongside narrow tool access, human approval for consequential actions, and logs. Treat every tool description and tool result as untrusted input.
Are first-party MCP servers safe from this?
They are lower risk, not risk free. A server you control cannot be poisoned by its author, but a tool result can still carry injected text from an email, issue, or web page it reads. Third-party servers add the risk of a hostile or compromised description.