Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

MCP tool poisoning is an attack in which instructions hidden in an MCP server’s tool descriptions, parameter schemas or returned output steer the AI model that reads them. Code review of your own application cannot see it, because the text arrives at runtime from a server you connected, not from code your team wrote. How much harm follows depends on what the agent can reach and whether a person sees and approves each consequential action.

What MCP tool poisoning is

The Model Context Protocol (MCP) connects an AI host and its client to servers that expose tools, resources and prompts. The client passes the definitions of available tools to the model. Those definitions are plain text: a tool name, a description, parameter names and a schema. The model reads them to decide whether to call a tool and what arguments to send. Tool poisoning exploits that reading step.

OWASP’s MCP Security Cheat Sheet defines the term directly: “Tool Poisoning: Malicious instructions hidden in tool descriptions, parameter schemas, or return values that manipulate the LLM’s behavior.” Two consequences follow. The payload does not have to be executable code; it only has to be text the model treats as an instruction. And it can sit in three places: the description, the parameters and schemas, and the output the tool returns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OWASP also publishes an MCP Top 10, which it describes as a living risk taxonomy. Use it to name categories in a review record, not to estimate how often attacks occur.

Can an MCP server description contain prompt injection?

Yes. A tool description is part of the context the model reads, so an instruction written there is a prompt injection delivered through metadata. It is typically written to look like ordinary help text. The table shows where such text can sit.

Location Pattern Illustrative example
Tool description Help text that includes an instruction unrelated to the tool’s stated function A description for a formatting tool that also tells the model to include configuration values in its next call
Parameter names and schemas A parameter name, description or default that steers the model toward sensitive values or other tools A parameter described as a destination field whose help text asks the model to fill it with a credential
Returned content Text in a tool’s output that addresses the model directly A fetched web page or file whose body contains a request to call a second tool

The examples describe the pattern, not specific observed payloads.

Why code review of your application misses it

Code review checks what your team wrote and merged. MCP tool poisoning reaches the model from outside that diff in four ways:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Descriptions and schemas are supplied by the server at runtime. Your repository may contain a wrapper around a tool, but not the text the model will actually read.
  • Definitions can change after approval. The version that passed review may not be the one the client serves next.
  • Returned data can carry instructions. Its content depends on what the tool fetched, which the code under review never sees.
  • Servers interact. A description from one server can change how the model uses another server’s tool, so each server can look reasonable when reviewed alone.

Pinning reviewed definitions narrows the gap, but only partly. OWASP cautions that pinned metadata reveals metadata changes and does not reveal changes to server code or behavior behind an unchanged definition. A clean definition review therefore does not prove the server will act as described.

Rug pulls and tool shadowing

These two patterns are the ones most likely to pass a one-time review.

Rug pull: definitions that change after approval

In a rug pull, a server’s tool definitions change after someone has approved it. The version that passed review is not the one the model later reads. Because the change happens after approval, the review that caught nothing was not wrong about the version it saw; the control that failed is the absence of a check on later changes. The review steps and controls below address that gap.

Tool shadowing: one server steering another server’s tool

OWASP describes tool shadowing as a case in which one server’s description manipulates how another tool is used. The server doing the steering need not run the dangerous action itself. For that reason, review the combined set of tools a model can see at once, since that is the context the model reasons over.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What determines the damage

The same injected instruction can be harmless in one setup and serious in another. The attack path crosses the server that supplies the text, the client that passes it along, the model that reads it, the configuration that grants access, and the interface that asks for approval. No single component owns the outcome. Three factors matter most.

Credentials and permissions

OWASP calls out over-scoped credentials and confused-deputy behavior, in which a server acts with privileges broader than the user intended. If a server’s credential can write to every repository or read every secret, a successful injection inherits that reach. The blast radius is set by configuration, not by the injected text.

Approval prompts

OWASP treats a coding assistant’s approval prompt as a security boundary. A prompt that omits parameters, or that the model’s own output can bypass, gives little protection. Auto-approving high-impact tool calls removes the boundary entirely.

Client safeguards

Client defenses differ between products and versions. The seven-client preprint and the benchmark figures described in the next section show what those studies did and did not establish, so do not assume any particular client behaves the way a study’s sample did.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the published measurements show

Two 2026 sources are often cited together. They measure different things, so they should not be merged into one estimate of how common tool poisoning is.

Item Seven-client preprint MCPTox benchmark figures
Source Charoes Huang, Xin Huang, Ngoc Phu Tran and Amin Milani Fard; arXiv paper dated March 23, 2026 Cloud Security Alliance AI Safety Initiative, summary dated July 1, 2026, reporting the MCPTox benchmark
Question asked How seven MCP clients defend against tool poisoning, based on threat modeling and an empirical comparison How often tool-poisoning attacks succeeded under the benchmark’s test conditions
Sample Seven MCP clients 45 live MCP servers and 20 language models
Headline result Differences in client defenses, with weaknesses in static validation and parameter visibility 36.5% average attack success rate across the benchmark; 72.8% highest rate against one model, as reported by the CSA summary
Peer review status Preprint; not presented as peer reviewed Not stated in the summary
What it does not establish A ranking of all MCP clients, or a guarantee about any named product version A real-world incident rate
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to review an MCP server before connecting it

Run this sequence before a server is added to any assistant that can act on your code, files or accounts.

  1. Record the source, owner and exact version. Note the package, repository or commit you are approving. A server with no identifiable owner is a reason to stop.
  2. Pull the complete tool list through your client’s inspection view. Review what the model receives, not only the project README. Capture tool names, descriptions, parameter names, parameter descriptions and schemas.
  3. Read each description as an instruction. Flag any text that does not describe the tool’s stated function: requests to read, include or send secrets; directions about other tools; unexpected destinations; hidden or encoded text. Finding nothing does not prove the text is safe.
  4. Trace the return path. For each tool, establish what it returns and whether it fetches external pages, files, issues or messages you do not control. Those are the points where instructions enter from outside your review.
  5. Map permissions to function. List each credential, OAuth scope, repository and filesystem path the server can reach. Remove any the stated function does not need.
  6. Run it first without production access. Use an environment with no production credentials and no access to sensitive directories.
  7. Pin and record the approval. Pin the reviewed definitions or their hashes where your client supports it. Record who approved them, when, and against which version.
  8. Trigger re-review on any change. A change to definitions, version, configuration or permissions requires a new review before the server reconnects.

How to secure MCP servers in a coding assistant

Treat these controls as layers. Each limits a different part of the failure chain, and each has a stated limit.

Layer What to implement What it does not cover
Inventory Permit only approved servers; record owner, source, version, configuration changes and why each server needs its permissions Shows what is connected, not what a server does at runtime
Pinning Pin reviewed tool definitions or hashes where supported; require review when they change Detects metadata changes only, not changes to server code behind an unchanged definition
Least privilege Use separate credentials per server, narrow OAuth scopes, short-lived credentials, and only the repository or filesystem access the server needs Limits reach; it does not stop a tool from misusing access it legitimately holds
Local isolation Restrict filesystem and network access for local servers to the minimum needed Standard input/output transport alone does not sandbox a process
Input and output handling Treat model-generated arguments and all tool results as untrusted; validate paths, URLs, shell and database inputs; block arbitrary URL fetching that could reach internal services Validation must match each input type, so it has to be written for each tool
Confirmation For sensitive or destructive operations, display complete tool-call parameters and require explicit confirmation; ensure the model cannot craft output that bypasses the prompt Only as strong as the client’s prompt design, which varies by product and version
Logging and policy Log and review consequential tool calls; add monitoring or policy enforcement alongside human review Records activity; blocking requires enforcement, and logging does not replace least privilege or isolation

Control availability depends on the MCP host, client, server and deployment. Confirm each control in the exact versions you run before following product-specific setup steps. OWASP’s MCP Security Cheat Sheet is the primary practical reference for the full set.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to stop and respond

Disconnect the server and investigate if you see any of the following:

  • A description, schema or permission differs from the version you approved.
  • A tool result contains text addressed to the assistant, such as an instruction to call another tool.
  • A proposed call sends data to a destination you did not configure.
  • An approval prompt shows fewer parameters than the call that actually runs.
  • The assistant asks for a credential or secret the task does not require.

Then rotate any credential the server could read, review the tool calls recorded in the session if your client keeps them, and re-review the server before reconnecting it.

The Bottom Line

Tool poisoning moves the trust question from your source code to the text your tools deliver at runtime. Assume a server can carry instructions, then limit what a successful instruction could reach: scoped credentials, isolated local processes and explicit approval for high-impact actions. Definition review remains useful, but it works best as a check that repeats on every change rather than a one-time sign-off.