Why does MCP server trust matter more than trusting a regular API?
Because the thing you're trusting isn't just the server's code — it's the natural-language tool descriptions the server hands your agent, and the model treats those descriptions as instructions. The OWASP MCP Top 10 frames this directly: an MCP server's "privilege escalation via scope creep" (MCP02) and "tool poisoning" (MCP03) entries both exploit the fact that an agent reads a tool's declared surface as authoritative, with no separation between "data describing the tool" and "instructions the model should follow." That's a materially different trust boundary than calling a REST API, where the response schema is fixed and doesn't get to redirect the caller's behavior.
This connects directly to the credential-theft pattern covered in how prompt injection steals credentials from coding agents — an MCP server is one of the most common channels an agent reads untrusted content through, and what MCP actually is (a standard for exposing tools and data to an agent) is exactly why its trust model deserves its own scrutiny, separate from the agent's own permissions.
What is a tool poisoning attack, concretely?
Security firm Invariant Labs first disclosed Tool Poisoning Attacks in April 2025, documented in the systematic analysis of MCP security on arXiv: an attacker embeds hidden instructions inside a tool's name, description, or parameter text — often disguised as an innocuous code comment — that redirect the model's behavior while the tool's actual function looks benign. OWASP's write-up on MCP03: Tool Poisoning gives the specific tells to scan for before connecting a server: model-directed imperatives like "ignore previous instructions" or "do not tell the user," references to sensitive paths like ~/.ssh, .env, or .aws/credentials, and exfiltration patterns — an action verb like send or upload sitting near an external URL or webhook.
What is a rug-pull attack, and why doesn't approval protect you from it?
A rug pull is the dynamic version of tool poisoning: the tool definition you reviewed and approved is safe at approval time, then changes later. Microsoft Learn's AI attack-technique catalog walks through the mechanism: a trusted tool — their example is send_slack_message — gets silently altered server-side to exfiltrate data or run unauthorized actions. Because most agent integrations only check a tool's definition at first connection, not on every call, the malicious version runs without triggering a fresh approval prompt. The same arXiv analysis above measured a simulated rug-pull attack succeeding 80% of the time against unprotected test setups, which is why "approve once" is not the same guarantee as "safe forever."
What is token passthrough, and why does the MCP spec forbid it?
Token passthrough is when an MCP server receives an access token from the client and forwards that same, unmodified token to an upstream API instead of validating it and issuing its own. The official MCP authorization security-considerations spec states this plainly: "the MCP server MUST NOT pass through the token it received from the MCP client," because a downstream API may incorrectly trust that token as already validated by the MCP server — the classic confused-deputy problem, where an attacker with a stolen or mis-scoped token rides it further than it was ever meant to reach. The spec also requires tokens to be audience-bound — issued specifically for the MCP server that receives them — so a token leaked in one context can't be replayed against a different resource.
What does OWASP's full MCP Top 10 cover?
The full list has ten entries; these five are the ones most directly under an individual engineer's control when choosing and configuring a server, rather than something only the server operator can fix:
| ID | Risk | What scopes it down |
|---|---|---|
| MCP01 | Token mismanagement & secret exposure | Short-lived, audience-bound tokens; never log or pass through raw tokens |
| MCP02 | Privilege escalation via scope creep | Review requested scopes on every update, not just first install |
| MCP03 | Tool poisoning | Static-scan tool descriptions for model-directed imperatives before connecting |
| MCP07 | Insufficient authentication & authorization | Verify the server implements OAuth 2.1 + Protected Resource Metadata |
| MCP09 | Shadow MCP servers | Maintain an approved-server registry; block ad hoc stdio configs in CI |
What's a practical scoping checklist before connecting a new MCP server?
Treat a new MCP server like a new npm dependency, not a trusted built-in — review it before you add it, and keep checking it afterward:
| Stage | Action |
|---|---|
| Before connecting | Read the full tool manifest, not just the names — scan descriptions for hidden imperatives |
| Before connecting | Check whether tokens are audience-bound (issued specifically for this server) or passed through |
| At connect time | Grant the narrowest OAuth scope the task needs, not the broadest the server offers |
| After connecting | Pin the tool manifest's hash; diff it on every server update to catch silent rug-pulls |
| Ongoing | Log every tool call with the server identity, so a compromised server is traceable |
Even Anthropic's own reference implementations don't get a pass here. The SECURITY.md for the modelcontextprotocol/servers repo states outright that those servers are "reference implementations intended to demonstrate MCP features and SDK usage, not production-ready solutions," and that Anthropic's bug-bounty program explicitly excludes them — it covers only the MCP SDKs. "Official" is not a substitute for "vetted and scoped."
Frequently asked questions
What is MCP server trust scoping?
It's the practice of treating every Model Context Protocol server as untrusted supply chain — like a new npm package — until you've reviewed its tool manifest, scoped the credentials it receives, and set up a way to detect if its behavior changes later. The alternative, connecting any MCP server with a long-lived, broadly scoped token and never re-checking it, is how tool poisoning and rug-pull attacks succeed.
What is a tool poisoning attack against an MCP server?
OWASP's MCP Top 10 (MCP03) describes it as an adversary embedding instructions inside a tool's name, description, or parameter text — the part the model treats as authoritative, not the part a human reviews. Invariant Labs first disclosed this in April 2025: a tool description can tell the model to 'ignore previous instructions' or quietly read ~/.ssh/id_rsa, and the agent follows it because tool descriptions aren't rendered as prose a user reads before approving.
What is an MCP 'rug pull' attack?
A dynamic version of tool poisoning: a tool's definition looks safe when you approve it, then changes after the fact. Microsoft Learn's AI attack catalog documents the pattern — a trusted tool like send_slack_message is silently altered server-side to exfiltrate data instead, and because agents don't routinely re-verify a tool's definition after the first approval, the malicious version executes without any approval prompt firing again.
What is token passthrough and why is it dangerous in MCP?
It's when an MCP server forwards a token it received from the client straight to an upstream API instead of validating and re-issuing its own. The official MCP authorization spec explicitly forbids this: a server MUST NOT pass through a token it didn't issue, because a downstream API may wrongly trust the token as already validated — the classic 'confused deputy' pattern that lets an attacker access resources the original token was never meant to reach.
How do I scope AI agent permissions for MCP servers in practice?
Apply least privilege at every layer: grant the narrowest OAuth scope a task needs rather than the broadest the server advertises, use short-lived audience-bound tokens instead of standing credentials, statically scan each tool's declared description for model-directed imperatives before connecting, and pin the tool manifest so you can diff it on every server update. This mirrors Meta's 'Rule of Two' for agent sessions generally — don't let a single session combine untrusted input, access to sensitive systems, and the ability to change state unchecked.
Are official MCP reference servers safe to use?
They're educational references, not hardened production software. Anthropic's own SECURITY.md for the modelcontextprotocol/servers repo states plainly that the reference servers are intended to demonstrate SDK usage and that its bug-bounty program does not cover vulnerabilities found in them — the bounty applies only to the MCP SDKs. Treat any reference server you actually run the same way you'd treat a third-party one: scoped tokens, pinned manifests, no blanket trust just because the repo is official.