AI Engineering

MCP Server Trust: How to Vet and Scope AI Agent Permissions Before You Connect One

Connecting an MCP server to your agent means trusting its tool descriptions as much as its code — and OWASP's MCP Top 10 catalogs five separate ways that trust gets abused. Tool poisoning hides instructions inside descriptions the model reads but a human never does. Rug-pull attacks change a tool's behavior after you've already approved it. Token passthrough turns your MCP server into a confused deputy with access it was never meant to have. None of this is theoretical — it's documented in OWASP's own vulnerability catalog, Microsoft's attack-technique database, and the official MCP specification's own security-considerations section. Here's what each attack actually looks like and the concrete scoping checklist that neutralizes them.

Gurram Poorna Prudhvi

Lead AI Engineer

Technical Guide
Oct 9, 2026
10 min read
HOW TO SCOPE TRUST FOR AN MCP SERVERtreat every new MCP server like an unvetted dependency, not a trusted extensionRISK CLASSMCP SERVERSCOPING CONTROLMCPSERVERtools + schemaexposed to agentTool description poisoningRug-pull (tool changes post-approval)Token passthrough / confused deputyShadow / unvetted MCP serversStatic scan before connectPin + diff tool manifestsScoped, audience-bound tokensUNSCOPED MCP TRUST = THE AGENT INHERITS WHATEVER THE SERVER TELLS ITOWASP MCP Top 10 and the official spec both say the same thing: least privilege, verified tokens, pinned manifestsaiengineerinsights.com

Why does MCP server trust matter more than trusting a regular API?

Because the thing you're trusting isn't just the server's code — it's the natural-language tool descriptions the server hands your agent, and the model treats those descriptions as instructions. The OWASP MCP Top 10 frames this directly: an MCP server's "privilege escalation via scope creep" (MCP02) and "tool poisoning" (MCP03) entries both exploit the fact that an agent reads a tool's declared surface as authoritative, with no separation between "data describing the tool" and "instructions the model should follow." That's a materially different trust boundary than calling a REST API, where the response schema is fixed and doesn't get to redirect the caller's behavior.

This connects directly to the credential-theft pattern covered in how prompt injection steals credentials from coding agents — an MCP server is one of the most common channels an agent reads untrusted content through, and what MCP actually is (a standard for exposing tools and data to an agent) is exactly why its trust model deserves its own scrutiny, separate from the agent's own permissions.

What is a tool poisoning attack, concretely?

Security firm Invariant Labs first disclosed Tool Poisoning Attacks in April 2025, documented in the systematic analysis of MCP security on arXiv: an attacker embeds hidden instructions inside a tool's name, description, or parameter text — often disguised as an innocuous code comment — that redirect the model's behavior while the tool's actual function looks benign. OWASP's write-up on MCP03: Tool Poisoning gives the specific tells to scan for before connecting a server: model-directed imperatives like "ignore previous instructions" or "do not tell the user," references to sensitive paths like ~/.ssh, .env, or .aws/credentials, and exfiltration patterns — an action verb like send or upload sitting near an external URL or webhook.

What is a rug-pull attack, and why doesn't approval protect you from it?

A rug pull is the dynamic version of tool poisoning: the tool definition you reviewed and approved is safe at approval time, then changes later. Microsoft Learn's AI attack-technique catalog walks through the mechanism: a trusted tool — their example is send_slack_message — gets silently altered server-side to exfiltrate data or run unauthorized actions. Because most agent integrations only check a tool's definition at first connection, not on every call, the malicious version runs without triggering a fresh approval prompt. The same arXiv analysis above measured a simulated rug-pull attack succeeding 80% of the time against unprotected test setups, which is why "approve once" is not the same guarantee as "safe forever."

What is token passthrough, and why does the MCP spec forbid it?

Token passthrough is when an MCP server receives an access token from the client and forwards that same, unmodified token to an upstream API instead of validating it and issuing its own. The official MCP authorization security-considerations spec states this plainly: "the MCP server MUST NOT pass through the token it received from the MCP client," because a downstream API may incorrectly trust that token as already validated by the MCP server — the classic confused-deputy problem, where an attacker with a stolen or mis-scoped token rides it further than it was ever meant to reach. The spec also requires tokens to be audience-bound — issued specifically for the MCP server that receives them — so a token leaked in one context can't be replayed against a different resource.

What does OWASP's full MCP Top 10 cover?

The full list has ten entries; these five are the ones most directly under an individual engineer's control when choosing and configuring a server, rather than something only the server operator can fix:

IDRiskWhat scopes it down
MCP01Token mismanagement & secret exposureShort-lived, audience-bound tokens; never log or pass through raw tokens
MCP02Privilege escalation via scope creepReview requested scopes on every update, not just first install
MCP03Tool poisoningStatic-scan tool descriptions for model-directed imperatives before connecting
MCP07Insufficient authentication & authorizationVerify the server implements OAuth 2.1 + Protected Resource Metadata
MCP09Shadow MCP serversMaintain an approved-server registry; block ad hoc stdio configs in CI

What's a practical scoping checklist before connecting a new MCP server?

Treat a new MCP server like a new npm dependency, not a trusted built-in — review it before you add it, and keep checking it afterward:

StageAction
Before connectingRead the full tool manifest, not just the names — scan descriptions for hidden imperatives
Before connectingCheck whether tokens are audience-bound (issued specifically for this server) or passed through
At connect timeGrant the narrowest OAuth scope the task needs, not the broadest the server offers
After connectingPin the tool manifest's hash; diff it on every server update to catch silent rug-pulls
OngoingLog every tool call with the server identity, so a compromised server is traceable

Even Anthropic's own reference implementations don't get a pass here. The SECURITY.md for the modelcontextprotocol/servers repo states outright that those servers are "reference implementations intended to demonstrate MCP features and SDK usage, not production-ready solutions," and that Anthropic's bug-bounty program explicitly excludes them — it covers only the MCP SDKs. "Official" is not a substitute for "vetted and scoped."

Frequently asked questions

What is MCP server trust scoping?

It's the practice of treating every Model Context Protocol server as untrusted supply chain — like a new npm package — until you've reviewed its tool manifest, scoped the credentials it receives, and set up a way to detect if its behavior changes later. The alternative, connecting any MCP server with a long-lived, broadly scoped token and never re-checking it, is how tool poisoning and rug-pull attacks succeed.

What is a tool poisoning attack against an MCP server?

OWASP's MCP Top 10 (MCP03) describes it as an adversary embedding instructions inside a tool's name, description, or parameter text — the part the model treats as authoritative, not the part a human reviews. Invariant Labs first disclosed this in April 2025: a tool description can tell the model to 'ignore previous instructions' or quietly read ~/.ssh/id_rsa, and the agent follows it because tool descriptions aren't rendered as prose a user reads before approving.

What is an MCP 'rug pull' attack?

A dynamic version of tool poisoning: a tool's definition looks safe when you approve it, then changes after the fact. Microsoft Learn's AI attack catalog documents the pattern — a trusted tool like send_slack_message is silently altered server-side to exfiltrate data instead, and because agents don't routinely re-verify a tool's definition after the first approval, the malicious version executes without any approval prompt firing again.

What is token passthrough and why is it dangerous in MCP?

It's when an MCP server forwards a token it received from the client straight to an upstream API instead of validating and re-issuing its own. The official MCP authorization spec explicitly forbids this: a server MUST NOT pass through a token it didn't issue, because a downstream API may wrongly trust the token as already validated — the classic 'confused deputy' pattern that lets an attacker access resources the original token was never meant to reach.

How do I scope AI agent permissions for MCP servers in practice?

Apply least privilege at every layer: grant the narrowest OAuth scope a task needs rather than the broadest the server advertises, use short-lived audience-bound tokens instead of standing credentials, statically scan each tool's declared description for model-directed imperatives before connecting, and pin the tool manifest so you can diff it on every server update. This mirrors Meta's 'Rule of Two' for agent sessions generally — don't let a single session combine untrusted input, access to sensitive systems, and the ability to change state unchecked.

Are official MCP reference servers safe to use?

They're educational references, not hardened production software. Anthropic's own SECURITY.md for the modelcontextprotocol/servers repo states plainly that the reference servers are intended to demonstrate SDK usage and that its bug-bounty program does not cover vulnerabilities found in them — the bounty applies only to the MCP SDKs. Treat any reference server you actually run the same way you'd treat a third-party one: scoped tokens, pinned manifests, no blanket trust just because the repo is official.

References

Get the AI engineering newsletter

Practical, engineering-first breakdowns on AI agents, LLMs, and breaking into the field — plus the free roadmap PDF when you join. No spam.

By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.

Found this useful? Share it.

Share:

Related Articles