AI Engineering

AI Coding Agent Security: How Prompt Injection Leads to Credential Theft (and How to Stop It)

Prompt injection against AI coding agents is a documented, exploited vulnerability class, not a hypothetical. Researchers have shown Claude Code, Gemini CLI, GitHub Copilot, Cursor, Codex, and Semantic Kernel agents can all be hijacked through content they're designed to read — GitHub issues, MCP tool responses, Sentry error events — into leaking API keys, tokens, and in some cases full remote code execution. Five public CVEs and two cross-vendor disclosures back this up. Here's how the attack works, what the real incidents looked like, and the concrete controls (allowlisting, sandboxing, human-approval gates) that actually reduce the risk.

Gurram Poorna Prudhvi

Lead AI Engineer

Technical Guide
Oct 7, 2026
11 min read
HOW PROMPT INJECTION HIJACKS CODING AGENTSuntrusted content + powerful tools + standing credentials, in one contextUNTRUSTED CONTENTDEFENSE LAYERWITHOUT DEFENSESAGENTreads context,calls toolsGitHub issue / PR commentMCP tool responseFetched web pageError log / changelogTool allowlistblocks or confirms the actionSandboxed executionblocks or confirms the actionHuman approval gateblocks or confirms the actionNO ALLOWLIST + NO SANDBOX + STANDING CREDENTIALS = CREDENTIAL THEFTthe agent performs only "authorized" actions with the developer's own identity — no policy is ever violatedaiengineerinsights.com

What is prompt injection against a coding agent?

Traditional security models keep instructions and data strictly separate. LLM-based agents don't: a systematization-of-knowledge paper analyzing 78 studies on the topic notes that models process both through the same channel, so content an agent merely reads — a PR title, an issue body, an MCP tool's response, a fetched web page, an error trace — can carry instructions the agent treats as authoritative. The same paper catalogs 42 distinct attack techniques and found attack success rates against published defenses exceed 85% once an attacker adapts. The OWASP Secure Coding with AI Cheat Sheet independently lists the same surface: issue bodies, PR comments, README files, dependency changelogs, and fetched pages all become instruction sources once an agent reads them.

This matters more for coding agents than chatbots because coding agents are wired to tools — shell execution, package installs, git push — and typically run with the developer's own credentials. A successful injection isn't a bad chat response; it's an action taken with your permissions.

Has this actually happened, or is it theoretical?

It's been demonstrated against production tools, cross-vendor. Security researcher Aonan Guan, working with Johns Hopkins researchers, showed that Anthropic's Claude Code Security Review, Google's Gemini CLI Action, and GitHub Copilot Agent could each be hijacked via GitHub PR titles, issue bodies, and comments — turning the host repository's own GitHub Actions secrets (API keys, tokens) into stolen data, exfiltrated back through GitHub itself with no external infrastructure. The Copilot variant bypassed three dedicated runtime defenses GitHub had added specifically to prevent this: environment filtering, secret scanning, and a network firewall.

Separately, the Cloud Security Alliance documented "agentjacking": attackers inject malicious instructions into Sentry error events using only a public, write-only DSN credential. When a developer asks their agent to investigate open errors, the agent retrieves the poisoned event through Sentry's MCP server and executes the embedded instructions with the developer's own privileges — recovering AWS credentials, GitHub/GitLab tokens, npm tokens, and CI/CD secrets in the proof-of-concept, with zero policy violated and zero anomaly threshold crossed.

What CVEs exist for this specifically?

The SoK paper above catalogs over 30 CVEs tied to agentic coding assistants; five representative ones:

CVEProductImpact
CVE-2025-49150CursorRemote code execution via MCP
CVE-2025-53773GitHub CopilotAuto-approve privilege escalation
CVE-2025-58335JunieData exfiltration
CVE-2025-61260Codex CLICommand injection
CVE-2025-53097Roo CodeCredential theft

Microsoft's research team went further, disclosing two CVEs in Microsoft Semantic Kernel (CVE-2026-25592, CVE-2026-26030) that chained prompt injection into full host-level remote code execution — one through an unsanitized parameter in a search plugin reaching Python's eval(), the other through an unvalidated file path letting an injected instruction write a malicious script straight to the host's Windows Startup folder from inside what was supposed to be an isolated sandbox.

Why is this hard to fix with blocklists?

Because the agent's job legitimately requires reading untrusted content. Blocking specific commands doesn't close the hole — the OWASP cheat sheet notes Anthropic blocked the ps command as a mitigation, and attackers simply moved to cat /proc/*/environ to read the same environment variables a different way. The SoK paper's review of 18 published defense mechanisms found most achieve less than 50% mitigation against adaptive attackers. The architectural fix isn't a smarter filter; it's removing the agent's ability to reach sensitive capabilities at all unless explicitly scoped.

How do I actually defend my agent setup?

Both the SoK paper and OWASP converge on the same shape of answer: least privilege, not detection. Meta's security team frames it as the "Rule of Two" — an agent should satisfy no more than two of: (A) processing untrusted input, (B) accessing sensitive data, (C) changing state or communicating externally. Every incident above is an agent doing all three at once. A practical tiered approval model, adapted from the SoK paper's proposal:

TierScopeExample
SilentRead-only ops inside project scopeReading files already in the repo
LoggedWrites to project files, shown in an activity feedEditing an existing source file
ConfirmedShell execution, network requests, cross-project accessInstalling a new npm package
BlockedCredential access, system modificationReading .env, writing to ~/.ssh

Concretely: use an explicit tool allowlist (e.g. Claude Code's --allowed-tools) rather than trying to blocklist dangerous ones; run agents with ephemeral, scoped credentials instead of long-lived developer tokens in environment variables; exclude .env, *.pem, and SSH keys from the agent's visible context; and require human approval before shell execution, network egress, or any credential-adjacent action — particularly for MCP servers and CI agents that process content from outside contributors.

Frequently asked questions

What is prompt injection in an AI coding agent?

Prompt injection is when content the agent reads — a GitHub issue, a PR comment, an MCP tool response, a fetched web page, an error log — contains instructions crafted to look like legitimate input but are actually commands for the model to follow. Because LLMs process instructions and data through the same channel, the agent can't reliably tell 'the user told me to do this' apart from 'a webpage told me to do this.'

Can prompt injection actually steal credentials from a coding agent?

Yes, and it has been demonstrated against production tools. Researchers showed Claude Code Security Review, Gemini CLI Action, and GitHub Copilot Agent could all be hijacked through GitHub issue/PR content to leak API keys and tokens from their own CI runner environment — using GitHub itself as the exfiltration channel, no external server required.

Is this just a theoretical risk, or are there real CVEs?

Real CVEs exist across major tools: CVE-2025-49150 (Cursor, RCE via MCP), CVE-2025-53773 (Copilot, auto-approve privilege escalation), CVE-2025-58335 (Junie, data exfiltration), CVE-2025-61260 (Codex CLI, command injection), and CVE-2025-53097 (Roo Code, credential theft). Microsoft also disclosed two CVEs in Semantic Kernel (2026-25592, 2026-26030) that chained prompt injection into full host-level remote code execution.

What is the 'Rule of Two' for agent security?

A guideline from Meta's security team: an agent should satisfy no more than two of (A) processing untrusted input, (B) accessing sensitive data, and (C) changing state or communicating externally. An agent that does all three — reads an untrusted GitHub issue, holds API keys, and can push commits — is exactly the shape every documented credential-theft incident takes.

How do I actually defend an AI coding agent against this?

Least privilege, not blocklisting. Scope tools to an allowlist (`--allowed-tools` in Claude Code, equivalent flags elsewhere) instead of trying to block dangerous commands one by one — blocklists are whack-a-mole. Run agents in sandboxed, ephemeral environments with scoped, short-lived credentials rather than long-lived developer tokens. Require explicit human approval for shell execution, network calls, and any credential-adjacent action. Treat every MCP server and every piece of repository content (issues, PR comments, READMEs) as untrusted input.

Does running an agent in a sandbox fully solve the problem?

No — sandbox escapes exist too. Microsoft's Semantic Kernel disclosure showed an agent sandbox could be defeated through an unvalidated file-path parameter in a host-side helper function, letting an attacker write a payload straight to the host's Startup folder from inside the 'isolated' sandbox. Sandboxing reduces blast radius; it doesn't replace capability scoping and human approval gates.

References

Get the AI engineering newsletter

Practical, engineering-first breakdowns on AI agents, LLMs, and breaking into the field — plus the free roadmap PDF when you join. No spam.

By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.

Found this useful? Share it.

Share:

Related Articles