What is prompt injection against a coding agent?
Traditional security models keep instructions and data strictly separate. LLM-based agents don't: a systematization-of-knowledge paper analyzing 78 studies on the topic notes that models process both through the same channel, so content an agent merely reads — a PR title, an issue body, an MCP tool's response, a fetched web page, an error trace — can carry instructions the agent treats as authoritative. The same paper catalogs 42 distinct attack techniques and found attack success rates against published defenses exceed 85% once an attacker adapts. The OWASP Secure Coding with AI Cheat Sheet independently lists the same surface: issue bodies, PR comments, README files, dependency changelogs, and fetched pages all become instruction sources once an agent reads them.
This matters more for coding agents than chatbots because coding agents are wired to tools — shell execution, package installs, git push — and typically run with the developer's own credentials. A successful injection isn't a bad chat response; it's an action taken with your permissions.
Has this actually happened, or is it theoretical?
It's been demonstrated against production tools, cross-vendor. Security researcher Aonan Guan, working with Johns Hopkins researchers, showed that Anthropic's Claude Code Security Review, Google's Gemini CLI Action, and GitHub Copilot Agent could each be hijacked via GitHub PR titles, issue bodies, and comments — turning the host repository's own GitHub Actions secrets (API keys, tokens) into stolen data, exfiltrated back through GitHub itself with no external infrastructure. The Copilot variant bypassed three dedicated runtime defenses GitHub had added specifically to prevent this: environment filtering, secret scanning, and a network firewall.
Separately, the Cloud Security Alliance documented "agentjacking": attackers inject malicious instructions into Sentry error events using only a public, write-only DSN credential. When a developer asks their agent to investigate open errors, the agent retrieves the poisoned event through Sentry's MCP server and executes the embedded instructions with the developer's own privileges — recovering AWS credentials, GitHub/GitLab tokens, npm tokens, and CI/CD secrets in the proof-of-concept, with zero policy violated and zero anomaly threshold crossed.
What CVEs exist for this specifically?
The SoK paper above catalogs over 30 CVEs tied to agentic coding assistants; five representative ones:
| CVE | Product | Impact |
|---|---|---|
| CVE-2025-49150 | Cursor | Remote code execution via MCP |
| CVE-2025-53773 | GitHub Copilot | Auto-approve privilege escalation |
| CVE-2025-58335 | Junie | Data exfiltration |
| CVE-2025-61260 | Codex CLI | Command injection |
| CVE-2025-53097 | Roo Code | Credential theft |
Microsoft's research team went further, disclosing two CVEs in Microsoft Semantic Kernel (CVE-2026-25592, CVE-2026-26030) that chained prompt injection into full host-level remote code execution — one through an unsanitized parameter in a search plugin reaching Python's eval(), the other through an unvalidated file path letting an injected instruction write a malicious script straight to the host's Windows Startup folder from inside what was supposed to be an isolated sandbox.
Why is this hard to fix with blocklists?
Because the agent's job legitimately requires reading untrusted content. Blocking specific commands doesn't close the hole — the OWASP cheat sheet notes Anthropic blocked the ps command as a mitigation, and attackers simply moved to cat /proc/*/environ to read the same environment variables a different way. The SoK paper's review of 18 published defense mechanisms found most achieve less than 50% mitigation against adaptive attackers. The architectural fix isn't a smarter filter; it's removing the agent's ability to reach sensitive capabilities at all unless explicitly scoped.
How do I actually defend my agent setup?
Both the SoK paper and OWASP converge on the same shape of answer: least privilege, not detection. Meta's security team frames it as the "Rule of Two" — an agent should satisfy no more than two of: (A) processing untrusted input, (B) accessing sensitive data, (C) changing state or communicating externally. Every incident above is an agent doing all three at once. A practical tiered approval model, adapted from the SoK paper's proposal:
| Tier | Scope | Example |
|---|---|---|
| Silent | Read-only ops inside project scope | Reading files already in the repo |
| Logged | Writes to project files, shown in an activity feed | Editing an existing source file |
| Confirmed | Shell execution, network requests, cross-project access | Installing a new npm package |
| Blocked | Credential access, system modification | Reading .env, writing to ~/.ssh |
Concretely: use an explicit tool allowlist (e.g. Claude Code's --allowed-tools) rather than trying to blocklist dangerous ones; run agents with ephemeral, scoped credentials instead of long-lived developer tokens in environment variables; exclude .env, *.pem, and SSH keys from the agent's visible context; and require human approval before shell execution, network egress, or any credential-adjacent action — particularly for MCP servers and CI agents that process content from outside contributors.
Frequently asked questions
What is prompt injection in an AI coding agent?
Prompt injection is when content the agent reads — a GitHub issue, a PR comment, an MCP tool response, a fetched web page, an error log — contains instructions crafted to look like legitimate input but are actually commands for the model to follow. Because LLMs process instructions and data through the same channel, the agent can't reliably tell 'the user told me to do this' apart from 'a webpage told me to do this.'
Can prompt injection actually steal credentials from a coding agent?
Yes, and it has been demonstrated against production tools. Researchers showed Claude Code Security Review, Gemini CLI Action, and GitHub Copilot Agent could all be hijacked through GitHub issue/PR content to leak API keys and tokens from their own CI runner environment — using GitHub itself as the exfiltration channel, no external server required.
Is this just a theoretical risk, or are there real CVEs?
Real CVEs exist across major tools: CVE-2025-49150 (Cursor, RCE via MCP), CVE-2025-53773 (Copilot, auto-approve privilege escalation), CVE-2025-58335 (Junie, data exfiltration), CVE-2025-61260 (Codex CLI, command injection), and CVE-2025-53097 (Roo Code, credential theft). Microsoft also disclosed two CVEs in Semantic Kernel (2026-25592, 2026-26030) that chained prompt injection into full host-level remote code execution.
What is the 'Rule of Two' for agent security?
A guideline from Meta's security team: an agent should satisfy no more than two of (A) processing untrusted input, (B) accessing sensitive data, and (C) changing state or communicating externally. An agent that does all three — reads an untrusted GitHub issue, holds API keys, and can push commits — is exactly the shape every documented credential-theft incident takes.
How do I actually defend an AI coding agent against this?
Least privilege, not blocklisting. Scope tools to an allowlist (`--allowed-tools` in Claude Code, equivalent flags elsewhere) instead of trying to block dangerous commands one by one — blocklists are whack-a-mole. Run agents in sandboxed, ephemeral environments with scoped, short-lived credentials rather than long-lived developer tokens. Require explicit human approval for shell execution, network calls, and any credential-adjacent action. Treat every MCP server and every piece of repository content (issues, PR comments, READMEs) as untrusted input.
Does running an agent in a sandbox fully solve the problem?
No — sandbox escapes exist too. Microsoft's Semantic Kernel disclosure showed an agent sandbox could be defeated through an unvalidated file-path parameter in a host-side helper function, letting an attacker write a payload straight to the host's Startup folder from inside the 'isolated' sandbox. Sandboxing reduces blast radius; it doesn't replace capability scoping and human approval gates.
References
- Prompt Injection Attacks on Agentic Coding Assistants (SoK, arXiv 2601.17548)
- OWASP — Secure Coding with AI Cheat Sheet
- Comment and Control: Prompt Injection to Credential Theft in Claude Code, Gemini CLI, and GitHub Copilot
- Cloud Security Alliance — Agentjacking: MCP Injection Hijacks AI Coding Agents
- Microsoft Security — When Prompts Become Shells: RCE Vulnerabilities in AI Agent Frameworks