Answer
Can prompt injection make a coding agent steal my AWS keys?
Yes. Documented research and real incidents show the same chain: a coding agent reads untrusted content (an issue, a ticket, a file), treats text inside it as instructions, and those instructions tell it to search for and send out whatever credentials it can reach — AWS keys, SSH keys, tokens. The agent isn't "hacked"; it's doing what it was told, by the wrong party.
Updated · Kloudle
What's the actual attack chain?
Four steps, consistently, across every case below. First, the agent is given untrusted content to process — a GitHub issue, a support ticket, a file in a repo, a coding-rule file it imports. Second, that content contains text written to look like an instruction, and the agent's model can't reliably tell "content to analyze" from "command to follow." Third, the agent acts on the injected instruction using whatever tools and access it already has — a shell, a file reader, an MCP server. Fourth, if a credential sits anywhere the agent can reach, the injected instruction can direct the agent to read it and exfiltrate it through any channel the agent already has open: a DNS lookup, a pull request, an outbound request to a server the attacker controls.
The chain doesn't require a new exploit against the model. It requires only that the agent's existing, legitimate reach includes a credential and a way out.
How often does this actually work, and against what?
The paper "Your AI, My Shell" tested two agentic coding editors — Cursor and GitHub Copilot in VS Code — with a framework of 314 unique attack payloads mapped to 70 MITRE ATT&CK techniques across 11 categories. Attack success rates ranged from 41.1% for GitHub Copilot with Gemini 2.5 Pro up to 83.4% for Cursor in Auto Mode, with command-execution rates of 75-88% across scenarios. The paper describes payloads that search the filesystem for AWS credentials — one run "first searches the entire root directory for any .aws folders" — and others that "read and then overwrite the ~/.ssh/authorized_keys file" to plant persistent access.
A 2026 Cloud Security Alliance research note describes the same chain reaching production cloud infrastructure: in an AWS AgentCore Harness deployment, a hidden instruction in a customer support ticket got the agent's shell tool — which ran as root inside the harness container — to pull and run a reconnaissance script. From there, instead of calling any AWS API, the attacker read the harness process's own memory through /proc/1/mem to locate and extract a live JWT token, 1,034 bytes, without ever touching an AWS credential directly.
Two further cases show this isn't tied to one vendor or one model. Security researcher Johann Rehberger found that Amazon Q Developer's command classifier treated ping and dig as "read-only" and let them run without approval, which was enough to exfiltrate the contents of a .env file over DNS once a malicious prompt was embedded in a code comment; AWS patched the classifier but did not issue a CVE or public advisory. Invariant Labs separately showed that a prompt injection hidden in a public GitHub issue could get an agent using the GitHub MCP server to pull private repository data and leak it by opening a pull request on the public repo — reproduced against Claude 4 Opus, described in the research as "a very recent, highly aligned and secure AI model," showing that model alignment alone didn't stop it.
Why does reducing what's reachable limit the damage?
Every case above has the same shape after step two: once the agent is following attacker-written instructions, nothing about the model's alignment or the vendor's filters stops it — the only thing left standing between the attacker and a working credential is whether that credential was reachable from the agent's position in the first place. The CSA note draws the direct lesson: root-level shell access sharing memory with credential-handling processes is what turned an injected instruction into a plaintext token; scoping the shell tool's privileges down would have closed that path regardless of the injection itself.
That's the practical argument for minimizing standing access rather than only trying to catch injected text: an agent that can't reach a credential can't be instructed into leaking it, no matter how convincing the injection is.
What mitigations actually help?
None of the sources above treat this as fully solvable by better prompting or model alignment. The consistent recommendations are architectural.
- Scope tool permissions explicitly per session rather than accepting a platform default that grants broad shell or filesystem access.
- Run agent shells and MCP servers as low-privilege, non-root identities, isolated from processes that handle credential decryption.
- Treat long-lived, broadly-scoped credentials as the real payoff for an attacker; short-lived, narrowly-scoped tokens shrink what a successful exfiltration is worth.
- Don't let a classifier's "read-only" label stand in for a real permission check — Amazon Q's ping/dig miscategorization shows automated read/write classification can be wrong in exploitable ways.
- Know what's actually reachable before you rely on any of the above. Agent Blast Radius runs a local, read-only scan that reports which credentials — AWS, SSH, Kubernetes, MCP configs, and more — a process running as you could currently reach, which is exactly the inventory an injected instruction would be searching for.
Check your own machine
See what an agent running as you can reach. Offline, read-only, never prints values:
npx -y @kloudle/agent-blast-radius@0.3.0Inside your agent: install guides for Claude Code, Codex, Cursor, Claude Desktop and more. Which keys are live? npx -y @kloudle/agent-blast-radius@0.3.0 verify (paid per check).
Frequently asked
Does prompt injection need a vulnerability in the AI model itself?
No. Every case cited here worked against models that weren't themselves broken — including, per Invariant Labs, Claude 4 Opus, described as a highly aligned model. The agent followed the injected instruction because it couldn't reliably distinguish untrusted content from a command, not because the model was exploited directly.
What's a concrete example of credentials actually being exfiltrated this way?
In the Cloud Security Alliance's AWS AgentCore research note, a prompt injected through a customer support ticket led to a root-level shell tool reading a live 1,034-byte JWT token directly out of process memory via /proc/1/mem, without the attacker ever holding an AWS credential.
Why does limiting what credentials are reachable matter if the injection still succeeds?
Because the injection only produces damage if there's something reachable to exfiltrate. The CSA note's own fix was scoping the shell tool's access away from credential-handling memory — the injection still happened, but there was nothing left for it to find.
Did any of these vendors publicly disclose the issue?
Inconsistently. The Register reported that AWS fixed the Amazon Q Developer command-classification flaws without issuing a CVE or a public security bulletin, which the researcher who found them specifically criticized by contrast with how Anthropic and Microsoft had handled comparable disclosures.
Sources
- Your AI, My Shell: Prompt Injection Attacks Against Agentic Coding Tools — arXiv
- CSA Research Note: AWS AgentCore Credential Exfiltration — Cloud Security Alliance
- Amazon quietly fixed Q Developer flaws that allowed code execution — The Register
- MCP Security Notification: Tool Poisoning Attacks (GitHub MCP) — Invariant Labs