The Allowlist Illusion: Why Command Approval Keeps Failing in Coding Agents
Three unrelated 2026 disclosures — Cursor, Semantic Kernel, and the wider prompt-injection numbers behind them — converge on the same gap: an allowlist checks what a command looks like, not what put it there.
Ask an AI coding agent to run terminal commands unattended and the standard mitigation is an allowlist: a fixed set of commands — git status, npm test, ls — that execute without a human in the loop, with everything else routed to approval. It is a sensible design on its face. The commands look safe individually, and a person still signs off on anything unfamiliar.
2026 produced a run of disclosures that take that design apart from different directions, and reading them together says something more specific than "prompt injection is bad." Each shows a different point where the allowlist's actual guarantee — this exact command runs, on this exact arguments — turned out not to be the guarantee anyone needed.
An allowlist verifies syntax. None of the failures below required forging a new command. They worked by changing what an approved command did once it ran, or by using a code path the allowlist was never positioned to see in the first place.
The allowlist never looked at built-ins
CVE-2026-22708 landed on Cursor. In Auto-Run mode with an allowlist configured, the check only inspected external binaries. Shell built-ins — export, alias, typeset, declare — ran unchecked regardless of allowlist state, because they are not separate processes the allowlist mechanism was watching for.
That gap is enough on its own. Prompt injection delivered through a README, a dependency file, or an issue comment — any text the agent reads while working — could instruct it to run export PATH=/tmp/evil:$PATH, or alias git to a script under attacker control. Neither line needs approval, because built-ins are invisible to a mechanism designed to gate binaries. The next allowlisted command the developer actually approves, something as unremarkable as git branch, then resolves to the attacker's binary instead of git.
The developer did everything the allowlist asked of them. They reviewed a command that looked exactly like the one that ran, and it did something else, because the environment underneath it had already been rewritten by a step nobody was screening. Cursor rated the bug High and patched it in version 2.3, but the underlying lesson outlives the patch: a review model built around individual command shape has no representation for state that persists across commands.
Some agents don't need a shell to get one
Microsoft's Semantic Kernel disclosure in May made a related point more starkly. There was no shell to bypass and no allowlist to route around — the vulnerable primitive was a tool function the framework had already, legitimately, exposed to the model.
CVE-2026-26030 hit the Python SDK: a crafted filter value walked Python's class hierarchy to reach the interpreter's import mechanism, giving the model a path to arbitrary code execution — Microsoft's proof of concept simply launched calc.exe on the host. CVE-2026-25592 hit the .NET SDK for a plainer reason: DownloadFileAsync on SessionsPythonPlugin had been marked with the [KernelFunction] attribute, which means the LLM could call it directly. Nobody bypassed a permission boundary in that second case, because the boundary had already granted the model that capability by decorator.
An allowlist has nothing to say about either bug, because there is no shell command in the trace to allow or deny. The attack surface here is not "what commands can this agent run" — it is "what did we mark as callable by the model," and in a framework with dozens of plugins that list grows without anyone re-reading each addition as a security-relevant grant. A [KernelFunction] attribute is not internal wiring; it is an API endpoint with the LLM as the only caller, and it deserves the review an API endpoint gets.
Single-attempt defense is answering the wrong question
Put a number on how often this actually works and the case for urgency gets stronger, not weaker. OWASP's 2026 GenAI Top Ten — led by prompt injection for the third year running, and for the first time weighted 25% on data drawn from 6,639 real incidents rather than expert opinion alone — reports indirect prompt-injection success in agentic coding environments at 4.7% for a single attempt, 33.6% at ten attempts, and 63.0% at a hundred.
4.7% reads like a comfortable margin in a one-off pen test. It is the wrong frame for how these agents actually run. A coding agent that reads the same repository on every CI trigger, every scheduled task, every re-run after a failed build, hands an attacker who has poisoned one README effectively unlimited attempts against the same payload. Evaluated at n=1, a defense can look adequate; evaluated at n=100, which is closer to how an unattended agent behaves over a project's lifetime, the same defense is failing on nearly two-thirds of runs.
Excessive Agency climbing to third on the same OWASP list is the mechanical consequence, not a coincidence. More autonomy per run means more chances for the same low per-attempt probability to compound into a near-certainty.
What changes when you design for this
Bound capability, not command shape.A process's reachable filesystem paths, network egress, and ability to mutate its own environment are properties you can scope with a sandbox or a restricted credential. They survive an attacker choosing different arguments than the ones you allowlisted; a list of approved strings does not.
Audit the sequence, not the final command. Cursor's gap existed because review happened at the moment of execution, with no visibility into the built-ins that ran earlier in the same session. Anything that can mutate PATH, aliases, or environment variables is reachable attack surface even when it never appears on an allowlist, because it changes what the allowlisted command resolves to.
Treat a tool decorator as a capability grant. Every function exposed to a model — a [KernelFunction], an MCP tool, a registered callable — is effectively public to whatever text the model has read that turn. Review each one for what it lets a caller do if the caller's intent is adversarial, not for what it does when called as intended.
Evaluate under repeated exposure, not a single attempt. If an agent re-reads the same repository, ticket queue, or inbox on a recurring trigger, model defenses against the number of times a real adversary gets to try, not the number of times a test suite tries once.