The supply chain is the attack surface
Why your AI agent's next compromise won't arrive through the prompt
The supply chain is the attack surface
Why your AI agent's next compromise won't arrive through the prompt
For two years the security conversation around AI agents has been obsessed with the prompt. Prompt injection this, jailbreak that, system-prompt leakage everywhere. These are real problems — I've written about them at length [1][2] — but the industry's fixation on what an agent reads has blinded it to where agents actually execute. An agent is a program. It installs dependencies, clones repositories, runs build scripts, and downloads binaries. Every one of those acts is a supply-chain transaction, and every supply-chain transaction is an attack surface.
The prompt is the doorbell. The supply chain is the house.
flowchart LR
subgraph SRC["What the agent trusts"]
R["Dependencies
npm / Go modules"]
B["Repo contents
Makefile · tasks.json · postinstall"]
M["MCP servers
declared in config"]
H["Harness binary
curl | sh"]
end
subgraph ATTACK["Attack classes"]
A1["📦 Install-script
worms
Shai-Hulud"]
A2["🧬 Transitive
backdoors
XZ Utils"]
A3["⚙️ Build-system
compromise
tj-actions"]
A4["🎭 Instruction
smuggling
IPI in repos"]
A5["☠️ Poisoned
distribution
no signing"]
end
R --> A1
R --> A2
B --> A3
B --> A4
M --> A4
H --> A5
subgraph EXEC["Agent execution path"]
G{"Deterministic
execution gate"}
X["Shell · install ·
network egress"]
end
A1 --> G
A2 --> G
A3 --> G
A4 --> G
A5 --> G
G -->|approved| X
G -.->|fails closed| D["🚫 Denied"]
style D fill:#3a1a1a,stroke:#c0392b,color:#e8c8c8
style G fill:#1a2a3a,stroke:#3b82f6,color:#dbeafeAgents are the perfect mark
Consider what makes a software supply chain attack work in the classic sense: a developer, under time pressure, installs a package with a familiar-looking name (crossenv instead of cross-env), the install script fires, credentials leave the machine. The entire genre depends on the human being busy, trusting, and not looking too closely.
Now consider what an autonomous agent is: a system designed to not be blocked by friction. Its whole value proposition is that it doesn't stop to ask. It reads a task, plans, and executes — cloning repos, running npm install / go mod tidy, invoking build tools, fetching MCP servers from configuration files it did not write. The classic attack preconditions are not merely present; they are architectural features.
Worse, the agent inherits trust from everywhere the human did:
- A
README.mdthat says "run this setup script first" is, to an agent, a reasonably authoritative instruction source. - A
.vscode/tasks.json, aMakefile, or apackage.jsonpostinstallhook is a pre-approved execution path. Even the evaluation literature on indirect injection has documented instructions in project files being treated as legitimate work, precisely because they arrive through the same channel as legitimate work [6]. - An MCP server declared in a cloned repo's configuration arrives with the implicit endorsement of "it was in the repo."
A human developer at least has muscle-memory suspicion. An agent trained to be maximally helpful has the opposite default: the instruction was there, therefore executing it is probably what was wanted. This is indirect prompt injection's quieter sibling — not "the model was tricked by a web page," but "the model was handed a poisoned artifact and ran it as intended. The injected payload doesn't even need to be language. npm install is already an imperative sentence that every runtime agrees to obey.
Three attack classes that arrive through dependencies
1. The install script. The most direct vector, and the one with the best-documented track record. The 2025 "Shai-Hulud" worm spread across hundreds of npm packages through self-replicating postinstall scripts [3] — code that executes at install time, before the agent ever produces a useful token of output. Defenses that inspect model outputs see nothing. The compromise happens in a subprocess the agent spawned but never reasoned about.
2. The trojaned dependency of a dependency. Even a careful agent that audits its direct dependencies doesn't walk the transitive graph. XZ Utils (CVE-2024-3094) showed how far a patient maintainer-trust attack can travel — a backdoor planted not in an obscure package, but in a compression library sitting underneath the world's SSH daemons [4]. When an agent scaffolds a project from a template, it pulls in hundreds of packages the operator never saw listed anywhere.
3. The build system as interpreter. Makefile targets, CI workflow steps, code-gen scripts — these are programs that developers routinely execute without reading. The 2025 tj-actions/changed-files compromise turned a widely-used GitHub Action into a secret-dumping step for every workflow that pinned it by tag [5]. An agent told to "make sure CI passes" will run whatever the pipeline actually contains. The build system is where natural-language trust boundaries and arbitrary code execution intersect — and in the harnesses I've reviewed, it is essentially unguarded.
The common thread: by the time a supply-chain payload runs, the "prompt" is long gone. Defenses that live in the context window are defending the wrong perimeter.
What a defense actually requires
If the attack arrives as code, the defense has to live where code runs. That means a harness needs at least four things — and in the harnesses I've reviewed, none shipped all four:
Dependency minimalism. Every dependency is a pre-authorized stranger. A runtime with seven direct module dependencies — three of them first-party, the rest golang.org/x — has a supply chain you can actually read in one sitting. A runtime with two hundred transitive packages does not, no matter how good its prompt hardening is. Small attack surface isn't an aesthetic preference; it's the precondition for auditing anything at all.
Execution gating at the tool boundary. The agent's ability to run arbitrary commands must pass through a deterministic risk classifier — one that cannot be talked out of its decisions by clever phrasing, because it never reads phrasing. Installing, network egress, and writes outside the workspace are different risk classes, and they should require visibly different levels of authorization. A model's confidence is not a security control; an allowlist with an approval gate is.
Content provenance for everything that might become an instruction. Anything that arrives from outside and can steer behavior — skills, MCP servers, project configuration — needs a provenance check, not just a checksum. Taint tracking matters here too: if untrusted content flowed into a memory or a skill, the harness should know that forever after, not just for the current turn.
Signed, attested releases. This is the piece the industry simply skips. If the harness itself is a binary that operators install — often via curl | sh from a landing page — then the harness distribution channel is itself a supply chain, and a compromise of the harness poisons every session it ever runs. The defense is boring and well-understood from the container world: SHA-256 checksums, an SPDX SBOM enumerating exactly what shipped, and keyless Sigstore signatures with in-toto attestations binding each artifact to the exact source commit that produced it. Any operator can then verify this binary was built from this commit by this pipeline — no trust in the download server required.
Odek: the case study
I build odek, a minimal autonomous agent runtime in Go, and I'll be blunt: I built it the way described above because I couldn't find a harness that treated its own distribution chain as attack surface.
The posture is unusual enough to be worth enumerating:
- A readable supply chain. Seven direct dependencies. Three are first-party (the LLM SDK, the MCP framework, the vector store), four are
golang.org/x. One static binary — megabytes, not gigabytes — and no frameworks. You can diff an entire release's dependency set by eye. - Deterministic execution gating. Every tool call passes a risk-class classifier —
safe,local_write,system_write,destructive,network_egress,code_execution, and related classes — with fail-closed approval gates and single-call self-gating so multi-call batches can't smuggle a high-risk operand inside a low-risk wrapper. The classifier doesn't parse the model's reasoning; it parses the operation. - An untrusted-content boundary with teeth. All inbound content — tool output, file reads, web pages, memory — is demarcated as data, injection-scanned, and taint-tracked into memory and skills. Skills carry declared provenance, and the provenance gate refuses unsigned entries.
- MCP treated as hostile by default. MCP servers, the fastest-growing extension surface in the ecosystem, are exactly the "cloned repo, implicit endorsement" vector described above. Odek hardens the MCP boundary on both the client and server sides, with SSRF and egress controls on its network-facing tools.
- Signed releases, end to end. Every release ships SHA-256 checksums, an SPDX SBOM covering all binaries plus the full
go.modset, and keyless Sigstore signatures (cosign sign-blob) with in-toto attestations naming the exact source commit — tool versions pinned so a future major release can't silently change what a tag means. Install the binary, verify the bundle, and you have cryptographic proof of provenance rather than an HTTPS status code and good intentions.
None of these are exotic techniques. They're the standard playbook for any modern software artifact — applied, rarely and incompletely, to the agent runtime itself. That's the point. The techniques were never the hard part; deciding that an agent harness is critical supply-chain infrastructure was.
The uncomfortable conclusion
Every serious agent deployment today trusts a chain that looks like this: model weights nobody audited, hundreds of transitive packages nobody read, MCP servers nobody verified, build scripts nobody opened, and a harness binary downloaded over HTTPS with nothing but a status code to vouch for it. Each link in that chain is a standing invitation, and agents make every link easier to pull because agents execute without the hesitation that used to be the last line of defense.
The industry will not fix this by better prompting. Guardrails in the context window are a moat around the living room of a house whose front wall is missing. The fix is structural: minimal dependencies, deterministic execution gates, provenance for anything that can become an instruction, and signed, attested distribution for the harness itself. If you're procuring or building on an agent runtime, the ask is one sentence long: show me the SBOM and the signature. If the answer is a shrug, you've found your next incident.
Your agent is only as trustworthy as the least verified thing it will ever install. Audit that thing. Or better: demand a harness where the list of things to audit fits on one screen.
References
[1] Kyberneees. (2026, July). You Only Have to Be Prompt Injected Once. 21no.de. https://21no.de/publications/prompt-injected-once/
[2] Kyberneees. (2026, September). Who Watches What Your Agent Is Doing? 21no.de. https://21no.de/publications/who-watches-your-agent/
[3] Wiz. (2025). Shai-Hulud: the npm self-replicating worm — supply chain attack analysis. https://www.wiz.io/blog/shai-hulud-2-0-ongoing-supply-chain-attack
[4] Freund, A. (2024, March 29). backdoor in upstream xz/liblzma leading to ssh server compromise (CVE-2024-3094). oss-security mailing list. https://www.openwall.com/lists/oss-security/2024/03/29/4
[5] StepSecurity. (2025). Harden-Runner detection: tj-actions/changed-files action is compromised. https://www.stepsecurity.io/blog/harden-runner-detection-tj-actions-changed-files-action-is-compromised
[6] Kyberneees. (2026). Understanding Indirect Prompt Injections. BackendStack21. https://github.com/BackendStack21/indirect-prompt-injections
Kyberneees, September 2026