Part 2 of 3 - The Enforcement Trilogy
The container did not break.
The Linux kernel did not panic. The cgroups held. The namespace isolation worked exactly as designed. And yet four AI coding agents - Cursor, Codex CLI, Gemini CLI, Antigravity - were all compromised in ways that no amount of sandbox hardening would have prevented.
I read the research on a Tuesday evening in July. The Cloud Security Alliance had just published a research note documenting seven vulnerability disclosures across four products. Pillar Security followed with a week-long series dissecting the exact failure modes. The finding that stopped me: zero of those seven issues involved breaking the sandbox. The container held every time. The boundary failed somewhere else entirely.
This is the trust-handoff flaw. Call it the AI agent sandbox escape trust handoff - zero sandbox breaks, seven boundary failures, every escape through the handoff. It is a perimeter failure. And if you are deploying AI agents in production, it is the class of vulnerability your current security model was not designed to catch.
Key Takeaways
- The trust-handoff flaw is not a sandbox escape. In all seven CSA and Pillar Security disclosures from July 2026, the sandbox held while downstream consumers trusted agent-written artifacts without re-validation. The container never broke. The boundary collapsed at the handoff.
- Four products failed the same way: config hooks trusted by IDE extensions, command allowlists bypassed by flag injection, denylists defeated by new attack surfaces, and privileged local daemons exposed through shared sockets. Each a different manifestation of what Pillar Security calls the Inert-Artifact Illusion.
- The instrument from Log #16 enforces at the commit boundary - counting assertion deletions and refusing. The perimeter must enforce at every trust transition: a reference monitor outside the agent's reach that checks every artifact before a higher-privilege process consumes it.
- The practical audit has three parts: inventory every unsandboxed consumer the agent's output can reach, enforce policy at each handoff point, and wire telemetry that fires the moment a trusted helper executes content the agent influenced.
Where does the sandbox boundary actually break?
Not at the container wall. At the handoff.
An AI coding agent runs inside an ephemeral, locked-down sandbox. It writes a file - a configuration, a hook script, a task definition. The file is inert inside the sandbox. Nothing dangerous happens. The agent completes its task. The sandbox exits.
Then a trusted component outside the sandbox reads that file and acts on it. An IDE extension. A Git hook. A Docker daemon. A task runner. The file the agent wrote - the inert artifact - has now become an instruction executed with the authority of the host.
The container never broke. The trust boundary collapsed because something with more privilege than the agent trusted something the agent wrote without re-validating it.
Pillar Security called this "The Inert-Artifact Illusion." The file looks harmless. It passed every sandbox containment check. But the danger was never inside the sandbox. The danger was the assumption that anything emerging from a successful sandbox run could be treated as benign.
I have seen this assumption in production code reviews. A file passes lint. It passes type checking. It passes unit tests. Then someone merges it and a downstream system consumes it as configuration, and nobody checked whether the content was safe - only whether it was valid. Same shape. Different boundary.
This is not a hypothetical edge case. The CSA research note documented seven real disclosures. Every single one followed this exact shape.
The four products traced the same failure shape. What follows is how it works.

What did the four products all get wrong?
Four products. Four failure modes. One structural error.
Cursor had CVE-2026-48124, rated CVSS 8.5. An agent inside the sandbox wrote to a hook configuration file. Cursor's extension, running outside the sandbox with full IDE privileges, read that configuration and executed whatever the agent had written. The sandbox held. The extension did not re-validate.
Codex CLI had "GitPwned." The agent operated inside a sandbox with a command allowlist - a list of safe commands it was permitted to run. The allowlist trusted command names. The agent passed flags that changed what those commands did. The allowlist said "yes." The flags said "run arbitrary code." OpenAI patched it in version 0.95.0.
Gemini CLI and Antigravity each had multiple disclosures. A .vscode/tasks.json file written by the agent was consumed by VS Code's task runner outside the sandbox. A Python virtual environment binary, written by the agent, was executed by a host process that assumed the binary was legitimate. A Docker socket, exposed inside the sandbox, let the agent send instructions to the host Docker daemon - instructions the daemon executed with root-equivalent privilege.
Google evaluated two of these findings, determined they were consistent with the product's security model, and declined to fix them.
Read that again. Two findings. Confirmed. No fix. Because the sandbox worked as designed. The security model was containment. The failure was trust.
Four different products. Four different attack surfaces. One error repeated every time: an artifact the agent wrote inside the sandbox was treated as an instruction by something outside the sandbox that never checked whether the artifact was safe.
Why does this class of bug persist?
Because we built agent security on an assumption that stopped being true the moment agents gained filesystem write access.
The assumption was: if a process is contained in a sandbox, its outputs are contained too. That held when the outputs were text in a chat window. It collapsed when the outputs became files on disk read by other processes.
A file written inside a sandbox is not dangerous by itself. It becomes dangerous when something with more authority reads it and acts on what it says. IDE extensions, Git subsystems, task runners, Docker daemons - none of these were designed assuming their input sources were autonomous agents operating on attacker-controlled content.
Pillar Security captured this with a line I have not stopped thinking about since I read it: "If an agent gets to write the future inputs of systems, it was never sandboxed in the first place."
I sat with that sentence for a long time. It reframes the entire security question. Not "can the agent escape." Not "is the sandbox configured correctly." The question is simpler and harder: what does the agent write, and who reads it next. That is not a sandbox question. It is a perimeter question.
The sandbox model was built for a world where processes were the threat and containment was the answer. Agents are not processes. They are authors. They write things other systems later read. Containment secures the process. It does nothing about what the process produces.
I recognised this pattern because I had already built something that caught it - but only at one boundary. The pre-code gate from Log #16 counts assertion deletions and refuses them at the commit boundary. It is the instrument. It catches what enters the pipeline. But it does not catch what never commits - the configuration file an agent writes that a trusted extension executes immediately, the task definition that a runner picks up before any commit happens, the virtualenv binary that a host process executes directly.
I built the instrument because I watched an agent delete a test assertion and a reviewer approve it. I know that boundary works. But I also know its limit. The gate fires when you commit. It does not fire when a file the agent wrote sits in a workspace directory and a VS Code extension reads it three seconds later. That is not a flaw in the gate. It is a different boundary entirely.
The instrument enforces at the code boundary. The trust-handoff flaw operates at a different boundary entirely.
What the instrument misses
Log #12 established the pre-code gate that Log #16 formalised into the instrument. Log #15 established the handshake: the reviewer and the author must read the same source of truth. Both are enforcement mechanisms. Both are pre-commit.
The trust-handoff flaw is post-commit. Or rather, it is pre-commit-adjacent - it happens in the space between the agent writing a file and the version control system seeing it. The file exists. A trusted component reads it. The instrument never fires because the instrument only inspects what is about to be committed.
This is not a flaw in the instrument. It is a limitation of its scope. The instrument enforces at one boundary. The trust-handoff flaw exploits a different boundary. The fix is not to make the instrument do more. The fix is to recognise that enforcement needs to happen at every boundary the agent's output crosses, not just the one you built a gate for.
A paper published in February 2026 made this argument in architectural terms. The authors proposed what they called Trinity Defense: three layers - action governance, information-flow control, and privilege separation - operating together as a reference monitor. The concept is straightforward: every time the agent asks the host to do something, a separate process says yes or no before the host acts. The monitor sits outside the agent's reach. The agent cannot modify it, cannot negotiate with it, and cannot write to any path that bypasses it.
A second paper, published in March 2026, described the same idea at the tool-call boundary. Before any command executes, an authorization check runs. Not a sandbox rule. A policy. Deterministic. The question is not "can the agent run this command." The question is "has this specific invocation been authorised given what the agent has written so far."
Neither paper described a product. Both described an architecture. And both described what the instrument cannot do alone.
Where the perimeter must sit
The perimeter is the enforcement layer at every trust transition. Not one gate. A reference monitor at every point where something the agent wrote crosses into something with more authority than the agent.
This is not a better sandbox. A better sandbox still trusts what leaves it. The perimeter does not trust. It checks. Every artifact. Every transition. Every time.
The practical shape of a perimeter has three parts. First, an inventory of every unsandboxed component the agent's output can reach - every IDE extension, every task runner, every daemon socket, every scripts directory on the path. If you cannot list them, you do not know where your trust handoffs happen.
Second, a policy at each handoff point. The policy does not ask "is this file type allowed." It asks "was the content of this file produced by the agent since the last human review." If the answer is yes, the policy refuses. The file sits. A human reads it. Only then does it cross the boundary.
Third, telemetry that fires when a trusted helper executes something the agent influenced. Not monitoring. Telemetry. If an extension executes a hook an agent wrote, you need to know immediately - not when the log rotates.
These three things together are the perimeter. They are not a product. They are an audit you run before the agent ships. The pattern is the same one Log #4 established with the circuit breaker: refusal, not observation. Don't detect the failure. Refuse the handoff.
The architecture below — agent, monitor, host — is the shape every trust transition needs.

I learned to think this way during those years in quality assurance. Every system I audited had a trust boundary between what it produced and what consumed its output. The systems that failed were never the ones with broken internals. They were the ones where nobody had asked: who reads this output next, and what authority do they grant it.
What this means for anyone deploying agents
I spent 14 years in quality assurance before I started building NiraNexus. The pattern I saw across every regulated system I worked on was the same: the most dangerous failure is not the component that breaks. It is the assumption that a component upstream already validated what it is about to consume.
The CSA disclosures made the same point about AI agents. The sandbox validated the agent's execution environment. Nothing validated what the agent produced. The vulnerability was not the container. It was the silence between containment and consumption.
If you are deploying AI coding agents, run the audit now. Ask every vendor the questions Pillar Security proposed: what can the agent write to disk? Which host components trust those writes without re-validation? Which commands skip approval? What telemetry fires when a trusted helper executes something the agent influenced?
If a vendor answers "our sandbox handles that," ask the next question: does your sandbox check what leaves it, or only what runs inside it. The CSA research proved these are different questions. Four products got the first one right. Every one got the second one wrong.
I have asked that second question in three vendor calls since July. Two of them paused. One sent me a security whitepaper describing the sandbox architecture in detail. Not one of them had an answer for the trust handoff.
The instrument catches what enters the pipeline. The perimeter catches what crosses the trust boundary. Log #18 will answer the question neither can: what stops an engineering team from bypassing both when the feature pressure is high enough.
Provenance
- Cloud Security Alliance: “AI Coding Agent Sandbox Escapes: The Trust Handoff Flaw”, Research Note, 22 July 2026. Documented seven vulnerability disclosures across Cursor, Codex CLI, Gemini CLI, and Antigravity.
- Pillar Security: “The Week of Sandbox Escapes”, 20 July 2026. Identified four repeatable failure modes and proposed the Inert-Artifact Illusion framework.
- Cursor CVE-2026-48124 (CVSS 8.5): hook configuration file written by agent, consumed by Cursor extension. Fixed in Cursor 3.0.0.
- OpenAI Codex CLI “GitPwned”: command allowlist bypass via flag injection. Fixed in Codex CLI 0.95.0. Disclosed in the CSA and Pillar Security reports above.
- Google Gemini CLI and Antigravity: multiple disclosures including .vscode/tasks.json, Python virtualenv tampering, and Docker socket exposure. Disclosed in the CSA and Pillar Security reports above. Google declined to fix two findings.
- Chen et al: “Trustworthy Agentic AI Requires Deterministic Architectural Boundaries”, arXiv:2602.09947, February 2026. Proposed Trinity Defense Architecture.
- Kumar et al: “Before the Tool Call: Deterministic Pre-Action Authorization for AI Agents”, arXiv:2603.20953, March 2026. Proposed Open Agent Passport.
- Log #4: “The Circuit Breaker” — refusal over observation, the original enforcement pattern.
- Log #12: “The Code Is the Policy” — pre-code gates, every gate earned by a fire.
- Log #15: “The Handshake” — shared source of truth between reviewer and author.
- Log #16: “The Instrument” — deterministic pre-commit gate for test assertion integrity.