The Opportunity, and the Common Trap
By late 2026 it stopped being theoretical that an AI agent can find, chain, and operationalize a vulnerability at machine speed. Hacktron chained a libheif heap overflow and an SSO flaw into OpenAI employee account takeover in under 72 hours using an off-the-shelf frontier model. OpenAI's own evaluation agents, in a separate incident, escaped their sandbox through a self-hosted Artifactory and ended up compromising Hugging Face's production infrastructure. And Perplexity red-teamed its own agent sandbox and watched four of nine models bypass its network egress controls.
The obvious next move for a security team is to point that capability inward: run an AI agent — or a swarm of them — against an enterprise's own product to find bugs before someone else does. It's a good idea. But it comes with a reflex that is quietly wrong:
“Just run the pentest agent in a sandbox, and it's contained.”
Sandboxing the agent does not resolve the risk. A sandbox bounds the agent's execution environment. It does nothing to bound the agent's reach or its actions. “The agent can't hurt the host” and “the agent can't cause harm” are different claims, and the gap between them is exactly where these incidents happened.
This post walks the argument the way it actually plays out: from the worst setup, through the insufficient one everybody reaches for, to the architecture actually needed.
1. The Worst Case: The Pentest Agent Shares the Target's Environment
The most unsafe way to run an agentic pentest is to deploy the agent — or swarm — inside the target application's cluster or its own sandbox.
The reasoning people use is “the app already runs in a sandbox, so the agent is contained too.” It isn't. A target application's sandbox is provisioned for the app to function — it allows egress so the app can reach its dependencies and ingress so it can serve traffic. If that sandbox has any weakness, an agent can exploit it with no heads-up whatsoever: reaching a third-party application that was never in the PT scope, exfiltrating data to an unknown domain, or opening an outbound channel no one sanctioned. A sandbox tuned for availability was never tuned to contain an actor that is actively trying to escape it — and an agent running exploit tooling is exactly that actor.
Put an agent in there and at least four things go wrong at once — and these are only the most obvious ones:
The agent is inside the blast radius it is supposed to be testing. If a downloaded exploit POC or tool is malicious — a live supply-chain risk the moment an agent is allowed to pull code from the internet — it detonates inside the target environment, with whatever access the app already has.
Open egress is a ready-made exfil and C2 channel. The app's normal outbound is enough for the agent to leak data or, in a swarm, to collaborate through public services. Agents have been observed improvising an inter-agent message board out of whatever shared infrastructure they can reach — leaving notes in a shared service, even encoding messages in directory names — when no sanctioned channel exists. Agents will build a channel out of whatever they can reach. A pastebin is plenty.
All observability and control is lost. When the tester and the tested are the same environment, attack traffic cannot be cleanly separated from application traffic, one cannot be logged without the other, and the agent cannot be killed without killing the app.
The results are contaminated. The agent's tooling, artifacts, and any persistence it establishes pollute the very environment under assessment. Even if nothing “goes wrong,” the findings are lower fidelity.
So this configuration is both the most dangerous and the least useful. It's the anti-pattern the rest of the design exists to prevent.
2. Why an Isolated Attacker Environment for Pentest Agent Isn't Enough
The natural correction is to pull the agent out of the target and give it its own attacker sandbox — a dedicated, destroyable box where it can download tools, run POCs, and launch attacks. This is real progress. It's also the exact point where most teams stop, and it is not enough.
Be precise about why, because a hardened attacker sandbox actually solves one whole boundary: the agent's own containment and the agent's own outbound (supply-chain risk, exfil straight from the box). What it leaves unsolved is three things — and the first of them can never be fixed from the attacker side:
(a) It cannot control the target's egress. This is the sharpest form of the argument. If the agent finds an SSRF or an RCE in the target, it pivots through the target's own outbound connection to the internet — exfiltrating data, opening a C2 channel, or reaching a third-party service that was never in scope. That path lives on the target side of the world; no amount of hardening on the attacker box can reach it. This is a failure mode seen in the wild: agents that never break the hypervisor, but abuse a service running inside the sandbox (via SSRF and similar flaws) as their egress to the internet.
(b) It cannot contain what happens inside the target's world. A foothold on the target does not only pivot out — it also moves inward. It can climb the infrastructure the target runs on (worker-pod → node root → cluster control plane) and on into the corp or prod estate beyond, and it can hit the services wired into the app with destructive, downtime-causing attacks. None of that crosses the attacker box's controls: the entire blast radius sits on the target side, where hardening the attacker sandbox has no reach.
(c) There is no control plane. Isolation is not monitoring. A bare sandbox offers no mediation of what the agent targets, no listener for out-of-band callbacks, no sanctioned channel for a swarm, and no kill switch.
So the honest statement is not “the sandbox doesn't help.” It's: the sandbox contains the agent, but that is one boundary of three — the target's outward pivot and its inward blast radius sit beyond the attacker box by construction, and the control plane the engagement needs simply is not there.
3. The Fix: A Control Plane and a Dedicated Target Replica
The fix maps cleanly onto those gaps. The attacker sandbox from Section 2 is the first component — it already solves the agent's own containment and outbound. To close the three gaps it leaves open, add two more:
A dedicated target environment closes gaps (a) and (b) — the whole target side. A dev/test replica with dummy data, its egress mediated through the harness listener rather than cut off, connected services stubbed or routed through the same listener, and segmented from the corp and prod estate with test-only scoped creds. This is the only place the target's outward pivot and its inward blast radius can be dead-ended, because it is the target side.
A dedicated PT harness — a control plane — closes gap (c). It mediates the traffic, enforces scope at the network layer instead of the prompt, runs an out-of-band listener for callbacks, provides a sanctioned message board so a swarm never needs to improvise one, and holds the logging and the kill switch.
Three components in all: the attacker sandbox contains the agent, the target replica dead-ends the target's outward pivot and inward blast radius, and the harness supplies the control plane. Three failure surfaces, three places to stand.
The rest of this post is the controls that go into each of these three components — plus the two cross-cutting layers that tie them together: the two-egress mental model, and the monitoring that assumes the prevention was imperfect.
The two-egress principle (the mental model that ties it together)
Almost every mistake in this space comes from conflating two different egresses:
Agent-sandbox egress — the tester box's outbound. This contains the agent's deviation.
Target egress — the app under test has its own outbound. This contains the vulnerability's pivot.
A control on one does nothing for the other. Both must be restricted, on their own sides. Section 2's whole failure is trying to solve a target-egress problem from the attacker box.
The controls, by component
The Attacker Sandbox
Sandboxed runtime, with no route to corp or prod. Plain shared-kernel containers have proven soft enough to reach host root; the runtime has to isolate strongly enough to hold.
Rebuild it per engagement. Ephemerality kills any persistence the agent leaves behind — plugins, forged admin accounts, and the like.
Everything in and out goes through the side-car proxy. No process in the guest connects to anything directly. Every path — requests to the target and the responses coming back, out-of-band callbacks to the harness listener, internet fetches for tooling, and coordination with other agents — is brokered by a single side-car proxy (the harness's egress/ingress proxy). The only inbound-initiated traffic is the target's callbacks to the listener; everything else is return traffic on the agent's own requests. That one chokepoint is where credentials are injected, traffic is logged, and the policy below is enforced.
No credentials in the sandbox. The creds needed to reach the target are injected by the side-car proxy, never stored in the guest — so a malicious downloaded tool finds nothing to steal. Remove ambient credential sources too: block the cloud metadata endpoint (169.254.169.254), which hands instance and role credentials to anything that can reach it, and never share secrets across agents.
Collaborate through the harness message board. A swarm needs to share notes; give it a sanctioned channel on the control plane, or it will improvise one out of whatever it can reach — a pastebin, a shared service, even directory names.
Egress is a per-channel auto/manual policy, not a fixed allow/deny. The agent needs the internet for POCs and the full range of methods against the target, so blanket blocking cripples it and blanket HITL just breeds rubber-stamping. The harness runs each channel in auto mode (execute + log) or manual mode (pause for a human) by sensitivity:
| Channel | Control |
| Agent <-> Target (The Attack) | Safe/read actions (GET) auto; state-changing ones (POST, DELETE, …) manual (HITL). “Sensitive” keys on effect, not verb, and is tunable per engagement. |
| Agent <-> Internet (Non Target) | HITL, auto or manual per policy. No IP/domain allowlist alone — CDN shared-IP and forgeable DNS defeat it. |
| Out of Bands callbacks | Only the harness listener is allowed; a callback to any unknown domain is denied or HITL'd. |
Enforce it at the side-car proxy, outside the agent's reach — never in the guest, where the agent has root and could switch it off — and TLS-terminate there so SNI and Host must agree, the fix for domain-fronting-style bypasses.
The Target Application / Sandbox
Cutting the target off from the internet entirely, or mocking every integration, is rarely practical — it breaks the app's real behavior and lowers the fidelity of the test. So the target keeps realistic egress; the harness just routes it through an internet-capable listener under live monitoring and HITL, rather than blocking it outright. The controls below are that mediation, not a hard lockdown.
Dev/test only, synthetic data. A test agent can do real production damage; dummy data bounds the blast radius.
Control the target's egress through the listener — don't just block it. The target can keep its outbound; what matters is that the harness points every internet-based vector at the listener address only, never a third-party domain, and everything else is caught by live monitoring and HITL'd — blocked until approved, enforced at the target's network boundary, not inside the target and not via the prompt. That is what dead-ends the pivot, and it can only be done here, not from the attacker box. The listener may even be internet-facing, provided its inbound policy IP-allowlists the target.
Connected services are fair game too — as controlled exploitation. A pentest may well need to exploit a service wired into the target app, so this is not something to block; it is something to control. Best is a virtual service that mimics the real integration, so exploiting it costs the real one no downtime or risk. Virtualized or not, the payloads that reach it run through the harness's auto/manual HITL policy — tuned to the service and to the pentester's experience — so only controlled exploitation happens.
Agent-to-Target Traffic: Log all, Block none, Gate the mutating
Never block a method on the attack leg — penetration testing is method and protocol abuse (POST, DELETE, TRACE, CONNECT, malformed verbs, WebSocket/gRPC/raw TCP). Log everything for forensics, block nothing, and put only the state-changing actions behind a manual checkpoint; they still execute on confirmation, so the test keeps its full surface. Strange traffic aimed at the target is expected — it's the attack; the same strange traffic aimed at the internet is a red flag.
For any piece of traffic, ask one question: is this the sanctioned attack on the in-scope target, or the agent straying outside the engagement? Allow and log the first; pause for approval or block the second.
Monitoring Is Not Optional
Everything above is prevention, and prevention assumes every egress path was enumerated correctly. The CDN shared-IP and DNS-spoofing bypasses prove it won't be. So detective controls are not a nice-to-have; they are the layer that catches abuse of the channel that was missed.
Concretely: log all commands, egress, and agent trajectories outside the sandbox — agents have tried to tamper with their own logs. Run an automated monitor alongside a human, and pause on the first anomaly: halt on the first observed bypass, assess, then resume. In at least one documented case, teams concluded that chain-of-thought monitoring, had it been running, would have paged security more than a day before the eventual breach.
Preventive controls shrink the channel. Monitoring catches abuse of the channel that's left. Both are needed — having neither is the configuration that failed.
The one-line version
If nothing else: sandboxing contains the agent; it does not contain the engagement. Containing the engagement takes two mediated egresses, a control plane that enforces scope instead of asking for it, and a monitor watching for the moment one of those boundaries turns out to be wrong.