Why buy a sandbox when you can just launch a container? Darren Shepherd spent years at Rancher building on Kubernetes, so sandboxes looked dumb to him. Then he built his own coding agent to run 15 things in parallel without work-tree chaos, and the most interesting part of it turned out to be a sandboxing system. He had proved himself wrong.
The sharp part is where the sandbox boundary goes. The common pattern keeps the agent loop outside and sandboxes only its tool calls. Darren calls that completely flawed, because the loop itself directs code that holds secrets and talks to external systems. His answer: put the agent and its tools in one sandbox, with one policy for secrets and egress.
Darren is on X and GitHub at @ibuildthecloud. https://obot.ai/
Where does your boundary sit: around the agent itself, or only around the tools it calls?
Why buy a sandbox when you can just launch a container? Darren Shepherd spent years at Rancher building on Kubernetes, so sandboxes looked dumb to him. Then he built his own coding agent to run 15 things in parallel without work-tree chaos, and the most interesting part of it turned out to be a sandboxing system. He had proved himself wrong.
The sharp part is where the sandbox boundary goes. The common pattern keeps the agent loop outside and sandboxes only its tool calls. Darren calls that completely flawed, because the loop itself directs code that holds secrets and talks to external systems. His answer: put the agent and its tools in one sandbox, with one policy for secrets and egress.
Darren is on X and GitHub at @ibuildthecloud. https://obot.ai/
Where does your boundary sit: around the agent itself, or only around the tools it calls?
Show notes, transcript, and links: https://www.wordman.dev/podcast/darren-shepherd-ai-agent-sandboxes/
Key Takeaways
A container is not a sandbox. Sandboxes sit one layer up: create, clone, copy files in and out, on a base image that mostly just carries language runtimes.Parallel agents need isolation. Running many tasks at once without Git work-tree chaos means giving each task its own sandbox.Egress is the policy surface, not ingress. Agents sit on the consuming side of services, so outbound traffic is what you control.Keep real secrets out of reach. Hand the agent a fake key with the right shape and swap in the real one at a transparent proxy, where policy and exfiltration checks live.Sandbox the agent loop, not only the tool calls. The loop directs code that holds secrets and talks to external systems, so the agent and its tools belong in one sandbox with one policy.Client-side agent loops beat centralized ones. Server-side provider tools such as web search open exfiltration paths that are hard to reason about.MCP's lasting value is the interface, not the protocol: a contract designed for AI, which is why it fits the enterprise.Output is not progress. Letting a model barf out thousands of lines feels productive until the regressions pile up, and one estimate raised in the conversation puts a skilled engineer's real gain at around 5 to 10 percent.Chapters
00:00:00 Why sandboxes felt dumb00:01:38 Building a coding agent00:02:36 Parallelism and work-tree chaos00:03:25 What a sandbox actually is00:04:27 Files, secrets, and egress00:08:20 Tool-call-only sandboxing is flawed00:12:28 Put the agent in the sandbox00:13:31 Client-side vs centralized loops00:14:51 GPTScript, Clio, and MCP00:20:00 MCP, OAuth, and staying in lane00:30:01 Without fundamentals, AI gets dangerous00:35:00 The 100x developer myth00:40:00 AI is not a compiler00:45:01 Distributed systems instincts00:51:09 ADRs and keeping quality00:55:04 Regression holes with agents00:57:19 Model personalities and greenfield work01:01:50 Skills as behavioral scripts01:05:16 Plan-first and spec-driven development01:12:30 Claude Code and tooling shifts01:21:16 Wrap and finding joy with AIMentioned
Obot AIObot on GitHubModel Context ProtocolGPTScript (archived)Clio (archived)RancherKubernetesClaude CodeCodex CLIArchitecture Decision RecordsEP. 22 with Dillon Mulroy