Skip to content

Blog

Dumb vs. Smart Sandboxes: Why AI Coding Agents Need More Than Isolation

By Doron Grinstein11 min read
Dumb vs Smart Sandboxes

Isolation keeps an agent’s code away from your infrastructure. A smart sandbox also lets the agent into the systems where real bugs live, without handing it the keys.

TL;DR

  • A dumb sandbox isolates an AI coding agent’s code. Connecting it to real systems is left to you, usually with long-lived credentials.
  • A smart sandbox adds two more things: a scoped workload identity that gets short-lived cloud credentials automatically, and private network reach to specific hosts and ports.
  • Control Plane Sandboxes are smart sandboxes. Each runs in a dedicated Firecracker microVM, and one identity object grants AWS, GCP, and Azure permissions plus private-network access. The agent’s code uses standard SDKs with no stored keys.

You isolated your coding agent. Now how does it reach your database?

Isolation keeps an AI agent’s code away from your infrastructure. But real engineering work happens inside that infrastructure. Debugging a slow production query, tracing a broken data pipeline, or calling an internal API all require the agent to get in.

The hard part isn’t running the agent’s code safely. It’s letting the agent in without handing it the keys.

What is a dumb sandbox?

A dumb sandbox is an isolated execution environment, usually a container or microVM, built to run untrusted AI-generated code safely. It gives the agent a filesystem, a shell, and a place to install packages and run tests without touching the host.

That isolation matters. “Dumb” doesn’t mean insecure. It means the sandbox solves one problem, execution, and leaves the next two to you: who the agent is when it calls a cloud API, and how it reaches a private database or service.

What is a smart sandbox?

A smart sandbox is an isolated execution environment that also carries a scoped workload identity and controlled private-network access. The agent can call real cloud services and private hosts with least-privilege permissions, and no long-lived credentials are stored in the sandbox.

A dumb sandbox decides where an agent can run. A smart sandbox also decides what it can safely reach.

Why isolation alone breaks down for AI coding agents

Isolation breaks down the moment an agent’s task depends on a real system. Three everyday tasks show it:

  • “Why did checkout queries get slow?” The agent needs the production database, which lives in a private VPC.
  • “Fix the nightly pipeline.” The agent needs to read the S3 bucket or BigQuery table the pipeline writes to.
  • “Why is the orders API returning 502s?” The agent needs an internal service that has no public endpoint.

Teams usually patch the gap in one of two ways. Both have costs.

Workaround 1: hand the agent credentials. An AWS access key, a database password, or an API token gets injected as an environment variable, a mounted file, or worse, pasted into the prompt. Now any code the agent runs can read that key, and so can anything that compromises the agent. Isolation protects the host. It does nothing for the S3 bucket that key unlocks.

Workaround 2: give it mocks. Copied databases and stubbed services are useful for tests, but production bugs live in production state, configuration, and traffic. An agent can pass every test against a mock and still miss the real cause.

The real requirement is narrower and harder: give the agent access to exactly the resources a task needs, for as long as it needs them, and nothing more.

The three pillars of a smart sandbox: isolation, identity, reach

A smart sandbox needs three controls, and each protects a different boundary. Isolation protects your hosts from the agent. Identity limits which cloud resources it can use. Reach limits which private endpoints it can connect to.

  • Isolation: A dedicated Firecracker microVM contains everything the agent runs.
  • Identity: Short-lived cloud credentials for AWS, GCP, and Azure, with no stored keys.
  • Reach: Private hosts by name, on specific ports only, over an outbound-only tunnel.

1. Isolation: contain what the agent runs

Every Control Plane Sandbox runs in its own dedicated Firecracker microVM, the virtualization technology AWS built for Lambda and Fargate. Each microVM has its own guest kernel, so agent code is separated from the host and from other workloads by a hardware-virtualization boundary, not just a shared kernel. Workloads can’t talk to each other by default: firewalls, mutual TLS client certificates, and proxies enforce that.

The agent gets a full Linux environment with a browser IDE, terminal, and SSH. It can install dependencies and run builds freely, inside its own boundary.

2. Identity: cloud access with no stored keys

Each sandbox can be assigned one identity, and that identity is its entire credential surface. With Universal Cloud Identity, the identity carries the cloud permissions for AWS, GCP, and Azure, at most one account per provider.

600+ cloud services, one identity. Reach services across AWS, GCP, and Azure with short-lived credentials and no stored keys.

Here’s what happens under the hood:

  1. Control Plane creates a principal in your own cloud account: an IAM role on AWS, a service account on GCP, or a managed identity on Azure.
  2. Inside the sandbox, Control Plane serves each provider’s standard instance-metadata endpoint.
  3. Unmodified SDKs and CLIs (aws, gcloud, az, boto3, and so on) find short-lived credentials exactly where they already look.

No access key goes into the prompt, an environment variable, or the container image. Widening or narrowing what the agent can do means editing the identity, not rebuilding the sandbox.

3. Reach: private hosts by name, specific ports only

Cloud Wormhole connects a sandbox to TCP and UDP endpoints in private networks: AWS, GCP, or Azure VPCs, on-prem data centers, even a developer laptop.

You run a small wormhole agent (a lightweight VM, not an AI agent) inside the private network. It opens an outbound-only tunnel to Control Plane, so nothing in that network has to be exposed to the internet. On the identity, you name the host, the agent that can see it, and the ports to open. The sandbox then dials that hostname as if it were local, on those ports only.

For services that should never leave a cloud provider’s network, the same identity can use AWS PrivateLink or GCP Private Service Connect instead.

Reachability is not authorization. The wormhole opens a network path. The database still checks its own user, password, and grants. That’s a feature: you can give the agent a read-only database role and a single port, and both layers have to agree.

Example: an AI agent debugging a slow production query

Here’s a full smart-sandbox setup for one realistic task. The database sits in a private VPC, and the agent also needs read access to CloudWatch logs.

The task: “Find out why checkout queries got slow after the last deploy.”

Step 1: Define one identity for the task

This YAML grants read-only AWS access and opens exactly one private host on exactly one port.

checkout-debugger.yamlname: checkout-debugger description: Read-only access for diagnosing checkout latency # Universal Cloud Identity: short-lived AWS credentials, no stored keys aws: cloudAccountLink: /org/acme/cloudaccount/aws-prod policyRefs: - aws::arn:aws:iam::123456789012:policy/CloudWatchLogsReadOnly # Cloud Wormhole: one private host, one port networkResources: - name: checkout-db agentLink: /org/acme/agent/prod-vpc IPs: ["10.0.1.100"] ports: [5432]

Step 2: Apply it and attach it to the sandbox

cpln apply -f checkout-debugger.yaml --gvc sandboxes

Then select checkout-debugger as the sandbox’s identity when you create it.

Step 3: Let the agent work with ordinary tools

Inside the sandbox, nothing needs configuring:# AWS CLI finds short-lived credentials automatically aws logs filter-log-events --log-group-name /ecs/checkout --filter-pattern "slow query" # The private database resolves by name, over the wormhole psql "host=checkout-db port=5432 user=agent_readonly dbname=orders"

The agent can now run EXPLAIN ANALYZE against real data, compare it with recent deploys, and propose a fix. Three layers bound what it can do:

  • Cloud: Read CloudWatch logs, nothing else.
  • Network: One host, port 5432, no other route into the VPC.
  • Database: agent_readonly can SELECT, not UPDATE or DROP.

No key was distributed, the database was never exposed to the internet, and revoking access means deleting one identity.

Dumb sandbox vs. smart sandbox: side-by-side comparison

The difference isn’t whether a sandbox can reach external systems. Most can, with enough glue. It’s whether identity and private connectivity are built in or left as separate projects.

Dumb sandbox vs. Control Plane smart sandbox

CapabilityDumb sandboxControl Plane smart sandbox
Isolated code executionYesYes
Isolation boundaryVaries: shared-kernel container or microVMDedicated Firecracker microVM
Cloud authenticationKeys injected as env vars or filesShort-lived credentials via workload identity
Long-lived keys in the sandboxUsuallyNone for AWS, GCP, Azure
Multi-cloud accessOne integration per providerOne identity covers all three, 600+ services
Private databases and APIsVPN, peering, or public exposureWormhole: named host, specific ports, outbound-only tunnel
Scoping accessSpread across IAM, VPN, and secret managerOne identity object
Revoking accessRotate every key that was handed outEdit or delete the identity
Execution time limitOften capped per taskNo execution timeout
Best forGenerating code, unit tests, local filesDebugging and operating real infrastructure

A dumb sandbox is the right tool when the agent only needs code and a test runner. Once it needs your databases, buckets, and internal services, identity and reach become the hard part.

How Control Plane Sandboxes work

A Control Plane Sandbox is a ready-to-code Linux environment in a dedicated Firecracker microVM. It runs as a standard Control Plane workload, so it inherits the platform’s identity, networking, and security model.

  • Create it your way: Provision sandboxes from the console, the cpln CLI, the API, or infrastructure as code, so agent pipelines can spin up environments on demand.
  • Toolbox images: Build an image from a template by picking runtimes, packages, and AI tools, or bring your own OCI image. Set an org default so every sandbox starts the same.
  • Right-sized compute: Choose from 2 to 16+ CPUs using size profiles your admins define.
  • No execution timeout: Long-running agent tasks keep going for as long as they need, within your resource limits.
  • Sleep when idle, wake fast: Idle sandboxes sleep automatically, so you pay for storage only, and they wake up quickly when work resumes.
  • Persistent state: /root lives on a persistent volume, so cloned repos, git credentials, and IDE settings survive sleep and restarts.
  • Connect any way: Use the browser IDE or terminal with zero installs, or run cpln sandbox connect to open VS Code, Cursor, or SSH.
  • Org-wide guardrails: Admins set default environment variables, map secrets to environment variables, and choose the default identity.

The Control Plane MCP Server and AI plugin let Claude Code, Codex, and other agents drive the platform directly.

Frequently asked questions

What is the best sandbox for AI coding agents?

It depends on the job. For generating code and running unit tests, any well-isolated sandbox works. For agents that debug or operate real systems, pick a smart sandbox that combines isolation, workload identity, and private network access, so the agent can reach real resources without long-lived keys.

How do I give an AI agent access to AWS without access keys?

Use workload identity. Assign the sandbox an identity tied to an IAM role, and the AWS SDK picks up short-lived credentials from the standard instance-metadata endpoint. On Control Plane, the same identity can hold GCP and Azure permissions too.

Can an AI agent connect to a private database without a VPN?

Yes. With Control Plane Cloud Wormhole, a wormhole agent inside your private network opens an outbound-only tunnel. The sandbox reaches the database by hostname on the ports you allow, and the database stays off the public internet. The database still enforces its own login and permissions.

Is it safe to let an AI coding agent access production?

It can be, with layered limits. Grant read-only cloud permissions, open only the specific hosts and ports the task needs, and use a read-only database role. Short-lived credentials shrink the exposure window, but they’re still usable while valid, so scope them tightly.

Does isolation still matter if the sandbox has cloud access?

Yes. The three controls protect different boundaries. Isolation, a dedicated Firecracker microVM on Control Plane, protects your hosts and other workloads from the agent’s code. Identity limits which cloud resources it can use. Network controls limit which endpoints it can reach.

What’s the difference between a sandbox and a smart sandbox?

A sandbox isolates code execution. A smart sandbox adds scoped workload identity and controlled private-network access, so the agent can work against real infrastructure under least privilege.

Which clouds do Control Plane Sandboxes support?

Universal Cloud Identity supports AWS, GCP, and Azure, with one account per provider per identity. Cloud Wormhole reaches TCP and UDP endpoints in any private network, including on-prem data centers.

How do I create a Control Plane Sandbox?

Use the console, the cpln CLI, the API, or infrastructure as code. Toolbox images standardize the tooling every sandbox starts with, and a new sandbox is ready in about 60 to 90 seconds.

Do Control Plane Sandboxes have a time limit?

No. There’s no execution timeout, so long-running agent tasks can run as long as they need, within your resource limits. Idle sandboxes sleep automatically and wake up quickly, with repos and settings intact.

How much compute does a sandbox get?

From 2 to 16+ CPUs, chosen from size profiles your admins define.

Your agent can write the fix. Can it reach the problem?

Give it a safe way into the systems where real bugs live: one identity, short-lived credentials, and exactly the hosts and ports you choose.

Create a sandbox (ready in about 90 seconds)

Further reading