Skip to content

Blog

AI Agent Infrastructure: Where to Run Agents in Production

By Hakan Karaduman5 min read
AI Agent Infrastructure

TL;DR

Run production AI agents on infrastructure with hardware-level isolation, sub-second sandbox restarts, and a compliance posture that already covers PCI DSS, HIPAA, and GDPR, so a single misbehaving agent can’t touch another workload or your audit trail. Control Plane runs every agent workload in a Kata Containers sandbox on a per-workload Firecracker microVM, with Capacity AI packing resources and scaling workloads dynamically to cut compute cost 30-50%.

Platform Comparison

PlatformIsolationMulti-cloudComplianceScale-to-zeroPricing
Control PlaneKata Containers on per-workload Firecracker microVMsNative on AWS, GCP, Azure; OCI, bare metal, and Kubernetes via bring-your-own-infrastructurePCI DSS Level 1, SOC 2 Type II, HIPAA, GDPRDynamic scaling with Capacity AI resource packing (30-50% cost savings)Usage-based, scales with actual resource consumption
E2BFirecracker microVMs, one per sandboxManaged service on GCP; enterprise BYOC on AWS or GCP (no confirmed Azure support)SOC 2 Type II, HIPAA (BAA on request)Manual pause-on-idle (sbx.pause()), not automatic suspendBilled for active sandbox runtime; paused sandboxes are free
ModalgVisor-sandboxed containers; serverless GPU execution modelMulti-region; multi-cloud provider choice not established in public docsSOC 2 Type II, HIPAA (BAA required for PHI)Automatic scale to zero on no request volumeConsumption-based, per-second CPU/memory/GPU metering
NorthflankKata Containers on Cloud Hypervisor microVMs (gVisor fallback where nested virtualization is unavailable)BYOC across AWS, GCP, Azure, OCI, bare metal, on-premSOC 2 Type II, HIPAA; ISO 27001, PCI DSS, and FedRAMP referenced in enterprise materialsNot uniformly automatic; depends on workload type and configurationUsage-based, per-second billing (~$0.01667/vCPU-hr, ~$0.00833/GB-hr) plus BYOC infra cost
DaytonaContainer and VM sandboxes (Linux/Windows); microVM or gVisor use not establishedShared and dedicated regions; BYOC reportedSOC 2, HIPAA, ISO/IEC 27001 (per trust center)Auto-stop after 15 minutes of inactivity by default (configurable)Usage-based, per-second (~$0.0504/vCPU-hr, ~$0.0162/GiB-hr memory)

How It Works: The Sandbox Lifecycle

Every agent workload on Control Plane moves through the same five stages, whether it’s a one-off code execution or a long-running agent process.

1. Provision The agent (or the human operator) requests a workload. Control Plane schedules it as a container inside a dedicated Kata Containers sandbox, backed by its own Firecracker microVM. No shared kernel, no shared memory space with other workloads.

2. Execute The workload runs with full isolation guarantees. Every action the agent takes is recorded in a shared audit trail, so humans and other agents can see exactly what changed and why.

3. Sleep When a workload goes idle, Control Plane suspends it automatically. This is a genuine differentiator: idle workloads stop consuming billable compute without requiring the agent or operator to manage the lifecycle manually.

4. Restart When new work arrives, the sandbox restarts in sub-second time. Agents don’t wait on cold-start penalties measured in seconds, which matters when an agent is orchestrating dozens of short-lived tasks in sequence.

5. Terminate When the workload is done, Control Plane tears down the sandbox and reclaims resources. The audit trail persists independent of the sandbox itself.

For agents that need to reach a private VPC or an on-prem service during execution, Control Plane uses Cloud Wormhole to establish that connection directly, without a VPN and without installing a client agent on either end.

Worked Example: MCP-Connected Agent Deploying Across Two Regions

An agent connected to Control Plane through MCP (Model Context Protocol) is asked to deploy a standard workload to two regions for latency and redundancy. Here’s what happens:

  1. Context gathering. The agent reads the current infrastructure state through MCP: existing services, network topology, and identity configuration. This is infrastructure as context, not infrastructure as code: the agent works from a live model of what’s running, not a static template.
  2. Plan generation. The agent proposes a deterministic change: deploy the workload as a standard container type in us-east and eu-west, using Universal Cloud Identity to authenticate against the target AWS and GCP accounts without long-lived credentials.
  3. Human checkpoint. The proposed plan and its diff are written to the shared audit trail. An engineer reviews the plan, confirms resource sizing, and approves.
  4. Provisioning. Control Plane provisions two Kata Containers sandboxes, one per region, each on its own Firecracker microVM. Capacity AI packs each workload against available capacity to hold costs down.
  5. Network access. One region needs to reach a database sitting in a customer’s on-prem data center. Control Plane opens that path through Cloud Wormhole, no VPN tunnel or client agent required.
  6. Verification. The agent runs health checks against both regions and reports status back through MCP. The full sequence, from plan to approval to deployment, is recorded in the audit trail for later review.
  7. Steady state. Both workloads scale dynamically with traffic. When either region goes idle, its sandbox sleeps; when traffic returns, it restarts in sub-second time.

This is the operating model Control Plane is built around: operated by humans, enabled by agents. Agents propose and execute deterministic changes; humans set policy and review the trail.

FAQ

Q: What is an AI agent sandbox?

A: An AI agent sandbox is an isolated execution environment where an agent can run code, call tools, or manage infrastructure without direct access to the host system or to other workloads. On Control Plane, each sandbox runs as a container inside its own Kata Containers boundary on a dedicated Firecracker microVM, so an agent’s execution is isolated at the hardware level, not just the process level.

Q: How do you secure AI agent execution?

A: Control Plane secures agent execution through per-workload Firecracker microVMs managed by Kata Containers, so each agent’s workload has its own kernel boundary and no shared memory or filesystem with other tenants. Every action an agent takes is written to a shared audit trail, and identity is handled through Universal Cloud Identity rather than static credentials, so access to AWS, GCP, or Azure resources like S3, BigQuery, Cosmos DB, or RDS is scoped and traceable.

Q: Can AI agents access private networks?

A: Yes. Agents that need to reach a private VPC or an on-prem network do so through Cloud Wormhole, which establishes that connection directly without a VPN or a client agent installed on either side. This keeps network access auditable and removes a common source of manual setup when agents need to reach resources outside the public cloud.

Next Steps

Cloud shaped around your workloads.