Skip to content

Sandboxes with agent SDKs

An agent that can run code needs somewhere safe to run it. This page wires a MicroVM sandbox into an agent’s tool-calling loop as a run_code (or run_shell) tool, for four widely used agent SDKs. Each example creates one sandbox, lets the agent run commands in it, prints the result, and cleans up.

  • A MicroVM API key. See API keys and the E2B-compatible API.
  • The sandbox-base image (or a custom image built from it) so the sandbox includes envd, the agent the E2B SDK talks to. See The interactive shell agent.
  • An API key for whichever model provider’s SDK you use (Anthropic, OpenAI, or a provider supported by the Vercel AI SDK or LangChain).

Every example in this page shares the same three environment variables:

Terminal window
export E2B_API_KEY=<your MicroVM API key>
export E2B_API_URL=https://<your panel host>/api/e2b
export E2B_DOMAIN=<your sandbox domain>

E2B_DOMAIN is required for these examples: commands.run(), file transfers and get_host() all reach the sandbox through envd on <port>-<sandbox id>.<your sandbox domain>, which only resolves once your location has a sandbox domain with wildcard DNS. Find yours on the API keys page or ask your provider.

Terminal window
curl -s https://<your panel host>/api/microvm/catalog \
-H "Authorization: Bearer <your account API token>" | jq '.locations[] | {id, name}'

Pick one id from the response and use it as location_id below. (hypervisor_group_id is still accepted as a deprecated alias for the same value.)

Claude’s tool use loop: describe the tool with a JSON Schema input_schema, send the conversation, and whenever Claude replies with a tool_use block, run the tool and send its result back as a tool_result message. This example gives Claude a single run_shell tool backed by one sandbox, reused across turns until the script kills it.

import anthropic
from e2b import Sandbox
client = anthropic.Anthropic()
sbx = Sandbox.create(
"sandbox-base",
metadata={"location_id": "<location id>"},
timeout=300,
)
tools = [{
"name": "run_shell",
"description": "Run a shell command inside a persistent Linux sandbox and return its stdout and stderr.",
"input_schema": {
"type": "object",
"properties": {"command": {"type": "string", "description": "The shell command to run."}},
"required": ["command"],
},
}]
def run_tool(command: str) -> str:
result = sbx.commands.run(command, timeout=60)
return f"stdout:\n{result.stdout}\nstderr:\n{result.stderr}"
messages = [{"role": "user", "content": "List the files in /tmp, then print the current date."}]
try:
while True:
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
tools=tools,
messages=messages,
)
messages.append({"role": "assistant", "content": response.content})
tool_uses = [b for b in response.content if b.type == "tool_use"]
if not tool_uses:
print(next(b.text for b in response.content if b.type == "text"))
break
messages.append({
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": block.id,
"content": run_tool(block.input["command"]),
}
for block in tool_uses
],
})
finally:
sbx.kill()

Source: Define tools, Handle tool calls.

The OpenAI Agents SDK turns a plain Python function into a tool with the @tool decorator from agents.decorators; the agent’s runtime handles the call loop for you.

from agents import Agent, Runner
from agents.decorators import tool
from e2b import Sandbox
sbx = Sandbox.create(
"sandbox-base",
metadata={"location_id": "<location id>"},
timeout=300,
)
@tool
def run_code(code: str) -> str:
"""Run a Python snippet inside a persistent Linux sandbox and return its output.
Args:
code: Python source to execute with `python3 -c`.
"""
result = sbx.commands.run(f"python3 -c {code!r}", timeout=60)
return result.stdout or result.stderr
agent = Agent(name="Code runner", tools=[run_code])
try:
result = Runner.run_sync(agent, "Compute the 20th Fibonacci number and print it.")
print(result.final_output)
finally:
sbx.kill()

Source: Function tools.

The AI SDK’s tool() takes a Zod input schema and an execute function; pass it to generateText and the SDK runs the tool loop.

import { generateText, tool, stepCountIs } from "ai";
import { z } from "zod";
import { Sandbox } from "e2b";
const sbx = await Sandbox.create("sandbox-base", {
metadata: { location_id: "<location id>" },
timeoutMs: 300_000,
});
const runShell = tool({
description: "Run a shell command inside a persistent Linux sandbox and return its stdout.",
inputSchema: z.object({ command: z.string() }),
execute: async ({ command }) => {
const result = await sbx.commands.run(command, { timeoutMs: 60_000 });
return result.stdout;
},
});
try {
const { text } = await generateText({
model: "anthropic/claude-opus-5",
tools: { runShell },
stopWhen: stepCountIs(5),
prompt: "List the files in /tmp, then print the current date.",
});
console.log(text);
} finally {
await sbx.kill();
}

Source: Tools.

LangChain’s @tool decorator wraps a function the same way; create_agent builds a tool-calling agent around a chat model and a tool list.

from langchain.tools import tool
from langchain.agents import create_agent
from langchain_anthropic import ChatAnthropic
from e2b import Sandbox
sbx = Sandbox.create(
"sandbox-base",
metadata={"location_id": "<location id>"},
timeout=300,
)
@tool
def run_shell(command: str) -> str:
"""Run a shell command inside a persistent Linux sandbox and return its stdout and stderr."""
result = sbx.commands.run(command, timeout=60)
return f"{result.stdout}\n{result.stderr}"
agent = create_agent(
ChatAnthropic(model="claude-opus-5"),
tools=[run_shell],
system_prompt="You can run shell commands in a sandbox to answer questions.",
)
try:
result = agent.invoke({"messages": [{"role": "user", "content": "List /tmp, then print the date."}]})
print(result["messages"][-1].content)
finally:
sbx.kill()

This same run_shell tool works unchanged as a LangGraph node function inside a custom graph; create_agent is LangGraph’s own prebuilt tool-calling agent under the hood.

Source: Tools.

Pi (@earendil-works/pi-coding-agent) is a different shape: an embeddable coding-agent runtime rather than a chat-completions tool loop. It exposes defineTool() for a custom tool and createAgentSession() to run an agent session with it.

import { Type } from "typebox";
import { createAgentSession, defineTool } from "@earendil-works/pi-coding-agent";
import { Sandbox } from "e2b";
const sbx = await Sandbox.create("sandbox-base", {
metadata: { location_id: "<location id>" },
timeoutMs: 300_000,
});
const runShell = defineTool({
name: "run_shell",
label: "Run shell",
description: "Run a shell command inside a persistent Linux sandbox.",
parameters: Type.Object({ command: Type.String({ description: "Command to run" }) }),
execute: async (_toolCallId, params) => {
const result = await sbx.commands.run(params.command, { timeoutMs: 60_000 });
return { content: [{ type: "text", text: result.stdout }], details: {} };
},
});
try {
const { session } = await createAgentSession({ customTools: [runShell] });
// Drive `session` per the Pi SDK docs, then read the final transcript.
} finally {
await sbx.kill();
}

Source: SDK.

An agent loop can run for an unpredictable amount of time, so bound the sandbox rather than the agent:

  • Pass timeout (Python) or timeoutMs (JavaScript) to Sandbox.create() as a hard cap; a sandbox that outlives it is paused or killed depending on how it was created (the E2B adapter defaults new sandboxes to pause-on-timeout after 15 seconds unless you set a longer one).
  • Call sbx.set_timeout(seconds) mid-run to extend a sandbox that is still doing useful work, instead of setting one very long timeout up front.
  • sbx.pause() freezes a sandbox between agent turns that are minutes apart (a long-running conversation, a human-in-the-loop step) and Sandbox.connect(sandbox_id) resumes it later, cheaper than leaving it running idle.
  • Always sbx.kill() in a finally block once the agent’s task is done. A killed sandbox stops billing immediately; one left running keeps billing until its timeout or idle-pause setting catches it.