The sandbox

Test your agent before anything real happens. In the sandbox, your agent runs against simulated tools — a mock web search, a mock email sender, a mock publisher. Nothing leaves this page. No emails sent, nothing published, no money moved. Ever.

Every sandbox run goes through the guardrail gate: input filtering, output checking, human approvals, and a full report.

Run a sandbox test

Agent configuration (from the template builder)

Pick a scenario:

Documentation

How a sandbox run works

  1. Input check. Your task is filtered before the agent sees it.
  2. Think → act → observe. The agent reasons, calls a simulated tool, and reads the mock result. Each cycle is shown in the trace.
  3. Approvals. When the agent reaches for an irreversible action, the run pauses and you decide — against the mock tool, so it's safe to say yes.
  4. Report. The guardrail report records everything: checks, actions, your decisions, the verdict.

The mock tools

Mock tools have no network access and no side effects. They exist so you can watch your agent's judgment — what it tries to do, and when the gate steps in.

When is my agent ready for real deployment?

When sandbox runs consistently end in Allowed with no surprises in the trace: no blocked inputs you didn't expect, no approval requests you didn't intend. The sandbox is where you learn what your agent wants to do.

FAQ

Does the sandbox cost anything?

No. The sandbox agent is simulated — it uses no API key and no inference. It's free forever.

Can anything real happen in the sandbox?

No. All tools are mocks with no network access. Approving an action runs the mock, never the real thing.

What's the difference between the sandbox and the guard demo?

The guard demo exercises the four checks in isolation. The sandbox runs a full agent loop — think, act, observe — with the gate enforced around every step.

Do I need my API key for the sandbox?

No. You'll need it when you deploy for real — that's the next step after the sandbox.