The guardrail gate

Every agent run passes through four checks. The gate is never skipped — there is no bypass, no backdoor, no "just this once".

Check 1

Input filtering

The task is screened before it reaches the agent: blocked topics, prompt-injection patterns, and personal data get flagged or stopped cold.

Check 2

Output checking

The agent's output is screened before it goes anywhere: same filters, applied to what the agent produced.

Check 3

Human approval

Irreversible actions — sending, publishing, paying, deleting — pause for your explicit decision. Approve or deny, one tap.

Check 4 · Guardrail report. Every run produces an audit trail: what was checked, what was blocked, what you approved. Yours to keep.

Try the gate

A live demo with a sample research agent. Type a task and watch all four checks run. (Demo mode: the agent step is simulated — connect your key to run a real agent.)

Agent configuration (from the template builder)

Documentation

What the gate catches

What it doesn't catch

Honest limits: client-side filtering is a baseline, not a proof. It won't catch novel jailbreak phrasing, subtle manipulation, or a compromised provider. The human-approval step exists precisely because filters are fallible — the final call is always yours.

The enforcement contract

Agent code runs through GuardedRunner — the single path from "agent wants to act" to "action happens". It has no skip flag and no bypass. The sandbox and deploy steps build on this contract.

FAQ

Can the gate be skipped?

No. GuardedRunner has no bypass — every run flows input check → agent → output check → approvals → report. Skipping isn't a feature that exists.

Where does the report go?

It's rendered for you after every run, and you can copy it as JSON. Reports live in your browser unless you save them.

What if my task is blocked by mistake?

Rephrase the task or narrow the agent's blocked topics in the template builder. The report tells you exactly which rule fired.

Does the gate call my API key?

The gate itself never calls your provider — it filters text locally. Your key is only used when the agent itself runs.