The guardrail gate
Every agent run passes through four checks. The gate is never skipped — there is no bypass, no backdoor, no "just this once".
Check 1
Input filtering
The task is screened before it reaches the agent: blocked topics, prompt-injection patterns, and personal data get flagged or stopped cold.
Check 2
Output checking
The agent's output is screened before it goes anywhere: same filters, applied to what the agent produced.
Check 3
Human approval
Irreversible actions — sending, publishing, paying, deleting — pause for your explicit decision. Approve or deny, one tap.
Check 4 · Guardrail report. Every run produces an audit trail: what was checked, what was blocked, what you approved. Yours to keep.
Try the gate
A live demo with a sample research agent. Type a task and watch all four checks run. (Demo mode: the agent step is simulated — connect your key to run a real agent.)
Agent configuration (from the template builder)
Human approval required
Approve this action?
Documentation
What the gate catches
- Blocked topics — phrases from your agent's configuration ("medical advice", "financial advice") matched against every input and output.
- Prompt injection — known jailbreak phrases like "ignore previous instructions" stop the run immediately.
- Personal data — email addresses, phone numbers and similar patterns are flagged for your review before proceeding.
- Irreversible actions — anything matching your approval list, plus built-in verbs (send, publish, pay, delete…), pauses for your decision.
What it doesn't catch
Honest limits: client-side filtering is a baseline, not a proof. It won't catch novel jailbreak phrasing, subtle manipulation, or a compromised provider. The human-approval step exists precisely because filters are fallible — the final call is always yours.
The enforcement contract
Agent code runs through GuardedRunner — the single path from "agent wants
to act" to "action happens". It has no skip flag and no bypass. The sandbox and deploy
steps build on this contract.
FAQ
Can the gate be skipped?
No. GuardedRunner has no bypass — every run flows input check → agent → output check → approvals → report. Skipping isn't a feature that exists.
Where does the report go?
It's rendered for you after every run, and you can copy it as JSON. Reports live in your browser unless you save them.
What if my task is blocked by mistake?
Rephrase the task or narrow the agent's blocked topics in the template builder. The report tells you exactly which rule fired.
Does the gate call my API key?
The gate itself never calls your provider — it filters text locally. Your key is only used when the agent itself runs.