Guardrails
Why it matters
A model is capable and occasionally wrong. The failure to fear is rarely a wrong answer sitting in a chat window. It is a confident wrong action with real consequences: a refund paid out, a customer emailed the wrong price, a record overwritten.
Guardrails cap the downside. They let you give an agent real work without betting the business on every single output.
They are also what makes more autonomy possible. A team that can show an agent has no way to move money, and that every customer-facing message passes a check, can safely widen what the agent does. A team without guardrails has to keep the agent small, or accept the risk.
Build them before the agent gets real reach into the business. Added afterwards, they are a repair, not a safeguard.
How to apply it
- Start every agent with read-only access and add each write action one at a time as trust is earned.
- Put a human approval step on anything that touches money, a customer directly or a record that is hard to undo.
- Write an explicit refusal list of requests the agent should decline.
- Check both directions: what the agent is asked and what it produces.
- Try to break the guardrails yourself before an edge case finds the gap.
What it is
Guardrails are controls placed around an AI agent so that a mistake stays small. They fall into three layers. Access limits decide what the agent can reach: read-only by default, write access added deliberately. Approval steps decide where a person signs off, such as before money moves or a customer is contacted. Checks look at what goes in and what comes out, for example blocking a request to disclose private data, or rejecting a reply that quotes a price not found in the price list.
An instruction written in the prompt ("never issue refunds") is a request, and a model can ignore it. A permission enforced in software is a rule: if the agent's account cannot issue refunds, it cannot, whatever it is told. Strong guardrails use both, but rely on the second.
Common mistakes
- Relying on the prompt alone. "Never issue refunds" in the instructions is a request. If the agent's account can issue a refund, one odd conversation can still produce one.
- Guarding only the output. An agent that reads a malicious email or web page can be steered by it. Limit what it can reach as well as what it can say.
- Starting with write access everywhere. Granting broad permissions "to save time" removes the main protection. Add each write action one at a time.
- Never testing them. A rule nobody has tried to break is an assumption. Run a few hostile and awkward requests before launch and after every change.
- Setting them once. The agent, its tools and the business change. Review the list each quarter, and after any incident.