Guardrails

Definition
The rules and checks that keep an AI agent inside safe bounds, what it can say, what it can do, and where it must stop and ask.

Why it matters

A model is capable and occasionally wrong. The failure to fear is rarely a wrong answer sitting in a chat window. It is a confident wrong action with real consequences: a refund paid out, a customer emailed the wrong price, a record overwritten.

Guardrails cap the downside. They let you give an agent real work without betting the business on every single output.

They are also what makes more autonomy possible. A team that can show an agent has no way to move money, and that every customer-facing message passes a check, can safely widen what the agent does. A team without guardrails has to keep the agent small, or accept the risk.

Build them before the agent gets real reach into the business. Added afterwards, they are a repair, not a safeguard.

How to apply it

  • Start every agent with read-only access and add each write action one at a time as trust is earned.
  • Put a human approval step on anything that touches money, a customer directly or a record that is hard to undo.
  • Write an explicit refusal list of requests the agent should decline.
  • Check both directions: what the agent is asked and what it produces.
  • Try to break the guardrails yourself before an edge case finds the gap.

What it is

Guardrails are controls placed around an AI agent so that a mistake stays small. They fall into three layers. Access limits decide what the agent can reach: read-only by default, write access added deliberately. Approval steps decide where a person signs off, such as before money moves or a customer is contacted. Checks look at what goes in and what comes out, for example blocking a request to disclose private data, or rejecting a reply that quotes a price not found in the price list.

An instruction written in the prompt ("never issue refunds") is a request, and a model can ignore it. A permission enforced in software is a rule: if the agent's account cannot issue refunds, it cannot, whatever it is told. Strong guardrails use both, but rely on the second.

Common mistakes

  • Relying on the prompt alone. "Never issue refunds" in the instructions is a request. If the agent's account can issue a refund, one odd conversation can still produce one.
  • Guarding only the output. An agent that reads a malicious email or web page can be steered by it. Limit what it can reach as well as what it can say.
  • Starting with write access everywhere. Granting broad permissions "to save time" removes the main protection. Add each write action one at a time.
  • Never testing them. A rule nobody has tried to break is an assumption. Run a few hostile and awkward requests before launch and after every change.
  • Setting them once. The agent, its tools and the business change. Review the list each quarter, and after any incident.
Worked example

Suppose a twelve-person services firm lets an AI agent draft replies to inbound enquiries. Its first version quoted a day rate that was not on the price list, in one of ten test replies, which is exactly the kind of error guardrails exist for. The team applies three layers. The agent gets read-only access to the price list and cannot send anything itself. A person approves every reply that mentions money. A check blocks any quoted price that does not match the list. The workflow is built in n8n, so the approval step sits between the draft and the outbox. Over two weeks of testing, no unchecked price reached a customer.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    Human-in-the-loop

    The guardrail of keeping a person approving certain steps.

  2. Article

    Autonomous agent

    The kind of agent that most needs guardrails, since it acts without approval by default.

  3. Article

    Hallucination

    One of the failures guardrails are built to catch before it reaches a customer.