Guardrails are the rules, limits and checks placed around an AI tool so that it stays within what a business has decided it may say and do. They include written instructions, restrictions on which data and systems the tool can reach, automatic filters on what it produces, and points where a person has to approve its work.
How guardrails work
No single control does the job. Guardrails work in layers, and each layer catches problems the others miss.
- What goes in. Rules about what staff may paste into a tool, such as no customer records or contract terms in a consumer chatbot, and filters that strip personal data before a request is sent.
- Instructions. A system prompt that sets tone, topics to refuse, claims never to make and when to hand over to a person.
- Permissions. What an AI agent can touch. It might read the CRM but not edit it, draft emails but not send them, or suggest budget changes but not apply them.
- What comes out. Checks on the output before anyone sees it: banned phrases, regulated claims, personal data, links to unknown sites, or answers that do not match an approved source.
- Approval and records. A human in the loop for anything public, costly or hard to undo, plus logs showing what the tool did and who signed it off.
Take a private dental practice in London adding a chatbot to its website. Sensible guardrails would limit answers to the practice’s own price list, opening hours and booking process; refuse clinical questions and pass them to reception; never describe treatment outcomes; and never ask for medical history in the chat window. Each of those is a separate control, and the chatbot is only as safe as the weakest one.
Why it matters
When an AI tool speaks in your name, what it says is your responsibility. A chatbot that promises a discount, a social post that makes a health claim or an email that quotes a price you do not offer is treated exactly as if a member of staff had written it. Under the CAP Code you need evidence for objective claims, and that does not change because a model produced the words.
Data is the second reason. Pasting customer details into a tool whose terms let the provider keep or train on that data can breach UK GDPR, so guardrails on inputs protect you as much as guardrails on outputs.
The third reason is money. Agents connected to ad accounts, shop platforms or email tools can spend, publish and send. A loose permission that nobody noticed can cost far more than the tool saves, and the risk of hallucination means some outputs will be confidently wrong.
Common mistakes
- Relying on the system prompt alone. Instructions can be overridden, including by text hidden in a web page or email the tool reads (prompt injection).
- Connecting agents with admin-level access because it was quicker to set up.
- Keeping no log, so nobody can say afterwards what the tool did or why.
- Writing the rules once and never testing them with awkward or hostile inputs.
- Making the rules so strict that staff give up and use personal accounts instead, which leaves you with no guardrails at all.
How to act on it
- List every AI tool in use across the business, who uses it and what data it touches.
- For each, write down in plain English what it may do, what it must never do and when it must stop and ask.
- Give each tool the least access that still lets it do the job, and add spending or sending limits where the platform allows.
- Put an approval step on anything public, anything that spends money and anything that cannot be reversed.
- Test it. Ask the tool the questions a difficult customer or a competitor would ask, and check what happens.
- Review logs and incidents monthly, and update the rules when the tool or your business changes.
If you want help deciding where AI fits in your marketing and what controls it needs, that is part of my digital marketing strategy and consulting work.
