AI agents
What are AI guardrails?
In short
AI guardrails are checks placed around a language model or AI agent that limit what goes in, what it can reach, what it can do and what comes out. They work at four points: input, retrieval, tool use and output. Guardrails reduce risk but do not remove it, so they work best alongside permissions, human approvals and an audit trail.
AI guardrails are the checks around a language model or AI agent that keep it inside safe limits. A plain chatbot only needs checks on the words going in and coming out. An AI agent reads company data and calls tools that change things, so it needs checks at more points.
This page explains the four places guardrails sit, common examples of each, how guardrails differ from permissions and human approvals, and what guardrails cannot do.
What are AI guardrails?
A guardrail is any rule or check that sits between the model and the world. It can block a request, change data before the model sees it, stop a tool call, or rewrite or reject an answer.
Guardrails exist because a large language model (LLM) does not follow rules reliably on its own. Instructions in a system prompt help, but a model can be confused, tricked or simply wrong. The OWASP Top 10 for LLM Applications lists the main risks, including prompt injection, sensitive information disclosure, improper output handling and excessive agency (an agent that can do more than its job needs). Most guardrails are aimed at one of those risks.
What are the main types of AI guardrails?
Think of an agent’s work as a path: a request comes in, the agent finds information, it may call tools, and it gives an answer. Guardrails sit at each step.
- Input guardrails check the request before the model acts on it. Examples: blocking off-topic or abusive requests, detecting known prompt injection phrases, and limiting message length.
- Retrieval guardrails control what data the agent can find. Examples: searching only the folders picked for this agent, and applying the asker’s own file permissions so they only get answers from documents they could open anyway.
- Tool guardrails control what the agent can do. Examples: giving an agent only the tools its job needs, making a database connection read-only, limiting it to certain tables, and requiring a person to approve risky actions.
- Output guardrails check the answer before a person or system uses it. Examples: masking personal data, checking that an answer cites a real source, validating structured output against a schema, and filtering unsafe content.
A fifth, less visible layer is data guardrails: masking sensitive values such as card numbers or e-mail addresses in documents and tool results before the model reads them. This limits what can leak even if every other check fails.
What do guardrails look like in practice?
| Stage | Risk it targets | Example guardrail | Typical method |
|---|---|---|---|
| Input | Prompt injection, misuse | Reject requests outside the agent’s job | Rules or a classifier model |
| Retrieval | Oversharing private files | Apply the asker’s permissions inside the search | Access control in the index |
| Data | Personal data reaching the model | Mask SSNs, card numbers, API keys | Pattern and detector matching |
| Tools | Excessive agency | Read-only by default, approval for writes | Tool allow lists and approvals |
| Output | Wrong or unsafe answers | Require citations, validate format | Rules, schemas or a checker model |
Two methods appear again and again:
- Deterministic rules: patterns, allow lists, schemas and permission checks. They are fast and predictable, but only catch what they were written for.
- Model-based checks: a second model judges whether an input or output is safe or on topic. They catch more varied cases, but they cost time and money and can be wrong too.
How are guardrails different from permissions and approvals?
These three ideas overlap, and people often mix them up. They answer different questions.
| Control | Question it answers | When it acts | Example |
|---|---|---|---|
| Permissions | What is this person or agent allowed to reach? | Before any data is found or any tool is offered | A sales rep’s agent cannot see HR folders |
| Guardrails | Is this specific input, data or output safe? | During the run, on each step | Mask a card number in a tool result |
| Human approvals | Should this specific action happen now? | Just before a risky action runs | A manager approves a refund over a set amount |
Permissions set the outer boundary. Guardrails filter what moves inside that boundary. Approvals put a person in charge of the few actions that are too risky to leave to software. A solid setup uses all three. For more on approvals, see human in the loop AI.
What can AI guardrails not do?
Honest limits matter, because over-trusting a guardrail is a risk of its own.
- They do not eliminate prompt injection. Attackers keep finding new wording. Treat any text the agent reads, including documents and tool results, as untrusted.
- Pattern-based masking misses things. A detector finds formats it knows, such as card numbers. It will not catch every name, address or piece of context that identifies a person.
- Model-based checks can be fooled by the same tricks that fool the main model.
- Approvals depend on people. A reviewer who clicks “approve” on everything adds no safety.
- Filtering after the fact is weaker than not fetching. If an agent retrieves a document the person should not see and a filter hides it later, the data has already been read. Applying access rules inside the search is safer. Our benchmark post shows why doing that well is harder than it looks.
The practical lesson: design so a guardrail failure is survivable. Limit what the agent can reach and do, so even a tricked agent cannot cause serious harm.
How do you choose guardrails for an AI agent?
Start from what the agent can do, not from a list of products.
- List the agent’s data and tools. Write down every source it reads and every action it can take.
- Remove what it does not need. The smallest set of tools and folders is the strongest guardrail.
- Sort actions by risk. Reads are usually low risk, writes medium, and deletes or payments high.
- Add approvals to high-risk actions and consider conditions, such as amounts over a limit.
- Mask sensitive data the agent does not need to see in plain form.
- Log everything, so you can review what happened and tune the rules.
- Test with bad inputs, including injected instructions hidden in documents.
Frameworks such as the NIST AI Risk Management Framework help you tie these choices to a wider risk process. For a full checklist, see AI agent governance.
How promptev handles AI guardrails
- The asker’s identity and groups are applied inside the search for Google Drive, SharePoint, OneDrive and Dropbox, so a file they cannot open is never fetched.
- Masking uses built-in detectors (e-mail, phone, US SSN, card numbers, IBAN, API keys) plus your own patterns, and can mask, hash or remove values in found documents and tool results before the model reads them.
- Each tool has an approval setting: by default reads need none, writes ask once and destructive actions always ask, and builders can change it.
- Registered HTTP, database and MCP tools can require approval every time or only when a condition is met, and an admin’s rule is a floor that agents cannot weaken.
- Database tools are read-only by default and limited to the tables you pick.
- Every run, approval and settings change is recorded in the audit trail. See governance for details.
Frequently asked questions
What is the difference between AI guardrails and LLM guardrails?
The terms are mostly used the same way. LLM guardrails usually means checks on the text going into and out of a model, while AI agent guardrails also cover what the agent can read and which tools it can call.
Can guardrails stop prompt injection?
No guardrail stops prompt injection completely. Input filters catch known patterns, but the stronger defence is limiting what an agent can reach and do, so a tricked agent cannot cause much harm.
Are guardrails the same as content moderation?
Content moderation is one kind of guardrail, usually applied to inputs and outputs. Guardrails also include data masking, access limits, tool restrictions and approval steps.
Do guardrails slow down an AI agent?
Some do. Extra model calls for classification add time, while simple rules such as pattern masking or tool allow lists add very little. Most teams put the heavier checks only on risky actions.
Who should own AI guardrails in a company?
Usually a shared group: security sets the baseline rules, the team that builds the agent tunes them for its job, and an admin can enforce a floor that individual builders cannot weaken.

Faisal Saeed is Founder & CEO of Promptev, building next-gen context engineering infrastructure that enables teams to orchestrate, scale, and deploy production-ready generative AI systems with confidence.