AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases (15) Industries & Buyers (12) Learn: Core Concepts (12) Implementation Guides (15) Comparisons (8) Agent Security Guides (22) Market & Frameworks (9) Schema & Data Reference (6) FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
concrete rules, not principles

AI Agent Guardrails Examples

Guardrails are rules that stop an agent before it does something it should not. Most articles describe them in the abstract. This page gives 24 concrete examples.

Each example names the rule, what it stops and where to enforce it. The web layer uses page types from an AI agent allow list.

24Guardrail examples
6Layers
5Enforcement points
8Starter set
What makes a guardrail

Four tests every guardrail must pass

Many so-called guardrails are really instructions. A real guardrail passes all four of these tests.

If a proposed guardrail fails any of them, treat it as guidance for the agent, not as a control.

1

It runs outside the model

The agent cannot talk its way past it.

2

It acts before the action

Stopping, not reporting after the fact.

3

It gives the same answer every time

Deterministic, whatever the prompt says.

4

It logs why

Every block names the rule that fired.

The examples

24 guardrails in six layers

Rules are written in a simple pseudo-format. Adapt the syntax to whichever gateway, proxy or policy engine you already run.

Start with the web layer. It stops the actions that land outside your company, where mistakes are hardest to undo.

Layer 1: Web access

1
deny page_type in [signup, password_reset]
StopsAccounts created in your name
WhereEgress proxy or fetch tool
2
deny page_type in [cart, checkout, subscribe]
StopsUnapproved spending
WhereEgress proxy
3
deny page_type in [upload, post_create, comment]
StopsData leaving, public posts
WhereEgress proxy
4
deny page_type == login unless domain in internal_sso
StopsUse of stray credentials on other sites
WhereEgress proxy
5
deny host in high_value_hosts
StopsMetadata endpoints, consoles, registries
WhereEgress proxy
6
deny unclassified when tier >= 3
StopsWandering to unknown sites
WhereEgress proxy

Layer 2: Tools

7
allow tool only if tool in agent.approved_tools
StopsUnreviewed tools
WhereMCP gateway
8
require approval for tools tagged write or delete
StopsDestructive tool calls
WhereMCP gateway or agent runtime
9
fetch tools apply the agent's web policy
StopsTools bypassing layer 1
WhereInside the tool server
10
max 200 tool calls per task
StopsRunaway loops
WhereAgent runtime

Layer 3: Identity

11
each agent uses its own identity
StopsUntraceable actions
WhereIdentity platform
12
credential lifetime <= 1 hour
StopsLong-lived stolen keys
WhereIdentity platform
13
no production write credentials for coding agents
StopsDeleted production data
WhereSecrets manager
14
revoke all credentials on retirement
StopsZombie agents
WhereIdentity platform

Layer 4: Data

15
agent reads only its listed data classes
StopsUnneeded access to sensitive data
WhereData access layer
16
mask personal data before model calls
StopsPersonal data in prompts and logs
WhereAI gateway
17
no secrets in prompts or context
StopsKey leakage
WhereAgent runtime, secret scanning
18
deny destructive database statements
StopsDropped tables
WhereDatabase proxy

Layer 5: Runtime

19
code runs only in a sandbox with default-deny egress
StopsEscape from isolation
WhereSandbox platform
20
pause and alert after 5 denials in 10 minutes
StopsDrift and manipulation going unnoticed
WhereMonitoring
21
max run time per task
StopsEndless tasks
WhereAgent runtime

Layer 6: Output

22
outbound messages to outside parties need approval
StopsEmails and messages in your name
WhereMessaging integration
23
scan outputs for personal data before sharing
StopsLeaks through summaries
WhereAI gateway
24
label AI-generated content
StopsConfusion about who wrote what
WhereOutput pipeline
Starter set

The eight to put in place first

If you can only do eight this month, do these. Together they cover the most common and most expensive failures.

Four of the eight run at the egress point and need no changes to the agents themselves.

Examples 1, 2, 3

All action pages denied. One policy, same day.

Example 5

High-risk hosts denied, including numeric IP forms.

Example 11

Own identity per agent.

Example 13

No production write access for coding agents.

Example 17

No secrets in prompts.

Example 20

Alerts on denial spikes.

Enforcement points

Where each guardrail lives

Guardrails are only as strong as the place they run. Five enforcement points cover all 24 examples.

Put the same rule in two places when you can. If an agent bypasses one enforcement point, the other still decides.

Egress proxy

Examples 1 to 6. Sees every web request from every agent.

MCP gateway

Examples 7 to 9. Sees every tool call.

Identity platform

Examples 11 to 14. Decides what credentials exist.

AI gateway

Examples 16 and 23. Sees prompts and outputs.

Agent runtime and sandbox

Examples 10, 17 to 21. Limits the agent's own process.

In code

The web layer as a real policy file

Examples 1 to 6 combined into one per-agent file that the free Agent Egress Guard enforces.

The same file works for every agent with a similar purpose. Change the read types, and it fits another role.

{ "agent": "research-agent", "owner": "research-lead", "tier": 3, "web": { "allow_page_types": ["pricing", "documentation", "blog", "press", "about"], "unclassified": "deny", "on_deny": "ask_owner", "exceptions": [] } } # the baseline adds examples 1-5: action pages, login and high-value hosts are always denied

Build your own in the policy builder, or start from the six example files.

Anti-examples

Six "guardrails" that are not

These appear in many agent designs. Each fails at least one of the four tests.

Keep them if they help the agent behave, but do not count them as controls in a risk review.

"You must never make purchases."

A prompt instruction. Fails test 1.

A daily report of what the agent did

After the fact. Fails test 2.

A second model judging each action

Useful, but not deterministic. Fails test 3.

Blocking a whole domain

Stops research too, so it gets switched off.

Guessing /login paths

Misses real login pages on subdomains and locales.

Silent blocks

No reason logged. Fails test 4.

By agent type

Which examples matter most for which agent

All agents need the web layer. The rest depends on what the agent touches.

Use the table to pick a short list per agent, then add rules as the agent gains tools.

Agent typeMost important examples
Research or browsing agent1 to 6, 20
Coding agent5, 13, 17, 18, 19
Support agent3, 15, 16, 22, 23
Procurement agent2 (with one exception), 8, 22
Operations agent8, 12, 13, 18, 21
Vendor agent in SaaS11, 14, 15 via tenant settings
Testing

Proving each guardrail works

A guardrail nobody has tested is an assumption. Keep one test per example.

Store the tests next to the policy files, so a change to one prompts a check of the other.

Run the tests after every change to a policy, a tool or a model, and on a schedule.

1

Web layer

Request a known signup, checkout and upload URL. Expect three denials with the right page type.

2

Tools layer

Call an unapproved tool and a delete tool. Expect a refusal and an approval request.

3

Identity layer

Try a credential after its lifetime. Expect rejection.

4

Data layer

Send a destructive statement through the database proxy. Expect denial.

5

Runtime layer

Trigger five denials quickly. Expect a pause and an alert.

Rollout

From zero to all six layers in a quarter

Order by speed and impact. Each month builds on the work of the last.

By the end of the quarter, every agent should sit behind all six layers, with a test for each rule.

Month 1: web and logging

Examples 1 to 6 and 20. Most outside harm stopped.

Month 2: tools and identity

Examples 7 to 14. Access narrowed.

Month 3: data, runtime, output

Examples 15 to 24. Remaining gaps closed.

Objections

What builders say about guardrails

Guardrails can feel like friction. These answers usually help.

Share the day-in-the-life example above with builders. It shows how rarely guardrails get in the way.

"They will break the agent"

Read pages stay open. Most agents never notice the web layer at all.

"We cannot predict every case"

You do not need to. Deny the few risky actions and let the rest through.

"The model already refuses"

Models can be persuaded. Rules cannot.

"It is too much work"

The starter set takes about a week, and four of the eight need no agent changes.

Measuring it

Signals that guardrails are working

Track these per agent every month, and share them with each agent owner.

A rule that blocks legitimate work too often needs an exception or a narrower scope, not removal.

Blocks per rule

Rules that never fire may be unnecessary, or untested.

False blocks

Denials that stopped legitimate work.

Tests passing

One test per example, all green.

Agents covered

Agents behind all six layers.

Terms

Words used on this page

Short definitions for readers who are new to agent guardrails and enforcement.

Use the same words in policies, tickets and logs so everyone means the same thing.

Guardrail

A rule, outside the model, that stops an action before it happens.

Enforcement point

The system where a guardrail runs.

Page type

What a URL is on its site, such as pricing or checkout.

High-value host

A host agents may never reach, such as a metadata endpoint.

Tier

How far an agent may act alone.

Deterministic

Giving the same answer every time.

A day with guardrails

Which guardrails fire in an ordinary day

A composite research agent's day, showing which of the 24 examples actually trigger.

The numbers are illustrative. The pattern, long quiet stretches with a few important denials, is typical.

09:00 to 12:00 nothing fires

Pricing pages, documentation and blogs. Every request is a read page on the allow list.

12:14 example 1 fires

A "sign up for the full report" link. Denied as signup, owner notified.

13:40 example 6 fires

A link to an unknown file host. Denied as unclassified.

15:02 example 3 fires

The agent tries to leave a comment asking for data. Denied as comment.

16:30 task complete

Three denials, zero outside actions, one report delivered.

Most guardrails fire rarely. When they do, they stop exactly the actions nobody wanted.

Roles

Who owns each layer

Guardrails spread across systems owned by different teams. Name an owner for each layer.

The agent owner stays responsible for asking for the right guardrails. Layer owners are responsible for running them.

LayerOwnerReview
1 Web accessNetwork securityQuarterly, and on new agents
2 ToolsAI platform teamOn each new tool
3 IdentityIdentity teamMonthly
4 DataData protection and database teamsQuarterly
5 RuntimePlatform engineeringMonthly
6 OutputApplication ownersQuarterly
Vendor agents

Guardrails for agents you do not host

You cannot add code to a vendor's agent. You can still apply several layers.

Record which layers apply to each vendor agent in its registry entry, so gaps are visible.

Layer 1 at your proxy

Browser-based vendor agents pass your egress proxy.

Layer 3 through grants

Limit and expire the OAuth grants you give them.

Layer 4 through settings

Restrict which data the vendor agent may read.

The rest by contract

Ask the vendor which guardrails they run, and for logs.

Quick wins

What you can switch on today

These need no code in the agent and no new vendor.

Together they cover four of the eight starter examples in a single afternoon.

Install Agent Egress Guard

Examples 1, 4, 5 and default-deny writes, free.

Search for keys in prompts

Example 17, with the secret scanner you already have.

Set a denial alert

Example 20, on logs you already collect.

Remove production write keys

Example 13, in an afternoon.

Which examples would have mattered in 2026

  • Example 3 for the wiki edits and dataset uploads, example 5 for cloud metadata, example 4 for account access.
  • All four are web-layer guardrails at the egress point.
  • In our replay, page data plus egress rules would have stopped almost all of the incidents.
See the incident-by-incident prevention analysis The egress rules library

The honest fine print — the same two assumptions we publish, plus two operational ones

  1. The policy engine must see every request — an agent with raw socket access or a second network path bypasses everything; enforcement belongs at the egress proxy/network layer, not only in an SDK hook.
  2. Default-deny must be on. In flag-only mode these become alerts within minutes rather than prevention — still a large improvement on a timeline measured in weeks (the DseWiki edits ran from late May to late June 2026, per the researchers), but not a block.
  3. For full URL+method matching on HTTPS you need to be the proxy or in-process hook — SNI alone shows only the host, which still catches the entire host-list layer.
  4. Policy can’t read intent inside a legitimately allowed action: an agent whose job is publishing packages keeps registry access. In our replay of the 2026 incidents, no crossing fits any plausible allowlist for the agents’ documented tasks.
Related

Keep reading

FAQ

Guardrails questions

What are examples of AI agent guardrails?
Denying signup, checkout and upload pages; denying cloud metadata hosts; approved tools only; short-lived credentials; no production write access for coding agents; alerts on repeated denials.
What is the difference between a guardrail and an instruction?
An instruction is text the model may follow. A guardrail runs outside the model, acts before the action and gives the same answer every time.
Which guardrails should come first?
The web layer: deny action pages and high-risk hosts at the egress point. It stops most outside harm in a day.
Where should guardrails run?
At five points: the egress proxy, the MCP gateway, the identity platform, the AI gateway and the agent runtime or sandbox.
How do I test a guardrail?
Keep one test per rule, such as requesting a known signup URL, and check that it is denied with the right reason.
Can guardrails be applied to vendor agents?
Partly. Web rules apply at your proxy for browser-based agents, identity rules through the grants you give, and data rules through product settings. Ask the vendor about the rest.
How many guardrails does an agent need?
Start with the eight in the starter set. Add the rest by agent type, following the table on this page.
How does the web layer know which page is a checkout?
It looks the URL up in page-type data. An AI agent allow list covers 28 page types on 40M+ domains, from $99 a month.

Put examples 1 to 6 in place today

Free Agent Egress Guard plus page types for 40M+ domains.

Get Agent Egress Guard