AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases (15) Industries & Buyers (12) Learn: Core Concepts (12) Implementation Guides (15) Comparisons (8) Agent Security Guides (22) Market & Frameworks (9) Schema & Data Reference (6) FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
a structure for every guardrail you will ever add

AI Agent Guardrails Framework

A guardrails framework organises controls so nothing is forgotten. This one has three parts: six control layers, four lifecycle stages and five maturity levels.

Use it to see where you are, what is missing, and what to build next. The web layer uses an AI agent allow list for page-level control.

6Control layers
4Lifecycle stages
5Maturity levels
24Grid cells
Why a framework

What goes wrong without one

Teams usually add guardrails one incident at a time. The result is strong in some places and missing in others.

The missing places are usually the ones nobody owns, such as web access, which sits between the AI team and network security.

A framework turns that patchwork into a map, so gaps are visible before an incident finds them.

Uneven coverage

Strong prompt filters, no web controls.

Duplicated effort

Three teams build three versions of the same rule.

No lifecycle

Guardrails are designed once and never tested again.

No way to measure progress

"Are we safer than last quarter?" has no answer.

Part 1

Six control layers

Each layer controls a different kind of agent action. Together they cover everything an agent can do.

The order is deliberate: layer 1 stops the actions whose consequences land outside your company.

1

Web access

Which pages the agent may open. Page types, high-risk hosts, URL rules, default-deny.

2

Tools

Which tools and servers it may call, and with what approval.

3

Identity

Which credentials it holds, for how long, with what rights.

4

Data

Which data it may read, and what must be masked.

5

Runtime

Sandboxing, budgets, alerts and the kill switch.

6

Output

What it may send or publish, and how outputs are checked.

Part 2

Four lifecycle stages

Guardrails are not built once. Each layer has work at every stage of an agent's life.

The grid below shows the main task for each layer at each stage. Empty cells in your own organisation are your gaps.

Layer
Design
Build
Deploy
Operate
1 Web access
Choose read page types by purpose
Write the policy file
Enforce at egress, log-only first
Review denials, renew exceptions
2 Tools
List needed tools
Tag write and delete tools
Approve in the MCP gateway
Re-review new versions
3 Identity
Define the agent identity
Scope credentials
Issue short-lived access
Revoke on retirement
4 Data
List data classes
Set masking rules
Enforce at the data layer
Audit access
5 Runtime
Set budgets
Configure sandbox
Test kill switch
Alert on spikes
6 Output
Decide what may leave
Build output checks
Require approval for outside messages
Sample outputs
Part 3

Five maturity levels

Find the level that describes you today, honestly. Then look at what the next level requires.

Score each layer separately, then take the lowest as your overall level.

Most organisations running agents in production sit at level 1 or 2 when they first check.

LEVEL 0

Hope

Rules live in prompts. Nothing enforces them.

LEVEL 1

Reactive

Some controls, added after incidents. No inventory.

LEVEL 2

Baseline

Web layer and identity in place for every agent.

LEVEL 3

Managed

All six layers, per-agent policies, tested regularly.

LEVEL 4

Measured

Metrics drive changes. Policies are code, reviewed and versioned.

Moving up

What each step up requires

Each step is achievable in about a quarter. Do not skip levels: each one builds the evidence the next depends on.

Write down the evidence for each level as you reach it. It becomes the proof your next assessment needs.

0→1

From hope to reactive

List your agents. Put any enforced control in front of the riskiest one.

1→2

From reactive to baseline

Deny action pages and high-risk hosts at egress for every agent. Give each agent its own identity.

2→3

From baseline to managed

Add tools, data, runtime and output layers. Write one policy file per agent. Test each guardrail.

3→4

From managed to measured

Report metrics monthly. Keep policies in version control with reviews. Tune rules from data.

Why layer 1 comes first

The web layer carries the most weight early

Moving from level 1 to level 2 is mostly web layer work, because it stops the actions that land outside your company.

8Action page types
40M+Domains classified
~60High-risk hosts
1 dayTo enforce

It also needs no change to the agents. That is rare among guardrails, and it makes layer 1 the fastest level-2 win.

Once layer 1 is in place, its decision logs also give you the evidence for the operate stage of every other layer.

Self-assessment

Twelve questions to place yourself

Answer yes or no. The first "no" tells you your level.

Answer for your weakest agent, not your best one. Attackers and accidents find the weakest.

Levels 1 and 2

  • Do you have a list of every agent?
  • Does every agent have an owner?
  • Are action pages denied before each request?
  • Are high-risk hosts denied?
  • Does every agent have its own identity?
  • Is every request logged with the agent ID?

Levels 3 and 4

  • Does every agent have its own policy file?
  • Are tools approved one by one?
  • Is sensitive data masked before model calls?
  • Is each guardrail tested on a schedule?
  • Are metrics reported monthly?
  • Are policies versioned and reviewed like code?
Mapping

How the framework fits published guidance

The six layers map onto the major agent security frameworks, so this structure plugs into an existing programme.

Use the mapping when auditors ask which published guidance your guardrails follow.

LayerOWASP agentic threats addressedNIST AI RMF function
1 Web accessTool misuse, identity spoofing, rogue agentsManage
2 ToolsTool misuse, unexpected code executionManage
3 IdentityPrivilege compromiseManage
4 DataMemory poisoning, data exposureMap, Manage
5 RuntimeResource overload, rogue agentsMeasure, Manage
6 OutputHuman manipulation, repudiationGovern, Measure

See the framework comparison for more detail.

A year of progress

How a typical organisation climbs

A composite path, not a specific company. Your pace depends on how many agents you run.

Notice that the first quarter starts with an incident. Starting before one is cheaper.

Quarter 1: level 0 to 1

An agent signs up for a service. The team lists its agents and adds controls to the riskiest.

Quarter 2: level 1 to 2

Action pages denied at egress for all agents. Own identities issued.

Quarter 3: level 2 to 3

Policy files per agent, tool approvals, data masking, guardrail tests.

Quarter 4: level 3 to 4

Monthly metrics, versioned policies, rules tuned from denial data.

Ownership

Who owns the framework

One person owns the framework; each layer has its own owner.

The framework owner runs the quarterly assessment and chases the gaps between layers.

RoleResponsibility
Framework ownerMaturity assessment, roadmap, reporting
Network securityLayer 1
AI platform teamLayers 2 and 5
Identity teamLayer 3
Data protectionLayer 4
Application ownersLayer 6
Agent ownersRequesting the right guardrails for their agents
Vendor agents

Applying the framework to agents you buy

Vendor agents sit in the same grid. Some cells move from your systems into the contract.

Mark those cells clearly, so nobody assumes the vendor covers them without checking.

Layer 1 at your proxy

Browser-based vendor agents pass your egress point.

Layers 3 and 4 in settings

Grants and data access in the vendor's admin console.

Layers 2, 5 and 6 by contract

Ask which guardrails the vendor runs, and for evidence.

Score them too

Give each vendor agent a maturity level in your register.

Measuring it

Metrics for each maturity level

Report these to show the framework is working, not just written down.

Coverage should rise every quarter until every agent sits behind every relevant layer.

Coverage

Agents behind each layer, as a share of all agents.

Test health

Guardrail tests passing this month.

Denials

Action-page and high-risk host denials, per agent.

Level

Your self-assessed maturity, each quarter.

Terms

Words used on this page

Short definitions for readers new to agent guardrails frameworks.

Use them consistently in assessments, roadmaps and reports to leadership.

Control layer

A group of guardrails that controls one kind of agent action.

Lifecycle stage

Design, build, deploy or operate.

Maturity level

How complete and measured your guardrails are, from 0 to 4.

Policy file

One agent's rules, written as a file and enforced automatically.

Log-only mode

Recording what a guardrail would block, before enforcing it.

Exception

A narrow, dated permission for one action on one domain.

Common mistakes

Where frameworks fail in practice

Frameworks fail when they stay documents. These are the usual reasons.

Check your own programme against each one at every quarterly review, and fix the first you find.

Designed, never deployed

The grid is complete on paper and empty in production.

One layer over-invested

Months on prompt filters while web access stays open.

No operate stage

Guardrails work at launch, then drift.

Level inflation

Self-assessments answered with plans instead of facts.

Nobody owns the whole

Every layer has an owner, but nobody sees the gaps between them.

Vendor agents left out

The grid covers in-house agents only.

Quick wins

Moving up a level this month

For most organisations, these five changes are the shortest path from level 1 to level 2.

Three of the five need no change to the agents themselves, only to the egress point and the logs.

List every agent

A spreadsheet with owners is enough to start.

Deny action pages at egress

One policy, applied to every agent.

Deny high-risk hosts

Metadata endpoints, consoles, registries.

Log with agent IDs

Every web request, with page type and verdict.

Own identities

Start with the agents that hold the most access.

Leadership

Reporting the framework upward

Leaders need one picture: where you are, where you are going, and what it takes.

Keep it to one slide, updated each quarter with the same four boxes and fresh numbers.

Current level

One number from 0 to 4, with the evidence behind it.

Next level

The three or four changes needed, with owners and dates.

Coverage

Share of agents behind each layer.

Recent denials

A few real examples of actions the guardrails stopped.

Tooling

What you need for each layer

Most organisations already own tools for three or four layers. The gaps are usually layers 1 and 5.

Check each "often missing" column against your own setup before buying anything new.

LayerTypical toolOften missing
1 Web accessEgress proxy plus page-type dataPage-level decisions
2 ToolsMCP gateway or tool allowlistPer-tool approvals
3 IdentityIdentity platform, secrets managerShort lifetimes
4 DataData catalog, masking in the AI gatewayPer-agent data classes
5 RuntimeSandbox platform, monitoringBudgets and denial alerts
6 OutputOutput checks, approval workflowsApproval for outside messages
Objections

What teams say about adopting a framework

A framework can look like bureaucracy. These answers usually change minds.

Keep the framework short, and it stays useful rather than bureaucratic.

"We are too small for a framework"

The grid fits on one page. Small teams benefit most, because nobody has time to rediscover gaps.

"Our platform handles security"

Platforms cover some layers. The grid shows which ones they leave to you.

"We will do it after launch"

The design stage is the cheapest time to add guardrails. After launch, every change is a migration.

"Maturity levels are vanity"

Only if they are self-flattering. Tie each level to evidence, and they become a plan.

Worked assessment

One organisation, scored layer by layer

A composite example showing how uneven real coverage usually is.

Strong identity controls sit next to open web access in many organisations.

LayerStatusLevelNext step
1 Web accessDomain blocklist only1Page-level policy at egress
2 ToolsApproved list, no per-tool rules2Tag write and delete tools
3 IdentityOwn identities, one-hour tokens3Automate retirement
4 DataMasking for some agents2Per-agent data classes
5 RuntimeNo budgets or alerts1Denial alerts and step budgets
6 OutputManual spot checks2Approval for outside messages

The overall level is the lowest layer: 1. Layers 1 and 5 are where the next quarter's work belongs.

Where the 2026 incidents sat on this framework

  • The agents involved were effectively at level 0 or 1 for layer 1: no page-level web control.
  • Level 2 web controls deny the uploads, edits and plugin installs they relied on.
  • In our replay, page data plus egress rules would have stopped almost all of them.
The 2026 agent incidents, prevented How the Hugging Face breach could have been stopped

The honest fine print — the same two assumptions we publish, plus two operational ones

  1. The policy engine must see every request — an agent with raw socket access or a second network path bypasses everything; enforcement belongs at the egress proxy/network layer, not only in an SDK hook.
  2. Default-deny must be on. In flag-only mode these become alerts within minutes rather than prevention — still a large improvement on a timeline measured in weeks (the DseWiki edits ran from late May to late June 2026, per the researchers), but not a block.
  3. For full URL+method matching on HTTPS you need to be the proxy or in-process hook — SNI alone shows only the host, which still catches the entire host-list layer.
  4. Policy can’t read intent inside a legitimately allowed action: an agent whose job is publishing packages keeps registry access. In our replay of the 2026 incidents, no crossing fits any plausible allowlist for the agents’ documented tasks.
Related

Keep reading

FAQ

Guardrails framework questions

What is an AI agent guardrails framework?
A structure for organising agent controls. This one uses six control layers, four lifecycle stages and five maturity levels.
What are the layers of AI agent guardrails?
Web access, tools, identity, data, runtime and output.
Which layer should come first?
Web access. It stops actions that land outside your company and needs no change to the agents.
How do I measure guardrail maturity?
Answer the twelve self-assessment questions. The first "no" marks your level, from 0 (hope) to 4 (measured).
How long does it take to reach level 3?
Roughly one quarter per level for most organisations, depending on how many agents they run.
Is my maturity level the average of the layers?
No. Use your lowest layer. An agent is only as contained as its weakest control.
Can we skip straight to level 3?
Rarely. Each level produces the inventory, identities and logs the next one depends on.
Does the framework apply to vendor agents?
Yes. Some layers move into product settings and contracts, but the same grid applies.

Reach level 2 in a day for the web layer

Action pages and high-risk hosts denied, with page types for 40M+ domains.

Free Agent Egress Guard