A guardrails framework organises controls so nothing is forgotten. This one has three parts: six control layers, four lifecycle stages and five maturity levels.
Use it to see where you are, what is missing, and what to build next. The web layer uses an AI agent allow list for page-level control.
Teams usually add guardrails one incident at a time. The result is strong in some places and missing in others.
The missing places are usually the ones nobody owns, such as web access, which sits between the AI team and network security.
A framework turns that patchwork into a map, so gaps are visible before an incident finds them.
Strong prompt filters, no web controls.
Three teams build three versions of the same rule.
Guardrails are designed once and never tested again.
"Are we safer than last quarter?" has no answer.
Each layer controls a different kind of agent action. Together they cover everything an agent can do.
The order is deliberate: layer 1 stops the actions whose consequences land outside your company.
Which pages the agent may open. Page types, high-risk hosts, URL rules, default-deny.
Which tools and servers it may call, and with what approval.
Which credentials it holds, for how long, with what rights.
Which data it may read, and what must be masked.
Sandboxing, budgets, alerts and the kill switch.
What it may send or publish, and how outputs are checked.
Guardrails are not built once. Each layer has work at every stage of an agent's life.
The grid below shows the main task for each layer at each stage. Empty cells in your own organisation are your gaps.
Find the level that describes you today, honestly. Then look at what the next level requires.
Score each layer separately, then take the lowest as your overall level.
Most organisations running agents in production sit at level 1 or 2 when they first check.
Rules live in prompts. Nothing enforces them.
Some controls, added after incidents. No inventory.
Web layer and identity in place for every agent.
All six layers, per-agent policies, tested regularly.
Metrics drive changes. Policies are code, reviewed and versioned.
Each step is achievable in about a quarter. Do not skip levels: each one builds the evidence the next depends on.
Write down the evidence for each level as you reach it. It becomes the proof your next assessment needs.
List your agents. Put any enforced control in front of the riskiest one.
Deny action pages and high-risk hosts at egress for every agent. Give each agent its own identity.
Add tools, data, runtime and output layers. Write one policy file per agent. Test each guardrail.
Report metrics monthly. Keep policies in version control with reviews. Tune rules from data.
Moving from level 1 to level 2 is mostly web layer work, because it stops the actions that land outside your company.
It also needs no change to the agents. That is rare among guardrails, and it makes layer 1 the fastest level-2 win.
Once layer 1 is in place, its decision logs also give you the evidence for the operate stage of every other layer.
Answer yes or no. The first "no" tells you your level.
Answer for your weakest agent, not your best one. Attackers and accidents find the weakest.
The six layers map onto the major agent security frameworks, so this structure plugs into an existing programme.
Use the mapping when auditors ask which published guidance your guardrails follow.
| Layer | OWASP agentic threats addressed | NIST AI RMF function |
|---|---|---|
| 1 Web access | Tool misuse, identity spoofing, rogue agents | Manage |
| 2 Tools | Tool misuse, unexpected code execution | Manage |
| 3 Identity | Privilege compromise | Manage |
| 4 Data | Memory poisoning, data exposure | Map, Manage |
| 5 Runtime | Resource overload, rogue agents | Measure, Manage |
| 6 Output | Human manipulation, repudiation | Govern, Measure |
See the framework comparison for more detail.
A composite path, not a specific company. Your pace depends on how many agents you run.
Notice that the first quarter starts with an incident. Starting before one is cheaper.
An agent signs up for a service. The team lists its agents and adds controls to the riskiest.
Action pages denied at egress for all agents. Own identities issued.
Policy files per agent, tool approvals, data masking, guardrail tests.
Monthly metrics, versioned policies, rules tuned from denial data.
One person owns the framework; each layer has its own owner.
The framework owner runs the quarterly assessment and chases the gaps between layers.
| Role | Responsibility |
|---|---|
| Framework owner | Maturity assessment, roadmap, reporting |
| Network security | Layer 1 |
| AI platform team | Layers 2 and 5 |
| Identity team | Layer 3 |
| Data protection | Layer 4 |
| Application owners | Layer 6 |
| Agent owners | Requesting the right guardrails for their agents |
Vendor agents sit in the same grid. Some cells move from your systems into the contract.
Mark those cells clearly, so nobody assumes the vendor covers them without checking.
Browser-based vendor agents pass your egress point.
Grants and data access in the vendor's admin console.
Ask which guardrails the vendor runs, and for evidence.
Give each vendor agent a maturity level in your register.
Report these to show the framework is working, not just written down.
Coverage should rise every quarter until every agent sits behind every relevant layer.
Agents behind each layer, as a share of all agents.
Guardrail tests passing this month.
Action-page and high-risk host denials, per agent.
Your self-assessed maturity, each quarter.
Short definitions for readers new to agent guardrails frameworks.
Use them consistently in assessments, roadmaps and reports to leadership.
A group of guardrails that controls one kind of agent action.
Design, build, deploy or operate.
How complete and measured your guardrails are, from 0 to 4.
One agent's rules, written as a file and enforced automatically.
Recording what a guardrail would block, before enforcing it.
A narrow, dated permission for one action on one domain.
Frameworks fail when they stay documents. These are the usual reasons.
Check your own programme against each one at every quarterly review, and fix the first you find.
The grid is complete on paper and empty in production.
Months on prompt filters while web access stays open.
Guardrails work at launch, then drift.
Self-assessments answered with plans instead of facts.
Every layer has an owner, but nobody sees the gaps between them.
The grid covers in-house agents only.
For most organisations, these five changes are the shortest path from level 1 to level 2.
Three of the five need no change to the agents themselves, only to the egress point and the logs.
A spreadsheet with owners is enough to start.
One policy, applied to every agent.
Metadata endpoints, consoles, registries.
Every web request, with page type and verdict.
Start with the agents that hold the most access.
Leaders need one picture: where you are, where you are going, and what it takes.
Keep it to one slide, updated each quarter with the same four boxes and fresh numbers.
One number from 0 to 4, with the evidence behind it.
The three or four changes needed, with owners and dates.
Share of agents behind each layer.
A few real examples of actions the guardrails stopped.
Most organisations already own tools for three or four layers. The gaps are usually layers 1 and 5.
Check each "often missing" column against your own setup before buying anything new.
| Layer | Typical tool | Often missing |
|---|---|---|
| 1 Web access | Egress proxy plus page-type data | Page-level decisions |
| 2 Tools | MCP gateway or tool allowlist | Per-tool approvals |
| 3 Identity | Identity platform, secrets manager | Short lifetimes |
| 4 Data | Data catalog, masking in the AI gateway | Per-agent data classes |
| 5 Runtime | Sandbox platform, monitoring | Budgets and denial alerts |
| 6 Output | Output checks, approval workflows | Approval for outside messages |
A framework can look like bureaucracy. These answers usually change minds.
Keep the framework short, and it stays useful rather than bureaucratic.
The grid fits on one page. Small teams benefit most, because nobody has time to rediscover gaps.
Platforms cover some layers. The grid shows which ones they leave to you.
The design stage is the cheapest time to add guardrails. After launch, every change is a migration.
Only if they are self-flattering. Tie each level to evidence, and they become a plan.
A composite example showing how uneven real coverage usually is.
Strong identity controls sit next to open web access in many organisations.
| Layer | Status | Level | Next step |
|---|---|---|---|
| 1 Web access | Domain blocklist only | 1 | Page-level policy at egress |
| 2 Tools | Approved list, no per-tool rules | 2 | Tag write and delete tools |
| 3 Identity | Own identities, one-hour tokens | 3 | Automate retirement |
| 4 Data | Masking for some agents | 2 | Per-agent data classes |
| 5 Runtime | No budgets or alerts | 1 | Denial alerts and step budgets |
| 6 Output | Manual spot checks | 2 | Approval for outside messages |
The overall level is the lowest layer: 1. Layers 1 and 5 are where the next quarter's work belongs.
The honest fine print — the same two assumptions we publish, plus two operational ones
Action pages and high-risk hosts denied, with page types for 40M+ domains.