Write each agent's rules in a file. Review it like code, test it in your pipeline, and enforce it before every request the agent sends.
This page shows the file format, the workflow and the tests, using an AI agent allow list for the web rules.
Most agent rules today live in a document or a prompt. Both fail in the same predictable ways.
Infrastructure teams solved the same problem years ago. Firewall rules and cloud permissions moved from tickets into files, and mistakes dropped because every change was reviewed and tested.
The policy says one thing. The agent's configuration says another.
Nobody notices until an incident.
"Never sign up for anything" is text the model may ignore.
Injected content can override it.
Who approved this agent's access, and when?
Without history, nobody can answer.
The file is the configuration, the history is in version control, and the enforcer reads the same file.
Plain JSON, so any language can read it. YAML works too when your enforcer supports it.
Keep files small. A good agent policy fits on one screen, so a reviewer can hold the whole thing in mind.
| Field | Meaning | Good default |
|---|---|---|
| agent, owner | Who the file is for, and who answers for it | Always set both |
| tier | How much the agent may do alone, 1 to 4 | Be honest; most real agents are 3 or 4 |
| review_by | Date the policy must be reviewed again | 90 days out |
| allow_page_types | Read page types this agent may open | Only what its purpose needs |
| deny_page_types | Extra read types to block | Empty unless there is a reason |
| allow_domains, deny_domains | Domain-level overrides | Short lists only |
| unclassified | What happens to pages nobody classified | deny for tiers 3 and 4 |
| on_deny | Block, or route to the owner for approval | ask_owner for tier 3, block for tier 4 |
| exceptions | One page type, one domain, an end date | As few as possible |
The enforcer applies five steps in a fixed order. Knowing the order makes every verdict easy to explain.
High-risk hosts, verified page types, URL rules and default-deny for unknown writes. Every agent gets these.
If the domain is on deny_domains, the request is denied whatever else is true.
A baseline denial can be lifted by a matching, unexpired exception, with or without owner approval.
With on_deny set to ask_owner, other denials become approval requests. High-risk hosts never do.
Read pages outside allow_page_types, types in deny_page_types, and unclassified pages under "deny" are blocked.
The file can only open narrow exceptions. It cannot switch off the baseline for high-risk hosts. That protects against a careless edit.
The order also makes verdicts predictable. Given the same file and the same URL, every enforcer returns the same answer.
"checkout exception added for supplier.example until March" is one line in a review.
Version control records who changed what, and who approved it.
A bad change is reverted like any other commit.
Keep a small list of URLs per agent with the verdict you expect. Run it on every change.
When a real denial surprises someone, add that URL to the list. The test suite grows from real experience.
A policy that blocks everything is safe and useless. Prove the agent can still do its job.
At least one URL per action type the agent could meet.
Inside the exception's domain, and just outside it.
Encoded paths and numeric IP forms of blocked hosts.
AgentPolicy.load("policy.json").check(Guard(), url) before each fetch.
An intercepting proxy loads the file for the agent's identity and checks every request.
The format is plain JSON. Apply the same five-step order in any language.
List only the read types the agent's purpose needs. A pricing monitor needs pricing, not blog.
Tier 4 agents deny unclassified pages. Add domains as they prove necessary.
One type, one domain, an end date. Renewal forces a second look.
Tier 3 agents send denials to their owner instead of silently failing.
Keep the risky-host and rule layers central. Agent files only hold what differs.
Allowing a domain also allows its signup and checkout pages.
It ends up as permissive as the most demanding agent.
Temporary access becomes permanent by default.
The model reads it; nothing enforces it.
A typo in a page type name silently changes behaviour.
| Question | Agent policy file | General-purpose policy engine |
|---|---|---|
| Who can read it? | Anyone: plain fields | Engineers who know the policy language |
| Knows page types? | Yes, built in | Only if you feed it the data |
| Flexibility | Fixed fields, fixed order | Anything you can express |
| Review effort | Low | Higher |
| Best for | Web access per agent | Complex rules across many systems |
Many teams use both. A general engine decides cross-system rules, and calls the page-type lookup when a rule depends on what a URL is.
Start with agent policy files. Move to a general engine only when rules start depending on other systems, such as ticket status or time of day.
Use the policy builder or agent-egress-guard policy init.
One folder, one file per agent, owners as code reviewers.
Fail the build on errors, comment on warnings.
Five to ten URLs per agent, with expected verdicts.
Record verdicts for a week, then switch to blocking.
The procurement team wants its research agent to place small orders with one supplier.
The agent owner opens a pull request adding a checkout exception for supplier.example, ending in March, owner approval per order.
Validation passes. The golden tests are updated: checkout on supplier.example now expects approval_required, checkout elsewhere still expects deny.
Security reads a six-line diff, asks for a lower spend limit in the purchasing system, and approves.
The enforcer reloads the file. The first order request reaches the owner for approval within minutes.
The exception ends on its date. The agent is denied again until someone renews it on purpose.
Every step left a record. An auditor can replay the whole story from the repository history.
Compare that with an email thread and a manual setting change, which leave nothing an auditor can trust.
| Change | Proposed by | Approved by |
|---|---|---|
| Add a read page type | Agent owner | AI platform team |
| Add an allowed domain | Agent owner | Security |
| Add an exception for an action page | Agent owner | Security and the budget or data owner |
| Change tier or unclassified setting | Agent owner | Security and risk |
| Change the shared baseline | Security | Security lead, with notice to all owners |
Encode these rules as required reviewers in your repository. Then the workflow enforces the governance, not a meeting.
Keep the list of approvers short. Two named people per change type is enough for most teams, and it keeps reviews fast.
No agent ships without its file. The pipeline refuses unregistered agents.
New tool, new task or new model means a new pull request.
A scheduled job opens an issue when review_by is close.
Deleting the file removes all access. The history stays in the repository.
Compare the repository against discovered agents each month.
Count files past their review_by date. The target is zero.
A growing exception list means the allow lists need rethinking.
A red build that caught a mistake is the system doing its job.
Every new page type or domain should trace back to the agent's written purpose.
One page type, one domain, a limit and an end date. Anything wider needs a stronger reason.
For money and publishing pages, a named person should approve each one.
A policy change without a test change is a warning sign.
New tools often push an agent up a tier without anyone saying so.
Short dates for wide permissions, longer ones for narrow read-only agents.
The rules every agent gets: risky hosts, verified page types, URL rules and default-deny for unknown writes.
A known request with the verdict you expect, run on every change.
A narrow, dated permission for one action page type on one domain.
How far the agent may act without a person deciding each step.
A verdict that pauses the request until the agent's owner decides.
A destination no layer can identify. Strict agents deny it.
The honest fine print — the same two assumptions we publish, plus two operational ones
Free builder, free enforcer, and page-type data when you need full coverage.