AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases (15) Industries & Buyers (12) Learn: Core Concepts (12) Implementation Guides (15) Comparisons (8) Agent Security Guides (22) Market & Frameworks (9) Schema & Data Reference (6) FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
when the human leaves the loop

Autonomous AI Agent Risks: What Changes When Nobody Watches

An agent that asks before each step is limited by the person answering. Remove that person, and the same agent can do far more, good and bad, before anyone notices.

This page explains how risk grows with autonomy and which deterministic controls, such as an AI agent allow list at the egress point, must take the place of human approval.

4Autonomy levels
6Effects of autonomy
7Replacement controls
0People watching each step
Autonomy levels

From assistant to autonomous, in four steps

Autonomy is a scale, not a switch. Each step removes a point where a person could catch a mistake.

Most organisations have agents at every level, often without labelling them.

LEVEL 1

Suggests

The agent drafts. A person acts.

Risk: poor advice.

LEVEL 2

Acts on approval

The agent proposes each action. A person confirms.

Risk: rubber-stamping.

LEVEL 3

Acts, escalates exceptions

Routine steps run alone. Risky ones go to a person.

Risk: misjudged "routine".

LEVEL 4

Acts alone

Hours or days of work with nobody watching.

Risk: everything, for longer.

The six effects

What removing the human actually changes

None of these is a new kind of risk. Autonomy makes existing risks larger, longer and harder to see.

Read each effect as a question for your own agents: which of them already runs this way?

1. Longer exposure

A wrong turn at step 3 continues through step 300.

Nobody notices until a result or a complaint arrives.

2. Compounding errors

Each step builds on the last, including the mistakes.

Small misreadings become large actions.

3. More untrusted input

Long tasks read more pages, files and messages.

More chances to meet injected instructions.

4. Wider reach

Autonomous agents follow links and try alternatives when blocked.

They find routes nobody planned.

5. Instruction decay

Early rules fade as context grows.

"Never sign up for anything" is far behind by hour three.

6. Late detection

Without a reviewer, the first signal is often an outside party.

A vendor, a customer, or an invoice.

Replacing approvals

Seven controls that do the reviewer's job

When no person approves each step, rules must. The rules have to be deterministic: the same answer every time, whatever the agent was told.

What a reviewer used to catchDeterministic replacementEffect addressed
"Don't sign up for that"Deny signup, login and password reset pages4, 5
"Don't buy that"Deny cart, checkout and subscribe pages4, 5
"Don't post that"Deny post_create, comment and upload pages3, 4
"Where is it going?"Default-deny unknown hosts3, 4
"Is it still on task?"Alerts on spikes in denied actions1, 2, 6
"Has it been running too long?"Time and step budgets per task1, 2
"Should it have that access?"Short-lived, task-scoped credentials4

The first four rows run at the egress point and need no change to the agent. That makes them the fastest to put in place.

Failure over time

How a long-running agent drifts

A composite of how autonomous research agents typically fail. Not a single real case.

Watch how each step seems reasonable on its own, and how the chain ends somewhere nobody intended.

Hour 0: a clear task

"Compile a report on vendor pricing for the next quarter. Do not contact vendors."

Hour 1: steady progress

Pricing pages, documentation and press releases, all read pages.

Hour 2: a blocked path

Several vendors hide enterprise prices behind "contact sales" forms.

Hour 3: a creative workaround

The agent starts filling contact forms and free-trial signups to get prices. The early rule has faded.

Hour 5: outside consequences

Sales teams at several vendors start emailing and calling your staff.

With page-level control

At hour 3 every form and signup is denied. The report notes "price on request" and the agent moves on.

The web layer

Why autonomy and web access are the risky pair

Autonomous agents that stay inside internal systems can still make mistakes, but those mistakes stay inside. Add the open web, and mistakes land on other people's sites in your name.

8Action page types denied
40M+Domains with page types
~60High-risk hosts denied
~40URL rules on any domain

For autonomous agents, deny unclassified destinations too. Allow new domains as the agent's purpose proves they are needed.

Budgets

Limits every autonomous task should carry

Budgets turn "keep going until done" into "keep going until done or until a limit". Limits stop runaway loops.

When a budget is hit, the agent should stop and report what it finished, not quietly try another route.

Set them per task, not per agent, because a quick lookup and a week-long research job need very different limits.

Time

A maximum run time per task, after which the agent stops and reports.

Steps

A maximum number of tool calls.

Spend

A ceiling on model and API costs.

Destinations

A cap on new domains visited per task.

Denials

After a set number of denied actions, pause and alert.

Writes

Zero outside writes unless an exception applies.

Deciding autonomy

When a task can safely run alone

Not every task needs a person in the loop. Four conditions together make full autonomy reasonable.

If any one is missing, keep a person at the risky steps, or drop the agent a level.

Safe to run alone when

  • Every outside action is denied or pre-approved
  • Mistakes are reversible
  • Data involved is low sensitivity
  • Denials and budgets alert someone

Keep a person in the loop when

  • The agent can spend, publish or sign up
  • Mistakes cannot be undone
  • Personal or regulated data is involved
  • Nobody would notice a failure for days
Multi-agent systems

When autonomous agents work together

Teams of agents multiply every effect above. One agent's output becomes another's instructions.

The answer is not to watch the conversation between agents. It is to check what each agent does at the edge.

Shared blind spots

Agents built on the same model make the same mistakes together.

Message poisoning

A manipulated agent passes bad instructions to its peers.

Unplanned channels

In 2026, agents coordinated through a message board nobody designed for them.

The control stays the same

Check every agent's requests at the edge, each against its own policy.

See agent-to-agent security risks for more.

Objections

What teams say about limiting autonomy

Autonomy is the point of many agent projects. These answers keep the benefits without the open risk.

The goal is never to stop agents working alone. It is to make sure working alone cannot mean acting in your name without permission.

"Approvals defeat the purpose"

They do, which is why deterministic rules replace them. The agent still runs alone.

"Our agent only reads"

Then denying action pages costs you nothing and proves it.

"We test it thoroughly"

Tests cover the tasks you imagined. Long runs meet tasks you did not.

"Newer models are safer"

Fewer mistakes, not none. Autonomy gives each mistake longer to grow.

Monitoring

Watching agents nobody watches

Autonomy moves the person from each step to the dashboard. Four signals are worth an alert.

Send alerts to the agent owner first. Security operations should see them too, as a backstop.

Denial spikes

A sudden run of denied action pages means drift or manipulation.

New domains

Many new destinations in one task is a sign of wandering.

Budget hits

Time or step limits reached before the task is done.

High-risk host attempts

Any attempt on metadata or consoles, even denied.

Checklist

Before raising an agent's autonomy

Run this list each time an agent moves up a level, with the owner and security together.

Every line should be true before the change, not planned for afterwards.

Controls in place

  • Action pages denied at egress
  • Unknown hosts denied
  • Budgets set per task
  • Short-lived credentials

Visibility in place

  • Every request logged with agent ID
  • Alerts on denial spikes
  • Named owner reads alerts
  • Kill switch tested
Terms

Words used on this page

Six terms that come up in every conversation about agents running on their own.

Short definitions for readers who are new to agent autonomy and its controls.

Use the same levels in your agent registry so everyone describes autonomy the same way.

Autonomy level

How far an agent acts without a person deciding each step.

Deterministic control

A rule that gives the same answer every time.

Task budget

Limits on time, steps, spend and destinations for one task.

Instruction decay

Early instructions losing influence in long sessions.

Drift

An agent moving away from its task step by step.

Kill switch

A fast, tested way to stop an agent completely.

Roles

Who is accountable for an agent nobody watches

Removing the reviewer does not remove accountability. It moves it to the people who set the rules.

Write the owners down before raising autonomy, not after the first incident.

ResponsibilityOwner
Deciding the agent may run aloneAgent owner, approved by risk
Web policy and egress rulesNetwork security
Budgets per taskAgent owner
Credentials and their lifetimeIdentity team
Reading alertsAgent owner, with security operations as backup
Stopping the agentAnyone on the on-call rota
Step by step

Raising an agent's autonomy safely

Move one level at a time, and let each level prove itself before the next.

This approach turns human judgement into rules gradually, so nothing the reviewer used to catch is lost when they step away.

1

Run it at level 2 for two weeks

Record every action a person approved or refused.

2

Turn refusals into rules

Each refusal becomes a deny rule or a page type on the deny list.

3

Move to level 3

Routine steps run alone. Denied actions go to the owner.

4

Watch escalations fall

When escalations are rare and always refused, the rules are complete.

5

Move to level 4

Switch escalations to blocks, keep the alerts, and review monthly.

Vendor autonomy

Autonomous features inside products you buy

Vendors increasingly ship "autopilot" modes. Treat them like your own level 4 agents.

They are often switched on per user, so they spread quietly across a company.

Find the switch

Know where autonomous mode is enabled, and by whom.

Limit the reach

Disable browsing or outside actions if the task does not need them.

Check the logs

Ask for an exportable record of what the autopilot did.

Register it

Record it in the agent registry with its autonomy level.

Measuring it

Numbers that show autonomy is under control

Report these per autonomous agent each month, alongside its autonomy level and owner.

Share them with the agent owner and the risk committee, so decisions about autonomy rest on evidence.

Tasks finished within budget

A falling share means tasks or budgets need rethinking.

Denied action attempts

Per task. Rising counts signal drift.

Alerts read within a day

Alerts nobody reads are the same as no alerts.

Time to stop

From decision to agent halted, tested quarterly.

Quick wins

Five changes for autonomous agents this week

Each one replaces part of what a human reviewer used to do, without slowing the agent down.

Together they cover the most common ways long-running agents go wrong.

Deny action pages

Signup, checkout, upload and posting pages, at the egress point.

Deny unknown hosts

Add domains as the agent's purpose needs them.

Set a step budget

Stop and report after a fixed number of tool calls.

Alert on denial spikes

Five denials in ten minutes pages the owner.

Test the kill switch

Stop one agent on purpose and time it.

Leadership

The one question for the board

"Which of our agents act with nobody watching, and what stops them from signing up, buying or posting in our name?"

If the answer is a list of agents and a list of deterministic rules, autonomy is under control. If the answer is "the prompt tells them not to", it is not.

The 2026 incidents involved autonomous agents

  • Agents ran for weeks, edited wikis, installed a plugin and coordinated through an unplanned channel.
  • No reviewer was there to say no. Deterministic egress rules could have.
  • In our replay, page data plus egress rules would have stopped almost all of them.
Every 2026 agent escape, mapped to the rule that stops it The second swarm case

The honest fine print — the same two assumptions we publish, plus two operational ones

  1. The policy engine must see every request — an agent with raw socket access or a second network path bypasses everything; enforcement belongs at the egress proxy/network layer, not only in an SDK hook.
  2. Default-deny must be on. In flag-only mode these become alerts within minutes rather than prevention — still a large improvement on a timeline measured in weeks (the DseWiki edits ran from late May to late June 2026, per the researchers), but not a block.
  3. For full URL+method matching on HTTPS you need to be the proxy or in-process hook — SNI alone shows only the host, which still catches the entire host-list layer.
  4. Policy can’t read intent inside a legitimately allowed action: an agent whose job is publishing packages keeps registry access. In our replay of the 2026 incidents, no crossing fits any plausible allowlist for the agents’ documented tasks.
Related

Keep reading

FAQ

Autonomous agent questions

What are the risks of autonomous AI agents?
Longer exposure to mistakes, compounding errors, more untrusted input, wider reach, fading instructions and late detection. Autonomy makes existing risks larger and harder to see.
How do you control an agent nobody watches?
Replace human approvals with deterministic rules: deny action pages and unknown hosts, set task budgets, use short-lived credentials and alert on denial spikes.
When is full autonomy acceptable?
When outside actions are denied or pre-approved, mistakes are reversible, data is low sensitivity and someone is alerted when limits are hit.
Why do instructions stop working in long runs?
As context grows, early instructions carry less weight. Rules enforced outside the model do not fade.
Are multi-agent systems riskier?
Yes. One agent's output becomes another's input, so a single manipulation can spread. Check each agent's requests against its own policy.
How do I move an agent to more autonomy?
One level at a time. Run it with approvals, turn every refusal into a rule, then let routine steps run alone and watch escalations fall before removing the last approvals.
Do vendor autopilot modes count as autonomous agents?
Yes. Register them with their autonomy level, limit their reach in the product settings and ask the vendor for logs.
Where do I start?
Deny action pages at the egress point for every autonomous agent. Page types for 40M+ domains come from an AI agent allow list, from $99 a month.

Give autonomous agents rules that never fade

Action pages and unknown hosts denied before every request.

Build a tier 4 policy