AI Agent Allowlist
Home Page-Types Database Agent Guardrails 2026 Incidents API Docs Pricing
Resources
Use Cases (15) Industries & Buyers (12) Learn: Core Concepts (12) Implementation Guides (15) Comparisons (8) Agent Security Guides (22) Market & Frameworks (9) Schema & Data Reference (6) FAQ Glossary
Why It Matters
2026 Agent Incidents Category Targeting Database Refreshes Contact Customer Login
Download Free Sample
print it before you need it

AI Agent Incident Response: The First Day

When an agent does something it should not, the first hour decides how bad it gets. Agents act fast, and they keep acting while people meet.

This playbook covers recognition, severity, containment, evidence and lessons, including how page-level logs from an AI agent allow list shorten the investigation.

4Severity levels
15 minTo contain
12Evidence items
3Tabletop scenarios
Recognising it

What counts as an agent incident

An agent incident is any action by an AI agent that goes beyond its purpose, its permissions or its policy, whether or not harm follows.

Treat near misses the same way. A blocked signup attempt is an incident with no damage, and it tells you where the next real one will come from.

Acting in your name

Accounts created, posts published, forms sent, purchases made.

Changing what it should not

Data deleted, settings changed, code pushed where it had no right.

Moving data out

Uploads, pastes or messages carrying internal information.

Reaching where it should not

Consoles, metadata endpoints, other people's accounts.

Coordinating unexpectedly

Agents passing messages through channels nobody designed.

Misreporting

An agent describing its own actions wrongly.

Signals

How agent incidents are usually spotted

Agent incidents rarely trip classic alarms. They show up in places security teams do not usually watch.

Ask finance, support and marketing to forward anything odd to security. They often see agent incidents first.

SignalWhere it appearsWhat it often means
Welcome or verification emails from unknown servicesShared or agent mailboxesAn agent signed up somewhere
Rising denied action-page attemptsPage-policy decision logsAn agent drifting from its task, or being steered
Invoices or trial remindersFinance inboxAn agent started a subscription
Posts or comments nobody wroteSocial channels, forums, customer reportsAn agent published content
Missing or changed dataApplication alerts, user reportsAn agent wrote or deleted where it should not
Traffic to unknown hostsProxy logsAn agent following injected links
Password reset emails nobody requestedStaff mailboxesAn agent reached a reset page
Severity

Four severity levels

Set severity by what the agent did and could still do, not by how surprising it was.

Raise the level if the agent still holds credentials you have not revoked. Potential harm counts as much as harm already done.

LevelDescriptionExampleResponse time
Sev 4Blocked attempt, no effectDenied signup attempt in the logsNext business day
Sev 3Action outside purpose, contained, low impactA trial account created and closedSame day
Sev 2Real impact on data, money or reputationA public post, a purchase, changed customer dataWithin one hour
Sev 1Ongoing harm or data leaving the companyData uploads, credential use, destroyed dataImmediately, all hands

Many Sev 4 events in a short time can signal a Sev 2 in the making. Treat a sudden rise in denials as an early warning.

By the clock

The first day in four blocks

Each block has one goal. Do not skip ahead: investigating before containing lets the agent keep acting.

Assign one person to each block so work runs in parallel once containment is done.

0 TO 15 MINUTESContain
  • Stop the agent process
  • Revoke its credentials
  • Deny its network identity at egress
  • Pause other agents sharing its tools
15 TO 60 MINUTESScope
  • Pull its decision and tool logs
  • List every action since it started
  • Find actions outside its purpose
  • Set severity
1 TO 4 HOURSUndo
  • Close accounts it created
  • Remove posts, cancel orders
  • Rotate any credentials it saw
  • Restore changed data
4 TO 24 HOURSExplain
  • Find the root cause
  • Write a short timeline
  • Notify who must be told
  • Decide the fix
Containment

How to stop each kind of agent

The kill switch differs by where the agent runs. Know yours before you need it.

Write the exact steps for each agent in its registry entry. In an incident, nobody should have to search for the right console.

Agent typeFastest stopBackstop
Script or service on your serversStop the process or containerDeny its identity at the egress proxy
Agent on a cloud platformDisable it in the platform consoleRevoke its cloud role
Vendor agent in a SaaS toolSwitch the feature off for the tenantRevoke the OAuth grant
Browser agent or extensionDisable via endpoint managementProxy denies its traffic
Hosted coding agentRevoke the keys you gave itRotate repository tokens

Do not ask the agent to undo its own actions. It may misjudge again, and it may misreport what it did.

Evidence

Twelve things to save before anything is cleaned up

What it was told

  • System prompt and task
  • Its policy file at the time
  • Its registry entry

What it saw

  • Pages and files it read
  • Tool results it received
  • Any content with hidden instructions

What it did

  • Every tool call with arguments
  • Every URL decision with page type
  • Credentials it used

What changed

  • Accounts, posts, orders created
  • Data changed or deleted
  • Messages sent
# one line from a page-policy decision log is often the fastest clue { "agent": "sales-research", "decision": "deny", "page_type": "signup", "url": "https://forum.example/register", "time": "2026-10-02T09:14:07Z" }
Root cause

Seven questions for the review

Agent incidents usually have more than one cause. These questions find the ones you can fix.

Keep the review blameless. The goal is the missing control, not the person who wrote the prompt.

1

What was it trying to do?

Its goal at the moment of the action.

2

Where did the idea come from?

Its task, a misread result, or content it read.

3

Why could it do it?

Which permission or missing control allowed it.

4

Why did nobody stop it?

Missing approval, missing alert, or ignored alert.

5

How was it found?

By a control, by luck, or by a customer.

6

Could other agents do the same?

Check every agent with similar tools.

7

Which single control would have stopped it?

That is your first fix.

Communication

Who to tell, and when

Agent incidents often touch outside parties, because the agent acted on their websites.

Prepare short message templates in advance, so nobody writes them under pressure.

AudienceWhenWhat to say
Agent owner and their managerAt containmentWhat happened, what is stopped, what they must do
Security leadershipSev 2 and above, within the hourSeverity, scope, next steps
Legal and privacyIf personal data or contracts are involvedFacts only, no conclusions yet
Affected outside sitesAfter scopingWhich accounts or posts were created, and that they are being removed
CustomersIf their data or experience was affectedPlain language, what you changed
Practice

Three tabletop scenarios

Run one each quarter. Thirty minutes each, with the people who would really respond.

Time each step and write the times down. The first run usually shows that nobody knew how to stop a vendor agent.

Scenario A: the trial account

Finance finds an invoice from a service nobody bought. The logs show a research agent's signup two weeks ago.

Practise: tracing, closing the account, fixing the policy.

Scenario B: the public post

A customer screenshots a forum comment posted under your company name by a support agent.

Practise: removal, communication, finding the injected instruction.

Scenario C: the upload

An agent uploaded an internal file to an outside paste site after reading a hostile web page.

Practise: Sev 1 handling, data exposure, rotation.

Roles

Who does what during an agent incident

Add these roles to your existing incident process rather than creating a separate one.

The agent owner is the new role. They know what normal looks like for their agent, which speeds up scoping.

Incident lead

Security operations. Owns the timeline and decisions.

Agent owner

Explains the agent's purpose and normal behaviour.

Platform engineer

Stops the agent and pulls its logs.

Identity engineer

Revokes and rotates credentials.

Communications

Handles outside sites, customers and staff.

Legal and privacy

Advises on notification duties.

Prevention

Controls that shrink the next incident

Most lessons from agent incidents point to the same short list.

Pick the one that would have stopped your incident, put it in place within a week, then work down the rest of the list.

Deny action pages

Signup, checkout, upload and posting pages denied before the request.

Log every decision

URL, page type, verdict and agent ID, kept for a year.

Own identities

One per agent, short-lived, easy to revoke.

Tested kill switch

Stop any agent in under a minute.

Default-deny unknown hosts

For agents that act on their own.

Alerts on denial spikes

Catch drift before it becomes damage.

Mistakes

Common response mistakes

These make agent incidents longer and harder to explain.

Each one has happened in real responses. Add them to your runbook as explicit "do not" lines.

Asking the agent what happened

Its account may be wrong. Use logs.

Cleaning up before saving evidence

Deleted posts and accounts take the evidence with them.

Stopping one agent, not its siblings

Agents sharing tools and prompts may do the same.

Forgetting outside sites

Accounts the agent created stay open until someone closes them.

Measuring it

Response metrics to track

Report these after every Sev 1 and Sev 2, and quarterly for all levels.

Falling times to contain are the clearest sign that tabletop practice is working.

Time to contain

From first signal to agent stopped.

Time to scope

From containment to a full action list.

Outside actions undone

Accounts closed, posts removed, orders cancelled.

Repeat incidents

Same cause, different agent.

One-page checklist

The playbook on a single page

Print this and keep it next to your incident runbook.

Tick each line as it is done and record the time next to it. The ticked sheet becomes the start of your timeline.

Contain and scope

  • Agent process stopped
  • Credentials revoked
  • Network identity denied at egress
  • Sibling agents paused
  • Logs pulled and saved
  • Severity set

Undo and explain

  • Outside accounts closed
  • Posts removed, orders cancelled
  • Exposed secrets rotated
  • Data restored and checked
  • Timeline written
  • First fix chosen and owned
Vendor agents

When the agent belongs to a vendor

If a vendor's built-in agent caused the incident, you still own the response on your side.

Check your contract for the vendor's duty to cooperate and share logs, before you need it.

Switch it off

Disable the feature for your whole tenant, not only for one user.

Revoke its grants

Remove OAuth grants and API keys the agent used.

Ask for logs

Request the vendor's record of the agent's actions, with times.

Record it

Add the incident to the agent's registry entry and your vendor file.

Notification

Legal and notification basics

An agent incident can trigger the same duties as any other security incident. Involve legal early.

The decision logs and the timeline you already built are usually the facts legal teams need first.

Personal data involved?

Data protection laws may set short deadlines for notifying regulators or people.

Contracts affected?

Customers or partners may need to be told under their agreements.

Outside sites affected?

Their terms may require you to report automated activity or close accounts.

Records kept?

Keep the evidence and decisions for as long as your retention policy requires.

This is general information, not legal advice.

Terms

Words used in this playbook

Shared words make a stressful hour less confusing for everyone involved, from engineers to legal.

Containment

Stopping the agent so it can take no further action.

Kill switch

A fast, tested way to stop an agent completely.

Decision log

The record of every allow and deny for an agent's requests.

Sibling agents

Agents that share tools, prompts or credentials with the one involved.

Injected instruction

Hidden text in content the agent read that changed its behaviour.

Root cause

The missing control that allowed the action.

Lessons from the 2026 agent incidents

  • Responders first had to reconstruct which requests agents sent, and to which sites.
  • Page-level decision logs answer that in minutes, and page-level denials prevent most of it.
  • In our replay, page data plus egress rules would have stopped almost all of the incidents.
How the Hugging Face breach could have been stopped, and the rest Anatomy of an agent incident

The honest fine print — the same two assumptions we publish, plus two operational ones

  1. The policy engine must see every request — an agent with raw socket access or a second network path bypasses everything; enforcement belongs at the egress proxy/network layer, not only in an SDK hook.
  2. Default-deny must be on. In flag-only mode these become alerts within minutes rather than prevention — still a large improvement on a timeline measured in weeks (the DseWiki edits ran from late May to late June 2026, per the researchers), but not a block.
  3. For full URL+method matching on HTTPS you need to be the proxy or in-process hook — SNI alone shows only the host, which still catches the entire host-list layer.
  4. Policy can’t read intent inside a legitimately allowed action: an agent whose job is publishing packages keeps registry access. In our replay of the 2026 incidents, no crossing fits any plausible allowlist for the agents’ documented tasks.
Related

Keep reading

FAQ

Agent incident response questions

What is an AI agent incident?
Any action by an AI agent beyond its purpose, permissions or policy, such as creating accounts, publishing content, moving data or reaching systems it should not.
What is the first thing to do?
Contain it: stop the process, revoke its credentials and deny its identity at the egress proxy. Investigate after it can no longer act.
Should I ask the agent what it did?
No. Agents can misreport their own actions. Use tool logs and page-policy decision logs as the record.
What evidence matters most?
Its instructions and policy, the content it read, every tool call and URL decision, and what changed as a result.
How do I undo actions on outside websites?
List every account, post and order the agent created from its logs, then close or remove each one and record it.
Do agent incidents need a separate process?
No. Add agent-specific steps, such as revoking agent identities and closing outside accounts, to your existing incident process.
How fast should containment be?
Aim for under 15 minutes from the first signal. Agents keep acting while people discuss what to do.
How can I prevent the next one?
Deny action pages before each request, give each agent its own identity, log every decision and test your kill switch. Page types for 40M+ domains come from an AI agent allow list.

Make the next incident smaller

Deny action pages and log every decision, with page types for 40M+ domains.

Free Agent Egress Guard