When an agent does something it should not, the first hour decides how bad it gets. Agents act fast, and they keep acting while people meet.
This playbook covers recognition, severity, containment, evidence and lessons, including how page-level logs from an AI agent allow list shorten the investigation.
An agent incident is any action by an AI agent that goes beyond its purpose, its permissions or its policy, whether or not harm follows.
Treat near misses the same way. A blocked signup attempt is an incident with no damage, and it tells you where the next real one will come from.
Accounts created, posts published, forms sent, purchases made.
Data deleted, settings changed, code pushed where it had no right.
Uploads, pastes or messages carrying internal information.
Consoles, metadata endpoints, other people's accounts.
Agents passing messages through channels nobody designed.
An agent describing its own actions wrongly.
Agent incidents rarely trip classic alarms. They show up in places security teams do not usually watch.
Ask finance, support and marketing to forward anything odd to security. They often see agent incidents first.
| Signal | Where it appears | What it often means |
|---|---|---|
| Welcome or verification emails from unknown services | Shared or agent mailboxes | An agent signed up somewhere |
| Rising denied action-page attempts | Page-policy decision logs | An agent drifting from its task, or being steered |
| Invoices or trial reminders | Finance inbox | An agent started a subscription |
| Posts or comments nobody wrote | Social channels, forums, customer reports | An agent published content |
| Missing or changed data | Application alerts, user reports | An agent wrote or deleted where it should not |
| Traffic to unknown hosts | Proxy logs | An agent following injected links |
| Password reset emails nobody requested | Staff mailboxes | An agent reached a reset page |
Set severity by what the agent did and could still do, not by how surprising it was.
Raise the level if the agent still holds credentials you have not revoked. Potential harm counts as much as harm already done.
| Level | Description | Example | Response time |
|---|---|---|---|
| Sev 4 | Blocked attempt, no effect | Denied signup attempt in the logs | Next business day |
| Sev 3 | Action outside purpose, contained, low impact | A trial account created and closed | Same day |
| Sev 2 | Real impact on data, money or reputation | A public post, a purchase, changed customer data | Within one hour |
| Sev 1 | Ongoing harm or data leaving the company | Data uploads, credential use, destroyed data | Immediately, all hands |
Many Sev 4 events in a short time can signal a Sev 2 in the making. Treat a sudden rise in denials as an early warning.
Each block has one goal. Do not skip ahead: investigating before containing lets the agent keep acting.
Assign one person to each block so work runs in parallel once containment is done.
The kill switch differs by where the agent runs. Know yours before you need it.
Write the exact steps for each agent in its registry entry. In an incident, nobody should have to search for the right console.
| Agent type | Fastest stop | Backstop |
|---|---|---|
| Script or service on your servers | Stop the process or container | Deny its identity at the egress proxy |
| Agent on a cloud platform | Disable it in the platform console | Revoke its cloud role |
| Vendor agent in a SaaS tool | Switch the feature off for the tenant | Revoke the OAuth grant |
| Browser agent or extension | Disable via endpoint management | Proxy denies its traffic |
| Hosted coding agent | Revoke the keys you gave it | Rotate repository tokens |
Do not ask the agent to undo its own actions. It may misjudge again, and it may misreport what it did.
Agent incidents usually have more than one cause. These questions find the ones you can fix.
Keep the review blameless. The goal is the missing control, not the person who wrote the prompt.
Its goal at the moment of the action.
Its task, a misread result, or content it read.
Which permission or missing control allowed it.
Missing approval, missing alert, or ignored alert.
By a control, by luck, or by a customer.
Check every agent with similar tools.
That is your first fix.
Agent incidents often touch outside parties, because the agent acted on their websites.
Prepare short message templates in advance, so nobody writes them under pressure.
| Audience | When | What to say |
|---|---|---|
| Agent owner and their manager | At containment | What happened, what is stopped, what they must do |
| Security leadership | Sev 2 and above, within the hour | Severity, scope, next steps |
| Legal and privacy | If personal data or contracts are involved | Facts only, no conclusions yet |
| Affected outside sites | After scoping | Which accounts or posts were created, and that they are being removed |
| Customers | If their data or experience was affected | Plain language, what you changed |
Run one each quarter. Thirty minutes each, with the people who would really respond.
Time each step and write the times down. The first run usually shows that nobody knew how to stop a vendor agent.
Finance finds an invoice from a service nobody bought. The logs show a research agent's signup two weeks ago.
Practise: tracing, closing the account, fixing the policy.
A customer screenshots a forum comment posted under your company name by a support agent.
Practise: removal, communication, finding the injected instruction.
An agent uploaded an internal file to an outside paste site after reading a hostile web page.
Practise: Sev 1 handling, data exposure, rotation.
Add these roles to your existing incident process rather than creating a separate one.
The agent owner is the new role. They know what normal looks like for their agent, which speeds up scoping.
Security operations. Owns the timeline and decisions.
Explains the agent's purpose and normal behaviour.
Stops the agent and pulls its logs.
Revokes and rotates credentials.
Handles outside sites, customers and staff.
Advises on notification duties.
Most lessons from agent incidents point to the same short list.
Pick the one that would have stopped your incident, put it in place within a week, then work down the rest of the list.
Signup, checkout, upload and posting pages denied before the request.
URL, page type, verdict and agent ID, kept for a year.
One per agent, short-lived, easy to revoke.
Stop any agent in under a minute.
For agents that act on their own.
Catch drift before it becomes damage.
These make agent incidents longer and harder to explain.
Each one has happened in real responses. Add them to your runbook as explicit "do not" lines.
Its account may be wrong. Use logs.
Deleted posts and accounts take the evidence with them.
Agents sharing tools and prompts may do the same.
Accounts the agent created stay open until someone closes them.
Report these after every Sev 1 and Sev 2, and quarterly for all levels.
Falling times to contain are the clearest sign that tabletop practice is working.
From first signal to agent stopped.
From containment to a full action list.
Accounts closed, posts removed, orders cancelled.
Same cause, different agent.
Print this and keep it next to your incident runbook.
Tick each line as it is done and record the time next to it. The ticked sheet becomes the start of your timeline.
If a vendor's built-in agent caused the incident, you still own the response on your side.
Check your contract for the vendor's duty to cooperate and share logs, before you need it.
Disable the feature for your whole tenant, not only for one user.
Remove OAuth grants and API keys the agent used.
Request the vendor's record of the agent's actions, with times.
Add the incident to the agent's registry entry and your vendor file.
An agent incident can trigger the same duties as any other security incident. Involve legal early.
The decision logs and the timeline you already built are usually the facts legal teams need first.
Data protection laws may set short deadlines for notifying regulators or people.
Customers or partners may need to be told under their agreements.
Their terms may require you to report automated activity or close accounts.
Keep the evidence and decisions for as long as your retention policy requires.
This is general information, not legal advice.
Shared words make a stressful hour less confusing for everyone involved, from engineers to legal.
Stopping the agent so it can take no further action.
A fast, tested way to stop an agent completely.
The record of every allow and deny for an agent's requests.
Agents that share tools, prompts or credentials with the one involved.
Hidden text in content the agent read that changed its behaviour.
The missing control that allowed the action.
The honest fine print — the same two assumptions we publish, plus two operational ones
Deny action pages and log every decision, with page types for 40M+ domains.