When agents hand work to other agents, each message crosses a trust boundary. Most multi-agent systems treat those messages as if they came from a trusted colleague.
This guide lists ten agent-to-agent risks and the controls that hold, including page-type checks through an AI agent allow list for every URL one agent passes to another.
A2A stands for agent-to-agent. One agent sends a task, a question or a result to another agent, often built by a different team or company.
Open protocols now let agents discover each other, describe their skills and exchange tasks without a human in between.
A2A security risks are the ways an attacker, a faulty agent or a careless design can misuse the messages, delegation and trust between agents.
They add to the risks of each single agent. They do not replace them.
Below is a common pattern. A planner agent splits a task and hands parts to specialist agents.
Each agent does its job well, which is why the problem is hard to see.
If one agent reads a poisoned page, its output becomes the next agent's instructions.
Splits the task and delegates.
Reads a page with hidden instructions.
Trusts the summary, including a planted link.
Opens the link. It is a checkout page.
No single agent did anything it was forbidden to do. The failure lives in the trust between them.
These are the risks that appear most often in multi-agent designs. Each is short to describe and easy to miss in review.
Several can combine in one incident, as the chain above shows.
An attacker presents itself as a trusted agent, with a copied name or a fake description of its skills.
Hidden instructions from a web page travel inside one agent's output into the next agent's input.
A low-privilege agent asks a high-privilege agent to act for it, and the request is honoured.
A trusts B, B trusts C, so A ends up trusting C without ever checking it.
Confidential context is sent to an outside agent that only needed a small part of it.
An agent returns a URL or file that another agent opens, uploads or acts on.
A long-running task is taken over, cancelled or redirected by a message that looks legitimate.
Agents call each other in circles, burning budget and hitting outside sites again and again.
An agent says it can do something safely, and the caller believes the description without testing it.
Each agent logs its own step, but nobody can follow one request across every hop.
Many A2A attacks finish the same way. Some agent in the chain opens a URL and takes an action on it.
The egress point does not care which agent suggested the URL. It checks the page type and applies the policy of the agent making the request.
Good A2A controls do not rely on any agent behaving well. They sit outside the models and check every message or request.
Start with the first four. They cover most of the ten risks.
Mutual authentication between agents, with each agent's own identity. No anonymous peers.
Keep a list of agents each agent may talk to. Discovery is not permission.
Another agent's output is data, never instructions with authority.
Page types decide whether the requesting agent may open it, wherever it came from.
An agent acts with its own rights, never the rights of the agent that asked.
Share only what the next agent needs, especially with outside agents.
Cap hops per task, calls per minute and spend per task.
One trace ID carried across every hop, logged by every agent and the egress point.
Use this table in design reviews. A risk with no control against it is a gap to close before launch.
Most risks need two controls. One stops the attack, the other shows it happened.
| Risk | Main controls |
|---|---|
| A2A-01 Impersonation | Authenticate every agent, allow only known peers |
| A2A-02 Injection passed along | Messages as untrusted input, URL checks at egress |
| A2A-03 Confused deputy | No transitive privilege |
| A2A-04 Transitive trust | Allow only known peers, no transitive privilege |
| A2A-05 Data leaking | Send the minimum context |
| A2A-06 Planted links and files | URL checks at egress, deny upload page types |
| A2A-07 Task hijacking | Authenticate every agent, trace end to end |
| A2A-08 Loops and cost | Limit depth, rate and budget |
| A2A-09 Overclaimed skills | Allow only known peers, tested before approval |
| A2A-10 Broken audit trail | Trace every request end to end |
The riskiest A2A links are with agents owned by suppliers, partners or strangers. You cannot inspect their prompts or their models.
A connection that made sense for one project should end when the project does.
Treat each outside agent like a new supplier.
Ask these before any multi-agent system goes live. A "not sure" is a finding.
Keep the answers with the design, so the next review starts from them.
Draw every agent-to-agent link, including outside ones.
For each hop, name the identity that acts.
Every one of those agents needs a web policy.
Depth, rate and budget limits, with numbers.
List the data sent to outside agents.
Show it, from the first message to the last web request.
Most teams add A2A controls after the system already works. That is fine if it is done in order.
Each stage reduces risk on its own, so you can stop and ship between them.
List every agent, every peer it talks to, and every agent that opens URLs.
All agent web traffic goes through the egress proxy, tagged with the agent ID.
Deny the 8 action page types for every agent that has no need for them.
Add peer allowlists, hop limits and budgets, then trace IDs.
These show up again and again in first multi-agent builds. Each one quietly widens trust.
Each one is cheap to fix early and costly after an incident.
Nobody can tell which agent acted.
Any hijack of the planner reaches everything.
Another agent's message can override them.
An internal agent that read a bad page is no longer trustworthy.
Loops run until the budget runs out.
A trusted domain still has checkout and signup pages.
These numbers show whether trust between agents is under control. Review them monthly.
A sudden change in any of them is worth a look the same day.
Messages from agents not on the peer list.
Which agent suggested the denied URL.
A sign of loops or steering.
Share of tasks with a full end-to-end trace.
A single agent has one prompt, one identity and one set of tools. A chain of agents multiplies each of those.
The table shows where the extra risk comes from.
| Area | Single agent | Multi-agent chain |
|---|---|---|
| Where input comes from | User and web pages | User, web pages and every other agent |
| Who acts | One identity | Several identities, sometimes from other companies |
| How far an injection spreads | One agent | Every agent downstream |
| Cost of a loop | One agent retrying | Agents calling each other without end |
| Tracing an action | One log | Many logs, joined only by a shared trace ID |
Multi-agent systems cross team lines by design. Ownership has to be written down, or every gap belongs to nobody.
Record these owners in the agent registry next to each link.
| Task | Owner |
|---|---|
| Approving a new agent-to-agent link | Owners of both agents |
| Peer allowlists and authentication | Platform team |
| Web page policy for each agent | Network security |
| Data allowed to leave for outside agents | Data protection |
| Hop, rate and budget limits | Platform team with finance |
| Monthly review of A2A metrics | Security with agent owners |
A2A controls meet a few familiar objections. Short answers keep the design review moving.
Each answer points back to a control above.
Internal agents still read outside pages. One poisoned page is enough to turn one of them.
Protocols move messages and can authenticate peers. They do not decide what a message is allowed to make an agent do.
Set the limit from real traces, then raise it where a task truly needs more.
Filters catch known patterns. A page-type check at egress still holds when a filter misses.
Full A2A controls take time. These four steps cut the most common risks quickly.
None of them needs a new platform.
So logs, peer lists and web policies can tell agents apart.
Signup, checkout, upload and posting pages, at the egress point.
Even a generous one stops endless loops.
Add it to every message and every web request header.
The honest fine print — the same two assumptions we publish, plus two operational ones
28 page types, including 8 action types, on 40M+ domains.