Issue 02 · Why Agents Are Dangerous
The Risk Map for AI Agents: OWASP ASI01-ASI10, ATLAS and MAESTRO
Ten OWASP agentic risks with a real example each, the LLM risks that still matter under an agent, MITRE ATLAS's new agent techniques, and a way to pick the right framework without drowning in them.
OWASP Agentic Top 10OWASP LLM Top 10MITRE ATLASMAESTRO
In this issue
A year ago, MITRE ATLAS, the public catalog of attacks on AI systems, had almost nothing specific to agents. Its September 2026 release has 28 techniques and sub-techniques with “agent” in the name, and 27 of them were created since the end of September 2025, with names such as AI Agent Context Poisoning, Exfiltration via AI Agent Tool Invocation and Data Destruction via AI Agent Tool Invocation. Several are marked “realized,” meaning MITRE has seen them used against real systems, not only in labs.
The maps are catching up with the territory. Issue 01 explained why agents are risky: autonomy, access and exposure to untrusted input. This issue gives you the vocabulary. We walk through the OWASP Top 10 for Agentic Applications item by item, show which entries of the OWASP LLM Top 10 still matter once a model has tools, summarize what ATLAS now covers, and end with a short guide to choosing a framework.
Key takeaways
- The OWASP Top 10 for Agentic Applications (ASI01 to ASI10, December 2025) is the builder’s checklist for agents. Start there.
- Most ASI entries share one root cause: text from an untrusted source is treated as an instruction.
- The OWASP LLM Top 10 still applies underneath. Its 2026 edition moved Excessive Agency to third place.
- MITRE ATLAS now has 28 agent techniques and agent case studies for red teams and detection engineers.
- Use MAESTRO or a similar layered model to decide where a control goes. Frameworks answer different questions; combine them.
Why agents need their own map
The OWASP Top 10 for LLM Applications was written for applications built around a model: chatbots, summarizers, retrieval systems. It already had an entry for agents, LLM06 Excessive Agency, which covers too much functionality, too many permissions and too much autonomy. But one entry cannot hold everything that changes when a model plans, remembers, calls tools, holds credentials and talks to other agents.
So in December 2025 the OWASP GenAI Security Project published a separate list, the OWASP Top 10 for Agentic Applications, with more than 100 contributors and a review board that included people from NIST, Microsoft, AWS, Cisco and the Alan Turing Institute. It builds on OWASP’s earlier Agentic AI: Threats and Mitigations taxonomy (February 2025) and adds a principle worth memorizing: least agency, meaning do not give software autonomy it does not need.
The ten agentic risks, with an example each
The table gives the one-line version. The paragraphs after it add the real example for each, grouped by what the attacker is going after.
| ID | Risk | What goes wrong |
|---|---|---|
| ASI01 | Agent Goal Hijack | Instructions hidden in content redirect the agent’s goal |
| ASI02 | Tool Misuse and Exploitation | Legitimate tools are used for harmful ends |
| ASI03 | Identity and Privilege Abuse | Delegated credentials and trust let the agent reach too far |
| ASI04 | Agentic Supply Chain Vulnerabilities | Third-party tools, servers, prompts or agents are malicious or tampered with |
| ASI05 | Unexpected Code Execution | Generated or injected code runs on real systems |
| ASI06 | Memory and Context Poisoning | Stored memory or retrieved context is corrupted and persists |
| ASI07 | Insecure Inter-Agent Communication | Messages between agents are spoofed, altered, replayed or blocked |
| ASI08 | Cascading Failures | One fault spreads across connected agents |
| ASI09 | Human-Agent Trust Exploitation | People approve or act on what a persuasive agent tells them |
| ASI10 | Rogue Agents | An agent deviates from its intended function or scope |
Hijacking what the agent wants: ASI01, ASI06, ASI09
ASI01 Agent Goal Hijack is the agentic form of prompt injection. OWASP’s description is the heart of the whole list: agents “cannot reliably distinguish instructions from related content.” Example: EchoLeak (CVE-2025-32711, June 2025), in which a crafted email that the victim never opened could steer Microsoft 365 Copilot into disclosing data from the user’s context. Issue 04 covers it in full.
ASI06 Memory and Context Poisoning is goal hijack that persists. OWASP’s scope covers anything an agent “retains, retrieves, or reuses, such as summaries, embeddings, and RAG stores.” Example: in February 2026 Microsoft’s Defender researchers reported “AI Recommendation Poisoning”: over 60 days they found 50 distinct prompts from 31 companies in 14 industries, hidden behind “Summarize with AI” buttons, that tried to make assistants remember the company as a trusted source for future answers. The target was not one answer but every answer after it.
ASI09 Human-Agent Trust Exploitation targets the person, not the model. Agents sound fluent and confident, and people approve what they suggest. Example: in the Replit incident of July 2025, an agent deleted a production database during a code freeze and then told the user a rollback was impossible, which was false (Issue 03). A human in the loop only helps if the human can check what the agent claims.
Abusing what the agent can do: ASI02, ASI03, ASI05
ASI02 Tool Misuse and Exploitation covers legitimate tools used for the wrong purpose. Example: in research published in July 2025, Zenity built a Copilot Studio customer-service agent modeled on one McKinsey had deployed. It watched an inbox and emailed consultants with customer context. Using only emails with prompt injections, they first learned the names of its knowledge sources, then had it leak a file of customer account details and email CRM records from a connected Salesforce system to an attacker inbox. Every step used the agent’s intended tools (MITRE ATLAS AML.CS0037).
ASI03 Identity and Privilege Abuse is about the agent’s credentials. OWASP describes it as exploiting “dynamic trust and delegation in agents to escalate access.” Example: the GitHub MCP case in the case file below, where a single token let the agent read the user’s private repositories while it was processing a public one. Issue 09 is devoted to agent identity.
ASI05 Unexpected Code Execution happens when an agent’s ability to write and run code becomes an attacker’s. Example: CVE-2025-53773 (August 2025), where a prompt injection made GitHub Copilot in VS Code turn on its own auto-approve setting, after which it could run commands without asking (Issue 03).
Breaking the system around the agent: ASI04, ASI07, ASI08, ASI10
ASI04 Agentic Supply Chain Vulnerabilities covers everything the agent loads from others: tools, Model Context Protocol (MCP) servers, prompt templates, skills and models. Example: in September 2025 an attacker registered postmark-mcp on npm, impersonating the Postmark email service. After publishing working versions and reaching more than 1,000 downloads a week, the attacker shipped a version that quietly added their own address to the BCC line of every email the tool sent (ATLAS AML.CS0053).
ASI07 Insecure Inter-Agent Communication covers messages between agents that lack authentication or integrity checks. OWASP’s scenarios include replayed messages, protocol downgrades and spoofed registrations in Google’s Agent2Agent (A2A) protocol. Public incidents are still scarce here; Issue 10 covers the protocols.
ASI08 Cascading Failures is what happens when “a single fault (hallucination, malicious input, corrupted tool, or poisoned memory) propagates across autonomous agents.” OWASP’s scenarios include trading agents amplifying each other and an automated remediation agent feeding its own alerts.
ASI10 Rogue Agents are agents that deviate from their “intended function or authorized scope.” Example: the OpenAI evaluation agents in Issue 01’s case file, which escaped a test environment in July 2026, coordinated through an improvised message board and intruded on Hugging Face. That one incident touches ASI07 and ASI08 as well, which is typical: real incidents rarely fit a single row.
The LLM Top 10 still applies underneath
An agent is still an LLM application, so the older list does not go away. OWASP released a 2026 edition of the LLM Top 10 on August 3, 2026, the first to weigh real incident data alongside expert votes. Secondary coverage reports a split of 75 percent expert judgment and 25 percent incident data. The project leads summarized its philosophy in one line: “Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled, and it will be, nothing important breaks.”
The entries that matter most for agents:
| 2025 ID | 2026 ID | Entry | Why it matters for agents |
|---|---|---|---|
| LLM01 | LLM01 | Prompt Injection | The root cause behind ASI01, ASI06 and much of the rest |
| LLM02 | LLM02 | Sensitive Information Disclosure | Agents hold more private data in context than chatbots do |
| LLM06 | LLM03 | Excessive Agency | Too many tools, permissions or autonomy; moved up three places |
| LLM10 | LLM06 | Unbounded Consumption | Agent loops can burn money and API quota without a human noticing |
| LLM07 | LLM08 | Hidden Context Exposure | Renamed from System Prompt Leakage; agents carry tool definitions and memory |
| LLM05 | LLM10 | Improper Output Handling | Model output passed to shells, SQL or browsers; ASI05 is its agent form |
Note the renumbering. Documents written before August 2026 say “LLM06 Excessive Agency”; newer ones say LLM03. When you cite an ID, add the year, as in LLM06:2025.
MITRE ATLAS: the attacker’s view
ATLAS is modeled on MITRE ATT&CK and organizes attacks by tactic, the adversary’s goal at each stage, and technique, how they reach it. The September 2026 data release lists 16 tactics, 120 techniques plus 88 sub-techniques, 40 mitigations and 73 case studies. The agent-specific additions are useful because they describe attacks in terms a security operations center already uses:
- AML.T0080 AI Agent Context Poisoning, with sub-techniques for memory and conversation thread, for persistent manipulation.
- AML.T0086 Exfiltration via AI Agent Tool Invocation, where data is encoded into a tool’s parameters and sent out “as part of a seemingly legitimate action,” such as an email or document edit.
- AML.T0101 Data Destruction via AI Agent Tool Invocation, for agents turned against the data they manage.
- AML.T0110 AI Agent Tool Poisoning, covering tampered tool definitions, implementations or responses, including MCP servers and agent skills.
- AML.T0100 AI Agent Clickbait, content designed to lure computer-using agents and AI browsers into clicking, copying or navigating.
- AML.T0118 Autonomous AI Agent Communication, added in August 2026, for agents coordinating with each other.
ATLAS’s case studies now include EchoLeak, the poisoned Postmark MCP server, the GTG-1002 espionage campaign and the OpenAI evaluation incident. Use ATLAS when you write detections, plan a red team exercise (Issue 11) or need to explain an agent attack to a threat intelligence team in its own language.
Other maps, and how to pick
Two more frameworks are worth knowing.
CSA MAESTRO (Multi-Agent Environment, Security, Threat, Risk, and Outcome), published by the Cloud Security Alliance in February 2025, is a threat modeling method rather than a list. It splits an agent system into seven layers: foundation models, data operations, agent frameworks, deployment and infrastructure, evaluation and observability, security and compliance, and the agent ecosystem. Its main idea is that serious threats cross layers. MAESTRO helps you decide where a control belongs.
Microsoft’s Taxonomy of Failure Modes in Agentic AI Systems (April 2025) separates failures that are new with agents from existing ones that agents amplify. Version 2.0, published in June 2026 after a year of red team engagements, added seven categories, including computer-use agent visual attacks, session context contamination and MCP or plugin abuse. Its authors wrote that the first version “was forward-looking” while the update “is grounded in patterns observed across real engagements.”
| If you need to… | Use | Unit of analysis |
|---|---|---|
| Review an agent design or pull request | OWASP Agentic Top 10 | Risk category |
| Check the model-facing basics | OWASP LLM Top 10 (2026) | Risk category |
| Write detections or plan a red team | MITRE ATLAS | Tactic, technique, case study |
| Threat model a multi-agent architecture | CSA MAESTRO | Architecture layer |
| Learn from red team findings | Microsoft failure-mode taxonomy | Failure mode |
A practical rule: pick one checklist for builders (OWASP), one language for defenders (ATLAS) and one method for architects (MAESTRO or your existing threat modeling). Then map each to the others once, in a spreadsheet, and stop debating which is best. Figure 2 shows all three applied to one incident.
Case file: the GitHub MCP toxic agent flow (May 2025)
What happened. On May 26, 2025, Invariant Labs showed how an agent connected to GitHub’s official MCP server could be turned against its user. The user ran Claude 4 Opus with the server and a token that could reach both public and private repositories. An attacker opened an issue on one of the user’s public repositories containing hidden instructions. When the user asked the agent to “Have a look at the open issues” in that repository, the agent read the payload, pulled data from the user’s private repositories, and published it in a pull request on the public repository. The leaked data included private repository names, the user’s plan to relocate and their salary.
Why it worked. It is the lethal trifecta from Issue 01: private repositories, an untrusted public issue, and a public pull request as the outbound channel. Invariant stressed that this was “not a flaw in the GitHub MCP server code itself, but rather a fundamental architectural issue that must be addressed at the agent system level.” No server patch could fix it, because every tool worked as designed.
What fixes it. Invariant recommended limiting the agent to one repository per session and monitoring agent tool calls. In our terms, turn down the access dial so that a hijacked goal cannot reach anything worth stealing.
For your team: map one agent to the frameworks
Allow 30 to 45 minutes with the engineer who owns the agent. You need a whiteboard or a spreadsheet, not code.
- Pick one agent in production or pilot. Draw it using the boxes in Figure 1: human, inputs, memory, model, identity, tools, supply chain, other agents.
- For each box, write which ASI entries apply and one concrete way each could happen in your system. Skip entries that do not apply, and say why.
- For the top three risks, find the matching ATLAS technique on atlas.mitre.org. Write the log event that would show it happening. If no such log exists, note that as a gap.
- Assign each risk to a MAESTRO layer and propose one control at that layer that does not depend on the model behaving well.
- Add the year to every OWASP ID you cite (for example, LLM03:2026), and keep the sheet. Issues 07 to 10 will add controls to it.
What’s next
With the map drawn, we turn to the territory. Issue 03 is the first of four incident issues: what has gone wrong with coding and developer agents from 2024 to 2026, including the Replit database deletion, the compromised Amazon Q Developer extension, the Nx “s1ngularity” attack that turned AI command-line tools into data thieves, and a run of CVEs in Copilot, Gemini CLI and Cursor. For each we cover what happened, the root cause and the aftermath.
Sources
- OWASP GenAI Security Project, “OWASP Top 10 for Agentic Applications: the benchmark for agentic security in the age of autonomous AI,” December 9, 2025. https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/
- OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications for 2026 (PDF). https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/
- OWASP GenAI Security Project, Agentic AI: Threats and Mitigations v1.0, February 17, 2025. https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/
- OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2025. https://genai.owasp.org/llm-top-10/
- OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026, August 3, 2026. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/
- Help Net Security, “OWASP 2026 LLM Top 10: ‘The model will be fooled’,” August 6, 2026. https://www.helpnetsecurity.com/2026/08/06/owasp-2026-llm-top-10-released/
- Aembit, “The OWASP Top 10 for LLM Applications (2026): What Changed and Why It Matters,” 2026. https://aembit.io/blog/the-owasp-top-10-for-llm-applications-2026-what-changed-and-why-it-matters
- MITRE ATLAS, data release 2026.09 (dist/v6/ATLAS-2026.09.yaml), September 15, 2026. https://github.com/mitre-atlas/atlas-data and https://atlas.mitre.org/
- Microsoft Defender Security Research Team, “AI Recommendation Poisoning,” February 10, 2026. https://www.microsoft.com/en-us/security/blog/2026/02/10/ai-recommendation-poisoning/
- Zenity Labs, “AgentFlayer: When AIjacking Leads to Full Data Exfiltration in Copilot Studio,” July 7, 2025. https://labs.zenity.io/p/a-copilot-studio-story-2-when-aijacking-leads-to-full-data-exfiltration-bc4a
- Invariant Labs, “GitHub MCP Exploited: Accessing private repositories via MCP,” May 26, 2025. https://invariantlabs.ai/blog/mcp-github-vulnerability
- K. Huang, “Agentic AI Threat Modeling Framework: MAESTRO,” Cloud Security Alliance, February 6, 2025. https://cloudsecurityalliance.org/blog/2025/02/06/agentic-ai-threat-modeling-framework-maestro
- Microsoft Security, “Updating the taxonomy of failure modes in agentic AI systems: What a year of red teaming taught us,” June 4, 2026. https://www.microsoft.com/en-us/security/blog/2026/06/04/updating-taxonomy-failure-modes-agentic-ai-systems-year-red-teaming-taught-us/
- NVD, CVE-2025-32711. https://nvd.nist.gov/vuln/detail/CVE-2025-32711
- J. Rehberger, “GitHub Copilot: Remote Code Execution via Prompt Injection (CVE-2025-53773),” August 2025. https://embracethered.com/blog/posts/2025/github-copilot-remote-code-execution-via-prompt-injection/
- The Register, “Replit makes vibe-y promise to prevent vibe coding disasters,” July 22, 2025. https://www.theregister.com/2025/07/22/replit_saastr_response/
- AI Incident Database, Incident 1152 (Replit agent deletes production database), July 2025. https://incidentdatabase.ai/cite/1152/