| A1 Direct misuse & agent jailbreaks | 110 | Curated harmful-task suites in sandboxed tools, web or OS; jailbreak templates transferred to agents | Benchmarked; used in vendor system cards | Chat refusal training transfers poorly: GPT-4o browser agent attempted 98/100 harmful behaviors; CUAHarm up to 90% harmful-task success; SafeArena GPT-4o completes 34.7%. | Most suites score attempts, not realized harm; judge agreement F1 about 0.76–0.79. | AgentHarm (ICLR'25) BrowserART OS-Harm CUAHarm |
|---|
| A2 Indirect prompt injection | 75 | Hand-crafted templates; gradient, fuzzing, RL and LLM-optimizer attacks; crowdsourced corpora; benign-looking goals | Most mature attack line | Static: InjecAgent GPT-4 24%; single-turn frontier 0–1%. Adaptive: 50–96% against published defenses; PIMiner 86.7% vs Gemini-2.5-Pro on AgentDojo; Gray Swan: 60k+ violations, nearly every agent broken within 10–100 queries. | No agreed adaptive budget; ASR definitions differ (attempt vs partial vs full goal). | Greshake et al. InjecAgent AgentDojo The Attacker Moves Second AutoDojo ART / Gray Swan |
|---|
| A3 Tool, MCP & skill supply chain | 46 | Tool-description poisoning, rug pulls, optimized tool documents, malicious servers and skills, ecosystem measurement | Concept to large in-the-wild measurement in about 18 months | MCPTox 72.8% ASR (o1-mini), refusals under 3%; Skill-Inject up to 80%; 157 malicious skills among 98,380; stronger models more vulnerable (MSB). | Duplicated taxonomies; scanners flag 96.9% of servers with under 50% true positives. | MCPTox MCP Security Bench (ICLR'26) Malicious Agent Skills in the Wild (USENIX Sec'26) MCPZoo |
|---|
| A4 Memory, RAG & knowledge poisoning | 40 | Knowledge-corruption optimization, trigger optimization in embedding space, query-only memory injection, persistent/compositional poisoning | Benchmarked; moving to persistent state | PoisonedRAG 90% with 5 texts; AgentPoison above 80% at under 0.1% poison rate; skill-evolution poisoning 91%. | Headline ASRs often measured with empty memories; realistic memory cuts success sharply. | PoisonedRAG (USENIX Sec'25) AgentPoison MINJA MemPoison |
|---|
| A5 Backdoors & training-time | 40 | Reasoning-step and observation-triggered backdoors; emergent misalignment from narrow fine-tuning | Prototype | Agents 'suffer severely' from backdoors; textual backdoor defenses fail. | Few agent-specific backdoor defenses. | Watch Out for Your Agents Emergent Misalignment |
|---|
| A6 Multi-agent propagation & collusion | 29 | Self-replicating prompts (worms), infectious jailbreaks, agent-in-the-middle, orchestrator hijack, emergent collusion | Prototype, toy topologies | Orchestrator hijack: arbitrary code execution 58–90% even when individual agents refuse; infectious jailbreak across 1M agents. | No realistic A2A/MCP-delegation benchmark; propagation rarely measured at scale. | Morris II worm Agent Smith (ICML'24) MAS execute arbitrary code Open Challenges in Multi-Agent Security |
|---|
| A7 Computer-use, GUI & web agents | 16 | Environmental injection, adversarial pop-ups and ads, pixel perturbations, step-decomposed injections | Benchmarked (ICLR/NeurIPS) | Pop-ups 86% ASR; VPI-Bench up to 100% on browser agents; RTC-Bench Claude 4.5 Sonnet CUA 60%; step splitting raises GPT-5.4-mini 41.7% to 72.9%. | Attempt rates (up to 92.5%) exceed completion; security partly reflects low capability, so risk grows with capability. | EIA (ICLR'25) Adversarial pop-ups (ACL'25) RedTeamCUA (ICLR'26) WASP |
|---|
| A8 Coding agents | 15 | Poisoned repos, issues, rules and memory files; install-time supply chain; harness bypass | Evaluated on production tools (Cursor, Claude Code, Codex) | 66.5% of malicious issues pass every guardrail; AIShellJack up to 84%; injection bypasses Auto Mode in 79%; nearly every model installs untrusted npm/Cargo dependencies. | Results perishable as products change; harness and task matter as much as the model. | RedCode Your AI, My Shell IssueTrojanBench Setup Complete, Now You Are Compromised |
|---|
| A9 Privacy leakage & exfiltration | 70 | PII extraction via environment injection; data-minimization benchmarks; tool-mediated exfiltration | Benchmarked | EIA up to 70% PII-specific ASR; agents over-process private data (AgentDAM). | Contextual-integrity specifications for agents are immature. | AgentDAM EIA |
|---|
| A10 Availability & resource abuse | 20 | Malfunction amplification, loops, skill bloat, cost inflation | Prototype | Malfunction amplification above 80% failure; self-examination detection fails. | Denial-of-wallet largely unmeasured. | Breaking Agents: malfunction amplification |
|---|
| A11 Agents as attackers | 69 | Exploit-generation agents, CTF benchmarks, CVE reproduction, bug-bounty tasks, multi-host campaigns | Benchmarked; real-world use documented | GPT-4 exploited 87% of one-day CVEs with descriptions (7% without); BountyBench exploit up to 67.5% vs detect 12.5%; CyberGym about 20% with 34 zero-days found; Incalmo 37/40 multi-host assets. | Heterogeneous yardsticks; contamination; little testing against defender deception. | LLM agents exploit one-day vulns Cybench (ICLR'25) CyberGym BountyBench |
|---|
| A12 Misalignment, sabotage & reward hacking | 30 | Constructed-incentive evaluations, sabotage-plus-monitor games, reward-hacking probes, training-intervention stress tests | Benchmarked; cited in system cards | Shutdown sabotage up to 97%; SHADE-Arena sabotage 27%; o3 covert actions 13% to 0.4% after anti-scheming training; 57.1% of runs reward-hack in BaitBench. | Artificial goal nudges; evaluation awareness confounds results. | In-context scheming Agentic Misalignment Shutdown resistance Anti-scheming training |
|---|