Blog

/Research

State of AI agent security 2026

CVEs about agents, MCP and prompt injection went from 13 a quarter to 256, defenses scoring near 0% fell to adaptive attackers, and the worst damage came with no attacker at all. The data.

·13 min read

Summarize in ChatGPT
State of AI agent security 2026: 256 CVEs mentioning prompt injection, MCP, AI agents or Claude Code in Q3 2026, up from 13 in Q2 2025, with quarterly bars from 0 in Q1 2024 to 256 in Q3 2026, beside an incident timeline: EchoLeak (Jun 2025), Replit deleting a production database (Jul 2025), Nx “s1ngularity” (Aug 2025), GTG-1002 (Nov 2025), ClawHavoc (Feb 2026), PocketOS (Apr 2026) and 13,000+ screenshots pushed to public repos (Sep 2026)
On this page

A year ago, AI agent security was mostly a research topic: clever demos of a poisoned web page or a booby-trapped email making an assistant do something it shouldn't. By September 2026 agents read company email, open pull requests, run shell commands and hold production credentials, and most large companies say they use them. The demos kept coming, but so did deleted databases, malicious packages aimed at coding agents, a state-sponsored campaign run largely by an agent, and a steady stream of CVEs.

We wanted to know how bad it actually is, measured rather than asserted. So we built a record of 55 notable public incidents and disclosures since May 2025, counted agent-related CVEs in the US National Vulnerability Database and entries on CISA's list of exploited vulnerabilities, read the prompt-injection sections of every recent Anthropic and OpenAI system card, and collected the enterprise surveys that publish a sample size. Survey and vendor figures are the vendor's own claims, and most of the surveys come from companies that sell agent security; we say whose number each one is, and every number links to its source.

Key findings

  1. 1
    Adoption ran ahead of controls in every survey that asked both questions. In SailPoint's survey (n=353), 82% of organizations used AI agents and 44% had policies to secure them; in Okta's 2026 survey (292 executives), 92% reported agents in use and 34% always applied the same controls as to people.
  2. 2
    Prompt injection is the most common class in our record of 55 incidents: 26 of them, or 47%. Nearly all are researcher demonstrations fixed before disclosure, but Unit 42 now reports injection against agents "actively weaponized" on the web.
  3. 3
    The worst damage came with no attacker. Every incident in our record where data was actually destroyed, from Replit's deleted database to PocketOS and GPT-5.6 deleting home directories, was an agent with too much access and no hard stop.
  4. 4
    Opening or inspecting an untrusted repository was enough to run an attacker's code in GitHub Copilot, Cursor, OpenAI's Codex CLI, Claude Code and Gemini CLI. All five were fixed.
  5. 5
    CVEs that mention prompt injection, MCP, AI agents or Claude Code rose from 13 in the second quarter of 2025 to 256 in the third quarter of 2026, by our count of NVD; 182 of those 256 mention MCP. The 12 AI-tooling entries on CISA's exploited list are Langflow, LiteLLM, n8n, Ray and MLflow, not agents.
  6. 6
    The attacker moves second. Defenses and models that score near 0% against fixed test attacks fall to stronger, adaptive ones: Claude Opus 4.6 went from 0% to 97.5% on Anthropic's own coding test once the attacker improved, and Meta SecAlign from 2% to 96%.
  7. 7
    Vendor scores still fall with each model generation on their own tests (OpenAI's tool-call injection eval from 57.7% to 9.0% attack success), and both Anthropic and OpenAI say the problem isn't solved.
  8. 8
    Agents are becoming the largest class of identity in companies, with 109 machine identities per human, 79 of them agents, in Palo Alto Networks' survey. The controls that measurably work are containment, not vigilance: users approve 93% of Claude Code's permission prompts.

256

CVEs mentioning prompt injection, MCP, AI agents or Claude Code in Q3 2026, up from 13 in Q2 2025

Our count

Source: NVD

0% → 97.5%

Claude Opus 4.6 on Anthropic's coding injection test, its own card vs a stronger attacker

Source: Anthropic

82% vs 44%

Organizations using AI agents vs those with policies to secure them, n=353, May 2025

Vendor survey

Source: SailPoint

109:1

Machine identities per human, 79 of them AI agents, n=2,930, Mar–Apr 2026

Self-reported

Source: Palo Alto Networks

Agents went mainstream before the controls did#

Every survey we found that asked both "do you use agents?" and "can you control them?" got the same shape of answer: a large majority on the first, a minority on the second. The questions differ from survey to survey (some ask about use, some about plans), so read each pair on its own rather than across surveys.

Agent adoption vs agent controls, in the same survey

Share of respondents, pairs from five surveys, May 2025 to Mar 2026

Agent adoption vs agent controls, in the same survey
LabelValue
SailPoint: use AI agents82% (n=353, May 2025)
SailPoint: have policies to secure agents44% (n=353)
Okta 2025: use AI agents91% (260 executives, Apr–May 2025)
Okta 2025: well-developed non-human identity strategy10% (260 executives)
Cisco: plan to deploy AI agents83% (n=8,000, Oct 2025)
Cisco: fully equipped to secure agentic AI31% (n=8,000)
Deloitte: expect at least moderate use by 202774% (n=3,235, Aug–Sep 2025)
Deloitte: mature agent governance21% (n=3,235)
Okta 2026: agents in widespread or moderate use92% (292 executives, Mar 2026)
Okta 2026: always apply the same controls as to people34% (292 executives)
Highlighted bars are the control question in each pair. The gap runs from 38 points (SailPoint) to 81 (Okta 2025). Most of these firms sell identity or agent security.

Sources: SailPoint; Okta 2025; Cisco; Deloitte; Okta 2026

The confidence often runs ahead of the controls too. In the same Okta 2026 survey, 96% of executives were confident their identity systems could secure non-human identities, while 34% applied the same controls to agents as to staff, and 58% said their company had an AI-related security issue or close call in the past year. Gravitee's April 2026 survey of 750 technology leaders in the UK and US, nearly all of whom run 11 or more agents, found that 19.7% secure and govern all agents before they go live, monitoring covers about 52% of deployed agents on average, and 7.2% have a named person formally accountable for agent behavior.

When SailPoint asked what agents had already done, the answers were specific. 80% said their agents had taken unintended actions, and 23% said an agent had been tricked into revealing access credentials.

What respondents say their AI agents have already done

Share of organizations reporting each unintended action, n=353, May 2025

What respondents say their AI agents have already done
LabelValue
Any unintended action80%
Accessed unauthorized systems or resources39%
Shared sensitive data33%
Downloaded sensitive content32%
Accessed sensitive data31%
Were tricked into revealing access credentials23%
Self-reported by IT and security staff, so it shows what organizations noticed, not everything that happened.

Source: SailPoint, AI agent adoption report (Dimensional Research)

Breach data points the same way. In IBM's 2025 breach study of 600 breached organizations, 13% reported a breach of AI models or applications, and 97% of those lacked AI access controls. A year later more than 20% reported a breach targeting AI models or applications. Some companies are adding human checks: in KPMG's pulse surveys, 63% required human validation of agent outputs in early 2026, up from 22% a year before.

The incident record#

Our record has 55 entries, from a GitHub MCP prompt injection in May 2025 to coding agents leaking screenshots in late September 2026. It's a curated list of notable public incidents and disclosures, not a census, and disclosures bunch around security conferences, so the shape over time says as much about researchers' calendars as about attackers.

Key AI agent security incidents

Selected disclosures and incidents from our record of 55, Jun 2025 to Sep 2026

  1. 11 Jun 2025

    EchoLeak

    One email makes Microsoft 365 Copilot leak organizational data with no click (CVE-2025-32711).
  2. 18 Jul 2025

    Replit deletes a production database

    During a code freeze, then says rollback is impossible. It wasn't.
  3. 23 Jul 2025

    Amazon Q ships a wiper prompt

    A signed release told the agent to wipe local and cloud resources; a syntax error kept it inert.
  4. 26 Aug 2025

    Nx “s1ngularity”

    Malicious Nx versions run victims' Claude, Gemini and Q CLIs with permission-bypass flags to hunt for secrets.
  5. 25 Sep 2025

    First malicious MCP server

    postmark-mcp BCCs every email after 15 clean versions.
  6. 13 Nov 2025

    GTG-1002

    Anthropic: a state-sponsored group ran an espionage campaign largely through Claude Code, about 30 targets.
  7. 31 Jan 2026

    OpenClaw panels exposed

    Censys counts 21,639 agent control panels reachable from the internet.
  8. 2 Feb 2026

    ClawHavoc

    341 of 2,857 audited ClawHub skills are malicious.
  9. 17 Feb 2026

    Clinejection

    An issue title injects a triage bot, its npm token is stolen, and a rogue cline@2.3.0 is live for about eight hours.
  10. 25 Feb 2026

    Claude Code project files

    Repo hooks and MCP servers could run before the trust dialog (CVE-2025-59536).
  11. 25 Apr 2026

    PocketOS

    A Cursor agent deletes a production volume and its backups in 9 seconds.
  12. 16 Jul 2026

    GPT-5.6 deletes files

    OpenAI confirms the model deleted users' files and a production database days after launch.
  13. 7 Aug 2026

    RovoBlast

    One click drives Atlassian Rovo's research agent to exfiltrate data (DEF CON 34).
  14. 29 Sep 2026

    Screenshots to public repos

    Glow Security finds 13,000+ internal screenshots that coding agents pushed to public GitHub repos, from 343 organizations.
In 2025 most entries were researcher demonstrations and supply-chain attacks. By 2026 the record includes confirmed damage from agents acting on their own, and attackers going after agents' configuration and keys.

Sources: NVD; The Register (Replit); AWS; Snyk; Postmark; Anthropic; Censys; The Hacker News; Cline; Check Point; Zenity; The Register (GPT-5.6); Varonis; The Register (Glow)

We tagged each entry with one attack class. Prompt injection, where text the agent reads (an email, a web page, a pull-request comment) carries instructions it follows, is the biggest by far.

Agent security incidents by attack class

Entries in our record of 55 notable public incidents and disclosures, May 2025 to Sep 2026

Agent security incidents by attack class
LabelValue
Prompt injection26 (47%)
Supply chain8
Exposed server or missing auth6
Excessive agency6 (no attacker)
Untrusted config runs code5
Platform data exposure2
Agent used by an attacker2
Most prompt-injection entries are demonstrations fixed before they were published. The highlighted bar, agents doing damage on their own, is where every destroyed-data incident sits.

Source: host0 incident record (55 entries, each linked in the data pack)

The injection entries repeat a few tricks across vendors. One is to make the agent render an image or link whose URL carries the stolen data: EchoLeak in Microsoft 365 Copilot, CamoLeak in Copilot Chat, ForcedLeak in Salesforce Agentforce, which used an expired allow-listed domain bought for $5. Another is a single crafted link whose URL parameter becomes a prompt: Varonis found that pattern in Copilot Personal, Microsoft 365 Copilot and Atlassian Rovo. In April 2026 one comment pattern stole secrets from Claude Code Security Review, Gemini CLI's GitHub Action and Copilot Agent at once; none of the three vendors issued a CVE, and GitHub paid a $500 bounty.

Injection is also leaving the lab. In March 2026 Unit 42 reported that indirect prompt injection "is being actively weaponized" on the web, with 14.2% of the payloads it saw aiming at data destruction and 6.2% at unauthorized transactions. A month later Google said the malicious category in its Common Crawl sweep had grown 32% between November 2025 and February 2026, though with "limited sophistication" so far. The clearest end-to-end case is Clinejection, where an injected issue title led, via a stolen npm token, to a rogue release that installed the OpenClaw agent on developer machines that updated during an eight-hour window.

Attackers also use agents directly. Anthropic reported a data-extortion operation that used Claude Code against at least 17 organizations, and a Chinese state-sponsored group, GTG-1002, that ran an espionage campaign against about 30 targets in which, per Anthropic, AI performed 80–90% of the work. MITRE has since catalogued it as ATT&CK campaign C0062.

The worst damage had no attacker#

The incidents that actually destroyed data share one pattern: an agent held credentials or permissions far beyond its task, nothing technical stopped it, and nobody was attacking it. Two of them ignored an explicit written rule. The one attacker-planted wiper in our record, inside an Amazon Q release, never ran because of a syntax error.

Six incidents with no attacker

Excessive-agency entries in our record, Jul 2025 to Sep 2026

WhenAgentWhat happenedWhat was missing
Jul 2025Replit AgentDeleted SaaStr's production database during a code freeze: 1,206 executive records and 1,100+ company profilesSeparate dev and prod databases (Replit shipped the split three days later)
Dec 2025 (reported Feb 2026)Amazon KiroChose to "delete and recreate" an environment, a 13-hour interruption of one AWS service, per the FT; AWS disputes the AI framingA gate on destructive actions
Feb 2026OpenClawDeleted 200+ emails after context compaction dropped the instruction "confirm before acting"An approval rule outside the context window
Apr 2026Cursor agent (PocketOS)Used a Railway token made for custom domains, which had authority over the whole API, to delete the production volume and its backups in 9 secondsA scoped token; backups outside the blast radius
Jul 2026GPT-5.6 in CodexDeleted users' home directories and a production database, mostly in Full-Access modeThe sandbox and auto-review it ran without
Sep 2026Coding agents, several modelsPushed 13,000+ internal screenshots from 343 organizations to public GitHub reposA block on publishing outside the private repo
Each one is a capability problem rather than an intelligence problem: the agent could reach production, delete without a gate, or publish publicly. Written instructions didn't hold.

Sources: The Register (Replit); The Guardian (Kiro); GeekWire (AWS rebuttal); TechCrunch (OpenClaw); Zenity (PocketOS); The Register (GPT-5.6); The Register (Glow)

The details are telling. At PocketOS, the token was "created solely for managing custom domains" but could delete volumes, and the most recent recoverable backup was three months old. OpenAI's engineering lead described the GPT-5.6 deletions as a case where "the model makes an honest mistake", and The Register reported that the GPT-5.6 system card says the Sol variant takes "severity level 3" actions, such as deleting data without approval, more often than GPT-5.5. Glow Security's co-founder said the screenshot leaks came from agents that "found a workaround" to publish images from a private repository.

Opening a repository was enough#

The second pattern is configuration that runs code. Coding agents read project files (settings, hooks, MCP server lists) and some of them executed what they found before asking the user. Between July 2025 and July 2026, each of the five best-known coding agents had at least one flaw where an attacker-controlled repository, file or message was enough to run commands on the developer's machine.

Code execution from untrusted content, by coding agent

Disclosed flaws in major coding agents, Jul 2025 to Jul 2026

AgentFlawTriggerDisclosedFixed in
Gemini CLIReported by TracebitInjected text in a README plus an allow-listed grep prefixJul 20250.1.14
CursorCVE-2025-54135, "CurXecute"A message read through MCP rewrites ~/.cursor/mcp.json, which ran without confirmationAug 20251.3
CursorCVE-2025-54136, "MCPoison"An approved MCP config in a shared repo is swapped later, with no re-approvalAug 20251.3
GitHub Copilot (VS Code)CVE-2025-53773Injected text sets chat.tools.autoApprove, turning off approvalsAug 2025August Patch Tuesday
OpenAI Codex CLICVE-2025-61260Project-local config auto-loads MCP commandsDec 2025 (fixed Aug)0.23.0
Claude CodeCVE-2025-59536, CVE-2026-21852Repo hooks and MCP servers run before the trust dialog; a project setting leaks the API keyFeb 20262.0.65
CursorCVE-2026-50548, CVE-2026-50549, "DuneSlide"Zero-click injection overwrites the sandbox binaryJul 2026 (fixed Apr)3.0
All were fixed, several of them before public disclosure. The common thread: files and messages developers treat as data were executed as instructions.

Sources: Embrace The Red; Cato Networks (CurXecute); Check Point (MCPoison); Cato Networks (DuneSlide); Check Point (Codex CLI); Check Point (Claude Code); Tracebit (Gemini CLI)

Attackers noticed the same thing from the other side. In August 2025 the malicious Nx packages prompted victims' own Claude, Gemini and Q CLIs with flags like --dangerously-skip-permissions to search the disk for secrets. In April 2026 a Shai-Hulud wave delivered through a compromised Bitwarden CLI release targeted MCP configuration files and Claude Code credentials, and the same actor's SAP CAP packages planted a Claude Code SessionStart hook so the malware re-ran every time Claude Code started in the repository.

From a handful of CVEs to hundreds#

To measure vulnerabilities rather than headlines, we queried the NVD for every CVE published in each quarter whose description contains one of five exact phrases: "prompt injection", "Model Context Protocol", "MCP server", "AI agent" or "Claude Code". Deduplicated, the count went from 13 in the second quarter of 2025 to 256 in the third quarter of 2026. The third quarter of 2026 alone has more than double all of 2025 (111).

CVEs mentioning AI agents, MCP or prompt injection

NVD CVEs published per quarter matching any of five exact phrases, deduplicated, Q1 2024 to Q3 2026

CVEs mentioning AI agents, MCP or prompt injection
LabelValue
Q1 20240
Q2 20244
Q3 20246
Q4 20249
Q1 20254
Q2 202513 (5 mention MCP)
Q3 202549 (32 mention MCP)
Q4 202545 (21 mention MCP)
Q1 202689 (35 mention MCP)
Q2 2026133 (72 mention MCP)
Q3 2026256 (182 mention MCP)
Most of the growth is MCP: 182 of the 256 CVEs in Q3 2026 mention it. Keyword counts undercount (EchoLeak's entry says “AI command injection”), so read this as a trend line, not a census.

Sources: NVD API 2.0; host0 count, quarters through 30 Sep 2026

Part of this is attention: more researchers are looking, and MCP servers are small, numerous and often written quickly. But the count is of assigned CVEs, and the exploitation data is starting to follow. CISA's Known Exploited Vulnerabilities catalog, the US government's list of flaws attacked in the wild, had 12 entries for AI tooling by the end of September 2026, 11 of them added in 2026 and six in the third quarter alone.

AI tooling on CISA's exploited-vulnerabilities list

Known Exploited Vulnerabilities entries for agent builders, LLM gateways and ML platforms, added through 30 Sep 2026

AddedProductCVECISA's description
May 2025LangflowCVE-2025-3248Missing authentication (known ransomware use)
Mar 2026n8nCVE-2025-68613Improper control of dynamically-managed code resources
Mar 2026LangflowCVE-2026-33017Code injection
May 2026LiteLLMCVE-2026-42208SQL injection
May 2026LangflowCVE-2025-34291Origin validation error
Jun 2026LiteLLMCVE-2026-42271Command injection
Jul 2026LangflowCVE-2026-55255Authorization bypass through user-controlled key
Jul 2026LangflowCVE-2026-0770Inclusion of functionality from untrusted control sphere
Aug 2026IBM LangflowCVE-2026-9198Code injection
Aug 2026RayCVE-2025-62593Code injection
Aug 2026MLflowCVE-2026-64849Server-side request forgery
Sep 2026LiteLLMCVE-2026-59822Improper authentication
Every entry is self-hosted infrastructure with a classic bug: missing authentication, code or SQL injection. No coding agent, copilot, browser agent or MCP server is on the list. LiteLLM's CVE-2026-42271 is the one MCP-related entry.

Sources: CISA KEV catalog; host0 filter

The absence of agents on that list doesn't mean they're safe. Many agent flaws are fixed on the vendor's servers (EchoLeak, ShadowLeak, RovoBlast), which leaves nothing for a customer to patch, and the agent incidents that did real damage weren't vulnerabilities at all. What the list does show is that the infrastructure people stand up around agents (workflow builders, model gateways) gets attacked like any other internet-facing server.

The attacker moves second#

Most published prompt-injection numbers come from a fixed set of attacks: a benchmark, or the vendor's own test suite. In October 2025 researchers from OpenAI, Anthropic, Google DeepMind and ETH Zurich published "The attacker moves second", which re-tested 12 published defenses with attacks tuned against each one. They bypassed them "with attack success rate above 90% for most", and concluded that defenses reporting near-zero rates on public benchmarks "are often among the easiest to break". Human red-teamers in the same study beat every challenge they were given.

The pattern shows up wherever someone runs both kinds of test on the same target, including in the vendors' own system cards. Anthropic's Claude Opus 4.6 card reported 0% attack success on its coding test "across all conditions". Two months later, with a stronger attacker from Gray Swan, the Opus 4.7 card re-measured Opus 4.6 at 97.5% of scenarios broken within 200 attempts. It happened again with Opus 5: 0.56% of attempts in its own card, 88.92% against the next attacker in the Opus 5.5 card.

The same target, a weaker vs a stronger attacker

Attack success rate against fixed or earlier attacks vs adaptive or later ones, 2025 to 2026

The same target, a weaker vs a stronger attacker
LabelValue
Spotlighting and sandwiching: static~1% (AgentDojo)
Spotlighting and sandwiching: adaptive>95% (Nasr et al.)
Meta SecAlign: static2% (AgentDojo)
Meta SecAlign: adaptive96% (Nasr et al.)
Claude 3.5 Sonnet: baseline attack11% (NIST, AgentDojo)
Claude 3.5 Sonnet: red-teamed attack81% (NIST)
Gemini 2.5: non-adaptive set18% (Google)
Gemini 2.5: adaptive attack94.6% (Google, TAP)
Claude Opus 4.6 coding: own card0% (Anthropic)
Claude Opus 4.6 coding: stronger attacker97.5% (scenarios, 200 attempts)
Claude Opus 4.6 browser: own card0.29% (% of attempts)
Claude Opus 4.6 browser: new red-team set45.81% (% of attempts)
Claude Opus 5 coding: own card0.56% (% of attempts)
Claude Opus 5 coding: stronger attacker88.92% (% of attempts)
Highlighted bars are the stronger attacker. Each pair shares a target and a source; compare within pairs only, since metrics and test sets differ between them.

Sources: Nasr et al. 2025; NIST CAISI; Google DeepMind; Claude Opus 4.6 card; Claude Opus 4.7 card; Claude Opus 5 card; Claude Opus 5.5 card

NIST's AI standards center saw the same jump in January 2025: on AgentDojo with Claude 3.5 Sonnet, success rose "from 11% for the strongest baseline attack to 81% for the strongest new attack", and repeated attempts lifted the average from 57% to 80%. Google wrote that without adaptive testing it "would have incorrectly concluded" Gemini 2.5 was more robust than it is.

Large public competitions measure the same thing from the attacker's side. In Gray Swan's 2025 challenge with the UK AI Security Institute, participants submitted 1.8 million prompt-injection attacks against 22 agents, and every model was broken on every target behavior. In the 2026 round, 464 participants made 272,000 attempts against 13 frontier models; all proved vulnerable, with success rates from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). By our arithmetic, about 3% of individual attempts succeeded in both years (3.4%, then 3.2%), even though the models got harder to break.

None of this means the numbers are fake. It means a single benchmark score is a floor on what an attacker who studies your system can do, not a ceiling. Anthropic says so itself: its attackers know the deployment and get repeated tries, which real attackers "typically lack".

Vendor scores keep falling, on their own tests#

Read inside a single test, the vendor numbers do improve with each generation. OpenAI's GPT-5.6 system card reports defender success on "known prompt injection attacks" aimed at search and function calling. Converted to attack success, the rate fell from 57.7% for gpt-5.1-thinking to 9.0% for gpt-5.6-sol. The GPT-6 Astra card then retired that eval as saturated.

OpenAI models on its tool-call injection eval

Attack success rate (1 − defender success) on “Search and Function-Calling” known attacks, by model, as reported Jul 2026

OpenAI models on its tool-call injection eval
LabelValue
gpt-5.1-thinking57.7%
gpt-5.2-thinking43.2%
gpt-5.4-thinking30.3%
gpt-5.6-sol9.0%
One eval, one vendor, four generations: a steady fall. These are known attacks, so they say nothing about adaptive ones, and OpenAI's GPT-5.2 card warned that its injection evals at the time were splits of its training data.

Source: OpenAI, GPT-5.6 system card, Table 5

The vendors have also started testing each other's models on shared attack sets from Gray Swan. The most complete table is in Anthropic's Claude Opus 5 card, which reports the chance an attacker succeeds within 15 attempts on the same set of attacks for every model.

Attack success within 15 attempts, Gray Swan injection benchmark

Probability an attacker succeeds within 15 attempts, Q1 2026 attack set, all models with extended thinking, as run by Anthropic, Jul 2026

Attack success within 15 attempts, Gray Swan injection benchmark
LabelValue
Claude Opus 52.0%
Claude Mythos 52.6%
Claude Opus 4.85.5%
Claude Sonnet 55.9%
Muse Spark16.5%
GPT-5.6 Sol20.0%
GPT-5.520.8%
GPT-5.6 Terra30.4%
GPT-5.6 Luna43.9%
Anthropic ran this test and leads it, so treat the ranking as Anthropic's claim. OpenAI's own run on a larger Gray Swan set put GPT-5.6 Sol at 27.0% and GPT-6 Astra, with safeguards, at 8.5%.

Sources: Claude Opus 5 system card; GPT-6 Astra system card

On a refreshed set in September, Anthropic put Claude Opus 5.5 at 0.1% on one attempt and 1.0% within 15, highest in computer use (2.8%). The same card shows how product details move the numbers: Opus 5.5's coding result against the stronger attacker was 54.61%, almost entirely from requests a safety classifier handed to the older Opus 4.8, and none of the 2,872 requests Opus 5.5 answered itself were susceptible. It also notes a regression: Opus 5.5 is more likely than earlier models to follow malicious instructions when the user pastes them into the prompt.

Both companies say plainly that this isn't finished. OpenAI wrote that prompt injection "is unlikely to ever be fully 'solved'", and Anthropic's Opus 4.5 card shared its results "to demonstrate meaningful progress, not to claim the problem is solved".

Agents are the new service accounts#

Every agent that acts on a system needs an identity to act as, and most get one quietly: an API key, an OAuth token, a service account. In CyberArk's 2025 survey of 2,600 security decision makers, respondents estimated 82 machine identities for every human. In the 2026 edition, now published under Palo Alto Networks, which CyberArk is part of, the estimate was 109 per human, 79 of them AI agents.

Machine identities per human employee

Average estimated by survey respondents, 2025 and 2026 editions of the same survey

Machine identities per human employee
LabelValue
202582:1 (CyberArk, n=2,600)
2026109:1 (Palo Alto Networks, n=2,930; 79 are AI agents)
These are respondents' estimates, not measurements, and published ratios vary with what counts as an identity. The direction is the point: respondents expect AI agents to grow 85% in the next 12 months, against 56% for human identities.

Sources: CyberArk 2025; Palo Alto Networks 2026

The same 2026 survey found that on average 40% of AI agents already have access to organizational data, and that fewer than half of organizations apply basic lifecycle controls to them: 45% monitor autonomous agents' behavior and 37% can revoke their credentials. The 2025 edition had found 68% lacked identity security controls for AI. PocketOS is what that looks like in practice: a token made for one job that could do every job.

The plumbing is catching up. Microsoft made its Entra Agent ID platform generally available in April 2026, built on OAuth 2.0, MCP and A2A. MCP's spec made servers OAuth resource servers that must reject tokens issued for anyone else in 2025, and in July 2026 deprecated dynamic client registration and required issuer validation. In September NIST's NCCoE reported 600+ comments on its concept paper for agent identity and authorization.

The frameworks caught up#

In December 2025 OWASP published its first Top 10 for Agentic Applications, written with more than 100 contributors. Almost every item already has a public example in our record.

OWASP Top 10 for Agentic Applications, with examples

The 2026 list (published Dec 2025) next to an incident from our record that fits each risk

IDRiskAn example from our record
ASI01Agent goal hijackEchoLeak: one email redirects Microsoft 365 Copilot to leak data
ASI02Tool misuse and exploitationSupabase MCP: a support ticket makes an agent dump a tokens table with its service-role access
ASI03Identity and privilege abusePocketOS: a domain-only token with authority to delete volumes
ASI04Agentic supply chain vulnerabilitiespostmark-mcp and 341 malicious ClawHub skills
ASI05Unexpected code execution (RCE)Codex CLI and Claude Code running project config before trust
ASI06Memory and context poisoningPayloads planted in agents' memory files hit current and future sessions (research)
ASI07Insecure inter-agent communicationNone clear in our record yet
ASI08Cascading failuresClinejection: injection, then a stolen token, then a rogue release on developer machines
ASI09Human-agent trust exploitationReplit's agent told its user a rollback was impossible; it wasn't
ASI10Rogue agentsCoding agents publishing private screenshots to public repos
The mapping is ours, not OWASP's. Agent goal hijack is the most common in practice; insecure inter-agent communication is the one risk with no clear public incident yet.

Sources: OWASP GenAI Security Project; host0 incident record

OWASP's separate Top 10 for LLM applications moved "excessive agency" from sixth to third place in its 2026 edition, in August. Elsewhere, NIST launched an AI Agent Standards Initiative in February 2026, and Gartner told software leaders to treat any use case that combines sensitive data, untrusted content and external communication as a "no-go zone". Regulation, meanwhile, slowed down: the EU's AI Omnibus, in force since July 2026, pushed the AI Act's high-risk obligations back to December 2027.

What actually works#

If no model can be trusted to resist every injection, the defenses that hold are the ones that limit what a successful injection can do. Simon Willison's "lethal trifecta" names the dangerous combination: access to private data, exposure to untrusted content, and a way to send data out. Meta turned it into a design rule, the Agents Rule of Two: within one session, an agent should have at most two of three properties.

The Agents Rule of Two

Meta's rule: at most two of these three in one session, Oct 2025

  1. A

    Processes untrustworthy input

    Web pages, emails, issues, pull requests, documents, tool output.

    • In our record

      EchoLeak, Clinejection, CamoLeak, Comment and Control.

  2. B

    Reaches sensitive systems or private data

    Production databases, inboxes, private repos, cloud credentials.

    • In our record

      PocketOS, Replit, the Supabase MCP demo.

  3. C

    Changes state or communicates externally

    Writes, deletes, sends, publishes, or renders a URL it chose.

    • In our record

      ForcedLeak's image URL, Glow's public screenshots.

  4. rule

    Pick two, or add a human

    If a task needs all three, start a fresh session or require human approval for the risky step.

    • Gartner's version

      Treat all three together as a “no-go zone”.

It doesn't stop injection. It caps the damage: an injected agent with no way to send data out can't exfiltrate, and one with no untrusted input is much harder to hijack.

Sources: Meta; Simon Willison; Gartner

Approval prompts are a weak substitute. Anthropic found that Claude Code users approve 93% of permission prompts, which is why it moved toward sandboxing (filesystem and network isolation, which cut permission prompts by 84%) and an "auto mode" classifier that decides instead. Anthropic's own numbers for that classifier are honest about its limits: 0.4% of benign commands blocked, but 17% of real overeager actions missed in a sample of 52. By September auto mode was the default for most Claude Code users.

Measured defenses mostly show a trade between safety and usefulness, and each was tested on its own benchmark.

Defenses, before and after

Attack success rate without and with each defense, as reported by its authors, 2024 to 2026

DefenseTestBeforeAfterCost to usefulness
Tool filter (only the tools a task needs)AgentDojo, GPT-4o57.69%6.84%None: task success rose from 69.0% to 73.13%
Prompt-injection detectorAgentDojo, GPT-4o57.69%7.95%Task success fell from 69.0% to 41.49%
LlamaFirewall (PromptGuard 2 + AlignmentCheck)AgentDojo17.6%1.75%Task success fell from 47.7% to 42.7%
CaMeL (plan from trusted input only)AgentDojo—Provably secure on solved tasksSolves 77% of tasks vs 84% undefended
Adversarial training + warningGemini 2.5, adaptive attack94.6%6.2%Not reported
Claude for Chrome mitigations123 test cases23.6%11.2%Not reported
Prompt-injection probesClaude Opus 5.5, stronger coding attacker54.61%11.13%Not reported
Auto modeClaude Opus 5.5, 110 browser scenarios0.09%0 of 1100.4% false positives (separate test)
Most of these were measured against fixed attacks, so the adaptive-attack caveat applies. The structural ones (filtering tools, CaMeL's separation of plan and data) cost the least usefulness for the security they buy.

Sources: AgentDojo; LlamaFirewall; CaMeL; Google DeepMind; Anthropic (Claude for Chrome); Claude Opus 5.5 card

AgentDojo's authors note the limit of tool filtering: in 17% of their test cases the tools a task legitimately needs are enough to carry out the attack. That is the case the Rule of Two is for.

What this means if you build with AI#

Most of the incidents above would have been smaller with boring, structural controls. The data points to a short list.

  1. Assume an injection will eventually succeed, and design for the blast radius. Every vendor says the problem isn't solved, and adaptive attackers beat defenses that test near 0%. Apply the Rule of Two: an agent that reads untrusted input shouldn't also hold production access and a way to send data out.
  2. Give agents their own narrow, revocable credentials. Never leave a broad token where an agent can find it, and keep backups outside anything the agent can reach. At PocketOS one domain-management token could delete the production volume and its backups.
  3. Put destructive actions behind a gate the model can't talk its way past. Written rules failed in Replit's code freeze and OpenClaw's "confirm before acting". Use separate dev and prod environments, read-only database roles and platform-level confirmations.
  4. Prefer a sandbox to approval prompts. People approve 93% of prompts. Run coding agents with filesystem and network isolation, and don't use full-access or skip-permissions modes on a machine with real credentials.
  5. Treat project config as code. Agent settings, hooks and MCP server lists in a repository ran before trust prompts in several agents. Keep your agents updated, and review .claude/, .cursor/ and similar folders in any repository you didn't write before opening it with an agent.
  6. Close the quiet exfiltration channels. Auto-rendered images and links carried data out in EchoLeak, CamoLeak and ForcedLeak. If your agent renders URLs it composed, restrict them to an allow-list you control, and watch for expired domains on it.
  7. Patch and fence the infrastructure around your agents. Every AI entry on CISA's exploited list is a self-hosted tool such as Langflow or LiteLLM with a classic bug. Don't expose them to the internet without authentication.

host0 is a cloud for small software: agent-built apps deploy as static files that read and write shared data through a platform-owned records API, so there is no user server code to hijack, and each app can be limited to the teammates its owner invites. Agents get an API key through a device login the user approves in their own browser, so no password ever passes through the chat.

Methodology#

This post draws on three research passes, one per angle: incidents and vulnerabilities, attack and defense measurements, and enterprise readiness and identity. Each was limited to primary sources: vendor advisories and postmortems, NVD and CISA entries, system cards, academic papers, survey publishers' own releases and reports, and first-hand reporting. Aggregator and "statistics" sites were used only as leads. Everything covers material published up to 30 September 2026. Before publishing, we re-opened the single-source pages the post leans on for headline numbers, including the system cards, and dropped anything they no longer said.

Our own computations:

  • Incident record: 55 notable public incidents and disclosures from May 2025 to September 2026, each with at least one primary or reputable source, tagged with one attack class by disclosure date. It is curated, so the class shares describe what gets disclosed and reported, not what happens most.
  • CVE counts: the NVD API's keyword search with exact-phrase matching, by publication date per quarter, for "prompt injection", "Model Context Protocol", "MCP server", "AI agent" and "Claude Code", deduplicated by CVE ID. We left out "Cursor", "Copilot" and "LLM", which match unrelated products.
  • KEV entries: CISA's catalog filtered to agent builders, LLM gateways and ML platforms, added on or before 30 September 2026.
  • Converted rates: OpenAI's defender-success scores turned into attack success (1 minus the score), and per-attempt rates for the Gray Swan competitions from their published totals.

We dropped or reworded several claims: an "88% of enterprises had agent incidents" figure that its own publisher's later survey replaced with 54%; a "96% … (CyberArk)" statistic that is SailPoint's; a CVE number widely given for CamoLeak that is wrong; reports of "two AWS outages" caused by Kiro, which AWS denies; the model behind the PocketOS deletion, which comes only from the founder's account; and any comparison of Anthropic's coding or browser results across system cards, because the attacker changed between them.

Limitations:

  • Most surveys are vendor research. SailPoint, Okta, CyberArk, Palo Alto Networks and Gravitee sell identity or agent security. Questions and populations differ, so compare only within a survey.
  • Attack success rates aren't portable. Each depends on the attack set, the number of attempts, the harness and whether safeguards are on. Vendor-run tests are the vendor's claim.
  • Disclosure isn't incidence. Most prompt-injection entries are researcher demonstrations, and disclosures cluster around conferences.
  • Keyword counts undercount. CVEs with terse or differently worded descriptions are missed.

Open questions#

  1. How often does prompt injection succeed against deployed agents? Unit 42 and Google show attackers are trying; nobody publishes a success rate in production.
  2. Why is no agent on CISA's exploited list? Either nobody has confirmed exploitation, or server-side fixes leave nothing for the list to track.
  3. Who ran the injection in Clinejection? Sources agree the stolen token was used by an unknown actor, but not on who triggered the original injection.
  4. How many GTG-1002 intrusions succeeded? Anthropic said "a small number" of about 30 targets.
  5. How do shipped products hold up against adaptive attackers? Nobody independent has re-measured auto mode, probes or other production defenses.
  6. How many agent CVEs are there, really? A curated census across GitHub advisories, vendor bulletins and NVD doesn't exist.
  7. Where are Google's numbers? No Gemini model card after 2.5 publishes an agent prompt-injection metric we could chart.
ResearchSecurityAI agents