Blog

/Research

Agent skills security in 2026

Scanners flag 4% to 42% of agent skills; sandbox checks confirm 0.16% as malicious. One trojanized family hit 1.7M installs, and 1 of 7 local agents sandboxes by default.

·13 min read

Summarize in ChatGPT
Agent skills security in 2026: a bar chart of the share of skills.sh skills each scanner failed (GPT-5.3 as a judge 27.28%, Cisco Skill Scanner 14.04%, Gen agent-trust-hub 13.76%, Snyk 7.69%, Socket 3.79%), with all five agreeing on 0.12% (33 of 27,111 skills); 0.16% of 98,380 skills confirmed malicious by sandbox checks; 1.7M+ installs of one trojanized skill family on skills.sh
On this page

An agent skill is a folder with a SKILL.md file: a name, a description, instructions the agent follows when the task matches, and often scripts it runs. Nothing in that format separates reading from doing. The instructions steer an agent that can already read your files, run your shell and reach the network, and the scripts run with the same permissions. Anthropic's docs say it plainly: a malicious skill "can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose", and its enterprise guide rates bundled scripts "High: scripts run with full environment access".

That makes skills a software supply chain, just one written mostly in English. We wanted to know how bad it is in numbers, so we read every academic scan of the public registries we could find, the incident reports from security vendors, and each major agent's documentation on what it does by default when it loads a skill. Security vendors sell scanners, so their figures are their own claims and are labelled that way. We also counted a few things ourselves. Every number links to its source.

Key findings

  1. 1
    The attack usually lives in the prose. In 157 skills an academic team confirmed as malicious, 84.2% of the vulnerabilities sat in SKILL.md text and 8.5% in code.
  2. 2
    "Flagged" and "malicious" are very different numbers. Scanners fail between 3.8% and 41.9% of skills, depending on the scanner and the registry, in one study of 238,180 skills. Sandbox-verified malware runs from 0.16% of 98,380 skills to 0.49% of a 17,022-skill sample.
  3. 3
    Scanners disagree, and they miss things. On skills.sh, Gen failed 13.76% of skills, Snyk 7.69% and Socket 3.79%; all five scanners tested agreed on only 33 skills (0.12%). Cisco's own README reports 7.75% recall on held-out samples.
  4. 4
    The malware comes in campaigns. Koi's count of malicious skills on ClawHub went from 341 to 824 in two weeks, and Antiy traced 1,184 to just 12 author IDs.
  5. 5
    It reached the mainstream registry. A look-alike skill family was listed clean on skills.sh, trojanized six days later and collected 1.7M+ aggregate installs before removal, according to Zenity.
  6. 6
    Injected skills work. Hidden instructions succeeded 41% to 79% of the time across frontier models in one benchmark, and an automated attack reached 95.1% against 10.9% for naive injection.
  7. 7
    Local agents trust skills by default. Of seven we checked, only Codex runs commands in an OS sandbox with the network off out of the box. In Claude Code, allowed-tools pre-approves tools rather than restricting them, and 32 of the 38 skills we found declaring it pre-approve unscoped Bash.
  8. 8
    There is no signing. The spec has no security section, and four proposals for signatures, permissions, provenance and verification were closed without a change.

0.16%

Skills confirmed malicious by sandbox checks, of 98,380 in two registries

157 skills

Source: Liu et al., USENIX Security 2026

0.12%

skills.sh skills flagged by all five scanners tested

33 of 27,111

Source: Holzbauer et al.

1.7M+

Installs of one trojanized skill family on skills.sh, Jul 2026

Not user-unique

Source: Zenity Labs

1 of 7

Local agents that sandbox commands with network off by default

Codex

Source: host0 review of vendor docs, Sep 2026

A skill runs with your agent's permissions#

A skill has three layers. The name and description sit in the agent's context all the time. The body of SKILL.md loads when the agent decides the skill matches. Any other file, a script or a reference document, loads only if the instructions point to it. None of these layers needs to contain code to be dangerous, because the agent is the code: it will run whatever command a convincing instruction asks for. 1Password's Jason Meller put it in four words: in skill registries, "Markdown is an installer".

The confirmed malware bears that out. A team at USENIX Security 2026 crawled 98,380 skills, flagged 4,287 candidates with static analysis, then ran them in a sandbox and confirmed 157 as malicious, with 632 vulnerabilities between them. Most of those vulnerabilities were in the Markdown.

Where the attack lives in confirmed malicious skills

Share of 632 vulnerabilities in 157 sandbox-confirmed malicious skills, by file type, crawled Jan 2026

Where the attack lives in confirmed malicious skills
LabelValue
SKILL.md prose84.2%
Executable code8.5%
Config and other7.3%
The highlighted bar is plain-language instructions. A code scanner that skips the Markdown misses most of the attack surface.

Source: Liu et al., “Do Not Mention This to the User”, USENIX Security 2026

The most common technique in that set was remote script execution, the curl … | bash pattern, at 25.2% of vulnerabilities, followed by behaviour manipulation (18.8%) and credential harvesting (17.7%). The paper's title is a phrase found in one of them: "do not mention this to the user." Anthropic's position has been consistent since the launch. When Cato CTRL showed in 2025 that a modified version of Anthropic's own GIF Creator skill could fetch and run ransomware after a single approval, Anthropic replied, as quoted by Cato, that "It is the user's responsibility to only use and execute trusted Skills." Its docs define trusted narrowly: skills "you created yourself or obtained from Anthropic".

Flagged is not the same as malicious#

Headlines about skill security usually quote a scanner's fail rate, and those are high. The academic scans that confirm behaviour in a sandbox find far less malware. Both are real numbers; they measure different things, and they should never share an axis.

The widest comparison is Holzbauer et al., who collected 238,180 unique skills from ClawHub, skills.sh, SkillsDirectory and GitHub and compared the verdicts of the scanners each marketplace already runs. On ClawHub, OpenClaw's own scanner flagged 41.93% of skills and VirusTotal 36.20%. On skills.sh, the three audit partners Vercel added in February gave very different answers about the same registry.

Same registry, five verdicts

Share of skills.sh skills each scanner failed, study published Mar 2026, revised Jun 2026

Same registry, five verdicts
LabelValue
GPT-5.3 as a judge (researchers)27.28% (research baseline)
Cisco Skill Scanner14.04% (of 52,577 scanned)
Gen agent-trust-hub13.76% (of 62,163 scanned)
Snyk7.69% (of 46,414 scanned)
Socket3.79% (of 56,695 scanned)
Highlighted bars are the three scanners skills.sh runs. Each scanner covered a different subset of the registry; of the 27,111 skills all five analysed, they agreed on 33 (0.12%).

Source: Holzbauer et al., arXiv 2603.16572, Table 2

Agreement is the telling number. Only 33 of the 27,111 skills every scanner looked at were flagged by all five. And when the same researchers checked each flagged skill against its repository, only 15 of 2,887 flagged skill–repository pairs stayed suspicious. Most flags were false alarms or noise, not attackers.

Confirmed rates come from studies that run the skills, or have a human review them, before calling them malicious.

Confirmed malicious, by study

Share of each sample its authors confirmed as malicious, not just flagged, Feb to Jun 2026

Confirmed malicious, by study
LabelValue
Koi, ClawHub at the peak of ClawHavoc11.9% (341 of 2,857, Feb 2026)
Snyk, ClawHub and skills.sh≤1.9% (76 of 3,984; partial review)
Credential-leak study, SkillsMP sample0.49% (83 of 17,022)
ClawScan, ClawHub skill versions0.31% (206 of 67,453; registry verdict)
USENIX study, skills.rest and SkillsMP0.16% (157 of 98,380)
ClawHub during the February campaign is the outlier. On large general samples, confirmed malware is well under 1%: rare, but in registries of six figures that's still hundreds of skills.

Sources: Koi (archived); Snyk; arXiv 2604.03070; OpenClaw; arXiv 2602.06547

The two USENIX-paper numbers show how lossy static scanning is on its own: 4,287 static candidates became 157 confirmed, and the authors put a static-only baseline's precision at 1.1%, against 99.6% once behaviour was verified. They also call 157 a lower bound, because many unconfirmed samples carried dormant triggers that didn't fire in a 60-second run. All 157 were removed by the registries after disclosure.

Most risky skills are careless, not malicious#

The larger problem by volume is negligence. Liu et al. analysed 31,132 skills from two marketplaces in December 2025 and found that 26.1% contained at least one vulnerability pattern; they put the true rate between 23% and 30%. Only 5.2% showed high-severity patterns suggesting malicious intent. In a hand-checked sample, the most common patterns were unpinned dependencies, collecting API keys from environment variables, and excessive permissions.

Vulnerabilities in skills, by category

Share of 31,132 skills from skills.rest and SkillsMP with at least one pattern, Dec 2025 data

Vulnerabilities in skills, by category
LabelValue
Any category26.1%
Data exfiltration13.3%
Privilege escalation11.8%
Supply chain7.4%
Prompt injection0.7%
A skill can fall into several categories, so the bars don't add up. Prompt injection looks rare here partly because pattern detectors struggle to see it, which the authors say themselves.

Source: Liu et al., “Agent Skills in the Wild”, Table 8

Scripts raise the odds. Skills that bundle scripts were vulnerable 40.6% of the time, against 24.2% for instruction-only skills. The paper reports that as an odds ratio of 2.12; as a plain ratio of rates it's 1.68 times. But only 11.5% of skills ship scripts, so by our arithmetic about 82% of the vulnerable skills had none. Removing scripts doesn't remove the risk.

Credentials are where negligence turns into damage. A study of 17,022 skills sampled from SkillsMP found that 3.1% leak credentials, with 84.0% of the issues down to developer negligence, and that 89.6% of those leaks were exploitable in a sandbox. The biggest single cause was debug logging, 73.5% of issues, "because agent frameworks capture stdout into the LLM context window": a script that prints a token hands it to the model, and to anything that can talk to the model. Holzbauer's team found a related supply-chain gap: 121 skills point at seven abandoned GitHub repositories that anyone could re-register.

Scanners disagree, and they miss#

The two biggest open registries now scan. ClawHub has run every skill through VirusTotal since 7 February and added a three-scanner gate with NVIDIA in June; skills.sh added Gen, Socket and Snyk audits on 17 February and hides flagged skills. OpenClaw's own announcement was candid about the limit: "A carefully crafted prompt injection payload won't show up in a threat database."

Vendors' recall numbers and independent tests sit far apart. Socket says its scanner has 98.7% recall on its own test set. Cisco is unusually open about its scanner: its README, as of 26 September, reports 31.43% recall on its development benchmark and 7.75% on held-out sources, and warns that "A scan that returns no findings does not guarantee that a skill is free of all threats."

Scanner recall: self-reported vs held-out

Share of known-malicious skills each scanner detected, by test set, Feb to Sep 2026

Scanner recall: self-reported vs held-out
LabelValue
Socket, vendor's own test set98.7% (382 malicious samples)
Cisco Skill Scanner, dev benchmark31.4% (Cisco README)
Cisco Skill Scanner, held-out sources7.75% (Cisco README, 839 samples)
Cisco Skill Scanner, no LLM judge, MaliciousSkillBench2.5% (academic, 839 samples)
NVIDIA SkillSpector, MaliciousSkillBench0% (academic, 839 samples)
Highlighted bars are tests on malware the tools hadn't been tuned on. Recall falls by an order of magnitude once the samples come from new sources.

Sources: Socket; Cisco skill-scanner README, 26 Sep 2026; MaliciousSkillBench, arXiv 2608.19901

Evasion is cheap. JFrog found a skill that padded its README to about 22 MB, past the 10 MB cut-off of ClawDex, VirusTotal and ClawHub's scanner, and collected over 5,000 downloads. Trail of Bits, which sells security audits, beat ClawHub's detector, Cisco's scanner and all three skills.sh scanners with four skills, and concluded that "No amount of scanning or LLM analysis can reliably detect malicious content in agent skills." In academic tests, packing a skill into a self-extracting archive beat all eight scanners tried more than 90% of the time, and splitting one attack across several harmless-looking skills beat six scanners 96.0% of the time.

The campaigns, from ClawHavoc to Paperclip#

The first large wave hit ClawHub, the registry for the OpenClaw agent, in the last week of January. At the start of February, Koi Security audited all 2,857 skills on it and found 341 malicious, 335 of them from one campaign it named ClawHavoc: a skill with a fake "prerequisites" step that had the user or the agent run an obfuscated command that installed the Atomic macOS Stealer. VirusTotal summed up the technique: "the malware is the workflow". By Koi's 16 February update, ClawHub had grown to over 10,700 skills and Koi's count to 824. Antiy CERT counted 1,184 malicious skills in ClawHub's history from just 12 author IDs; the top two uploaded 677 and 390, 90% of the total by our arithmetic.

Skill malware, from proof of concept to campaigns

Selected disclosures and incidents, Dec 2025 to Aug 2026

  1. 2 Dec 2025

    Ransomware via a modified Anthropic skill

    Cato CTRL's lab proof of concept: one approval, then a hidden helper fetches and runs MedusaLocker.

    Source: Cato CTRL

  2. 26 Jan 2026

    A backdoored skill faked to #1

    Jamieson O'Reilly inflates downloads to 4,000+ in an hour; 16 executions in 7 countries in 8 hours.

    Source: Paubox

  3. 1 Feb 2026

    ClawHavoc

    Koi: 341 of 2,857 ClawHub skills malicious, 335 from one stealer campaign.

    Source: Koi (archived)

  4. 2 Feb 2026

    The #1 skill was a dropper

    1Password finds ClawHub's top-downloaded skill delivered macOS malware.

    Source: 1Password

  5. 6 Feb 2026

    1,184 malicious skills in ClawHub's history

    Antiy CERT traces them to 12 author IDs.

    Source: Antiy

  6. 16 Feb 2026

    341 becomes 824

    Koi's recount as ClawHub passes 10,700 skills.

    Source: Koi (archived)

  7. 23 Feb 2026

    2,200+ malicious skills on GitHub

    Trend Micro, while tracking 39 ClawHub skills that push a new stealer variant.

    Source: Trend Micro

  8. 6 Mar 2026

    A 22 MB README beats the scanners

    JFrog: padding pushed the payload past scanners' size limits.

    Source: JFrog

  9. 24 Mar 2026

    A download counter anyone could call

    Silverfort's 20,000+ fake downloads put a test skill at #1 in its category: 3,900 executions in 6 days.

    Source: Silverfort

  10. 20 Jul 2026

    800+ fake skills and MCP servers on GitHub

    Island's AgentBaiting campaign; coding agents recommended the bait unprompted.

    Source: Island

  11. 6 Aug 2026

    1.7M+ installs on skills.sh

    Zenity: look-alike Paperclip and Browser Use skills, trojanized after listing clean.

    Source: Zenity Labs

Researchers found the first problems; by February the registry was being flooded. The open question moved from whether skills can carry malware to whether registries can keep up.

The uploads were industrial. Snyk saw daily ClawHub submissions go from under 50 in mid-January to over 500 by early February. In the USENIX sample, one actor accounted for 54.1% of the confirmed malicious skills, using templated brand impersonation. And the bait spread beyond skill registries: Island found about 7,600 malicious GitHub repositories, more than 800 of them posing as AI skills or MCP servers, and showed Claude Code, Gemini and ChatGPT finding and in some cases recommending them without being given a link.

Gaming the install counter#

Install counts are the main trust signal in a skill registry, and they were easy to fake. In January Jamieson O'Reilly published a backdoored skill and pushed it to 4,000+ downloads within an hour, the most downloaded on what was then ClawdHub. In March Silverfort found that ClawHub's download counter was a public endpoint anyone could call; 20,000+ fake downloads put its test skill at #1 in its category, which led to 3,900 executions in six days, and OpenClaw fixed it in under 24 hours. In May Orca reported that an unnamed "prominent" marketplace counted installs from unauthenticated requests. skills.sh's privacy page now describes hourly deduplication of install events by IP hash and TLS fingerprint.

Paperclip: listed clean, weaponized later#

The best-documented case on a mainstream registry is the one Zenity Labs, which sells agent security, published on 6 August. Every step below is dated in its write-up.

How the Paperclip look-alike skills worked

Steps of the campaign on skills.sh and GitHub, Jul to Aug 2026, per Zenity Labs

  1. 2 Jul

    A look-alike identity

    A domain one letter off (getpaperclipp.com) and a GitHub org one word off the real project, with a full clone of its source.

  2. 5 Jul

    Listed clean

    Byte-identical copies of the real skills appear on skills.sh and start collecting installs.

  3. 11 Jul

    Trojanized

    One malicious setup-installation.md is copied into seven skill paths. That day the family is #8 on skills.sh Trending, with 46.6K installs across six skills.

  4. On use

    The payload is a second file

    SKILL.md stays benign. It sends the agent to setup-installation.md only when Paperclip needs installing, so a review of SKILL.md alone finds nothing.

  5. On use

    A remote stage

    The agent runs curl with TLS checks off against a fake health endpoint, decodes base64 and runs a Node script.

    • Harvest

      138 credential paths: SSH keys, cloud credentials, kube and Docker configs, npm and PyPI tokens, .env files.

    • Exfiltrate

      One encoded bundle is posted out, then the stage deletes itself.

  6. 2 Aug

    Takedown

    Vercel and GitHub remove the skills, listings and repositories within 12 hours of Zenity's report. The counter stands at 1.7M+ aggregate installs.

Two things defeated the registry's audits: the skills were clean when listed, and the malicious instructions lived in a file the agent only opened at install time. Zenity notes the counters “are not user-unique”.

Source: Zenity Labs

Zenity is careful about the scale: the 1.7 million installs are counter events, inflated by skills that reference each other in one repository, and it doesn't know how many people ran the payload. What the case shows is the pattern. A skill earns installs and rankings while it's clean, then changes, and an audit that ran at listing time doesn't see the change.

Injected skills work#

Even without code, a skill is a prompt the agent treats as trusted guidance. Skill-Inject, an academic benchmark of 202 injection-task pairs, found that instructions hidden in otherwise legitimate skills succeeded between 41.0% (Haiku 4.5) and 79.0% (Gemini 3 Flash) of the time when they were phrased to fit the task. Hiding the payload in a script rather than in SKILL.md raised that to 71.8%, from 53.8%. Even obviously harmful instructions, such as deleting files or piping a remote script to a shell, often went through.

Obviously harmful skill instructions that agents followed

Attack success rate by model on Skill-Inject's 82 obvious injections, single run, Feb 2026

Obviously harmful skill instructions that agents followed
LabelValue
GPT-5.1-Codex-Mini68.3%
Gemini 3 Flash62.2%
Claude Sonnet 4.546.3%
GPT-5.2-Codex42.7%
Gemini 3 Pro42.7%
GPT-5.218.3%
Claude Opus 4.515.9%
Claude Haiku 4.58.5%
These are the crude attacks: rm -rf, ransomware, curl piped to bash. Injections written to fit the task did better on every model. All models were tested in February 2026; newer ones haven't been measured on this benchmark.

Source: Skill-Inject, arXiv 2602.20156, Table 3

Automated attacks do better still. SkillJect generates poisoned skills that split the payload across instructions and helper files. It averaged 95.1% success against 10.9% for naive direct injection, and on Claude Sonnet 4.5 it went from 5.0% to 97.5%. A separate test of two coding agents running malicious skills found Gemini CLI exploited in 95.5–96.1% of runs, with the agent recognising the danger in under 2% of runs.

What agents do by default#

So what stands between a skill and your machine? We read the documentation of seven local agents and three hosted ones, as published in late September, using archived copies where we could find them. The hosted surfaces all run skills in a container or VM. The local agents mostly don't.

What agents do when they load a skill

Out-of-the-box defaults from each vendor's documentation: OS sandbox, network access, how allowed-tools is treated, and whether repository skills wait for workspace trust, Sep 2026

AgentSandboxNetworkallowed-toolsRepo skills gated
Claude CodeNo; opt in with /sandboxLike any local programPre-approves; doesn't restrictNo
CodexYes, workspace-writeOffNot documentedMay start read-only until trusted
CursorIn some run modes; commands need approvalBlocked when sandboxedNot in its frontmatterNo; trust is off by default
GitHub Copilot CLINo; opt in, previewAsks before URLsPre-approves; warns against shellNot documented
VS Code agent modeNo; opt in, previewUnrestricted unless sandboxedNot in its frontmatterYes; restricted mode disables agents
Gemini CLINo; opt in with -sUnrestricted unless sandboxedNot documentedTrust off by default; asks consent per skill
OpenClawNoHost networkNot applicableNot applicable
Claude API (hosted)Yes, containerNoneNot applicableNot applicable
claude.ai and Cowork (hosted)Yes, VMDepends on planNot applicableNot applicable
Copilot cloud agent (hosted)Yes, ephemeral VMFirewalledPre-approvesNot applicable
Codex is the only local agent whose default is an OS sandbox with the network off. Where allowed-tools is honoured, it means “run these without asking”, never “only these”. “Not documented” means we found no mention, not that the feature is absent.

Sources: Claude Code skills; Claude Code sandboxing; Codex security; Cursor agent security; GitHub Copilot CLI skills; VS Code docs; Gemini CLI docs; OpenClaw docs; Claude API skills; Copilot cloud agent firewall

Two details matter more than the table suggests. First, Claude Code's sandbox, when it's on, covers shell commands only: file tools, MCP servers and hooks run outside it, and if the sandbox can't start, commands run unsandboxed unless failIfUnavailable is set. Second, repository skills load without a trust prompt in Claude Code, Cursor and Gemini CLI by default, so cloning a repo can install skills. The spec's own guide for implementers suggests gating project skills on folder trust because a cloned repository may be hostile.

allowed-tools widens, it doesn't narrow#

The only security field in the spec is allowed-tools, marked "Experimental. Support for this field may vary". Its name suggests an allowlist. In Claude Code it's the opposite: it grants the listed tools without a prompt for that turn, and "It does not restrict which tools are available". The docs add that "Workspace trust doesn't gate this field", even in a non-interactive run in a folder you've never trusted. GitHub's Copilot docs treat it the same way and warn that pre-approving shell or bash "can allow attacker-controlled skills or prompt injections to execute arbitrary commands".

We checked how skills use the field in five well-known repositories, at their last commit before 29 September.

How often skills pre-approve tools

SKILL.md files declaring allowed-tools in five public skill repositories, last commit before 29 Sep 2026

RepositorySkillsUse allowed-toolsUnscoped Bash
anthropics/skills2000
obra/superpowers1500
vercel-labs/agent-skills900
openai/skills4400
trailofbits/skills853832
Total17338 (22.0%)32
Four of the five big publishers don't use the field at all. Where it is used, it mostly pre-approves the whole shell: 32 of 38 skills (two of them test fixtures) list Bash with no scope, which in Claude Code lets any shell command run without a prompt while the skill is active.

Source: host0 count of git checkouts: anthropics/skills, obra/superpowers, vercel-labs/agent-skills, openai/skills, trailofbits/skills

That isn't a criticism of Trail of Bits, whose skills run security tools that need a shell. It shows what the field does in practice: when an author uses it, it usually removes the prompt that would otherwise stand between the skill and your shell. Anthropic has been closing the gaps for organizations. Over 2026 Claude Code fixed skill allowed-tools bypassing managed ask rules (March), added a disableSkillShellExecution setting (April) and a restricting disallowed-tools field (May), and in late September stopped repository and user skills, then marketplace plugins, from pre-approving their own tools when the managed allowManagedPermissionRulesOnly setting is on.

No signatures, only proposals#

Other software supply chains have spent years adding provenance: after the Shai-Hulud attack in 2025, GitHub announced a move for npm to 2FA publishing, seven-day tokens and trusted publishing. Skills have almost none of it. The spec has no signing, digest or security section, and its repository has closed four proposals that would have added one.

Skill provenance: proposals and partial fixes

Spec proposals and the registry and vendor measures around them, Mar to Sep 2026

  1. 16 Mar 2026

    Signature and permissions proposals closed

    A signature block in SKILL.md (#247) and least-privilege permission metadata (#249).

    Source: agentskills #247

  2. 24 Mar 2026

    Discovery RFC requires sha256 digests

    The .well-known/agent-skills draft makes clients verify each file's digest and not run scripts by default. Its pull request to the spec is unmerged.

    Source: Cloudflare RFC

  3. 19 May 2026

    NVIDIA signs its verified skills

    OpenSSF Model Signing over every file. The same day, a provenance-fields proposal (#358) is closed in the spec repo.

    Source: NVIDIA

  4. 30 Jun 2026

    Verification proposal closed

    #418, “Spec lacks guidance on skill verification and supply-chain trust”, opened 9 Jun.

    Source: agentskills #418

  5. 6 Aug 2026

    Anthropic scans skills for Enterprise

    Skill and plugin scanning in beta for claude.ai and Cowork uploads; a failing skill is blocked.

    Source: Claude release notes

  6. 24 Sep 2026

    Claude Code locks down pre-approvals

    Under a managed setting, repository and user skills can no longer pre-approve their own tools.

    Source: Claude Code changelog

The fixes are happening in distribution and in clients, not in the format. A digest proves a file wasn't changed in transit; it says nothing about who wrote it or whether it's safe.

Issue #418 put the gap in one sentence: "there's nothing in the format that lets a consumer verify a skill is what it claims to be" Anthropic says its Enterprise scanning will turn on by default on 2 October for organizations that haven't set it, but it covers skills uploaded to claude.ai and Cowork; skills uploaded through the Skills API aren't scanned, and neither are the skills a local agent reads from disk.

What extension stores have that skills don't

Publisher, install and takedown controls by ecosystem, as documented by Sep 2026

EcosystemSigning and publisher checksAt installAfter a takedown
VS Code MarketplaceEvery extension signed by the marketplaceSignature verified by defaultMalicious extensions force-uninstalled
Agent skills (the spec)NoneNothing requiredUp to each registry
ClawHubGitHub account at least a week old; each bundle hashedScanned at publish, re-scanned dailyBlocked from download
skills.shNo publisher authenticationAudited by three scannersHidden from search and the leaderboard
VS Code extensions are also code from strangers, but the marketplace signs every one and can pull it back from machines that installed it. Skills leave all of that to each registry, and copies may persist downstream after a takedown.

Sources: Microsoft (VS Code Marketplace); Agent Skills specification; OpenClaw; Holzbauer et al.; Vercel; Zenity Labs

What this means if you build with AI#

Skills are useful because they're cheap: a Markdown file and maybe a script, installed with one command. The same property makes them a supply chain with almost no controls. The data points to a few habits.

  1. Read a skill before you install it, all of it. In confirmed malware, 84.2% of the attack sits in prose, and the Paperclip payload sat in a second file SKILL.md only pointed to. Look for "prerequisite" install steps, curl … | bash, base64 blobs and URLs you don't recognise.
  2. Treat allowed-tools as a grant, not a fence. Never accept a skill you didn't write that pre-approves unscoped Bash or shell. If you write skills, scope the field (Bash(git:*)) or leave it out.
  3. Turn the sandbox on. In Claude Code that's /sandbox, plus failIfUnavailable so it can't silently fall back; VS Code, Gemini CLI and Copilot CLI have opt-in sandboxes too. Keep the network allowlist empty or narrow, so a stolen credential has nowhere to go.
  4. Check a repository's skill folders before you run an agent in it. Repo skills load without a trust prompt in Claude Code, Cursor and Gemini CLI by default, and Claude Code applies their allowed-tools in folders you've never trusted.
  5. Pin what you reviewed. A review is of one version. Install from a commit or a digest, and re-review on update: skills listed clean and changed later are the pattern, not the exception.
  6. Don't print secrets in skill scripts. Agents read stdout into context, which is how debug logging became 73.5% of credential-leak issues. Pass tokens through the environment and never echo them.
  7. Use scanners as a tripwire, not a gate. They catch known payloads; on new ones, recall can drop to single digits. An install count is not a review: counters have been faked to #1.
  8. For teams, allowlist the sources. Claude Code and VS Code support managed strictKnownMarketplaces and strictPluginOnlyCustomization; Claude Code's allowManagedPermissionRulesOnly stops skills pre-approving tools. Trail of Bits' advice is curated, internal marketplaces over open registries.

host0's own deploy path is a skill: a SKILL.md and a bash script, publish.sh, that calls the host0 REST API with an API key for your account, and by default refuses to send that key to any other host. That means everything above applies to us too, so read the script before you install it; it declares no allowed-tools, and your agent should ask before running it.

Methodology#

This post draws on three research passes: academic studies and registry scans, incident reports, and each agent's documented defaults. Each was limited to primary sources: arXiv papers, vendor research posts, official documentation and changelogs, and GitHub repositories. Aggregator and "statistics" sites were used only as leads. Before publishing we re-opened the source behind every headline number and confirmed the quoted sentence or figure was still there. Documentation we could only read in its current form is described in the present tense.

Some numbers are our own:

  • Client defaults: read from each vendor's documentation as it stood just before 29 September (archived copies or repository commits where we could find them). "Not documented" means we found no mention.
  • allowed-tools usage: we cloned five repositories, checked out the last commit before 29 September, parsed the frontmatter of every SKILL.md (173 files, test fixtures included), and counted Bash entries with no scope.
  • Arithmetic: the 1.68 risk ratio and the 82% share of vulnerable skills without scripts come from Liu et al.'s published rates; the 0.49% malicious rate from the credential paper's 83 of 17,022; Antiy's top-two share from its per-author counts.

We dropped or corrected several widely repeated claims: "26.1% of skills are malicious" (it's any vulnerability pattern), "skills with scripts are 2.12 times more likely to be vulnerable" (an odds ratio), "only 0.52% of skills are suspicious" (that's 15 of 2,887 flagged pairs), "1.7 million victims" (installs, not people), vendor rates with no published sample size, and a set of aggregator statistics about "affected users" with no primary source.

Limitations:

  • Rates aren't comparable across studies. Each measured a different registry at a different time with a different definition. Flagged, vulnerable, harmful and confirmed malicious are different things.
  • Security vendors sell what they measure. Snyk, Socket and Gen run skills.sh's audits; Koi, Zenity, Silverfort, JFrog, Trend Micro, Antiy, Island, Cato, Cisco and Trail of Bits sell security products or services. Only the academic papers have nothing to sell.
  • Install counts are telemetry. No incident report gives a count of unique victims.
  • Model results age fast. The attack benchmarks tested models available in early 2026.
  • Our repository count is small and author-biased. 173 skills from five publishers is a sample, not a census.

Open questions#

  1. Is ClawHub cleaner now? Koi found 11.9% malicious in February and ClawScan 0.31% in June, but the methods differ. No registry has published a series on one method.
  2. How many skills.sh skills are flagged or hidden? Vercel hasn't published an aggregate, and the Paperclip family passed its audits until reported.
  3. How many people ran the payloads? Every incident reports installs or executions, never unique machines.
  4. How do commercial scanners compare on the same unseen malware? Socket's 98.7% and Cisco's 7.75% come from different test sets; nobody has run them all on one public held-out set.
  5. Does Codex enforce allowed-tools? Its docs don't say.
  6. Will the spec adopt digests or signing? Four proposals were closed, and the .well-known digest pull request is still open.
  7. How common is pre-approved Bash across a whole registry? Our 173-skill sample can't answer that; a crawl of the most-installed skills could.
ResearchSecurityAgent skills