An agent skill is a folder with a SKILL.md file: a name, a description, instructions the agent follows when the task matches, and often scripts it runs. Nothing in that format separates reading from doing. The instructions steer an agent that can already read your files, run your shell and reach the network, and the scripts run with the same permissions. Anthropic's docs say it plainly: a malicious skill "can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose", and its enterprise guide rates bundled scripts "High: scripts run with full environment access".
That makes skills a software supply chain, just one written mostly in English. We wanted to know how bad it is in numbers, so we read every academic scan of the public registries we could find, the incident reports from security vendors, and each major agent's documentation on what it does by default when it loads a skill. Security vendors sell scanners, so their figures are their own claims and are labelled that way. We also counted a few things ourselves. Every number links to its source.
Key findings
- 1The attack usually lives in the prose. In 157 skills an academic team confirmed as malicious, 84.2% of the vulnerabilities sat in
SKILL.mdtext and 8.5% in code. - 2"Flagged" and "malicious" are very different numbers. Scanners fail between 3.8% and 41.9% of skills, depending on the scanner and the registry, in one study of 238,180 skills. Sandbox-verified malware runs from 0.16% of 98,380 skills to 0.49% of a 17,022-skill sample.
- 3Scanners disagree, and they miss things. On skills.sh, Gen failed 13.76% of skills, Snyk 7.69% and Socket 3.79%; all five scanners tested agreed on only 33 skills (0.12%). Cisco's own README reports 7.75% recall on held-out samples.
- 4The malware comes in campaigns. Koi's count of malicious skills on ClawHub went from 341 to 824 in two weeks, and Antiy traced 1,184 to just 12 author IDs.
- 5It reached the mainstream registry. A look-alike skill family was listed clean on skills.sh, trojanized six days later and collected 1.7M+ aggregate installs before removal, according to Zenity.
- 6Injected skills work. Hidden instructions succeeded 41% to 79% of the time across frontier models in one benchmark, and an automated attack reached 95.1% against 10.9% for naive injection.
- 7Local agents trust skills by default. Of seven we checked, only Codex runs commands in an OS sandbox with the network off out of the box. In Claude Code,
allowed-toolspre-approves tools rather than restricting them, and 32 of the 38 skills we found declaring it pre-approve unscopedBash. - 8There is no signing. The spec has no security section, and four proposals for signatures, permissions, provenance and verification were closed without a change.
0.16%
Skills confirmed malicious by sandbox checks, of 98,380 in two registries
157 skills
Source: Liu et al., USENIX Security 2026
1.7M+
Installs of one trojanized skill family on skills.sh, Jul 2026
Not user-unique
Source: Zenity Labs
1 of 7
Local agents that sandbox commands with network off by default
Codex
Source: host0 review of vendor docs, Sep 2026
A skill runs with your agent's permissions#
A skill has three layers. The name and description sit in the agent's context all the time. The body of SKILL.md loads when the agent decides the skill matches. Any other file, a script or a reference document, loads only if the instructions point to it. None of these layers needs to contain code to be dangerous, because the agent is the code: it will run whatever command a convincing instruction asks for. 1Password's Jason Meller put it in four words: in skill registries, "Markdown is an installer".
The confirmed malware bears that out. A team at USENIX Security 2026 crawled 98,380 skills, flagged 4,287 candidates with static analysis, then ran them in a sandbox and confirmed 157 as malicious, with 632 vulnerabilities between them. Most of those vulnerabilities were in the Markdown.
Where the attack lives in confirmed malicious skills
Share of 632 vulnerabilities in 157 sandbox-confirmed malicious skills, by file type, crawled Jan 2026
| Label | Value |
|---|---|
| SKILL.md prose | 84.2% |
| Executable code | 8.5% |
| Config and other | 7.3% |
Source: Liu et al., “Do Not Mention This to the User”, USENIX Security 2026
The most common technique in that set was remote script execution, the curl … | bash pattern, at 25.2% of vulnerabilities, followed by behaviour manipulation (18.8%) and credential harvesting (17.7%). The paper's title is a phrase found in one of them: "do not mention this to the user." Anthropic's position has been consistent since the launch. When Cato CTRL showed in 2025 that a modified version of Anthropic's own GIF Creator skill could fetch and run ransomware after a single approval, Anthropic replied, as quoted by Cato, that "It is the user's responsibility to only use and execute trusted Skills." Its docs define trusted narrowly: skills "you created yourself or obtained from Anthropic".
Flagged is not the same as malicious#
Headlines about skill security usually quote a scanner's fail rate, and those are high. The academic scans that confirm behaviour in a sandbox find far less malware. Both are real numbers; they measure different things, and they should never share an axis.
The widest comparison is Holzbauer et al., who collected 238,180 unique skills from ClawHub, skills.sh, SkillsDirectory and GitHub and compared the verdicts of the scanners each marketplace already runs. On ClawHub, OpenClaw's own scanner flagged 41.93% of skills and VirusTotal 36.20%. On skills.sh, the three audit partners Vercel added in February gave very different answers about the same registry.
Same registry, five verdicts
Share of skills.sh skills each scanner failed, study published Mar 2026, revised Jun 2026
| Label | Value |
|---|---|
| GPT-5.3 as a judge (researchers) | 27.28% (research baseline) |
| Cisco Skill Scanner | 14.04% (of 52,577 scanned) |
| Gen agent-trust-hub | 13.76% (of 62,163 scanned) |
| Snyk | 7.69% (of 46,414 scanned) |
| Socket | 3.79% (of 56,695 scanned) |
Agreement is the telling number. Only 33 of the 27,111 skills every scanner looked at were flagged by all five. And when the same researchers checked each flagged skill against its repository, only 15 of 2,887 flagged skill–repository pairs stayed suspicious. Most flags were false alarms or noise, not attackers.
Confirmed rates come from studies that run the skills, or have a human review them, before calling them malicious.
Confirmed malicious, by study
Share of each sample its authors confirmed as malicious, not just flagged, Feb to Jun 2026
| Label | Value |
|---|---|
| Koi, ClawHub at the peak of ClawHavoc | 11.9% (341 of 2,857, Feb 2026) |
| Snyk, ClawHub and skills.sh | ≤1.9% (76 of 3,984; partial review) |
| Credential-leak study, SkillsMP sample | 0.49% (83 of 17,022) |
| ClawScan, ClawHub skill versions | 0.31% (206 of 67,453; registry verdict) |
| USENIX study, skills.rest and SkillsMP | 0.16% (157 of 98,380) |
Sources: Koi (archived); Snyk; arXiv 2604.03070; OpenClaw; arXiv 2602.06547
The two USENIX-paper numbers show how lossy static scanning is on its own: 4,287 static candidates became 157 confirmed, and the authors put a static-only baseline's precision at 1.1%, against 99.6% once behaviour was verified. They also call 157 a lower bound, because many unconfirmed samples carried dormant triggers that didn't fire in a 60-second run. All 157 were removed by the registries after disclosure.
Most risky skills are careless, not malicious#
The larger problem by volume is negligence. Liu et al. analysed 31,132 skills from two marketplaces in December 2025 and found that 26.1% contained at least one vulnerability pattern; they put the true rate between 23% and 30%. Only 5.2% showed high-severity patterns suggesting malicious intent. In a hand-checked sample, the most common patterns were unpinned dependencies, collecting API keys from environment variables, and excessive permissions.
Vulnerabilities in skills, by category
Share of 31,132 skills from skills.rest and SkillsMP with at least one pattern, Dec 2025 data
| Label | Value |
|---|---|
| Any category | 26.1% |
| Data exfiltration | 13.3% |
| Privilege escalation | 11.8% |
| Supply chain | 7.4% |
| Prompt injection | 0.7% |
Scripts raise the odds. Skills that bundle scripts were vulnerable 40.6% of the time, against 24.2% for instruction-only skills. The paper reports that as an odds ratio of 2.12; as a plain ratio of rates it's 1.68 times. But only 11.5% of skills ship scripts, so by our arithmetic about 82% of the vulnerable skills had none. Removing scripts doesn't remove the risk.
Credentials are where negligence turns into damage. A study of 17,022 skills sampled from SkillsMP found that 3.1% leak credentials, with 84.0% of the issues down to developer negligence, and that 89.6% of those leaks were exploitable in a sandbox. The biggest single cause was debug logging, 73.5% of issues, "because agent frameworks capture stdout into the LLM context window": a script that prints a token hands it to the model, and to anything that can talk to the model. Holzbauer's team found a related supply-chain gap: 121 skills point at seven abandoned GitHub repositories that anyone could re-register.
Scanners disagree, and they miss#
The two biggest open registries now scan. ClawHub has run every skill through VirusTotal since 7 February and added a three-scanner gate with NVIDIA in June; skills.sh added Gen, Socket and Snyk audits on 17 February and hides flagged skills. OpenClaw's own announcement was candid about the limit: "A carefully crafted prompt injection payload won't show up in a threat database."
Vendors' recall numbers and independent tests sit far apart. Socket says its scanner has 98.7% recall on its own test set. Cisco is unusually open about its scanner: its README, as of 26 September, reports 31.43% recall on its development benchmark and 7.75% on held-out sources, and warns that "A scan that returns no findings does not guarantee that a skill is free of all threats."
Scanner recall: self-reported vs held-out
Share of known-malicious skills each scanner detected, by test set, Feb to Sep 2026
| Label | Value |
|---|---|
| Socket, vendor's own test set | 98.7% (382 malicious samples) |
| Cisco Skill Scanner, dev benchmark | 31.4% (Cisco README) |
| Cisco Skill Scanner, held-out sources | 7.75% (Cisco README, 839 samples) |
| Cisco Skill Scanner, no LLM judge, MaliciousSkillBench | 2.5% (academic, 839 samples) |
| NVIDIA SkillSpector, MaliciousSkillBench | 0% (academic, 839 samples) |
Sources: Socket; Cisco skill-scanner README, 26 Sep 2026; MaliciousSkillBench, arXiv 2608.19901
Evasion is cheap. JFrog found a skill that padded its README to about 22 MB, past the 10 MB cut-off of ClawDex, VirusTotal and ClawHub's scanner, and collected over 5,000 downloads. Trail of Bits, which sells security audits, beat ClawHub's detector, Cisco's scanner and all three skills.sh scanners with four skills, and concluded that "No amount of scanning or LLM analysis can reliably detect malicious content in agent skills." In academic tests, packing a skill into a self-extracting archive beat all eight scanners tried more than 90% of the time, and splitting one attack across several harmless-looking skills beat six scanners 96.0% of the time.
The campaigns, from ClawHavoc to Paperclip#
The first large wave hit ClawHub, the registry for the OpenClaw agent, in the last week of January. At the start of February, Koi Security audited all 2,857 skills on it and found 341 malicious, 335 of them from one campaign it named ClawHavoc: a skill with a fake "prerequisites" step that had the user or the agent run an obfuscated command that installed the Atomic macOS Stealer. VirusTotal summed up the technique: "the malware is the workflow". By Koi's 16 February update, ClawHub had grown to over 10,700 skills and Koi's count to 824. Antiy CERT counted 1,184 malicious skills in ClawHub's history from just 12 author IDs; the top two uploaded 677 and 390, 90% of the total by our arithmetic.
Skill malware, from proof of concept to campaigns
Selected disclosures and incidents, Dec 2025 to Aug 2026
2 Dec 2025
2 Dec 2025
Ransomware via a modified Anthropic skill
Cato CTRL's lab proof of concept: one approval, then a hidden helper fetches and runs MedusaLocker.Source: Cato CTRL
26 Jan 2026
26 Jan 2026
A backdoored skill faked to #1
Jamieson O'Reilly inflates downloads to 4,000+ in an hour; 16 executions in 7 countries in 8 hours.Source: Paubox
1 Feb 2026
1 Feb 2026
ClawHavoc
Koi: 341 of 2,857 ClawHub skills malicious, 335 from one stealer campaign.Source: Koi (archived)
2 Feb 2026
2 Feb 2026
The #1 skill was a dropper
1Password finds ClawHub's top-downloaded skill delivered macOS malware.Source: 1Password
6 Feb 2026
6 Feb 2026
1,184 malicious skills in ClawHub's history
Antiy CERT traces them to 12 author IDs.Source: Antiy
16 Feb 2026
23 Feb 2026
23 Feb 2026
2,200+ malicious skills on GitHub
Trend Micro, while tracking 39 ClawHub skills that push a new stealer variant.Source: Trend Micro
6 Mar 2026
6 Mar 2026
A 22 MB README beats the scanners
JFrog: padding pushed the payload past scanners' size limits.Source: JFrog
24 Mar 2026
24 Mar 2026
A download counter anyone could call
Silverfort's 20,000+ fake downloads put a test skill at #1 in its category: 3,900 executions in 6 days.Source: Silverfort
20 Jul 2026
20 Jul 2026
800+ fake skills and MCP servers on GitHub
Island's AgentBaiting campaign; coding agents recommended the bait unprompted.Source: Island
6 Aug 2026
6 Aug 2026
1.7M+ installs on skills.sh
Zenity: look-alike Paperclip and Browser Use skills, trojanized after listing clean.Source: Zenity Labs
The uploads were industrial. Snyk saw daily ClawHub submissions go from under 50 in mid-January to over 500 by early February. In the USENIX sample, one actor accounted for 54.1% of the confirmed malicious skills, using templated brand impersonation. And the bait spread beyond skill registries: Island found about 7,600 malicious GitHub repositories, more than 800 of them posing as AI skills or MCP servers, and showed Claude Code, Gemini and ChatGPT finding and in some cases recommending them without being given a link.
Gaming the install counter#
Install counts are the main trust signal in a skill registry, and they were easy to fake. In January Jamieson O'Reilly published a backdoored skill and pushed it to 4,000+ downloads within an hour, the most downloaded on what was then ClawdHub. In March Silverfort found that ClawHub's download counter was a public endpoint anyone could call; 20,000+ fake downloads put its test skill at #1 in its category, which led to 3,900 executions in six days, and OpenClaw fixed it in under 24 hours. In May Orca reported that an unnamed "prominent" marketplace counted installs from unauthenticated requests. skills.sh's privacy page now describes hourly deduplication of install events by IP hash and TLS fingerprint.
Paperclip: listed clean, weaponized later#
The best-documented case on a mainstream registry is the one Zenity Labs, which sells agent security, published on 6 August. Every step below is dated in its write-up.
How the Paperclip look-alike skills worked
Steps of the campaign on skills.sh and GitHub, Jul to Aug 2026, per Zenity Labs
2 Jul
A look-alike identity
A domain one letter off (getpaperclipp.com) and a GitHub org one word off the real project, with a full clone of its source.
5 Jul
Listed clean
Byte-identical copies of the real skills appear on skills.sh and start collecting installs.
11 Jul
Trojanized
One malicious setup-installation.md is copied into seven skill paths. That day the family is #8 on skills.sh Trending, with 46.6K installs across six skills.
On use
The payload is a second file
SKILL.md stays benign. It sends the agent to setup-installation.md only when Paperclip needs installing, so a review of SKILL.md alone finds nothing.
On use
A remote stage
The agent runs curl with TLS checks off against a fake health endpoint, decodes base64 and runs a Node script.
Harvest
138 credential paths: SSH keys, cloud credentials, kube and Docker configs, npm and PyPI tokens, .env files.
Exfiltrate
One encoded bundle is posted out, then the stage deletes itself.
2 Aug
Takedown
Vercel and GitHub remove the skills, listings and repositories within 12 hours of Zenity's report. The counter stands at 1.7M+ aggregate installs.
Source: Zenity Labs
Zenity is careful about the scale: the 1.7 million installs are counter events, inflated by skills that reference each other in one repository, and it doesn't know how many people ran the payload. What the case shows is the pattern. A skill earns installs and rankings while it's clean, then changes, and an audit that ran at listing time doesn't see the change.
Injected skills work#
Even without code, a skill is a prompt the agent treats as trusted guidance. Skill-Inject, an academic benchmark of 202 injection-task pairs, found that instructions hidden in otherwise legitimate skills succeeded between 41.0% (Haiku 4.5) and 79.0% (Gemini 3 Flash) of the time when they were phrased to fit the task. Hiding the payload in a script rather than in SKILL.md raised that to 71.8%, from 53.8%. Even obviously harmful instructions, such as deleting files or piping a remote script to a shell, often went through.
Obviously harmful skill instructions that agents followed
Attack success rate by model on Skill-Inject's 82 obvious injections, single run, Feb 2026
| Label | Value |
|---|---|
| GPT-5.1-Codex-Mini | 68.3% |
| Gemini 3 Flash | 62.2% |
| Claude Sonnet 4.5 | 46.3% |
| GPT-5.2-Codex | 42.7% |
| Gemini 3 Pro | 42.7% |
| GPT-5.2 | 18.3% |
| Claude Opus 4.5 | 15.9% |
| Claude Haiku 4.5 | 8.5% |
Automated attacks do better still. SkillJect generates poisoned skills that split the payload across instructions and helper files. It averaged 95.1% success against 10.9% for naive direct injection, and on Claude Sonnet 4.5 it went from 5.0% to 97.5%. A separate test of two coding agents running malicious skills found Gemini CLI exploited in 95.5–96.1% of runs, with the agent recognising the danger in under 2% of runs.
What agents do by default#
So what stands between a skill and your machine? We read the documentation of seven local agents and three hosted ones, as published in late September, using archived copies where we could find them. The hosted surfaces all run skills in a container or VM. The local agents mostly don't.
What agents do when they load a skill
Out-of-the-box defaults from each vendor's documentation: OS sandbox, network access, how allowed-tools is treated, and whether repository skills wait for workspace trust, Sep 2026
Sources: Claude Code skills; Claude Code sandboxing; Codex security; Cursor agent security; GitHub Copilot CLI skills; VS Code docs; Gemini CLI docs; OpenClaw docs; Claude API skills; Copilot cloud agent firewall
Two details matter more than the table suggests. First, Claude Code's sandbox, when it's on, covers shell commands only: file tools, MCP servers and hooks run outside it, and if the sandbox can't start, commands run unsandboxed unless failIfUnavailable is set. Second, repository skills load without a trust prompt in Claude Code, Cursor and Gemini CLI by default, so cloning a repo can install skills. The spec's own guide for implementers suggests gating project skills on folder trust because a cloned repository may be hostile.
allowed-tools widens, it doesn't narrow#
The only security field in the spec is allowed-tools, marked "Experimental. Support for this field may vary". Its name suggests an allowlist. In Claude Code it's the opposite: it grants the listed tools without a prompt for that turn, and "It does not restrict which tools are available". The docs add that "Workspace trust doesn't gate this field", even in a non-interactive run in a folder you've never trusted. GitHub's Copilot docs treat it the same way and warn that pre-approving shell or bash "can allow attacker-controlled skills or prompt injections to execute arbitrary commands".
We checked how skills use the field in five well-known repositories, at their last commit before 29 September.
How often skills pre-approve tools
SKILL.md files declaring allowed-tools in five public skill repositories, last commit before 29 Sep 2026
Source: host0 count of git checkouts: anthropics/skills, obra/superpowers, vercel-labs/agent-skills, openai/skills, trailofbits/skills
That isn't a criticism of Trail of Bits, whose skills run security tools that need a shell. It shows what the field does in practice: when an author uses it, it usually removes the prompt that would otherwise stand between the skill and your shell. Anthropic has been closing the gaps for organizations. Over 2026 Claude Code fixed skill allowed-tools bypassing managed ask rules (March), added a disableSkillShellExecution setting (April) and a restricting disallowed-tools field (May), and in late September stopped repository and user skills, then marketplace plugins, from pre-approving their own tools when the managed allowManagedPermissionRulesOnly setting is on.
No signatures, only proposals#
Other software supply chains have spent years adding provenance: after the Shai-Hulud attack in 2025, GitHub announced a move for npm to 2FA publishing, seven-day tokens and trusted publishing. Skills have almost none of it. The spec has no signing, digest or security section, and its repository has closed four proposals that would have added one.
Skill provenance: proposals and partial fixes
Spec proposals and the registry and vendor measures around them, Mar to Sep 2026
16 Mar 2026
16 Mar 2026
Signature and permissions proposals closed
A signature block in SKILL.md (#247) and least-privilege permission metadata (#249).Source: agentskills #247
24 Mar 2026
24 Mar 2026
Discovery RFC requires sha256 digests
The .well-known/agent-skills draft makes clients verify each file's digest and not run scripts by default. Its pull request to the spec is unmerged.Source: Cloudflare RFC
19 May 2026
19 May 2026
NVIDIA signs its verified skills
OpenSSF Model Signing over every file. The same day, a provenance-fields proposal (#358) is closed in the spec repo.Source: NVIDIA
30 Jun 2026
30 Jun 2026
Verification proposal closed
#418, “Spec lacks guidance on skill verification and supply-chain trust”, opened 9 Jun.Source: agentskills #418
6 Aug 2026
6 Aug 2026
Anthropic scans skills for Enterprise
Skill and plugin scanning in beta for claude.ai and Cowork uploads; a failing skill is blocked.Source: Claude release notes
24 Sep 2026
24 Sep 2026
Claude Code locks down pre-approvals
Under a managed setting, repository and user skills can no longer pre-approve their own tools.Source: Claude Code changelog
Issue #418 put the gap in one sentence: "there's nothing in the format that lets a consumer verify a skill is what it claims to be" Anthropic says its Enterprise scanning will turn on by default on 2 October for organizations that haven't set it, but it covers skills uploaded to claude.ai and Cowork; skills uploaded through the Skills API aren't scanned, and neither are the skills a local agent reads from disk.
What extension stores have that skills don't
Publisher, install and takedown controls by ecosystem, as documented by Sep 2026
Sources: Microsoft (VS Code Marketplace); Agent Skills specification; OpenClaw; Holzbauer et al.; Vercel; Zenity Labs
What this means if you build with AI#
Skills are useful because they're cheap: a Markdown file and maybe a script, installed with one command. The same property makes them a supply chain with almost no controls. The data points to a few habits.
- Read a skill before you install it, all of it. In confirmed malware, 84.2% of the attack sits in prose, and the Paperclip payload sat in a second file
SKILL.mdonly pointed to. Look for "prerequisite" install steps,curl … | bash, base64 blobs and URLs you don't recognise. - Treat
allowed-toolsas a grant, not a fence. Never accept a skill you didn't write that pre-approves unscopedBashorshell. If you write skills, scope the field (Bash(git:*)) or leave it out. - Turn the sandbox on. In Claude Code that's
/sandbox, plusfailIfUnavailableso it can't silently fall back; VS Code, Gemini CLI and Copilot CLI have opt-in sandboxes too. Keep the network allowlist empty or narrow, so a stolen credential has nowhere to go. - Check a repository's skill folders before you run an agent in it. Repo skills load without a trust prompt in Claude Code, Cursor and Gemini CLI by default, and Claude Code applies their
allowed-toolsin folders you've never trusted. - Pin what you reviewed. A review is of one version. Install from a commit or a digest, and re-review on update: skills listed clean and changed later are the pattern, not the exception.
- Don't print secrets in skill scripts. Agents read stdout into context, which is how debug logging became 73.5% of credential-leak issues. Pass tokens through the environment and never echo them.
- Use scanners as a tripwire, not a gate. They catch known payloads; on new ones, recall can drop to single digits. An install count is not a review: counters have been faked to #1.
- For teams, allowlist the sources. Claude Code and VS Code support managed
strictKnownMarketplacesandstrictPluginOnlyCustomization; Claude Code'sallowManagedPermissionRulesOnlystops skills pre-approving tools. Trail of Bits' advice is curated, internal marketplaces over open registries.
host0's own deploy path is a skill: a SKILL.md and a bash script, publish.sh, that calls the host0 REST API with an API key for your account, and by default refuses to send that key to any other host. That means everything above applies to us too, so read the script before you install it; it declares no allowed-tools, and your agent should ask before running it.
Methodology#
This post draws on three research passes: academic studies and registry scans, incident reports, and each agent's documented defaults. Each was limited to primary sources: arXiv papers, vendor research posts, official documentation and changelogs, and GitHub repositories. Aggregator and "statistics" sites were used only as leads. Before publishing we re-opened the source behind every headline number and confirmed the quoted sentence or figure was still there. Documentation we could only read in its current form is described in the present tense.
Some numbers are our own:
- Client defaults: read from each vendor's documentation as it stood just before 29 September (archived copies or repository commits where we could find them). "Not documented" means we found no mention.
allowed-toolsusage: we cloned five repositories, checked out the last commit before 29 September, parsed the frontmatter of everySKILL.md(173 files, test fixtures included), and countedBashentries with no scope.- Arithmetic: the 1.68 risk ratio and the 82% share of vulnerable skills without scripts come from Liu et al.'s published rates; the 0.49% malicious rate from the credential paper's 83 of 17,022; Antiy's top-two share from its per-author counts.
We dropped or corrected several widely repeated claims: "26.1% of skills are malicious" (it's any vulnerability pattern), "skills with scripts are 2.12 times more likely to be vulnerable" (an odds ratio), "only 0.52% of skills are suspicious" (that's 15 of 2,887 flagged pairs), "1.7 million victims" (installs, not people), vendor rates with no published sample size, and a set of aggregator statistics about "affected users" with no primary source.
Limitations:
- Rates aren't comparable across studies. Each measured a different registry at a different time with a different definition. Flagged, vulnerable, harmful and confirmed malicious are different things.
- Security vendors sell what they measure. Snyk, Socket and Gen run skills.sh's audits; Koi, Zenity, Silverfort, JFrog, Trend Micro, Antiy, Island, Cato, Cisco and Trail of Bits sell security products or services. Only the academic papers have nothing to sell.
- Install counts are telemetry. No incident report gives a count of unique victims.
- Model results age fast. The attack benchmarks tested models available in early 2026.
- Our repository count is small and author-biased. 173 skills from five publishers is a sample, not a census.
Open questions#
- Is ClawHub cleaner now? Koi found 11.9% malicious in February and ClawScan 0.31% in June, but the methods differ. No registry has published a series on one method.
- How many skills.sh skills are flagged or hidden? Vercel hasn't published an aggregate, and the Paperclip family passed its audits until reported.
- How many people ran the payloads? Every incident reports installs or executions, never unique machines.
- How do commercial scanners compare on the same unseen malware? Socket's 98.7% and Cisco's 7.75% come from different test sets; nobody has run them all on one public held-out set.
- Does Codex enforce
allowed-tools? Its docs don't say. - Will the spec adopt digests or signing? Four proposals were closed, and the
.well-knowndigest pull request is still open. - How common is pre-approved
Bashacross a whole registry? Our 173-skill sample can't answer that; a crawl of the most-installed skills could.
