Blog

/Research

Skills vs MCP: what the data says

Five MCP servers load ~55K tokens before you type; five skills, ~500. Yet a skill went unused in 56% of Vercel's evals. Context, evals and adoption, side by side.

·10 min read

Summarize in ChatGPT
Skills vs MCP, split black and white down the middle: tokens loaded at session start for five of each, ~500 vs ~55K; days until OpenAI, Google, Microsoft, GitHub and Cursor all adopted, 98 vs 135; reference repo GitHub stars at about one year, 178,481 vs 72,958; npm downloads in the 30 days to 28 Sep 2026, 28.9M vs 220.0M
On this page

Two standards that started at Anthropic now shape how most AI agents reach tools and learn procedures. The Model Context Protocol, released in November 2024, connects an agent to tools and data in other systems. Agent Skills, released in October 2025, are folders with a SKILL.md file of instructions and scripts that an agent loads only when a task calls for them. From the day the second one shipped, developers argued about whether it made the first one obsolete. Simon Willison predicted a "Cambrian explosion in Skills" that would make the MCP rush look pedestrian, and by February an essay called "MCP is dead. Long live the CLI." was on the front page of Hacker News.

Most of that argument ran on anecdotes. We wanted the numbers: what each one costs in context before the user types anything, what happens to accuracy as tools pile up, what the head-to-head evaluations measured, how fast each spread when you line them up by age, and what the security scans found. So we collected every published eval and measurement we could find, pulled download counts, star histories and client lists ourselves, and checked the claims against their sources. Vendors that sell one side or the other are labelled, and every number links to where it came from. The short version: the data doesn't crown a winner, and in September the two standards merged.

Key findings

  1. 1
    The cost gap is real but comes from a default, not the protocol. Five MCP servers loaded about 55K tokens of tool definitions before the conversation started in Anthropic's example. Five skills cost about 500, at the ~100 tokens per skill Anthropic's docs give. Loading tools on demand closes most of the gap: Tool Search cut tokens by 85%.
  2. 2
    More tools make agents pick worse. With every tool in the prompt, one benchmark found 13.62% tool-selection accuracy against 43.13% when tools were retrieved first, and Claude-4-Sonnet fell from 62.4% to 44.2% when all 65 servers were mounted at once.
  3. 3
    Head to head, quality comes out even and cost doesn't. In Arize's 500-trial eval, MCP and skills scored 0.834 and 0.833 on correctness, but on the hardest tasks MCP cost more than six times as much and took five times as long. A well-designed MCP server matched a CLI on cost.
  4. 4
    Skills fail differently: they don't fire. In Vercel's Next.js evals the skill was never invoked in 56% of cases. The best result anyone has published pairs the two: Supabase's MCP plus a skill beat MCP alone for all four agent and model pairs it tested.
  5. 5
    Skills spread faster at the top; MCP is bigger everywhere usage shows. Skills reached OpenAI, Google, Microsoft, GitHub and Cursor in 98 days against MCP's 135, and the skills reference repo had 178,481 GitHub stars at 345 days old, about 2.4 times what MCP's reference repo had at 360 days. But the MCP SDK was downloaded 7.6 times as often as the skills CLI on npm in the month to 28 September, and MCP's official client list was twice as long at the same age.
  6. 6
    Their risks sit in different places. A skill acts with whatever the agent can already do, and an academic scan flagged 26.1% of 31,132 skills. MCP brings its own OAuth and its own network surface, and a separate scan flagged 7.2% of 1,899 servers, plus 5.5% for tool poisoning. The methods differ, so neither number makes one safer.
  7. 7
    The argument ended in a merge. On 13 September SEP-2640 became a final MCP extension that serves skills as MCP resources under skill://, so a server can ship the instructions for its own tools.

55K vs 500

Tokens loaded before the first message: five MCP servers vs five skills

Our arithmetic

Sources: Anthropic; Claude docs

>6×

MCP's cost vs skills on Arize's hardest tasks, at the same correctness

Source: Arize, May 2026

56%

Vercel eval cases where the skill was never invoked, Jan 2026

Vercel runs skills.sh

Source: Vercel

7.6×

npm downloads, MCP SDK vs the skills CLI, 30 days to 28 Sep 2026

Our count

Source: npm downloads API

Two layers, not two rivals#

The two are different kinds of thing. A skill is text and, often, scripts: a name and a one-line description the agent always sees, and a body it reads when the description matches the task. MCP is a protocol: a server, local or remote, that exposes tools an agent can call, with its own transport, versioning and authorization. Anthropic's own one-line summary, from its guide to how the pieces fit, is that "MCP connects Claude to data; Skills teach Claude what to do with that data."

OpenAI's plugin docs draw the same line as a rule. "Use skills when instructions and the tools already available to the model are enough to complete the task", and build an MCP server "when your plugin must connect to a service, expose a controlled set of tools, authenticate users, or run behavior on infrastructure you operate." A plugin can carry both.

Agent Skills and MCP, dimension by dimension

What each standard is and how it works, as of Sep 2026

Agent SkillsMCP
What it isA folder: SKILL.md with a name, a description and instructions, plus optional scripts and reference filesA JSON-RPC protocol: servers expose tools, resources and prompts to clients
Where it runsInside the agent's own environmentA separate process, local (stdio) or remote (HTTP)
AuthNone of its own; it acts with the agent's permissionsOAuth 2.1 for remote servers (optional); local servers read credentials from the environment
Context costAbout 100 tokens per skill always loaded; the body only when triggeredEvery tool's schema loaded up front by default; tool search defers it
Required fields2 (name, description)3 in a registry entry (name, description, version)
VersioningNo version field; updates mean copying the folder againVersioned packages; remote servers change in place
RegistryNone in the spec; the largest index is Vercel's skills.shThe official MCP Registry, in preview since Sep 2025
Governanceagentskills.io open standard, stewarded by AnthropicLinux Foundation's Agentic AI Foundation since Dec 2025
Best for, per the vendorsProcedures, house style, domain knowledgeLive data, actions on a service, signed-in users
Skills are lighter in every row: two required fields, no transport, no auth of their own. The cost of that lightness is that a skill can do only what the agent around it can already do.

Sources: Agent Skills specification; MCP specification; Claude docs; MCP Registry; OpenAI plugins

What each costs before you type#

Every MCP tool a client loads comes with a name, a description and a JSON schema, and by default all of it goes into the model's context at the start of the session. Anthropic measured a common five-server setup (GitHub, Slack, Sentry, Grafana and Splunk) at "58 tools consuming approximately 55K tokens before the conversation even starts", and says it has seen tool definitions reach 134K tokens internally. Mario Zechner measured Playwright's MCP server at 13.7K tokens and Chrome DevTools' at 18K, and replaced both with a README for his own browser scripts that came to 225 tokens.

A skill costs much less up front because only its name and description are loaded. Anthropic's docs put that at "~100 tokens per Skill", with the body, under 5K tokens, loaded only when the skill is triggered, and bundled files costing nothing until they're read. Armin Ronacher, who moved all his MCP servers to skills in December, put the difference in one line: "skills do not actually load a tool definition into the context."

Put those per-unit figures into a simple model and the gap is two orders of magnitude.

Tokens loaded at session start, N MCP servers vs N skills

Tokens in context before the first message, by number of servers or skills installed; worked example

Tokens loaded at session start, N MCP servers vs N skills
Seriesxy
MCP, all tools loaded111K
MCP, all tools loaded555K (~55K)
MCP, all tools loaded10110K
MCP, all tools loaded20220K (220K)
MCP with tool search13.5K
MCP with tool search53.5K
MCP with tool search103.5K
MCP with tool search203.5K (~3.5K)
Skills, metadata only00
Skills, metadata only1100
Skills, metadata only5500 (~500)
Skills, metadata only101K
Skills, metadata only202K (~2K)
At five each, eagerly loaded MCP costs about 55,000 tokens against 500 for skills, 110 times as much; even after one skill's body loads (5,500 tokens) it's 10 times. Twenty servers wouldn't fit in a 200K window. With tool search, MCP's first-task cost stays flat at about 3,500 tokens, in the same range as skills.

Sources: Anthropic, advanced tool use; Claude docs, Agent Skills; host0 arithmetic: 11,000 tokens per server (Anthropic's 55K ÷ 5), 100 per skill, ~500 + ~3,000 with tool search

The model has two soft spots. The 100 tokens is a budget in the docs, not a measurement. And skills have a ceiling of their own: Claude Code's docs say its skill listing is capped at 1% of the context window, dropping descriptions when it overflows, so a large skill collection degrades by losing descriptions rather than by growing.

The MCP side answered by borrowing the idea. Anthropic's Tool Search defers tool definitions until the model searches for them, "an 85% reduction in token usage", and Claude Code turned it on automatically in January for setups whose tool descriptions would take more than 10% of the context. Cursor's version, loading MCP tool descriptions only when needed, cut total agent tokens by 46.9% in runs that called an MCP tool. Anthropic's code execution with MCP pattern, which presents tools as files the agent reads on demand, went from 150,000 tokens to 2,000 in its example, and the same post notes that "adding a SKILL.md file to these saved functions creates a structured skill." By August the protocol's maintainers conceded the point on their roadmap: with a hundred tools, "the model pays for that entire surface before the user has asked a single question."

More tools, worse choices#

Tokens are the visible cost. The less visible one is accuracy: the more tools a model can see, the more often it picks the wrong one. Anthropic says the most common failures "are wrong tool selection and incorrect parameters, especially when tools have similar names", and its docs say Claude's tool choice degrades past 30–50 tools. Three studies measured the gap between showing every tool and showing only the relevant ones.

Accuracy with every tool loaded vs a narrowed set

Tool-selection accuracy or task success, by study, setup and model, 2025

Accuracy with every tool loaded vs a narrowed set
LabelValue
RAG-MCP: every tool in the prompt13.62%
RAG-MCP: tools retrieved first43.13%
MCPVerse, Claude-4-Sonnet: all 65 servers44.2% (550+ tools)
MCPVerse, Claude-4-Sonnet: relevant servers62.4%
Anthropic eval, Opus 4: all tools loaded49%
Anthropic eval, Opus 4: Tool Search74%
Anthropic eval, Opus 4.5: all tools loaded79.5%
Anthropic eval, Opus 4.5: Tool Search88.1%
Highlighted bars are the narrowed setups; each pair is one study's own measure, so compare within a pair, not across. Stronger models lose less, but every pair moves the same way.

Sources: RAG-MCP, arXiv 2505.03275; MCPVerse, arXiv 2508.16260; Anthropic

The MCPVerse result has a twist worth keeping. Its authors note that agentic models like Claude-4-Sonnet did slightly better with a handful of relevant servers (62.4%) than with only the exact tools a task needed (62.3%), so a few extra tools don't hurt a strong model. Mounting all 65 servers, an action space of more than 140K tokens, did, and only Claude-4-Sonnet and Gemini-2.5-Pro could run that setting at all. IBM's LongFuncEval saw the same direction across five models: "a performance drop of 7% to 85% as the number of tools increases", from 49 tools to 741.

This is where skills have a structural edge. A skill doesn't add a tool to the menu; it teaches the agent to use tools it already has, usually a shell. The menu stays short however many skills are installed.

Head to head#

Only a handful of public evaluations put the two approaches on the same tasks, and none of them compares a skill alone with MCP alone on a real service with sign-in. Most compare MCP with a command-line tool, with or without a skill that explains it.

Every head-to-head eval we found

Public evaluations comparing MCP with skills, CLIs or both, Aug 2025 to May 2026

EvalDateComparedResultCaveat
Mario ZechnerAug 2025One terminal tool as MCP vs CLI, Claude Code, 120 runsBoth 100% success; MCP $19.45 vs CLI $19.95One well-designed server
VercelJan 2026Skill vs an always-loaded docs index, Next.js 16Skill unused in 56% of cases; 53% pass by default, 79% when told to use it; docs index 100%Vercel runs skills.sh
ScalekitMar 2026GitHub MCP vs gh CLI vs CLI + skill, Claude Sonnet 4, 75 runsMCP used 8.8× the tokens of the CLI; 18 of 25 MCP runs succeeded vs 25 of 25Sells agent auth; read-only tasks
SupabaseApr 2026No tools vs MCP vs MCP + skill, four agent and model pairsMCP + skill best for all four; Claude Code with Sonnet 4.6: 46% → 58% → 71%Ships both; six scenarios, LLM judge
ArizeMay 2026MCP vs two skills (CLI), Opus 4.6, 500 trialsCorrectness 0.834 vs 0.833 and 0.826; hardest tier, MCP cost >6× and took 5× as longTest repo was synthetic
Where MCP met a CLI-plus-skill setup, accuracy was a tie or worse for MCP, and MCP cost more. Against a bare CLI, a well-designed server broke even. The one with the best scores used both. Note who ran each test: two of five sell something on one side.

Sources: Mario Zechner; Vercel; Scalekit; Supabase; Arize

The cleanest cost comparison is Scalekit's, with the caveat that Scalekit sells agent authentication. It ran five read-only GitHub tasks through GitHub's MCP server, the gh command-line tool, and gh plus a skill that explains it. The MCP arm's median token use ranged from 4 to 32 times the CLI's, depending on the task, and all seven of its failures were TCP timeouts reaching GitHub's hosted server.

Tokens to finish five GitHub tasks, by approach

Sum of median tokens per task across Scalekit's five read-only tasks, Claude Sonnet 4, Mar 2026

Tokens to finish five GitHub tasks, by approach
LabelValue
gh CLI26,159 (25/25 succeeded)
gh CLI + skill32,717 (25/25 succeeded)
GitHub MCP server230,254 (18/25 succeeded)
The highlighted MCP bar is 8.8 times the bare CLI and 7.0 times the CLI with a skill. The skill added about 25% to the CLI's tokens: that's the price of the instructions. Scalekit is a vendor; the sums are ours.

Sources: Scalekit; host0 arithmetic

Mario Zechner's earlier test is the counterweight. When he wrapped one terminal tool both ways and ran it 120 times through Claude Code, both versions succeeded every time and MCP came out slightly cheaper and faster, which led him to conclude that "MCP vs CLI truly is a wash" and that the problem with many servers is design: "badly designed wrappers that dump unnecessary JSON everywhere." GitHub, for its part, moved the data fetching in its own agentic workflows from its MCP server to gh steps that run before the agent starts.

Arize's eval is the largest. Over 500 trials, correctness was a tie, and Claude with no guidance at all scored 0.845 because it already knows the GitHub CLI from training. The difference was on the hardest tier, where MCP made about 12 tool calls per task to the skills' 5, cost more than six times as much, and the agent gave up on the MCP tools often enough that their "tool fidelity" fell to 0.33. Arize's conclusion wasn't a winner: "MCP plus the command line."

Skills have their own failure mode, and it's quieter: the agent doesn't use them. Vercel found that "in 56% of eval cases, the skill was never invoked", so the skill added nothing by default; telling the agent to use it raised the pass rate from 53% to 79%, and an 8KB docs index the agent always sees scored 100%. Academic benchmarks agree that the skill itself matters: curated skills raised pass rates by about 16 points on SkillsBench, while on SWE-Skills-Bench "39 of 49 skills yield zero pass-rate improvement."

Adoption, aligned by age#

MCP is nearly 11 months older, so comparing the two on the same day flatters it. On 28 September 2026 MCP was 672 days old and Agent Skills 347. We lined both up by days since launch where the data allowed it.

The first signal is attention. Anthropic's skills repository, the reference collection most people installed first, picked up stars far faster than MCP's reference servers repository did at the same age.

GitHub stars on each reference repository, by days since launch

anthropics/skills (from 16 Oct 2025) and modelcontextprotocol/servers (from 25 Nov 2024), stars at Wayback Machine snapshots, first 360 days

GitHub stars on each reference repository, by days since launch
Seriesxy
anthropics/skills023
anthropics/skills1574
anthropics/skills35.2K
anthropics/skills713.1K
anthropics/skills1715K
anthropics/skills4118.4K
anthropics/skills6422.3K
anthropics/skills9141.7K
anthropics/skills10961.1K
anthropics/skills12068.9K
anthropics/skills15194.4K
anthropics/skills180116.8K
anthropics/skills212135.4K
anthropics/skills240150.2K
anthropics/skills270160.8K
anthropics/skills301168.7K
anthropics/skills334176.5K
anthropics/skills345178.5K
modelcontextprotocol/servers043
modelcontextprotocol/servers1857
modelcontextprotocol/servers42.6K
modelcontextprotocol/servers83.6K
modelcontextprotocol/servers315.7K
modelcontextprotocol/servers607.3K
modelcontextprotocol/servers9310.8K
modelcontextprotocol/servers12024.4K
modelcontextprotocol/servers15140.6K
modelcontextprotocol/servers18048.5K
modelcontextprotocol/servers21155.6K
modelcontextprotocol/servers24361.5K
modelcontextprotocol/servers27365.6K
modelcontextprotocol/servers30568.9K
modelcontextprotocol/servers33171K
modelcontextprotocol/servers36073K
The skills repo passed 13,000 stars in its first week and had 178,481 by day 345; MCP's servers repo had 72,958 at day 360, about 2.4 times fewer. Stars measure attention, not use, and the skills repo also hosts Anthropic's own document skills. Both curves bend at the same point in the story: the skills repo took off in January, when OpenAI, Google, Microsoft and Cursor shipped skills; MCP's did in March and April 2025, when its big adopters arrived.

Sources: Wayback Machine snapshots of github.com/anthropics/skills and github.com/modelcontextprotocol/servers; host0 count

The second signal is client support, and here the order flips once you look past the biggest names. Skills reached all five of OpenAI, Google, Microsoft, GitHub and Cursor faster, partly because MCP had already done the persuading and a client only needs a file loader to read a SKILL.md.

Days from launch to each big adopter

First shipped support by vendor, in days since each standard launched (MCP 25 Nov 2024, Skills 16 Oct 2025)

VendorSkills: first supportDayMCP: first supportDay
OpenAICodex CLI, experimental47Agents SDK121
GitHubCopilot coding agent and CLI63GitHub MCP Server preview130
GoogleGemini CLI preview83Gemini pledge (Gemini CLI on day 212)135
MicrosoftVS Code 1.10884VS Code 1.99129
CursorCursor 2.498Cursor 0.45~66
All five98135
All five had skills by day 98; all five had committed to MCP by day 135, though Google's first shipping MCP client came on day 212. Cursor was MCP's first big adopter and skills' last.

Sources: Cursor 2.4; GitHub changelog; VS Code 1.108; VS Code 1.99; Codex PR #7412; Gemini CLI; TechCrunch (OpenAI); TechCrunch (Google); GitHub MCP Server; Cursor forum

Past the big five, MCP's client list grew much longer. Its official clients page, which took community pull requests, listed 93 clients at day 343. The agentskills.io showcase, which Anthropic curates, listed 46 at day 347 and hadn't changed since August.

Official client lists, by days since launch

Clients listed on modelcontextprotocol.io (community-maintained) and agentskills.io (curated), counted from each page's git history

Official client lists, by days since launch
Seriesxy
MCP clients page03
MCP clients page55
MCP clients page207
MCP clients page3510
MCP clients page4612
MCP clients page6713
MCP clients page8118
MCP clients page8920
MCP clients page10822
MCP clients page12227
MCP clients page13329
MCP clients page14130
MCP clients page15645
MCP clients page17050
MCP clients page18451
MCP clients page20156
MCP clients page21459
MCP clients page23273
MCP clients page24876
MCP clients page26278
MCP clients page29279
MCP clients page30884
MCP clients page31889
MCP clients page34092
MCP clients page34393
MCP clients page37094
Agent Skills showcase639
Agent Skills showcase8211
Agent Skills showcase8412
Agent Skills showcase10326
Agent Skills showcase11327
Agent Skills showcase12931
Agent Skills showcase14032
Agent Skills showcase16533
Agent Skills showcase18036
Agent Skills showcase18637
Agent Skills showcase21641
Agent Skills showcase23942
Agent Skills showcase26744
Agent Skills showcase29746
Agent Skills showcase34746
Skills led until about day 140; MCP pulled away in spring 2025, when OpenAI, Google and Microsoft arrived within a month of each other. The skills list starts at day 63, when the open standard was published. MCP's page peaked at 114 clients and was deleted in May 2026.

Sources: modelcontextprotocol/modelcontextprotocol; agentskills/agentskills; host0 count of git history

On every number shaped like usage, MCP is bigger, by a margin that depends on what's counted. Downloads are the clearest, with the caveat that they compare different things: the MCP SDK is a library that every server and many clients depend on, while the skills package is Vercel's installer, which counts every npx skills run.

Skills and MCP side by side, late September 2026

Matched measures for each standard, dated on or before 28 Sep 2026

MeasureAgent SkillsMCPMCP ÷ Skills
npm downloads, 30 days to 28 Sep (skills CLI vs @modelcontextprotocol/sdk)28.9M220.0M7.6×
PyPI downloads, 1–28 Sep (skills-ref vs mcp)268,172212.0M~790×
New GitHub repos by topic, 1–28 Sep (agent-skills vs mcp)6,28712,3502.0×
Official client list at ~day 34546932.0×
Reference repo stars at ~day 350178,481 (day 345)72,958 (day 360)0.4×
Index size1M skills on skills.sh (Vercel, 25 Sep)≥ 37,060 servers in the official registrydifferent units
MCP leads every usage measure, from 2 times (new GitHub repos) to about 790 times (the official Python libraries). The two index sizes count different things, a folder of instructions vs a server you run or call, so don't compare them.

Sources: npm downloads API; ClickHouse public PyPI dataset; GitHub search API; Vercel; MCP Registry API; host0 counts

The registry row needs its caveats. Vercel, which runs skills.sh, said on 25 September that the registry "grew to one million agent skills and recorded nearly 280 million installs" in seven months, and that 48% of skills were installed exactly once. For MCP we rebuilt the official registry's history from its public API: at least 37,060 server names still listed today had first been published by 28 September, a lower bound because deleted entries don't show. Agent Skills has no official registry at all.

Security and operations#

The two standards put the risk in different places. A skill has no permissions of its own: it runs inside the agent, with whatever the agent can reach. In Claude Code, Anthropic's docs say skills "have the same network access as any other program on the user's computer", warn that malicious skills "could lead to data exfiltration, unauthorized system access, or other security risks", and advise using skills "only from trusted sources". The risk is the supply chain and prompt injection, and the defence is vetting what you install.

MCP is a network service, so it has a network service's problems and tools. Remote servers can use OAuth, which Arize called MCP's real enterprise advantage: "OAuth lets administrators control who and what gets access to which resources. CLI auth usually does not." But authorization is optional in the spec, local servers read credentials from the environment, and the specification itself says that "MCP itself cannot enforce these security principles at the protocol level." Exposed servers, auth bugs and tool poisoning are the result.

What security studies found in each ecosystem

Selected scans and benchmarks, each with its own sample and method, Jun 2025 to Apr 2026

StudyDateWhat was examinedFinding
Liu et al. (academic)Jan 202631,132 skills from two marketplaces26.1% had at least one vulnerability; 5.2% looked malicious; skills with scripts 2.12× as likely to be vulnerable
Snyk ToxicSkills (sells scanning)Feb 20263,984 skills from ClawHub and skills.sh13.4% had a critical issue; 76 malicious payloads
Skill-Inject (academic)Feb 2026202 injection tasks hidden in skillsUp to 80% attack success on frontier models
Hasan et al. (academic)Jun 20251,899 open-source MCP servers7.2% had general vulnerabilities; 5.5% MCP-specific tool poisoning
MCPTox (academic)Aug 202545 live servers, 353 tools, 20 agentsBest tool-poisoning attack succeeded 72.8% of the time
Trend MicroJul 2025 → Apr 2026Internet-exposed servers with no client auth492 → 1,467, same method
CensysApr 2026Internet-reachable MCP services12,520 services on 8,758 IP addresses
Read each row on its own. The skill scans counted marketplace folders; the MCP scans counted repositories, live endpoints or benchmark attacks. No study has run one method over both, so these rates can't rank the two.

Sources: Liu et al.; Snyk; Skill-Inject; Hasan et al.; MCPTox; Trend Micro; Censys

Operations follow the same split. MCP servers are versioned packages or hosted endpoints, and a remote server can change its tools without the user doing anything, which Ronacher counted against it: servers "keep changing their tool descriptions at will". Skills have the opposite problem. The spec has no version field, updating means copying the folder again, and Anthropic's own surfaces don't sync them: a skill uploaded to claude.ai isn't available in the API or in Claude Code.

From debate to merge#

The argument ran for about a year, and it moved from "which one" to "both, together".

Skills vs MCP: the debate

Positions, evals and standards work, Jul 2025 to Sep 2026

  1. 3 Jul 2025

    Ronacher: code beats tool calls

    MCP “isn't truly composable”; most composition happens through inference.
  2. 15 Aug 2025

    Zechner: “a wash”

    One tool as MCP and as a CLI: both 100% success, similar cost.
  3. 16 Oct 2025

    Agent Skills launch

    Simon Willison expects a “Cambrian explosion in Skills”.
  4. 2 Nov 2025

    “What if you don't need MCP at all?”

    Zechner swaps 13.7K–18K tokens of MCP for a 225-token README.
  5. 13 Nov 2025

    Anthropic draws the line

    “MCP connects Claude to data; Skills teach Claude what to do with that data.”
  6. 24 Nov 2025

    Tool Search

    MCP tools load on demand; 85% fewer tokens in Anthropic's example.
  7. 13 Dec 2025

    Ronacher moves all his MCP servers to skills

    Including Sentry's, which cost him about 8K tokens up front.
  8. 13 Jan 2026

    First proposal to put skills in MCP

    SEP-2076 proposes a new primitive; closed unmerged 42 days later.
  9. 1 Feb 2026

    Skills Over MCP group forms

    An interest group inside MCP governance; a working group from April.
  10. 28 Feb 2026

    “MCP is dead. Long live the CLI.”

    Eric Holmes's essay reaches the Hacker News front page.
  11. 23 Apr 2026

    SEP-2640 opened

    Serve skills over MCP's existing Resources primitive.
  12. 18 Jun 2026

    “Horrible Agent Experience”

    Angie Jones on the AAIF blog: a bare tool list isn't enough; ship instructions with servers.
  13. 6 Aug 2026

    Agent Plugins 1.0

    Amazon, Cursor, Microsoft, OpenAI and Vercel package MCP config and skills together.
  14. 22 Aug 2026

    MCP roadmap concedes the cost

    Progressive tool discovery becomes a priority.
  15. 13 Sep 2026

    SEP-2640 final

    Skills become an official MCP extension, served as resources under skill://.
  16. 23 Sep 2026

    One marketplace, two standards

    Anthropic's Claude Marketplace lists plugins built on MCP and Agent Skills.
The sharpest takes came early. By mid-2026 the vendors, Anthropic, OpenAI, Google, Microsoft and the MCP maintainers, were all describing the two as complements, and the standards work turned to shipping them together.

Sources: Armin Ronacher; Mario Zechner; Anthropic; MCP working group; AAIF; SEP-2640

Google's framing, when it launched its own skills repository in April, captures where most vendors ended up: heavy MCP use causes "context bloat", and skills are the condensed expertise that fixes it. Not a replacement; a layer on top.

How skills travel over MCP#

SEP-2640 names the problem it solves plainly: "A server and the skill that teaches an agent to use it are versioned, discovered, and installed separately." Its fix adds no new primitive. Each file in a skill folder becomes an MCP resource, conventionally under a skill:// URI, and the skill format itself is left entirely to the Agent Skills spec. The proposal took 143 days from pull request to merge, on 13 September 2026.

How SEP-2640 serves a skill over MCP

The io.modelcontextprotocol/skills extension, final 13 Sep 2026

  1. Declare

    The server declares the io.modelcontextprotocol/skills extension and must implement skills/list and skills/get.

  2. List

    The host lists skills. Each entry carries the SKILL.md URI and every file's digest, or is marked dynamic.

  3. Load on demand

    Only when a skill is used does the host read skill://<name>/SKILL.md through resources/read, checking it against the digest. Supporting files are fetched when read, never ahead of need.

  4. security

    Treat as untrusted

    Skill content is untrusted model input, tagged with the server it came from.

    • No silent execution

      No host-side code runs without explicit per-skill approval.

    • No impersonation

      Names resolve per server, so a server can't shadow a popular skill or a local one.

  5. Re-approve on change

    Approval binds to the file digests. If the server changes the skill, the host must re-prompt before loading or running it.

The extension reuses MCP's existing Resources primitive and the Agent Skills format unchanged. The security rules are stricter than for a local skill: the spec says hosts must treat MCP-served skills as a higher-risk surface than remote tool calls.

Source: SEP-2640: Skills Extension

Two details matter for anyone shipping a server. The digests are unsigned and come from the same server as the content, so the spec says plainly that "a match proves the two are consistent, not that either is trustworthy." And client support is still thin: Arcade, which sells MCP tooling, reports that OpenAI imports skills over MCP as snapshots taken when a plugin is submitted, so later updates on the server don't reach users. We found no OpenAI statement confirming it.

What this means if you build with AI#

The data doesn't pick a winner. It does point to a sensible default for each job.

  1. Write a skill when the agent already has the tools. If the job is a procedure, house style, or how to use a CLI the agent can already run, a skill costs about 100 tokens until it's needed and adds nothing to the tool menu. That's OpenAI's rule too.
  2. Build an MCP server when someone else's agent needs your service. Live data, actions on an account and per-user sign-in are what MCP is for, and OAuth is the advantage a CLI lacks. If you control both ends, a CLI plus a skill was cheaper in every cost comparison we found but one.
  3. Ship both for anything non-trivial. The best published scores came from MCP plus a skill, and SEP-2640 now lets a server carry its own instructions.
  4. Keep the tool menu short. Accuracy falls as tools pile up, past 30–50 tools by Anthropic's account. Expose a small default toolset, and use clients with tool search or deferred loading.
  5. Write descriptions that trigger. A skill that never loads does nothing, which is what happened in 56% of Vercel's cases. Say in the description what the skill does and when to use it, and test that it fires. For knowledge the agent needs on every task, put it where it's always loaded.
  6. Vet skills like software, and lock down servers like services. A skill runs with the agent's permissions, so install only from sources you trust and read the scripts. An MCP server is an endpoint: put authentication in front of it and pin the versions you depend on.
  7. Measure your own setup. Every number above comes from someone else's tasks. Count the tokens your servers load and check whether your skills fire before you decide.

host0 uses both: agents deploy apps through the host0 skill and our REST API, and an MCP server manages the apps once they're live. It's the same split the vendors describe, at the scale of a small product.

Methodology#

This post draws on three research passes: context cost and evaluations, positions and security, and adoption. Each was limited to primary sources: specifications, vendor documentation and engineering posts, academic papers, and the evaluations' own write-ups. The post is dated 28 September 2026, and nothing cited postdates it. Before publishing we re-opened every single-source page behind a headline number and confirmed it still said what we quote. Two things changed on that pass: MCPVerse's abstract says Claude-4-Sonnet improved with more tools, which is true from its smallest to its middle setting, so we report all three; and Arize's eval had a fourth, no-guidance arm, which we added.

Our own numbers:

  • Worked example: 11,000 tokens per MCP server is Anthropic's five-server total divided by five; 100 tokens per skill is the docs' budget; tool search is Anthropic's ~500-token search tool plus about 3,000 tokens for the tools it finds.
  • Scalekit sums: we added Scalekit's five per-task medians for each approach.
  • Stars: the star count on each repository page in Internet Archive snapshots, because the GitHub stargazers API returned errors for these repositories.
  • Client lists: entries counted in each version of the official pages, from their git history.
  • Downloads: npm's downloads API and the public ClickHouse PyPI dataset, cut at 28 September 2026.
  • Registry: a full pull of every version in the official MCP Registry, taking each name's earliest publish date; deleted names are missing, so it's a lower bound.

Limitations:

  • The evals are small and few. Five head-to-heads, two from companies with a stake, none with a skill-only arm on a real service with sign-in.
  • The context model is a model. It uses averages and a docs budget, not a measurement of any one setup, and clients now defer MCP tools by default.
  • Downloads are not users. The skills CLI and the MCP SDK are different kinds of package.
  • Security numbers don't compare. Each study had its own sample, scanner and definition.
  • Live docs are undated. Where we cite Anthropic's or OpenAI's current documentation, we say "the docs say".

Open questions#

  1. Skill alone vs MCP alone. Nobody has published an eval of each on its own against a real API that needs sign-in.
  2. One scanner, both ecosystems. No study has run the same security method over skills and MCP servers.
  3. What deferred MCP loading really costs. Clients don't publish how many tokens a deferred tool's name and server instructions take.
  4. How big a skill's metadata is in practice. The 100 tokens is a budget; nobody has published a measured average.
  5. How often skills fire outside one eval. Vercel's 56% hasn't been replicated, and no vendor publishes trigger rates.
  6. Who implements SEP-2640. Which clients list and load skills from MCP servers today, and does OpenAI's import track updates?
  7. Comparable usage. OpenAI says 26.6% of weekly Codex users invoked a skill in June; no vendor publishes a matching share for MCP tool calls.
ResearchAgent skillsMCP