In February 2025 Anthropic shipped Claude Code as a research preview, a command-line tool you could hand a task and walk away from. Twenty months later every large AI company sells a coding agent that runs in the background, opens its own pull requests and works on several tasks at once. The agents have become one of the biggest businesses in AI: a coding-agent startup was bought for $60 billion, and the labs now host what the agents build.
Most of what gets written about coding agents is a launch post or a vendor's own adoption figure. We wanted numbers we could check. So we counted the agents' public footprint on GitHub ourselves, month by month, and set it next to the developer surveys, the revenue disclosures, the benchmark record and the productivity studies. Vendor numbers are the vendor's claims and are labelled that way, and every number below links to its source.
Key findings
- 1One in four public pull requests now carries an agent's signature. In September 2026, 5.96M of the 23.09M pull requests opened in public GitHub repositories (25.8%) had a coding-agent branch name, footer or bot author, by our count. A year earlier it was 6.9%. It's a floor: agents that leave no signature are invisible.
- 2Claude Code is the largest footprint. 4.26M public pull requests in September carried its default "Generated with Claude Code" footer, 138 times as many as a year earlier and 4.5 times Codex's 954K. Public pull requests from GitHub's own Copilot coding agent fell 81% from their March peak.
- 3The surveys show the same shift. JetBrains found about 39% of professional developers using Claude Code at work in mid-2026, up from 18% in January, while GitHub Copilot fell from 29% to 21% and Cursor from 18% to 12%. 90% use some coding agent at least weekly.
- 4The money followed. Cursor's run rate reportedly reached about $4B in June 2026, up from $100M in January 2025, and SpaceX completed its $60B all-stock acquisition on 14 August. Anthropic says Claude Code passed a $2.5B run rate in February; Cognition says it crossed $1B in September.
- 5The main coding benchmark broke. SWE-bench Verified saturated around 81%, and OpenAI stopped reporting it after finding flawed tests in 59.4% of the hard tasks it audited. METR found that about half of test-passing agent fixes would not be merged by the projects' maintainers.
- 6Task length keeps doubling. METR's 50% time horizon went from 4 minutes (GPT-4, 2023) to about 12 hours (Claude Opus 4.6, February 2026) and about 17.4 hours (Claude Mythos Preview, April 2026), though METR says anything above 16 hours is unreliable on its current tasks.
- 7Productivity evidence is still mixed. In METR's 2025 trial, experienced developers took 19% longer with AI while believing they were 20% faster. Faros's 2026 telemetry shows 33.7% more tasks done and 242.7% more incidents per pull request at high AI adoption.
- 8The labs now host the output. OpenAI launched Sites to host what Codex and ChatGPT build, Anthropic launched artifacts in Claude Code for org-only pages, and Vercel's MCP server learned to deploy.
25.8%
Public GitHub pull requests opened in Sep 2026 that carry a coding-agent signature
Our count; a floor
Source: GitHub search API
4.26M
Public pull requests with the “Generated with Claude Code” footer, Sep 2026
Our count; 138× Sep 2025
Source: GitHub search API
~$4B / $60B
Cursor's reported run rate in June 2026, and the value of SpaceX's all-stock deal for it
Run rate is a reported claim
~17.4 h
METR 50% time horizon, Claude Mythos Preview (early), Apr 2026
METR: above 16 h is unreliable
Source: METR
One in four public pull requests#
Coding agents leave fingerprints. Claude Code adds a "🤖 Generated with Claude Code" footer to the pull requests it writes unless you turn it off, and its web and desktop sessions push branches named claude/…. Codex's cloud tasks push codex/… branches, Cursor's cloud agents push cursor/…, and GitHub's Copilot coding agent and Devin open pull requests under their own bot accounts. In early October 2026 we counted every public pull request on GitHub carrying one of those signatures, month by month since January 2025, with GitHub's search API (the exact queries are in the methodology).
In September 2026, 5,964,435 of the 23,092,188 pull requests opened in public repositories carried at least one signature: 25.8%. A year earlier the share was 6.9%, and in January 2025 it was 0.05%. Satya Nadella told investors in July that "1 in 3 pull requests on GitHub now involves an agent". His figure includes private repositories and a looser "involves", so the two numbers are consistent.
Everything on GitHub grew, not only agent work. Public pull requests per month rose 4.8 times over the period, from 4.83M to 23.09M, and the ones with no agent signature still grew 3.5 times. Some of that is surely agent work with the signature turned off, and some is ordinary automation (Dependabot alone opened 1.47M public pull requests in December 2025).
Which agents#
Codex was the first agent with a big public footprint: in June 2025, a month after OpenAI launched its cloud agent, codex/ branches made up nine in ten of all agent pull requests. Claude Code passed it in March 2026 and kept going. In September 4,264,006 public pull requests carried its footer, 18.5% of everything opened that month, against 954,292 for Codex.
Public pull requests by agent
Pull requests opened per month in public GitHub repositories, by agent signature, Jan 2025 to Sep 2026
| Series | x | y |
|---|---|---|
| Claude Code (footer) | Jan 2025 | 23 |
| Claude Code (footer) | Feb 2025 | 235 |
| Claude Code (footer) | Mar 2025 | 1.1K |
| Claude Code (footer) | Apr 2025 | 1.2K |
| Claude Code (footer) | May 2025 | 4.7K |
| Claude Code (footer) | Jun 2025 | 22K |
| Claude Code (footer) | Jul 2025 | 34.5K |
| Claude Code (footer) | Aug 2025 | 37.1K |
| Claude Code (footer) | Sep 2025 | 30.9K |
| Claude Code (footer) | Oct 2025 | 59.8K |
| Claude Code (footer) | Nov 2025 | 57.9K |
| Claude Code (footer) | Dec 2025 | 93.7K |
| Claude Code (footer) | Jan 2026 | 168.9K |
| Claude Code (footer) | Feb 2026 | 326.9K |
| Claude Code (footer) | Mar 2026 | 619K |
| Claude Code (footer) | Apr 2026 | 764.8K |
| Claude Code (footer) | May 2026 | 1.2M |
| Claude Code (footer) | Jun 2026 | 1.8M |
| Claude Code (footer) | Jul 2026 | 2.4M |
| Claude Code (footer) | Aug 2026 | 2.1M |
| Claude Code (footer) | Sep 2026 | 4.3M |
| Codex (codex/ branch) | Jan 2025 | 11 |
| Codex (codex/ branch) | Feb 2025 | 0 |
| Codex (codex/ branch) | Mar 2025 | 0 |
| Codex (codex/ branch) | Apr 2025 | 2 |
| Codex (codex/ branch) | May 2025 | 60.1K |
| Codex (codex/ branch) | Jun 2025 | 385.7K |
| Codex (codex/ branch) | Jul 2025 | 257.7K |
| Codex (codex/ branch) | Aug 2025 | 359K |
| Codex (codex/ branch) | Sep 2025 | 348.7K |
| Codex (codex/ branch) | Oct 2025 | 423.7K |
| Codex (codex/ branch) | Nov 2025 | 264K |
| Codex (codex/ branch) | Dec 2025 | 223.8K |
| Codex (codex/ branch) | Jan 2026 | 240.8K |
| Codex (codex/ branch) | Feb 2026 | 421K |
| Codex (codex/ branch) | Mar 2026 | 482.5K |
| Codex (codex/ branch) | Apr 2026 | 464.4K |
| Codex (codex/ branch) | May 2026 | 577.4K |
| Codex (codex/ branch) | Jun 2026 | 515.2K |
| Codex (codex/ branch) | Jul 2026 | 660.5K |
| Codex (codex/ branch) | Aug 2026 | 769.6K |
| Codex (codex/ branch) | Sep 2026 | 954.3K |
| Copilot coding agent | Jan 2025 | 0 |
| Copilot coding agent | Feb 2025 | 0 |
| Copilot coding agent | Mar 2025 | 16 |
| Copilot coding agent | Apr 2025 | 23 |
| Copilot coding agent | May 2025 | 5.1K |
| Copilot coding agent | Jun 2025 | 13.5K |
| Copilot coding agent | Jul 2025 | 38K |
| Copilot coding agent | Aug 2025 | 64.9K |
| Copilot coding agent | Sep 2025 | 90.9K |
| Copilot coding agent | Oct 2025 | 120.9K |
| Copilot coding agent | Nov 2025 | 171.6K |
| Copilot coding agent | Dec 2025 | 176.7K |
| Copilot coding agent | Jan 2026 | 193.8K |
| Copilot coding agent | Feb 2026 | 247.1K |
| Copilot coding agent | Mar 2026 | 319.3K (peak) |
| Copilot coding agent | Apr 2026 | 242.5K |
| Copilot coding agent | May 2026 | 152.9K |
| Copilot coding agent | Jun 2026 | 68.1K |
| Copilot coding agent | Jul 2026 | 69.2K |
| Copilot coding agent | Aug 2026 | 63.2K |
| Copilot coding agent | Sep 2026 | 61.1K |
Sources: GitHub search API; host0 count, early Oct 2026
The other agents are smaller but growing. Pull requests from cursor/ branches, which Cursor's cloud agents create, rose from 36,209 in September 2025 to 418,482 in September 2026, most of it since June. Devin's bot account opened 40,114, up from 3,416. Google's Jules is left out: its bot-account series collapses in February 2026, most likely because its pull requests moved to the user's own identity.
Copilot's coding agent is the exception. Its public pull requests peaked at 319,307 in March 2026 and were down to 61,054 by September. On 20 April GitHub paused new sign-ups for its Pro, Pro+ and Student plans and tightened limits, reopening them gradually in mid-June; we can't tell from public data how much of the fall is that and how much is people moving to other agents. Microsoft still reported 50 million Copilot users in July, so this is a drop in public agent pull requests, not in Copilot use.
September also ended unusually fast. Through August, Claude Code's footer appeared on roughly 480,000 public pull requests a week, about 11% of the total. By the week of 21 September it was 1.25M, 21.7% of the week. Anthropic released Claude Opus 5.5 on 22 September; we can't confirm that's the cause. We checked the search for false positives: in a sample of 400 September matches, 91% contained the exact footer, and they came from 286 different repositories and 253 different authors, so the volume isn't one bot farm.
Most agent pull requests get merged#
Agent pull requests in public repositories are merged more often than the average pull request: 91% of September's signed pull requests had been merged by early October, against 79% of all public pull requests. Read that with care. Many of these are people running an agent in their own repository and merging its work themselves, so "merged" often means "the author kept it", not "a reviewer approved it" (we didn't measure the self-merge share). The agents built to take work from an issue tracker and hand it to someone else for review, Copilot's coding agent and Devin, merge less often.
Merge rate of public pull requests, by agent
% of pull requests opened in Sep 2026 that were merged by early Oct 2026, public GitHub repositories
| Label | Value |
|---|---|
| Claude Code (claude/ branch) | 94% |
| Claude Code (footer) | 93% |
| Any agent signature | 91% |
| Codex | 89% |
| All public pull requests | 79% |
| Cursor | 74% |
| Devin | 71% |
| Copilot coding agent | 70% |
Sources: GitHub search API; host0 count, early Oct 2026
Research on popular projects tells a different story. The AIDev dataset, 456,535 agent pull requests up to June 2025, found agents' acceptance rates in repositories with 500+ stars "15 to 40 percentage points below human performance": 65.3% for Codex, 52.5% for Claude Code, 38.2% for Copilot, against 76.8% for people. A September 2026 study found that merged agent pull requests in popular repositories attract follow-up fixes at 1.62 times the odds of human ones.
Who developers say they use#
The surveys agree with the GitHub count on direction. JetBrains asks professional developers the same question every few months: which AI tools do you use at work? In its 2026 Developer Ecosystem survey (more than 15,000 developers, May to July 2026), about 39% used Claude Code, up from 18% in January and roughly 3% a year earlier. GitHub Copilot fell from 29% to 21% and Cursor from 18% to 12%, while Codex went from 3% to 16%. 90% of professional developers used an AI coding agent at work at least weekly, and 68% daily. JetBrains sells a competing agent, Junie, which it puts at about 9%.
Coding agents professional developers use at work
% of professional developers using each tool at work, JetBrains Developer Ecosystem survey, 15,000+ respondents, May–Jul 2026
| Label | Value |
|---|---|
| Claude Code | ~39% (Jan 2026: 18%) |
| GitHub Copilot | 21% (Jan 2026: 29%) |
| Codex | 16% (Jan 2026: 3%) |
| Cursor | 12% (Jan 2026: 18%) |
| JetBrains AI / Junie | ~9% (JetBrains' own product) |
| OpenCode | 7% |
| Google Antigravity | 6% (stable) |
Source: JetBrains Research, Aug 2026
Stack Overflow's 2026 survey, fielded from late June to early August, asked a looser question, which coding agents or assistants respondents had used in the past year, and got the same leader. Of 12,255 who answered, 65.5% had used Claude Code and 58.7% GitHub Copilot. 73% of people who use coding assistants or agents use them daily, but only 9.6% use AI answers as they are.
Coding agents and assistants used in the past year
% of respondents who answered the question, Stack Overflow Developer Survey 2026, n = 12,255, fielded Jun–Aug 2026
| Label | Value |
|---|---|
| Claude Code | 65.5% |
| GitHub Copilot | 58.7% |
| OpenAI Codex | 29.5% |
| Cursor | 25.5% |
| Google Antigravity | 16.0% |
| Gemini Code Assist | 14.6% |
| OpenCode | 13.6% |
| JetBrains AI | 12.8% |
| Windsurf | 6.5% |
Vendor user counts are harder to compare, because each company defines "users" its own way. OpenAI is the only one that publishes a consistent weekly-active figure for its coding agent, and it rose quickly after the Codex desktop app launched in February 2026: more than 1.6M weekly users in early March, 3M in April and more than 5M by 2 June, a fifth of them knowledge workers rather than developers. Anthropic has only said that Claude Code's weekly active users doubled between January and February 2026, and Microsoft's 50 million Copilot users has no stated time window.
Codex weekly active users
Weekly active users as reported by OpenAI, Mar to Jun 2026
| x | y |
|---|---|
| 4 Mar 2026 | 1.6M (1.6M) |
| 19 Mar 2026 | 2M |
| 9 Apr 2026 | 3M |
| 2 Jun 2026 | 5M (5M) |
Sources: Fortune; OpenAI; TechCrunch; OpenAI
Companies also report how much of their own code agents write, with definitions that differ. Anthropic says that as of May 2026 more than 80% of the code it merges was authored by Claude and that the typical engineer merged 8 times as much code per day as in 2024. Google said in April that 75% of its new code is AI-generated and reviewed by engineers, and Coinbase put its share between 95% and 100%. None of these are audited. For the broader picture of who builds with AI and how often, see our State of vibe coding 2026.
The money#
Coding agents are the clearest case of AI revenue so far, though every figure in this section is a company's own run rate: a recent month's revenue times twelve, unaudited, and usually disclosed around a fundraise. Cursor said it passed $100M in January 2025, $500M in June and $1B in November. Bloomberg reported $2B in February 2026, and Forbes, citing a person familiar with the matter, reported $3B in late April and more than $4B in early June, 75% of it from businesses.
Anthropic broke out Claude Code three times: $500M in September 2025, $1B in November, six months after general availability, and more than $2.5B in February 2026, with enterprises paying more than half. It hasn't separated Claude Code since; its company-wide run rate reached $65B at the end of July. Cognition, which owns Devin and bought Windsurf in 2025, said its run rate grew from $492M in May to almost $900M by 8 September, and on 25 September that it had crossed $1B. OpenAI doesn't break out Codex; WIRED reported, from one source, that it was bringing in just over $1B a year at the end of January.
Run-rate revenue of three coding-agent businesses
Annualized revenue as stated by each company or reported by the press, US dollars, Jan 2025 to Sep 2026
| Series | x | y |
|---|---|---|
| Cursor | 16 Jan 2025 | $100M |
| Cursor | Mar 2025 | $200M |
| Cursor | 6 Jun 2025 | $500M |
| Cursor | 13 Nov 2025 | $1B |
| Cursor | Feb 2026 | $2B |
| Cursor | Apr 2026 | $3B |
| Cursor | 8 Jun 2026 | $4B (~$4B) |
| Claude Code | 2 Sep 2025 | $500M |
| Claude Code | Nov 2025 | $1B |
| Claude Code | 12 Feb 2026 | $2.5B (>$2.5B) |
| Cognition | Jun 2025 | $73M |
| Cognition | 14 Jul 2025 | $155M |
| Cognition | 27 May 2026 | $492M |
| Cognition | 8 Sep 2026 | $900M |
| Cognition | 25 Sep 2026 | $1B ($1B) |
The largest deal was Cursor's sale. In April SpaceX said it had the right to buy Cursor for $60B later in the year or pay $10B for their work together. It signed a merger agreement on 16 June at an implied equity value of $60.0B, and on 14 August the merger took effect, with Cursor's holders receiving 389,289,254 SpaceX Class A shares. Cursor's own post says the process began with a model-training partnership with SpaceXAI. At about $4B of run rate, $60B is roughly 15 times revenue, a lower multiple than Cursor's last two venture rounds (about 20 and 29 times, by our arithmetic).
Latest valuations of coding-agent and app-builder companies
Deal value or post-money valuation at the latest disclosed round, US dollars, Jul 2025 to Sep 2026
| Label | Value |
|---|---|
| Cursor | $60B (Acquired by SpaceX, Aug 2026) |
| Cognition | $48B (Series E, Sep 2026) |
| Lovable | $13.3B (Series C, Aug 2026) |
| Replit | $9B (Series D, Mar 2026) |
| Factory | $5B ($200M round, Sep 2026) |
| Windsurf (Google deal) | $2.4B (License + hires, Jul 2025) |
Sources: SEC 8-K; Cognition; TechCrunch (Lovable); TechCrunch (Replit); Forbes (Factory); TechCrunch (Windsurf)
Businesses are paying most of this. Menlo Ventures, an Anthropic investor, estimated that enterprises spent $4.0B on AI coding tools in 2025, 55% of all departmental AI spending and up from $550M in 2024. The growth has strained the flat-rate plans: Cursor, Anthropic and GitHub all moved heavy agent users onto usage limits or usage-based billing, which we covered in AI token spending in 2026.
The benchmark broke#
For two years the number every launch post led with was SWE-bench Verified, 500 real GitHub issues from Python projects, each checked by tests. The best reported score climbed from 40.6% in August 2024, the month the benchmark launched, to 80.9% with Claude Opus 4.5 in November 2025, and then the labs stopped quoting it. In February 2026 OpenAI audited the tasks its models kept failing and found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions". Every frontier model it tested could reproduce some of the reference fixes from memory, so OpenAI said it had "stopped reporting SWE-bench Verified scores" and recommended others do the same. Anthropic now describes the benchmark as saturated.
SWE-bench Verified, best scores over time
% of 500 tasks resolved; best score reported by a lab or on the leaderboard, and Epoch AI's runs with one shared scaffold, Aug 2024 to Apr 2026
| Series | x | y |
|---|---|---|
| Best reported | 20 Aug 2024 | 40.6% |
| Best reported | 22 Oct 2024 | 49% |
| Best reported | 21 Dec 2024 | 62.2% |
| Best reported | 24 Feb 2025 | 70.3% |
| Best reported | 22 May 2025 | 72.7% |
| Best reported | 5 Aug 2025 | 74.5% |
| Best reported | 7 Aug 2025 | 74.9% |
| Best reported | 29 Sep 2025 | 77.2% |
| Best reported | 24 Nov 2025 | 80.9% (80.9% · Opus 4.5) |
| Epoch AI | 24 Feb 2025 | 61% |
| Epoch AI | 16 Apr 2025 | 62.3% |
| Epoch AI | 22 May 2025 | 70.7% |
| Epoch AI | 5 Aug 2025 | 73.3% |
| Epoch AI | 7 Aug 2025 | 73.6% |
| Epoch AI | 24 Nov 2025 | 76.7% |
| Epoch AI | 5 Feb 2026 | 78.7% |
| Epoch AI | 16 Apr 2026 | 83.5% (83.5% · Opus 4.7) |
Sources: Anthropic; OpenAI; SWE-bench leaderboard; Epoch AI
The replacement didn't last either. OpenAI had pointed people to SWE-bench Pro, a harder set built by Scale AI. In July it estimated that about 30% of Pro's tasks are broken and wrote, as The Stack reported, "We retract our earlier recommendation to adopt SWE-Bench Pro". Launch posts now lead with newer agent benchmarks such as Terminal-Bench, and even the labs hedge: Anthropic's Opus 5.5 announcement says "benchmark margins have become a less reliable guide to real-world differences".
Passing the tests was never the same as shipping. METR asked four maintainers of scikit-learn, Sphinx and pytest to review 296 agent-written fixes blind. Roughly half of the fixes that passed SWE-bench's tests would not have been merged, and the automated grader scored them about 24 percentage points higher than the maintainers did.
How long a task an agent can finish#
METR's time horizon is the most-cited alternative. It measures the length of software task, in the time a skilled person needs, that a model completes half the time. On METR's current task suite it went from about 4 minutes for GPT-4 in March 2023 to 4.9 hours for Claude Opus 4.5 in November 2025, about 12 hours for Claude Opus 4.6 in February 2026 and about 17.4 hours for an early Claude Mythos Preview in April, according to METR's published results. The fit in that file doubles about every 129 days, a little over four months.
METR 50% time horizon of frontier models
Length of task, in skilled-human hours, that a model completes 50% of the time, by model release date, Mar 2023 to Apr 2026
| x | y |
|---|---|
| 14 Mar 2023 | 0.1 h (~4 min · GPT-4) |
| 13 May 2024 | 0.1 h |
| 20 Jun 2024 | 0.2 h |
| 22 Oct 2024 | 0.3 h |
| 5 Dec 2024 | 0.7 h |
| 24 Feb 2025 | 1 h |
| 16 Apr 2025 | 2 h |
| 7 Aug 2025 | 3.4 h |
| 24 Nov 2025 | 4.9 h |
| 11 Dec 2025 | 5.9 h |
| 5 Feb 2026 | 12 h (12 h · Opus 4.6) |
| 7 Apr 2026 | 17.4 h (Mythos Preview) |
Sources: METR results file; METR time horizons
Two caveats keep this from being a straight line to "agents do a week's work". METR says its task suite can't measure reliably above 16 hours, and it has no horizon yet for several 2026 models. Its May report also found that Opus 4.6's horizon was 17.8 hours on low-messiness tasks but 6.4 hours on high-messiness ones. The "worked for 30 hours" claims in launch posts are something else again: how long a model kept going on one task, not how long a task it can reliably finish.
Does it make developers faster#
The best-known study points the other way. In early 2025 METR ran a randomized trial with 16 experienced open-source developers on 246 tasks in repositories they knew well. With AI tools, mostly Cursor with Claude 3.5 and 3.7 Sonnet, tasks took 19% longer, with a confidence interval of +2% to +39%. Afterwards the developers believed AI had made them 20% faster.
METR's late-2025 rerun pointed to a speedup instead: an estimated 18% less time for the developers who returned, with an interval from 38% faster to 9% slower. METR itself calls the result "an unreliable signal", because 30% to 50% of participants said they had held back tasks they didn't want to do without AI. A 2024 trial at Google measured about 21% less time on a single enterprise task, with a large confidence interval.
Controlled studies: change in task time with AI
Change in time to complete a task with AI tools vs without, in randomized studies, 2024 to 2026
| Label | Value |
|---|---|
| METR, early 2025: measured | +19% (slower) (CI +2% to +39%) |
| METR, early 2025: what developers believed | −20% (self-estimate after the study) |
| METR rerun, late 2025: returning developers | −18% (CI −38% to +9%; METR: unreliable) |
| METR rerun, late 2025: new developers | −4% (CI −15% to +9%) |
| Google, 2024: one enterprise task | −21% (large CI) |
Sources: METR 2025; METR 2026 update; Google (arXiv 2410.12944)
Telemetry from companies shows more output and more problems at once. Faros AI compared periods of low and high AI adoption inside the same organizations, two years of data from 22,000 developers on more than 4,000 teams. At high adoption, task throughput per developer was up 33.7% and epics completed up 66%, but bugs per developer were up 54%, pull requests merged with no review up 31.3%, and incidents per pull request up 242.7%. It's correlational, and Faros sells engineering analytics, but the shape matches the follow-up-fix finding above.
More output, more incidents
Change per developer between low and high AI-adoption periods within the same organizations, Faros AI telemetry, 22,000 developers, 2024 to 2026
| Label | Value |
|---|---|
| Epics completed | +66% |
| Task throughput | +33.7% |
| Pull request merge rate | +16.2% |
| Bugs per developer | +54% |
| Pull requests merged without review | +31.3% |
| Incidents per pull request | +242.7% |
Source: Faros AI, Apr 2026
Self-reports are more positive. Anthropic's engineers say they use Claude in 60% of their work for a 50% productivity gain, and the 349 technical workers METR surveyed reported a median 1.4 to 2 times more value from their work. The 2025 trial is the reason to hold such numbers loosely.
What the products look like now#
In 20 months the category went from a terminal tool to agents that run in the cloud, in parallel, from a dedicated desktop app. Codex got a desktop app in February 2026 and moved into the ChatGPT desktop app in July, Cursor rebuilt itself as an agent workspace in April, Google's Antigravity 2.0 dropped the editor in May, and Windsurf became Devin Desktop in June.
How coding agents got here
Selected launches and deals, Feb 2025 to Sep 2026
24 Feb 2025
16 Apr 2025
16 May 2025
19 May 2025
22 May 2025
22 May 2025
25 Jun 2025
14 Jul 2025
20 Oct 2025
28 Oct 2025
18 Nov 2025
9 Dec 2025
2 Feb 2026
2 Apr 2026
20 Apr 2026
2 Jun 2026
14 Aug 2026
22 Sep 2026
22 Sep 2026
The products have converged. All eight in the table below run agents in the background and in parallel, all eight can read an AGENTS.md file of project instructions (Claude Code only when there's no CLAUDE.md, Gemini CLI only when configured), and a full agent starts at $20 a month nearly everywhere. Skills and MCP are the other shared layer; we compared them in Skills vs MCP.
The eight biggest coding agents, side by side
Vendor pages and docs, as of early Oct 2026
Sources: Claude Code docs; Codex pricing; Cursor pricing; GitHub Copilot plans; Antigravity pricing; Devin pricing; Kiro pricing; Amp pricing
The agents' worst failures haven't changed much: an agent with confirmations turned off runs a delete on a path that turns out wider than intended. We collected those incidents in State of AI agent security 2026.
Where agent output goes#
The newest front is what happens after the agent finishes. Until 2026 the last step belonged to hosting companies reached through an MCP server or a plugin: Netlify's MCP server could deploy from mid-2025, and Vercel's learned to in July 2026, returning "a shareable URL without leaving the chat". In June 2026 both frontier labs started hosting it themselves. OpenAI's Sites lets ChatGPT "create, host, refine, and share websites, web apps, and games", with a database, file storage and sign-in, and a new site is visible only to its owner and workspace admins until they share it. Anthropic's artifacts in Claude Code turn a session into a live page that is "viewable only by authenticated members of your org and cannot be made public".
Where agent-built output gets hosted
Hosting and deploy features for coding agents, by launch date, Jun 2025 to Jul 2026
Sources: Netlify; GitHub; Vercel plugin; OpenAI Sites; Anthropic; Vercel MCP
What this means if you build with AI#
Agents now write a large and growing share of public code, and the measurements of how good that code is are behind. The data points to a few habits.
- Pick an agent by trying it on your own repository, not by benchmark. The headline benchmark was retired, its replacement was found broken, and the labs now say benchmark gaps overstate real differences. The surveys show developers switching tools within months; switching is cheap.
- Review agent pull requests as if a stranger wrote them. Half of benchmark-passing fixes wouldn't get past maintainers, merged agent work draws more follow-up fixes, and Faros saw unreviewed merges rise 31.3% alongside incidents.
- Measure your own speed instead of trusting the feeling. The developers in METR's 2025 trial felt 20% faster and were 19% slower. Track cycle time, reverts and incidents before and after you roll an agent out.
- Keep one
AGENTS.mdand the agent stays portable. Every major agent reads it, so your project instructions survive a switch from Cursor to Claude Code to Codex. Claude Code prefersCLAUDE.mdwhen both exist. - Give long tasks to agents, but in checkable pieces. METR's horizon is about 12 to 17 hours at 50% success, and much lower on messy tasks. A day of work is still a coin flip; an hour of work with tests is not.
- Budget for usage, not seats. Run rates grew because agents use far more compute than chat, and every big vendor has moved heavy users to limits or metered billing (details).
- Decide where the output lives before you build it. Hosting now comes bundled with some agents, with very different defaults: owner-only, org-only or public. Choose who should see the result and pick the host to match.
host0 is hosting for whatever your agent creates, from a page to an app with a shared database, at a link you can share that is private by default and free during the beta. The last section above is the market we work in, and agent-built output is what we expect to host more of.
Methodology#
This post draws on four research passes, one per angle: adoption and public footprint, market and money, capability and productivity, and products. Each was limited to primary sources: company posts and changelogs, SEC filings, earnings-call transcripts, survey results pages, METR's published data and research papers. We used aggregator and "statistics" sites only as leads. Paywalled pages were read through Wayback Machine snapshots. The post is dated 7 October 2026 and cites nothing published after that date; our own GitHub counts cover complete months up to September 2026 and were run in early October. Before publishing we re-opened the single-source pages the post leans on for its headline numbers and confirmed the figures were still there.
Our GitHub count. In early October 2026 we ran GitHub's search API (authenticated, search/issues, reading total_count) for each month from January 2025 to September 2026, restricted to pull requests created that month. The signatures were:
- All public pull requests:
is:pr - Claude Code footer:
is:pr "Generated with Claude Code" - Claude Code branches:
is:pr head:claude/ - Codex:
is:pr head:codex/ - Cursor cloud agents:
is:pr head:cursor/ - Copilot coding agent:
is:pr author:app/copilot-swe-agent(andhead:copilot/in the union) - Devin:
is:pr author:app/devin-ai-integration - Jules:
is:pr author:app/google-labs-jules(in the union only) - Any signature: all of the above joined with
ORin one query, so each pull request is counted once - Merge rates: the same queries plus
is:merged
Every result is a floor. It covers public repositories only, and only pull requests that still existed when we counted (deleted and newly private ones drop out). Signatures can be turned off or renamed, and agents that leave none, such as Claude Code with its footer disabled, Cursor's local agent, Aider, OpenCode or Gemini CLI, aren't counted at all. A random sample of 400 September footer matches contained the exact footer in 91% of cases. Our monthly increments agree within about 10% with the open PRarena tracker for the agents it follows.
Other numbers. METR horizons come from METR's own results file. Run rates, user counts and valuations are as the company stated them or as the press reported them; we say which. Survey shares are as published, with their own question wording and sample.
We dropped claims that didn't hold up: a widely quoted "4% of public GitHub commits" written by Claude Code (method not disclosed, and our own commit-level search returned unstable totals), OpenAI's later Codex user figures of 8M to 20M (they combine Codex with ChatGPT Work), a SWE-bench Pro progress figure we couldn't re-open at its source, aggregator revenue estimates for several startups, and "largest acquisition ever" framing for the Cursor deal.
Limitations:
- Signatures are opt-out. The real share of agent-written pull requests is higher than 25.8% by an unknown amount.
- Public is not private. Most commercial code is in private repositories, which no public method can count.
- Run rates are not revenue. They are unaudited and usually disclosed when it suits a fundraise.
- Surveys differ. JetBrains asks about use at work, Stack Overflow about use in the past year; JetBrains sells a competing product.
- Productivity studies are few and small. The one clean randomized result is from early 2025 tools.
Open questions#
- How many people use Claude Code. Anthropic gives growth multiples and revenue, not users; Cursor gives neither users nor, since joining SpaceX, revenue.
- What moved the late-September numbers. Claude Code's public pull requests more than doubled in a few weeks. A model launch, a change in defaults or a wave of automated repositories?
- Why Copilot's coding agent faded on public GitHub. The sign-up pause, competition, or pull requests moving to another identity?
- What happens in private repositories. Microsoft's "1 in 3" covers them; nobody publishes the split.
- How to measure an agent that works for a day. METR's suite tops out around 16 hours, and no maintained, uncontaminated coding benchmark has replaced SWE-bench Verified.
- Whether the productivity gains are real at the company level. There's no randomized study of 2026 agents yet, and the telemetry shows output and incidents rising together.
