Blog

/Research

State of AI coding agents 2026

One in four public GitHub PRs now carries a coding-agent signature, 4.26M of them Claude Code's in September, and SpaceX paid $60B for Cursor. The data.

·18 min read

Summarize in ChatGPT
State of AI coding agents 2026 cover: a large 25.8% (share of public GitHub PRs carrying a coding-agent signature in Sep 2026, about 1 in 4) above monthly bars rising from 0.05% in Jan 2025 to 25.8% in Sep 2026, with tiles for 4.26M PRs with Claude Code's footer in Sep 2026, a 17.4 h METR time horizon (Claude Mythos Preview) and developers 19% slower in METR's 2025 RCT.
On this page

In February 2025 Anthropic shipped Claude Code as a research preview, a command-line tool you could hand a task and walk away from. Twenty months later every large AI company sells a coding agent that runs in the background, opens its own pull requests and works on several tasks at once. The agents have become one of the biggest businesses in AI: a coding-agent startup was bought for $60 billion, and the labs now host what the agents build.

Most of what gets written about coding agents is a launch post or a vendor's own adoption figure. We wanted numbers we could check. So we counted the agents' public footprint on GitHub ourselves, month by month, and set it next to the developer surveys, the revenue disclosures, the benchmark record and the productivity studies. Vendor numbers are the vendor's claims and are labelled that way, and every number below links to its source.

Key findings

  1. 1
    One in four public pull requests now carries an agent's signature. In September 2026, 5.96M of the 23.09M pull requests opened in public GitHub repositories (25.8%) had a coding-agent branch name, footer or bot author, by our count. A year earlier it was 6.9%. It's a floor: agents that leave no signature are invisible.
  2. 2
    Claude Code is the largest footprint. 4.26M public pull requests in September carried its default "Generated with Claude Code" footer, 138 times as many as a year earlier and 4.5 times Codex's 954K. Public pull requests from GitHub's own Copilot coding agent fell 81% from their March peak.
  3. 3
    The surveys show the same shift. JetBrains found about 39% of professional developers using Claude Code at work in mid-2026, up from 18% in January, while GitHub Copilot fell from 29% to 21% and Cursor from 18% to 12%. 90% use some coding agent at least weekly.
  4. 4
    The money followed. Cursor's run rate reportedly reached about $4B in June 2026, up from $100M in January 2025, and SpaceX completed its $60B all-stock acquisition on 14 August. Anthropic says Claude Code passed a $2.5B run rate in February; Cognition says it crossed $1B in September.
  5. 5
    The main coding benchmark broke. SWE-bench Verified saturated around 81%, and OpenAI stopped reporting it after finding flawed tests in 59.4% of the hard tasks it audited. METR found that about half of test-passing agent fixes would not be merged by the projects' maintainers.
  6. 6
    Task length keeps doubling. METR's 50% time horizon went from 4 minutes (GPT-4, 2023) to about 12 hours (Claude Opus 4.6, February 2026) and about 17.4 hours (Claude Mythos Preview, April 2026), though METR says anything above 16 hours is unreliable on its current tasks.
  7. 7
    Productivity evidence is still mixed. In METR's 2025 trial, experienced developers took 19% longer with AI while believing they were 20% faster. Faros's 2026 telemetry shows 33.7% more tasks done and 242.7% more incidents per pull request at high AI adoption.
  8. 8
    The labs now host the output. OpenAI launched Sites to host what Codex and ChatGPT build, Anthropic launched artifacts in Claude Code for org-only pages, and Vercel's MCP server learned to deploy.

25.8%

Public GitHub pull requests opened in Sep 2026 that carry a coding-agent signature

Our count; a floor

Source: GitHub search API

4.26M

Public pull requests with the “Generated with Claude Code” footer, Sep 2026

Our count; 138× Sep 2025

Source: GitHub search API

~$4B / $60B

Cursor's reported run rate in June 2026, and the value of SpaceX's all-stock deal for it

Run rate is a reported claim

Sources: Forbes; SEC 8-K

~17.4 h

METR 50% time horizon, Claude Mythos Preview (early), Apr 2026

METR: above 16 h is unreliable

Source: METR

One in four public pull requests#

Coding agents leave fingerprints. Claude Code adds a "🤖 Generated with Claude Code" footer to the pull requests it writes unless you turn it off, and its web and desktop sessions push branches named claude/…. Codex's cloud tasks push codex/… branches, Cursor's cloud agents push cursor/…, and GitHub's Copilot coding agent and Devin open pull requests under their own bot accounts. In early October 2026 we counted every public pull request on GitHub carrying one of those signatures, month by month since January 2025, with GitHub's search API (the exact queries are in the methodology).

In September 2026, 5,964,435 of the 23,092,188 pull requests opened in public repositories carried at least one signature: 25.8%. A year earlier the share was 6.9%, and in January 2025 it was 0.05%. Satya Nadella told investors in July that "1 in 3 pull requests on GitHub now involves an agent". His figure includes private repositories and a looser "involves", so the two numbers are consistent.

Share of public GitHub pull requests with a coding-agent signature

Pull requests opened per month in public repositories, % carrying an agent branch prefix, footer or bot author, Jan 2025 to Sep 2026

Share of public GitHub pull requests with a coding-agent signature
xy
Jan 20250.1%
Feb 20250.1%
Mar 20250.1%
Apr 20250.1%
May 20251.3%
Jun 20256.8%
Jul 20255.7%
Aug 20257.8%
Sep 20256.9% (6.9%)
Oct 20258.4%
Nov 20259.9%
Dec 20259.1%
Jan 202610.4%
Feb 202612.7%
Mar 202615.2%
Apr 202615.2%
May 202618.4%
Jun 202620%
Jul 202621.7%
Aug 202620.5%
Sep 202625.8% (25.8%)
The first jump, in June 2025, is Codex's cloud agent; the steady climb since late 2025 is mostly Claude Code. A floor: agents that leave no branch prefix, footer or bot author aren't counted.

Sources: GitHub search API; host0 count, early Oct 2026

Everything on GitHub grew, not only agent work. Public pull requests per month rose 4.8 times over the period, from 4.83M to 23.09M, and the ones with no agent signature still grew 3.5 times. Some of that is surely agent work with the signature turned off, and some is ordinary automation (Dependabot alone opened 1.47M public pull requests in December 2025).

Which agents#

Codex was the first agent with a big public footprint: in June 2025, a month after OpenAI launched its cloud agent, codex/ branches made up nine in ten of all agent pull requests. Claude Code passed it in March 2026 and kept going. In September 4,264,006 public pull requests carried its footer, 18.5% of everything opened that month, against 954,292 for Codex.

Public pull requests by agent

Pull requests opened per month in public GitHub repositories, by agent signature, Jan 2025 to Sep 2026

Public pull requests by agent
Seriesxy
Claude Code (footer)Jan 202523
Claude Code (footer)Feb 2025235
Claude Code (footer)Mar 20251.1K
Claude Code (footer)Apr 20251.2K
Claude Code (footer)May 20254.7K
Claude Code (footer)Jun 202522K
Claude Code (footer)Jul 202534.5K
Claude Code (footer)Aug 202537.1K
Claude Code (footer)Sep 202530.9K
Claude Code (footer)Oct 202559.8K
Claude Code (footer)Nov 202557.9K
Claude Code (footer)Dec 202593.7K
Claude Code (footer)Jan 2026168.9K
Claude Code (footer)Feb 2026326.9K
Claude Code (footer)Mar 2026619K
Claude Code (footer)Apr 2026764.8K
Claude Code (footer)May 20261.2M
Claude Code (footer)Jun 20261.8M
Claude Code (footer)Jul 20262.4M
Claude Code (footer)Aug 20262.1M
Claude Code (footer)Sep 20264.3M
Codex (codex/ branch)Jan 202511
Codex (codex/ branch)Feb 20250
Codex (codex/ branch)Mar 20250
Codex (codex/ branch)Apr 20252
Codex (codex/ branch)May 202560.1K
Codex (codex/ branch)Jun 2025385.7K
Codex (codex/ branch)Jul 2025257.7K
Codex (codex/ branch)Aug 2025359K
Codex (codex/ branch)Sep 2025348.7K
Codex (codex/ branch)Oct 2025423.7K
Codex (codex/ branch)Nov 2025264K
Codex (codex/ branch)Dec 2025223.8K
Codex (codex/ branch)Jan 2026240.8K
Codex (codex/ branch)Feb 2026421K
Codex (codex/ branch)Mar 2026482.5K
Codex (codex/ branch)Apr 2026464.4K
Codex (codex/ branch)May 2026577.4K
Codex (codex/ branch)Jun 2026515.2K
Codex (codex/ branch)Jul 2026660.5K
Codex (codex/ branch)Aug 2026769.6K
Codex (codex/ branch)Sep 2026954.3K
Copilot coding agentJan 20250
Copilot coding agentFeb 20250
Copilot coding agentMar 202516
Copilot coding agentApr 202523
Copilot coding agentMay 20255.1K
Copilot coding agentJun 202513.5K
Copilot coding agentJul 202538K
Copilot coding agentAug 202564.9K
Copilot coding agentSep 202590.9K
Copilot coding agentOct 2025120.9K
Copilot coding agentNov 2025171.6K
Copilot coding agentDec 2025176.7K
Copilot coding agentJan 2026193.8K
Copilot coding agentFeb 2026247.1K
Copilot coding agentMar 2026319.3K (peak)
Copilot coding agentApr 2026242.5K
Copilot coding agentMay 2026152.9K
Copilot coding agentJun 202668.1K
Copilot coding agentJul 202669.2K
Copilot coding agentAug 202663.2K
Copilot coding agentSep 202661.1K
Claude Code's footer passed Codex's branches in March 2026 and was 4.5× Codex by September. Copilot's agent fell 81% from its March peak, starting the month GitHub paused new individual sign-ups.

Sources: GitHub search API; host0 count, early Oct 2026

The other agents are smaller but growing. Pull requests from cursor/ branches, which Cursor's cloud agents create, rose from 36,209 in September 2025 to 418,482 in September 2026, most of it since June. Devin's bot account opened 40,114, up from 3,416. Google's Jules is left out: its bot-account series collapses in February 2026, most likely because its pull requests moved to the user's own identity.

Copilot's coding agent is the exception. Its public pull requests peaked at 319,307 in March 2026 and were down to 61,054 by September. On 20 April GitHub paused new sign-ups for its Pro, Pro+ and Student plans and tightened limits, reopening them gradually in mid-June; we can't tell from public data how much of the fall is that and how much is people moving to other agents. Microsoft still reported 50 million Copilot users in July, so this is a drop in public agent pull requests, not in Copilot use.

September also ended unusually fast. Through August, Claude Code's footer appeared on roughly 480,000 public pull requests a week, about 11% of the total. By the week of 21 September it was 1.25M, 21.7% of the week. Anthropic released Claude Opus 5.5 on 22 September; we can't confirm that's the cause. We checked the search for false positives: in a sample of 400 September matches, 91% contained the exact footer, and they came from 286 different repositories and 253 different authors, so the volume isn't one bot farm.

Most agent pull requests get merged#

Agent pull requests in public repositories are merged more often than the average pull request: 91% of September's signed pull requests had been merged by early October, against 79% of all public pull requests. Read that with care. Many of these are people running an agent in their own repository and merging its work themselves, so "merged" often means "the author kept it", not "a reviewer approved it" (we didn't measure the self-merge share). The agents built to take work from an issue tracker and hand it to someone else for review, Copilot's coding agent and Devin, merge less often.

Merge rate of public pull requests, by agent

% of pull requests opened in Sep 2026 that were merged by early Oct 2026, public GitHub repositories

Merge rate of public pull requests, by agent
LabelValue
Claude Code (claude/ branch)94%
Claude Code (footer)93%
Any agent signature91%
Codex89%
All public pull requests79%
Cursor74%
Devin71%
Copilot coding agent70%
The highlighted bar is the baseline. Not a quality measure: in a one-person repository, merged usually means the author kept the change. September's cohort is still partly open, so its rates will rise a little.

Sources: GitHub search API; host0 count, early Oct 2026

Research on popular projects tells a different story. The AIDev dataset, 456,535 agent pull requests up to June 2025, found agents' acceptance rates in repositories with 500+ stars "15 to 40 percentage points below human performance": 65.3% for Codex, 52.5% for Claude Code, 38.2% for Copilot, against 76.8% for people. A September 2026 study found that merged agent pull requests in popular repositories attract follow-up fixes at 1.62 times the odds of human ones.

Who developers say they use#

The surveys agree with the GitHub count on direction. JetBrains asks professional developers the same question every few months: which AI tools do you use at work? In its 2026 Developer Ecosystem survey (more than 15,000 developers, May to July 2026), about 39% used Claude Code, up from 18% in January and roughly 3% a year earlier. GitHub Copilot fell from 29% to 21% and Cursor from 18% to 12%, while Codex went from 3% to 16%. 90% of professional developers used an AI coding agent at work at least weekly, and 68% daily. JetBrains sells a competing agent, Junie, which it puts at about 9%.

Coding agents professional developers use at work

% of professional developers using each tool at work, JetBrains Developer Ecosystem survey, 15,000+ respondents, May–Jul 2026

Coding agents professional developers use at work
LabelValue
Claude Code~39% (Jan 2026: 18%)
GitHub Copilot21% (Jan 2026: 29%)
Codex16% (Jan 2026: 3%)
Cursor12% (Jan 2026: 18%)
JetBrains AI / Junie~9% (JetBrains' own product)
OpenCode7%
Google Antigravity6% (stable)
Claude Code roughly doubled in five months while Copilot and Cursor lost share. Developers can use several tools, so the bars don't add up to 100%.

Source: JetBrains Research, Aug 2026

Stack Overflow's 2026 survey, fielded from late June to early August, asked a looser question, which coding agents or assistants respondents had used in the past year, and got the same leader. Of 12,255 who answered, 65.5% had used Claude Code and 58.7% GitHub Copilot. 73% of people who use coding assistants or agents use them daily, but only 9.6% use AI answers as they are.

Coding agents and assistants used in the past year

% of respondents who answered the question, Stack Overflow Developer Survey 2026, n = 12,255, fielded Jun–Aug 2026

Coding agents and assistants used in the past year
LabelValue
Claude Code65.5%
GitHub Copilot58.7%
OpenAI Codex29.5%
Cursor25.5%
Google Antigravity16.0%
Gemini Code Assist14.6%
OpenCode13.6%
JetBrains AI12.8%
Windsurf6.5%
“Used in the past year” is a lower bar than JetBrains' “use at work”, so the shares are higher. The question changed from 2025, so don't compare it with last year's survey.

Source: Stack Overflow Developer Survey 2026

Vendor user counts are harder to compare, because each company defines "users" its own way. OpenAI is the only one that publishes a consistent weekly-active figure for its coding agent, and it rose quickly after the Codex desktop app launched in February 2026: more than 1.6M weekly users in early March, 3M in April and more than 5M by 2 June, a fifth of them knowledge workers rather than developers. Anthropic has only said that Claude Code's weekly active users doubled between January and February 2026, and Microsoft's 50 million Copilot users has no stated time window.

Codex weekly active users

Weekly active users as reported by OpenAI, Mar to Jun 2026

Codex weekly active users
xy
4 Mar 20261.6M (1.6M)
19 Mar 20262M
9 Apr 20263M
2 Jun 20265M (5M)
Each point is “more than” the figure shown. OpenAI's later 8M–20M figures count Codex and ChatGPT Work together, with no time window, so they aren't on this line.

Sources: Fortune; OpenAI; TechCrunch; OpenAI

Companies also report how much of their own code agents write, with definitions that differ. Anthropic says that as of May 2026 more than 80% of the code it merges was authored by Claude and that the typical engineer merged 8 times as much code per day as in 2024. Google said in April that 75% of its new code is AI-generated and reviewed by engineers, and Coinbase put its share between 95% and 100%. None of these are audited. For the broader picture of who builds with AI and how often, see our State of vibe coding 2026.

The money#

Coding agents are the clearest case of AI revenue so far, though every figure in this section is a company's own run rate: a recent month's revenue times twelve, unaudited, and usually disclosed around a fundraise. Cursor said it passed $100M in January 2025, $500M in June and $1B in November. Bloomberg reported $2B in February 2026, and Forbes, citing a person familiar with the matter, reported $3B in late April and more than $4B in early June, 75% of it from businesses.

Anthropic broke out Claude Code three times: $500M in September 2025, $1B in November, six months after general availability, and more than $2.5B in February 2026, with enterprises paying more than half. It hasn't separated Claude Code since; its company-wide run rate reached $65B at the end of July. Cognition, which owns Devin and bought Windsurf in 2025, said its run rate grew from $492M in May to almost $900M by 8 September, and on 25 September that it had crossed $1B. OpenAI doesn't break out Codex; WIRED reported, from one source, that it was bringing in just over $1B a year at the end of January.

Run-rate revenue of three coding-agent businesses

Annualized revenue as stated by each company or reported by the press, US dollars, Jan 2025 to Sep 2026

Run-rate revenue of three coding-agent businesses
Seriesxy
Cursor16 Jan 2025$100M
CursorMar 2025$200M
Cursor6 Jun 2025$500M
Cursor13 Nov 2025$1B
CursorFeb 2026$2B
CursorApr 2026$3B
Cursor8 Jun 2026$4B (~$4B)
Claude Code2 Sep 2025$500M
Claude CodeNov 2025$1B
Claude Code12 Feb 2026$2.5B (>$2.5B)
CognitionJun 2025$73M
Cognition14 Jul 2025$155M
Cognition27 May 2026$492M
Cognition8 Sep 2026$900M
Cognition25 Sep 2026$1B ($1B)
Each line stops at the last figure its company disclosed. Cognition's first point is Devin alone; its July 2025 point adds Windsurf's $82M, our arithmetic. Cursor's 2026 points are press reports.

Sources: Cursor; TechCrunch; Forbes; Anthropic; Cognition; Cognition

The largest deal was Cursor's sale. In April SpaceX said it had the right to buy Cursor for $60B later in the year or pay $10B for their work together. It signed a merger agreement on 16 June at an implied equity value of $60.0B, and on 14 August the merger took effect, with Cursor's holders receiving 389,289,254 SpaceX Class A shares. Cursor's own post says the process began with a model-training partnership with SpaceXAI. At about $4B of run rate, $60B is roughly 15 times revenue, a lower multiple than Cursor's last two venture rounds (about 20 and 29 times, by our arithmetic).

Latest valuations of coding-agent and app-builder companies

Deal value or post-money valuation at the latest disclosed round, US dollars, Jul 2025 to Sep 2026

Latest valuations of coding-agent and app-builder companies
LabelValue
Cursor$60B (Acquired by SpaceX, Aug 2026)
Cognition$48B (Series E, Sep 2026)
Lovable$13.3B (Series C, Aug 2026)
Replit$9B (Series D, Mar 2026)
Factory$5B ($200M round, Sep 2026)
Windsurf (Google deal)$2.4B (License + hires, Jul 2025)
Cognition went from $26B in May to $48B in September. Windsurf's $2.4B bought Google a license and its leadership, not the company; Cognition bought the rest for an undisclosed sum.

Sources: SEC 8-K; Cognition; TechCrunch (Lovable); TechCrunch (Replit); Forbes (Factory); TechCrunch (Windsurf)

Businesses are paying most of this. Menlo Ventures, an Anthropic investor, estimated that enterprises spent $4.0B on AI coding tools in 2025, 55% of all departmental AI spending and up from $550M in 2024. The growth has strained the flat-rate plans: Cursor, Anthropic and GitHub all moved heavy agent users onto usage limits or usage-based billing, which we covered in AI token spending in 2026.

The benchmark broke#

For two years the number every launch post led with was SWE-bench Verified, 500 real GitHub issues from Python projects, each checked by tests. The best reported score climbed from 40.6% in August 2024, the month the benchmark launched, to 80.9% with Claude Opus 4.5 in November 2025, and then the labs stopped quoting it. In February 2026 OpenAI audited the tasks its models kept failing and found that "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions". Every frontier model it tested could reproduce some of the reference fixes from memory, so OpenAI said it had "stopped reporting SWE-bench Verified scores" and recommended others do the same. Anthropic now describes the benchmark as saturated.

SWE-bench Verified, best scores over time

% of 500 tasks resolved; best score reported by a lab or on the leaderboard, and Epoch AI's runs with one shared scaffold, Aug 2024 to Apr 2026

SWE-bench Verified, best scores over time
Seriesxy
Best reported20 Aug 202440.6%
Best reported22 Oct 202449%
Best reported21 Dec 202462.2%
Best reported24 Feb 202570.3%
Best reported22 May 202572.7%
Best reported5 Aug 202574.5%
Best reported7 Aug 202574.9%
Best reported29 Sep 202577.2%
Best reported24 Nov 202580.9% (80.9% · Opus 4.5)
Epoch AI24 Feb 202561%
Epoch AI16 Apr 202562.3%
Epoch AI22 May 202570.7%
Epoch AI5 Aug 202573.3%
Epoch AI7 Aug 202573.6%
Epoch AI24 Nov 202576.7%
Epoch AI5 Feb 202678.7%
Epoch AI16 Apr 202683.5% (83.5% · Opus 4.7)
Vendor scores use different scaffolds and sometimes subsets of the tasks; Epoch's line runs every model the same way. The reported line ends in November 2025: from February 2026 the labs stopped quoting the benchmark.

Sources: Anthropic; OpenAI; SWE-bench leaderboard; Epoch AI

The replacement didn't last either. OpenAI had pointed people to SWE-bench Pro, a harder set built by Scale AI. In July it estimated that about 30% of Pro's tasks are broken and wrote, as The Stack reported, "We retract our earlier recommendation to adopt SWE-Bench Pro". Launch posts now lead with newer agent benchmarks such as Terminal-Bench, and even the labs hedge: Anthropic's Opus 5.5 announcement says "benchmark margins have become a less reliable guide to real-world differences".

Passing the tests was never the same as shipping. METR asked four maintainers of scikit-learn, Sphinx and pytest to review 296 agent-written fixes blind. Roughly half of the fixes that passed SWE-bench's tests would not have been merged, and the automated grader scored them about 24 percentage points higher than the maintainers did.

How long a task an agent can finish#

METR's time horizon is the most-cited alternative. It measures the length of software task, in the time a skilled person needs, that a model completes half the time. On METR's current task suite it went from about 4 minutes for GPT-4 in March 2023 to 4.9 hours for Claude Opus 4.5 in November 2025, about 12 hours for Claude Opus 4.6 in February 2026 and about 17.4 hours for an early Claude Mythos Preview in April, according to METR's published results. The fit in that file doubles about every 129 days, a little over four months.

METR 50% time horizon of frontier models

Length of task, in skilled-human hours, that a model completes 50% of the time, by model release date, Mar 2023 to Apr 2026

METR 50% time horizon of frontier models
xy
14 Mar 20230.1 h (~4 min · GPT-4)
13 May 20240.1 h
20 Jun 20240.2 h
22 Oct 20240.3 h
5 Dec 20240.7 h
24 Feb 20251 h
16 Apr 20252 h
7 Aug 20253.4 h
24 Nov 20254.9 h
11 Dec 20255.9 h
5 Feb 202612 h (12 h · Opus 4.6)
7 Apr 202617.4 h (Mythos Preview)
On a linear scale the curve is almost flat until 2025. Confidence intervals are wide (Mythos Preview: about 8.5 to 55 hours), and METR says measurements above 16 hours are unreliable with its current tasks.

Sources: METR results file; METR time horizons

Two caveats keep this from being a straight line to "agents do a week's work". METR says its task suite can't measure reliably above 16 hours, and it has no horizon yet for several 2026 models. Its May report also found that Opus 4.6's horizon was 17.8 hours on low-messiness tasks but 6.4 hours on high-messiness ones. The "worked for 30 hours" claims in launch posts are something else again: how long a model kept going on one task, not how long a task it can reliably finish.

Does it make developers faster#

The best-known study points the other way. In early 2025 METR ran a randomized trial with 16 experienced open-source developers on 246 tasks in repositories they knew well. With AI tools, mostly Cursor with Claude 3.5 and 3.7 Sonnet, tasks took 19% longer, with a confidence interval of +2% to +39%. Afterwards the developers believed AI had made them 20% faster.

METR's late-2025 rerun pointed to a speedup instead: an estimated 18% less time for the developers who returned, with an interval from 38% faster to 9% slower. METR itself calls the result "an unreliable signal", because 30% to 50% of participants said they had held back tasks they didn't want to do without AI. A 2024 trial at Google measured about 21% less time on a single enterprise task, with a large confidence interval.

Controlled studies: change in task time with AI

Change in time to complete a task with AI tools vs without, in randomized studies, 2024 to 2026

Controlled studies: change in task time with AI
LabelValue
METR, early 2025: measured+19% (slower) (CI +2% to +39%)
METR, early 2025: what developers believed−20% (self-estimate after the study)
METR rerun, late 2025: returning developers−18% (CI −38% to +9%; METR: unreliable)
METR rerun, late 2025: new developers−4% (CI −15% to +9%)
Google, 2024: one enterprise task−21% (large CI)
Bar length is the size of the change; the label gives its sign. Negative means faster. Only the highlighted bar is a clean measurement of a slowdown; the 2025 developers believed the opposite of what was measured.

Sources: METR 2025; METR 2026 update; Google (arXiv 2410.12944)

Telemetry from companies shows more output and more problems at once. Faros AI compared periods of low and high AI adoption inside the same organizations, two years of data from 22,000 developers on more than 4,000 teams. At high adoption, task throughput per developer was up 33.7% and epics completed up 66%, but bugs per developer were up 54%, pull requests merged with no review up 31.3%, and incidents per pull request up 242.7%. It's correlational, and Faros sells engineering analytics, but the shape matches the follow-up-fix finding above.

More output, more incidents

Change per developer between low and high AI-adoption periods within the same organizations, Faros AI telemetry, 22,000 developers, 2024 to 2026

More output, more incidents
LabelValue
Epics completed+66%
Task throughput+33.7%
Pull request merge rate+16.2%
Bugs per developer+54%
Pull requests merged without review+31.3%
Incidents per pull request+242.7%
Highlighted bars are costs, the rest output. Correlational: Faros compares each organization's own low- and high-adoption periods.

Source: Faros AI, Apr 2026

Self-reports are more positive. Anthropic's engineers say they use Claude in 60% of their work for a 50% productivity gain, and the 349 technical workers METR surveyed reported a median 1.4 to 2 times more value from their work. The 2025 trial is the reason to hold such numbers loosely.

What the products look like now#

In 20 months the category went from a terminal tool to agents that run in the cloud, in parallel, from a dedicated desktop app. Codex got a desktop app in February 2026 and moved into the ChatGPT desktop app in July, Cursor rebuilt itself as an agent workspace in April, Google's Antigravity 2.0 dropped the editor in May, and Windsurf became Devin Desktop in June.

How coding agents got here

Selected launches and deals, Feb 2025 to Sep 2026

  1. 24 Feb 2025

    Claude Code research preview

    A command-line agent, launched with Claude 3.7 Sonnet.
  2. 16 Apr 2025

    Codex CLI

    OpenAI's open-source terminal agent.
  3. 16 May 2025

    Codex cloud agent

    Many tasks in parallel, each in its own cloud sandbox.
  4. 19 May 2025

    GitHub Copilot coding agent

    Assign an issue, get a pull request (preview).
  5. 25 Jun 2025

    Gemini CLI

    Open source, with a free tier of 1,000 requests a day.
  6. 14 Jul 2025

    Windsurf splits

    Google pays $2.4B for a license and its leaders; Cognition buys the rest.
  7. 20 Oct 2025

    Claude Code on the web

    Parallel cloud sessions.
  8. 28 Oct 2025

    GitHub Agent HQ

    Third-party agents inside GitHub.
  9. 18 Nov 2025

    Google Antigravity

    An agent-first IDE, free in preview.
  10. 9 Dec 2025

    Agentic AI Foundation

    MCP, AGENTS.md and goose move to the Linux Foundation.
  11. 2 Feb 2026

    Codex desktop app

    Parallel agent threads with worktrees.
  12. 2 Apr 2026

    Cursor 3

    Rebuilt as a workspace for agents.
  13. 20 Apr 2026

    Copilot pauses sign-ups

    New individual plans paused for about two months.
  14. 2 Jun 2026

    OpenAI Sites; Devin Desktop

    OpenAI hosts what Codex builds; Windsurf is renamed.
  15. 14 Aug 2026

    SpaceX completes the Cursor deal

    $60B, all stock.
  16. 22 Sep 2026

    Claude Opus 5.5

The first year was about getting agents out of the editor and into the cloud; 2026 was about desktop apps for running many of them, and consolidation.

Sources: Anthropic; OpenAI Codex changelog; GitHub; Cursor; Cognition; SEC 8-K

The products have converged. All eight in the table below run agents in the background and in parallel, all eight can read an AGENTS.md file of project instructions (Claude Code only when there's no CLAUDE.md, Gemini CLI only when configured), and a full agent starts at $20 a month nearly everywhere. Skills and MCP are the other shared layer; we compared them in Skills vs MCP.

The eight biggest coding agents, side by side

Vendor pages and docs, as of early Oct 2026

AgentWhere it runsOpen sourceDefault modelEntry priceCloud and parallel agents
Claude Code (Anthropic)Terminal, IDEs, desktop, webNoOpus 5.5 on paid plansPro $20/moSubagents, agent teams, cloud sessions
Codex (OpenAI)CLI, IDE, ChatGPT desktop app, cloudCLI only (Apache-2.0)GPT-6 Sol on paid plansFree tier; Plus $20/moCloud tasks, parallel threads
Cursor (SpaceX)IDE, agent workspace, CLI, cloudNoAuto; Grok 4.7 and Composer 2.5 in its own poolFree tier; Pro $20/moLocal and cloud agents
GitHub Copilot (Microsoft)IDEs, CLI, GitHub cloud agentChat extension (MIT)Auto selectionFree tier; Pro $10/moCloud agent on paid plans
Antigravity + Gemini CLI (Google)Desktop app, CLI, IDEGemini CLI only (Apache-2.0)Gemini 3.xFree for individualsAsync and scheduled agents
Devin (Cognition)Cloud, Devin Desktop, CLINoSWE-2 plus third-party modelsFree tier; Pro $20/moParallel cloud Devins
Kiro (AWS)IDE, CLI, webNoClaude and GPT models on paid plansFree tier; Pro $20/moAutonomous agent
Amp (Amp Frontier)CLI, desktop, webNoVaries by roleFree with your own keys; $20/moRemote “orbs”, subagents
Defaults and prices change often; this is a snapshot. “Open source” means an OSI license on the agent itself.

Sources: Claude Code docs; Codex pricing; Cursor pricing; GitHub Copilot plans; Antigravity pricing; Devin pricing; Kiro pricing; Amp pricing

The agents' worst failures haven't changed much: an agent with confirmations turned off runs a delete on a path that turns out wider than intended. We collected those incidents in State of AI agent security 2026.

Where agent output goes#

The newest front is what happens after the agent finishes. Until 2026 the last step belonged to hosting companies reached through an MCP server or a plugin: Netlify's MCP server could deploy from mid-2025, and Vercel's learned to in July 2026, returning "a shareable URL without leaving the chat". In June 2026 both frontier labs started hosting it themselves. OpenAI's Sites lets ChatGPT "create, host, refine, and share websites, web apps, and games", with a database, file storage and sign-in, and a new site is visible only to its owner and workspace admins until they share it. Anthropic's artifacts in Claude Code turn a session into a live page that is "viewable only by authenticated members of your org and cannot be made public".

Where agent-built output gets hosted

Hosting and deploy features for coding agents, by launch date, Jun 2025 to Jul 2026

LaunchedFeatureWhat it doesWho can see the result
Jun 2025Netlify MCP serverThe agent creates a site and deploys itA public Netlify URL
Jul 2025GitHub SparkPrompt to full-stack app, with hosting, data and authA Spark URL
Mar 2026Vercel plugin for coding agentsA /deploy command in Claude Code and CursorA Vercel URL
Jun 2026OpenAI SitesBuilds and hosts sites and apps, with a database and storageOwner only by default; can be shared up to public
Jun 2026Claude Code artifactsA session becomes a live page at a claude.ai linkMembers of the organization only
Jul 2026Vercel MCP deploysShips files and returns a linkWhoever has the link
The two labs took opposite defaults: OpenAI hosts full apps that can be shared up to the public web, while Claude Code's artifacts never leave the organization.

Sources: Netlify; GitHub; Vercel plugin; OpenAI Sites; Anthropic; Vercel MCP

What this means if you build with AI#

Agents now write a large and growing share of public code, and the measurements of how good that code is are behind. The data points to a few habits.

  1. Pick an agent by trying it on your own repository, not by benchmark. The headline benchmark was retired, its replacement was found broken, and the labs now say benchmark gaps overstate real differences. The surveys show developers switching tools within months; switching is cheap.
  2. Review agent pull requests as if a stranger wrote them. Half of benchmark-passing fixes wouldn't get past maintainers, merged agent work draws more follow-up fixes, and Faros saw unreviewed merges rise 31.3% alongside incidents.
  3. Measure your own speed instead of trusting the feeling. The developers in METR's 2025 trial felt 20% faster and were 19% slower. Track cycle time, reverts and incidents before and after you roll an agent out.
  4. Keep one AGENTS.md and the agent stays portable. Every major agent reads it, so your project instructions survive a switch from Cursor to Claude Code to Codex. Claude Code prefers CLAUDE.md when both exist.
  5. Give long tasks to agents, but in checkable pieces. METR's horizon is about 12 to 17 hours at 50% success, and much lower on messy tasks. A day of work is still a coin flip; an hour of work with tests is not.
  6. Budget for usage, not seats. Run rates grew because agents use far more compute than chat, and every big vendor has moved heavy users to limits or metered billing (details).
  7. Decide where the output lives before you build it. Hosting now comes bundled with some agents, with very different defaults: owner-only, org-only or public. Choose who should see the result and pick the host to match.

host0 is hosting for whatever your agent creates, from a page to an app with a shared database, at a link you can share that is private by default and free during the beta. The last section above is the market we work in, and agent-built output is what we expect to host more of.

Methodology#

This post draws on four research passes, one per angle: adoption and public footprint, market and money, capability and productivity, and products. Each was limited to primary sources: company posts and changelogs, SEC filings, earnings-call transcripts, survey results pages, METR's published data and research papers. We used aggregator and "statistics" sites only as leads. Paywalled pages were read through Wayback Machine snapshots. The post is dated 7 October 2026 and cites nothing published after that date; our own GitHub counts cover complete months up to September 2026 and were run in early October. Before publishing we re-opened the single-source pages the post leans on for its headline numbers and confirmed the figures were still there.

Our GitHub count. In early October 2026 we ran GitHub's search API (authenticated, search/issues, reading total_count) for each month from January 2025 to September 2026, restricted to pull requests created that month. The signatures were:

  • All public pull requests: is:pr
  • Claude Code footer: is:pr "Generated with Claude Code"
  • Claude Code branches: is:pr head:claude/
  • Codex: is:pr head:codex/
  • Cursor cloud agents: is:pr head:cursor/
  • Copilot coding agent: is:pr author:app/copilot-swe-agent (and head:copilot/ in the union)
  • Devin: is:pr author:app/devin-ai-integration
  • Jules: is:pr author:app/google-labs-jules (in the union only)
  • Any signature: all of the above joined with OR in one query, so each pull request is counted once
  • Merge rates: the same queries plus is:merged

Every result is a floor. It covers public repositories only, and only pull requests that still existed when we counted (deleted and newly private ones drop out). Signatures can be turned off or renamed, and agents that leave none, such as Claude Code with its footer disabled, Cursor's local agent, Aider, OpenCode or Gemini CLI, aren't counted at all. A random sample of 400 September footer matches contained the exact footer in 91% of cases. Our monthly increments agree within about 10% with the open PRarena tracker for the agents it follows.

Other numbers. METR horizons come from METR's own results file. Run rates, user counts and valuations are as the company stated them or as the press reported them; we say which. Survey shares are as published, with their own question wording and sample.

We dropped claims that didn't hold up: a widely quoted "4% of public GitHub commits" written by Claude Code (method not disclosed, and our own commit-level search returned unstable totals), OpenAI's later Codex user figures of 8M to 20M (they combine Codex with ChatGPT Work), a SWE-bench Pro progress figure we couldn't re-open at its source, aggregator revenue estimates for several startups, and "largest acquisition ever" framing for the Cursor deal.

Limitations:

  • Signatures are opt-out. The real share of agent-written pull requests is higher than 25.8% by an unknown amount.
  • Public is not private. Most commercial code is in private repositories, which no public method can count.
  • Run rates are not revenue. They are unaudited and usually disclosed when it suits a fundraise.
  • Surveys differ. JetBrains asks about use at work, Stack Overflow about use in the past year; JetBrains sells a competing product.
  • Productivity studies are few and small. The one clean randomized result is from early 2025 tools.

Open questions#

  1. How many people use Claude Code. Anthropic gives growth multiples and revenue, not users; Cursor gives neither users nor, since joining SpaceX, revenue.
  2. What moved the late-September numbers. Claude Code's public pull requests more than doubled in a few weeks. A model launch, a change in defaults or a wave of automated repositories?
  3. Why Copilot's coding agent faded on public GitHub. The sign-up pause, competition, or pull requests moving to another identity?
  4. What happens in private repositories. Microsoft's "1 in 3" covers them; nobody publishes the split.
  5. How to measure an agent that works for a day. METR's suite tops out around 16 hours, and no maintained, uncontaminated coding benchmark has replaced SWE-bench Verified.
  6. Whether the productivity gains are real at the company level. There's no randomized study of 2026 agents yet, and the telemetry shows output and incidents rising together.
ResearchCoding agents