Blog

/Research

AI token spending in 2026

Tokens got ~10× cheaper a year, yet Google went from 9.7T to 3.2 quadrillion a month and Claude Code's cost per developer doubled. Where the money goes, in data.

·15 min read

Summarize in ChatGPT
Cover: AI token spending in 2026. Three stats: 3.2 quadrillion tokens a month across Google in May 2026, ~30,000× OpenRouter token growth since Aug 2023, and $13 Claude Code cost per developer per active day, up from $6. Below, a full-width area chart of Google's tokens per month: 9.7T in May 2024, 480T in May 2025, 980T in Jul 2025, 1.3Q in Oct 2025 and 3.2Q in May 2026.
On this page

The price of a token keeps falling. For the same level of capability, inference gets about ten times cheaper every year, by a16z's estimate, and Epoch AI measured declines of 9× to 900× a year depending on the benchmark. If you only read the price lists, AI should be getting cheap.

It isn't, at least not in total. Google now processes over 3.2 quadrillion tokens a month, OpenRouter's volume has grown about 30,000-fold since 2023, Anthropic and OpenAI together report well over $100B in annualised revenue, and companies have started capping what each engineer can spend. We pulled the public numbers on volume, prices, total spend and per-developer cost, and the measured effect of the levers that cut a bill, all as of late September 2026. Vendor figures are the vendor's own claims and are labelled that way. Every number links to its source.

Key findings

  1. 1
    Tokens got cheaper and we bought far more of them. The same capability costs about 10× less each year, while Google's monthly tokens rose from 9.7 trillion to over 3.2 quadrillion in two years, roughly 330×.
  2. 2
    OpenRouter, the one neutral multi-model gauge, grew about 30,000× to a 4.5 quadrillion-token annual run rate in August 2026, doubling every 11 weeks. A natural experiment there shows the mechanism: when OpenAI discounted GPT-5.6 Luna, its daily tokens rose 13.8×, and demand stayed after the discount ended.
  3. 3
    The money followed. Menlo Ventures puts enterprise spend on model APIs at $12.5B in 2025, up from $3.5B in its November 2024 estimate. Anthropic says its run rate went from about $1B at the start of 2025 to $65B at the end of July 2026, and OpenAI's ARR was nearing $70B by late September.
  4. 4
    Agents are the multiplier. Anthropic says an agent uses about 4× the tokens of a chat and a multi-agent system about 15×. Agentic coding uses about 1,000× the tokens of code chat, and almost all of it is input: 99.5% of the tokens in Claude Code's own sample session.
  5. 5
    What a developer costs doubled in a year. Anthropic's guidance for Claude Code went from $6 per developer per day to $13 per active day. Jellyfish puts the median developer at $52 a month and the 90th percentile at $691, and Uber capped spend at $1,500 per employee per tool after using its annual AI budget in four months.
  6. 6
    The frontier price came back up. After falling about 11× from GPT-4 to GPT-5, a new top tier at $10 input and $50 output per million tokens arrived with Claude Fable 5 and GPT-6 Astra. The most expensive model to run Artificial Analysis's index went from $1.78 a task in June to $7.63 in September.
  7. 7
    The levers are large and measured. In Anthropic's tests prompt caching cut agent-loop cost 2.7× to 5.3×. In our worked example, caching, batch and a smaller model take one agent session from $2.20 to $0.16.

3.2 quadrillion

Tokens a month across Google's products and APIs, May 2026 (9.7 trillion two years earlier)

Source: Google I/O 2026

~30,000×

OpenRouter token growth since Aug 2023, to a 4.5 quadrillion a year run rate

Aug 2026

Source: Menlo Ventures

10× a year

How fast the price of the same capability falls

Epoch AI: 9× to 900×

Source: a16z

$13

Claude Code's average cost per developer per active day in 2026, up from $6 a day

Anthropic's guidance

Source: Claude Code docs

Same capability, a tenth of the price each year#

The long-run trend is not in dispute. In November 2021, a model good enough to score 42 on the MMLU benchmark (GPT-3) cost $60 per million tokens. Three years later Llama 3.2 3B did it for $0.06, a thousandfold drop that a16z's Guido Appenzeller summed up as "LLMflation": for equivalent performance, "the cost is decreasing by 10x every year". Epoch AI's analysis finds the same direction with a wide range, 9× to 900× a year depending on which score you hold fixed.

Flagship list prices show it too, up to a point. OpenAI's flagship went from $37.50 per million tokens (GPT-4, March 2023, blending input and output 3:1) to $3.44 (GPT-5, August 2025), about 11× cheaper in 29 months. Anthropic's top model fell from $30 (Claude 3 Opus) to $10 (Opus 4.5) in November 2025. Then both lines turned up.

Price of each lab's top model

Blended list price per 1M tokens (3 input : 1 output), US dollars, Mar 2023 to Sep 2026

Price of each lab's top model
Seriesxy
OpenAI flagshipMar 2023$37.5 (GPT-4)
OpenAI flagshipNov 2023$15
OpenAI flagshipMay 2024$7.5
OpenAI flagshipAug 2024$4.38
OpenAI flagshipApr 2025$3.5
OpenAI flagshipAug 2025$3.44 (GPT-5)
OpenAI flagshipSep 2026$20 (GPT-6 Astra)
Anthropic top modelMar 2024$30 (Claude 3 Opus)
Anthropic top modelNov 2025$10 (Opus 4.5)
Anthropic top modelJun 2026$20 (Fable 5)
Anthropic top modelSep 2026$20
Prices of the top general model fell about 11× at OpenAI between GPT-4 and GPT-5, then a premium tier came back in 2026 at $10 input and $50 output per million. OpenAI's intermediate 2026 models are left out because their launch dates aren't confirmed.

Sources: OpenAI; Anthropic; Claude pricing; host0 blended-price arithmetic

The labs now sell two tiers. At the top, Claude Fable 5 (June 2026) and GPT-6 Astra (September) both list at $10 per million input tokens and $50 per million output. Below them sits a workhorse tier that keeps getting cheaper: Claude Opus 5.5 at $4 and $20, 20% below Opus 5, and GPT-6.1 Sol at $2 and $10, which TechCrunch reports OpenAI pitches as near-Astra intelligence at one-fifth of the price. Google announced Gemini 4 Argon on 30 September at an introductory $2 and $10, rising to $4 and $20.

The price card at the end of September 2026

Blended list price per 1M tokens (3 input : 1 output), US dollars, models available by 30 Sep 2026

The price card at the end of September 2026
LabelValue
GPT-6 Astra$20 ($10 in / $50 out)
Claude Fable 5.1$20 ($10 in / $50 out)
Claude Opus 5.5$8 ($4 / $20)
GPT-6.1 Sol$4 ($2 / $10)
Claude Sonnet 5.5$4 ($2 / $10)
Claude Haiku 4.5$2 ($1 / $5)
Gemini 3.8 Flash$1.50 ($0.75 / $3.75, introductory)
DeepSeek V4.1-Flash$0.53 ($0.30 / $1.20, peak hours)
GPT-6 Luna$0.20 ($0.10 / $0.50)
Highlighted bars are the new top tier. The cheapest model on the card costs 1/100 of the most expensive one per token.

Sources: OpenAI pricing; Claude pricing; Gemini pricing; DeepSeek pricing

What companies actually pay sits far below list. Ramp, which tracks what businesses on its platform spend, measured an effective price of $0.68 per million tokens in early September, down 41% from a 2026 peak of $1.15 in March, as customers moved toward cheaper standard models: frontier models (Opus, Fable, Sol) were 45% of token share, down from 53% in August. One more quiet change cuts the other way. Anthropic's tokenizer since Opus 4.7 produces about 30% more tokens for the same text, so a lower price per token isn't always a lower price per page.

Tokens grew faster than prices fell#

If prices fall 10× a year and usage grows less than that, spend shrinks. It didn't. At Google I/O in May 2026, Sundar Pichai said Google was processing 9.7 trillion tokens a month across its products two years earlier, about 480 trillion a year earlier, and "over 3.2 quadrillion" now. That's roughly 50× in the first year and 7× in the second.

Tokens Google processes a month

Across Google's products and APIs, quadrillions of tokens per month (1 quadrillion = 1,000 trillion), May 2024 to May 2026

Tokens Google processes a month
xy
May 20240.0097 (9.7 trillion)
May 20250.48
Jul 20250.98
Oct 20251.3
May 20263.2 (3.2 quadrillion)
About 330× in two years. Google stopped giving a monthly total after I/O 2026 and now reports API tokens per minute instead.

Sources: Google I/O 2026; Google I/O 2025; Alphabet Q2 2025; Alphabet Q3 2025

Developer traffic alone tells the same story. Google's API went from 7 billion tokens a minute in October 2025 to about 22 billion in July 2026, which is roughly 0.95 quadrillion a month by our arithmetic. OpenAI's API was at 6 billion a minute at DevDay 2025 and more than 15 billion by March 2026, its latest public figure. Both companies say they are short of capacity: Google said in July it continues "to be supply constrained", and Microsoft expects to stay constrained at least through 2026.

API tokens per minute, Google and OpenAI

Tokens per minute through each company's developer APIs, billions, Oct 2025 to Jul 2026

API tokens per minute, Google and OpenAI
Seriesxy
Google29 Oct 20257B
Google4 Feb 202610B
Google29 Apr 202616B
Google19 May 202619B
Google22 Jul 202622B
OpenAI6 Oct 20256B
OpenAI31 Mar 202615B
Google's API roughly tripled in three quarters. OpenAI hasn't published a figure since spring, so the gap at the right edge is a gap in disclosure, not a ranking.

Sources: Alphabet earnings remarks; Google I/O 2026; TechCrunch (DevDay 2025); OpenAI

OpenRouter is the closest thing to a neutral meter, since it routes traffic to 400+ models from many labs. When it rebranded in August 2023 it handled about 3 billion tokens a week. In May 2026 it reported 25 trillion a week, and when Stripe agreed to buy it in August, its investor Menlo Ventures put it at "~30,000x to a 4.5+ quadrillion-token annual run rate", doubling every 11 weeks. Our own conversion of those figures (86.5 trillion a week against 3 billion) gives about 28,800×, so the round number holds.

OpenRouter's annual token run rate

Tokens a year at each date's pace, Aug 2023 to Aug 2026

OpenRouter's annual token run rate
LabelValue
Aug 2023~0.16T (3B a week × 52)
Mid-2024~10T
Jun 2025100T+
Nov 2025~260T (5T a week × 52)
May 20261.5 quadrillion
Aug 20264.5 quadrillion+
The first bars are too small to see at this scale, which is the point: about 30,000× in three years. Two bars are weekly figures we multiplied by 52.

Sources: Menlo Ventures (Aug 2026); Menlo Ventures (May 2026); Menlo Ventures (Jun 2025); Business Wire

China reports in tokens a day, and its numbers are as steep. ByteDance's Doubao went from 4 trillion a day in December 2024 to 180 trillion in June 2026. China's data administration said national consumption rose from 100 billion tokens a day at the start of 2024 to more than 140 trillion in March 2026, over a thousandfold.

Tokens a day in China

Average daily tokens, trillions, ByteDance's Doubao and China's national total, Jan 2024 to Jun 2026

Tokens a day in China
Seriesxy
DoubaoDec 20244T
DoubaoMar 202512.7T
DoubaoSep 202530T
DoubaoDec 202550T
DoubaoApr 2026120T
DoubaoJun 2026180T
China, nationalJan 20240.1T
China, nationalDec 2025100T
China, nationalMar 2026140T
Doubao is most of China's total, according to CEIBS, which is why the two lines sit close together. The national series has only three public points.

Sources: SCMP; China Daily; TMTPost; SCIO; CEIBS

A price cut, measured#

Economists call this the Jevons paradox: make a resource cheaper to use and total use rises enough to swamp the saving. Satya Nadella invoked it during the DeepSeek week in January 2025 (NPR). This summer OpenRouter got to test it. OpenAI ran a discount on two GPT-5.6 models and then cut their list prices, for an effective discount of 90% on Luna and 60% on Terra, while GPT-5.6 Sol stayed at full price. OpenRouter's analysis found daily Luna tokens rose 13.8× and Terra 5.6×, against 1.11× for Sol. OpenAI's share of all OpenRouter tokens went from 7.1% to 12.4%, and the period after the programme saw 1.38× the tokens of the discount period itself.

What a price cut did to demand

Rise in daily tokens on OpenRouter during OpenAI's GPT-5.6 discount, multiple of the period before, summer 2026

What a price cut did to demand
LabelValue
GPT-5.6 Luna13.8× (90% effective discount)
GPT-5.6 Terra5.6× (60% effective discount)
GPT-5.6 Sol1.11× (no discount (control))
The bigger the discount, the bigger the jump. Demand didn't fall back when prices went up again: usage after the programme averaged 1.38× the discount period.

Source: OpenRouter, GPT 5.6 discounts and Jevons paradox

The supply side compounds it. Microsoft said in July 2025 that software optimisation alone was delivering 90% more tokens per GPU than a year earlier. Cheaper tokens to produce, cheaper tokens to buy, and every cut so far has been met with more use.

Where the money went#

Menlo Ventures, which surveys enterprise buyers, estimated that companies spent $3.5B on model APIs in November 2024 and $8.4B by mid-2025. Its full-year figure for 2025 was $12.5B for foundation model APIs, out of $37B of enterprise generative AI spend, itself up 3.2× from 2024. Gartner, with a different method, forecasts spend on foundation models of $23.4B in 2026, double its 2025 estimate.

Enterprise spend on model APIs

US dollars; Menlo Ventures estimates and a Gartner forecast, Nov 2024 to 2026

Enterprise spend on model APIs
LabelValue
Nov 2024$3.5B (Menlo)
Mid-2025$8.4B (Menlo)
2025$12.5B (Menlo, full year)
2026$23.4B (Gartner forecast, different method)
Highlighted bars come from one source; the faded one is a forecast from another with a different definition, shown for direction only.

Sources: Menlo Ventures (mid-2025); Menlo Ventures (2025); Gartner via CXOToday

Most of that went to three labs. By Menlo's estimate, Anthropic took 40% of enterprise LLM API spend in 2025, OpenAI 27% and Google 21%, and in coding Anthropic led OpenAI 54% to 21%. Ramp's card data points the same way in 2026: by August, 43.8% of US businesses paid Anthropic for subscriptions or tokens, against 39.8% for OpenAI.

The labs' own revenue figures are bigger than any API-spend estimate, because they include subscriptions, consumer apps and coding tools. Anthropic says its run rate went from $87M at the start of 2024 to about $1B at the start of 2025, $14B in February 2026 and over $30B in April; CNBC reported $65B at the end of July. OpenAI's CFO put its annualised revenue at over $20B for 2025, and Axios reported its ARR nearing $70B on 29 September.

Annualised revenue, Anthropic and OpenAI

Run-rate or annual recurring revenue as reported, US dollars, Dec 2023 to Sep 2026

Annualised revenue, Anthropic and OpenAI
Seriesxy
AnthropicJan 2024$87M
AnthropicJan 2025$1B
AnthropicAug 2025$5B
AnthropicDec 2025$9B
AnthropicFeb 2026$14B
AnthropicApr 2026$30B
AnthropicMay 2026$47B
AnthropicJul 2026$65B
OpenAIDec 2023$2B
OpenAIDec 2024$6B
OpenAIDec 2025$20B
OpenAIAug 2026$40B
OpenAISep 2026$70B (~$70B)
Run rates annualise the latest month; they aren't audited revenue. Anthropic's prospectus shows about $4.6B of revenue for calendar 2025, while its run rate ended that year near $9B. Anthropic's last point is end of July; OpenAI's is late September.

Sources: Anthropic; CNBC; Reuters; Bloomberg; Axios

Selling tokens is also expensive to do. Anthropic's IPO prospectus, as reported by Reuters, shows $7.33B of compute and infrastructure cost in 2025, more than half its operating expenses, and an operating loss of over $8B on about $4.6B of revenue. Two customers made up nearly a quarter of that revenue.

Agents are why#

A chat is a question and an answer. An agent is a loop: it reads files, calls tools, reads the results, and sends the whole growing conversation back to the model on every turn. Anthropic's engineers put a number on it in 2025: agents use about 4× the tokens of a chat, multi-agent systems about 15×, and token usage alone explained 80% of the performance differences on their browsing benchmark. A Stanford-affiliated study of eight models on SWE-bench found agentic coding uses about 1,000× the tokens of code chat, "with input tokens rather than output tokens driving the overall cost", and that runs of the same task differ by up to 30× without more tokens buying more accuracy.

How many more tokens agents use

Multiples reported by each source, against the baseline each one names, 2025 to 2026

SetupTokensBaselineSource
Agent~4×A chat interactionAnthropic, Jun 2025
Claude Code agent team, plan mode~7×A standard Claude Code sessionClaude Code docs, 2026
Multi-agent system~15×A chat interactionAnthropic, Jun 2025
Agentic coding (SWE-bench)~1,000×Code chat or code reasoningBai et al., Apr 2026
Same task, same model, different runsup to 30×The cheapest runBai et al., Apr 2026
Average prompt per request on OpenRouter~4×Early 2025 (1.5K → 6K+ tokens)OpenRouter, Dec 2025
The baselines differ, so read each row on its own. All of them point the same way: work done in a loop costs an order of magnitude or more over a single exchange.

Sources: Anthropic; Claude Code docs; Bai et al.; OpenRouter State of AI

It shows up in the market totals. On OpenRouter, programming went from about 11% of tokens in early 2025 to over half by the end of the year. OpenRouter suggests 6 February 2026 may have been the last day humans used more tokens than agents; from then to August, agent token volume grew about 14×, from 0.51 trillion to 7.3 trillion, while human volume grew 2.8×. Google says its own internal developer tools went from half a trillion tokens a day in March 2026 to more than three trillion in May.

Where an agent session's tokens go#

Claude Code's documentation includes a sample session that makes the shape concrete. On Sonnet 4.6, it used 1.2k input tokens, 5.3k output, 940k cache reads and 50k cache writes, for $0.55. Almost everything the model read was its own context, re-sent each turn and served from cache.

One Claude Code session, token by token

The sample session in Anthropic's Claude Code docs, Sonnet 4.6 list prices, US dollars

Token typeTokensPrice per 1MCostShare of cost
Cache reads940,000$0.30$0.28251%
Cache writes (5 min)50,000$3.75$0.18834%
Output5,300$15.00$0.08014%
Fresh input1,200$3.00$0.004under 1%
Total996,500$0.55
99.5% of the tokens are input-side, 187 for every token the model wrote. Cache reads are the biggest cost even at a tenth of the input price. Without caching, the same session would cost about $3.05.

Sources: Claude Code docs, Sep 2026 snapshot; Claude pricing; host0 arithmetic

That's why the agent era hits input pricing and caching hardest. Manus, which builds a general agent, reported an input-to-output ratio of about 100:1 and called the cache hit rate "the single most important metric" for a production agent. The Stanford-affiliated SWE-bench study found that cache reads dominate both token volume and dollar cost in every phase of a run.

What a developer costs now#

Anthropic's own guidance tracks the shift. Through early 2026, the Claude Code docs said the average cost was $6 per developer per day, with 90% of users under $12. In April 2026 that changed to "around $13 per developer per active day and $150-250 per developer per month, with costs remaining below $30 per active day for 90% of users" across enterprise deployments. The average rose 2.2× and the 90th percentile 2.5×, while list prices per token were falling.

Independent data puts numbers on the spread. Jellyfish, which measures engineering teams, looked at 12,000 developers at 200 companies in the first quarter of 2026. The median developer used about 51 million tokens a month and spent $52.38; at the 90th percentile it was about 380 million tokens and $691.14. Its mid-year report says per-developer consumption rose about 18.6× in nine months, with a mean of $248 a month and a 99th percentile of $2,452.

What a developer spends on tokens a month

US dollars per developer per month, by percentile and by guideline, 2026

What a developer spends on tokens a month
LabelValue
Jellyfish median$52.38 (Q1 2026)
Claude Code average (Anthropic)$150–250 (enterprise deployments)
Jellyfish 75th percentile$226.58 (Q1 2026)
Jellyfish mean$248 (H1 2026 report)
Jellyfish 90th percentile$691.14 (Q1 2026)
Uber's cap per tool$1,500 (per employee, from Jun 2026)
Jellyfish 99th percentile$2,452 (H1 2026 report)
The distribution has a long tail: Jellyfish's 99th percentile is $2,452 a month, against a median of $52 in its first-quarter sample. Uber's cap sits between the 90th and 99th percentiles.

Sources: Jellyfish (Apr 2026); Jellyfish (H1 2026); Claude Code docs; TechCrunch

More tokens bought more output, but at a falling rate. In Jellyfish's data the heaviest users merged about twice as many pull requests for about ten times the tokens, and the cost per merged PR rose from $0.28 in the lowest usage tier to $89.32 in the highest.

Companies noticed. Uber had told staff to use AI as much as possible and ranked usage on internal leaderboards, then, after it blew through its annual AI budget in four months, set a monthly cap of $1,500 per employee per agentic coding tool. The CEO of Faros AI told TechCrunch one of his engineers spent $40,000 on tokens in a month, and Priceline's Cursor renewal came back 4–5× more expensive. Meta's Adam Mosseri said the burn rate of a strong engineer might equal their salary. Gartner predicts AI coding costs will exceed the average developer's salary by 2028, and says 23% of tech leaders already spend $200 to $500 per developer per month on tokens.

Pricing moved to match. In July 2025 Anthropic said one Claude Max subscriber had consumed "tens of thousands" of dollars of model usage on a $200 plan and introduced weekly limits. Cursor switched its Pro plan to $20 of included usage and refunded surprised users. And on 1 June 2026 GitHub moved every Copilot plan to usage-based billing, because under a flat per-request price "a quick chat question and a multi-hour autonomous coding session can cost the user the same amount".

The frontier got expensive again#

Two things push the cost of a hard task up even as tokens get cheaper: the new premium tier, and models that think longer at maximum effort. Artificial Analysis prices a full run of its Intelligence Index for each model. In June the most expensive model was Claude Opus 4.8 at $1.78 per task. In September, Claude Fable 5.1 at maximum effort cost $7.63 per task, using 78K output tokens, and GPT-6 Astra reached the same score for $3.26 with 27K. The cheap end barely moved: DeepSeek V4 Pro ran at $0.04 per task.

Cost to run one task of the Artificial Analysis index

US dollars per task at each model's maximum effort, Intelligence Index v4.1, Jun and Sep 2026

Cost to run one task of the Artificial Analysis index
LabelValue
Claude Fable 5.1 (Sep)$7.63 (78K output tokens a task)
GPT-6 Astra (Sep)$3.26 (27K output tokens a task)
Claude Opus 4.8 (Jun)$1.78 (most expensive in June)
DeepSeek V4 Pro (Jun)$0.04
The top of the chart rose about 4.3× in three months as $10/$50 models arrived. In June the cheapest leading model already cost about 1/45 of the most expensive one.

Sources: Artificial Analysis (Jun 2026); Artificial Analysis (Sep 2026)

For a fixed task, though, the cost of getting it done keeps falling. In 2024 the SWE-agent paper reported GPT-4 Turbo resolving 12.47% of SWE-bench issues at $1.59 per resolved issue. On the official SWE-bench Verified leaderboard with the same simple agent, Claude 4.6 Opus resolves 75.6% for about $0.73 per resolved issue by our arithmetic, and several open-weight models get within a few points for a few cents.

Cost per resolved coding issue

US dollars per resolved SWE-bench Verified issue with mini-SWE-agent (total cost ÷ issues solved), share solved in notes, Jul 2025 to Feb 2026

Cost per resolved coding issue
LabelValue
GPT-4o$7.08 (21.6% solved)
Claude 4 Opus$1.67 (67.6%)
Gemini 3 Pro$1.38 (69.6%)
Claude 4.5 Opus (high)$0.98 (76.8%)
Claude 4.6 Opus$0.73 (75.6%)
GPT-5.2 (high)$0.65 (72.8%)
Gemini 3 Flash (high)$0.47 (75.8%)
Kimi K2.5 (high)$0.21 (70.8%)
MiniMax M2.5 (high)$0.10 (75.8%)
DeepSeek V3.2 Reasoner$0.05 (60.0%)
Same agent, same 500 tasks. Better models fail less and loop less, so the cost per solved issue falls faster than the price per token. Submissions price at their providers' rates, so treat small gaps as noise.

Sources: SWE-bench leaderboard data; host0 arithmetic

Labs now market this directly. Anthropic says Opus 5.5 will cost 40% less than Opus 5 on typical workloads at default settings, partly from a lower price and partly from using fewer tokens. Its own cost guide shows the same effect can run backwards: Fable 5.1 matched Fable 5's coding score for 43% less per solved task, but on a long research benchmark the same upgrade cost 41% more per task, because the newer model did more work.

The levers that cut the bill#

The best public evidence on what works is Anthropic's guide to optimizing for cost and intelligence, which reports measured runs for each lever (the numbers below are from its September version). Its first finding is about why agents get expensive: "A 40-turn task sends its first turn 40 times, so task cost grows with roughly the square of turn count."

Caching comes first. A cache read costs a tenth of the input price on Anthropic, OpenAI's newest models and Google's, and 2% at DeepSeek. Across a day of real traffic, Anthropic says agent loops read a median 84% of their input from cache, and the top tenth 94% or more. In its measurements caching cut agent-loop cost by a factor of 2.7 to 5.3, and a small triage agent's bill by 83%. On a deep research benchmark, Fable 5.1 fell from $37.94 to $7.12 per task with caching. Breaking the cache mid-session costs real money too: changing the effort setting and adding a tool took one triage session from $0.81 to $0.95.

Batch is the second free lever. Anthropic, OpenAI and Google all take 50% off for work that can wait up to 24 hours, and it stacks with caching. DeepSeek went further and made off-peak hours half price. Speed now goes the other way: OpenAI's Fast mode and Anthropic's fast mode on Opus 5.5 cost twice the standard price.

Then trim the context. Tool definitions are input on every turn. When Anthropic grew a test agent's catalog to 502 tools, run cost nearly doubled from $0.55 to $1.02; with tool search it stayed at $0.56, with the same accuracy. Cursor's on-demand tool loading cut total agent tokens by 46.9% in runs that used an MCP tool, and GitHub's token work cut one agentic workflow's effective tokens by 62%. Data matters as much as tools: a roughly 91,000-token table pasted into every request let Sonnet 5 answer 6 of 25 questions for $5.01, while uploading the file and letting the model run code answered all 25 for $0.40.

What each lever saved, as measured

Reduction in cost or tokens on each source's own workload, 2026

What each lever saved, as measured
LabelValue
Upload data, use code execution92% ($5.01 → $0.40, Anthropic)
Prompt caching, triage agent83% (Anthropic)
Token work, GitHub Auto-Triage62% (effective tokens, 109 runs)
On-demand MCP tools, Cursor46.9% (total agent tokens)
Tool search at 502 tools45% (Anthropic)
Prune stale tool results39% (long run, Anthropic)
Compaction32% (long run, Anthropic)
Context editing0% (long run; +74% on a short one)
Each bar is one measurement on one workload, so they don't add up and aren't directly comparable. Context editing, which Anthropic launched in 2025 with an 84% saving, saved nothing on its 2026 long run.

Sources: Anthropic cost guide (Sep 2026); GitHub; Cursor

Choose the model by cost per solved task, not per token. On a coding benchmark, Anthropic found Fable 5.1 at low effort solved 88.6% of tasks for $0.54 each, against 77.4% at $0.84 for Sonnet 5: "11 more points for 35% less per solved task, despite a per-token price five times higher." Running everything at low effort and re-running only the failures held the pass rate at about half the cost. Splitting work between an orchestrator and cheap worker models saved time but saved money in only two of the situations it measured. Routing research points the same way: LMSYS's RouteLLM cut cost by over 85% on one benchmark while keeping 95% of GPT-4's quality, by sending only hard queries to the big model.

A worked example: one agent session, priced#

To put the levers side by side, take one agent session with a million input tokens and 20,000 output tokens, a 50:1 ratio that is if anything light on input next to Manus's 100:1. In the cached case, 90% of the input is read from cache and 10% written to it. Prices are each vendor's list prices at the end of September 2026.

One agent session, lever by lever

Cost of 1M input + 20K output tokens on Claude Sonnet 5, then Claude Haiku 4.5, US dollars, list prices, Sep 2026

One agent session, lever by lever
LabelValue
Sonnet 5, no cache$2.20
Sonnet 5, 90% cached$0.63
Sonnet 5, cached + batch$0.32
Haiku 4.5, cached + batch$0.16
Caching does most of the work (−71%), batch halves what's left, and a smaller model halves it again: −93% in total, if the smaller model can do the job.

Sources: Claude pricing; host0 arithmetic

The same session on nine models

Cost of 1M input + 20K output tokens, US dollars, list prices at the end of September 2026

ModelNo cache90% cachedCached + batch
Claude Fable 5.1$11.00$2.48$1.24
GPT-6 Astra$11.00$3.15$1.58
Claude Opus 5.5$4.40$1.08$0.54
Claude Sonnet 5$2.20$0.63$0.32
GPT-6.1 Sol$2.20$0.54$0.27
Claude Haiku 4.5$1.10$0.32$0.16
Gemini 3.8 Flash (introductory)$0.83$0.22$0.11
DeepSeek V4.1-Flash (peak)$0.32$0.06$0.03 off-peak
GPT-6 Luna$0.11$0.03$0.016
From the most expensive uncached run to the cheapest cached-and-batched one, the same session spans about 700×. Within any one model, caching cuts 71% to 82%. Gemini and DeepSeek bill newly cached tokens at the normal input price.

Sources: Claude pricing; OpenAI pricing; Gemini pricing; DeepSeek pricing; host0 arithmetic

The formula is short enough to run on your own numbers: without caching, cost = input tokens × input price + output tokens × output price; with caching, replace the input term with 90% at the cache-read price plus 10% at the cache-write price. Sonnet 5 cached, for example, is 0.9 × $0.20 + 0.1 × $2.50 + 0.02 × $10 = $0.63.

What this means if you build with AI#

Cheaper tokens won't make your bill smaller on their own. The habits that do, from the data above:

  1. Turn on caching before anything else, and keep the prefix stable. It's the largest measured lever by a wide margin. Put the system prompt and tools first, append rather than edit, and don't change models, effort or tools mid-session.
  2. Watch your cache hit rate, not your token count. A median agent loop reads 84% of its input from cache. If yours is far below that, something is breaking the prefix.
  3. Batch anything nobody is waiting for. Evals, backfills and scheduled jobs get 50% off on top of caching. Don't pay for fast modes unless latency is the product.
  4. Load tools and data on demand. Use tool search for large catalogs, ship small default toolsets, and give the model files and code execution instead of pasting tables into the prompt, which in Anthropic's test was about 12× cheaper and more accurate.
  5. Measure cost per solved task. A model that costs five times more per token can be 35% cheaper per solved task. Start at low effort, re-run failures at higher effort, and send simple subagent work to a small model.
  6. Budget per developer, with a ceiling. Anthropic's own planning number is $150–250 a month, but the tail is long: Jellyfish's 99th percentile spends $2,452 a month. Set alerts and caps before the bill does it for you, as Uber learned.
  7. Don't put a model where a lookup will do. Every request that runs through a model in production is a recurring token cost. If an app only needs to show and store data, serve it without one.

That last point is where host0 fits: it's a cloud for small software, where an agent builds an app and host0 gives it a URL and a shared database. A static frontend on host0 with platform records answers every visit without calling a model, so the only token bill is the one you paid to build it.

Methodology#

This post draws on four research passes, one per angle: token volume, prices and total spend, agents and per-developer cost, and savings levers. Each was limited to primary sources: earnings remarks and company posts, pricing pages and changelogs, investor reports, research papers and the archived documentation of Claude Code and the Claude API. Statistics and aggregator sites were used only as leads. The post is dated 1 October 2026 and uses nothing published after that date. Where a page might have changed since, we cite a Wayback Machine snapshot from before it (the Claude Code costs page from 28 September 2026 and Anthropic's cost guide from 15 September). Before publishing we re-opened every single-source page the post leans on for a headline number and confirmed the quoted sentence was still there; one wording, Anthropic's "40% less" for Opus 5.5, was corrected to match the page.

Several numbers are our own arithmetic:

  • Blended prices weight input and output 3:1, Epoch AI's convention: (3 × input + output) ÷ 4.
  • OpenRouter run rates convert weekly figures to a year by multiplying by 52 (two bars), and back for the ~30,000× check.
  • Google's API tokens per month multiply tokens per minute by 43,200 minutes in a 30-day month.
  • The session cost and the worked example use each vendor's list prices: input, output, cache read and cache write per million tokens.
  • Cost per resolved SWE-bench issue divides each submission's average cost per task by its share of tasks solved.

Limitations:

  • Companies count tokens differently. Google's monthly total covers all its products, its per-minute figure only the API; OpenAI's covers only the API; China reports daily averages. Never add them up.
  • Run rates aren't revenue. They annualise the latest month, and most 2026 figures for Anthropic and OpenAI are press reports of investor updates, not filings.
  • Market sizes depend on method. Menlo changed its method between its 2024 and 2025 estimates, and Gartner uses another. We chart them with labels, not as one series.
  • Vendor savings are measured on the vendor's workloads. Anthropic's cost guide, Cursor's and GitHub's results are their own tests; yours will differ.
  • Per-developer data is a sample. Jellyfish covers its own customers; Anthropic's guidance covers enterprise deployments of Claude Code.

Open questions#

  1. How many tokens Anthropic serves. It publishes revenue and, via its prospectus, compute cost, but no token volume.
  2. OpenAI's API volume since spring. Its last public figure is "more than 15 billion" tokens a minute in March and April 2026.
  3. How revenue splits between API and subscriptions at Anthropic and OpenAI. Claude Code's last official run rate was over $2.5B in February 2026.
  4. Real cache hit rates across the industry. Anthropic's 84% median is the only published aggregate.
  5. What a typical agent session costs in the wild. Beyond Claude Code's sample session and Jellyfish's monthly percentiles, nobody publishes per-session token counts for coding agents.
  6. How much the new tokenizers raised bills. Anthropic says its tokenizer since Opus 4.7 makes about 30% more tokens for the same text, but not how that varies between code and prose.
  7. A 2026 enterprise API spend estimate on the same method as 2025's $12.5B. Gartner's forecast is the only 2026 figure, and it measures something different.
ResearchTokensAgents