The price of a token keeps falling. For the same level of capability, inference gets about ten times cheaper every year, by a16z's estimate, and Epoch AI measured declines of 9× to 900× a year depending on the benchmark. If you only read the price lists, AI should be getting cheap.
It isn't, at least not in total. Google now processes over 3.2 quadrillion tokens a month, OpenRouter's volume has grown about 30,000-fold since 2023, Anthropic and OpenAI together report well over $100B in annualised revenue, and companies have started capping what each engineer can spend. We pulled the public numbers on volume, prices, total spend and per-developer cost, and the measured effect of the levers that cut a bill, all as of late September 2026. Vendor figures are the vendor's own claims and are labelled that way. Every number links to its source.
Key findings
- 1Tokens got cheaper and we bought far more of them. The same capability costs about 10× less each year, while Google's monthly tokens rose from 9.7 trillion to over 3.2 quadrillion in two years, roughly 330×.
- 2OpenRouter, the one neutral multi-model gauge, grew about 30,000× to a 4.5 quadrillion-token annual run rate in August 2026, doubling every 11 weeks. A natural experiment there shows the mechanism: when OpenAI discounted GPT-5.6 Luna, its daily tokens rose 13.8×, and demand stayed after the discount ended.
- 3The money followed. Menlo Ventures puts enterprise spend on model APIs at $12.5B in 2025, up from $3.5B in its November 2024 estimate. Anthropic says its run rate went from about $1B at the start of 2025 to $65B at the end of July 2026, and OpenAI's ARR was nearing $70B by late September.
- 4Agents are the multiplier. Anthropic says an agent uses about 4× the tokens of a chat and a multi-agent system about 15×. Agentic coding uses about 1,000× the tokens of code chat, and almost all of it is input: 99.5% of the tokens in Claude Code's own sample session.
- 5What a developer costs doubled in a year. Anthropic's guidance for Claude Code went from $6 per developer per day to $13 per active day. Jellyfish puts the median developer at $52 a month and the 90th percentile at $691, and Uber capped spend at $1,500 per employee per tool after using its annual AI budget in four months.
- 6The frontier price came back up. After falling about 11× from GPT-4 to GPT-5, a new top tier at $10 input and $50 output per million tokens arrived with Claude Fable 5 and GPT-6 Astra. The most expensive model to run Artificial Analysis's index went from $1.78 a task in June to $7.63 in September.
- 7The levers are large and measured. In Anthropic's tests prompt caching cut agent-loop cost 2.7× to 5.3×. In our worked example, caching, batch and a smaller model take one agent session from $2.20 to $0.16.
3.2 quadrillion
Tokens a month across Google's products and APIs, May 2026 (9.7 trillion two years earlier)
Source: Google I/O 2026
~30,000×
OpenRouter token growth since Aug 2023, to a 4.5 quadrillion a year run rate
Aug 2026
Source: Menlo Ventures
$13
Claude Code's average cost per developer per active day in 2026, up from $6 a day
Anthropic's guidance
Source: Claude Code docs
Same capability, a tenth of the price each year#
The long-run trend is not in dispute. In November 2021, a model good enough to score 42 on the MMLU benchmark (GPT-3) cost $60 per million tokens. Three years later Llama 3.2 3B did it for $0.06, a thousandfold drop that a16z's Guido Appenzeller summed up as "LLMflation": for equivalent performance, "the cost is decreasing by 10x every year". Epoch AI's analysis finds the same direction with a wide range, 9× to 900× a year depending on which score you hold fixed.
Flagship list prices show it too, up to a point. OpenAI's flagship went from $37.50 per million tokens (GPT-4, March 2023, blending input and output 3:1) to $3.44 (GPT-5, August 2025), about 11× cheaper in 29 months. Anthropic's top model fell from $30 (Claude 3 Opus) to $10 (Opus 4.5) in November 2025. Then both lines turned up.
Price of each lab's top model
Blended list price per 1M tokens (3 input : 1 output), US dollars, Mar 2023 to Sep 2026
| Series | x | y |
|---|---|---|
| OpenAI flagship | Mar 2023 | $37.5 (GPT-4) |
| OpenAI flagship | Nov 2023 | $15 |
| OpenAI flagship | May 2024 | $7.5 |
| OpenAI flagship | Aug 2024 | $4.38 |
| OpenAI flagship | Apr 2025 | $3.5 |
| OpenAI flagship | Aug 2025 | $3.44 (GPT-5) |
| OpenAI flagship | Sep 2026 | $20 (GPT-6 Astra) |
| Anthropic top model | Mar 2024 | $30 (Claude 3 Opus) |
| Anthropic top model | Nov 2025 | $10 (Opus 4.5) |
| Anthropic top model | Jun 2026 | $20 (Fable 5) |
| Anthropic top model | Sep 2026 | $20 |
Sources: OpenAI; Anthropic; Claude pricing; host0 blended-price arithmetic
The labs now sell two tiers. At the top, Claude Fable 5 (June 2026) and GPT-6 Astra (September) both list at $10 per million input tokens and $50 per million output. Below them sits a workhorse tier that keeps getting cheaper: Claude Opus 5.5 at $4 and $20, 20% below Opus 5, and GPT-6.1 Sol at $2 and $10, which TechCrunch reports OpenAI pitches as near-Astra intelligence at one-fifth of the price. Google announced Gemini 4 Argon on 30 September at an introductory $2 and $10, rising to $4 and $20.
The price card at the end of September 2026
Blended list price per 1M tokens (3 input : 1 output), US dollars, models available by 30 Sep 2026
| Label | Value |
|---|---|
| GPT-6 Astra | $20 ($10 in / $50 out) |
| Claude Fable 5.1 | $20 ($10 in / $50 out) |
| Claude Opus 5.5 | $8 ($4 / $20) |
| GPT-6.1 Sol | $4 ($2 / $10) |
| Claude Sonnet 5.5 | $4 ($2 / $10) |
| Claude Haiku 4.5 | $2 ($1 / $5) |
| Gemini 3.8 Flash | $1.50 ($0.75 / $3.75, introductory) |
| DeepSeek V4.1-Flash | $0.53 ($0.30 / $1.20, peak hours) |
| GPT-6 Luna | $0.20 ($0.10 / $0.50) |
Sources: OpenAI pricing; Claude pricing; Gemini pricing; DeepSeek pricing
What companies actually pay sits far below list. Ramp, which tracks what businesses on its platform spend, measured an effective price of $0.68 per million tokens in early September, down 41% from a 2026 peak of $1.15 in March, as customers moved toward cheaper standard models: frontier models (Opus, Fable, Sol) were 45% of token share, down from 53% in August. One more quiet change cuts the other way. Anthropic's tokenizer since Opus 4.7 produces about 30% more tokens for the same text, so a lower price per token isn't always a lower price per page.
Tokens grew faster than prices fell#
If prices fall 10× a year and usage grows less than that, spend shrinks. It didn't. At Google I/O in May 2026, Sundar Pichai said Google was processing 9.7 trillion tokens a month across its products two years earlier, about 480 trillion a year earlier, and "over 3.2 quadrillion" now. That's roughly 50× in the first year and 7× in the second.
Tokens Google processes a month
Across Google's products and APIs, quadrillions of tokens per month (1 quadrillion = 1,000 trillion), May 2024 to May 2026
| x | y |
|---|---|
| May 2024 | 0.0097 (9.7 trillion) |
| May 2025 | 0.48 |
| Jul 2025 | 0.98 |
| Oct 2025 | 1.3 |
| May 2026 | 3.2 (3.2 quadrillion) |
Sources: Google I/O 2026; Google I/O 2025; Alphabet Q2 2025; Alphabet Q3 2025
Developer traffic alone tells the same story. Google's API went from 7 billion tokens a minute in October 2025 to about 22 billion in July 2026, which is roughly 0.95 quadrillion a month by our arithmetic. OpenAI's API was at 6 billion a minute at DevDay 2025 and more than 15 billion by March 2026, its latest public figure. Both companies say they are short of capacity: Google said in July it continues "to be supply constrained", and Microsoft expects to stay constrained at least through 2026.
API tokens per minute, Google and OpenAI
Tokens per minute through each company's developer APIs, billions, Oct 2025 to Jul 2026
| Series | x | y |
|---|---|---|
| 29 Oct 2025 | 7B | |
| 4 Feb 2026 | 10B | |
| 29 Apr 2026 | 16B | |
| 19 May 2026 | 19B | |
| 22 Jul 2026 | 22B | |
| OpenAI | 6 Oct 2025 | 6B |
| OpenAI | 31 Mar 2026 | 15B |
Sources: Alphabet earnings remarks; Google I/O 2026; TechCrunch (DevDay 2025); OpenAI
OpenRouter is the closest thing to a neutral meter, since it routes traffic to 400+ models from many labs. When it rebranded in August 2023 it handled about 3 billion tokens a week. In May 2026 it reported 25 trillion a week, and when Stripe agreed to buy it in August, its investor Menlo Ventures put it at "~30,000x to a 4.5+ quadrillion-token annual run rate", doubling every 11 weeks. Our own conversion of those figures (86.5 trillion a week against 3 billion) gives about 28,800×, so the round number holds.
OpenRouter's annual token run rate
Tokens a year at each date's pace, Aug 2023 to Aug 2026
| Label | Value |
|---|---|
| Aug 2023 | ~0.16T (3B a week × 52) |
| Mid-2024 | ~10T |
| Jun 2025 | 100T+ |
| Nov 2025 | ~260T (5T a week × 52) |
| May 2026 | 1.5 quadrillion |
| Aug 2026 | 4.5 quadrillion+ |
Sources: Menlo Ventures (Aug 2026); Menlo Ventures (May 2026); Menlo Ventures (Jun 2025); Business Wire
China reports in tokens a day, and its numbers are as steep. ByteDance's Doubao went from 4 trillion a day in December 2024 to 180 trillion in June 2026. China's data administration said national consumption rose from 100 billion tokens a day at the start of 2024 to more than 140 trillion in March 2026, over a thousandfold.
Tokens a day in China
Average daily tokens, trillions, ByteDance's Doubao and China's national total, Jan 2024 to Jun 2026
| Series | x | y |
|---|---|---|
| Doubao | Dec 2024 | 4T |
| Doubao | Mar 2025 | 12.7T |
| Doubao | Sep 2025 | 30T |
| Doubao | Dec 2025 | 50T |
| Doubao | Apr 2026 | 120T |
| Doubao | Jun 2026 | 180T |
| China, national | Jan 2024 | 0.1T |
| China, national | Dec 2025 | 100T |
| China, national | Mar 2026 | 140T |
Sources: SCMP; China Daily; TMTPost; SCIO; CEIBS
A price cut, measured#
Economists call this the Jevons paradox: make a resource cheaper to use and total use rises enough to swamp the saving. Satya Nadella invoked it during the DeepSeek week in January 2025 (NPR). This summer OpenRouter got to test it. OpenAI ran a discount on two GPT-5.6 models and then cut their list prices, for an effective discount of 90% on Luna and 60% on Terra, while GPT-5.6 Sol stayed at full price. OpenRouter's analysis found daily Luna tokens rose 13.8× and Terra 5.6×, against 1.11× for Sol. OpenAI's share of all OpenRouter tokens went from 7.1% to 12.4%, and the period after the programme saw 1.38× the tokens of the discount period itself.
What a price cut did to demand
Rise in daily tokens on OpenRouter during OpenAI's GPT-5.6 discount, multiple of the period before, summer 2026
| Label | Value |
|---|---|
| GPT-5.6 Luna | 13.8× (90% effective discount) |
| GPT-5.6 Terra | 5.6× (60% effective discount) |
| GPT-5.6 Sol | 1.11× (no discount (control)) |
The supply side compounds it. Microsoft said in July 2025 that software optimisation alone was delivering 90% more tokens per GPU than a year earlier. Cheaper tokens to produce, cheaper tokens to buy, and every cut so far has been met with more use.
Where the money went#
Menlo Ventures, which surveys enterprise buyers, estimated that companies spent $3.5B on model APIs in November 2024 and $8.4B by mid-2025. Its full-year figure for 2025 was $12.5B for foundation model APIs, out of $37B of enterprise generative AI spend, itself up 3.2× from 2024. Gartner, with a different method, forecasts spend on foundation models of $23.4B in 2026, double its 2025 estimate.
Enterprise spend on model APIs
US dollars; Menlo Ventures estimates and a Gartner forecast, Nov 2024 to 2026
| Label | Value |
|---|---|
| Nov 2024 | $3.5B (Menlo) |
| Mid-2025 | $8.4B (Menlo) |
| 2025 | $12.5B (Menlo, full year) |
| 2026 | $23.4B (Gartner forecast, different method) |
Sources: Menlo Ventures (mid-2025); Menlo Ventures (2025); Gartner via CXOToday
Most of that went to three labs. By Menlo's estimate, Anthropic took 40% of enterprise LLM API spend in 2025, OpenAI 27% and Google 21%, and in coding Anthropic led OpenAI 54% to 21%. Ramp's card data points the same way in 2026: by August, 43.8% of US businesses paid Anthropic for subscriptions or tokens, against 39.8% for OpenAI.
The labs' own revenue figures are bigger than any API-spend estimate, because they include subscriptions, consumer apps and coding tools. Anthropic says its run rate went from $87M at the start of 2024 to about $1B at the start of 2025, $14B in February 2026 and over $30B in April; CNBC reported $65B at the end of July. OpenAI's CFO put its annualised revenue at over $20B for 2025, and Axios reported its ARR nearing $70B on 29 September.
Annualised revenue, Anthropic and OpenAI
Run-rate or annual recurring revenue as reported, US dollars, Dec 2023 to Sep 2026
| Series | x | y |
|---|---|---|
| Anthropic | Jan 2024 | $87M |
| Anthropic | Jan 2025 | $1B |
| Anthropic | Aug 2025 | $5B |
| Anthropic | Dec 2025 | $9B |
| Anthropic | Feb 2026 | $14B |
| Anthropic | Apr 2026 | $30B |
| Anthropic | May 2026 | $47B |
| Anthropic | Jul 2026 | $65B |
| OpenAI | Dec 2023 | $2B |
| OpenAI | Dec 2024 | $6B |
| OpenAI | Dec 2025 | $20B |
| OpenAI | Aug 2026 | $40B |
| OpenAI | Sep 2026 | $70B (~$70B) |
Selling tokens is also expensive to do. Anthropic's IPO prospectus, as reported by Reuters, shows $7.33B of compute and infrastructure cost in 2025, more than half its operating expenses, and an operating loss of over $8B on about $4.6B of revenue. Two customers made up nearly a quarter of that revenue.
Agents are why#
A chat is a question and an answer. An agent is a loop: it reads files, calls tools, reads the results, and sends the whole growing conversation back to the model on every turn. Anthropic's engineers put a number on it in 2025: agents use about 4× the tokens of a chat, multi-agent systems about 15×, and token usage alone explained 80% of the performance differences on their browsing benchmark. A Stanford-affiliated study of eight models on SWE-bench found agentic coding uses about 1,000× the tokens of code chat, "with input tokens rather than output tokens driving the overall cost", and that runs of the same task differ by up to 30× without more tokens buying more accuracy.
How many more tokens agents use
Multiples reported by each source, against the baseline each one names, 2025 to 2026
Sources: Anthropic; Claude Code docs; Bai et al.; OpenRouter State of AI
It shows up in the market totals. On OpenRouter, programming went from about 11% of tokens in early 2025 to over half by the end of the year. OpenRouter suggests 6 February 2026 may have been the last day humans used more tokens than agents; from then to August, agent token volume grew about 14×, from 0.51 trillion to 7.3 trillion, while human volume grew 2.8×. Google says its own internal developer tools went from half a trillion tokens a day in March 2026 to more than three trillion in May.
Where an agent session's tokens go#
Claude Code's documentation includes a sample session that makes the shape concrete. On Sonnet 4.6, it used 1.2k input tokens, 5.3k output, 940k cache reads and 50k cache writes, for $0.55. Almost everything the model read was its own context, re-sent each turn and served from cache.
One Claude Code session, token by token
The sample session in Anthropic's Claude Code docs, Sonnet 4.6 list prices, US dollars
Sources: Claude Code docs, Sep 2026 snapshot; Claude pricing; host0 arithmetic
That's why the agent era hits input pricing and caching hardest. Manus, which builds a general agent, reported an input-to-output ratio of about 100:1 and called the cache hit rate "the single most important metric" for a production agent. The Stanford-affiliated SWE-bench study found that cache reads dominate both token volume and dollar cost in every phase of a run.
What a developer costs now#
Anthropic's own guidance tracks the shift. Through early 2026, the Claude Code docs said the average cost was $6 per developer per day, with 90% of users under $12. In April 2026 that changed to "around $13 per developer per active day and $150-250 per developer per month, with costs remaining below $30 per active day for 90% of users" across enterprise deployments. The average rose 2.2× and the 90th percentile 2.5×, while list prices per token were falling.
Independent data puts numbers on the spread. Jellyfish, which measures engineering teams, looked at 12,000 developers at 200 companies in the first quarter of 2026. The median developer used about 51 million tokens a month and spent $52.38; at the 90th percentile it was about 380 million tokens and $691.14. Its mid-year report says per-developer consumption rose about 18.6× in nine months, with a mean of $248 a month and a 99th percentile of $2,452.
What a developer spends on tokens a month
US dollars per developer per month, by percentile and by guideline, 2026
| Label | Value |
|---|---|
| Jellyfish median | $52.38 (Q1 2026) |
| Claude Code average (Anthropic) | $150–250 (enterprise deployments) |
| Jellyfish 75th percentile | $226.58 (Q1 2026) |
| Jellyfish mean | $248 (H1 2026 report) |
| Jellyfish 90th percentile | $691.14 (Q1 2026) |
| Uber's cap per tool | $1,500 (per employee, from Jun 2026) |
| Jellyfish 99th percentile | $2,452 (H1 2026 report) |
Sources: Jellyfish (Apr 2026); Jellyfish (H1 2026); Claude Code docs; TechCrunch
More tokens bought more output, but at a falling rate. In Jellyfish's data the heaviest users merged about twice as many pull requests for about ten times the tokens, and the cost per merged PR rose from $0.28 in the lowest usage tier to $89.32 in the highest.
Companies noticed. Uber had told staff to use AI as much as possible and ranked usage on internal leaderboards, then, after it blew through its annual AI budget in four months, set a monthly cap of $1,500 per employee per agentic coding tool. The CEO of Faros AI told TechCrunch one of his engineers spent $40,000 on tokens in a month, and Priceline's Cursor renewal came back 4–5× more expensive. Meta's Adam Mosseri said the burn rate of a strong engineer might equal their salary. Gartner predicts AI coding costs will exceed the average developer's salary by 2028, and says 23% of tech leaders already spend $200 to $500 per developer per month on tokens.
Pricing moved to match. In July 2025 Anthropic said one Claude Max subscriber had consumed "tens of thousands" of dollars of model usage on a $200 plan and introduced weekly limits. Cursor switched its Pro plan to $20 of included usage and refunded surprised users. And on 1 June 2026 GitHub moved every Copilot plan to usage-based billing, because under a flat per-request price "a quick chat question and a multi-hour autonomous coding session can cost the user the same amount".
The frontier got expensive again#
Two things push the cost of a hard task up even as tokens get cheaper: the new premium tier, and models that think longer at maximum effort. Artificial Analysis prices a full run of its Intelligence Index for each model. In June the most expensive model was Claude Opus 4.8 at $1.78 per task. In September, Claude Fable 5.1 at maximum effort cost $7.63 per task, using 78K output tokens, and GPT-6 Astra reached the same score for $3.26 with 27K. The cheap end barely moved: DeepSeek V4 Pro ran at $0.04 per task.
Cost to run one task of the Artificial Analysis index
US dollars per task at each model's maximum effort, Intelligence Index v4.1, Jun and Sep 2026
| Label | Value |
|---|---|
| Claude Fable 5.1 (Sep) | $7.63 (78K output tokens a task) |
| GPT-6 Astra (Sep) | $3.26 (27K output tokens a task) |
| Claude Opus 4.8 (Jun) | $1.78 (most expensive in June) |
| DeepSeek V4 Pro (Jun) | $0.04 |
Sources: Artificial Analysis (Jun 2026); Artificial Analysis (Sep 2026)
For a fixed task, though, the cost of getting it done keeps falling. In 2024 the SWE-agent paper reported GPT-4 Turbo resolving 12.47% of SWE-bench issues at $1.59 per resolved issue. On the official SWE-bench Verified leaderboard with the same simple agent, Claude 4.6 Opus resolves 75.6% for about $0.73 per resolved issue by our arithmetic, and several open-weight models get within a few points for a few cents.
Cost per resolved coding issue
US dollars per resolved SWE-bench Verified issue with mini-SWE-agent (total cost ÷ issues solved), share solved in notes, Jul 2025 to Feb 2026
| Label | Value |
|---|---|
| GPT-4o | $7.08 (21.6% solved) |
| Claude 4 Opus | $1.67 (67.6%) |
| Gemini 3 Pro | $1.38 (69.6%) |
| Claude 4.5 Opus (high) | $0.98 (76.8%) |
| Claude 4.6 Opus | $0.73 (75.6%) |
| GPT-5.2 (high) | $0.65 (72.8%) |
| Gemini 3 Flash (high) | $0.47 (75.8%) |
| Kimi K2.5 (high) | $0.21 (70.8%) |
| MiniMax M2.5 (high) | $0.10 (75.8%) |
| DeepSeek V3.2 Reasoner | $0.05 (60.0%) |
Sources: SWE-bench leaderboard data; host0 arithmetic
Labs now market this directly. Anthropic says Opus 5.5 will cost 40% less than Opus 5 on typical workloads at default settings, partly from a lower price and partly from using fewer tokens. Its own cost guide shows the same effect can run backwards: Fable 5.1 matched Fable 5's coding score for 43% less per solved task, but on a long research benchmark the same upgrade cost 41% more per task, because the newer model did more work.
The levers that cut the bill#
The best public evidence on what works is Anthropic's guide to optimizing for cost and intelligence, which reports measured runs for each lever (the numbers below are from its September version). Its first finding is about why agents get expensive: "A 40-turn task sends its first turn 40 times, so task cost grows with roughly the square of turn count."
Caching comes first. A cache read costs a tenth of the input price on Anthropic, OpenAI's newest models and Google's, and 2% at DeepSeek. Across a day of real traffic, Anthropic says agent loops read a median 84% of their input from cache, and the top tenth 94% or more. In its measurements caching cut agent-loop cost by a factor of 2.7 to 5.3, and a small triage agent's bill by 83%. On a deep research benchmark, Fable 5.1 fell from $37.94 to $7.12 per task with caching. Breaking the cache mid-session costs real money too: changing the effort setting and adding a tool took one triage session from $0.81 to $0.95.
Batch is the second free lever. Anthropic, OpenAI and Google all take 50% off for work that can wait up to 24 hours, and it stacks with caching. DeepSeek went further and made off-peak hours half price. Speed now goes the other way: OpenAI's Fast mode and Anthropic's fast mode on Opus 5.5 cost twice the standard price.
Then trim the context. Tool definitions are input on every turn. When Anthropic grew a test agent's catalog to 502 tools, run cost nearly doubled from $0.55 to $1.02; with tool search it stayed at $0.56, with the same accuracy. Cursor's on-demand tool loading cut total agent tokens by 46.9% in runs that used an MCP tool, and GitHub's token work cut one agentic workflow's effective tokens by 62%. Data matters as much as tools: a roughly 91,000-token table pasted into every request let Sonnet 5 answer 6 of 25 questions for $5.01, while uploading the file and letting the model run code answered all 25 for $0.40.
What each lever saved, as measured
Reduction in cost or tokens on each source's own workload, 2026
| Label | Value |
|---|---|
| Upload data, use code execution | 92% ($5.01 → $0.40, Anthropic) |
| Prompt caching, triage agent | 83% (Anthropic) |
| Token work, GitHub Auto-Triage | 62% (effective tokens, 109 runs) |
| On-demand MCP tools, Cursor | 46.9% (total agent tokens) |
| Tool search at 502 tools | 45% (Anthropic) |
| Prune stale tool results | 39% (long run, Anthropic) |
| Compaction | 32% (long run, Anthropic) |
| Context editing | 0% (long run; +74% on a short one) |
Sources: Anthropic cost guide (Sep 2026); GitHub; Cursor
Choose the model by cost per solved task, not per token. On a coding benchmark, Anthropic found Fable 5.1 at low effort solved 88.6% of tasks for $0.54 each, against 77.4% at $0.84 for Sonnet 5: "11 more points for 35% less per solved task, despite a per-token price five times higher." Running everything at low effort and re-running only the failures held the pass rate at about half the cost. Splitting work between an orchestrator and cheap worker models saved time but saved money in only two of the situations it measured. Routing research points the same way: LMSYS's RouteLLM cut cost by over 85% on one benchmark while keeping 95% of GPT-4's quality, by sending only hard queries to the big model.
A worked example: one agent session, priced#
To put the levers side by side, take one agent session with a million input tokens and 20,000 output tokens, a 50:1 ratio that is if anything light on input next to Manus's 100:1. In the cached case, 90% of the input is read from cache and 10% written to it. Prices are each vendor's list prices at the end of September 2026.
One agent session, lever by lever
Cost of 1M input + 20K output tokens on Claude Sonnet 5, then Claude Haiku 4.5, US dollars, list prices, Sep 2026
| Label | Value |
|---|---|
| Sonnet 5, no cache | $2.20 |
| Sonnet 5, 90% cached | $0.63 |
| Sonnet 5, cached + batch | $0.32 |
| Haiku 4.5, cached + batch | $0.16 |
Sources: Claude pricing; host0 arithmetic
The same session on nine models
Cost of 1M input + 20K output tokens, US dollars, list prices at the end of September 2026
Sources: Claude pricing; OpenAI pricing; Gemini pricing; DeepSeek pricing; host0 arithmetic
The formula is short enough to run on your own numbers: without caching, cost = input tokens × input price + output tokens × output price; with caching, replace the input term with 90% at the cache-read price plus 10% at the cache-write price. Sonnet 5 cached, for example, is 0.9 × $0.20 + 0.1 × $2.50 + 0.02 × $10 = $0.63.
What this means if you build with AI#
Cheaper tokens won't make your bill smaller on their own. The habits that do, from the data above:
- Turn on caching before anything else, and keep the prefix stable. It's the largest measured lever by a wide margin. Put the system prompt and tools first, append rather than edit, and don't change models, effort or tools mid-session.
- Watch your cache hit rate, not your token count. A median agent loop reads 84% of its input from cache. If yours is far below that, something is breaking the prefix.
- Batch anything nobody is waiting for. Evals, backfills and scheduled jobs get 50% off on top of caching. Don't pay for fast modes unless latency is the product.
- Load tools and data on demand. Use tool search for large catalogs, ship small default toolsets, and give the model files and code execution instead of pasting tables into the prompt, which in Anthropic's test was about 12× cheaper and more accurate.
- Measure cost per solved task. A model that costs five times more per token can be 35% cheaper per solved task. Start at low effort, re-run failures at higher effort, and send simple subagent work to a small model.
- Budget per developer, with a ceiling. Anthropic's own planning number is $150–250 a month, but the tail is long: Jellyfish's 99th percentile spends $2,452 a month. Set alerts and caps before the bill does it for you, as Uber learned.
- Don't put a model where a lookup will do. Every request that runs through a model in production is a recurring token cost. If an app only needs to show and store data, serve it without one.
That last point is where host0 fits: it's a cloud for small software, where an agent builds an app and host0 gives it a URL and a shared database. A static frontend on host0 with platform records answers every visit without calling a model, so the only token bill is the one you paid to build it.
Methodology#
This post draws on four research passes, one per angle: token volume, prices and total spend, agents and per-developer cost, and savings levers. Each was limited to primary sources: earnings remarks and company posts, pricing pages and changelogs, investor reports, research papers and the archived documentation of Claude Code and the Claude API. Statistics and aggregator sites were used only as leads. The post is dated 1 October 2026 and uses nothing published after that date. Where a page might have changed since, we cite a Wayback Machine snapshot from before it (the Claude Code costs page from 28 September 2026 and Anthropic's cost guide from 15 September). Before publishing we re-opened every single-source page the post leans on for a headline number and confirmed the quoted sentence was still there; one wording, Anthropic's "40% less" for Opus 5.5, was corrected to match the page.
Several numbers are our own arithmetic:
- Blended prices weight input and output 3:1, Epoch AI's convention: (3 × input + output) ÷ 4.
- OpenRouter run rates convert weekly figures to a year by multiplying by 52 (two bars), and back for the ~30,000× check.
- Google's API tokens per month multiply tokens per minute by 43,200 minutes in a 30-day month.
- The session cost and the worked example use each vendor's list prices: input, output, cache read and cache write per million tokens.
- Cost per resolved SWE-bench issue divides each submission's average cost per task by its share of tasks solved.
Limitations:
- Companies count tokens differently. Google's monthly total covers all its products, its per-minute figure only the API; OpenAI's covers only the API; China reports daily averages. Never add them up.
- Run rates aren't revenue. They annualise the latest month, and most 2026 figures for Anthropic and OpenAI are press reports of investor updates, not filings.
- Market sizes depend on method. Menlo changed its method between its 2024 and 2025 estimates, and Gartner uses another. We chart them with labels, not as one series.
- Vendor savings are measured on the vendor's workloads. Anthropic's cost guide, Cursor's and GitHub's results are their own tests; yours will differ.
- Per-developer data is a sample. Jellyfish covers its own customers; Anthropic's guidance covers enterprise deployments of Claude Code.
Open questions#
- How many tokens Anthropic serves. It publishes revenue and, via its prospectus, compute cost, but no token volume.
- OpenAI's API volume since spring. Its last public figure is "more than 15 billion" tokens a minute in March and April 2026.
- How revenue splits between API and subscriptions at Anthropic and OpenAI. Claude Code's last official run rate was over $2.5B in February 2026.
- Real cache hit rates across the industry. Anthropic's 84% median is the only published aggregate.
- What a typical agent session costs in the wild. Beyond Claude Code's sample session and Jellyfish's monthly percentiles, nobody publishes per-session token counts for coding agents.
- How much the new tokenizers raised bills. Anthropic says its tokenizer since Opus 4.7 makes about 30% more tokens for the same text, but not how that varies between code and prose.
- A 2026 enterprise API spend estimate on the same method as 2025's $12.5B. Gartner's forecast is the only 2026 figure, and it measures something different.
