AIFEATURE

AI Inference Prices Are Falling About 10x a Year — What the Token-Cost Data Actually Shows

N
NathanTechnology Editor · Technical Lead
Published · Updated
Equivalent-capability AI inference prices have fallen roughly 10x per year across multiple benchmarks, with OpenAI's GPT-4o mini priced about 99% below its 2022 predecessor and DeepSeek-R1 undercutting OpenAI's o1 by roughly 27x. Yet Google now processes about 50 times more tokens monthly than a year earlier, and Gartner projects 2025 global generative-AI spending at $644 billion, up 76.4% — proof that cheaper tokens are fueling more total usage and spend, not less.

How fast is inference cost actually falling — and do different benchmarks agree?

Multiple benchmarks agree that equivalent-capability inference pricing has fallen at a pace close to 10x per year, though the exact multiple depends on which vendor, capability tier, and time window is measured. Venture firm a16z tracked the cost of reaching GPT-3-level capability and found it fell from $60 per million tokens at GPT-3's late-2021 debut to about $0.06 per million tokens by the end of 2024 — a roughly 1,000x decline over three years, which a16z calls "LLMflation" and translates to an underlying pace of about 10x per year for a fixed capability levelCITE:E1. Using a different yardstick — the lowest price across vendors for GPT-4-level capability (MMLU scores of 83 or above) — a16z found only a 62x decline between GPT-4's March 2023 launch and the end of 2024CITE:E2. OpenAI's own account of its single product line tells a third story: CEO Sam Altman stated that the cost of a fixed capability level falls roughly 10x every 12 months, and cited an approximately 150x drop in per-token price from GPT-4's early-2023 release to GPT-4o by mid-2024CITE:E3.

Comparison basisTime spanPrice declineSource
Same-capability tier, GPT-3 level (a16z "LLMflation")Late 2021 → end 2024 (~3 years)~1,000x ($60 → $0.06 per million tokens)CITE:E1
GPT-4-level capability, lowest price across vendorsMarch 2023 → end 2024~62xCITE:E2
OpenAI's own line, GPT-4 → GPT-4oEarly 2023 → mid-2024~150xCITE:E3

These three figures are not interchangeable: a16z's 1,000x and 10x-per-year figures describe a cross-market, same-capability benchmark over three years; its 62x figure is a narrower cross-vendor lowest-price comparison over roughly 21 months; and Altman's 150x figure describes only OpenAI's internal product progression over about 17 months. All three point in the same direction — steep, rapid decline — but the size of the number depends entirely on which slice of the market is being measured.

How is market pricing being pushed down further by competition?

Both incumbent repricing and low-cost new entrants are compounding the decline documented above. OpenAI itself supplied evidence of active repricing: its GPT-4o mini, launched in July 2024, is priced at $0.15 per million input tokens and $0.60 per million output tokens, which OpenAI states is more than 60% cheaper than GPT-3.5 Turbo and about 99% cheaper per token than its 2022 text-davinci-003 modelCITE:E4. Separately, DeepSeek's R1 model is priced at roughly $0.55 per million input tokens and $2.19 per million output tokens — about 3.6% of OpenAI o1's $15 input / $60 output pricing, or roughly 27x cheaperCITE:E5.

ModelInput price (per million tokens)Output price (per million tokens)Note
GPT-4o mini$0.15$0.60>60% cheaper than GPT-3.5 Turbo; ~99% cheaper than 2022 text-davinci-003CITE:E4
DeepSeek-R1$0.55$2.19~3.6% of OpenAI o1's price, ~27x cheaperCITE:E5
OpenAI o1 (reference point)$15$60Baseline used in the DeepSeek-R1 comparisonCITE:E5

Read together, these two data points show pricing pressure coming from two directions at once: OpenAI cutting its own prices well below its earlier models, and a lower-cost entrant like DeepSeek pricing an entire order of magnitude below an established closed-source model.

Why does falling per-token cost coincide with rising total inference spending?

Falling unit prices are coinciding with a sharp rise in total usage and total spending, not a decline in either. Google stated at its May 2025 I/O keynote that its products now process 480 trillion tokens per month, which it described as about 50 times the volume of a year earlierCITE:E6. At the same time, Gartner forecast that global generative-AI spending would reach $644 billion in 2025, up 76.4% year over year, while noting that roughly 80% of that spending is hardware — devices and servers — rather than API inference billsCITE:E7.

MetricValueSource
Google monthly token processing (May 2025)480 trillion tokens, ~50x year-over-yearCITE:E6
2025 global generative-AI spending (Gartner forecast)$644 billion, +76.4% year-over-yearCITE:E7
Share of that spending that is hardware (devices/servers)~80%CITE:E7

The two figures describe different parts of the same system: Google's number tracks token volume actually processed, while Gartner's number tracks total dollars spent across the generative-AI stack, most of which is capital equipment rather than usage-based API charges.

What this means

The evidence assembled here shows two trends running in parallel rather than in opposition. Per-token prices for equivalent capability have fallen by roughly 10x a year on a20 multi-year viewCITE:E1, further compressed by both incumbent repricingCITE:E4 and low-cost entrantsCITE:E5. At the same time, Google's reported 50x jump in monthly token volumeCITE:E6 and Gartner's 76.4% spending growth to $644 billionCITE:E7 show that cheaper tokens have not reduced the amount of money moving through the generative-AI market — they have coincided with more of it. Because Gartner's total is roughly 80% hardware rather than inference API chargesCITE:E7, the token-price collapse and the spending increase are not contradictory: one measures the cost of a unit of intelligence, the other measures how much capacity and usage the market is now buying at that lower price.

📊 Evidence

FAQ

How fast is inference cost actually falling — and do different benchmarks agree?

Multiple benchmarks agree that equivalent-capability inference pricing has fallen at a pace close to 10x per year, though the exact multiple depends on which ve…

How is market pricing being pushed down further by competition?

Both incumbent repricing and low-cost new entrants are compounding the decline documented above.

Why does falling per-token cost coincide with rising total inference spending?

Falling unit prices are coinciding with a sharp rise in total usage and total spending, not a decline in either.

📎 Sources

  1. a16z.com
  2. blog.samaltman.com
  3. openai.com
  4. artificialanalysis.ai
  5. blog.google
  6. gartner.com

Related data

Author's TakeNathan

The '10x a year' headline hides three measurements this data forces apart: a16z's 1,000x-in-three-years figure is a best-case, same-capability benchmark, its own 62x figure is a narrower cross-vendor comparison, and OpenAI's 150x figure covers only its own product line over a shorter window — treating them as one number overstates how fast any single buyer's bill falls. The more consequential pairing is Google's 50x token-volume growth against Gartner's 76.4% spending increase: falling unit price is not shrinking aggregate cost, it is funding more inference. Because Gartner's $644 billion figure is roughly 80% hardware rather than API spend, the metric worth watching next is whether that hardware share compresses as token prices keep falling — that would signal the market shifting from capacity build-out toward pure usage-based economics.

N
NathanTechnology Editor · Technical Lead

Related

FEATURE

China's AI Playbook Under Export Controls: Ahead on Models, a Generation Behind on Chips

China's AI industry shows a split scorecard: on models, DeepSeek-V3's open-weight, MIT-licensed mixture-of-experts architecture and DeepSeek-R1's benchmark parity with OpenAI's o1 preceded a $593 billion single-day NVIDIA sell-off, while the widely cited $5.576 million cost figure covers only the final training run. On hardware, Huawei's Ascend 910C still delivers about 60% of NVIDIA H100's per-chip performance, and HBM supply constraints cap actual 2025 shipments far below estimated capacity, even as US export controls kept tightening through January 2026.

Nathan ·
FEATURE

Why Are Tech Giants Racing to Buy Nuclear Power? Can SMRs Solve AI Data Centers' Electricity Crisis?

Facing a projected jump in global data center electricity demand from about 415 TWh in 2024 to about 945 TWh in 2030, Microsoft, Google, Amazon, and Meta have all signed nuclear power deals — but most rely on small modular reactors (SMRs) that won't deliver meaningful capacity until 2030 or later, and NuScale's 2023 project cancellation shows the economics are still unproven.

Nathan ·
FEATURE

Private Credit's Climb to $2.5 Trillion: Why AI Data Centers Are Turning to Non-Bank Lenders

Private credit has grown into a global asset class exceeding $2.5 trillion, driven by banks retreating from riskier lending and investor returns climbing toward 12%. A handful of firms — Apollo, Blackstone, Ares, KKR, and Blue Owl — now dominate the market, and Blue Owl has extended it into AI infrastructure with a $27 billion data center financing for Meta, even as the IMF and the Federal Reserve disagree on how much systemic risk this concentration carries.

林紀旭 James Lin ·