Equivalent-capability AI inference prices have fallen roughly 10x per year across multiple benchmarks, with OpenAI's GPT-4o mini priced about 99% below its 2022 predecessor and DeepSeek-R1 undercutting OpenAI's o1 by roughly 27x. Yet Google now processes about 50 times more tokens monthly than a year earlier, and Gartner projects 2025 global generative-AI spending at $644 billion, up 76.4% — proof that cheaper tokens are fueling more total usage and spend, not less.
How fast is inference cost actually falling — and do different benchmarks agree?
Multiple benchmarks agree that equivalent-capability inference pricing has fallen at a pace close to 10x per year, though the exact multiple depends on which vendor, capability tier, and time window is measured. Venture firm a16z tracked the cost of reaching GPT-3-level capability and found it fell from $60 per million tokens at GPT-3's late-2021 debut to about $0.06 per million tokens by the end of 2024 — a roughly 1,000x decline over three years, which a16z calls "LLMflation" and translates to an underlying pace of about 10x per year for a fixed capability levelCITE:E1. Using a different yardstick — the lowest price across vendors for GPT-4-level capability (MMLU scores of 83 or above) — a16z found only a 62x decline between GPT-4's March 2023 launch and the end of 2024CITE:E2. OpenAI's own account of its single product line tells a third story: CEO Sam Altman stated that the cost of a fixed capability level falls roughly 10x every 12 months, and cited an approximately 150x drop in per-token price from GPT-4's early-2023 release to GPT-4o by mid-2024CITE:E3.
| Comparison basis | Time span | Price decline | Source |
|---|
| Same-capability tier, GPT-3 level (a16z "LLMflation") | Late 2021 → end 2024 (~3 years) | ~1,000x ($60 → $0.06 per million tokens) | CITE:E1 |
| GPT-4-level capability, lowest price across vendors | March 2023 → end 2024 | ~62x | CITE:E2 |
| OpenAI's own line, GPT-4 → GPT-4o | Early 2023 → mid-2024 | ~150x | CITE:E3 |
These three figures are not interchangeable: a16z's 1,000x and 10x-per-year figures describe a cross-market, same-capability benchmark over three years; its 62x figure is a narrower cross-vendor lowest-price comparison over roughly 21 months; and Altman's 150x figure describes only OpenAI's internal product progression over about 17 months. All three point in the same direction — steep, rapid decline — but the size of the number depends entirely on which slice of the market is being measured.
How is market pricing being pushed down further by competition?
Both incumbent repricing and low-cost new entrants are compounding the decline documented above. OpenAI itself supplied evidence of active repricing: its GPT-4o mini, launched in July 2024, is priced at $0.15 per million input tokens and $0.60 per million output tokens, which OpenAI states is more than 60% cheaper than GPT-3.5 Turbo and about 99% cheaper per token than its 2022 text-davinci-003 modelCITE:E4. Separately, DeepSeek's R1 model is priced at roughly $0.55 per million input tokens and $2.19 per million output tokens — about 3.6% of OpenAI o1's $15 input / $60 output pricing, or roughly 27x cheaperCITE:E5.
| Model | Input price (per million tokens) | Output price (per million tokens) | Note |
|---|
| GPT-4o mini | $0.15 | $0.60 | >60% cheaper than GPT-3.5 Turbo; ~99% cheaper than 2022 text-davinci-003CITE:E4 |
| DeepSeek-R1 | $0.55 | $2.19 | ~3.6% of OpenAI o1's price, ~27x cheaperCITE:E5 |
| OpenAI o1 (reference point) | $15 | $60 | Baseline used in the DeepSeek-R1 comparisonCITE:E5 |
Read together, these two data points show pricing pressure coming from two directions at once: OpenAI cutting its own prices well below its earlier models, and a lower-cost entrant like DeepSeek pricing an entire order of magnitude below an established closed-source model.
Why does falling per-token cost coincide with rising total inference spending?
Falling unit prices are coinciding with a sharp rise in total usage and total spending, not a decline in either. Google stated at its May 2025 I/O keynote that its products now process 480 trillion tokens per month, which it described as about 50 times the volume of a year earlierCITE:E6. At the same time, Gartner forecast that global generative-AI spending would reach $644 billion in 2025, up 76.4% year over year, while noting that roughly 80% of that spending is hardware — devices and servers — rather than API inference billsCITE:E7.
| Metric | Value | Source |
|---|
| Google monthly token processing (May 2025) | 480 trillion tokens, ~50x year-over-year | CITE:E6 |
| 2025 global generative-AI spending (Gartner forecast) | $644 billion, +76.4% year-over-year | CITE:E7 |
| Share of that spending that is hardware (devices/servers) | ~80% | CITE:E7 |
The two figures describe different parts of the same system: Google's number tracks token volume actually processed, while Gartner's number tracks total dollars spent across the generative-AI stack, most of which is capital equipment rather than usage-based API charges.
What this means
The evidence assembled here shows two trends running in parallel rather than in opposition. Per-token prices for equivalent capability have fallen by roughly 10x a year on a20 multi-year viewCITE:E1, further compressed by both incumbent repricingCITE:E4 and low-cost entrantsCITE:E5. At the same time, Google's reported 50x jump in monthly token volumeCITE:E6 and Gartner's 76.4% spending growth to $644 billionCITE:E7 show that cheaper tokens have not reduced the amount of money moving through the generative-AI market — they have coincided with more of it. Because Gartner's total is roughly 80% hardware rather than inference API chargesCITE:E7, the token-price collapse and the spending increase are not contradictory: one measures the cost of a unit of intelligence, the other measures how much capacity and usage the market is now buying at that lower price.