← Back to Research
AI Compute Economics

Global AI Compute and Token Economics: The Inference Era

FutureX Research · AI Lab · 2026.03.15 · 26 pp · preview 7 pp

🎧

Listen · Audio Summary

5-8 min · AI narration in English · abstract + all key findings

Abstract

(Data updated through 2026-10-02) Inference now accounts for roughly two-thirds of global AI compute, and Goldman Sachs projects monthly token consumption to grow 24x to 120 quadrillion by 2030. Hard events in late July reinforced the "inference era" thesis: Moonshot AI open-sourced the 2.8-trillion-parameter Kimi K3 weights on July 26 as promised, then closed a $3.5B round at a $35B post-money valuation; Anthropic released Opus 5 on July 24 ahead of its October listing push; OpenAI cut GPT-5.6 Luna prices by 80% on July 30 and DeepSeek shipped V4-Flash on July 31, igniting a full-scale token price war; the four hyperscalers now plan roughly $725B of combined 2026 capex, and Alphabet's stock fell despite raising guidance to up to $205B; DeepSeek's second funding round (~$71B valuation) was reportedly paused on July 25. This report is based on public information and does not constitute investment advice.

Key Findings

  • 01Inference now takes roughly two-thirds of global AI compute (Deloitte TMT Predictions 2026; ~1/3 in 2023, ~1/2 in 2025); Goldman Sachs' May report "Decoding the Agentic Economy" projects monthly token consumption growing 24x to 120 quadrillion by 2030; agentic token flow on OpenRouter surpassed direct human usage around early February 2026, with agent requests consuming ~15x the tokens of human chat.
  • 02Open source and share: Moonshot AI released Kimi K3 weights on July 26 (2.8T parameters, 1M context — the largest open-weight release in history, one day ahead of schedule); Chinese models hit a record 58% of OpenRouter token share in July (briefly 63% in the first week), holding all top-five spots; Bloomberg reported July 29 that Moonshot closed a $3.5B round at a $35B post-money valuation and plans a pre-IPO round in August at up to $50B.
  • 03A token price war erupted in the last week of July: Anthropic released Opus 5 on July 24 ($5/$25 per 1M tokens with an adjustable effort dial); OpenAI cut GPT-5.6 Luna by 80% to $0.20/$1.20 and Terra by 20% to $2/$12 on July 30; DeepSeek shipped V4-Flash 0731 ($0.14/$0.28) on July 31; Alibaba's Qwen3.7 Flash (July 27) and Google's Gemini 3.6 Flash (July 21) landed in the same window.
  • 04Dual constraints of power and capex: xAI's $1B+ acquisition of APR Energy (1GW+ of mobile generation) won final clearance, but the NAACP and environmental groups filed a Clean Air Act suit over up to 35 allegedly unpermitted gas turbines (allegations pending), and TechCrunch reported July 31 that SpaceX will not remove all related turbines for another year; Alphabet's Q2 capex hit $44.9B with full-year guidance raised to $195-205B and free cash flow turning negative at -$5.9B; Meta guided $130-145B; the four hyperscalers plan ~$725B combined 2026 capex (~+77% YoY).
  • 05Endpoints and agent commercialization: OpenAI's first hardware, the $230 Codex Micro, sold out in ~12 hours on July 15, first deliveries July 24, with eBay resales around $800 and listings up to $1,850; Apple sued OpenAI and io for trade-secret theft on July 10 (allegations denied by OpenAI); SoftBank became Sierra's exclusive Japan partner on July 14, with pilot resolution rates rising from 83% to 97% as enterprise token spend shifts to outcome-based pricing.
  • 06Capital-market repricing: Anthropic confidentially filed its S-1 on June 1 targeting an October Nasdaq listing led by Goldman Sachs, JPMorgan and Morgan Stanley, expected to raise over $60B (May Series H valuation ~$965B); Zhipu placed new H-shares at HK$1,588 on July 9, raising ~HK$31.4B, and launched its "Touch High" plan on July 11; DeepSeek, after a record $7B first round in June, reportedly paused its second round (~$71B valuation) on July 25 per Bloomberg while preparing a STAR Market listing.

1. Overview: Tokens as the Yardstick of the Inference Era

The center of gravity in AI compute has shifted from training to inference. Deloitte's TMT Predictions 2026 estimates inference at roughly two-thirds of global AI compute in 2026, versus about one-third in 2023 and half in 2025. Goldman Sachs' May report "Decoding the Agentic Economy" draws a steeper curve: monthly token consumption reaching 120 quadrillion by 2030, a 24x increase, with consumer agents contributing 12x growth by 2030 and enterprise agents 55x by 2040. The driver is the always-on nature of agents: on OpenRouter, agentic token flow surpassed direct human usage around early February 2026, and a single agent request consumes roughly 15x the tokens of a human chat. Parallel to surging volume is collapsing unit price — per-token inference costs are falling 60-70% per year at the silicon level, and the price war in the last week of July (Opus 5 at $5/$25, GPT-5.6 Luna cut 80%, DeepSeek V4-Flash at $0.14/$0.28) compressed price cuts from an annual event into a weekly one. This volume-price scissors is the through-line of our value-chain analysis.

2. The Open-Source Chase: Kimi K3 and Chinese Models' Token Share

Kimi K3 (a 2.8-trillion-parameter MoE with a 1M-token context), released July 16, honored its open-source pledge on July 26 — dropping weights a day ahead of the promised July 27 — becoming the largest open-weight model in history, with day-0 hosting from Together AI and Modal. Share data also set records: Chinese models took 58% of OpenRouter token share in July (briefly touching 63% in the first week), holding all top-five spots by usage — Xiaomi's MiMo-V2.5 first, followed by DeepSeek, MiniMax, Qwen and Kimi — versus under 10% share in early 2025. Price is the core driver: open Chinese models run 60-90% cheaper than leading closed US APIs, and in agentic workloads where usage is amplified 15x, the unit-price gap is amplified by the same factor. Capital followed: Bloomberg reported July 29 that Moonshot closed a $3.5B round at a $35B post-money valuation — far above its initial $1-2B target — and plans a pre-IPO round in August at up to $50B ahead of a Hong Kong listing. Open weights have turned from a technical gesture into a valuation engine.

3. Physical Constraints: xAI's APR Energy Deal and 'Power Is Compute'

The bottleneck in compute is shifting from chips to power, and July's developments made this thread more concrete. xAI's $1B+ acquisition of mobile-generation firm APR Energy (1GW+ of gas turbines) to power the Memphis Colossus supercomputer has won final clearance; but environmental and compliance risks are rising in parallel: the NAACP and environmental groups filed a Clean Air Act lawsuit over up to 35 allegedly unpermitted gas turbines (allegations pending in court, not accepted by xAI), and TechCrunch reported July 31 that SpaceX will not remove all related turbines for another year. The other face of the power constraint is capex: Alphabet's July 22 report showed Q2 capex of $44.9B, full-year guidance raised to $195-205B, and quarterly free cash flow turning negative at -$5.9B — the stock fell about 5% the next day; Meta raised full-year capex guidance to $130-145B on July 29 with its stock also under pressure; Microsoft's AI business within Azure hit a $37B annualized run rate, up 123% YoY. The four hyperscalers plan roughly $725B of combined 2026 capex, up ~77% from ~$410B in 2025 — 'power is compute' is turning from slogan into a hard line item on income and cash-flow statements.

August Update · Verified (data current as of 2026-08-19): The Cheap-Token Era Ends as Compute Becomes an Asset Class

Reported (financial media, mid-August): DeepSeek announced an across-the-board increase to its API pricing. This is the first time a leading Chinese lab has raised prices at the same moment its capability jumped — an inflection in the two-year pattern of buying share with rock-bottom token prices.

Verified (SpaceX AI, August 12): in the same week Grok 4.6 entered at $2 per million input tokens and $6 per million output, roughly half of comparable frontier models.

Verified (Nvidia company statement, August 10): Nvidia signed memoranda of understanding with Apollo Global Management, Blackstone, BlackRock, Brookfield, Goldman Sachs and KKR to build an independent financing platform intended to mobilize more than $500 billion of third-party capital over time for data-center construction and hardware purchases by cloud providers and AI companies. Jensen Huang stated publicly that AI-factory compute is becoming an investable asset class, with demand rooted in real commercial workloads and each project independently diligenced. Nvidia shares closed down 2.86% that day.

Implication for this report's model: if compute can genuinely be collateralized and securitized, the decline in unit token cost is no longer driven by process improvements and competition alone — it also depends on the cost of capital, so rates and credit spreads transmit directly into inference pricing. This report modelled unit prices falling as supply expanded; financing cost now has to enter that model. Application-layer businesses whose gross margin assumes upstream prices only fall should re-run that stress test.

A Structural Shift in Demand: People Are No Longer the Main Consumers of Tokens

Verified (a16z, published August 10, 2026, data from openrouter.ai/rankings): on OpenRouter, agentic token consumption has reached 7.3 trillion on a seven-day average — roughly five times what human users consume directly. The agentic line first crossed above human usage on February 6, 2026, and grew about 14x over the following six months.

Verified (OpenRouter data, disclosed via a16z): more than 85% of agentic token consumption comes from cached prompts, and cached tokens account for nearly all of the category's relative growth. The cause is the structural difference between an agent and a one-off chat: an agent repeatedly runs a read-write-execute loop and preserves context across operations, so the initial prompt — loaded with a codebase style guide, company policy or task brief — is counted again and again.

Our estimate (read off the a16z chart's scale, not a disclosed figure): human and mixed usage each sat in the 1.3-1.4 trillion range over the same period. This is included only to convey relative magnitude and should not be cited as precise data.

Three revisions to this report's model:

First, the growth curve for inference demand should be decoupled from user counts. This report modelled token demand from DAU and calls per user; with agents now the dominant consumer, the real drivers are agent runtime, loop count and context length. Forecasting compute demand from a user-growth model will systematically understate it.

Second, surging token volume does not translate proportionally into model-vendor revenue. Cached prompts are priced well below fresh input tokens, and when more than 85% of the increment is cached, inferring revenue from total token volume overstates it badly. Monetization analysis should separate cached from non-cached volume.

Third, and most consequential for infrastructure: cache-heavy workloads keep the KV cache resident for long periods, shifting the bottleneck from FLOPs toward memory capacity and bandwidth. That runs against the instinct to answer a compute crunch by adding GPUs — the binding constraint is more likely on the memory side. This should be read together with this report's compute-infrastructure conclusions.

The Full Picture of Tiered Pricing: Cuts and Increases on the Same Price List

Verified (CNBC, Axios, VentureBeat and others, July 30, 2026): OpenAI repriced the GPT-5.6 family — Luna cut 80%, from $1 to $0.20 per million input tokens and $6 to $1.20 output; Terra cut about 20%, input $2.50 to $2 and output $15 to $12; while flagship Sol held at $5 input and $30 output. The repricing came just three weeks after the family launched on July 9.

Verified (official API docs and releases, Aug 12–13): DeepSeek shipped the V4 Pro release build while signalling across-the-board price increases; Grok 4.6 entered at $2 input and $6 output per million tokens.

Deepening this report's call: this report previously summarized August as one lab raising prices while another cut them. That summary was too coarse. Folding in the July 30 repricing, the real structure is that the split runs inside a single vendor: the cheap tier lost 80% of its price in three weeks while the flagship tier did not move at all. Downward pressure is therefore not spread evenly across frontier capability — it concentrates on the most substitutable tier. Capability gaps among small models have converged to the point where users choose on price, while the flagship tier retains enough irreplaceability to keep its pricing power intact.

Implications and falsifiable tests: if this structure holds, three things should follow — first, low-tier prices continue converging across vendors while flagship spreads widen; second, vendor disclosure shifts from model capability toward per-tier use cases, to justify the tiering; third, application-layer cost curves become tightly coupled to model selection, so forecasting cost from a single average token price will misstate it. These three serve as tests that would confirm or falsify this section.

Early-September Update · Verified (data current as of 2026-09-09): Frontier token prices flat for three months; competition shifts to long-term compute contracts and same-price upgrades

Confirmed (Google blog, September 2, 2026): Google released Gemini 3.8 Flash and 3.8 Flash Cyber, a restricted-access variant for cybersecurity work. API pricing keeps the introductory rate used for 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through December 31, 2026; from January 1, 2027, the standard rate of $1.50 and $7.50 applies. Workhorse models are now iterated at the same price with an expiry date attached; the price itself did not move.

Public market data (BenchLM Token Price Index, updated September 3, 2026): the frontier token price index reads 16 for September (March 2023 = 100), unchanged from July and August, the third straight month without a decline. The median blended price across 21 frontier models is $6.00 per million tokens, 84% below the base period. The index covers 40 models (21 frontier, 11 mid-tier, 8 budget). Two changes were recorded in the period: Claude Sonnet 5's blended price rose from $4 to $6, a 50% increase, and Claude Fable 5.1 entered the frontier tier at $20 blended.

Reported (Wall Street Journal, with Yahoo Finance follow-ups on September 1 and 2, 2026): Anthropic signed a $35 billion, six-year compute agreement with Nvidia-backed cloud provider Lambda, covering roughly 350 MW at Hut 8's Beacon Point campus in Nueces County, Texas, with capacity expected online in the first quarter of 2027. Nvidia holds the 15-year site lease, and Lambda buys and installs the GPUs and resells the compute to Anthropic; Lambda closed a $926 million senior secured term loan on August 27. Hut 8 has disclosed that the two-phase lease totals 704 MW with a base contract value of $19.6 billion. Together with the earlier $45 billion, six-year Nscale agreement (460 MW), Anthropic's named compute contracts total at least $80 billion. Note on figures: the September 1 report said term and capacity were undisclosed; the six-year term and 350 MW come from the September 2 report.

Implications for this report's conclusions: two judgments are reinforced and one needs revision. Three flat months in the price index reinforce the view that the era of cheap tokens has ended; frontier-tier cuts have been replaced by same-price upgrades and selective increases. The three-layer structure of the Lambda deal, with Nvidia holding the lease, Lambda borrowing to deploy, and the model company paying, reinforces the view that compute is being turned into an asset: supply now comes with a lease, a loan, and a price. The phrase "cuts and increases on the same price list" needs revision: September data show the frontier tier as a whole has reached a plateau, and the divergence has moved to selective increases and premium entries within the frontier tier, plus expiry-dated introductory pricing for workhorse models, with time as the new pricing tool.

Mid-to-Late-September Update · Verified (data current as of 2026-09-25): Compute contracts stretch to seven years, CPUs join the inference bill, and architecture keeps cutting per-token cost

Verified (DeepSeek official announcement, September 10, 2026): DeepSeek released V4.1-Flash, a 552-billion-parameter MoE model built on a new Causal Encoder-Decoder architecture that activates 8 billion parameters on input and 16 billion on output. KV-cache demand on HBM falls to one quarter of the previous generation and on SSD to one eighth. The company cut API prices with the release and introduced peak/off-peak pricing, with off-peak rates at half of peak. The selling point of this generation is inference cost: the memory bill for long-context and multi-turn agent calls is lowered by the structure itself.

Company disclosure (Akamai press release, September 24, 2026): Akamai signed a seven-year, $11.6 billion compute agreement with Anthropic to serve Anthropic's growing CPU workloads, expandable by a further $9 billion to a total of roughly $20 billion. Akamai plans about $5.5 billion of capital expenditure tied to the commitment and issued Anthropic a warrant for up to about 5% of its common stock at $111.33 per share, of which about 2% vests with the current commitment and about 3% vests on expansion milestones.

Reported (Bloomberg, September 10, 2026; relayed by Dataconomy on September 11): Microsoft plans to raise global data center capacity from about 12 gigawatts today to more than 38 gigawatts by 2032. The figure covers owned and leased sites and excludes capacity rented from neoclouds such as CoreWeave. The report cites people familiar with the plans; Microsoft has not confirmed the number in any filing or earnings call, and the report says the roadmap could still change.

Reported (Zhongtai Securities media and internet team, September 11, 2026): As of September 10, weekly token calls on OpenRouter were up 1,863% from the start of the year; the three most-called models were Hy4-preview, GPT-5.6 Luna and GLM 5.3 Flash. Volume on a third-party router is a direct reading of demand, and it moves in the same direction as supply-side price cuts.

Impact on this report's conclusions: These events reinforce the August finding that compute is being priced as a long-lived asset; a seven-year CPU contract and a capacity plan running to 2032 treat compute as balance-sheet material. They refine one point of granularity: the inference bill is no longer GPU-only, and CPUs enter long-term contracts at a ten-billion-dollar scale for the first time, so inference cost models need separate GPU and CPU lines. The driver of falling per-token cost is shifting from price wars to architecture; V4.1-Flash's KV-cache compression shows the supply-side cost curve still falling, while OpenRouter's volume growth shows demand responding to lower prices.

Late-September to Early-October Update · Verified (data current as of 2026-10-02): OpenRouter weekly token volume hits 146 trillion with Chinese models at 63% of the top 20, and Anthropic's draft prospectus reportedly shows about $518 billion in compute commitments

Reported (Yilan Business, September 28, 2026): In the week to September 27, weekly token consumption on OpenRouter reached 146.0 trillion, up 13.2% week on week and another record. Chinese models took 10 of the top 20 slots, consuming 71.87 trillion tokens combined, or 63.0% of the top-20 total. The top three were DeepSeek V4.1 Flash (19.60 trillion), GLM 5.3 Flash (16.30 trillion) and the anonymous "Space Bunny" model (13.90 trillion); tokenizer tests suggest it may come from MiniMax, which is unconfirmed.

Reported (draft prospectus seen by Reuters and the Financial Times, relayed by Fortune and The Decoder, September 28-29, 2026): Anthropic has about $518 billion in commitments for cloud services, compute and infrastructure over the coming years. Its 2025 compute and infrastructure spending was $7.33 billion, roughly triple the prior year.

Per company disclosure (Anthropic website and VentureBeat, September 28, 2026): Anthropic released Sonnet 5.5, priced at $2 per million input tokens and $10 per million output tokens, half the price of Opus 5.5 ($4/$20). The company says total cost per task can fall by up to 30%, mainly because the model uses fewer tokens and fewer tool calls.

Implications for this report: This strengthens the inference-era thesis and adds one new signal: Anthropic is marketing its price advantage as total cost per task rather than only price per million tokens, so the unit of price competition may start to shift. On the demand side, OpenRouter's latest weekly volume rose 13.2% week on week to another record. On the supply side, a single model company reportedly carries compute commitments on the order of $500 billion, which further supports our view that compute is being locked up in long-term contracts. Thresholds: if OpenRouter weekly growth stays below 5% for four straight weeks, or the Chinese share of the top 20 falls below 50%, our views on token growth and open-model share should be revisited.

Key Questions

What share of global AI compute does inference take in 2026, and how fast will token consumption grow?

Inference now accounts for roughly two-thirds of global AI compute (Deloitte TMT Predictions 2026; about 1/3 in 2023, 1/2 in 2025). Goldman Sachs' May report projects monthly token consumption growing 24x to 120 quadrillion by 2030. Agentic traffic on OpenRouter surpassed direct human usage around early February 2026, with agent requests using about 15x the tokens of human chat. Data through 2026-08-01.

Which models cut prices in the July 2026 token price war, and to what levels?

The price war erupted in the last week of July 2026: OpenAI cut GPT-5.6 Luna by 80% to $0.20/$1.20 per 1M tokens and Terra by 20% to $2/$12 on July 30; DeepSeek shipped V4-Flash 0731 ($0.14/$0.28) on July 31; Anthropic released Opus 5 ($5/$25, adjustable effort dial) on July 24; Alibaba's Qwen3.7 Flash (July 27) and Google's Gemini 3.6 Flash (July 21) landed in the same window.

How are Chinese models performing in global open-source and token-share terms, and what did Kimi K3 open-source?

Moonshot AI open-sourced Kimi K3 weights on July 26, 2026: 2.8 trillion parameters and 1M-token context, the largest open-weight release in history, one day ahead of schedule. Chinese models hit a record 58% of OpenRouter token share in July (briefly 63% in the first week), holding all top-five spots. Bloomberg reported July 29 that Moonshot closed a $3.5B round at a $35B post-money valuation, planning an August pre-IPO round at up to $50B.

Watch & Listen

In China: search WeChat Channels for 「倩姐投AI」; full library → Qian on AI

Sourcing and standards

Compiled from public sources; data current as of 2026.03.15. The text separates verified facts, reported claims, our own estimates and disputed points, and states the derivation behind every estimate. When we get something wrong, the correction is written into the report body with the original call left visible, and logged publicly.

Research standards & corrections →

📄 Full Report

Full report: 26 pages · provided to professional investors & partners only

This is the public preview. The full report includes the sections below. For compliance reasons it isn't posted publicly or offered as a free download. To request a copy, contact the FutureX team.

  • 🔒4. The Endpoint War: Codex Micro's 12-Hour Sellout, 8x Resale Premium, and Apple's Trade-Secret Lawsuit
  • 🔒5. Agent Commercialization: The SoftBank-Sierra Exclusive and the Outcome-Based Pricing Revolution
  • 🔒6. Capital Markets: Anthropic's October IPO Sprint with Opus 5 and the Near-Trillion-Dollar Benchmark
  • 🔒7. Private Markets: Moonshot at $35B Post-Money, Zhipu's 'Touch High', and What DeepSeek's Paused Round Signals
  • 🔒8. Value-Chain Map: Winners and Pressured Segments of the Inference Era (A Neutral Framework)
  • 🔒9. Risks and Uncertainties: Power Litigation, Price Wars, Capex, and Valuation Reversal
Request the full report →

Where we stand on this

Questions people ask next

Building in this space, or want to discuss this report? Write to us. We usually reply within 48 hours →

Related Research

Industry research from FutureX Capital's AI Lab, compiled from public information; not investment advice; contains no fund performance, AUM, or offer to raise capital.