← Back to Research
Compute Infrastructure

AI Compute Infrastructure 2026

FutureX Research · AI Lab · 2026.05.12 · 24 pp · preview 6 pp

🎧

Listen · Audio Summary

5-8 min · AI narration in English · abstract + all key findings

Abstract

(Data as of 2026-10-02) In late July, all four axes of AI compute infrastructure tightened at once. Moonshot AI released the full weights of its 2.8-trillion-parameter Kimi K3 on July 27, open-sourcing its training infrastructure and setting Hugging Face's fastest-growth record. Late-July earnings from Google, Microsoft and Amazon pushed combined 2026 Big Tech capex above $700 billion, with markets now pricing free-cash-flow discipline rather than growth alone. TSMC raised full-year capex to $60-64 billion with CoWoS still sold out, and reportedly launched an "EMIB-like" packaging route. SpaceX broke below its IPO price within a month of listing, SK Hynix sold off despite record results, and Anthropic's October IPO window approaches. The securitization of model companies and "power as the binding constraint" now define the second half. This refresh updates all data within the original value-chain framework.

Key Findings

  • 01Moonshot AI released the full weights of the 2.8-trillion-parameter Kimi K3 on July 27 (modified MIT license), also open-sourcing its MoonEP/FlashKDA/AgentEnv training infrastructure; it topped Hugging Face trending within 30 minutes — the platform's fastest launch ever — and leads WebDev Arena at 1679 Elo, above Claude Fable 5 (1631) and GPT-5.6 Sol (1618), with Day-0 support from Huawei Ascend 950 and Alibaba Cloud. Giant open-weight models are spilling inference demand across the industry.
  • 02Late-July earnings confirmed demand: Alphabet's Q2 capex hit $44.9B (roughly doubling YoY) with full-year guidance raised to $195-205B; Microsoft spent $41B in the quarter (+70%) as Azure passed $100B in annual revenue and the stock jumped 8%; Amazon lifted 2026 capex to ~$220B; Meta's Q2 free cash flow reportedly fell to $784M. Combined Big Tech outlays exceed $700B this year (Reuters), and markets now price free-cash-flow discipline.
  • 03TSMC's July 16 Q2: revenue $40.2B (+36% YoY in NT$), net income +77.4% — a fifth straight record quarter; full-year USD revenue growth guidance raised above 40% and capex lifted from $52-56B to $60-64B; CoWoS remains sold out with 52-78 week lead times; on July 30 TSMC reportedly launched an 'EMIB-like' packaging route, sending its ADRs up over 7%. Advanced packaging is still the binding constraint.
  • 04The power bottleneck hardened: Musk-affiliated entities reportedly acquired mobile-power firm APR Energy for up to ~$1B (Axios); SpaceX (which absorbed xAI in February) said on July 31 that Memphis Colossus runs on 69 unpermitted gas turbines to be removed by July 2027 as it transitions to a permanent 1.2GW captive gas plant, with $2.8B of further turbine purchases disclosed in its IPO filing; the House Energy Committee demanded records by August 11.
  • 05Model-company securitization entered a stress test: SpaceX raised $75B on June 12 in the largest US IPO ever, then broke below its $135 issue price in mid-July, shedding over $1.2 trillion from peak value (reported); SK Hynix listed on Nasdaq July 10 and its July 29 Q2 (operating profit KRW 60.54T, +550% YoY) still missed estimates, sinking the stock; Anthropic ($965B valuation after its May round) began investor meetings July 15 targeting an October listing, while OpenAI reportedly slips to 2027.
  • 06China primary market and applications refreshed: Moonshot reportedly plans a final pre-IPO round in August targeting a $50B pre-money valuation (current round at $31.5B pre-money, per NBD); Zhipu (HK-listed since January) completed a ~HK$31.4B placement on July 13; SoftBank became Sierra's exclusive Japan partner on July 14 (LINEMO resolution rate lifted to 97%); Apple sued OpenAI on July 10 over trade secrets, which OpenAI denies — application rollout and big-tech clashes are heating up in parallel.

Introduction: Four Lines Converge in Two Weeks — Demand, Supply, Power, Capital

Our previous edition closed on July 18 with 'three lines converging in a single week: demand, supply, power.' The following two weeks grew denser still, and capital markets became the fourth line: on July 22 Alphabet raised full-year capex guidance to $195-205B; on July 27 Moonshot released Kimi K3's full weights, setting Hugging Face's fastest-growth record; on July 29 Microsoft and SK Hynix reported on the same day — the former jumped 8% on demand strength while the latter, despite operating profit up more than 5.5x, sank on a miss; on July 31 SpaceX laid out the timetable for a 1.2GW captive power plant in Memphis. Demand has not slowed, supply remains constrained, and power is now a first-order bottleneck — all three theses strengthened. The genuinely new variable: public markets now price AI narratives with free-cash-flow discipline, and capex guidance has replaced revenue growth as the swing factor for stocks.

Demand Side: Giant Open Models and the Agent Race Have Not Slowed

The loudest demand signal came from the open-source camp. Kimi K3 (2.8T-parameter MoE, ~104B active, 1M-token context) released its full weights on July 27 under a modified MIT license — unusually also open-sourcing its MoonEP, FlashKDA and AgentEnv training infrastructure — and topped Hugging Face trending within 30 minutes. It leads WebDev Arena at 1679 Elo, above Claude Fable 5 and GPT-5.6 Sol, the first open model to top that board, though it still trails the two closed flagships on composite intelligence indices. Huawei Ascend 950, Alibaba Cloud's M890, Cursor and Cognition shipped Day-0 support; Chinese models reportedly account for ~60% of US token usage on OpenRouter. The funding evidence is equally dense: late-July earnings from Google, Microsoft, Amazon and Meta push combined 2026 Big Tech capex above $700B, and SoftBank's Masayoshi Son says AI will need $5 trillion per year by 2040. The race has not cooled: xAI is reportedly training a 2T-parameter Grok 4.6 aimed squarely at surpassing K3.

Supply Side: TSMC's Record Quarter Confirms the Cycle Has Not Peaked

TSMC delivered a fifth consecutive record quarter on July 16: revenue $40.2B (+36% YoY in NT$), net income +77.4%, gross margin 67.7%, with HPC at 66% of revenue and 2nm beginning to ramp. Management raised full-year USD revenue growth guidance from above 30% to above 40% and lifted capex from $52-56B to $60-64B — the supply side voting for demand with real money. The constraint remains packaging: CoWoS is sold out with 52-78 week lead times, Nvidia reportedly holding ~60% of capacity; on July 30 TSMC was reported to be developing an 'EMIB-like' packaging route to supplement CoWoS, sending its ADRs up more than 7%. Memory is equally tight: SK Hynix posted Q2 operating profit of KRW 60.54T, up more than 5.5x YoY, with HBM4 shipping since Q2 and its CEO forecasting the worst memory shortage in history for 2027 — yet the results still missed supercharged estimates and the stock sank. 'Records that are not enough' is itself the twin footnote to this cycle's heat and fragility.

August Update · Verified (data current as of 2026-08-19): A $500bn Financing Platform Turns GPUs Into Collateral

Verified (Nvidia company statement, August 10): Nvidia signed memoranda of understanding with six asset managers — Apollo Global Management, Blackstone, BlackRock, Brookfield, Goldman Sachs and KKR — to build an independent financing platform intended to mobilize more than $500 billion of third-party capital over time, specifically for data-center construction and hardware purchases by cloud providers and AI companies. Jensen Huang publicly answered circular-financing criticism, arguing that AI-factory compute is becoming an investable asset class, that demand comes from real commercial workloads, and that each project is independently diligenced. Nvidia closed down 2.86% that day.

Addition to this report's framework: this report previously set the binding constraints on compute infrastructure as capex capacity plus power and land supply. If the financing platform works, the constraint must be rewritten as asset-financing capacity — whether GPUs and full racks can be rated, pledged and securitized. Three consequences become observable: the marginal funder of new compute shifts from technology-company balance sheets to institutional capital; depreciation assumptions and residual values face outside creditor scrutiny for the first time rather than being set by the buyer; and if the secondary market disagrees on GPU residuals, financing costs will move before chip prices do.

Counter-evidence worth noting: the market's same-day reaction was negative. The new-asset-class reading and the circular-financing reading remain unresolved. This report draws no conclusion on that dispute and suggests treating the first cohort's independent diligence outcomes and actual lending terms as the test.

Demand Composition: 85% of Agent Tokens Are Cached, Moving the Bottleneck From Compute to Memory

Verified (a16z, published August 10, 2026, data from openrouter.ai/rankings): agentic token consumption on OpenRouter reached 7.3 trillion on a seven-day average, roughly five times human usage, growing about 14x in the six months after overtaking humans on February 6, 2026.

Verified (OpenRouter data, disclosed via a16z): more than 85% of agentic token consumption comes from cached prompts, which account for nearly all of the category's relative growth. Agents repeatedly run read-write-execute loops and preserve context across operations, so the initial prompt carrying task context is counted again and again.

Addition to this report's infrastructure call: this report reduced the demand side to growth in inference calls and derived GPU and rack demand from it. That simplification needs correcting. Cache-heavy workloads require the KV cache to stay resident for long periods, which lowers compute intensity per token while materially raising memory capacity and bandwidth occupancy. The direct implication is that the marginal bottleneck is more likely to appear in HBM supply, memory bandwidth and interconnect than in FLOPs themselves.

Testable implications: if this structure persists, three things should follow — high-bandwidth memory tightness arriving earlier and running hotter than GPU supply overall; memory-to-compute ratios rising in new platform generations; and inference providers drawing sharper pricing distinctions between cached and non-cached tokens. These serve as tests that would confirm or falsify this section.

Uncertainty to state alongside: this structure rests on a single platform's data. OpenRouter's workload mix need not represent the whole industry, and in particular need not represent hyperscalers' self-operated inference clusters. Treat this as directional, not as an industry-wide quantitative conclusion.

Late-August Update · Verified (data current as of 2026-08-31): Anthropic Retreats from Buying a Chip Company to Partnering — the Real Bar for In-House Silicon

Verified (Reuters exclusive, August 27–28, 2026; widely syndicated): Anthropic held talks to acquire AI chip startup MatX for roughly $7 billion, then abandoned the acquisition in favour of discussing a partnership. MatX, founded by former Google TPU engineers, is building a chip for training large models; it is reported to be raising a new round at about a $4 billion valuation. Reuters could not determine what ended the active negotiations; Anthropic declined to comment and MatX did not respond.

Why this matters more than another M&A rumour: a $7bn bid for a company valued at $4bn — nearly a 75% premium — signals extreme urgency around in-house silicon. Retreating to a partnership signals that the barrier is not money. The real constraints on custom silicon are three: long-cycle commitments for leading-edge foundry capacity, the migration cost of the software stack (compilers, kernel libraries, framework support), and the calendar time of each tape-out. None of the three shortens because $7bn was paid.

Amendment to this report's core conclusion: we previously argued that in-house silicon is the inevitable endpoint of frontier labs' compute cost structure. A qualifier is now needed — inevitable, but not necessarily self-built. Google's TPU is a decade of accumulation; Amazon's Trainium rests on its own cloud. Anthropic has neither fabs nor a cloud of its own, so its realistic path is deep custom partnership rather than vertical ownership. The economics differ entirely: partnership buys design optimisation and cost reduction, but not supply-chain exclusivity.

Implication for private markets: the exit path for AI chip startups is shifting from outright acquisition by a hyperscaler toward strategic partnership plus a minority stake. For early investors that lowers the exit valuation ceiling — but it also lets the company stay independent and serve multiple customers. The two outcomes imply completely different pricing models and must be distinguished in diligence.

Outside Views: Two More Sets of Eyes on the Shortage (updated 2026-09-02)

Dylan Patel (founder, SemiAnalysis; Aug 17; verified): the supply-demand gap is widening, not narrowing, so cost per token keeps climbing; memory is in a multi-year structural shortage — capacity grows 20–30% a year against demand that doubles. His late-June claim bears directly on this report's MatX amendment: co-designing silicon, kernels and models yields 100x; optimizing any single layer yields 8x — the real logic behind frontier labs' custom-silicon push, and why a design partnership is still worth having.

Gavin Baker (CIO, Atreides; Aug 23; verified): his firm's internal AI spend is roughly 100x higher in August than in March and still doubling monthly; he calls this a shortage running through 2028, not a bubble, and challenges bears to "name one metric in your business that is getting worse." Reportedly he also estimates large B200 cluster rents rose 50–60% in seven months — contract repricing at renewal, the very dial this report tracks.

Both men point the same way, and both sit on the buy side (one sells research, one runs long positions) — read them with that in mind. The bears' hardest number is Meta's Q2 free cash flow, down 91% year on year. The falsifier for the shortage story is simple: rent repricing stops.

FutureX Position · Open Source Breaks the Deadlock: Constraints Move, Debt Sits in the Middle Layer (Xiamen keynote, 2026-09-03)

The Xiamen keynote reinforces this report's conclusion and supplies the mechanism behind shifting constraints. FutureX's view: the binding constraint keeps moving. Chips came first, then power, then memory; today the tightest one is data. Each move opens an investment window.

On FutureX's five-layer map, infrastructure carries the most debt. Circular financing, credit sliding down the stack and GPU depreciation all land in this one layer. Our reading is a maturity mismatch: financing bonds typically run 7 to 10 years, book life is 5 to 6 years, economic life may be only 2 to 3. This report left the circular-financing question open. FutureX's position is to stay out of debt-financed compute rental.

Energy has not entered the narrative yet. FutureX's view: the ceiling on compute is written on the grid.

At the chip layer, our reading: the bottleneck is moving from raw compute to bandwidth and energy efficiency. Compute efficiency rose about 1,000x from 2012 to 2023. Quantization contributed 32x, instruction sets 12.5x, and process from 22nm to 4nm only 3x. The next 1,000x will come from silicon photonics, compute-in-memory and 3D integration, not from process nodes. This supports the report's conclusion: custom silicon is inevitable, building it yourself is not. The gains sit in architecture.

Full deck: /reports/open-source-breakthrough (first 5 pages public).

Early-September Update · Verified (data current as of 2026-09-09): Custom-chip orders accelerate, Nvidia buys the open-model hub, and neocloud debt gets priced on its own

Reported (CNBC, Broadcom press release, September 2, 2026): Broadcom reported fiscal Q3 2026 results (quarter ended August 2) after the close: revenue of $29.6 billion, up 86% year over year; AI semiconductor revenue of $16.7 billion, up 221%; adjusted EPS of $3.32 against a $3.24 consensus. Guidance calls for Q4 AI semiconductor revenue of $21.7 billion, up 236%, and on the call management outlined roughly $58 billion of AI revenue in fiscal 2026, $115 billion in fiscal 2027 and $230 billion in fiscal 2028. Note on scope: figures are taken from CNBC and Seeking Alpha coverage of the release and call; Broadcom's investor relations page did not load during this review and will be checked against the original next round.

Company disclosure (NVIDIA blog, September 3, 2026): Nvidia announced it will acquire Hugging Face for $12.93 billion. According to the company, the platform serves 18 million developers and researchers, hosts 3 million models, 500,000 datasets and 1 million applications, and is used by more than 200,000 companies; Nvidia itself has published over 500 models and 250 open datasets there. Jensen Huang said developers will keep choosing their own models, frameworks, clouds and inference providers, and that Nvidia compute will not be required to build on or deploy through Hugging Face.

Reported (The Motley Fool, September 3, 2026): Neocloud debt is being priced on its own. CoreWeave carries $35 billion of debt, guides 2026 capital expenditure to $35-39 billion, and paid $640 million of interest last quarter, 2.4 times the year-earlier figure; Nebius spent $8 billion on capex in the first six months of 2026, plans more than $20 billion for the full year, and has just raised $5.75 billion through convertible notes. Note on scope: these figures are the article's compilation of the two companies' earlier public disclosures; the pricing date of the Nebius convertible was not verified in this pass.

Impact on this report's conclusions: Two findings are reinforced and one premise is revised. First, demand did not slow after the July earnings season; Broadcom extended custom-silicon visibility from one quarter to three fiscal years. Second, capital is now priced by who is borrowing to buy chips as much as by who is buying them, so neocloud debt loads and hyperscaler free cash flow are read separately. The premise to revise sits in the open-source section: the largest distribution point for open models now belongs to the largest compute vendor. The view that the advantage sits in the middle layer still holds, but ownership of that layer is concentrating, and the platform-neutrality commitment is the item to track.

Mid-to-Late-September Update · Verified (data current as of 2026-09-25): Demand is now on the order book; the constraint sits in interconnect and power

Company disclosure (Oracle fiscal Q1 2027 results, September 10, 2026): Remaining performance obligations rose $209 billion year over year to $664 billion, and Oracle booked more than $30 billion of new AI cloud contracts in the quarter. Total revenue was $19.3 billion, up 30%; cloud revenue was $11.6 billion, up 62%, of which cloud infrastructure revenue was $7.4 billion, up 121%. Quarterly capital expenditure was $28.5 billion, about 3.4 times the $8.5 billion a year earlier, and the company brought 850 MW of additional data center capacity online during the quarter.

Verified (Tencent News report dated September 18, 2026, on Huawei Connect 2026, held September 17 in Shanghai): Huawei deputy chairman and rotating chairman Wang Tao unveiled the Ascend 960 SuperPoD, the first to use near-packaged optics (NPO). A single node scales to 4,096 cards with 8 EFLOPS of FP8 compute and 1 PB of HBM; 5,500 Hi-ONE units replace 48,000 800G optical modules, cutting power draw by more than 550 kW and lifting model FLOPS utilization 2.75x over a conventional cluster. The Ascend 960DT is scheduled for Q1 2027, three quarters ahead of the original plan; the 970 and 980 are slated for 2028 and 2029. Huawei says more than 1,000 Ascend SuperPoDs are deployed across more than 370 customers. The same report says DeepSeek plans to deploy at least 160,000 Huawei AI accelerators at a data center in Inner Mongolia (as reported).

Company disclosure (NVIDIA newsroom, September 9, 2026): NVIDIA announced work with eight Australian partners — Firmus, Sharon AI, IREN, ResetData, Megaport, CDC, NEXTDC and AirTrunk — to deliver up to 2 GW of AI compute by 2027. CDC says it operates more than 550 MW across Australia and New Zealand with a further 800 MW under construction; Sharon AI plans to deploy up to 68,000 NVIDIA GPUs. No dollar figures were disclosed.

Effect on this report's conclusions: Two judgments are reinforced. First, demand has moved from the capex line of cloud providers' earnings to signed customer contracts; Oracle's $664 billion backlog and 850 MW of delivery in a single quarter make on-time delivery the main variable. Second, "the constraint is migrating": Huawei building optics into the SuperPoD and NVIDIA placing a 2 GW plan in Australia both point to inter-card interconnect and grid access replacing the chip itself as the bottleneck. One correction: the late-August section used Anthropic to stress the difficulty of in-house silicon; Huawei pulling the 960DT forward by three quarters shows that schedule depends first on having a committed buyer, and a full-stack vendor with orders in hand can compress it.

Late-September to Early-October Update · Verified (data current as of 2026-10-02): Micron books $37.7 billion in fiscal fourth-quarter net income, and Anthropic’s draft prospectus reportedly lists about $518 billion in compute commitments

Per company disclosure (Micron earnings, via CNBC, September 30, 2026): Micron reported fiscal fourth-quarter revenue of $54.23 billion, up from $11.32 billion a year earlier, and net income of $37.7 billion versus $3.2 billion. DRAM revenue rose 343% to $39.8 billion, or 73% of total sales. The company guided next-quarter revenue to about $61.5 billion, above the $57 billion analysts expected. Micron is investing $250 billion to build two new HBM campuses in New York and Idaho.

Reported (Reuters, September 29, 2026): A non-public draft Anthropic prospectus seen by Reuters shows roughly $518 billion in planned cloud, compute and infrastructure commitments over the coming years, most of which cannot be canceled. Anthropic spent about $7.33 billion on compute and infrastructure in 2025, roughly three times the prior year, on revenue of nearly $4.6 billion.

Reported (Bloomberg, investingLive, September 25, 2026): Goldman Sachs expects hyperscaler AI capital spending of around $800 billion in 2026, rising about 50% to roughly $1.2 trillion in 2027.

Implications for this report: These facts reinforce two of our calls: demand is being locked into contract backlogs, and the bottleneck is moving from compute to memory. Micron's net margin of nearly 70% shows memory suppliers are still collecting a shortage premium. Anthropic's multi-year compute commitments are more than 110 times its 2025 revenue, so the center of demand-side risk is starting to shift from whether orders exist to whether buyers can keep financing the payments. Two thresholds to watch: whether Micron's next-quarter revenue lands near $61.5 billion (a clear miss would weaken the memory-shortage call), and whether hyperscalers' full-year 2027 capex guidance in early 2027 sums to close to $1.2 trillion. A combined figure below $1 trillion would move our supply-tightness scenario back toward the base case.

Key Questions

How much are Google, Microsoft and Amazon spending on AI capex in 2026?

Over $700 billion combined (Reuters). Alphabet's Q2 capex hit $44.9B with full-year guidance raised to $195-205B; Microsoft spent $41B in the quarter (+70%) as Azure passed $100B annual revenue; Amazon lifted 2026 capex to ~$220B. Markets now price free-cash-flow discipline — Meta's Q2 FCF reportedly fell to just $784M.

Is TSMC's CoWoS advanced packaging still sold out, and what are lead times?

Yes — CoWoS remains sold out with 52-78 week lead times. TSMC's July 16 Q2 showed $40.2B revenue (+36% YoY in NT$) and net income +77.4%, a fifth straight record; full-year capex was raised from $52-56B to $60-64B. On July 30 TSMC reportedly launched an 'EMIB-like' packaging route, lifting its ADRs over 7%. Advanced packaging is still the binding constraint.

What did Moonshot open-source with Kimi K3, and does it beat Claude and GPT?

On July 27 Moonshot released the full weights of the 2.8-trillion-parameter Kimi K3 under a modified MIT license, plus its MoonEP/FlashKDA/AgentEnv training infrastructure. It topped Hugging Face trending within 30 minutes — the platform's fastest launch ever — and leads WebDev Arena at 1679 Elo, above Claude Fable 5 (1631) and GPT-5.6 Sol (1618), with Day-0 support from Huawei Ascend 950 and Alibaba Cloud.

Watch & Listen

In China: search WeChat Channels for 「倩姐投AI」; full library → Qian on AI

Sourcing and standards

Compiled from public sources; data current as of 2026.05.12. The text separates verified facts, reported claims, our own estimates and disputed points, and states the derivation behind every estimate. When we get something wrong, the correction is written into the report body with the original call left visible, and logged publicly.

Research standards & corrections →

📄 Full Report

Full report: 24 pages · provided to professional investors & partners only

This is the public preview. The full report includes the sections below. For compliance reasons it isn't posted publicly or offered as a free download. To request a copy, contact the FutureX team.

  • 🔒The Power Bottleneck: From APR Energy to a 1.2GW Captive Plant — the Cost of 'Bring Your Own Power' and the Regulatory Backlash
  • 🔒Capital Markets: SpaceX Below Issue, SK Hynix Sold Off on Records — Pricing Anthropic's October IPO
  • 🔒Primary Market: Moonshot's $50B Pre-Money Push and Zhipu's HK$31.4B Placement — Falsifiable Premiums
  • 🔒Applications: SoftBank x Sierra Exclusivity and the Unit Economics of Agent Customer Service
  • 🔒Clash of Giants: Apple v. OpenAI and the Latest in the AI Hardware Talent War
  • 🔒Value-Chain Map, Risk Matrix and Scenario Analysis
Request the full report →

Where we stand on this

Questions people ask next

Building in this space, or want to discuss this report? Write to us. We usually reply within 48 hours →

Related Research

Industry research from FutureX Capital's AI Lab, compiled from public information; not investment advice; contains no fund performance, AUM, or offer to raise capital.