Core theses · Updated 2026-09-25
Value lands in the application layer and at the open-source chokepoints; the model layer cannot hold it
01The claim
This cycle's AI value lands in two places: ARR in the application layer, and the open-source chokepoints every model passes through. The model layer cannot hold it: large models have no network effects and no switching costs, so pricing power dissolves as capability converges. The bubble sits in the middle layer, which carries the most debt and lives on refinancing.
02Mechanism: why it happens
The model layer cannot hold value because large models have no network effects and no switching costs. One more user does not make the model better for the next user; swapping models means changing one endpoint. Revenue therefore follows the strongest model, and once capability converges, pricing power goes with it. FutureX Capital's reading: the best US and Chinese models sit 2.7% apart, the two leading labs swapped ARR rankings within a year, wholesale tokens have become a gross-margin war, and model companies are themselves moving up into applications.
Case: on LM Arena's March 2026 reading, the best US and Chinese models sit 2.7% apart (public statistics). On August 26, 2026 Zhipu open-sourced GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output, about 1/40th of Claude Opus 4.8 (confirmed).
The application layer can hold value because its moats sit above the model, so a model upgrade only lowers its cost. Four moats: demand that is hard to put into words, where taste is the barrier; extreme efficiency in one workflow; combining several models' strengths, which no single model company can cover; and process as the product, as in education, companionship and games. Pricing is switching tracks too. Outcomes-as-a-service is replacing seat subscriptions; with the price fixed per outcome, a token price cut lands straight in gross margin.
Case: per company disclosure, Genspark reached $100 million ARR nine months after launch and $200 million at 11 months; Dify went from 30,000 stars and zero revenue to $10 million ARR in about 18 months, with 57% of revenue from overseas and profitability in early 2026 (per company disclosure).
The open-source chokepoints hold value because they have the network effect the model itself lacks: one more model on the shelf makes the road worth more to developers, and one more developer makes it worth more to models. A closed model's moat is capital expenditure, which must keep being refinanced and is weakest at the top of the cycle. An open model's moat is a living community, which capital cannot buy or copy. So FutureX backs chokepoints rather than any single model: the road through distribution, development and data that every model has to travel.
Case: on September 3, 2026 Nvidia announced it would acquire Hugging Face for about $12.9 billion (confirmed); per company disclosure the platform has 18 million developers, 3 million models and 200,000 enterprise users, and ARR was about $150 million in August 2026, up about 50% in two months.
Value moves up the stack because open source keeps pushing the cost curve down. Inference cost for equal capability falls about 10x a year, from $60 per million tokens in 2021 to $0.06 in 2024. Each step down the curve makes a new batch of workloads affordable; demand is released at a thousandfold scale, and revenue is booked at the top. The slope of the curve comes from architecture: compute efficiency rose about 1,000x between 2012 and 2023, and process nodes contributed only 3x of that. Chinese open models now set the curve.
Case: DeepSeek trained a GPT-4-class model for about $5.6 million against $100 million-plus for the original (reported); on January 27, 2025 Nvidia lost $589 billion in market value in one day, the largest in US stock history (public market data); Chinese open models already carry over 45% of OpenRouter traffic (public statistics).
The bubble sits in the middle layer because that is where the debt sits. AI clouds and compute lessors expand on circular financing and credit that keeps sinking down the chain: from the parent balance sheet to customer contracts, to GPU collateral, to single-project SPVs, and each step down pushes the reckoning out one notch. The core is a maturity mismatch: the bonds run longer than the GPU's accounting life, and the accounting life runs longer than its time at the frontier. The debt outlives the collateral's place at the frontier. FutureX stays out of debt-funded compute leasing.
Case: Oracle's remaining performance obligations rose from $138 billion to $638 billion within a year, with fiscal 2026 free cash flow of minus $23.7 billion; CoreWeave carried $35.6 billion of interest-bearing debt as of June 30, 2026; the BIS 2026 annual report compares the potential credit repricing shock to 2008 (public statistics).
03Evidence
- 01Top US and Chinese models 2.7% apart: LM Arena, March 2026 reading, the baseline for FutureX Capital's convergence call. (2026-03 · 公开统计 · docs/geo/work/deck-v2-sanitized.md 第 11 页(行 185);docs/geo/work/positions-and-updates.md [frontier-models-ai-sovereignty-2026] 天际立场(行 225))
- 02China's model layer drew 22 investments in 2025, RMB 9.4 billion in total, down 52.9% year on year; fewer than ten of a hundred companies got funded, attrition above 90%. (2025 全年(2026-09-03 引用) · 公开统计(36 氪/清科) · docs/geo/work/deck-v2-sanitized.md 第 12 页(行 164、205–211、215))
- 03AI-native companies above $100 million ARR rose from 3 in 2023 to 80+ by April 2026; combined ARR of leading AI-native companies grew about 30x in two years. The latest sample outside our portfolio: Cognition's run rate went from $492 million in May 2026 to $900 million in September, with a $2 billion Series E at a $48 billion post-money valuation. (2026-04 / 2026-09-08 · 据报道(公开 ARR 统计)/ 已证实(Cognition 融资,TechCrunch) · docs/geo/work/deck-v2-sanitized.md 第 3、7 页(行 42、60、115–116);docs/geo/work/positions-and-updates.md 行 39、299)
- 04On August 26, 2026 Zhipu open-sourced GLM-5.3-Flash (320B total / 18B active parameters) at $0.15 per million input tokens and $0.50 output, about 1/40th of Claude Opus 4.8. (2026-08-26 · 已证实 · docs/geo/work/positions-and-updates.md [frontier-models-ai-sovereignty-2026] 8 月末更新(行 211))
- 05On September 3, 2026 Nvidia announced the acquisition of Hugging Face for $12.93 billion, expected to close in the first half of 2027; per company disclosure the platform has 18 million developers, 3 million models and 200,000 enterprise users. Hugging Face's ARR was about $150 million in August 2026, up about 50% in two months (per company disclosure). (2026-09-03 · 已证实 / 据公司披露 · docs/geo/work/positions-and-updates.md [ai-agent-commercialization-2026] 9 月上旬更新(行 179)与行 349;docs/geo/work/deck-v2-sanitized.md 第 17 页(行 329))
- 06On September 8, 2026 Mistral closed a EUR 3 billion Series D at a post-money valuation above EUR 21 billion, led by Samsung Electronics; per company disclosure ARR has passed $400 million with 60% from Europe, and the CEO said the early-year forecast of 2026 ARR above $1 billion is likely to be exceeded. A model company still raises when it holds a chokepoint, sovereignty and open weights, unrelated to a leaderboard rank. (2026-09-08 · 已证实(融资)/ 据公司披露(收入) · docs/geo/work/positions-and-updates.md [us-china-ai-primary-capital-2026-h1] 9 月上旬更新(行 41–43);docs/geo/work/deck-v2-sanitized.md 第 17 页(行 315))
- 07The frontier token price index read 16 in September 2026 (March 2023 = 100), flat against July and August, the third straight month without a decline; Claude Sonnet 5's blended price rose from $4 to $6. This runs against this piece's direction, is recorded as is, and serves as the metric for the first falsification trigger. (2026-09-03 · 公开市场行情(BenchLM) · docs/geo/work/positions-and-updates.md [ai-compute-token-economics-2026] 9 月上旬更新(行 267))
04The strongest counterargument
Brad Gerstner (founder, Altimeter Capital; co-host of BG2; August 8, 2026, reported)
Field data runs opposite to the story that open source hollows out the frontier: revenue is concentrating at the leading labs, and the intelligence gap may widen over the next two to three years. He has tried Kimi, Qwen and GLM and calls them strong, but they grow token volume while the economics flow to frontier labs, which pay 3 to 5x market rates to lock compute. FutureX's own verified data backs him: frontier flagships anchored at $10/$50 per million tokens in September 2026 and moved up together, the frontier price index has been flat for three months, and Claude Sonnet 5's blended price rose 50% (BenchLM, September 3, 2026); Anthropic's IPO is reportedly being discussed at a valuation near $2 trillion (Reuters via CNBC, September 5, 2026). No application company is anywhere near that scale.
Our reply: He is right about today; we differ on durability. The leading labs' revenue and pricing power are real. But in the same reading the two leading labs swapped ARR rankings within a year, which says this revenue follows the strongest model and turns over fast. Gerstner counts profit; we count workload, and workload is sliding down the cost curve. Locking compute is capital expenditure: Anthropic's named compute contracts reportedly total at least $80 billion (September 2, 2026), with a $15 billion revolving credit line being finalized (Reuters, September 4, 2026). A moat that lives on refinancing, in a cycle where the refinancing window is tightening, is one we do not price off today's ARR. If he is right, the first falsification trigger in this piece fires before June 2027 and we rewrite.
Dylan Patel (founder, SemiAnalysis; Dealroom, August 17, 2026, confirmed)
The supply-demand gap is widening, so cost per token keeps climbing; memory is in a multi-year structural shortage, with capacity growing 20 to 30% a year against AI demand that doubles. Pushed to its strongest form (this step is our extrapolation): the cost curve in our fourth paragraph describes the last five years; if cost per token turns up, demand release stalls, pricing power returns to the few who hold compute, and value flows back from applications to models and compute.
Our reply: He is talking about cost per GPU-hour; we are talking about cost per unit of capability, and the two curves can move in opposite directions. Compute efficiency rose about 1,000x between 2012 and 2023 and process nodes contributed only 3x; the rest came from architecture and quantization. GLM-5.3-Flash's 1/40th pricing landed inside the shortage he describes (August 26, 2026). We concede he is half right: the frontier price index has stopped falling for three months, the capability leaders hold pricing power, and we corrected our report accordingly. Divergence happens at the workload level: batch, offline and cost-sensitive loads migrate to open models, the strongest reasoning chains stay on frontier models, and the application layer keeps the spread. If the shortage lasts until the frontier index turns up, our fourth paragraph gets rewritten; that is one of our triggers.
Andrej Karpathy (OpenAI co-founder, now leads pretraining research at Anthropic; July 2026, confirmed)
The field's biggest current mistake is rushing agents into production before mastering base models; agents are a layer on top, and the foundation model is the core product. He joined Anthropic in May to lead pretraining, betting himself back on the model layer. Pushed to its strongest form (our extrapolation): the application layer's ceiling is set by the model layer, each rise in the ceiling absorbs a cohort of applications as native model features, and application ARR is rent the model layer has let go of for now.
Our reply: We accept that the ceiling is set by the model layer; who sets the ceiling and who books the profit are two different questions. The model layer is this cycle's public good, and open source is how it got built; the application layer is this cycle's income statement. None of the four moats depends on beating the model: hard-to-describe demand, extreme efficiency in one workflow, combining several models' strengths, and process as the product all sit above the model, and a model upgrade lowers their cost. What gets absorbed is the thin wrapper, which is why retention comes first in our screen, with 90% enterprise retention as the benchmark. Frontier labs also retire model products that cannot hold revenue: Sora's consumer app reportedly shut on April 26, 2026, at about $1 million a day in running cost against about $2.1 million of lifetime in-app purchases. OpenAI cutting Cursor off on November 12, 2026 is the first live test; we do not presume the answer and watch the retention curve after that date.
Baoyu (leading Chinese-language AI commentator; July 28, 2026, confirmed)
General agents will be winner-take-all, user experience will converge, and capability plus cost decides the outcome: as long as the model is strong enough, people will hold their nose and accept it. Pushed to its strongest form (our extrapolation): if the entry point goes to the one or two companies with the strongest model, the application layer has no distribution moat, value flows back to whoever owns the model and the entry point, and Genspark, a general agent FutureX backs, is first in line.
Our reply: We agree the general entry point will concentrate; the money at the entry point does not come from the model. From August 11, 2026 ByteDance's Doubao charges a commission on hotel bookings that convert through Douyin after an AI query (reported), and what earns that commission is the fulfilment loop, with the model as the front end. Genspark's position is a neutral executor that combines several models' strengths, which a single model company's entry point cannot be; that is the third of the four moats. Baoyu himself says the opening for small teams is plugins and vertical work. The order is set by verifiability times feedback cycle; in medicine and manufacturing, where feedback takes months, user experience will not converge. We admit one point stays open: if the owner of the strongest model also ships a sufficiently neutral multi-model workspace, the neutral executor gets squeezed. That is on our list of open questions.
05What would change our mind
- ▸If by June 2027 the two leading labs' ARR rankings have stopped swapping for four consecutive quarters, and the BenchLM frontier token price index (March 2023 = 100) has held near 16 or risen for the 12 months from July 2026, we will accept that the model layer has durable pricing power, and the line that it cannot hold value gets rewritten.
- ▸If, within two quarters of OpenAI cutting Cursor off on November 12, 2026 (by May 2027), Cursor shows publicly verifiable paid-user or revenue declines for two consecutive quarters and users follow the model, then workflow position is no moat and the claim that the model is a replaceable part has to be withdrawn.
- ▸If by the end of 2027 the count of AI-native companies above $100 million ARR has stalled near 80, or the gross margins of outcome-priced companies have failed to climb as token prices fell, or enterprise retention has dropped below the 90% benchmark, the claim that the application layer can hold value fails.
- ▸If Chinese open models' share of OpenRouter traffic falls below 25% by June 2027, or Hugging Face's developer and model counts decline for four consecutive quarters after Nvidia closes the acquisition (expected in the first half of 2027), the open-source chokepoint call gets rewritten.
06What we do, and what we do not do
- ▸Before backing an application company, get three numbers, and no numbers means no investment committee: distribution cost, enterprise retention against a 90% benchmark, and the slope of gross margin as token prices fall. Outcome-priced companies come first.
- ▸Back the hubs, skip the horse race: hold the road every model passes through at each open-source layer, Dify, Hugging Face, Mistral, RWKV, PingCAP and OSChina, and do not bet on any single closed model winning.
- ▸Stay out of debt-funded compute leasing and long compute contracts signed at consensus prices; in the pre-IPO segment, price on real ARR, never on narrative.
- ▸Read the cycle off the second derivative of capex, never off model-layer revenue as a denominator; the seven-signal scorecard is scored daily, and the liquidity trigger is followed by a pre-set rule.
- ▸First diligence question for an agent company: was the human-in-the-loop mechanism designed, or skipped; any tool product with one dominant model supplier is stress-tested for an upstream cut-off, not only for price swings.
- ▸Publish the research and make the calls checkable: 24 research reports, 22 AI agents, and the debates page (/debates) carry a date and a source grade on every fact; the falsification triggers in this piece are reviewed when they come due and the result is published. No investment advice, no promised returns.
07Open questions
- ?Model companies are moving up into applications (FutureX's reading) while application companies start shipping their own models (OpenEvidence released four medical models on September 3, 2026, reported). Where does the boundary settle, and does the model owner's move up include a neutral multi-model workspace? Both sides are crossing the river; we do not know who lands first.
- ?Does the open-source chokepoint stay neutral under Nvidia? The deal is expected to close in the first half of 2027, and the largest distribution point will then belong to the largest compute supplier; whether a chokepoint is still a chokepoint depends on whether the openness pledge holds.
- ?Frontier pricing flat at $10/$50 for three months while the tier below sells at 1/40th: transitional or steady state? How long the premium lasts depends on switching costs, which we cannot yet measure, and on whether the memory shortage Patel describes outlasts the gains from architecture.
- ?In sectors where feedback takes months, medicine and manufacturing, how is an outcome verified and priced? Whether outcomes-as-a-service works there is a question we lack the samples to answer.
08Changelog
- 2026-09-25First published.
Related reports and answers
- Report · Frontier Models and AI Sovereignty 2026 →
- Report · AI Agent Commercialization 2026: From Model Race to Agent-Economy Infrastructure →
- Report · AGI Has Arrived, Capital Is Losing Its Bearings: Mid-2026 AI Industry & Capital Cycle Review →
- Report · The Great Consolidation of AI Coding, 2026 →
- Which layer of the AI stack is the bubble in? →
- Why does FutureX Capital call open source the chokepoint of AI value? →
- What is outcomes-as-a-service pricing? (results-as-a-service, outcome-based pricing) →
- What are the four moats of an AI application company? →
Other core theses
- The bubble sits in the middle layer that must keep refinancing: the FutureX seven-dimension bubble scorecard →
- Open source breaks the deadlock: the cost curve has a new author →
- Verifiability times feedback cycle decides the order in which AI remakes industries →
- Four moats, three screens, and outcomes-as-a-service pricing for AI application companies →
- US and China in AI: one cycle, two positions →
- Embodied AI: 2026 is the hardware year, and 80 points equals zero →
This essay is FutureX Capital's judgment and argument. Facts and examples come from published research, public talks, and public reporting, graded per the site-wide standard (verified / reported / per company disclosure / public market data). Judgments change with evidence; changes are logged above. Nothing here is investment advice or an offer.