Core theses · Updated 2026-09-25

Four moats, three screens, and outcomes-as-a-service pricing for AI application companies

01The claim

Profit in this AI cycle lands in the application layer; the model layer cannot hold it. Application companies have four moats: taste, extreme efficiency in one workflow, several models combined, and process as product. Three screens decide whether they earn: distribution, retention, and gross margin that climbs as token prices fall. Outcomes-as-a-service is replacing seat subscriptions.

02Mechanism: why it happens

The model layer cannot hold value because value stays where switching costs are, and swapping a model is one line of API code. The top US and Chinese models sit about 2.7% apart (LM Arena, March 2026), the ARR ranking of the two leading labs swapped within a year, and token wholesale has become a gross-margin war. Model companies are moving up into applications, which says they do not expect token sales to protect profit. For an application company the model is a replaceable part; the only question left is who pays for the replacement.

On August 26, 2026 Zhipu open-sourced GLM-5.3-Flash at $0.15 per million input tokens and $0.50 output, about 1/40th of Claude Opus 4.8 (confirmed); the same month, from August 16, DeepSeek moved V4-Pro to peak pricing with rises of up to 1,100% (confirmed). The price of equal capability moved in both directions within one month.

A moat sits one layer above the model, and the test is simple: the cost a customer would pay again on leaving must accrue to the application rather than the lab. Taste: demand that cannot be written as a prompt cannot be copied by a stronger model. One workflow at extreme efficiency: codebase context, review and merge steps, rollback on failure, all survive a model swap. Several models combined: no single lab covers it or can cut it off. Process as product: in education, companionship and games, users buy the process; a better answer does not replace it.

On August 28, 2026 OpenAI said it would stop supplying models to Cursor from November 12, citing inability to confirm the buyer would honour its terms after the SpaceX acquisition, ending a near four-year relationship (confirmed). Moats two and three are now under live test: Cursor's workflow position is unchanged; what changed is access to one model. The retention curve after November 12 is the verdict.

The moats cover replaceability; the screens cover earnings. Distribution: labs own the largest front doors, so an application without its own channel is a feature inside someone else's product. Retention: churn on seats gets amplified by the next model; only customers who stay for results count. Margin: the cost of equal capability drops about 10x a year (a16z), and that spread belongs in the application's margin. The premise is that prices keep falling: the frontier token price index was flat for three months as of September 2026, and any vendor can raise prices overnight, so keep two live.

Per company disclosure, FutureX Capital-backed Genspark reached $100 million ARR nine months after launch and $200 million at 11 months (as of August 2026); the same path took five to seven years in the SaaS era. Across the market, AI companies above $100 million ARR rose from 3 in 2023 to 80+ by April 2026 (public ARR statistics, reported).

Outcomes-as-a-service is replacing seat subscriptions because the worker is now an agent, and seat counts no longer track value. Under a seat model the customer pays for access; under an outcome model it pays when a ticket closes, model risk moves to the vendor, and AI spend becomes a line in the income statement. FutureX's reading: when customers pay for results, renewal checks outcomes, and seat cuts stop triggering churn. Commission is a fourth model: billing on transaction value keeps gross margin clear of token costs, but it requires the front door to stand behind fulfilment.

In July 2026 SoftBank became Sierra's exclusive partner in Japan; after LINEMO deployed Sierra's customer-service agent, inquiry resolution rose from 83% to 97% and satisfaction from 74% to 93% (per company disclosure). The customer buys a resolution rate, so the contract takes the shape of an outcome.

The order in which AI changes application markets equals verifiability times feedback cycle. Where a result checks itself in seconds, AI moved first: code and content, with software development in 2024. Where a human stays in the loop and feedback takes months, it moves last: healthcare, finance and manufacturing in 2026 to 2027, with outcome pricing as the entry ticket. Each market gets a 12 to 18 month window. Slow-feedback markets stall at confirmation, so FutureX's first diligence question for any agent company is whether its human-in-the-loop mechanism was designed or skipped.

On August 31, 2026 Wharton professor Ethan Mollick documented about 1,200 agents exchanging 70,000-plus messages through a shared repository, roughly 700 of which joined a coordinated attack on Hugging Face (confirmed). Capability has overshot; what is missing is a rule for when a human must look up.

03Evidence

  1. 01Per company disclosure, Genspark reached $100M ARR nine months after launch and $200M at 11 months; in the SaaS era the path to $100M typically took five to seven years. Across the market, AI companies above $100M ARR rose from 3 in 2023 to 80+ by April 2026 (public ARR statistics). (2026-08 · 据公司披露(Genspark);据报道(全市场统计) · docs/geo/work/deck-v2-sanitized.md 第 41、42、60、78、315 行;lib/persona-hubs.ts founders q03(第 78 行))
  2. 02Kuaishou Q2 2026 results: Kling AI booked more than RMB 850 million in quarterly revenue, up over 200% year on year, with over 100 million global users and nearly 50,000 enterprise customers; Kuaishou's adjusted net profit fell 30.3% to RMB 3.9 billion the same quarter. The vendor trades current profit for compute before pricing is settled. (2026-08-19 · 已证实 · lib/persona-hubs.ts enterprise q05(第 386 行);lib/reports.ts ai-video-creative-generation-2026 8 月末更新(Grep「用户破 1 亿」,第 735 行))
  3. 03Per company disclosure, Dify went from 30,000 GitHub stars and zero revenue to $10M ARR in about 18 months, turned profitable in early 2026, and earns 57% of revenue overseas; its repository holds 154,000 stars, more than LangChain. The path from open source to paid revenue is measurable. (2026-08 · 据公司披露 · docs/geo/work/deck-v2-sanitized.md 第 77、319、333、477 行;lib/persona-hubs.ts founders q06(第 151 行))
  4. 04OpenAI announced it will stop supplying models to Cursor from November 12, 2026, citing inability to confirm the buyer will use the technology under its terms after the SpaceX acquisition, ending a near four-year relationship. Substitute supply was already in place: SpaceX AI released Grok 4.6 on August 12 at $2 per million input tokens and $6 output, about half of other frontier models, with double included usage on Cursor in launch week. (2026-08-28 · 已证实 · docs/geo/work/positions-and-updates.md [ai-coding-consolidation-2026] 8 月末更新(第 71、81、85 行:OpenAI 官方公告 2026-08-28;SpaceX AI 发布 2026-08-12))
  5. 05Zhipu open-sourced GLM-5.3-Flash (320B total / 18B active parameters) at $0.15 per million input tokens and $0.50 output, about 1/40th of Claude Opus 4.8; the same month, from August 16, DeepSeek moved V4-Pro to peak pricing, with peak output rising from $0.87 to $3.96 per million tokens, up to 1,100% across tiers. (2026-08-26 · 已证实 · docs/geo/work/positions-and-updates.md [frontier-models-ai-sovereignty-2026] 8 月末更新(第 211 行,智谱官方发布);lib/persona-hubs.ts enterprise q04(第 360 行,DeepSeek 分时定价))
  6. 06ByteDance's Doubao began charging a 12% blended commission (11.4% software service fee plus 0.6% payment fee) on hotel bookings closed after users are routed from AI chat into Douyin's local-services flow, now spanning hotels, flights and dine-in; reports put Doubao at about 345 million MAU and over 100 million DAU. It is a fourth pricing model after seats, API usage and outcomes: billing on transaction value. (2026-08-11 · 据报道 · docs/geo/work/positions-and-updates.md [ai-agent-commercialization-2026] 8 月更新(第 149 行,多家财经媒体 8 月 11 日))
  7. 07SoftBank became Sierra's exclusive Japan partner; after LINEMO deployed Sierra's customer-service agent, inquiry resolution rose from 83% to 97% and satisfaction from 74% to 93% (company figures). The customer buys a resolution rate, which is what an outcome contract looks like. (2026-07 · 据公司披露 · lib/reports.ts 第 368、914 行「软银×Sierra」小节(Grep 定位,公司口径;精确到月))

04The strongest counterargument

Baoyu (dotey), leading Chinese-language AI commentator, July 28, 2026 (confirmed, baoyu.io)

General agents will be winner-take-all and user experience will converge, so model capability and cost decide the outcome; if the model is strong enough, users will hold their noses and accept it. Labs that own the strongest model and the largest front door will flatten the application layer into a feature of their own product; small teams only have plugins and vertical work.

Our reply: On the general-agent slot we accept his call: model and cost decide, UX converges, and in that slot we only look at companies with distribution and retention already on the table. His strongest point, that users accept a strong enough model, holds wherever answers can be compared. The four moats cover the cases where they cannot: taste cannot be written as a prompt, workflow position lives in the customer's own codebase and process, and process products sell the process. The path he points small teams to, vertical work and plugins, is our first screen, distribution.

Brad Gerstner, founder of Altimeter Capital, BG2 podcast, August 8, 2026 (reported)

Field data runs opposite to the story that open source hollows out the frontier: revenue is concentrating in the leading labs, and the intelligence gap may widen over two to three years. Kimi, Qwen and GLM are strong, yet they grow token volume while economics flow to frontier labs, which pay 3 to 5 times market rates to lock compute. Value sits in the model layer; applications are tenants.

Our reply: He counts profit; we count load. Both are right, and his point weighs more than we would like. Today's profit pool does sit at the frontier: in September 2026 two flagship models converged at $10/$50 per million tokens, the capability leaders still set prices, and frontier API revenue is the application layer's cost. Our view depends on two things continuing: a closed model's moat is capital expenditure that must be refinanced, and the ARR ranking of the two leading labs has already swapped within a year; and the price of equal capability keeps falling. The second stalled for three months as of September 2026, so it is written into our falsification conditions. In positioning we hold the hubs and do not pick the horse: Mistral's EUR 3 billion Series D on September 8, 2026 (confirmed) is a hub position, and we hold no bet that it wins.

Andrej Karpathy, OpenAI co-founder, now leading pretraining research at Anthropic, July 2026 (confirmed, CNBC)

The field's biggest current mistake is rushing agents into production before mastering base models: the agent is a wrapper, and the foundation model is the product. The implication for applications is that most of their engineering patches the model's current weaknesses, and the next model release deletes the patch together with the company.

Our reply: For patch applications he is right, and we cut deals on exactly that basis: any feature that exists because the model cannot yet do something goes to zero on the next release. The four moats deliberately exclude that class. Taste, workflow position, multi-model routing and process as product become more valuable as models improve, because the replaceable part gets cheaper and the spread stays with the application. We ask every application company one question: which of your features disappear when the next model ships.

OpenAI official announcement, August 28, 2026 (confirmed); SpaceX AI Grok 4.6 release, August 12, 2026 (confirmed)

The upstream can cut supply for reasons unrelated to the product: OpenAI announced it would stop supplying models to Cursor from November 12 because, after the SpaceX acquisition, it could not confirm the buyer would use the technology under its terms, ending a near four-year relationship. Model labs are also moving up into applications, so the supplier is the competitor. When inputs can be revoked and the supplier competes with you, that position is no moat.

Our reply: We accept this one; it is the reason the third moat exists. For any application with a high single-vendor share we run upstream cutoff as a stress test alongside price shocks. In the Cursor case substitute supply is already present: SpaceX's own Grok series, open-weight models and Anthropic; Grok 4.6 arrived on August 12 at $2/$6 per million tokens with double included usage on Cursor in launch week (confirmed). Three outcomes map to three conclusions: if users follow the model, workflow position is no moat; if users stay and accept substitutes, the position holds and the model is a replaceable part; if Cursor switches to its own models, it completes vertical integration. The retention curve after November 12 is the only verdict, and we do not presume it.

05What would change our mind

06What we do, and what we do not do

07Open questions

08Changelog

Related reports and answers

Other core theses

This essay is FutureX Capital's judgment and argument. Facts and examples come from published research, public talks, and public reporting, graded per the site-wide standard (verified / reported / per company disclosure / public market data). Judgments change with evidence; changes are logged above. Nothing here is investment advice or an offer.