Core theses · Updated 2026-09-25
Verifiability times feedback cycle decides the order in which AI remakes industries
01The claim
The order in which AI remakes industries equals verifiability times feedback cycle. Where a result checks itself in seconds, AI moves first; where a human stays in the loop and feedback takes months, it moves last. Each sector gets a 12-to-18-month window after capability crosses its threshold and before enterprise budgets arrive.
02Mechanism: why it happens
Verifiability decides where capability gets proven first, and the threshold is proven capability. A model improves by being told it is wrong; a buyer pays only after seeing the result. Where being told is free, both happen at once, and capability gets priced there first. Code passes or fails its tests in seconds and content shows clicks the same day, so those two sectors sit at the front. A diagnosis waits for the course of an illness. Being told is slow and expensive there, and however strong the model, nobody can cheaply prove it is good enough.
On September 8, 2026 Cognition announced a $2 billion Series E at a $48 billion post-money valuation (confirmed, Bloomberg and Reuters); annualized revenue rose from $492 million in May to nearly $900 million (reported, Investing.com). Software entered its window in 2024 and is still accelerating two years later.
Feedback cycle decides how fast that proof repeats; in the formula it counts as a frequency, so a shorter cycle means a larger multiplier. With feedback in seconds a team runs tens of thousands of trials a year; in months, a handful. Three orders of magnitude in trials become three orders of magnitude in compounding. This is why knowledge work trails software by only a year: a report needs a person to judge it, so verifiability is weaker than for code, but the person judges the same day and feedback stays fast. Fast feedback compensates for weak verification.
Genspark reached $200 million ARR eleven months after launch: nine months from zero to $100 million, then a doubling in two (per company disclosure as of August 2026, unaudited; cited in the September 3, 2026 Xiamen keynote). In the SaaS era the same distance took five to seven years.
At the slow end capability is already sufficient; what jams is confirmation. In medicine, finance and manufacturing a person or a regulator confirms the result months later, and the organization lacks a rule for when a human must look up. Until that mechanism exists, pilots stay outside the profit and loss statement, and extra capability is only extra cost. This is where FutureX Capital's first diligence question for any agent company comes from: was the human-in-the-loop mechanism designed, or skipped.
On September 8, 2026 TechTarget cited a new study finding that fewer than 1% of FDA-cleared AI medical devices have validated clinical benefit (reported); the same week, STAT reported on September 3 that the FDA launched a pilot letting generative AI devices reach patients before formal authorization (reported). Access is loosening; the evidence gap now has a number.
The window runs 12 to 18 months because enterprise budgets run on annual cycles. In the quarter when capability crosses the threshold, only founders and early users can see it. Budgets arrive at scale in the next annual cycle, incumbents then ship the feature as a default, and regulators fill in behind. The moment the three move together, the window closes and prices return to consensus. So we date each sector at the threshold; the day budgets arrive does not count.
At its 20,000-person user meeting on August 17 to 20, 2026, Epic put AI on a platform footing, with Agent Factory offering more than 120 prebuilt agents ahead of a 2027 general release (confirmed, TechTarget and Advisory Board). The incumbent is making AI a default feature of the EHR; healthcare's window is arriving on its 2026-to-2027 date.
Robotics sits late in the order for the same reason: verification. Training data can come from human video; verification can only happen on a real machine, one trial after another. Whether a grasp worked is known only after the machine runs it, and the loop closes in days. Slow verification means no technical route can yet prove itself, while valuations have already converged. So FutureX Capital's call is that 2026 is the hardware year for humanoids and the intelligence year has not arrived; 80 points equals zero. We back dexterous hands and data and stay out of the full-body arms race.
Agility Robotics' S-4, filed with the US SEC on September 4, 2026, shows 2025 net sales of $1,781,967 against a net loss of $138,086,332, a loss about 77 times revenue (reported, via Humanoids Daily on September 7). After years of commercial pilots, deployments have yet to convert into revenue.
03Evidence
- 01The length of tasks AI can complete on its own doubled every 7 months from 2019 to 2025 and every 4.3 months after 2023. The curve is measured on tasks that can be verified automatically: where verification is cheap, capability is read first. (2026-01 · 已证实 · METR Time Horizon 1.1(2026-01),引自 docs/geo/work/deck-v2-sanitized.md 第 3 页)
- 02Duolingo's Q2 results on August 5, 2026: revenue $298.5 million (+18%), DAU 58.7 million (+23%), and the inference cost of one AI video call down from $0.30 to under $0.01. Education measures results daily and its costs fell by the quarter; it entered its window in 2025. (2026-08-05 · 已证实 · docs/geo/work/reports-context.json · ai-education-2026 keyFindings(公司财报))
- 03On September 8, 2026 Cognition announced a $2 billion Series E at a $48 billion post-money valuation, close to double the $26 billion of May; annualized revenue rose from $492 million in May to nearly $900 million. Software's window opened in 2024 and the revenue curve has not slowed. (2026-09-08 · 已证实(融资);据报道(ARR) · docs/geo/work/positions-and-updates.md · [us-china-ai-primary-capital-2026-h1] 与 [agi-capital-cycle-2026] 9 月上旬更新(彭博社、路透社、TechCrunch、Investing.com))
- 04Genspark reached $200 million ARR eleven months after launch, nine months from zero to $100 million and two more to double. Knowledge work measures its result the same day, and the revenue curve tracks the feedback curve. (2026-08 · 据公司披露 · docs/geo/work/deck-v2-sanitized.md 第 5 页、第 17 页(截至 2026 年 8 月,未经独立审计))
- 05MIT NANDA's 2025 report: 95% of enterprise generative-AI pilots have produced no measurable P&L impact. Sufficient capability and missing confirmation coexist; that is the normal state at the slow end of the order. (2025 · 已证实 · docs/geo/work/deck-v2-sanitized.md 第 4 页(MIT NANDA,2025))
- 06On September 3 the FDA launched a pilot admitting four generative AI devices and letting them reach patients before formal authorization (STAT); on September 8 TechTarget cited a new study finding fewer than 1% of FDA-cleared AI devices have validated clinical benefit. Access is outrunning verification. (2026-09-03 / 2026-09-08 · 据报道 · docs/geo/work/positions-and-updates.md · [ai-healthcare-inflection-2026] 与 [ai-mental-health-2026] 9 月上旬更新(STAT 9 月 3 日;TechTarget 9 月 8 日))
- 07Agility Robotics' S-4 filed with the SEC on September 4, 2026: 2025 net sales of $1,781,967, net loss of $138,086,332, a loss about 77 times revenue; its $300 million multi-year order comes from a related-party customer. (2026-09-04 · 据报道 · docs/geo/work/positions-and-updates.md · [embodied-ai-humanoid-robots-2026] 9 月上旬更新(Humanoids Daily 9 月 7 日转述))
04The strongest counterargument
Aaron Levie (Co-founder and CEO, Box), July 16, 2026, reported
After rounds with enterprise IT leaders, his conclusion: the bottleneck is change management and data readiness, and the models are already good enough; playbooks are narrowing from try-everything to roughly ten of the highest-value workflows. The strongest form: capability has crossed the threshold in most industries already, so the order is set by the buyer's capacity to change and by budget cycles. Verifiability is a property of the task; the buyer decides. The same verifiable task stalls in pilot at a company whose data is not ready.
Our reply: He is half right, and we wrote that half into the mechanism: in slow sectors capability is often already sufficient. What our formula ranks is the cost of proving it sufficient, and change management and data readiness are what that cost is spent on. A workflow whose result can be counted is cheap to prove, which is how it made his list of ten; one whose result cannot be counted never reaches the profit and loss statement, however much change management it gets. Inside a sector, the organization decides who deploys first; across sectors, verifiability prices the cost of proof. The two variables sit at different levels, so we rank sectors with one and pick companies with the other. If his view holds across sectors as well, see the second falsifier.
New York City Public Schools' procurement decision, September 2, 2026 (Chalkbeat New York, confirmed)
Eight days before the school year, New York City announced that pre-K through grade 8 would stop using student-facing generative AI for 2026-27; AI features in 38 existing edtech contracts were switched off, the pause runs a year, and about 600,000 students, roughly two thirds of the city's enrollment, are affected. Education sits at 2025 on our timetable. The strongest form: in sectors run by public procurement, one buyer's memo can shut the window, verifiability has no say, and the procurement cycle is the real timetable.
Our reply: The buyer can delay; we accept that and count it as error. For K-12 and hospital systems a purchaser can hold the window shut for a year, and our dates carry that error bar. Something else happened the same day: McGraw Hill acquired Teachally, a five-person teacher-tools team (September 2, 2026, company disclosure). Procurement changed who the customer is: the student side closed while the teacher side was buying. Our dates mark the moment capability crosses the threshold; a budget delay lengthens the window and does not cancel it. If the pause spreads to most large districts, see the third falsifier.
The US FDA's pre-authorization pilot, via STAT, September 3, 2026 (reported)
The FDA launched a pilot admitting four generative AI devices and letting companies put them in front of patients before marketing authorization, after publishing a discussion paper on such devices on August 18 with comments due October 19 (confirmed). The strongest form: regulators are actively shortening medicine's cycle, access is outrunning evidence, and medicine may change earlier than the 2026-to-2027 date we wrote. The regulator's calendar sets the order; verifiability merely follows.
Our reply: We grant that regulation can move the date. A second number from the same week shows what it moves: fewer than 1% of FDA-cleared AI devices have validated clinical benefit (TechTarget, September 8, reported). Looser access puts products in front of patients sooner; it does not get results judged sooner. Our date marks when vertical agents enter the profit and loss statement, and a ticket in does not count. If a regulator compresses the verification cycle itself, say by making clinical-benefit validation an automated premarket step, that is the feedback term in the formula shrinking: medicine moves earlier and the formula stands.
Jim Fan (Director of Robotics Research, NVIDIA; head of the GEAR team), May 8, 2026, confirmed
He asked the room to hold a moment of silence for teleoperation: one robot caps at 24 hours a day and yields about 3, which cannot scale. The replacement is egocentric human video plus world action models, with supply that can reach 10 million hours a year, and a physical Turing test passed within two to three years. The strongest form: the robotics bottleneck is data supply, data supply is decoupling from physical deployment and growing at internet scale, and once the reason for placing robotics late disappears, robotics jumps the queue.
Our reply: We accept the point on data supply and changed the mechanism accordingly: training data can come from video. Figure's crowdsourced platform Index had collected 16 million human videos across 108 countries and paid contributors $15 million by August 26; on September 3 Figure signed a $3.5 billion first-phase compute agreement with Nscale, with the first GPUs due online in the second half of 2027 (reported, eWeek, September 4). That solves the learning end. The proving end is unchanged: whether a move learned from video works is known only after a real machine runs it, and that loop still closes in days. His two-to-three-year timetable lands after 2028, which matches our call that 2026 is the hardware year and the intelligence year has not arrived. So we back the very route he describes: data (World Engine) and dexterous hands (Dexmate).
05What would change our mind
- ▸If, before June 2027, a company in healthcare, finance or manufacturing goes from zero to $100 million ARR inside 12 months while its results are still confirmed by people or regulators on a monthly cycle, the formula is wrong: feedback cycle did not constrain revenue, and we will rewrite both the dates and the ranking mechanism for those three sectors.
- ▸If a follow-up to the MIT NANDA survey shows during 2027 that the share of enterprise generative-AI pilots with no measurable P&L impact has fallen from 95% to below 50%, and those enterprises got there without outcome pricing or automated result verification, on change management alone, then Levie is right and we are wrong: organizational capacity is independent of verifiability and should be the first ranking variable.
- ▸If by September 2027 five or more of the ten largest US school districts have followed New York in pausing student-facing generative AI, and K-12 AI procurement spending is down year on year, then the procurement cycle overrode verifiability in education. We will move education off the 2025 row and list public procurement as a third variable outside the formula.
- ▸If by June 2027 GGII's count of 2026 humanoid shipments in China comes in below half its 62,500 forecast, and the leading full-body makers with disclosed financials still report under $10 million of 2026 revenue, then the robotics and logistics row was dated too early and should move back at least a year. We will change the date in public. The reverse holds too: if before then a full-body maker turns moves learned from human video into weekly on-machine iteration and audited revenue, robotics moves earlier.
06What we do, and what we do not do
- ▸We date each sector at the moment capability crosses the threshold and enter inside the window. Once incumbents ship the feature as a default and budgets arrive at consensus prices, we do not chase. On pre-IPO rounds at 50 to 100 times sales, we stay cautious.
- ▸In slow-feedback sectors, healthcare, finance and manufacturing, we back only companies that shorten the feedback cycle: outcome pricing, result verification built into the product, ownership of the data. A company that only plugs a model into an existing process, we pass on. First diligence question: was the human-in-the-loop mechanism designed, or skipped.
- ▸In robotics we back dexterous hands and data only, such as Dexmate and World Engine, and stay out of the full-body arms race. 2026 is the hardware year; the intelligence year has not arrived; 80 points equals zero.
- ▸We screen companies on three things: distribution, retention, and whether gross margin climbs as token prices fall. For open-source projects we add the path from stars to paid revenue; Dify took about 18 months from 30,000 stars and zero revenue to $10 million ARR (per company disclosure). FutureX Capital has screened 1,500+ AI projects with the same three tests.
- ▸What we do not do: we do not bet on which model wins; we hold the hubs (Hugging Face, Mistral, Dify) that every winning model has to pass through, and we do not touch debt-financed compute rental. For enterprise readers, one thing: put budget first into applications whose outcomes can be counted, see /enterprise.
- ▸We publish the timetable with its falsifiers and change dates in public when a trigger fires. The full deck is at /reports/open-source-breakthrough, the Q&A at /answers/twelve-to-eighteen-month-window and /answers/which-industries-ai-changes-first, and the opposing views at /debates.
07Open questions
- ?Does publishing a timetable of windows push prices at the tail of each window to consensus? On August 27 SoftBank moved to buy control of 1X at about $6 billion, below the roughly $10 billion it sought the year before (confirmed, The Information); Unitree's September 2 close was 50.4% below its opening price on debut (public market data). Both fall inside the robotics window we dated. Whether the timetable itself is one cause of the crowding, we cannot separate out.
- ?When verification is done by another model rather than a test suite or a person, does the sector count as fast or slow? Content and knowledge work already lean on model judges. Whether they inherit the judge's speed or the judge's error rate, we do not know.
- ?Is finance one sector? Trading gets feedback in days, credit in quarters, and we wrote the whole block at 2026 to 2027. It may split into two rows more than a year apart.
- ?After the window closes, can the company that entered inside it hold its position? Epic's 2027 general release of Agent Factory is the test in healthcare; OpenAI announced on August 28, 2026 that it will stop supplying models to Cursor from November 12, and Cursor's retention curve after the cutoff is software's first answer sheet.
08Changelog
- 2026-09-25First published.
Related reports and answers
- Report · AI Agent Commercialization 2026: From Model Race to Agent-Economy Infrastructure →
- Report · The AI Healthcare Investment Inflection 2026 →
- Report · Embodied AI & Humanoid Robots Investment 2026 (Mid-Year Refresh) →
- Report · AI Education 2026: Duolingo's 23% DAU Growth, Chegg's 99% Collapse, and Who Captures Value in the Chain →
- Which industries will AI change first? →
- What is the 12-to-18-month window in AI sectors? →
- Why do AI agents stall inside enterprises? →
- What is outcomes-as-a-service pricing? (results-as-a-service, outcome-based pricing) →
Other core theses
- Value lands in the application layer and at the open-source chokepoints; the model layer cannot hold it →
- The bubble sits in the middle layer that must keep refinancing: the FutureX seven-dimension bubble scorecard →
- Open source breaks the deadlock: the cost curve has a new author →
- Four moats, three screens, and outcomes-as-a-service pricing for AI application companies →
- US and China in AI: one cycle, two positions →
- Embodied AI: 2026 is the hardware year, and 80 points equals zero →
This essay is FutureX Capital's judgment and argument. Facts and examples come from published research, public talks, and public reporting, graded per the site-wide standard (verified / reported / per company disclosure / public market data). Judgments change with evidence; changes are logged above. Nothing here is investment advice or an offer.