This note is mean to put forward a multi year view of how compute supply should play out. There are far too many how are far too emotional about their favorite neocloud stock(s) and are oblivious to the capital cycle and industry structure. For the sake of this note, we leave out the scenario where AI compute demand falls short of expectations for a meaningful period, as this would result in a far left tail event wiping equity of many and then consolidation at great prices.
THIS IS A 30-MINUTE PLUS READ… SO IF YOU DON’T HAVE THE TIME, SKIM THE TITLES AND BOLD TEXT AND CHARTS. WELCOME FEEDBACK AND DEBATE.
My argument is that the current market structure is a manufactured disequilibrium that will decay overtime. The path there is important and will become more clear likely in the coming months as we understand whether Astra-class models and agents are the next OpenClaw-style inflection in demand we saw in early 2026. Regardless, the path is toward consolidation for a long list of reasons and neoclouds would be well-served to monetize sooner rather than later as we roll into a very favorable pricing environment in 2027.
The endgame is concentration. Compute should be owned mostly by the parties closest to demand, with the lowest cost of capital, and without the fragmented merchant tier between them. Owners of scarce, hard-to-replicate physical assets (permitted, powered, interconnected land backed by long-dated investment-grade offtake) are acquisition targets that clear at a strategic premium, not fire-sale prices, but only if the strike deals in a strong market environment. Pure GPU-rental operators carrying 10-15 year lease liabilities against a 2-5 year pricing window are the equity-wipe candidates if they do not execute. Nebius is a category error to lump - it owns capacity and runs a software stack and has marquee partnerships, which is why the market already pays it a different multiple. The highest-confidence three-year expressions are to own the integrated demand-owners who capture the rent, weighted toward those that own competitive silicon rather than merely cheap funding, and to rent the upstream chokepoints that get paid regardless of which neocloud wins. The merchant tier itself is event-driven, not compounding, and should be underwritten to an exit rather than to a terminal multiple.
Structure… few buyers and many sellers
Frontier compute demand is set by roughly seven entities vs. 30+ neoclouds and converted-miner landlords, all leveraged, all buying from the same chip vendor, all chasing the same handful of contracts. When concentrated buyers face a fragmented, capital-intensive supplier base, price-setting power sits with the buyers the moment scarcity relaxes. That is an oligopsony.
Demand
Forur captive builders plus three merchant labs matter. Everyone else, Mistral, Cohere, SSI, the inference resellers, the enterprise long tail, is small, lower-credit, and price-taking. Annual capex frames the scale as big four hyperscalers guided to roughly $720 to 745B for calendar 2026, up from about $410B in 2025. Add Oracle at roughly $70B for FY27 and the top of the demand stack is near $835B a year. On the merchant-lab side the flows are stocks not annual runs, but they are enormous: Anthropic committed north of $275B of cloud compute across five deals in 2026 alone; the Oracle to OpenAI arrangement is a similar order of magnitude.
Those seven are already integrating
The seven buyers are not passive tenants… each builds its own DCs... 6/7 have custom silicon in production or in tape-out... 3 of them run enterprise software and commercial security businesses that already sell the trust layer a lab needs.
Thomas Kurian, who runs Google Cloud, described the end state without meaning to: some labs use Google’s TPUs and Gemini, others use the TPUs and then buy Google’s cybersecurity protection for their models. That is the full stack sold in pieces, by the party that owns all the pieces.
The neocloud owns none of these columns and simply will not be able to compete here over time.
Next demand names uncertain
If the buyer set expands, it does not expand into a long tail of startups. It expands into a small number of entities with balance sheets and proprietary workloads. House ranking, in order of expected signable external demand:
None of these is a durable neocloud customer. Apple would build and own. Sovereigns are buying their own national infrastructure. Quant finance is the one genuinely additive merchant customer, and it is small relative to the labs. The buyer set widens slightly and gets richer, but it does not get more fragmented, which is what the merchant tier would need.
Weighing both sides
Two defects of the neocloud model
First, the merchant tier an income-statement problem via a permanent cost-to-serve disadvantage. Second, a balance-sheet problem with a liability that outlives the revenue that services it.
Unit economics currently masked
Standalone merchant compute carries a structural cost-to-serve disadvantage against the integrated owners on four counts, none cyclical:
Cost of capital: hyperscaler funds near its cost of cash and a neocloud at high-yield or GPU-collateralized rates… this delta compounds rapidly over the years
Dual-use flexing: integrated owner moves capacity between internal and external work vs. merchant binary idle risk
Captive silicon: TPU and Trainium cut inference cost per token below a merchant Nvidia fleet
Subsidy capacity: search, ads, cloud and enterprise software fund compute at deliberately thin margins for lock-in
None of this matters today, or until demand meaningfully breaks, because the merchant tier delivers the single scarcest thing in the industry… energized capacity now. Buyers whose own power queues run years out pay a premium to intermediaries who already hold sockets. That premium is a rent on a scarce factor, not an arbitrage. An arbitrage is riskless and closes cleanly - this is not that at all.
Liability outlives the rent
The second defect is what turns poor returns into zeroes. Premium pricing that justifies the build lasts one GPU generation. The lease and the debt that financed it last three to five. Stack this on the wedge above and the sequence is clear: the rent decays, the cost disadvantage surfaces, and the fixed obligation is still there for another decade.
Who owns what
Operator quality against asset quality dispersion
Treating this cohort as one asset class an error. Asset quality is the land, the power, the interconnect position, the permits. Operator quality is whether the megawatts actually produce sellable tokens: goodput, uptime, fabric design, cluster orchestration, time to energize. These are not correlated. A converted miner can hold superb land and be a poor operator. A pure tenant can be an excellent operator on borrowed shells.
Architecture trap… operational sophistication destroys acquisition value
A key risk is that hyperscalers will not buy merchant neoclouds, because they run proprietary reference designs and will not inherit someone else’s bespoke ops and infra. What a hyperscaler mostly likely wants from an acquisition is raw, clean, conformable megawatts: permitted sites, long-dated interconnects, pipeline capacity it can shape into its own internal design. Or they want a great price!
If what the buyer wants is conformable capacity rather than a company, it does not need the equity at all. It can buy the land, buy the development pipeline, or simply sign a lease. That weakens the takeout case for operators specifically and concentrates it on landbank holders with uncommitted capacity.
One correction to the objection, which matters for who bids as we’ve seen two distinct acquirer types. Hyperscalers buy capacity and conform it. AI-native buyers buy stacks, and they have been active, with CoreWeave acquiring Weights & Biases, OpenPipe and Monolith, and Nvidia absorbing Groq into rack-scale LPX in about eight months. So sophistication is not unsaleable but requires a different set of circumstances.
Roughly a dozen listed miners have signed decade-plus leases this year, and CoinShares put the sector at up to 70% of revenue from AI by December 2026, from about 30% earlier in the year, on more than $70B of announced contracts. That is a genuine transfer of scarce, energized land into AI service. It is also the clearest evidence for the thesis: when a $10M market-cap shell can sign a headline billion-dollar contract, the scarce thing is not the operator. It is the interconnect queue position, and that is exactly the asset that stops being scarce once the queue clears.
The one exception, and why it is not a contradiction
Among the pure-plays the answer is one name. Nebius is in a league of its own, and it should not be grouped with the generic GPU landlords or the leveraged sublease vehicles.
The separation rests on a verifiable asset fact, not positioning:
Owns >75% of its capacity, against a cohort defined by leasing shells
Built the layers between silicon and enterprise workload (orchestration, managed services, confidential compute), converting commodity hardware into something with retention rather than spot churn
~$46B backlog including the Microsoft agreement; the market already marks the difference, Nebius +130% ytd vs CoreWeave +7%
This is where the architecture trap stops being a contradiction. Sophistication is only self-defeating for an operator that needs an exit, and Nebius is building toward not needing one.
When does scarcity ease?
Scarcity eases unevenly by region and by procurement route. The binding constraint is energization. Interconnection queues in the busiest US markets run four to seven years. Queue-to-commercial-operation has stretched about 60% since 2017 and now averages over 2,100 days. Large power transformers run roughly 128 weeks and generator step-up units 144, with substation transformers now past 160 weeks, against a forecast 30% supply deficit through 2027; the same unit shipped in four to six weeks in 2020. Large-frame turbines from GE Vernova, Siemens Energy and Mitsubishi are booked through 2028. ERCOT’s large-load queue went from 63 GW to 226 GW in a single year with roughly nine of ten requests from data centers. PJM’s capacity auction came in 6.6 GW below its reliability requirement for 2027-28 at a record clearing price, with data centers driving about 94% of projected load growth through 2030. Dominion has said it cannot take additional large-load interconnection in Northern Virginia through 2030. Of 16 GW targeted for delivery in 2026, only about 5 GW entered active construction, and 30 to 50% of the remainder slips.
A second, independent constraint sits alongside it, and it needs to be scoped precisely rather than waved at. DRAM is a 2027 constraint on incremental capacity above current plans. Builds already in the plan of record have memory allocated against them. What memory gates is the upside case: the additional gigawatts a builder would add if it could, and the higher-specification configurations it would prefer. That is visible in Rubin Ultra de-speccing from planned 16-hi HBM4E toward 8-hi HBM4, which the house read treats as supply rationing rather than demand destruction. The practical effect is that memory does not stop the current build; it caps the response to any 2027 demand surprise, which is precisely the mechanism that keeps pricing firm.
The energization gate is firmly in place through 2027, and plausibly persists into 2028 depending on permitting, transformer and turbine progress. Both power and DRAM likely ease into 2028 rather than clearing. The merchant tier has sometime to execute, but any softening in the market (demand or bottlenecks) could prove disastrous for their equity.
Margin path
What matters is where each type of player’s margins are heading over the comin years. And again, this does not bode well for neoclouds who are simply disadvantaged.
NVDA: the counter that matters most
Nvidia sits on five sides at once… chip supplier, equity holder, lease-holder, rent-back guarantor, and now owner of the model-distribution layer via Hugging Face. The hyperscalers backstop offtake to keep capex off their own balance sheets. This is engineered solvency, and it can persist far longer than a naive short thesis assumes. It cuts both ways, both confirming these are financial vehicles rather than franchises, and it explaining why they stay bid.
Backstop is wildly important
The backstop is simultaneously what keeps the merchant tier alive and the clearest evidence it is not a business. In August 2026 Nvidia announced financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilize over $500B of third-party capital, and is backstopping up to 25% of a site’s residual value. Jensen Huang’s framing is that AI factories are an investable asset class… they produce revenue, serve a broad market, improve with CUDA over time, and can be redeployed.
Ben Thompson’s read is the one to carry: the guarantee is a price cut in disguise. Nvidia is putting its own profit at risk in order to lower the cost of capital for people building with its chips. Follow that logic and three things fall out.
Nvidia’s pricing power is already peaking. A monopolist with unchallenged pricing power raises price, it does not subsidize its customers’ financing. Thompson locates the cause precisely in the guarantee downstream of Google’s and Amazon’s aggressiveness. If capital is the binding constraint on a new data center, a cheaper chip up front beats a more efficient one, which is exactly the opening TPU and Trainium exploit. Nvidia’s own new reporting split, separating hyperscaler sales where it is fighting commoditization from everyone else where it runs the whole stack, is the company conceding the two-tier structure in its disclosure.
The fragmented tier is being financed, not validated. Nvidia is not making these companies profitable. It is making them fundable. This is a concentration hedge expressed in the capital structure. Remove the backstop and a large part of the tail cannot clear its return hurdle at any contract price the buyers will pay.
Risk is being moved somewhere worse. The big four raised $194B by early July 2026 versus $108B in all of 2025; spreads are widening, 86% of this year’s bonds trade above their issue yield, and cover has fallen below 2x from 5x in February. The marginal dollar is now moving past IG debt into insurance floats and pension money via the asset managers. Thompson’s analogy (and Nadella’s, on the MSFT call) is Jay Cooke financing Northern Pacific into the Panic of 1873… adjusted for economy size, that boom’s railway bonds were ~$600B a year, the same order as today’s tech capex.
When fragmentation unwinds
The fragmentation is a manufactured disequilibrium, held up by three props that are all decaying… a chip monopolist managing its own customer concentration, a physical shortage minting temporary scarcity rents, and the speed premium of an unbundled supply chain during a land grab. When those props go, the structure reverts toward vertical integration on the balance sheets closest to demand and with the lowest cost of capital.
Nvidia’s worst outcome is a world where four buyers take most of its output. Those four (Amazon, Google, Microsoft, Meta) all design competing silicon, all have buyer power, and all want Nvidia’s margin. So Nvidia seeded the alternative.
Training and inference are forking
It is now clear that the highest-end GPU is a training requirement, and not an inference one. Inference, which is where the volume and the recurring revenue go, runs well on heterogeneous and custom silicon… TPU, Trainium, Inferentia, and depreciated prior-generation GPUs. Amazon and Google pay thinner margins to Broadcom, Marvell and Alchip than Nvidia charges, and still resell compute at a lower cost per token. As inference becomes the majority of the workload, value migrates from Nvidia’s newest silicon toward custom silicon that mostly sits on integrated balance sheets.
This leads straight into the neocloud problem through the GPU’s second life. A demand-owner cascades a chip… bleeding edge for training, then the same chip redeployed for internal inference for years. A standalone neocloud cannot cascade. It has to re-lease the depreciated asset into a falling spot market. Same silicon, very different lifetime value, and it is worth more inside the integrated owner.
Bundle, unbundle, re-bundle
This is a classic technology cycle, not a one-off. Value integrates, then disaggregates for speed, then re-integrates once specs stabilize and scale justifies owning more of the stack. Today the AI silicon chain is fully unbundled: TSMC fabs, Nvidia designs, an OEM assembles, a neocloud operates, a lab consumes. Every hop leaks margin. Unbundling won the land-grab phase because specialization let everyone scale fast. Re-bundling wins the maturity phase, because at 800B+ of annual spend the margin leakage is intolerable and the integrated owner captures it. Google is the Apple of this cycle… its own chip (TPU), its own data centers, its own model (Gemini), its own distribution. That is why the fully integrated player is best positioned, and why the middle of the chain gets squeezed from both ends.
Rent extraction will be temporary
The value a buyer extracts per token is rising faster than the cost of producing a token is falling. While compute is short, the data-center layer can push its price up into that widening gap, or at least not let it stretch so much, and neoclouds that are recontracting or repricing right now capture it. That is real, and it is the entire bull case for owning them today.
But the DC layer’s pricing power is a shortage rent, not a value-add. When the shortage ends, cost deflation resumes and dominates, the compute price rolls back toward cost, and the widening surplus accrues to the model and application layers instead. The window is defined by the shortage, and its width depends entirely on the compute demand curve.
WACC and proximity to demand decide the owner
In steady state, compute should be owned by whoever can fund it most cheaply and flex it against internal demand. That is not a neocloud. Two structural edges compound (both priced in Figs 12a to 12e):
Cost of capital: on a 15-year depreciating asset, the WACC gap alone lets the mega-cap outbid everyone for the same asset and still clear its hurdle
Dual-use optionality: those with internal demand burn compute when that return is highest and sells externally when it is not… a pure neocloud is structurally short an option only the integrated owner can realize
The same asset is therefore worth more inside Google or Amazon than as standalone neocloud equity… the textbook condition for consolidation.
Cost of capital comps
Alphabet’s size-weighted coupon across $62.5bn of issuance is 5.14%. Amazon is 5.04% on $77bn, Meta 5.41% on $55bn, Oracle 5.61% on $43bn; Microsoft has issued essentially nothing and funds from free cash flow. The neocloud cohort spans 5.90% to 9.88%. So the gap is 99bp to 474bp, not the 700bp Fig 12 carries, and it is a range rather than a number.
Converts distort comps
The instrument that makes this cohort look cheaply funded is the one that most distorts the comparison. Across $24.8bn of convertible paper from CoreWeave, Nebius, IREN, TeraWulf, Cipher, Core Scientific and Bitdeer, the size-weighted headline coupon is 1.43%, and four of the fifteen tranches carry a zero coupon.
A convertible is a bond plus a call option written on the issuer’s own equity, and on a stock running 65-80% volatility that option is expensive. Valuing each one at issue and annualizing it over the life of the note gives a true economic cost of 8.2% to 9.5%. The headline understates by roughly 680 to 810bps, which means reported interest expense understates the true cost of capital across the entire cohort.
Carrying it through to returns
GCP sells complete TPU systems externally at $35bn per GW gross at a low-30s EBIT margin, so Google’s own cost to build a gigawatt of TPU is about $23.8bn. TPU v7 TCO per chip runs roughly 44% below a GB200 server, which grosses the equivalent Nvidia deployment to about $42.5bn. Add $11bn of shell and power to each and the stacks are $34.8bn against $53.5bn per GW. Revenue comes off the rental curve: at 3.5kW per GPU including PUE a gigawatt carries about 286,000 GPUs, so at 90% utilization the cost-based floor of $4.92 per GPU-hour yields $11.1bn per GW per year, today’s shortage pricing around $7.50 yields $16.9bn, and the value-based ceiling of $9.63 yields $21.7bn. Depreciate over six years, carry $1.5bn of cash operating cost, and the bridge closes.
The finding that reorders this note
Decomposing the gap says cost of capital is the smaller half of the story. Between Google on TPUs and a neocloud on Nvidia GPUs at shortage pricing there is a 20 point difference in spread. Hold capital constant and take away the TPU and the spread falls from +22 to +7.0: that 15 points is the silicon advantage. Hold silicon constant and take away the balance sheet and it falls from +7 to +2: that 4 points is the capital advantage.
Nvidia allocation stops being a moat
The original neocloud edge was privileged access to scarce Nvidia silicon. That edge is decaying on two fronts. Chips are getting easier than shells. TSMC has doubled CoWoS advanced packaging repeatedly; the acute allocation crunch of the early buildout has eased. A buyer with a signed power agreement and a data hall can now get accelerators on a shorter timeline than it can get a grid connection. The scarce thing moved from the chip to the socket to plug it into. An intermediary whose franchise was chip allocation is holding the asset that stopped being scarce.
Inference silicon is proliferating. As the mix shifts toward inference, purpose-built parts undercut the general-purpose GPU, and almost all of them are captive to a demand-owner. Note what OpenAI’s Jalapeño demonstrates beyond its specs: concept to tape-out in about nine months, with OpenAI’s own models assisting the design. That is the cost curve bending under AI’s own hand, and it is a template others will copy. Each new captive part removes a buyer from the merchant GPU market permanently.
Fewer labs perhaps, but more integration, and more captured rent
The final loop closes on the demand side. Inference, not training, is where model-company margin now accrues, and inference is a fundamentally different business that is more stable, repeatable, and continuously tunable. A lab running a large inference book can grind its cost per token down every quarter through kernels, caching, quantization, batching, scheduling and silicon choice. Those gains compound and they favor scale. Blended inference margins have already moved from under 40% to over 70% inside a year while the sticker price per token fell.
In commodity markets marginal cost determines not just profitability but viability. The frontier labs have the superior cost structure, and they are widening it by converting compute from a marginal cost into a capital cost, which is exactly what Anthropic did by buying TPUs outright rather than renting them. So the lab field narrows. As it narrows, economic profit concentrates in fewer hands, and those hands have both the cash and the motive to integrate further down the stack, capturing the rent that currently leaks to intermediaries.
The loop is bad for the merchant tier twice over. A consolidated lab field means fewer buyers, which worsens the oligopsony in Fig 1. And better-capitalized labs means more self-build, which shrinks the merchant market those fewer buyers address.
The open-weight objection, and why it fails
The most common rescue offered for the merchant tier is that open-weight models (Llama, DeepSeek, Mistral and successors) break the oligopsony by decentralizing the buyer universe from seven names into thousands of enterprises.
Model IP does decentralize, but serving capital does not. Enterprise adoption of open weights does not produce a Fortune 500 that each build data centers, rather it produces a Fortune 500 that queries hosted open models through managed endpoints. The question then becomes who serves those tokens at the lowest cost, and in a commodity market the low-cost producer wins. That is the integrated owner.
A hyperscaler can serve open models at or near zero margin, because serving is a customer-acquisition channel for database, security and application software it also sells. A merchant neocloud has to extract a living margin from the compute spread itself, because it has nothing to cross-sell.
Hyperscalers have a marginal funding is cheap and effectively unconstrained, while a neocloud funds at high-yield or GPU-collateralized rates unless it can pledge an investment-grade offtake contract.
Security culls the tail
One more force pushes the same direction. Model weights are the crown jewels, worth billions and increasingly the core IP of the labs. Handing them to a leveraged third party running multi-tenant hardware is a large attack surface and an IP risk. Security therefore raises the bar and pulls compute either fully in-house or toward a small number of certified, scaled, trusted providers. It is a tax the long tail cannot pay. Ben Thompson’s read on agentic security sharpens the point in a way that helps the integrated players specifically: incentives favor offense, and the only viable defense is for defenders to have access to the best models too. That makes the security layer inseparable from the model layer, and it is a product only a handful of companies can sell.
Which is exactly what is already happening. Thomas Kurian describes labs that use Google’s TPUs and then buy Google’s cybersecurity protection for their models. That is the integrated stack monetized twice off one customer, and it is revenue no neocloud can book. Of the seven buyers, only Google, Microsoft and Amazon own commercial security franchises, which is another reason the integration matrix concentrates rather than spreads.
Conclusion
Fragmentation is held up by Nvidia's channel strategy, a physical shortage, and the speed premium of unbundling. All three will decay. What remains is a bundling cycle turning back toward integration, a widening value-to-cost spread that pays the data-center layer only while supply is short, and a WACC-and-proximity gravity that makes the assets worth more inside the demand-owners. Fewer players will capture more margin on lower-cost-of-capital balance sheets. The neoclouds' job is to build, contract, and be absorbed on the way there.
The structure is a manufactured disequilibrium reverting toward integration, and that reversion is likely over a three year horizon or less. Own it through the consolidators, through the upstream, and selectively through Nebius.
DISCLOSURE
Position: At the time of publication, the author holds a positions in GOOGL and NBIS.
Trading Policy: The author will not materially alter this position within 48 hours of publication. After this period, the author may buy, sell, or otherwise adjust the position without further notice. Changes to the author’s view or position will be reflected in subsequent publications when material.
Conflicts: The author has received no compensation from the issuer or any party with a financial interest in this security.
Forward-Looking Statements: This report contains the author’s opinions, estimates, and projections, including price targets derived from financial models. These are forward-looking statements subject to substantial uncertainty. If these assumptions prove incorrect, the actual value may differ materially, including scenarios of significant loss or total impairment. The price target represents the author’s estimate of fair value under the stated assumptions, not a prediction of where the stock will trade.
This report is provided for informational purposes only and does not constitute investment advice. See the full Terms & Disclosures for additional important information.




























