The Economics of Open-Weight Inference

AI compute research and GPU market publications

Ornn publishes first-party research on GPU rental prices, compute markets, depreciation, procurement, memory, open-weight inference, and compute futures. This hub collects original summaries and the open-weight inference paper; the full essays remain on Substack.

Featured paper

The Economics of Open-Weight Inference

Ornn Data

Abstract

GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.

Read the paper

On Sovereign Compute

In the world of compute, a lot has changed in the last twelve months.

Advanced AI chips have crossed from commercial input to sovereign commodity. Export licensing, rare-earth retaliation, national subsidy races, and a wave of AI-sovereignty pledges now treat GPU capacity the way states once treated oil, rare earths, and LNG: once a resource is contested, governments secure access and permanently change how the market clears. Fabrication, high-bandwidth memory, and advanced packaging sit in a handful of jurisdictions and firms, so a single policy or geographic shock can reprice fleets in days. The essay’s thesis is that compute is already on that historical arc, yet it still lacks the benchmarks and hedging tools those earlier commodities eventually required. The practical market question is how operators, buyers, and lenders should budget and finance GPU capacity when jurisdiction—not just hardware generation—can move the price overnight.

Read the full essay on Substack

The Basis Gap in Compute Acquisition Structures

Why Basis Risk Defines Every Procurement Decision

GPU procurement is not a search for the cheapest hour. Buying, leasing, and reserving capacity each package three exposures differently: replacement-price risk, unused-capacity risk, and residual value. The essay’s thesis is that the hidden variable is basis—the gap between what a contract promises and what a workload later needs. A cheap commitment without reserved machines can be right on rate and wrong on access; a reservation without a matching spec can strand capacity while the buyer still pays for a second, better-fit block. At thousands of GPUs those mismatches become project-finance problems, because collateral values and contracted cash flows move on different clocks. The market question is which risks a fleet owner should keep, and which to push to a lessor or lender, when hardware, utilization, or demand diverge from the term sheet.

Read the full essay on Substack

Compute in Space. Why, How, and When?

The emergence of AI has led to tremendous demand for compute.

Ground datacenters cannot expand as fast as AI demand because power interconnects, water for cooling, and suitable land all take years. The essay argues that orbit is becoming a real alternative, not a slogan: launch costs have fallen sharply, solar collection is far stronger above the atmosphere, and radiative cooling can reject heat without consuming rivers. Early orbital H100 hardware and reusable-rocket economics are the proof points. Radiation, servicing, and extra latency still decide whether space is a hedge or a science project, especially for delay-sensitive inference. The practical question for buyers and financiers is when orbital capacity should enter a terrestrial fleet’s cost and obsolescence model—and how to hedge the chance that ground hardware is stranded if compute migrates off-planet.

Read the full essay on Substack

On GPU Depreciation

The last couple of months there has been strong commentary from investors — many of whom we admire — surrounding the current depreciation schedules of GPUs.

Investor alarm that five-year GPU lives overstate earnings treats an accounting schedule as if it measured physical wear. The essay’s thesis is that this conflates two different things. GAAP depreciation allocates capex, moves reported earnings, and creates a tax shield. Economic loss of GPU value is driven mainly by new architectures and training regimes, not by parts failure. Older accelerators can remain cheap for batched, throughput-heavy work long after they look obsolete on a latency benchmark, and hyperscaler retirement histories imply useful lives closer to seven to nine years. The market question for operators and lenders is how to report and finance useful life without using a depreciation table as a residual-value forecast—and whether economic obsolescence should be hedged rather than modeled as linear wear.

Read the full essay on Substack

Memory: How It Works and Why It Sets the Cost of Compute

The AI systems of today are incredibly powerful - they can write code better than most software engineers, analyze documents, generate images, and can reason through complex problems to reach coherent solutions.

Frontier models wait on memory more than they wait on arithmetic. Prefill and decode keep weights and a growing key-value cache in high-bandwidth memory, so HBM capacity and bandwidth set how much context, occupancy, and therefore usable compute a GPU can deliver. The essay’s thesis is that memory, not transistors, is now the price-setting input. HBM cannot expand smoothly: stacking yield, CoWoS packaging, and a three-vendor allocation market reprice at contract resets rather than on a continuous spot curve. Because memory already dominates the bill of materials on leading accelerators, a discrete HBM step change reprices whole clusters at once. The practical question is whether buyers and lenders can treat HBM as a scarce, hedgeable commodity instead of a background component cost.

Read the full essay on Substack

The AI Chip Wars: From Supplier Dominance to Contested Equilibrium

Recent reporting that OpenAI has begun running a portion of its training workloads on Amazon’s Trainium has been widely interpreted as a challenge to NVIDIA’s technical leadership.

Headlines about OpenAI training on Amazon Trainium, or NVIDIA licensing Groq inference designs, invite a winner-take-all reading of the chip race. The essay’s thesis is narrower: large buyers are building credible outside options so they can contest pricing power even if NVIDIA remains the performance leader. Hyperscaler silicon does not need to win every workload; it needs to be viable enough to change contract terms, optionality, and where margins sit in the stack. Competition will appear first in cloud pricing and flexibility, not in merchant-GPU share. The market question is how to budget and hedge GPU-hours when the price of compute fragments across architectures, regions, and bargaining positions rather than remaining a single supplier quote.

Read the full essay on Substack

Circular Financing in AI

How Compute is Pulled Forward - and Risk is Reallocated

Cross-investments, take-or-pay contracts, prepayments, and utilization guarantees among chip vendors, labs, clouds, and builders look like a closed money loop. The essay argues they are a familiar infrastructure workaround: they pull future demand into the present so physical capacity can be financed before transparent markets exist. That stopgap concentrates utilization, construction, funding, and obsolescence risk on a small set of counterparties and hides true prices. The structure works only while demand assumptions hold. Markets mature when contracts standardize, benchmarks appear, and risks can be unbundled. The practical question for lenders and operators is which balance sheet absorbs the shock if spending, power timelines, or GPU generations break those loops.

Read the full essay on Substack

Compute Futures

How to hedge compute - Part 2

Capital is pouring into AI infrastructure while the value of a GPU-hour remains hard to trade. The essay’s thesis is that a cash-settled future, marked to a daily compute index and paying an Asian-style average over the tenor, matches how compute is actually used: as a flow, not a single expiry print. Incremental daily settlement plus margin lets consumers, producers, and financiers hedge without taking delivery or operational SLA risk. Oil-style terminal settlement would misalign the hedge with the period of use. The market question is whether operators and investors can lock in a period-average GPU-hour cost—and transfer counterparty risk through collateral—instead of managing exposure with opaque bilateral contracts.

Read the full essay on Substack

Compute as a commodity isn't oil. It's electricity.

How to hedge compute - Part 1

The first hedge design for compute copied oil: lock a future price for a storable barrel. The essay’s thesis is that this analogy fails at the unit of account. Compute is quoted per GPU-hour, and unused capacity cannot be recovered, so it is a flow good. Electricity is the better analog: an unused megawatt-hour vanishes, and power markets already price location, real-time utilization, and forward delivery around that temporal structure. Compute markets are beginning to show the same features. The practical question for anyone hedging GPU cost or revenue is whether the instrument should settle like a stock commodity at expiry or like power over the period of use.

Read the full essay on Substack