The Economics of Open-Weight Inference

AI compute research and GPU market publications

Ornn publishes first-party research on GPU rental prices, compute markets, depreciation, procurement, memory, open-weight inference, and compute futures. This hub collects original summaries and the open-weight inference paper; the full essays remain on Substack.

Featured paper

The Economics of Open-Weight Inference

Ornn Data

Abstract

GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.

Read the paper

The Basis Gap in Compute Acquisition Structures

Why Basis Risk Defines Every Procurement Decision

GPU procurement is not a search for the cheapest hour. Buying, leasing, and reserving capacity each package three exposures differently: replacement-price risk, unused-capacity risk, and residual value. The essay’s thesis is that the hidden variable is basis—the gap between what a contract promises and what a workload later needs. A cheap commitment without reserved machines can be right on rate and wrong on access; a reservation without a matching spec can strand capacity while the buyer still pays for a second, better-fit block. At thousands of GPUs those mismatches become project-finance problems, because collateral values and contracted cash flows move on different clocks. The market question is which risks a fleet owner should keep, and which to push to a lessor or lender, when hardware, utilization, or demand diverge from the term sheet.

Read the full essay on Substack

Compute in Space. Why, How, and When?

The emergence of AI has led to tremendous demand for compute.

Ground datacenters cannot expand as fast as AI demand because power interconnects, water for cooling, and suitable land all take years. The essay argues that orbit is becoming a real alternative, not a slogan: launch costs have fallen sharply, solar collection is far stronger above the atmosphere, and radiative cooling can reject heat without consuming rivers. Early orbital H100 hardware and reusable-rocket economics are the proof points. Radiation, servicing, and extra latency still decide whether space is a hedge or a science project, especially for delay-sensitive inference. The practical question for buyers and financiers is when orbital capacity should enter a terrestrial fleet’s cost and obsolescence model—and how to hedge the chance that ground hardware is stranded if compute migrates off-planet.

Read the full essay on Substack

On GPU Depreciation

The last couple of months there has been strong commentary from investors — many of whom we admire — surrounding the current depreciation schedules of GPUs.

Investor alarm that five-year GPU lives overstate earnings treats an accounting schedule as if it measured physical wear. The essay’s thesis is that this conflates two different things. GAAP depreciation allocates capex, moves reported earnings, and creates a tax shield. Economic loss of GPU value is driven mainly by new architectures and training regimes, not by parts failure. Older accelerators can remain cheap for batched, throughput-heavy work long after they look obsolete on a latency benchmark, and hyperscaler retirement histories imply useful lives closer to seven to nine years. The market question for operators and lenders is how to report and finance useful life without using a depreciation table as a residual-value forecast—and whether economic obsolescence should be hedged rather than modeled as linear wear.

Read the full essay on Substack

Memory: How It Works and Why It Sets the Cost of Compute

The AI systems of today are incredibly powerful - they can write code better than most software engineers, analyze documents, generate images, and can reason through complex problems to reach coherent solutions.

Frontier models wait on memory more than they wait on arithmetic. Prefill and decode keep weights and a growing key-value cache in high-bandwidth memory, so HBM capacity and bandwidth set how much context, occupancy, and therefore usable compute a GPU can deliver. The essay’s thesis is that memory, not transistors, is now the price-setting input. HBM cannot expand smoothly: stacking yield, CoWoS packaging, and a three-vendor allocation market reprice at contract resets rather than on a continuous spot curve. Because memory already dominates the bill of materials on leading accelerators, a discrete HBM step change reprices whole clusters at once. The practical question is whether buyers and lenders can treat HBM as a scarce, hedgeable commodity instead of a background component cost.

Read the full essay on Substack

The AI Chip Wars: From Supplier Dominance to Contested Equilibrium

Recent reporting that OpenAI has begun running a portion of its training workloads on Amazon’s Trainium has been widely interpreted as a challenge to NVIDIA’s technical leadership.

Headlines about OpenAI training on Amazon Trainium, or NVIDIA licensing Groq inference designs, invite a winner-take-all reading of the chip race. The essay’s thesis is narrower: large buyers are building credible outside options so they can contest pricing power even if NVIDIA remains the performance leader. Hyperscaler silicon does not need to win every workload; it needs to be viable enough to change contract terms, optionality, and where margins sit in the stack. Competition will appear first in cloud pricing and flexibility, not in merchant-GPU share. The market question is how to budget and hedge GPU-hours when the price of compute fragments across architectures, regions, and bargaining positions rather than remaining a single supplier quote.

Read the full essay on Substack