The Economics of Open-Weight Inference

AI compute research and GPU market publications

Ornn publishes first-party research on GPU rental prices, compute markets, depreciation, procurement, memory, open-weight inference, and compute futures. This hub collects original summaries and the open-weight inference paper; the full essays remain on Substack.

The Basis Gap in Compute Acquisition Structures

Why Basis Risk Defines Every Procurement Decision

GPU procurement is not a search for the cheapest hour. Buying, leasing, and reserving capacity each package three exposures differently: replacement-price risk, unused-capacity risk, and residual value. The essay’s thesis is that the hidden variable is basis—the gap between what a contract promises and what a workload later needs. A cheap commitment without reserved machines can be right on rate and wrong on access; a reservation without a matching spec can strand capacity while the buyer still pays for a second, better-fit block. At thousands of GPUs those mismatches become project-finance problems, because collateral values and contracted cash flows move on different clocks. The market question is which risks a fleet owner should keep, and which to push to a lessor or lender, when hardware, utilization, or demand diverge from the term sheet.

Read the full essay on Substack

Compute in Space. Why, How, and When?

The emergence of AI has led to tremendous demand for compute.

Ground datacenters cannot expand as fast as AI demand because power interconnects, water for cooling, and suitable land all take years. The essay argues that orbit is becoming a real alternative, not a slogan: launch costs have fallen sharply, solar collection is far stronger above the atmosphere, and radiative cooling can reject heat without consuming rivers. Early orbital H100 hardware and reusable-rocket economics are the proof points. Radiation, servicing, and extra latency still decide whether space is a hedge or a science project, especially for delay-sensitive inference. The practical question for buyers and financiers is when orbital capacity should enter a terrestrial fleet’s cost and obsolescence model—and how to hedge the chance that ground hardware is stranded if compute migrates off-planet.

Read the full essay on Substack

On GPU Depreciation

The last couple of months there has been strong commentary from investors — many of whom we admire — surrounding the current depreciation schedules of GPUs.

Investor alarm that five-year GPU lives overstate earnings treats an accounting schedule as if it measured physical wear. The essay’s thesis is that this conflates two different things. GAAP depreciation allocates capex, moves reported earnings, and creates a tax shield. Economic loss of GPU value is driven mainly by new architectures and training regimes, not by parts failure. Older accelerators can remain cheap for batched, throughput-heavy work long after they look obsolete on a latency benchmark, and hyperscaler retirement histories imply useful lives closer to seven to nine years. The market question for operators and lenders is how to report and finance useful life without using a depreciation table as a residual-value forecast—and whether economic obsolescence should be hedged rather than modeled as linear wear.

Read the full essay on Substack

Circular Financing in AI

How Compute is Pulled Forward - and Risk is Reallocated

Cross-investments, take-or-pay contracts, prepayments, and utilization guarantees among chip vendors, labs, clouds, and builders look like a closed money loop. The essay argues they are a familiar infrastructure workaround: they pull future demand into the present so physical capacity can be financed before transparent markets exist. That stopgap concentrates utilization, construction, funding, and obsolescence risk on a small set of counterparties and hides true prices. The structure works only while demand assumptions hold. Markets mature when contracts standardize, benchmarks appear, and risks can be unbundled. The practical question for lenders and operators is which balance sheet absorbs the shock if spending, power timelines, or GPU generations break those loops.

Read the full essay on Substack