Featured paper
The Economics of Open-Weight Inference
Abstract
GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.