The Economics of Open-Weight Inference
How open-weight demand can support the useful life of NVIDIA GPU families
Ornn Data
Abstract
GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.
Key findings
- Hosted open-weight models were cheaper at several sampled common score thresholds on the Artificial Analysis Intelligence Index, not at every threshold. Closed models remain the cheaper qualifying option at some scores, and the closed frontier exceeds the open sample at the top of the range.
- Self-hosting on rented hardware produced compute-only costs of $0.12 to $0.35 per million output tokens at full utilization in the printed sample.
- On gpt-oss-120b, the A100 produced output more cheaply than the H100 at spot and at the three- and five-year term prices. The A100/H100 sparse result combines different third-party serving setups. The dense A100 row is estimated.
- The five-year A100 term price retains 80.2% of the one-month price, versus 43.7% to 59.8% for Hopper and 53.8% for Blackwell. Forward marks are analyst-produced indicators, not executable quotes.
- A100 occupancy rose from 74% to 90% as listed capacity rose 13% and the spot index rose 20%. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior.
Selected tables
| GPU | Dense full use | Dense base case | Sparse full use | Sparse base case |
|---|---|---|---|---|
| A100 SXM4 | $0.28 | $0.75 | $0.12 | $0.29 |
| H100 SXM | $0.26 | $0.68 | $0.27 | $0.64 |
| H200 | $0.32 | $0.81 | $0.35 | $0.98 |
| B200 | $0.16 | $0.39 | $0.15 | $0.47 |
Source: Ornn Data calculation from published throughput and 1 September 2026 spot rents (paper Table 7). Full use takes Offline throughput with no headroom. The base case takes Server throughput where available and otherwise rescales the GPUStack baseline; the latter does not establish a Server service level. Dense A100 is estimated. The A100/H100 sparse rows are third-party measurements, not MLPerf results. Costs are compute-only USD per million output tokens.
| Family | Occupancy, 1 Mar 2026 | Occupancy, 1 Sep 2026 | Occupancy, Mar–Sep mean | Listed capacity, Mar–Sep | Spot, Mar–Sep |
|---|---|---|---|---|---|
| A100 SXM4 | 74% | 90% | 80% | +13% | +20% |
Source: Ornn occupancy and listed-capacity series and the settled spot index (paper Table 8), retrieved 2 September 2026. Occupancy is rented capacity divided by listed capacity across tracked on-demand providers in the global region. Listed capacity measures tracked on-demand supply, not the installed base.
| Family | 6M / 1M | 1Y / 1M | 3Y / 1M | 5Y / 1M | Fwd 37–60 / 1M | Age at 5Y end (yrs) |
|---|---|---|---|---|---|---|
| A100 SXM4 | 95.0% | 93.1% | 84.2% | 80.2% | 74% | 11.3 |
| H100 SXM | 95.0% | 86.2% | 68.2% | 59.8% | 47% | 9.4 |
| H200 | 92.4% | 80.2% | 51.0% | 43.7% | 33% | 7.8 |
| B200 | 99.1% | 97.9% | 69.1% | 53.8% | 31% | 7.5 |
| B300 | 93.8% | 81.3% | 60.9% | 53.8% | 43% | 6.5 |
Source: Ornn reported term marks (paper Table 9). A100, H100, H200, and B200 marks were published 13 August 2026; B300 on 1 September 2026. Retention and implied-forward ratios are calculations from those marks. Ages use an illustrative 1 September 2026 start measured from family announcement dates. Forward marks are analyst-produced indicators, not executable quotes.
Open-weight portability
Closed models are available only through deployments authorized by their developers. Open weights allow independent deployment and, where compatible endpoints exist, a choice of providers serving the same checkpoint. Reported September 2026 OpenRouter snapshots listed eighteen to twenty-two providers for several widely served open models, with highest-to-lowest output-price ratios between 1.8 and 5.6. Those figures illustrate price variation, not interchangeable service: quantization, capacity, reliability, and migration costs can limit substitution.
The same choice extends to hardware. Memory, numerical format, software, and licensing determine which deployments are feasible. gpt-oss-120b uses MXFP4 expert weights and has been served on a single 80 GB A100. Operators can therefore consider older hardware where the model fits, the software supports it efficiently, and the license permits the intended use. The economic distinction is who can make the deployment decision.
GPU useful life
Ornn publishes a settled daily rental index and term-price curves at one-month, six-month, one-year, three-year, and five-year tenors. A term price is a flat GPU-hour rate over a stated contract tenor. An implied forward differences cumulative term costs to obtain a rate for the interval between two tenors. The marks price a rental service; an individual device’s resale value and operating life are separate quantities.
The A100’s five-year term price retains four-fifths of the one-month price for a contract that would run to September 2031, when the family will be more than eleven years old. The Hopper and Blackwell curves are steeper. Rising occupancy on an expanding A100 listing base means the family’s spot strength cannot be attributed solely to fewer tracked GPUs being available for rent. A new generation does not, by itself, make its predecessors economically obsolete.
Self-hosting economics
Self-hosting has no posted per-token price. Its compute-only cost is the GPU-hour rent on Ornn’s spot index divided by the throughput the operator achieves, adjusted for productive utilization and reserved headroom. The paper constructs that cost for two workloads: dense Llama-2-70B and sparse gpt-oss-120b.
The ordering reverses across those workloads. The dense comparison favors newer hardware. For gpt-oss-120b, the A100’s base-case cost is $0.29 per million output tokens, less than half the H100’s $0.64. Those A100 and H100 sparse inputs come from different third-party serving setups. The dense A100 row is estimated from a vendor throughput ratio in the absence of a matched MLPerf submission. Electricity at nameplate GPU power is a few percent of spot rent and does not reverse the printed rankings.
Interpretation
Long-running agents, batch evaluation, and parts of reinforcement learning offer flexibility over when and where work runs. Workloads that tolerate latency can route to any cost-efficient compatible hardware. Older hardware does not need to win every workload to retain an economic role. Newer GPUs can offer lower costs on demanding tasks while older capacity serves work suited to its memory, software, and rental price.
The right measure of hardware usefulness is the cost of the work it can serve and the demand for that work, not simply the age of the chip. Evaluating an individual investment still requires acquisition costs, operating cash flows, and the alternatives available to the owner. Financing and contract terms belong in that assessment because they help determine the price of a long commitment.
Limitations and disclosures
Forward marks are analyst-produced indicators, not executable quotes or verified averages of comparable executed contracts at every tenor. Differencing term marks does not identify expected future spot rent. August term marks and September spot observations are different vintages.
Occupancy records rental status across tracked on-demand providers, not the installed base, tenant workloads, or realized token throughput. Family ages run from NVIDIA’s announcement, not a device’s commissioning date. Strong percentage retention can also follow earlier declines in the starting rent. The marks do not establish individual-device survival, profitable operating life, residual value, or the appropriateness of an owner’s accounting depreciation schedule.
Dense A100 throughput is estimated; alternative precision assumptions can reverse its cost ranking. The A100/H100 sparse result combines different third-party serving setups. A common Index threshold screens models but does not establish equal task success. The two self-hosted examples do not establish the performance of the higher-scoring models in the hosted comparison.
The rental observations do not identify the contribution of open-weight demand. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior. Device retirements, dense inference, non-LLM work, operating withdrawal costs, and financing or scarcity effects remain alternative explanations.
Ornn Data publishes this paper and the rental index, occupancy series, and term marks used in it. It licenses data commercially and also offers GPU rentals through Ornn Compute. These activities create a commercial interest in both the interpretation and adoption of its data. There is no external funding to report. Nothing in this paper is an offer to trade or investment, accounting, tax, or legal advice.
Are open-weight models cheaper than closed models?
In the paper sample, the cheapest qualifying open-weight model, standardized by intelligence on the Artificial Analysis Intelligence Index, completes a task at roughly one fifth the cost of a comparable closed model. Open-weight models are cheaper only at several sampled common score thresholds, not at every threshold. Closed models remain the cheaper qualifying option at some scores, and the closed frontier exceeds the open sample at the top of the range.
Is the A100 cheaper than the H100 for self-hosted open-weight inference?
On gpt-oss-120b, a sparse mixture-of-experts model with 5.1 billion active parameters, the computed A100 cost is $0.12 per million output tokens at full use and $0.29 in the base case, below the H100 figures of $0.27 and $0.64. That A100/H100 sparse result combines different third-party serving setups. The dense Llama-2-70B comparison favors newer hardware, and the dense A100 row is estimated rather than taken from a matched MLPerf submission.
What are Ornn forward term-price marks?
Ornn term-price curves report flat GPU-hour rates at one-month through five-year tenors. The published five-year A100 mark retains 80.2% of the one-month term price, versus 43.7% to 59.8% for Hopper and 53.8% for Blackwell families. Forward marks are analyst-produced indicators, not executable quotes or verified averages of comparable executed contracts at every tenor.
Did open-weight demand cause A100 occupancy and rental-price behavior?
The paper reports A100 occupancy rising from 74% on 1 March 2026 to 90% on 1 September 2026, with mean occupancy of 80%, listed capacity up 13%, and the spot index up 20%. Those rental observations do not identify the contribution of open-weight demand. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior. Device retirements, dense inference, non-LLM work, operating withdrawal costs, and financing or scarcity effects remain alternative explanations.
How long can older NVIDIA GPUs keep earning after a new generation ships?
The paper argues that a new generation does not, by itself, make its predecessors economically obsolete. Older GPUs retain a role when useful workloads remain competitive on them and operators remain free to deploy those workloads. A five-year A100 contract beginning 1 September 2026 would end when the Ampere family is 11.3 years old, while still retaining 80.2% of the one-month term price. That is a family-level rental-service observation, not a finding about individual-device residual value or accounting depreciation.
Citation
Ornn Data. 2026. “The Economics of Open-Weight Inference.” 7 September. https://data.ornn.com/publications/the-economics-of-open-weight-inference.