Ornn paper: open-weight models cost 80% less than closed rivals
Study shows older GPUs retain value serving cheaper open-source inference workloads, challenging the assumption that new hardware obsoletes the old.
What to know
- Open-weight models cost roughly 80% less than equivalent closed-model APIs for equivalent intelligence, and self-hosting lowers costs to $0.12–$0.35 per million tokens.
- Older NVIDIA GPUs (A100) retain economic value on long-term contracts—the five-year A100 rental maintains 80% of its one-month price while newer families drop to 44–60% retention.
- Latency-tolerant workloads (batch inference, reinforcement learning, agents) create price-elastic demand that directs compute to cheaper hardware rather than forcing upgrades to newer generations.
- NVIDIA's acquisition of Hugging Face on September 3, 2026, gives the company direct control over a major open-weight repository as the open-source inference market grows.
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
1
A100 rental prices retain 80% value on five-year contracts
Ornn's rental market data shows the five-year A100 contract price holds 80% of its one-month rate, compared to 44–60% retention for newer Hopper and Blackwell families. This economic durability reflects continued demand from compute-intensive, latency-tolerant workloads like long-running agents and reinforcement learning.
“These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively.”
— Ornn Data · source -
background
Study shows open-weight models cost ~80% less than closed equivalents — Comparing across the Artificial Analysis Intelligence Index, the paper finds the cheapest open-weight model (gpt-oss-120b with 5.1 billion active parameters) completes tasks at roughly one-fifth the cost of comparable closed models, and that A100 GPUs produce output more cheaply than newer H100 and Hopper families on longer-term rental contracts.
-
2
Ornn publishes economics paper on open-weight inference costs
Ornn Data releases a research paper examining how open-weight AI models affect the economic lifespan of older GPU families. The study compares pricing across eleven open-weight and eight closed models, finding that self-hosted open-weight inference costs $0.12 to $0.35 per million output tokens—substantially cheaper than closed-model APIs.
-
first by HN Frontpage, 3d ago
-