conv.

All stories
AIQuiet 4d · day 6

Ornn paper: open-weight models cost 80% less than closed rivals

Study shows older GPUs retain value serving cheaper open-source inference workloads, challenging the assumption that new hardware obsoletes the old.

What to know

  • Open-weight models cost roughly 80% less than equivalent closed-model APIs for equivalent intelligence, and self-hosting lowers costs to $0.12–$0.35 per million tokens.
  • Older NVIDIA GPUs (A100) retain economic value on long-term contracts—the five-year A100 rental maintains 80% of its one-month price while newer families drop to 44–60% retention.
  • Latency-tolerant workloads (batch inference, reinforcement learning, agents) create price-elastic demand that directs compute to cheaper hardware rather than forcing upgrades to newer generations.
  • NVIDIA's acquisition of Hugging Face on September 3, 2026, gives the company direct control over a major open-weight repository as the open-source inference market grows.

Ornn Data Research organizationNVIDIA GPU manufacturer

Ornn paper: open-weight models cost 80% less than closed rivals
data.ornn.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 2 pieces in two hours at Sep 22, 9 AM; 4 pieces over 6 days (1 article · 3 posts) Sep 20, 3 PM — 1 piece · 1 post — Hacker News 1Sep 20, 5 PM — quietSep 20, 7 PM — quietSep 20, 9 PM — quietSep 20, 11 PM — quietSep 21, 1 AM — quietSep 21, 3 AM — quietSep 21, 5 AM — quietSep 21, 7 AM — quietSep 21, 9 AM — quietSep 21, 11 AM — quietSep 21, 1 PM — quietSep 21, 3 PM — quietSep 21, 5 PM — quietSep 21, 7 PM — quietSep 21, 9 PM — quietSep 21, 11 PM — quietSep 22, 1 AM — quietSep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — 2 pieces · 1 article · 1 post — Hacker News 1, Newswires 1Sep 22, 11 AM — 1 piece · 1 post — Mastodon 1Sep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — quietSep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — quietYesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quiet 1–2
Sep 21Sep 22Sep 23Sep 24yesterdaynow · 8:29 AM ET
  1. 1

    A100 rental prices retain 80% value on five-year contracts

    Ornn's rental market data shows the five-year A100 contract price holds 80% of its one-month rate, compared to 44–60% retention for newer Hopper and Blackwell families. This economic durability reflects continued demand from compute-intensive, latency-tolerant workloads like long-running agents and reinforcement learning.

    “These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively.”
    — Ornn Data · source
  2. background

    Study shows open-weight models cost ~80% less than closed equivalents — Comparing across the Artificial Analysis Intelligence Index, the paper finds the cheapest open-weight model (gpt-oss-120b with 5.1 billion active parameters) completes tasks at roughly one-fifth the cost of comparable closed models, and that A100 GPUs produce output more cheaply than newer H100 and Hopper families on longer-term rental contracts.

  3. 2

    Ornn publishes economics paper on open-weight inference costs

    Ornn Data releases a research paper examining how open-weight AI models affect the economic lifespan of older GPU families. The study compares pricing across eleven open-weight and eight closed models, finding that self-hosted open-weight inference costs $0.12 to $0.35 per million output tokens—substantially cheaper than closed-model APIs.

    1. first by HN Frontpage, 3d ago