Essay "Tokens Too Cheap to Meter" sparks debate over AI cost collapse
A blog post arguing AI inference costs are falling by orders of magnitude a year draws pickup on Hacker News and Lobsters, with commenters split on whether the trend is real or subsidized.
What to know
- Essay claims AI inference cost-per-task has fallen roughly two orders of magnitude within 2026 alone, driven by GPU power-efficiency doubling every two years and rapidly improving inference software.
- It predicts LLMs will become embedded computing infrastructure within one to two years and run at frontier quality on local commodity hardware within three to six years.
- Commenters dispute whether the cited price drops reflect genuine hardware efficiency or are being subsidized/marketed by AI labs, since providers control the benchmarks cited.
- Others counter with independent benchmark data and profitable third-party inference providers as evidence the trend is real and ongoing, citing newer cheaper model releases since publication.
The dispute Whether the essay's price-decline data reflects real hardware/software efficiency gains or is inflated/subsidized by AI labs' own claimed benchmarks. · positions read across 46 posts and comments
The claimed price/efficiency drops are supported by independent benchmarks and market evidence, and the trend is continuing with newer model releases.
-
“Since this post was written, there have actually been some more price-efficiency improvements from the U.S. LLM companies... For now, the trend continues.”
gulbanana · Lobsters ↗
It's unclear whether the reported price declines reflect true costs, since they may be subsidized by AI labs or just taken on the labs' word.
-
“How do we know that the price per token is going down? Are we just taking the AI labs word for this?”
roflmaoqwerty · Lobsters ↗
The practical implications are more nuanced: tool execution costs may fall similarly, and profitable third-party providers already show the economics work.
-
“Why wouldn't a tool's execution get cheaper on a similar curve to inference costs?”
mysteriouspants · Lobsters ↗
jyn.dev (blog author) Author of the essayroflmaoqwerty Lobsters commentergulbanana Lobsters commentertimthelion Lobsters commenter
How it unfolded 5 developments, newest first · click a bar or a number to jump postscomments
-
5
Commenters cite hardware benchmarks and provider economics as evidence
Other commenters counter the skepticism by pointing to independent benchmark data on cost-per-intelligence and to third-party inference providers profitably serving open-source models on commodity GPUs, arguing the efficiency gains are grounded in hardware trends rather than marketing.
“This isn't about marketing claims, this is just basic hardware benchmarking. The inference hardware is getting more efficient per watt.”
— timthelion -
Because there are third party inference providers serving large open source models without large investments and they seem to make s profit.
2 more of the top 3 · 41 posts in this stretch
-
A few years is hardly forever and the state of the world here indicates a lot of low hanging fruit still exists.An LLM can certainly be cheaper than grep, because it’s an approximation, while a grep is deterministic and must examine every byte in what can be a relatively complex state machine for a regex based grep. There are other scales to…
-
H
Tokens Too Cheap to Meter L: https:// jyn.dev/tokens-too-cheap-to-me ter/ C: https:// news.ycombinator.com/item?id=4 9813482 posted on 2026.09.23 at 05:21:17 (c=0, p=7)
-
-
4
Commenter notes newer model releases continue the price-efficiency trend
A Lobsters commenter points out that since the essay was written, new versions of Claude Opus, GPT Luna and GPT Sol have shipped, each claimed to be cheaper and more capable than its predecessor, suggesting the trend described in the post is ongoing.
“Since this post was written, there have actually been some more price-efficiency improvements from the U.S. LLM companies... For now, the trend continues.”
— gulbanana -
The author gives an example of using [jgrep](https://github.com/keltokhy/jgrep) rather than grep. If I wanted to search a corpus of text for references to sphynx cats, I (or an LLM) could write a regex to search that corpus with grep. But if I want it to also catch references like "sphinx cats", "them bald cats", or "那种没有毛的猫的品种", then I'm suddenly…
1 more of the top 2 · 2 posts in this stretch
-
Since this post was written, there have actually been some more price-efficiency improvements from the U.S. LLM companies. New versions of Claude Opus, GPT Luna and GPT Sol have been reached, each (according to claimed benchmarks) cheaper and more capable than the previous versions of those models. For now, the trend continues.
-
-
3
Commenters question whether price declines are real or subsidized
Several Lobsters commenters push back on the essay's premise, asking whether reported per-token price drops reflect true costs or are being subsidized by AI labs, and whether the industry's own claims can be trusted.
“How do we know that the price per token is going down? Are we just taking the AI labs word for this?”
— roflmaoqwerty -
LLMs are evaluated on benchmarks, which show this large drop in cost per intelligence: https://sanand0.github.io/llmpricing/intelligence.html This aligns with my experiences too.
2 more of the top 3 · 3 posts in this stretch
-
How do we know that the price per token is going down? Are we just taking the AI labs word for this?
-
> once models are cheaper than tools Why wouldn’t a tool’s execution get cheaper on a similar curve to inference costs?
-
-
2
Essay reposted to Lobsters and Hacker News, gains traction
The same essay was submitted separately to Lobsters and, a second time, to Hacker News within about nine hours, accumulating tens of upvotes and a discussion thread of roughly a dozen comments.
“This is an increase in efficiency that we haven't seen since Moore's Law in the 1960s.”
— jyn.dev (blog author), Essay author · source -
1
Essay argues AI token costs are collapsing toward zero
A blog post at jyn.dev claims machine learning inference costs are falling by several orders of magnitude per year, driven by GPU power-efficiency gains (doubling roughly every two years) and improving inference engines, and predicts LLMs will become embedded computing infrastructure within one to two years and run locally at frontier quality within three to six.
“The price of using machine learning intelligence is decreasing by several orders of magnitude a year and shows no signs of slowing.”
— jyn.dev (blog author)
Also covered reported alongside — the timeline has no entry for these yet
-
2 outlets Tokens too cheap to meter
first by HN Best, 1d ago · also HN Frontpage
What people are saying 15 voices from 3 sites · best of 46 · verbatim
- How do we know the cost per intelligence isn't being heavily subsidized by the labs?
- How do we know that the price per token is actually going down rather than just following labs' own claims?
- Yesterday
-
I don't know about the malleable software claim. Sure, people can build out their own thing. Malleable software in itself requires software that is designed to be malleable in the first place.And perhaps people will not be willing to accept the initial friction that malleable software brings (see people who complain about Emacs or Salesforce or…
-
N
Tokens too cheap to meter: https:// jyn.dev/tokens-too-cheap-to-me ter/ Discussion: http:// news.ycombinator.com/item?id=4 9813482
-
I'm always reminded on Orwell's quote about the then new atomic bomb and his prescience on how it would all work out:"Had the atomic bomb turned out to be something as cheap and easily manufactured as a bicycle or an alarm clock, it might well have plunged us back into barbarism, but it might, on the other hand, have meant the end of national…
-
4-5 orders of magnitude is huge. Assuming an order of base 10, it's 10000x-100000x. So a call to grep may return in 1s on a typical PC. That means a GPT call takes equivalent energy of 10000-100000 PCs to do the same in 1s. That's a difference that can't be equalized with scaling. It would require a revolutionary breakthrough.I also don't…
-
I just want to rant about these Artificial Analysis charts that you see everywhere:The "most attractive quadrant" is completely meaningless. The whole point of a Pareto curve is that each point on the curve is better than everything else on at least one dimension, and that you can make these comparisons without placing a value judgement on the…
-
I agree with OP that we will continue to see improvements, but there are also some serious bottlenecks ahead of us:- Energy is not infinite, neither energy efficiency is. - Datacentres neither. - Benchmarks are an abstraction of real world problems!On top, there is an overall "economic" aspect that most of the people miss: every change carries a…
-
It's true that LLMs "want" to be be local, but they won't shift broadly to being local until there's a sufficiently large supply of VRAM or (at least) "unified" memory from the manufacturers. (I'm also assuming here that radical regulatory changes like government bans of local models aren't going to happen.) So (AFAICS—I am no expert) the future…
-
This isn't what "subsidized" means. I think we can make some useful analysis about how much inference actually costs some providers, but I *don't* think it's unreasonable to say that we have very little idea how much inference costs OpenAI and Anthropic. It's entirely possible that they are asking customers for less than it costs them to serve…
-
A much deeper analysis on the falling price per task was published yesterday by Epoch AI [1]. It's a real statistical analysis and comes to more defensible and grounded conclusions. The headline takeaway is:The cost of a given level of performance often falls fastest right after that level is first achieved, that is, when it is state of the art…
-
> Tokens become cheaper than tool callsThe author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop."…
-
I found the OP insightful and worth a read. Thank you for sharing it on HN.The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they…
-
> NVIDIA will still boomI think Nvidia is under the same pressure as Anthropic/OpenAI. Nvidia will dominate research and probably keep dominating training, but the real volume is in inference. And for inference Nvidia's lead is only a few months, similar to the lead frontier labs have over open source. Nvidia will sell a lot of Rubin CPX's, but…
-
Open weight models can be self-hosted on your own hardware, with your own electricity, and you can calculate your own unsubsidised price per token. This used to be less meaningful when the frontier closed labs were far ahead of toy open models, but this year we've got large open weight models that are competitive with the best closed models from…
-
Check how similar the cost of Deepseek V4.1 Flash is around the world on different providers. Calculate how many users you can serve with 4xH200. And the quality is around GPT 5.6 Sol for programming and agentic tasks.
-
I guess my real question is how do we know the cost per intelligence isn't being heavily subsidized by the labs?