Analyst argues LLMs remain limited tools despite frontier lab advances
A lengthy technical post contends that most enterprises will need humans to oversee AI systems, not autonomous replacements, favoring cheaper open models.
What to know
- Author contends LLMs will function as assisted tools (not autonomous agents) for most enterprises due to structural architectural limits, not capability gaps.
- Cheap open models running in parallel swarms will likely outcompete expensive frontier models for most real-world applications, undermining frontier lab valuations.
- Only three narrow classes of firms can realistically deploy fully autonomous LLMs; the rest need heavy human oversight and guardrails.
- Disagreement hinges on whether current limitations are fundamental to transformer architecture or temporary gaps that frontier models will soon overcome.
The dispute Whether current LLM limitations are fundamental architectural constraints (supporting the author's thesis) or temporary capability gaps that frontier labs are rapidly closing (contradicting it). · positions read across 62 posts and comments
Current limitations are architectural and structural, favoring cheap open models and human-orchestrated swarms over frontier labs.
-
“Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.”
knuppar · Hacker News ↗
The author tested outdated models and ignores tool use; frontier models are advancing rapidly and constraints are temporary, not structural.
-
“The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.”
joefourier · Hacker News ↗
The goalposts for "limitation" keep shifting; frontier models keep improving on tasks that once seemed impossible.
-
“As models advance, we shift the goalpost for what "simplest task" means. Before, "simplest task" meant "write a coherent English sentence." Now, "simplest task" means autonomously fix, review, and merge a bugfix.”
someguynamedq · Hacker News ↗
jaykru Technical analyst, blogger
How it unfolded 2 developments, newest first · click a bar or a number to jump articlespostscomments
-
2
Hacker News commenters debate the premises and implications
Discussion fragments across multiple objections: some argue the author's valuation math is wrong, others challenge the chess benchmarking as irrelevant, and several push back on the claim that frontier models lack autonomous capability. Disagreement centers on whether current limitations are architectural or merely task-specific.
“As models advance, we shift the goalpost for what "simplest task" means.”
— someguynamedq · source -
I find these points interesting > the labor costs of rigorous specification can greatly exceed that of direct implementation of an informal specification. the hardware engineering world presents a great case study on this That tracks with my experience > navier-stokes and statements in pure mathematics like it are the absolute best case scenario…
2 more of the top 3 · 62 posts in this stretch
-
Yea I get the bearishness from my own personal experience.Personally, I use LLMs for a lot of things. Oftentimes, I'm a think out loud type of person so even having something that feels like a rubber duck, but more competent, is already amazing for me. And LLMs are a lot more competent than a rubber duck.But especially sometimes I've noticed that…
-
wake me when it wins a FIFA fields medal
-
-
1
jaykru publishes bearish LLM analysis after Navier-Stokes
A technical post argues that despite recent advances, LLMs will remain limited tools requiring heavy human oversight for most enterprise use cases. The author contends that current architectures have structural limitations preventing full autonomy, and that cheap open models running in parallel swarms will outcompete expensive frontier models for most applications.
“for most domains LLMs will continue to look like a cracked intern: quick and effective in the hands of an adult but not given run of the place.”
— jaykru -
first by HN Best, 12d ago · also HN Frontpage
-
What people are saying 21 voices from 2 sites · best of 62 · verbatim
- Sep 17
-
Hm, this feels like it remains true! I have several colleagues in areas like finance who write loads of SQL but don’t otherwise write much code. These days they lean on LLMs a lot for this, but this has been true since I started here a couple years ago, when the LLMs weren’t nearly as good.
-
> When Ray and I were designing Sequel in 1974, we thought that the predominant use of the language would be for ad-hoc queries by planners and other professionals whose domain of expertise was not primarily database management. We wanted the language to be simple enough that ordinary people could ‘‘walk up and use it’’ with a minimum of training…
-
I didn't know that, and I'm interested in that kind of thing! Can you share a link you think covers that history well?
-
I wouldn't qualify this as bearish. All things considered LLMs are and will be a groundbreaking tool. But fair points
-
One thing I rarely see mentioned to relativise the impressiveness of these maths achievements, is that openai and anthropic employ a lot of excellent mathematicians. Some just got out of uni, some have more experience. This makes the big picture shift from "autonomous AI solves anything" to "top-level mathematicians with near unlimited material…
-
People forget, but SQL was originally intended for _business people_. That lasted like 5 minutes, lol.
- Sep 16
-
It is at the root of: COBOL: "business people can use this English-like language to specify their problems easily, no more programmers" AppleScript: "English-like almost natural language can be used to write programs manipulating other programs" Inform: "English-like language can be used to write interactive fiction like Zork (this one is…
-
I think using AI for customer service is really, really ineffective, and it's plain to anybody that has interacted with it.It's a shame that we've had decades of shitty customer service from companies that have the most responsibility and resources to do it, and people in tech have shrugged their shoulders saying "it's unreasonable to expect…
-
This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out.> the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three:1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping…
-
How could you even doubt the work done when it just received a (just-made-up-by-a-trump-backer-meat-proxy-crypto-billionnaire-scammer) "prize" ?
-
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).On that note, I actually had…
-
Navier-Stokes is a well defined problem, or "a hard technical problem". Most problems in the business world lack a good definition and tacit knowledge is required to solve them. As far as I've seen, AI lacks any kind of tacit knowledge, strategic thinking, etc what so ever.Take a customer service person, that as soon as AI agents replaced was…
-
What happens when you ask those same frontier models to write a chess-playing program?I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get…
-
>> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data; but…
-
Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct.I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can…
-
> Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal…
-
The last post on HN I read was about someone using LLMs to reverse engineer an Apple GPU driver for linux in a month. The top comment points out how the poster must have had specialist internal domain specific contact with Apple. But then the thread concludes that wasn't the case and that this would take domain experts years to do.> "current…
-
The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something.It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on…
- Sep 15
-
"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, "No, they're really not.They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be.And that they…
-
The premise in the very first point seems off:> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers...Even assuming this is how the AI companies are being valued (they're not), the numbers are off.The "value" of most…
-
I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things…