Hacker News commenters debate the premises and implications
2 Sep 15 8:21 PM · 12d ago · 2 posts · 16 comments · 2 sources · development 2 of 2
Discussion fragments across multiple objections: some argue the author's valuation math is wrong, others challenge the chess benchmarking as irrelevant, and several push back on the claim that frontier models lack autonomous capability. Disagreement centers on whether current limitations are architectural or merely task-specific.
“As models advance, we shift the goalpost for what "simplest task" means.”
someguynamedq · hn ↗jaykru Technical analyst, blogger
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 24 voices · best of 62 · verbatim
-
I find these points interesting > the labor costs of rigorous specification can greatly exceed that of direct implementation of an informal specification. the hardware engineering world presents a great case study on this That tracks with my experience > navier-stokes and statements in pure mathematics like it are the absolute best case scenario…
-
Yea I get the bearishness from my own personal experience.Personally, I use LLMs for a lot of things. Oftentimes, I'm a think out loud type of person so even having something that feels like a rubber duck, but more competent, is already amazing for me. And LLMs are a lot more competent than a rubber duck.But especially sometimes I've noticed that…
-
wake me when it wins a FIFA fields medal
-
>> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data; but…
-
It is at the root of: COBOL: "business people can use this English-like language to specify their problems easily, no more programmers" AppleScript: "English-like almost natural language can be used to write programs manipulating other programs" Inform: "English-like language can be used to write interactive fiction like Zork (this one is…
-
Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct.I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can…
-
One thing I rarely see mentioned to relativise the impressiveness of these maths achievements, is that openai and anthropic employ a lot of excellent mathematicians. Some just got out of uni, some have more experience. This makes the big picture shift from "autonomous AI solves anything" to "top-level mathematicians with near unlimited material…
-
This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out.> the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three:1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping…
-
How could you even doubt the work done when it just received a (just-made-up-by-a-trump-backer-meat-proxy-crypto-billionnaire-scammer) "prize" ?
-
> Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal…
-
People forget, but SQL was originally intended for _business people_. That lasted like 5 minutes, lol.
-
What happens when you ask those same frontier models to write a chess-playing program?I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get…
-
I wouldn't qualify this as bearish. All things considered LLMs are and will be a groundbreaking tool. But fair points
-
The premise in the very first point seems off:> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers...Even assuming this is how the AI companies are being valued (they're not), the numbers are off.The "value" of most…
-
I didn't know that, and I'm interested in that kind of thing! Can you share a link you think covers that history well?
-
Navier-Stokes is a well defined problem, or "a hard technical problem". Most problems in the business world lack a good definition and tacit knowledge is required to solve them. As far as I've seen, AI lacks any kind of tacit knowledge, strategic thinking, etc what so ever.Take a customer service person, that as soon as AI agents replaced was…
-
> When Ray and I were designing Sequel in 1974, we thought that the predominant use of the language would be for ad-hoc queries by planners and other professionals whose domain of expertise was not primarily database management. We wanted the language to be simple enough that ordinary people could ‘‘walk up and use it’’ with a minimum of training…
-
The last post on HN I read was about someone using LLMs to reverse engineer an Apple GPU driver for linux in a month. The top comment points out how the poster must have had specialist internal domain specific contact with Apple. But then the thread concludes that wasn't the case and that this would take domain experts years to do.> "current…
-
Hm, this feels like it remains true! I have several colleagues in areas like finance who write loads of SQL but don’t otherwise write much code. These days they lean on LLMs a lot for this, but this has been true since I started here a couple years ago, when the LLMs weren’t nearly as good.
-
"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, "No, they're really not.They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be.And that they…
-
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).On that note, I actually had…
-
The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something.It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on…
-
I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things…
-
I think using AI for customer service is really, really ineffective, and it's plain to anybody that has interacted with it.It's a shame that we've had decades of shitty customer service from companies that have the most responsibility and resources to do it, and people in tech have shrugged their shoulders saying "it's unreasonable to expect…
All 2 developments of Analyst argues LLMs remain limited tools despite frontier… →
Hacker NewsMastodonNewswiresLobsters