Developer challenges article without empirical counterexamples
7 Yesterday 2:12 PM · 1d ago · 2 comments · 1 source · development 7 of 7
A commenter reports personally validating AI confidence scores and seeing linear accuracy scaling, claiming the article makes critiques without providing concrete failure examples to back them up.
“Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.”
benjaminsky2ihatethefuture.com author Article authoroli5679 AI systems engineer (ops automation)adamddev1 Commenterbenjaminsky2 Developer with validation experience
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 8 voices · verbatim
-
Reading this article put me in mind of how everyone says "things aren't made like they used to be." Someone was just telling me this above a washing machine. I think we've seen that cheap and fast won and the lower quality became normal. I think it's fair to say that there is a fear software is heading the same way. AI is going to help people ship…
-
I build ai systems for ops automation and don’t understand the author’s pessimism.I agree with that Evals are a scarce commodity rn. A business needs to define what good looks like. This is a laborious, and sometimes politically controversial, process.Given good Evals, frontier llms are a magical tool that can automate tasks and do them more…
-
> Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.I have definitely seen more bafflingly-poor OSS software that just plain doesn't work frequently now than before.But it's mostly software that wouldn't have existed before because it's trying to do super-niche things. So on the "hobby"…
-
> they can always shrug and say "well, AI makes mistakes." Error budgets? Failure modes? Test sets? All of those can be handled later.This is just the complete opposite in my experience. Tests are the first thing the AI writes, especially in low coverage or unknown domain situation.These systems have been trained for "generations" to oneshot…
-
That was awfully specific. Apple has managed to normalize inexplicable failures of things that had been working before for over a decade, across all of their apps and the OS. And they didn't even need AI for that. But AI will definitely speed up the rate of blunder everywhere.
-
I've seen people opt to refactor their entire codebase from one lang to another just because AI made it so much easier. Sure there were problems before, but now the new code with the better language is not readable and needs another refactor once this entire thing is done.
-
Amount of APIs failing for no reason and the answer is just retry these days is insaneI don't mind doing it but why is this the norm
-
"Appeal to adult animated show" is lame enough on Reddit but it's particularly odious blogspam here.
All 7 developments of Article on AI-driven software failures sparks debate over… →
NewswiresHacker NewsMastodonLobsters