Article on AI-driven software failures sparks debate over quality standards
A piece examining how AI agents are normalizing inexplicable software failures prompts disagreement over whether the tradeoffs are acceptable.
What to know
- An article critiques how AI agents enable and normalize opaque software failures, particularly in infrastructure and libraries—layers whose reliability affects all downstream systems.
- Defenders argue that with proper evaluation frameworks, AI tools improve on human-level work; critics counter that most organizations lack rigor and rationalize persistent failure rates instead of fixing them.
- The deeper debate centers on misaligned incentives: whether AI makes a pre-existing speed-over-quality problem worse, and whether confidence scores mislead stakeholders into false confidence.
The dispute Whether the article identifies a genuine systemic risk (failures in critical layers) or merely critiques poor implementation practices by organizations that lack rigor—and whether developers' production results validate AI tools or hide cascading problems. · positions read across 18 posts and comments
AI tools can be safe and productive when paired with rigorous evaluation and testing frameworks.
-
“Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.”
oli5679 · Hacker News ↗
Accepting AI-driven failures in infrastructure and libraries cascades into systemic unreliability affecting all users.
-
“But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.”
adamddev1 · Hacker News ↗
The real problem is perverse incentives toward speed over quality; AI is enabling but not causing this.
-
“The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather rare now - thirty years ago, Microsoft won the office wars with crap that had features.”
joe_the_user · Hacker News ↗
Confidence scores mislead stakeholders and obscure actual failure modes through false authority.
-
“An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal…”
WorldMaker · Hacker News ↗
ihatethefuture.com author Article authoroli5679 AI systems engineer (ops automation)adamddev1 Commenterbenjaminsky2 Developer with validation experience
How it unfolded 7 developments, newest first · click a bar or a number to jump articlespostscomments
-
7
Developer challenges article without empirical counterexamples
A commenter reports personally validating AI confidence scores and seeing linear accuracy scaling, claiming the article makes critiques without providing concrete failure examples to back them up.
“Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.”
— benjaminsky2 -
Reading this article put me in mind of how everyone says "things aren't made like they used to be." Someone was just telling me this above a washing machine. I think we've seen that cheap and fast won and the lower quality became normal. I think it's fair to say that there is a fear software is heading the same way. AI is going to help people ship…
2 more of the top 3 · 8 posts in this stretch
-
I build ai systems for ops automation and don’t understand the author’s pessimism.I agree with that Evals are a scarce commodity rn. A business needs to define what good looks like. This is a laborious, and sometimes politically controversial, process.Given good Evals, frontier llms are a magical tool that can automate tasks and do them more…
-
> Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.I have definitely seen more bafflingly-poor OSS software that just plain doesn't work frequently now than before.But it's mostly software that wouldn't have existed before because it's trying to do super-niche things. So on the "hobby"…
-
-
6
Commenter frames the core issue as perverse incentives toward speed
A voice argues the underlying problem predates AI—that market incentives favor speed over quality—but AI tools are now enabling this trend to accelerate and worsen the outcomes.
“The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather rare now - thirty years ago, Microsoft won the office wars with crap that had features.”
— joe_the_user -
Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.Yes but that's the big thing, now isn't it? These are nice tools, used wisely. But their unwise use, oh boy...The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather…
-
-
5
Commenter criticizes acceptance of persistent failure rates
A voice argues that developers and organizations are increasingly willing to accept and rationalize persistent failure rates (around 10%) rather than addressing root causes, treating AI as an excuse.
“most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them.”
— lokar -
I don’t think the author (or many people) doubt that one can (and some will) find a way that does not “suck”But it’s pretty clear that most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them…
1 more of the top 2 · 2 posts in this stretch
-
I’ve validated Jev’s confidence score. Accuracy scales linearly with confidence for the 3 use cases I tested. >.9 it matched a human labeler. I immediately discovered a user behavior I didn’t expect for ~$3. I can now mitigate in real-time due to low cost and latency. This may have a major positive financial impact for all our customers.Not sure…
-
-
4
Commenters identify confidence score interpretation as a core problem
Multiple voices highlight how AI confidence scores mislead business stakeholders into false assumptions about reliability, linking the issue to broader misuse of statistics in ML deployment decisions.
“An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal percentage".”
— WorldMaker -
"Confidence scores" have always implied an anthopocentric meaning that doesn't exist. An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal percentage". I still…
1 more of the top 2 · 2 posts in this stretch
-
I am big on reproducibility (nix aficionado) and determinism (flagging test failures are a red-alert, all-hands-on-deck situation in my world) and correctness.I am also big on testing (the correct things). And nine-nines (big on Elixir).And... I'm also big on agent-assisted dev. Which requires pretty much every check in the book to stay productive…
-
-
3
Concern raised about normalizing failures in infrastructure layers
A commenter warns that accepting AI-generated failures at the library and infrastructure level could cascade, degrading reliability for all downstream systems and users.
“But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.”
— adamddev1 -
Excellent post. People always defend agentic/LLM-driven development by saying, "Well it's good enough", or "It works most of the time."That may be tolerable for some user-facing app. But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows…
1 more of the top 2 · 2 posts in this stretch
-
> This leads to a normalization of inexplicability.It’s also tightly connected to a normalization of lack of accountability.> This isn't "getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" -- you still have to do the hard part.This is probably losing the younger portion of the audience…
-
-
2
Developer reports positive results using AI for ops automation
A commenter working on AI systems for operations automation pushes back on the article's pessimism, describing successful use of frontier LLMs for task automation with proper evaluation frameworks in place.
“Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.”
— oli5679 -
The "normalization of inexplicability" is indeed infuriating. It has always been bad when it comes to computer software, and it's increasingly creeping into other consumer products that depend on embedded software.I bought a new electric car recently. For the most part I've been quite happy with it. Shortly after I bought it, it started popping up…
-
-
1
Commenters challenge author's framing of cloud service reliability
Readers begin questioning the article's premises, with some arguing the author underestimates cloud services' actual accountability structures and the prevalence of existing debugging tools.
“I am betting author does not use cloud services much. It is not just "users"...”
— theamk -
> When a button breaks on a website, I have a model about what should have happened. Somewhere a contract got broken. [...] I might not have access to debug just an HTTP status 500, but I expect there to be somebody whose job is to understand why the endpoint is 500ing. The ownership is well-defined albeit opaque³.> For many users, however, the…
1 more of the top 2 · 2 posts in this stretch
-
If you spend more time with a product, you’re more likely to choose to do it again, even if it’s because of failure or annoyance. You’d justify it somehow (I have a leg up now or something). This actually applies to looking at things as well; a brightly colored box on the supermarket shelf is simply more likely to be chosen because you look at it…
-
-
background
Article published examining AI-driven software failure patterns — An article titled "The Normalization of Inexplicable Failures" is published, arguing that AI agents are enabling and normalizing software failures that lack clear explanations or debugging trails. The piece drew 150+ points and 41 comments on Hacker News, indicating significant reader engagement.
Also covered reported alongside — the timeline has no entry for these yet
-
first by HN Best, 1d ago · also HN Frontpage
What people are saying 5 voices from 1 site · best of 18 · verbatim
- Are developers who report success with AI tools measuring the right metrics, or are they missing silent failures and downstream damage?
- Should the onus be on individual organizations to implement rigorous testing, or is the scale of AI adoption creating market pressure that systematically undercuts quality?
- Yesterday
-
> they can always shrug and say "well, AI makes mistakes." Error budgets? Failure modes? Test sets? All of those can be handled later.This is just the complete opposite in my experience. Tests are the first thing the AI writes, especially in low coverage or unknown domain situation.These systems have been trained for "generations" to oneshot…
-
Amount of APIs failing for no reason and the answer is just retry these days is insaneI don't mind doing it but why is this the norm
-
I've seen people opt to refactor their entire codebase from one lang to another just because AI made it so much easier. Sure there were problems before, but now the new code with the better language is not readable and needs another refactor once this entire thing is done.
-
That was awfully specific. Apple has managed to normalize inexplicable failures of things that had been working before for over a decade, across all of their apps and the OS. And they didn't even need AI for that. But AI will definitely speed up the rate of blunder everywhere.
-
"Appeal to adult animated show" is lame enough on Reddit but it's particularly odious blogspam here.