conv.

All stories
TechActive · 34h

Article on AI-driven software failures sparks debate over quality standards

A piece examining how AI agents are normalizing inexplicable software failures prompts disagreement over whether the tradeoffs are acceptable.

What to know

  • An article critiques how AI agents enable and normalize opaque software failures, particularly in infrastructure and libraries—layers whose reliability affects all downstream systems.
  • Defenders argue that with proper evaluation frameworks, AI tools improve on human-level work; critics counter that most organizations lack rigor and rationalize persistent failure rates instead of fixing them.
  • The deeper debate centers on misaligned incentives: whether AI makes a pre-existing speed-over-quality problem worse, and whether confidence scores mislead stakeholders into false confidence.

The dispute Whether the article identifies a genuine systemic risk (failures in critical layers) or merely critiques poor implementation practices by organizations that lack rigor—and whether developers' production results validate AI tools or hide cascading problems. · positions read across 18 posts and comments

many voices

AI tools can be safe and productive when paired with rigorous evaluation and testing frameworks.

  • “Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.”

    oli5679 · Hacker News ↗
many voices

Accepting AI-driven failures in infrastructure and libraries cascades into systemic unreliability affecting all users.

  • “But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.”

    adamddev1 · Hacker News ↗
some voices

The real problem is perverse incentives toward speed over quality; AI is enabling but not causing this.

  • “The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather rare now - thirty years ago, Microsoft won the office wars with crap that had features.”

    joe_the_user · Hacker News ↗
some voices

Confidence scores mislead stakeholders and obscure actual failure modes through false authority.

  • “An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal…”

    WorldMaker · Hacker News ↗

ihatethefuture.com author Article authoroli5679 AI systems engineer (ops automation)adamddev1 Commenterbenjaminsky2 Developer with validation experience

How it unfolded 7 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 4 pieces in one half hour at Yesterday, 1 PM; 24 pieces over 35 hours (2 articles · 4 posts · 18 comments) Yesterday, 11:13 AM — 3 pieces · 2 articles · 1 post — Newswires 2, Hacker News 1Yesterday, 11:43 AM — 1 piece · 1 comment — Hacker News 1Yesterday, 12:13 PM — 3 pieces · 3 comments — Hacker News 3Yesterday, 12:43 PM — 3 pieces · 3 comments — Hacker News 3Yesterday, 1:13 PM — 2 pieces · 1 post · 1 comment — Hacker News 1, Mastodon 1Yesterday, 1:43 PM — 4 pieces · 4 comments — Hacker News 4Yesterday, 2:13 PM — quietYesterday, 2:43 PM — quietYesterday, 3:13 PM — quietYesterday, 3:43 PM — 2 pieces · 1 post · 1 comment — Hacker News 1, Mastodon 1Yesterday, 4:13 PM — quietYesterday, 4:43 PM — 1 piece · 1 comment — Hacker News 1Yesterday, 5:13 PM — quietYesterday, 5:43 PM — 2 pieces · 2 comments — Hacker News 2Yesterday, 6:13 PM — quietYesterday, 6:43 PM — quietYesterday, 7:13 PM — quietYesterday, 7:43 PM — quietYesterday, 8:13 PM — quietYesterday, 8:43 PM — quietYesterday, 9:13 PM — quietYesterday, 9:43 PM — 1 piece · 1 comment — Hacker News 1Yesterday, 10:13 PM — quietYesterday, 10:43 PM — quietYesterday, 11:13 PM — quietYesterday, 11:43 PM — quietToday, 12:13 AM — quietToday, 12:43 AM — quietToday, 1:13 AM — quietToday, 1:43 AM — quietToday, 2:13 AM — quietToday, 2:43 AM — quietToday, 3:13 AM — quietToday, 3:43 AM — quietToday, 4:13 AM — quietToday, 4:43 AM — quietToday, 5:13 AM — quietToday, 5:43 AM — quietToday, 6:13 AM — quietToday, 6:43 AM — quietToday, 7:13 AM — quietToday, 7:43 AM — quietToday, 8:13 AM — quietToday, 8:43 AM — quietToday, 9:13 AM — quietToday, 9:43 AM — quietToday, 10:13 AM — quietToday, 10:43 AM — quietToday, 11:13 AM — 1 piece · 1 post — Lobsters 1Today, 11:43 AM — quietToday, 12:13 PM — quietToday, 12:43 PM — quietToday, 1:13 PM — quietToday, 1:43 PM — quietToday, 2:13 PM — quietToday, 2:43 PM — quietToday, 3:13 PM — quietToday, 3:43 PM — quietToday, 4:13 PM — quietToday, 4:43 PM — quietToday, 5:13 PM — quietToday, 5:43 PM — quietToday, 6:13 PM — 1 piece · 1 comment — Hacker News 1Today, 6:43 PM — quietToday, 7:13 PM — quietToday, 7:43 PM — quietToday, 8:13 PM — quietToday, 8:43 PM — quietToday, 9:13 PM — quietToday, 9:43 PM — quiet 1–7
4 PMtoday8 AM4 PMnow · 10:13 PM ET
  1. 7

    Developer challenges article without empirical counterexamples

    A commenter reports personally validating AI confidence scores and seeing linear accuracy scaling, claiming the article makes critiques without providing concrete failure examples to back them up.

    “Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.”
    — benjaminsky2
    • Reading this article put me in mind of how everyone says "things aren't made like they used to be." Someone was just telling me this above a washing machine. I think we've seen that cheap and fast won and the lower quality became normal. I think it's fair to say that there is a fear software is heading the same way. AI is going to help people ship…

      BrkAway21Hacker News3h agoview on Hacker News ↗
    2 more of the top 3 · 8 posts in this stretch
    • I build ai systems for ops automation and don’t understand the author’s pessimism.I agree with that Evals are a scarce commodity rn. A business needs to define what good looks like. This is a laborious, and sometimes politically controversial, process.Given good Evals, frontier llms are a magical tool that can automate tasks and do them more…

      oli5679Hacker News1d agoview on Hacker News ↗
    • > Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.I have definitely seen more bafflingly-poor OSS software that just plain doesn't work frequently now than before.But it's mostly software that wouldn't have existed before because it's trying to do super-niche things. So on the "hobby"…

      majormajorHacker News1d agoview on Hacker News ↗
    all of them →
  2. 6

    Commenter frames the core issue as perverse incentives toward speed

    A voice argues the underlying problem predates AI—that market incentives favor speed over quality—but AI tools are now enabling this trend to accelerate and worsen the outcomes.

    “The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather rare now - thirty years ago, Microsoft won the office wars with crap that had features.”
    — joe_the_user
    • Thing is, the unreliable-software situation was already untenable before agents (in poor hands) made it worse.Yes but that's the big thing, now isn't it? These are nice tools, used wisely. But their unwise use, oh boy...The problem is one needs to be in a situation where the incentive is towards quality rather than speed. But that situation rather…

      joe_the_userHacker News1d agoview on Hacker News ↗
  3. 5

    Commenter criticizes acceptance of persistent failure rates

    A voice argues that developers and organizations are increasingly willing to accept and rationalize persistent failure rates (around 10%) rather than addressing root causes, treating AI as an excuse.

    “most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them.”
    — lokar
    • I don’t think the author (or many people) doubt that one can (and some will) find a way that does not “suck”But it’s pretty clear that most people are not. For whatever reasons (mgmt pressure, trying to get ahead, skill issues, etc) they half ass it, accept the 10% (silent) fail rate and blame the bad outcomes on the AI as if that absolves them…

      lokarHacker News1d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • I’ve validated Jev’s confidence score. Accuracy scales linearly with confidence for the 3 use cases I tested. >.9 it matched a human labeler. I immediately discovered a user behavior I didn’t expect for ~$3. I can now mitigate in real-time due to low cost and latency. This may have a major positive financial impact for all our customers.Not sure…

      benjaminsky2Hacker News1d agoview on Hacker News ↗
    all of them →
  4. 4

    Commenters identify confidence score interpretation as a core problem

    Multiple voices highlight how AI confidence scores mislead business stakeholders into false assumptions about reliability, linking the issue to broader misuse of statistics in ML deployment decisions.

    “An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal percentage".”
    — WorldMaker
    • "Confidence scores" have always implied an anthopocentric meaning that doesn't exist. An algorithm doesn't have "confidence" in the way that a person has confidence, but as soon you put something with that name in front of a business person they assume the number is always a meaningful "letter grade curve" or "universal percentage". I still…

      WorldMakerHacker News1d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • I am big on reproducibility (nix aficionado) and determinism (flagging test failures are a red-alert, all-hands-on-deck situation in my world) and correctness.I am also big on testing (the correct things). And nine-nines (big on Elixir).And... I'm also big on agent-assisted dev. Which requires pretty much every check in the book to stay productive…

      pmarreckHacker News1d agoview on Hacker News ↗
    all of them →
  5. 3

    Concern raised about normalizing failures in infrastructure layers

    A commenter warns that accepting AI-generated failures at the library and infrastructure level could cascade, degrading reliability for all downstream systems and users.

    “But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows EVERYTHING and EVERYONE down.”
    — adamddev1
    • Excellent post. People always defend agentic/LLM-driven development by saying, "Well it's good enough", or "It works most of the time."That may be tolerable for some user-facing app. But what if we start normalizing failures in the libraries, the infrastructure, and the compilers? Everything descends into a mess of unreliability, and that slows…

      adamddev1Hacker News1d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • > This leads to a normalization of inexplicability.It’s also tightly connected to a normalization of lack of accountability.> This isn't "getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" -- you still have to do the hard part.This is probably losing the younger portion of the audience…

      layer8Hacker News1d agoview on Hacker News ↗
    all of them →
  6. 2

    Developer reports positive results using AI for ops automation

    A commenter working on AI systems for operations automation pushes back on the article's pessimism, describing successful use of frontier LLMs for task automation with proper evaluation frameworks in place.

    “Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.”
    — oli5679
    • The "normalization of inexplicability" is indeed infuriating. It has always been bad when it comes to computer software, and it's increasingly creeping into other consumer products that depend on embedded software.I bought a new electric car recently. For the most part I've been quite happy with it. Shortly after I bought it, it started popping up…

      teraflopHacker News1d agoview on Hacker News ↗
  7. 1

    Commenters challenge author's framing of cloud service reliability

    Readers begin questioning the article's premises, with some arguing the author underestimates cloud services' actual accountability structures and the prevalence of existing debugging tools.

    “I am betting author does not use cloud services much. It is not just "users"...”
    — theamk
    • > When a button breaks on a website, I have a model about what should have happened. Somewhere a contract got broken. [...] I might not have access to debug just an HTTP status 500, but I expect there to be somebody whose job is to understand why the endpoint is 500ing. The ownership is well-defined albeit opaque³.> For many users, however, the…

      theamkHacker News1d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • If you spend more time with a product, you’re more likely to choose to do it again, even if it’s because of failure or annoyance. You’d justify it somehow (I have a leg up now or something). This actually applies to looking at things as well; a brightly colored box on the supermarket shelf is simply more likely to be chosen because you look at it…

      hyperhelloHacker News1d agoview on Hacker News ↗
    all of them →
  8. background

    Article published examining AI-driven software failure patterns — An article titled "The Normalization of Inexplicable Failures" is published, arguing that AI agents are enabling and normalizing software failures that lack clear explanations or debugging trails. The piece drew 150+ points and 41 comments on Hacker News, indicating significant reader engagement.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by HN Best, 1d ago · also HN Frontpage

What people are saying 5 voices from 1 site · best of 18 · verbatim

Still unanswered
  • Are developers who report success with AI tools measuring the right metrics, or are they missing silent failures and downstream damage?
  • Should the onus be on individual organizations to implement rigorous testing, or is the scale of AI adoption creating market pressure that systematically undercuts quality?