conv.

All stories
AIFading · day 5

Users report Fable 5 performance decline weeks after launch

AI model shows degraded reasoning and task performance in August, sparking debate over deliberate degradation versus infrastructure optimization.

What to know

  • Multiple Fable 5 and frontier model users report clear performance degradation 2–8 weeks after launch, despite models performing well initially.
  • Users and commenters debate whether degradation is deliberate cost-saving via compute optimization or accidental side effect of model changes.
  • Counter-argument: metrics like token efficiency show optimization (fewer tokens for same results) rather than genuine capability loss; Anthropic has denied intentional performance manipulation.
  • Broader concern: if cycle of release-then-degrade is industry practice, regulatory scrutiny similar to product standards may be warranted.

The dispute Whether observed performance decline represents intentional cost-saving degradation, accidental optimization side effects, or measurement artifacts from token efficiency improvements. · positions read across 27 posts and comments

most voices

AI companies deliberately degrade models after launch to cut compute costs and force subscription upgrades.

  • “I for sure felt that this was the case for a while now, but couldn't explain it. Newly released feels great for the first couple of weeks, but then it starts to get worse.”

    zerof1l · Hacker News ↗
some voices

Performance changes reflect optimization (fewer tokens, better prompting) rather than degradation; newer models outperform on benchmarks and hard problems.

  • “Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the 'right' things or structure.”

    rcr-anti · Hacker News ↗
some voices

The pattern is endemic industry scam: vendors boost early metrics with over-provisioning, then gradually degrade to baseline while releasing marginally better replacements.

  • “The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too. Then weeks later people find out that they have been duped…”

    r2-129 · Hacker News ↗
some voices

Lack of transparency about model changes warrants regulatory oversight similar to product standards.

  • “AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected…”

    jesse_dot_id · Hacker News ↗

Waterluvian Fable 5 userjotato gpt-5.6-luna usermlmonkey Frontier model researchertheplumber Claude userrcr-anti Claude Code tracker

Users report Fable 5 performance decline weeks after launch
twitter.com

How it unfolded 6 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 13 pieces in two hours at Sep 21, 11 AM; 31 pieces over 5 days (2 articles · 2 posts · 27 comments) Sep 21, 11 AM — 13 pieces · 2 articles · 1 post · 10 comments — Hacker News 11, Newswires 2Sep 21, 1 PM — 4 pieces · 4 comments — Hacker News 4Sep 21, 3 PM — 3 pieces · 1 post · 2 comments — Hacker News 2, Mastodon 1Sep 21, 5 PM — 5 pieces · 5 comments — Hacker News 5Sep 21, 7 PM — 2 pieces · 2 comments — Hacker News 2Sep 21, 9 PM — 1 piece · 1 comment — Hacker News 1Sep 21, 11 PM — quietSep 22, 1 AM — 1 piece · 1 comment — Hacker News 1Sep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — 1 piece · 1 comment — Hacker News 1Sep 22, 11 AM — quietSep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — quietSep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — 1 piece · 1 comment — Hacker News 1Yesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quietToday, 9 AM — quiet 1–6
Sep 22Sep 23Sep 24yesterdaynow · 10:56 AM ET
  1. 6

    talon8635 proposes deliberate degradation cycle theory

    A commenter suggests a cycle in which AI companies release a new model, slowly degrade it over months, then release a marginally better replacement to create perceived improvement despite stagnant actual progress—a strategy that could sustain an industry facing frontier stagnation.

    “Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that's marginally if at all better than the original to create a perceived improvement…”
    — talon8635
    • > to create a perceived improvementIn addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more…

      mrandishHacker News4d agoview on Hacker News ↗
    2 more of the top 3 · 19 posts in this stretch
    • > Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open…

      AmazingTurtleHacker News4d agoview on Hacker News ↗
    • I've subjectively detected this in previous codex releases where the 2 days before release of a new model the agent went from great to me pulling my hair out yelling at it. I think it's just a win-win for them. They need to ramp up basic capacity and usage on the new model, what better way to free up capacity than to reduce the effort with the…

      sporklandHacker News17h agoview on Hacker News ↗
    all of them →
  2. 5

    jesse_dot_id calls for regulatory oversight of AI model consistency

    A commenter invokes the historical Office of Weights and Measures, arguing AI companies should face similar regulatory scrutiny to prevent selling inconsistent products to consumers.

    “AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected…”
    — jesse_dot_id
    • I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you…

      rcr-antiHacker News4d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is…

      jesse_dot_idHacker News4d agoview on Hacker News ↗
    all of them →
  3. 4

    Waterluvian reports immediate performance degradation versus Friday evening

    A commenter describes noticing something wrong with Fable 5 compared to Friday evening, citing a specific incident where the model duplicated a method it was asked to delete, then acknowledged the error. The user also notes increased "thinking" time for previously simple tasks.

    “I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.”
    — Waterluvian
    • I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of…

      WaterluvianHacker News4d agoview on Hacker News ↗
  4. 3

    jotato describes model degradation over 2-3 weeks

    A user reports that gpt-5.6-luna, which performed as well as the previous model version in its first week, has become noticeably worse in the last 2–3 weeks, requiring explicit prompting for tasks it previously handled implicitly.

    “over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.”
    — jotato
    • Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.For example, I used to be able to prompt "Check the system logs on <server>…

      jotatoHacker News4d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new…

      theplumberHacker News4d agoview on Hacker News ↗
    all of them →
  5. background

    theplumber attributes performance changes to Anthropic cost management — A user reports clear differences between new Claude releases and their state 3–4 weeks later, attributing this to Anthropic applying "auto degradation" to save compute on tasks deemed non-critical, and suggesting new accounts receive an intelligence boost.

  6. 2

    mlmonkey reports massive performance drop week 1 to week 8

    A heavy user of frontier models describes a dramatic performance decline from week 1 to week 8 after launch, with the model transitioning from capable research assistant to eager but less capable tool.

    “the drop in performance from, say, week 1 to week 8 is often massive…”
    — mlmonkey
    • Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for…

      mlmonkeyHacker News4d agoview on Hacker News ↗
  7. 1

    rcr-anti notes fewer tokens required for same results in Claude Code

    A tracker of Claude Code performance disputes the degradation narrative, observing that newer versions require fewer tokens to achieve the same or better results, suggesting optimized prompt design and tool ergonomics rather than genuine capability loss.

    “Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the 'right' things or structure.”
    — rcr-anti
    • Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.Buy decent coffee instead of your $200 subscription and…

      r2-129Hacker News4d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).I wonder what their official explanation for this behavior is.

      alexjplantHacker News4d agoview on Hacker News ↗
    all of them →
  8. background

    Users report Fable 5 median thinking declined in August — A post titled "Fable 5 – Median thinking declined in August" surfaces on Hacker News, drawing attention to reports that the model's performance has degraded significantly since its August launch. Multiple users share anecdotal experiences of reduced capability.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by HN Best, 4d ago · also HN Frontpage

What people are saying 13 voices from 1 site · best of 27 · verbatim

Still unanswered
  • Do AI companies have an official explanation for the documented week-to-week performance variations?
  • Are benchmarks capturing real-world capability loss, or only token-count metrics that may hide optimization benefits?
  • Is this pattern industry-wide or specific to certain vendors, and if universal, does it suggest systemic business incentive?