Users report Fable 5 performance decline weeks after launch
AI model shows degraded reasoning and task performance in August, sparking debate over deliberate degradation versus infrastructure optimization.
What to know
- Multiple Fable 5 and frontier model users report clear performance degradation 2–8 weeks after launch, despite models performing well initially.
- Users and commenters debate whether degradation is deliberate cost-saving via compute optimization or accidental side effect of model changes.
- Counter-argument: metrics like token efficiency show optimization (fewer tokens for same results) rather than genuine capability loss; Anthropic has denied intentional performance manipulation.
- Broader concern: if cycle of release-then-degrade is industry practice, regulatory scrutiny similar to product standards may be warranted.
The dispute Whether observed performance decline represents intentional cost-saving degradation, accidental optimization side effects, or measurement artifacts from token efficiency improvements. · positions read across 27 posts and comments
AI companies deliberately degrade models after launch to cut compute costs and force subscription upgrades.
-
“I for sure felt that this was the case for a while now, but couldn't explain it. Newly released feels great for the first couple of weeks, but then it starts to get worse.”
zerof1l · Hacker News ↗
Performance changes reflect optimization (fewer tokens, better prompting) rather than degradation; newer models outperform on benchmarks and hard problems.
-
“Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the 'right' things or structure.”
rcr-anti · Hacker News ↗
The pattern is endemic industry scam: vendors boost early metrics with over-provisioning, then gradually degrade to baseline while releasing marginally better replacements.
-
“The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too. Then weeks later people find out that they have been duped…”
r2-129 · Hacker News ↗
Lack of transparency about model changes warrants regulatory oversight similar to product standards.
-
“AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected…”
jesse_dot_id · Hacker News ↗
Waterluvian Fable 5 userjotato gpt-5.6-luna usermlmonkey Frontier model researchertheplumber Claude userrcr-anti Claude Code tracker
How it unfolded 6 developments, newest first · click a bar or a number to jump articlespostscomments
-
6
talon8635 proposes deliberate degradation cycle theory
A commenter suggests a cycle in which AI companies release a new model, slowly degrade it over months, then release a marginally better replacement to create perceived improvement despite stagnant actual progress—a strategy that could sustain an industry facing frontier stagnation.
“Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that's marginally if at all better than the original to create a perceived improvement…”
— talon8635 -
> to create a perceived improvementIn addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more…
2 more of the top 3 · 19 posts in this stretch
-
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open…
-
I've subjectively detected this in previous codex releases where the 2 days before release of a new model the agent went from great to me pulling my hair out yelling at it. I think it's just a win-win for them. They need to ramp up basic capacity and usage on the new model, what better way to free up capacity than to reduce the effort with the…
-
-
5
jesse_dot_id calls for regulatory oversight of AI model consistency
A commenter invokes the historical Office of Weights and Measures, arguing AI companies should face similar regulatory scrutiny to prevent selling inconsistent products to consumers.
“AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected…”
— jesse_dot_id -
I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you…
1 more of the top 2 · 2 posts in this stretch
-
The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is…
-
-
4
Waterluvian reports immediate performance degradation versus Friday evening
A commenter describes noticing something wrong with Fable 5 compared to Friday evening, citing a specific incident where the model duplicated a method it was asked to delete, then acknowledged the error. The user also notes increased "thinking" time for previously simple tasks.
“I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.”
— Waterluvian -
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of…
-
-
3
jotato describes model degradation over 2-3 weeks
A user reports that gpt-5.6-luna, which performed as well as the previous model version in its first week, has become noticeably worse in the last 2–3 weeks, requiring explicit prompting for tasks it previously handled implicitly.
“over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.”
— jotato -
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.For example, I used to be able to prompt "Check the system logs on <server>…
1 more of the top 2 · 2 posts in this stretch
-
It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new…
-
-
background
theplumber attributes performance changes to Anthropic cost management — A user reports clear differences between new Claude releases and their state 3–4 weeks later, attributing this to Anthropic applying "auto degradation" to save compute on tasks deemed non-critical, and suggesting new accounts receive an intelligence boost.
-
2
mlmonkey reports massive performance drop week 1 to week 8
A heavy user of frontier models describes a dramatic performance decline from week 1 to week 8 after launch, with the model transitioning from capable research assistant to eager but less capable tool.
“the drop in performance from, say, week 1 to week 8 is often massive…”
— mlmonkey -
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for…
-
-
1
rcr-anti notes fewer tokens required for same results in Claude Code
A tracker of Claude Code performance disputes the degradation narrative, observing that newer versions require fewer tokens to achieve the same or better results, suggesting optimized prompt design and tool ergonomics rather than genuine capability loss.
“Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the 'right' things or structure.”
— rcr-anti -
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.Buy decent coffee instead of your $200 subscription and…
1 more of the top 2 · 2 posts in this stretch
-
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).I wonder what their official explanation for this behavior is.
-
-
background
Users report Fable 5 median thinking declined in August — A post titled "Fable 5 – Median thinking declined in August" surfaces on Hacker News, drawing attention to reports that the model's performance has degraded significantly since its August launch. Multiple users share anecdotal experiences of reduced capability.
Also covered reported alongside — the timeline has no entry for these yet
-
first by HN Best, 4d ago · also HN Frontpage
What people are saying 13 voices from 1 site · best of 27 · verbatim
- Do AI companies have an official explanation for the documented week-to-week performance variations?
- Are benchmarks capturing real-world capability loss, or only token-count metrics that may hide optimization benefits?
- Is this pattern industry-wide or specific to certain vendors, and if universal, does it suggest systemic business incentive?
- Sep 22
-
Not sure on the methodology here, but could this also just because people's projects mature on the new model which results in less thinking?I've personally seen people complain about new models getting slow after a week or so but I think it's mostly because the new model is able to briefly overcome the context collapse of their poorly managed…
-
I stopped reading when I realized the article reminds me of my own Claude-generated solutions at work - just an endless maze of special business logic on top of special business logic. You need an LLM to understand it. You need an LLM help write the documentation. You need an LLM to help read the documentation.
- Sep 21
-
All llm api providers should be compelled to return a checksum-like proof of quantization level of the model that served the request. Basic transparency should be the bare minimum.
-
AI just solved a millennium problem two weeks ago. "The pace is insane. And there is no reason to be this fast." to quote Terence Tao word by word.HN: Well, must be a stagnant industry...
-
There was a coding horror story I read some years ago where a developer bragged that he improved performance by artificially increasing iterations on some critical path in an app and then lowering the iterations occasionally while bragging to management about squeezing out more performance.Kind of reminds me of that, but with more smoke and mirrors
-
Definitely noticed this before, but this parti6 time was very noticeable. I'm convinced it's to get the benchmarks in, then lower cost and prepare for the next release to look better relatively to users.
-
> For an industry that’s stagnant in progressSurely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
-
No, what they are doing is trying to optimize inference to increase margins which leads to degradations. Model deployment is not like websites, you can continuously tune performance based on usage, new memory optimizations, etc.
-
I for sure felt that this was the case for a while now, but couldn’t explain it. Newly released feels great for the first couple of weeks, but then it starts to get worse.
-
This is exactly what I have been experiencing and the difference is night and day! We have been advertised and given a taste of what Fable was and after that been served an exteme watered down version. It is so bad that sometimes chatgpt feels better.
-
> to create a perceived improvement when in reality there isn’t really one?This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
-
There must be some benefit if all the providers are doing it independently.GPT5.6-Sol on Max thinking just became regarded as of a few days ago.The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).The cycle repeats.
-
This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason