Grok, ChatGPT, Claude, and Gemini suffer coordinated outages
Multiple major AI services went down simultaneously on Wednesday, sparking speculation about shared infrastructure failures.
What to know
- Four major AI services—Grok, ChatGPT, Claude, and Gemini—went down simultaneously on September 3, 2026, at around 15:26Z–15:45Z.
- Root cause remains unclear; users speculate about shared datacenter infrastructure, cascading traffic surges as users switched between services, or shared cloud dependencies.
- The outage exposed the concentration risk of LLM infrastructure and triggered discussion about distributed compute alternatives.
The dispute Whether the outages stem from shared infrastructure/cascading load (technical) versus systemic AI risk and centralized compute fragility (structural), with some questioning the reliability of AI platforms as utilities. · positions read across 85 posts and comments
Shared infrastructure or cascading load surge caused the outages.
-
“Maybe all of them use the same datacenter. Which has just experienced issues.”
pbasista · Hacker News ↗
Services secretly depend on each other (OpenAI queries Claude, Grok queries ChatGPT), creating a chain-of-failures.
-
“It's interesting that ChatGPT is down as well. SpaceX rents to Anthropic but not to OpenAI. Just a coincidence or something more to it?”
bluecalm · Hacker News ↗
The outage reflects broader infrastructure problems and workforce degradation in tech.
-
“Is it happening... meaning these tech companies have laid off too many experienced people, and their services are starting to degrade?”
goache · Hacker News
xAI Grok service providerOpenAI ChatGPT service providerAnthropic Claude service providerGoogle Gemini service provider
The record 1 articles and posts · last 25 days
What people are saying 24 voices from 1 site · best of 85 · verbatim
- What is the actual root cause of the coordinated outage?
- Do these AI services share underlying datacenter or cloud infrastructure?
- What does this reveal about the resilience and centralization risk of large AI platforms?
- Sep 4
-
In the future (ha ha ha) when everyone has become totally LLM dependent a service outage will cause all of them to stop walking and talking and otherwise functioning and their heads will simply droop down while they stand or sit silently in place.
- Sep 3
-
Posted by @SpaceXAI 1h ago: https://x.com/SpaceXAI/status/2095597264043717014> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.> All systems have now been restored and are functioning nominally.---It's…
-
nobody is paying engineers for "capable of", they're paying for performance and results..if there was no difference in performance or results, nobody would be using LLMs or agent harnesses like claude code
-
Claude and Codex were down too. A 3 month old github issue that was the top result for the error that Codex threw (a 404 of all things) briefly became a live chat for bored and annoyed AI users.
-
A lot of times you run into senior developers are even staff engineers who have no idea how the big picture works like they don’t know what an IP address is.I bet AI coding leaves them (and even the rest of us!) productive in areas where we’re in way over our heads, and we could just completely sink without it.
-
> as if you're supposed to buy like 8 H200s or something or code by hand is the solutionPeople like GGP who have been coding for 20 years presumably are perfectly capable of doing so by hand. Or should be.
-
It’s not even a sensible comparison.You run way better models, whatever is newest, for the five years it takes for the “spark cluster” to even reach cost parity with far worse performance.Those things aren’t spectacular at inference… it’s not really why you drop that kind of money on them.
-
;-)as if you're supposed to buy like 8 H200s or something or code by hand is the solutionI don't get how this is a gotcha, there simply is no world where everyone has their own self hosted frontier model
-
Microsoft had extensive issues earlier in the week. Anecdotally it feels like internet has been noticeably slower this week across sites, networks and devices. Now this. IDK, has anyone else noticed anything this week?
-
I don't think regular people did any 180 degree turns. Rather big corporations again demonstrated their total hypocrisy where they on one hand stomp on people for "infringing their intellectual property" while at the same time scrapping all data they can get their tendrils on, with total disregards for the wishes of the authors of the data. Even…
-
Some napkin math - total Internet bandwidth is estimated to be 3-4 petabits per second. Let's say 1% of that is inference traffic - 35 terabits - then that's ~4.375 trillion characters per second at 8 bits per character. A single keyboard press might require 2 mJ, so you need roughly 8.75 GW (electricity required for ~9 million homes) of…
-
I still don't understand why people are giving money to a white supremacist whose GenAI tool has been reported to happily produce questionable images and naming itself after a certain Austrian painter.
-
So you spent the cost of my entire cluster in one day on Sol tokens lol. I have no need for that many tokens, a few million a day is perfectly acceptable. But if you do, then yeah Sol is probably your bet. Have fun :)
-
It has nothing to do with "their work", and the opposition to AI copyright usage comes much more from non-devs than devs. The actual answer is that people are anti-megacorporation and pro-individual. They are fine with copyright that protects individual rights and opposed to copyright that enshrines megacorporation rights. It's not that difficult…
-
The mainframe idea is the problem.Whenever I think about AI taking over the world I think about this really neat 1970 movie 'Colossus: The Forbin Project',
-
Since we are talking for other people, "devs" wouldn't have supported someone downloading all those copyrighted works from the pirate bay and profited by, for example, selling them. I don't see any 180.
-
Sure you can. It'll just take a little while to shake the rust off. It's like riding a bike.Or maybe it's all gone forever and we're all brain damaged now. Spooky! I think there's a pretty low probability of that, and even if it was true, worrying doesn't help!
-
> stole the entire written output of the entire human speciesDevs used to love public domain works, hated copyright, vehemently opposed software patents, held the pirates side during the MP3 wars, are very supportive of ThePirateBay and insist the correct term is "copyright infringement", not "stealing", and that "stealing" is a egregious form of…
-
> "loyalty"? weird word choice.All I mean is that none of the frontier LLMs are significantly better than the others for the vast majority of work, so devs are free to move between providers freely.
-
"loyalty"? weird word choice.Devs switch because all of the magic of "AI" is LLMs + we stole the entire written output of the entire human species, with some remaining % coming from various tricks we've learned over the last 3 years like thinking and agent-toolcall loops, which weren't hard for literally everyone to copy. One of the consequences…
-
Did the society find the Reset Nexus?I mean, yeah it's probably a data center outage that cascaded to non-related providers due to companeis switching when their preferred was down.But that openai society definitely would have spun up their own agents if they had access to do so.Clearly, it is centralized AI that is the biggest risk to…
-
Seems like openai is recovering. But even if it's only 30 minutes, it's a pretty interesting experiment.I'm only half joking when I say that I'll write code by hand for money.
-
It's not the singularity.One of the major LLM providers goes down for some reason. Traffic shifts to the other providers because devs have no loyalty. LLMS are commodities. The other providers can't handle the increase in traffic and they go down as well.
-
ChatGPT here has been giving me Cloudflare errors, and Cloudflare reports that it is having issues: https://www.cloudflarestatus.com/Not sure if they are linked.