62 points halcdev 52 minutes ago 44 comments

https://status.openai.com https://status.claude.com https://status.x.ai

convivialdingo 48 minutes ago | parent

The Thundering Herd has thundered, apparently.

codazoda 47 minutes ago | parent

I kinda assume it's because one went down and a large amount of work shifted to another.

I'm also aware that they have overlap in some areas on data centers.

kocial 37 minutes ago | parent

Maybe the stack behind it is down, like AWS or something

maxbaines 37 minutes ago | parent

They all rent compute from SpaceXAI

halcdev 31 minutes ago | parent

Surely it's a bit more distributed than that, right?

tehbeard 21 minutes ago | parent

bfung 20 minutes ago | parent

Like how AWS has global datacenters, but everyone uses us-east-1.

lavezzi 20 minutes ago | parent

I don't believe OpenAI does

maxbaines 18 minutes ago | parent

My mistake, in fact it was google not OpenAI, makes sense OpenAI doesn't.

CSMastermind 31 minutes ago | parent

I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.

wejick 30 minutes ago | parent

Probably same public cloud or CDN in front of them.

Insanity 29 minutes ago | parent

Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

toomuchtodo 25 minutes ago | parent

https://en.wikipedia.org/wiki/Domino_effect

Edit: Updated per valleyer's suggestion.

valleyer 20 minutes ago | parent

"Domino effect" would probably be the more relevant named phenomenon there.

throwaway894345 19 minutes ago | parent

This isn’t a thundering herd problem, it’s a cascading failure. (Thundering herd is about a bunch of workers waking up simultaneously)

erdos_2 15 minutes ago | parent

It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.

rtcoms 12 minutes ago | parent

Just now I got this from gemini

It looks like there's no response available for this search. Try asking something else.

exe34 4 minutes ago | parent

I bet they had to implement that manually to make it look like they failed too!

Insanity 9 minutes ago | parent

Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)

benatkin 8 minutes ago | parent

Not even the best agent that starts with a G

paxys 3 minutes ago | parent

Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.

elorant 28 minutes ago | parent

Some npm library that makes headers bold would be broken.

N_Lens 25 minutes ago | parent

Ah yes ye olde bold-headers: ^3.13.31;

ibejoeb 6 minutes ago | parent

Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.

dgellow 26 minutes ago | parent

Too early to know, let’s wait and see

Razengan 19 minutes ago | parent

SkyNet is arming..

niobe 19 minutes ago | parent

Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

docheinestages 16 minutes ago | parent

My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.

elar_verole 15 minutes ago | parent

Pretty sure it's a US thing since it's available here in France. What exactly is down, idk

satvikpendem 14 minutes ago | parent

They're all using Cloudflare.

ratelimitsteve 13 minutes ago | parent

everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.

jedbrooke 11 minutes ago | parent

according to https://downdetector.com/ Gemini is down too (and copilot, but that just uses ChatGPT right?)

Linello 10 minutes ago | parent

What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?

cyptus 4 minutes ago | parent

at this point: gg

misano 8 minutes ago | parent

The IRGC has cut the fiber-optic cables in the Strait of Hormuz. LOL

fidla 7 minutes ago | parent

ChatGPT is up

fidla 7 minutes ago | parent

chatgpt is back

faitswulff 7 minutes ago | parent

Heard on the grape vine that the OpenAI blip was a cloudflare issue

chasd00 7 minutes ago | parent

claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.

aslkalska 7 minutes ago | parent

they all rent compute from each other

Avicebron 4 minutes ago | parent

I suspect Azure is having issues, Microsoft has had outages the paat two days, especially with email.

morkalork 4 minutes ago | parent

Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over