ChatGPT, Claude and Grok Went Down Together. Was It Azure, AGI or Something Else?
At 13:26 UTC on Thursday, September 3, 2026, Anthropic’s status page began reporting elevated errors on its newest Claude models. Within the hour Grok stopped answering on X, and at 14:43 UTC OpenAI’s status page turned red for ChatGPT and Codex. For roughly ninety minutes three of the leading AI assistants in the United States were failing at once, and by the time the last of them recovered the internet had produced a full set of explanations: a Microsoft Azure region had collapsed, Cloudflare had done it again, the new GPT-6 model had eaten its own data centre, or the machines had simply woken up and decided they had had enough.
Five days later, none of the three companies has published a root-cause report, and the public record still contains three different explanations rather than one. That is worth spelling out, because the wrong version of this story has been repeated by outlets that should know better, and because the right version is more useful than any of the theories. A simultaneous outage is not a shared outage. On September 3, three companies failed inside the same window for three different stated reasons, and nothing on the public record joins them.
Here is what each theory claims, what the evidence says, and why the question matters more than whichever answer turns out to be true.
What Actually Happened on September 3?
The cleanest sources are the status pages, because they are timestamped by the companies themselves.
| Service | First status update (UTC) | Stated cause | Resolved (UTC) |
|---|---|---|---|
| Claude (Anthropic) | 13:26, "elevated errors" on Claude Mythos 5.1, Fable 5.1 and Opus 5 | "An infrastructure issue" (spokesperson, no detail) | 16:23 (impact ended 16:16) |
| Grok (xAI / SpaceX) | About 13:30 per xAI's status record; Grok models degraded in GitHub Copilot from 14:17 | "An outage at our Memphis compute center" (SpaceX) | About 17:05; Copilot's Grok models restored 17:11 |
| ChatGPT and Codex (OpenAI) | 14:43, "Elevated errors across ChatGPT and Codex", 19 components | "A routing error starting around 7:43 am PT" (spokesperson) | 16:55 (mitigation applied 15:17) |
Anthropic’s incident log says the company “identified the cause” at 13:41 UTC, fifteen minutes in, then spent two and a half hours working on a fix before closing the incident. The cause was never named. Its spokesperson told The Register that “an infrastructure issue caused a partial outage across Claude.ai, Claude Code, Claude Cowork, and the Claude API” and that “service was restored at 16:16 UTC”.
OpenAI’s incident page lists 15 ChatGPT components and 4 Codex components, from login and conversations to voice mode and file uploads. The company’s spokesperson, Kathleen Chaykowski, gave Wired and The Register the same sentence: “A routing error starting around 7:43 am PT on Thursday, September 3 made ChatGPT and Codex unavailable for some users across platforms. As of about 8:17 am PT on Thursday, a solution was successfully implemented and is continuing to be monitored.” In UTC that is 14:43 to 15:17, a 34-minute fault that sat entirely inside the Anthropic and xAI windows.
Grok’s account came from SpaceX, which absorbed xAI earlier this year, in a social post quoted by The Register: “We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.” Elon Musk added that the company was “taking corrective action to ensure this does not happen again”, per Engadget. xAI’s status page, as quoted by Wired, opened its incident at 6:30 am Pacific and closed it at 10:05, a three-hour, 35-minute outage. GitHub’s status page recorded Grok models as degraded inside Copilot from 14:17 to 17:11 UTC “due to an issue with an upstream model provider”, and Cursor, the coding editor, logged a degradation of “All Grok Models” from 13:41 to 17:07 UTC. Those are the clearest third-party timestamps we have for the Grok incident, and they show how the dependency chain runs: when a model provider fails, the products built on it fail a few minutes later.
The scale, in crowd terms: Decrypt put Downdetector’s peaks at about 38,000 reports for ChatGPT and roughly 1,400 each for Claude and Grok. Perplexity, Mistral and DeepSeek showed no incident on their status pages during the outage window, and Z.ai posted “We’re still up.”
Gemini is the interesting absence. Ars Technica counted it as a fourth affected service on the strength of Downdetector reports rising from about 23 to 412 and a StatusGator note of a “likely Gemini API outage” between 10:45 and 11:15 am Eastern. Google’s Workspace status dashboard, which covers the Gemini app, recorded nothing; the only acknowledgement was a note on the Google AI Studio status page, quoted by LADbible, that the Gemini API had been having “problems serving recently created API keys, including through OpenAI-compatible libraries”. 9to5Google, which watches Google closely, wrote that “Gemini appears to be unaffected”. Keep that in mind, because Gemini becomes the control group in the first and loudest theory.
Did an Azure Outage Take Down ChatGPT, Claude and Grok?
This is the theory that reached headlines. Tech Times ran it on the day as “Gemini Survived When ChatGPT, Claude, and Grok Collapsed: Azure Is at Fault”. Computing followed with “Azure failure likely brought down ChatGPT, Claude and Grok”, built on StatusGator and Downdetector signals of “ingress failures” in Azure’s East US region. Shattered.io turned it into a 90-minute narrative, and in France, developpez.com told its readers that a failure in Azure East US “was the common denominator of this cascading outage”.
The theory is attractive for a reason. OpenAI and Anthropic both run substantial workloads on Azure under multi-billion-dollar agreements, Gemini runs on Google’s own cloud and stayed up, and everyone remembers October 29, 2025, when a configuration fault in Azure Front Door took Microsoft 365, the Azure portal and Alaska Airlines’ booking systems offline for more than eight hours. A shared cloud dependency is the obvious first suspect, and the argument’s logic, “the one on a different cloud survived”, sounds like science.
It fails on the evidence in four places.
First, Microsoft’s own record. Azure’s status history, checked on September 8, lists no incident at all for September 2026; the most recent post-incident review is dated July 23. Wired reported that Cloudflare, Amazon Web Services and Microsoft Azure “did not report outages on Thursday”, and 9to5Google, which had raised the Azure possibility itself, appended a one-line denial without naming a spokesperson: “Microsoft says that this is not the case”. Ars Technica noted that Downdetector reports “did spike somewhat” for AWS, Azure and Cloudflare that morning, which is exactly what crowd-report sites do during any large outage: people whose AI tool has failed report every service they can think of. Microsoft 365 Copilot Chat, the assistant that actually lives on Azure, logged Microsoft’s only official Copilot incident of the day at 17:36 UTC, after all three AI incidents had closed.
Second, the evidence itself. Trace the “Azure East US ingress failures” back through Tech Times and you arrive at a single user-submitted note on StatusGator, a third-party monitor: “EAST US ingress down since 10:26AM PT. Working with MSFT support, they reported upgrades at that time on their side.” That is one anonymous customer’s ticket, it no longer appears on StatusGator’s live page, and the time it gives, 10:26 am PT or 17:26 UTC, falls after Claude had been restored for over an hour, after ChatGPT’s incident had closed, and after Grok’s models were back in Copilot.
Third, Grok. SpaceX placed the Grok failure in its own Colossus site in Memphis, Tennessee, not in any Azure region. On the companies’ own account the theory cannot reach Grok, so at best it would explain two of the three, and those two have been explained differently by the companies concerned.
Fourth, and least noticed, the companies did name causes. “A routing error” and “an infrastructure issue” are vague, but they are the providers’ own words and neither points at a cloud region. Our rule, and it should be everyone’s, is that an outage has the cause its operator states until the operator states otherwise.
Verdict: not supported. The Azure story is a crowd-report artefact dressed up as a root cause.
Was It Cloudflare, DNS or Another Shared Pipe?
Cloudflare was the second suspect, for a good technical reason. ChatGPT’s error pages carried Cloudflare “cf-ray” identifiers, and in a Hacker News thread that ran past 700 comments, developers spotted the airport codes in those IDs and drew the natural conclusion. The precedent was fresh: on November 18, 2025, a Cloudflare fault, a bot-management configuration file that doubled in size and crashed the proxy, really did take ChatGPT offline along with X and a large slice of the web from 11:20 UTC.
The difference is what Cloudflare did each time. In November it published a detailed post-mortem the same day. On September 3 it said the opposite, on the record: “Cloudflare is not experiencing any significant service disruptions at this time. Our services are operating normally, and any reporting that deviates from this is incorrect.” Cloudflare’s status history shows nothing during the outage window but two scheduled maintenances in Montreal and Chicago. A company that admitted to breaking half the internet ten months earlier has no incentive to lie about a second time, and the cf-ray IDs prove only that Cloudflare sits in front of ChatGPT, which it always does. As one commenter in the thread put it, “I still don’t get why chatgpt.com would show a 404 because of an AI datacentre outage”: a 404 is an answer from a server, not a missing network.
“It’s probably DNS”, the other reflex, got its usual airing and its usual lack of evidence. No DNS operator reported an incident and no provider mentioned resolution failures.
Verdict: not supported, on Cloudflare’s explicit statement and the absence of any other pipe reporting a fault.
Is Memphis the Missing Link Between Claude and Grok?
Here is the thread that got the least attention relative to its evidence, and it does not involve Azure at all. Futurism’s follow-up the next day pointed at it; the contract behind it deserves spelling out.
On May 6, 2026, Anthropic announced that it had contracted “more than 300 megawatts of new capacity (over 220,000 NVIDIA GPUs)” at SpaceX’s Colossus 1 data center to serve Claude Pro and Max subscribers. Colossus 1 is in Memphis. On September 3, SpaceX blamed Grok’s failure on “an outage at our Memphis compute center” and apologised “to our impacted compute partners”. Anthropic is the best-documented of those partners, its incident started four minutes before xAI’s, and its status page said only that the cause had been identified.
That is a real, documented, physical dependency shared by two of the three services. If the Memphis event is what Anthropic’s “infrastructure issue” refers to, the Claude and Grok outages had one cause, and a boring one: a data centre had a bad morning. Anthropic has not said so, and Wired reported that it “declined to comment beyond its status page”. Until it does, the link is a plausible inference, not a fact, and we present it as one.
What Memphis cannot do is reach OpenAI. ChatGPT does not run there, OpenAI’s fault started more than an hour after the Memphis incident, and OpenAI described a routing error. As AI Chat Daily put it, the Memphis explanation “does not obviously cover OpenAI”.
Verdict: a documented dependency exists, neither company has connected it to September 3, and it cannot reach the third.
Did One Outage Knock Over the Others?
The cascade theory is the most respectable of the folk explanations. When one assistant fails, its users pile onto the next, and the surge takes that one down too. It has a precedent. On June 4, 2024, ChatGPT, Claude and Perplexity all went down within hours of each other. Perplexity’s error message said it outright: “We’re getting a lot of questions right now and have reached our capacity.” Nobody ever confirmed a shared cause that day either.
The theory had believers inside the industry this time. Grok’s error message, per The Verge, read “This model is overloaded right now. Please try again shortly or pick a different model”, and an OpenAI engineer posted on X that “rumor has it that when we go down, so much traffic has to be absorbed by the rest that they all go down”.
For September 3 the order of events is the problem. Claude failed first, Grok minutes later, and ChatGPT more than an hour after that. A cascade would have to have run towards ChatGPT, the service with by far the most users, from two much smaller ones, and OpenAI attributed its fault to routing, not load. The reverse direction, where a ChatGPT outage floods Claude and Grok, is the one that would make sense, and the clock rules it out.
Verdict: consistent with history, inconsistent with this timeline.
Was It GPT-6 Astra, AGI Waking Up, or Skynet?
This was the day’s other big story, and the two collided. While ChatGPT was down, the ChatGPT account posted “The stars are almost aligned”, and that afternoon OpenAI unveiled GPT-6 Astra, a model it said was trained on “more than 100,000 GPUs at our Stargate site in Texas”. Its president, Greg Brockman, told reporters: “It’s not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it’s reasonable.” Axios headlined the briefing “Welcome to the AGI era”. Fireship, whose first-look videos are where many working developers form their opinion of a new model, put the question straight in its September 4 title: “Did OpenAI actually build AGI?”
So the internet did what it does. “The system goes online September 3rd, 2026… Astra begins to learn at a geometric rate,” wrote one Hacker News user, quoting Terminator. “SkyNet is arming,” said another. Futurism collected the rest: “Finally I can see the Sun!”, Paris Marx’s “for a brief moment, millions of people had to use their brains again”, and ThePrimeagen’s complete analysis, four words long: “All the AI is down.” The more serious version, “Astra is being released today. Probably not a coincidence,” was the thread’s most repeated thought.
It is a coincidence, and an easy one to check. Astra reached only a limited group of enterprise customers in OpenAI’s Daybreak cybersecurity programme on September 3; paying subscribers got it the following day, after which Sam Altman wrote “First, sorry for the messy rollout”. Whatever load the launch put on OpenAI’s systems, it could not have touched Anthropic’s or xAI’s, and those two failed first. OpenAI’s own incident, per its spokesperson, was a 34-minute routing error, and a Hacker News user identifying as the incident commander for the day wrote in the thread: “We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don’t comment on other providers’ outages.” A model in limited preview does not have hands. It cannot reach into a Memphis data centre or a competitor’s routing table. If the AGI era began on September 3, it began by taking the morning off.
Verdict: the best jokes of the day, and nothing else.
Solar Storms, Cyberattacks and Plain Coincidence
The remaining theories are quick to dispose of. No provider mentioned an attack, and companies that suffer one tend to say so, because an attack is a better story than a routing error. The solar-storm version fails on the instruments: NOAA’s planetary K index, the standard measure of geomagnetic disturbance, peaked at 2.33 on September 3, checked on September 8, where a minor storm begins at 5. A geomagnetic storm strong enough to disturb data centres would in any case have disturbed power grids and satellites first.
That leaves coincidence, which is the least satisfying answer and the one the providers’ own statements leave you with. Each of these companies logs incidents every month; their status pages are long documents. Anthropic had already opened and closed a separate incident on Claude Sonnet 5 between 12:37 and 12:56 UTC that same morning, before any of this began. OpenAI’s page carried a single incident from 06:36 to 22:00 UTC on June 10, 2025, and Gemini, the survivor of September 3, was down for seven hours on June 10, 2026 because of “extreme read contention” in a Google database. Three incidents in one Thursday morning of the US working week, with one of them possibly shared through Memphis, is unusual, but the June 2024 triple outage shows it has happened before without a common cause ever surfacing. “The evidence supports three overlapping provider incidents, with different public explanations and incomplete root-cause disclosure,” concluded one careful reconstruction of the timeline, and that is where the honest account ends.
Why the Question Matters More Than the Answer
Whichever version is true, the day exposed the same thing: an enormous number of people and companies now depend on a handful of providers, and those providers depend on a handful of data-centre operators and, for the most part, one chip supplier.
The numbers are public. Amazon, Microsoft and Google took 28%, 20% and 15% of a cloud infrastructure market that reached $143 billion in the second quarter of 2026, according to Synergy Research Group. OpenAI has contracted an incremental $250 billion of Azure services, a $38 billion AWS agreement since expanded by $100 billion, up to 10 gigawatts of Nvidia systems and 6 gigawatts of AMD GPUs. Anthropic calls AWS its primary training partner, has contracted up to one million Google TPUs, committed $30 billion to Azure, and since May rents all of Colossus 1 in Memphis. xAI built Colossus with 100,000 Nvidia Hopper GPUs in 122 days and then doubled it. Nvidia sits behind most of those contracts, and every frontier model people reach for on a weekday morning lives on one of three clouds or on one campus in Tennessee. ChatGPT alone reported 900 million weekly users in February.
The dependency is now measurable. In a survey of 1,000 senior executives across 16 countries published in June 2026, IBM’s Institute for Business Value found that 71% said switching their primary AI vendor or model would be difficult, 91% said they do not fully understand their organisation’s dependencies across AI vendors, models and infrastructure, and 81% said a seven-day vendor outage would cause severe or critical disruption. Seventy-three percent described their AI estate as intentionally multi-vendor; only 7% operated at what IBM called an advanced level of control. “AI has introduced new forms of dependency that evolve faster than traditional governance, procurement, or technology cycles were designed to handle,” said IBM’s Ana Paula Assis.
Forrester’s Charlie Dai drew the operational conclusion the day after the outage, in ITPro: “When multiple major providers experience overlapping failures without a clearly established common cause, enterprises cannot accurately assess systemic risk, dependency concentration, or recurrence likelihood.” His prescription was “multi-model strategies, fallback workflows, and business continuity plans instead of assuming frontier AI services will always be available.”
Regulators are moving the same way. On July 13, 2026, the United Kingdom brought Amazon Web Services, Google Cloud, Microsoft and Oracle under direct supervision as critical third parties to the financial system. “When the same providers serve thousands of firms, a single failure can reverberate across the financial system,” said Nikhil Rathi, chief executive of the Financial Conduct Authority. The FCA’s policy statement PS26/2 adds mandatory reporting of operational incidents and material third-party arrangements from March 18, 2027, a regime written for exactly the kind of dependency that September 3 made visible.
That is the useful reading of the outage. The Azure theory was wrong, but the anxiety behind it was right: the systems people now use to write, code, search and decide are concentrated in a very small number of buildings, and nobody outside those buildings can see how they connect.
Is Decentralised AI a Real Alternative?
The case for decentralising AI used to be a crypto argument. September 3 made it an availability argument. If three companies with three separate stated causes can fail inside one window, the risk is not any single data centre but the shape of the industry: a few models, on a few clouds, mostly on one chip supplier, and everyone downstream of all of them at once.
Gonka is one of the projects built on that argument, and the one whose founders make it most bluntly. It describes itself as “a decentralized network for high-efficiency AI compute” and went live in August 2025 after incubation by Product Science, the Los Angeles company of David and Daniil Liberman, whose earlier startup was acquired by Snap. Independent operators contribute Nvidia GPUs and are paid in the network’s token; developers call open-weight models such as DeepSeek, Kimi and MiniMax through an OpenAI-compatible endpoint, with prices that move with utilisation rather than with a vendor’s price list. Its architecture documentation makes the relevant promise in one line: “The system is decentralized, with no single point directing inference requests to the network nodes.” Its whitepaper names the risk it exists to address: “Concentrating computational resources within a few dominant providers presents significant risks related to censorship and centralized control.” Bitfury committed $50 million to the network in December 2025, when it reported compute equivalent to more than 6,000 Nvidia H100s; by February 2026 the figure the project gave was about 14,000 H100-equivalents across roughly 20 countries.
The founders’ word for the goal is sovereignty. “If you do not control compute, your AI policy is a request rather than a strategy,” the Liberman brothers said in July, calling the alternative “GPU feudalism, a future where people become tenants on someone else’s compute estate”. In a Fortune op-ed in May they described hyperscaler concentration as a “single point of failure or control”. The claim is that a country, a university or a company can have AI capacity nobody else can switch off without building a hyperscale cloud, by pooling GPUs it already owns into a network with no owner.
Two caveats belong next to that, and Gonka’s own documents supply one of them. Decentralised networks trade a single point of failure for higher variance: “a decentralized network inherently possesses higher reliability variance than a dedicated data center,” as one 2026 sector analysis put it, and none of them yet offers the enterprise service guarantees a hyperscaler contract does. Gonka’s security analysis says that “at present, no successful cheating strategies against this combination of defenses are known”, which is an honest sentence rather than a warranty. And the network’s token has lost most of its value since January, which says nothing about whether the GPUs serve inference but a great deal about where the sector’s attention still goes. Bittensor, Akash, io.net and Prime Intellect are running variations of the same experiment.
What the experiment has shown so far is narrower than its marketing, and still worth having: open models can be served from thousands of GPUs that no single company controls, and a routing error at one vendor or a bad morning in Memphis does not take all of them down together. On a day like September 3, that is the only property that matters.
What to Do When Your AI Assistant Goes Down
For an individual, the September 3 playbook is short.
Check the status page before anything else. status.openai.com, status.claude.com and status.x.ai all showed the incident within minutes. If the page is red, nothing on your side will help; wait, or switch.
Keep a second assistant signed in. Gemini stayed up on September 3, most Claude models were back by 15:25 UTC while ChatGPT was still down, and on June 4, 2024 the pattern was different again. Two providers on two clouds is the cheapest resilience there is. Developers should go further and build the fallback into the code, so that a 34-minute routing error at one vendor becomes a log line rather than an incident.
Learn to tell an outage from a block. This is the case where a VPN matters. Both OpenAI and Anthropic publish lists of supported countries, and OpenAI’s page warns that “accessing or offering access to our services outside of the countries and territories listed below may result in your account being blocked or suspended”. Regulators block too: Italy’s data protection authority ordered ChatGPT to stop processing Italian users’ data in an order announced on March 31, 2023, and the service went dark in Italy until April 28. When the status page is green but the assistant refuses you, the problem is your location, not the provider. A connection through a server in your home country, one of Le VPN’s 100+ locations, restores the service you pay for while you travel, the same way it does for banking or television. Our guide to bypassing internet censorship covers the state-level version of the same problem, and this older post explains why a VPN can route around a regional network failure but never around a provider’s own.
If your AI agent needs a fixed exit country, give it one. Autonomous agents fail the location test too, and they cannot open a support ticket. Le VPN’s account-less x402 passes exist for that case: a WireGuard config bought inside an HTTP request, for a day, a week or a month on a single server in France, Germany or the UK.
None of this brings a dead model back. That is the point of the September 3 story. The tools have become infrastructure, the infrastructure is concentrated, and the only defence available to a user today is to depend on more than one piece of it.
About the author
Le VPN Blog Editor
Alan Summers has been writing and editing for the Le VPN blog for years, covering online privacy, cyber security, and the best ways to get the most out of a VPN. He keeps a close eye on the news that affects internet freedom around the world and turns it into practical advice for Le VPN readers.
Articles by Alan Summers →