Independent reporting on American politics
STATE BEACON

Four AI giants hit by overlapping outages, exposing shared‑infrastructure fragility

On September 3, 2026, OpenAI’s ChatGPT, Anthropic’s Claude, xAI’s Grok and Google’s Gemini all suffered service interruptions that overlapped for nearly two hours – the first recorded simultaneous outage across the AI frontier.

By State Beacon·
Server rack in a colocation data‑center that hosts the compute nodes for multiple frontier AI models

On September 3, 2026, the four leading cloud‑based artificial‑intelligence models – OpenAI’s ChatGPT, Anthropic’s Claude, xAI’s Grok and Google’s Gemini – all suffered service interruptions that overlapped for a window of roughly two hours. Ars Technica confirmed the overlapping windows and noted that this is the first recorded simultaneous outage across the AI frontier.

Timeline of the September 3 outage

The sequence of events, as reported by Ars Technica, began early in the morning and stretched into the early afternoon:

  • 9:23 am ET – Anthropic announced a partial outage for Claude.
  • 9:45 am ET – DownDetector logged 1,365 user‑reported incidents for xAI’s Grok.
  • 10:30 am ET – DownDetector showed a modest spike of roughly 23 reports for Google’s Gemini.
  • 10:43 am ET – OpenAI reported elevated error rates across ChatGPT and its Codex model.
  • 10:45 am ET – StatusGator recorded a likely outage for the Gemini API lasting until 11:15 am.
  • 12:16 pm ET – Anthropic confirmed that the Claude outage was resolved.
  • 12:55 pm ET – OpenAI marked the ChatGPT/Codex issue as resolved.

All four models were therefore unavailable, either partially or fully, during the overlapping window of roughly 10:45 am to 12:55 pm Eastern.

Outage windows for the four leading AI models on 3 Sept 2026 (Eastern Time)
Model Outage Start Outage End Duration Source
Claude (Anthropic) 9:23 am 12:16 pm ≈2h 53m Ars Technica
ChatGPT / Codex (OpenAI) 10:43 am 12:55 pm ≈2h 12m Ars Technica
Grok (xAI) ≈9:00 am (reported) ongoing at time of writing ≥4h Ars Technica
Gemini (Google) 10:45 am 11:15 am ≈30 min Ars Technica

Source: Ars Technica.

Scale of the disruptions

DownDetector, a crowdsourced outage‑monitoring service, recorded 1,365 incident reports for Grok by 9:45 am. The same platform logged 412 reports for Gemini around 11 am. By contrast, the two most mature models – Claude and ChatGPT – have historically high availability. Over the last 90 days, Claude posted a 99.4 percent uptime, while ChatGPT posted a 99.63 percent uptime, both figures supplied by Ars Technica.

These percentages illustrate that, under normal conditions, each model is expected to be reachable for roughly 99.5 percent of the time, equating to about 3.6 hours of downtime per month. The September 3 event pushed the combined downtime for the four services into a single, overlapping window, a scenario that had not been observed before.

Company backgrounds and stakes

OpenAI, founded on 11 December 2015, is headquartered in San Francisco and employs roughly 4,500 people according to Wikidata (Q21708200). Its chief executive, Sam Altman, has steered the firm through rapid scaling of its flagship ChatGPT model, which now serves hundreds of millions of users worldwide.

Anthropic, established on 26 January 2021, also operates out of San Francisco with an employee headcount of about 2,500 (Wikidata Q116758847). Dario Amodei serves as chief executive. Claude, Anthropic’s flagship conversational model, has been positioned as a safety‑first alternative to OpenAI’s offering.

xAI, formally known as XAI Floating Rate & Alternative Income Trust (ticker XFLT), is listed on the NYSE. While the research packet does not provide a chief executive or employee count for the trust, the outage involved its Grok model, which is marketed as a high‑performance, open‑source‑friendly alternative.

Google’s AI division, part of Alphabet Inc., runs Gemini from its Mountain View headquarters. Gemini’s API is a core component of Google Cloud’s AI suite, and the model’s uptime is critical for enterprise customers that embed Gemini into production workflows.

All four firms rely heavily on cloud infrastructure, much of which is hosted on shared data‑center providers, including Amazon Web Services, Microsoft Azure and Google Cloud itself. The overlapping outage suggests that a common point of failure – whether a network backbone, a DNS provider, or a third‑party monitoring service – may have impacted multiple providers simultaneously.

Implications for the AI ecosystem

The incident underscores a growing concern among enterprise users: as AI models become embedded in mission‑critical applications – from customer‑service chatbots to automated code generation – the tolerance for downtime shrinks dramatically. A two‑hour window where four leading models are simultaneously unavailable could halt business processes that depend on any one of them.

From a risk‑management perspective, the outage highlights the need for multi‑model redundancy. Companies that have built fallback mechanisms – for example, routing requests to an alternative model when the primary service degrades – may have mitigated the impact. However, the research packet does not provide data on how many customers had such safeguards in place.

Regulators are also watching. The commission’s “why now” note points to heightened scrutiny of AI reliability as firms embed these services into critical workflows. While no regulatory action has been announced as of the time of writing, the event could accelerate discussions about service‑level agreements (SLAs) for AI APIs and the transparency of outage reporting.

Finally, the overlapping nature of the outage raises technical questions about shared‑infrastructure dependencies. If a single upstream provider suffered a network partition, it could cascade across multiple AI platforms that rely on the same backbone. The packet does not identify the root cause, and the companies have not released post‑mortem analyses beyond the timing statements cited.

What remains unknown

Several key details are still missing:

  • The precise technical trigger for each provider’s outage – whether it was a hardware failure, a software bug, or a network issue – has not been disclosed.
  • Whether any of the four firms have formally coordinated with each other to investigate shared‑infrastructure risks.
  • The financial impact on each company. The packet contains no revenue or profit figures tied to the outage, and the firms have not reported any loss of business.
  • How many enterprise customers experienced actual disruption versus those who have redundancy built into their stacks.

Future statements from OpenAI, Anthropic, xAI and Google will be essential to flesh out the full picture.

In the meantime, the September 3 event serves as a reminder that the AI frontier, while rapidly expanding, remains dependent on the same underlying internet and cloud infrastructure that powers the rest of the digital economy. As AI services become ever more woven into everyday business processes, the industry will need to address the fragility exposed by this first simultaneous outage.