Your three LLM vendors went down the same morning: multi-model is not a resilience architecture

Your three LLM vendors went down the same morning: multi-model is not a resilience architecture

On Thursday, September 3, 2026, ChatGPT, Claude and Grok all became unavailable within tens of minutes of each other. OpenAI told The Register that a routing error starting around 7:43 am Pacific made ChatGPT and Codex unavailable for some users, with a fix applied around 8:17 am, and its status page incident listed nineteen degraded components. Anthropic confirmed to the same outlet "an infrastructure issue caused a partial outage" across Claude.ai, Claude Code, Claude Cowork and the API, with service restored at 16:16 UTC after a little more than three hours. xAI opened an incident on Grok around 6:30 am Pacific, and SpaceXAI later posted an apology attributing the failure to "an outage at our Memphis compute center", adding an apology to "our impacted compute partners". Axios also reported issues with Gemini, without confirmation from Google.

The instinctive reading is that the day validates multi-model: three vendors, three outages, three times as many reasons to route across several APIs. The day demonstrates the opposite. Multi-model protects you against a vendor that raises prices, retires a model, changes its terms or degrades its quality. It does not protect availability when the vendors you picked share the same compute layer, because in that case the correlation between their failures is not zero, it is close to one. On September 3, the vendor that many architectures had chosen as their third fallback was renting its GPUs to the vendor it was supposed to replace.

The September 1 piece on model deprecation notice periods and the September 3 piece on substitution clauses in SaaS contracts dealt with contractual dependency, the kind you can read in a document and negotiate. This one appears nowhere in your contracts: it lives in your vendors' compute subcontracting chain, a layer almost no vendor assessment grid asks to see.

Three vendor cards resting on a single foundation crossed by a fracture line.
Three separate model brands can rest on a single compute floor, invisible from the contract.

What the September 3 timeline actually shows

The honest reading of that day is that it mixes a coincidence with a real dependency, and the distinction matters if you want a usable conclusion.

On the OpenAI side, the stated cause is an internal routing error with no established link to a third party, and the outage window is short. On the Grok and Claude side, the public evidence converges on a common point. SpaceXAI named its Memphis compute center and apologized to its compute partners. And SpaceXAI announced on May 6, 2026 that it had signed an agreement giving Anthropic access to Colossus 1, the Memphis supercomputer built around more than 220,000 NVIDIA GPUs, capacity the announcement said would serve Claude Pro and Claude Max subscribers. Neither company has publicly established that the Claude outage originated there, and Anthropic's wording stays deliberately generic. The link is plausible and documented at both ends, it is not confirmed, and an article asserting it would go further than the facts allow.

The remaining press hypotheses hold up less well. Axios pointed to Azure disruptions as a possible contributor to all three outages; The Register observed the same day that the AWS, Google Cloud and Microsoft Azure status pages showed no relevant difficulties, while noting that those pages do not always reflect problems in a timely manner. Cloudflare, used in various capacities by all three companies, firmly denied any significant disruption to its services.

The vocabulary of dependency

Compute substrate: the data centers, accelerators and networks a model is actually served from, independent of the model's commercial brand.

Outage correlation: the probability that two components fail at the same time. A backup architecture is only worth something when this correlation is low between the primary component and its replacement.

Degraded mode: product behavior defined in advance for when a capability is unavailable, as opposed to failover, which assumes an equivalent capability exists elsewhere.

Register of information: the regulatory inventory of contracts with IT service providers required of EU financial entities under DORA, which must cover subcontracting chains supporting a critical function.

The layer your vendor grid does not describe

The blind spot is not vendor quality, it is the granularity of what you assess. A typical assessment grid covers the model, the API, data processing terms, the published SLA and the version roadmap. It stops at the brand. It never asks which data centers serve inference, which accelerators carry it, or which providers sit in the second tier of subcontracting.

Anthropic's case shows how fast that layer moves. The company serves Claude from a genuinely multi-vendor fleet, with AWS as its primary training partner, large Google TPU capacity ramping through 2026, NVIDIA GPUs, and Colossus 1 capacity since May. That is not a weakness, it is the company's explicit strategy. But it has a direct consequence for you: the composition of the floor under a single model name changes several times a year, without customer notification, and your correlation analysis from six months ago has expired without anything telling you so.

The scale of the problem is structural rather than anecdotal. The Dataiku survey run with Harris Poll among 600 enterprise CIOs reports that 81% expect to rely on two or more LLM providers in 2026, and that 55% have already switched providers at least once, cost reduction being the leading driver. The dominant motivation is per use case performance and price, not availability. Multi-model spread for good reasons, and it is then presented in risk committees as a continuity control it was never designed to provide.

Diagram contrasting three independent fallback paths with three paths converging on a shared substrate.
Failover assumes two independent paths; it is worth nothing when both paths meet again at ground level.

In the field, the symptom is nearly always the same. The multi-vendor switch exists, it is wired into the gateway, it is documented in the architecture diagram, and it has never run outside a unit test. The first time it runs for real is during the incident, and that is the day the team discovers that the backup vendor's quotas are sized for testing, that output formats differ enough to break downstream parsing, and that the switch triggers on an HTTP error code that does not occur when a vendor answers slowly instead of answering wrong.

Regulators already ask for this map

For part of this audience, that map is not an optional good practice. DORA requires EU financial entities to maintain a register of information covering IT service contracts, including the subcontracting chains that support a critical or important function, and to explicitly assess concentration risk. Article 30 sets the mandatory clauses: audit and access rights, exit strategy, subcontractor control, and service levels covering data location and computing capacity among others.

The designation on November 18, 2025 of the first nineteen critical ICT third-party providers placed under direct oversight by the European Supervisory Authorities shows what the regulator is looking at. The list is largely made up of infrastructure providers, including AWS, Microsoft, Google Cloud, Oracle, IBM, SAP, Deutsche Telekom, Equinix and Swift, selected against four criteria that include sector-wide concentration of reliance and substitutability of the service. The oversight logic targets the floor, not the application brand sitting on top of it. A financial entity that declares three LLM vendors in its register without documenting the compute carrying them is describing a diversification its own regulator will not recognize as one.

Design a degraded mode instead of a failover

The practical conclusion is not to abandon multi-model, which remains justified on cost, per use case quality and negotiating leverage. It is to stop counting it twice, once as resilience and once as performance, and to treat availability as a separate product problem.

A degraded mode is defined feature by feature, by answering one question: what does this screen, this journey, this batch job do when no model is reachable for three hours. Acceptable answers exist and are designed calmly, in advance: queue with replay and a committed deferred deadline, fall back to a weaker local model that is good enough for classification, revert to a deterministic rule for the most frequent subset of cases, or show an explicit message and run manually on purpose. Only one answer is unacceptable, the one that displays a technical error and lets the user retry in a loop, because it turns a three-hour vendor outage into a trust incident that outlasts it.

What to put in motion this week

Ask each of your model vendors in writing for the list of regions and infrastructure subcontractors serving your traffic, along with their notification policy when that changes. The answer, including a partial answer or a refusal, is itself an assessment data point worth recording. An entity subject to DORA already holds a contractual right to ask.

Recompute your outage correlation from the answers you get, not from the brands. If your three paths converge on one or two infrastructure operators, your architecture has a single point of failure presented under three names, and your continuity plan overstates your availability by an order of magnitude.

Run a total outage exercise in a pre-production environment, cutting access to every vendor at once rather than one after another. Testing with a single vendor unavailable validates the gateway; only testing with zero vendors validates the product. Measure what the user sees, what is lost, and what replays on recovery.

Separate the model SLA from the underlying infrastructure SLA, in your internal commitments and in your customer contracts. A vendor can only guarantee what it controls, and the part it does not control is exactly the part that produced the September 3 outage.

Write the degraded mode for your three most critical journeys, validate it with the business rather than with the platform team, and keep it to one page. A degraded mode that has not been accepted by whoever will answer the customer does not exist.

Conclusion

September 3 did not reveal new fragility among model vendors, whose availability rates remain comparable to the rest of the software industry. It revealed that a large share of multi-model architectures cannot answer, within the hour, whether their vendors are independent of one another. Until that question has a written and dated answer, the "vendor redundancy" line in your continuity plan describes a commercial intention, not a technical property of your system.


Sources: As of September 2026