AI API incidents — outages measured from 4 regions

In the last 72 hours 1 outage(s) could be attributed to a provider with confidence. Most recent: openai failed from 3 of 4 regions at once on 2026-07-25 09:20 UTC, for at least 14 minutes, the host itself answering HTTP 503.

How do you know it was the provider and not my network?

Every provider is probed from 4 independent regions (Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central)) every five minutes, and all 45 of them are probed on the same schedule. Two questions follow from that, and an incident has to survive both.

Did it fail from more than one place? If a request fails from one region while the others succeed, the likely cause is the path between that region and the provider — a peering problem, a route change, one bad edge node. We do not call that an outage.

Was it the only thing failing? If several unrelated APIs stop answering in the same few minutes, the shared cause sits above all of them: a transit route, a CDN both sit behind, or our own probes. Blaming one provider for that would be guessing. Those events are listed separately below and attributed to nobody. What is left — failing from several regions at once, alone — is the provider, and when its own host answers 5xx rather than going silent, the provider has said so itself. Neither a status page nor a crowd-reporting site can make this distinction: one sees only what the provider admits, the other only what users happen to report.

Which AI APIs had confirmed outages in the last 72 hours?

Started (UTC)ProviderRegions hitWhichDurationFailed probesDid they say so?
2026-07-25 09:20openai3 of 4Asia (Tokyo), Europe (Germany), South America (São Paulo)≥14 min4Yes — “Elevated error rates”, 3 min before us

Did the provider admit it?

The last column checks each failure against the provider’s own status page. Only 10 of the 45 providers tracked here publish a machine-readable status feed at all; for the rest there is nothing to check against, which is itself worth knowing before you plan around one.

Where a feed exists, both outcomes are informative. When it carries a matching entry, two independent records agree and the timing shows which noticed first. When it carries none, we recorded a failure their page does not mention — and that is not an accusation. Status pages have thresholds, and a short burst of 5xx can sit legitimately below one. We report the discrepancy; explaining it is the provider’s to do.

Which failures could not be attributed to anyone?

These hit several regions, but other providers were failing in the same minutes, so a shared upstream cause is at least as likely as a coincidence of separate outages. They are published because omitting them would flatter the data, and left unattributed because that is what the evidence supports.

Started (UTC)ProviderRegions hitWhichDurationFailed probesDid they say so?
2026-07-24 11:00novita4 of 4Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central)≥6 min9No machine-readable status page
2026-07-24 11:00together4 of 4Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central)≥6 min8No machine-readable status page

What isolated failures were recorded?

These failed from a single region only. They are published because hiding them would misrepresent the data, but they are not evidence that the provider was down.

Started (UTC)ProviderRegions hitWhichDurationFailed probesDid they say so?
2026-07-25 11:01hunyuan1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-25 09:10google1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-25 06:09hunyuan1 of 4South America (São Paulo)<5 min1No machine-readable status page
2026-07-25 05:47doubao1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-25 05:21iflytek1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-25 03:42google1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 23:51sambanova1 of 4Europe (Germany)<5 min1No entry in their status feed
2026-07-24 22:55sambanova1 of 4Europe (Germany)<5 min1No entry in their status feed
2026-07-24 20:19google1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 18:08sensenova1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 17:34sensenova1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 17:34stepfun1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 15:20doubao1 of 4US (Central)<5 min1No machine-readable status page
2026-07-24 15:01featherless1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-24 14:03sensenova1 of 4US (Central)<5 min1No machine-readable status page
2026-07-24 13:07sambanova1 of 4South America (São Paulo)<5 min1No entry in their status feed
2026-07-24 11:10replicate1 of 4Asia (Tokyo)<5 min1No entry in their status feed
2026-07-24 09:50yi-01ai1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-24 09:50baichuan1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-24 01:13sambanova1 of 4Europe (Germany)≥11 min2No entry in their status feed
2026-07-23 23:59stepfun1 of 4Europe (Germany)<5 min1No machine-readable status page
2026-07-23 15:11iflytek1 of 4US (Central)<5 min1No machine-readable status page
2026-07-23 15:11baichuan1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-23 15:11yi-01ai1 of 4Asia (Tokyo)<5 min1No machine-readable status page
2026-07-23 15:06sambanova1 of 4South America (São Paulo)<5 min1No entry in their status feed
2026-07-23 15:01featherless1 of 4South America (São Paulo)<5 min1No machine-readable status page
2026-07-23 12:17deepseek1 of 4Europe (Germany)<5 min1No machine-readable status page

What counts as a failure here?

A probe fails when the connection times out, is refused, or the API host answers with a 5xx status. An authentication error does not count: the probes are unauthenticated by design, so a 401 or 403 means the endpoint answered correctly and is recorded as up. Latency alone — even very high latency — is not recorded as an incident; it appears in the p95 figures instead.

Why is this window only 72 hours?

Because that is how much can be stated precisely. The underlying series is retained for 90 days and a longer view will be published once it covers enough time to distinguish a pattern from an accident. Extending the window before then would turn two coincidences into a trend, which is the opposite of the point.

Method and limitations: methodology · Machine-readable: incidents.json · Licence: CC BY 4.0