AI API incidents — outages measured from 4 regions
In the last 72 hours 1 outage(s) could be attributed to a provider with confidence. Most recent: openai failed from 3 of 4 regions at once on 2026-07-25 09:20 UTC, for at least 14 minutes, the host itself answering HTTP 503.
How do you know it was the provider and not my network?
Every provider is probed from 4 independent regions (Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central)) every five minutes, and all 45 of them are probed on the same schedule. Two questions follow from that, and an incident has to survive both.
Did it fail from more than one place? If a request fails from one region while the others succeed, the likely cause is the path between that region and the provider — a peering problem, a route change, one bad edge node. We do not call that an outage.
Was it the only thing failing? If several unrelated APIs stop answering in the same few
minutes, the shared cause sits above all of them: a transit route, a CDN both sit behind, or our own probes.
Blaming one provider for that would be guessing. Those events are listed separately below and attributed to
nobody. What is left — failing from several regions at once, alone — is the provider, and when its
own host answers 5xx rather than going silent, the provider has said so itself. Neither a status
page nor a crowd-reporting site can make this distinction: one sees only what the provider admits, the other
only what users happen to report.
Which AI APIs had confirmed outages in the last 72 hours?
| Started (UTC) | Provider | Regions hit | Which | Duration | Failed probes | Did they say so? |
|---|---|---|---|---|---|---|
| 2026-07-25 09:20 | openai | 3 of 4 | Asia (Tokyo), Europe (Germany), South America (São Paulo) | ≥14 min | 4 | Yes — “Elevated error rates”, 3 min before us |
Did the provider admit it?
The last column checks each failure against the provider’s own status page. Only 10 of the 45 providers tracked here publish a machine-readable status feed at all; for the rest there is nothing to check against, which is itself worth knowing before you plan around one.
Where a feed exists, both outcomes are informative. When it carries a matching entry, two
independent records agree and the timing shows which noticed first. When it carries none, we recorded
a failure their page does not mention — and that is not an accusation. Status
pages have thresholds, and a short burst of 5xx can sit legitimately below one. We report
the discrepancy; explaining it is the provider’s to do.
Which failures could not be attributed to anyone?
These hit several regions, but other providers were failing in the same minutes, so a shared upstream cause is at least as likely as a coincidence of separate outages. They are published because omitting them would flatter the data, and left unattributed because that is what the evidence supports.
| Started (UTC) | Provider | Regions hit | Which | Duration | Failed probes | Did they say so? |
|---|---|---|---|---|---|---|
| 2026-07-24 11:00 | novita | 4 of 4 | Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central) | ≥6 min | 9 | No machine-readable status page |
| 2026-07-24 11:00 | together | 4 of 4 | Asia (Tokyo), Europe (Germany), South America (São Paulo), US (Central) | ≥6 min | 8 | No machine-readable status page |
What isolated failures were recorded?
These failed from a single region only. They are published because hiding them would misrepresent the data, but they are not evidence that the provider was down.
| Started (UTC) | Provider | Regions hit | Which | Duration | Failed probes | Did they say so? |
|---|---|---|---|---|---|---|
| 2026-07-25 11:01 | hunyuan | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-25 09:10 | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page | |
| 2026-07-25 06:09 | hunyuan | 1 of 4 | South America (São Paulo) | <5 min | 1 | No machine-readable status page |
| 2026-07-25 05:47 | doubao | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-25 05:21 | iflytek | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-25 03:42 | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page | |
| 2026-07-24 23:51 | sambanova | 1 of 4 | Europe (Germany) | <5 min | 1 | No entry in their status feed |
| 2026-07-24 22:55 | sambanova | 1 of 4 | Europe (Germany) | <5 min | 1 | No entry in their status feed |
| 2026-07-24 20:19 | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page | |
| 2026-07-24 18:08 | sensenova | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 17:34 | sensenova | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 17:34 | stepfun | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 15:20 | doubao | 1 of 4 | US (Central) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 15:01 | featherless | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 14:03 | sensenova | 1 of 4 | US (Central) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 13:07 | sambanova | 1 of 4 | South America (São Paulo) | <5 min | 1 | No entry in their status feed |
| 2026-07-24 11:10 | replicate | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No entry in their status feed |
| 2026-07-24 09:50 | yi-01ai | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 09:50 | baichuan | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-24 01:13 | sambanova | 1 of 4 | Europe (Germany) | ≥11 min | 2 | No entry in their status feed |
| 2026-07-23 23:59 | stepfun | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
| 2026-07-23 15:11 | iflytek | 1 of 4 | US (Central) | <5 min | 1 | No machine-readable status page |
| 2026-07-23 15:11 | baichuan | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-23 15:11 | yi-01ai | 1 of 4 | Asia (Tokyo) | <5 min | 1 | No machine-readable status page |
| 2026-07-23 15:06 | sambanova | 1 of 4 | South America (São Paulo) | <5 min | 1 | No entry in their status feed |
| 2026-07-23 15:01 | featherless | 1 of 4 | South America (São Paulo) | <5 min | 1 | No machine-readable status page |
| 2026-07-23 12:17 | deepseek | 1 of 4 | Europe (Germany) | <5 min | 1 | No machine-readable status page |
What counts as a failure here?
A probe fails when the connection times out, is refused, or the API host answers with a 5xx status. An authentication error does not count: the probes are unauthenticated by design, so a 401 or 403 means the endpoint answered correctly and is recorded as up. Latency alone — even very high latency — is not recorded as an incident; it appears in the p95 figures instead.
Why is this window only 72 hours?
Because that is how much can be stated precisely. The underlying series is retained for 90 days and a longer view will be published once it covers enough time to distinguish a pattern from an accident. Extending the window before then would turn two coincidences into a trend, which is the opposite of the point.
Method and limitations: methodology · Machine-readable: incidents.json · Licence: CC BY 4.0